feat(lessons): the document shape is the stored record, and it travels (#3730)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / integration (push) Successful in 51s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m31s
CI & Build / Build & push image (push) Successful in 25s
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / integration (push) Successful in 51s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m31s
CI & Build / Build & push image (push) Successful in 25s
Milestone 385 step 3 — the step where the kind either works or is cosmetic.
THE DOCUMENT, and why there is no `lesson_document()` beside `rule_document()`
in embeddings. The step expected one. The difference is where the sharp shape
LIVES. A rule keeps its trigger in a column and its title is a plain name, so
`{title} — {trigger}` has to be synthesised at embed time and exists nowhere
else. A snippet — the only sharp record in the corpus by #2485's measurement,
0.153 top-to-second against 0.010–0.023 — gets there the other way: its STORED
title is already the join and its stored body already opens with the trigger,
so the ordinary `title\nbody` join IS the sharp document. Step 1 chose the
snippet route and step 2 built it, so `lessons.lesson_document` composes what
is STORED and the generic chunker does the rest.
The consequence the step asked about: `chunk_document` is untouched, so
CHUNKER_VERSION does not move and NOTHING re-embeds. The step's "Re-embed"
section describes a change this design does not make.
THE NARRATIVE stays in the body, departing from the step's instruction to keep
it out. `rule_document` excludes `why` because long dated narrative made
sixteen dev-logs land on the centroid of "development" — but that finding
predates chunking (#280). A body over budget is now split, and every chunk is
prefixed with the title, which for a lesson carries the trigger. The story
occupies its own vectors instead of averaging itself into the trigger's, and
each of those is still anchored to when the lesson applies. A guard asserts
exactly that. Holding the story out would cost the reader the only part that
explains the insight, to buy a sharpness the chunker already provides.
GLOBAL IN THE SEARCH is the real new code: `GLOBAL_NOTE_TYPES` and
`include_global_kinds` on `semantic_search_notes`, widening the PROJECT filter
alone. Off by default, because two callers depend on that filter holding — the
near-duplicate gate compares a record only against its own project on purpose,
and a globally visible kind there would let a lesson block an unrelated note's
create on a project its author never touched. It composes with `note_type`
rather than overriding it, so narrowing to snippets does not quietly acquire
lessons, and it changes nothing about the ACL: `notes_visibility_clause` still
gates every row.
Wired into the explicit MCP search only — the operator asked, and there is no
budget to spend. The unasked-for injection arms are step 5's subject (#3732)
and the legibility of a lesson appearing on a foreign project is step 7's
(#3734), so neither is turned on here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
@@ -140,6 +140,9 @@ async def search(
|
||||
enter_project) — otherwise this searches across ALL projects and
|
||||
bleeds unrelated work into the result set. 0 = search everything
|
||||
(use only when you genuinely want a cross-project sweep).
|
||||
A LESSON is the exception and arrives whatever the scope: the kind
|
||||
records an insight that transfers, so it is reachable from a
|
||||
project it was not written on.
|
||||
system_id: Narrow to records tagged to one System (a named
|
||||
subsystem/area — enter_project lists them). Use when investigating
|
||||
a specific subsystem: it cuts the candidates to records someone
|
||||
@@ -167,6 +170,14 @@ async def search(
|
||||
uid, q, limit=limit, is_task=is_task,
|
||||
project_id=project_id or None,
|
||||
system_id=system_id or None,
|
||||
# A LESSON is reachable from any project (milestone 385). The kind
|
||||
# exists to carry an insight to the next project, so a project filter
|
||||
# that hid it would hide it precisely where it is worth having. Only
|
||||
# the project filter widens — everything else about the scoping holds,
|
||||
# and a caller narrowing by `content_type` still gets what it asked
|
||||
# for. This is the explicit search, where the operator asked; the
|
||||
# unasked-for arms decide their own budget separately.
|
||||
include_global_kinds=True,
|
||||
# An explicit search reaches everything the operator may read, including
|
||||
# records shared with them one-to-one.
|
||||
scope="read",
|
||||
|
||||
@@ -508,6 +508,25 @@ async def upsert_note_embedding(
|
||||
logger.warning("Failed to persist embedding for note %d", note_id, exc_info=True)
|
||||
|
||||
|
||||
# Kinds that belong to no single project, and are therefore reachable from a
|
||||
# project-scoped search of a DIFFERENT project when a caller asks for them
|
||||
# (milestone 385 step 3).
|
||||
#
|
||||
# A lesson is the whole reason this exists. "A better way to think about this
|
||||
# problem" is not true only where it was learned, and a lesson confined to its
|
||||
# origin project would be unreachable exactly where it is most useful — on the
|
||||
# next project, which is the case the kind was created for (#3727).
|
||||
# `semantic_search_rules` has always had this property: it scopes by OWNERSHIP
|
||||
# rather than by what binds a given project, because "is there a rule about
|
||||
# this" is a question asked across a whole rulebook. A lesson asks the same
|
||||
# kind of question.
|
||||
#
|
||||
# Spelled as a literal rather than imported from `services.lessons`, which
|
||||
# imports `trigger_title` from this module and would close a cycle. A guard
|
||||
# pins the two equal instead.
|
||||
GLOBAL_NOTE_TYPES: tuple[str, ...] = ("lesson",)
|
||||
|
||||
|
||||
# Both searches rank WITHOUT the threshold and apply it in Python, so the best
|
||||
# rejected score stays observable (#3670). The qualifying set is provably
|
||||
# unchanged: rows arrive ordered by distance ascending, so every above-bar row
|
||||
@@ -536,6 +555,7 @@ async def semantic_search_notes(
|
||||
note_type: str | Sequence[str] | None = None,
|
||||
task_kind: str | Sequence[str] | None = None,
|
||||
orphan_only: bool = False,
|
||||
include_global_kinds: bool = False,
|
||||
scope: str = "own",
|
||||
demote_superseded: bool = True,
|
||||
system_id: int | None = None,
|
||||
@@ -570,6 +590,19 @@ async def semantic_search_notes(
|
||||
alone can express it. With `note_type="note", task_kind="issue"` a caller
|
||||
gets fixed problems and durable notes without the open to-do list.
|
||||
|
||||
`include_global_kinds` lets a project-scoped search ALSO reach the kinds in
|
||||
GLOBAL_NOTE_TYPES — records that belong to no single project — so a lesson
|
||||
written on one project is found from another. It widens the PROJECT filter
|
||||
only: a caller that also passes `note_type` still gets exactly the kinds it
|
||||
asked for, so narrowing to snippets does not quietly acquire lessons.
|
||||
|
||||
Off by default, because two callers depend on the project filter holding.
|
||||
The near-duplicate gate compares a record only against its own project on
|
||||
purpose, and a globally-visible kind would let a lesson block an unrelated
|
||||
note's create on a project its author never touched. Ordinary note recall
|
||||
is project-scoped for the same reason — the point of the carve-out is that
|
||||
ONE kind escapes, not that scoping is weaker.
|
||||
|
||||
`scope` ("own" | "browse" | "read", see access.notes_visibility_clause)
|
||||
decides how far this may see. It exists because this one function serves
|
||||
three different kinds of act: an explicit search, which should reach
|
||||
@@ -631,7 +664,12 @@ async def semantic_search_notes(
|
||||
if orphan_only:
|
||||
stmt = stmt.where(Note.project_id.is_(None))
|
||||
elif project_id is not None:
|
||||
stmt = stmt.where(Note.project_id == project_id)
|
||||
in_project = Note.project_id == project_id
|
||||
if include_global_kinds:
|
||||
in_project = or_(
|
||||
in_project, Note.note_type.in_(GLOBAL_NOTE_TYPES)
|
||||
)
|
||||
stmt = stmt.where(in_project)
|
||||
# Narrow to records tagged to one System (subsystem/area). An
|
||||
# association filter, not a ranking signal — membership in the
|
||||
# candidate set, decided before scoring, like project_id above.
|
||||
|
||||
@@ -136,3 +136,61 @@ def compose_title(what: str, when_to_apply: str = "") -> str:
|
||||
from scribe.services.embeddings import trigger_title
|
||||
|
||||
return trigger_title(what, when_to_apply)
|
||||
|
||||
|
||||
def compose_body(insight: str, when_to_apply: str = "") -> str:
|
||||
"""The lesson body — the trigger line first, the insight after.
|
||||
|
||||
The mirror of `compose_title` on the other half of the document, and the
|
||||
reason the pair is what makes a lesson findable: `chunk_document` joins
|
||||
them as `{title}\\n{body}`, so a lesson composed here states WHEN IT
|
||||
APPLIES in the title and again in the first line of the body. That is the
|
||||
twice-in-a-short-document shape note #2485 measured as the only sharp one
|
||||
in the corpus, reached the way a snippet reaches it — by being in the text
|
||||
— rather than by a second document builder at embed time.
|
||||
|
||||
`**When to apply:**` rather than plain text: the body is the READABLE
|
||||
form, `data` is the queryable mirror, and `_BODY_TRIGGER_RE` reads this
|
||||
line back when the mirror is missing. Its markdown must therefore match
|
||||
what that pattern expects, which is why neither is written by hand
|
||||
anywhere else.
|
||||
|
||||
The insight goes in the body rather than being held out of the document.
|
||||
`rule_document` excludes a rule's `why` because long dated narrative made
|
||||
sixteen dev-logs land on the centroid of "development" — but that finding
|
||||
predates chunking (#280). A body over the budget is now split into several
|
||||
chunks, EACH prefixed with the title, so a lesson's story no longer
|
||||
averages itself into its trigger: it occupies its own vectors, and every
|
||||
one of them still carries the trigger in its prefix. Holding it out would
|
||||
cost the reader the only part that explains the insight and would buy a
|
||||
sharpness the chunker already provides.
|
||||
"""
|
||||
lines = []
|
||||
trigger = (when_to_apply or "").strip()
|
||||
if trigger:
|
||||
lines.append(f"**When to apply:** {trigger}")
|
||||
insight = (insight or "").strip()
|
||||
if insight:
|
||||
lines.append(insight)
|
||||
return "\n\n".join(lines)
|
||||
|
||||
|
||||
def lesson_document(
|
||||
what: str, when_to_apply: str = "", insight: str = "",
|
||||
) -> tuple[str, str]:
|
||||
"""The (title, body) a lesson is STORED — and therefore embedded — as.
|
||||
|
||||
One call so the two halves cannot be composed apart. A lesson whose title
|
||||
carried the trigger and whose body did not would embed as an ordinary
|
||||
note wearing a label, and nothing would report it: the record would look
|
||||
right in every listing and simply never be retrieved at the moment it
|
||||
applies.
|
||||
|
||||
Deliberately returns what is STORED, not a separate embed-time shape.
|
||||
Rules need `rule_document` because a rule keeps its trigger in a column
|
||||
and its title is a plain name, so the sharp document has to be synthesised
|
||||
for the ranker and exists nowhere else. A lesson follows the snippet
|
||||
instead — the stored record IS the sharp document — which is why nothing
|
||||
re-embeds and `CHUNKER_VERSION` does not move.
|
||||
"""
|
||||
return compose_title(what, when_to_apply), compose_body(insight, when_to_apply)
|
||||
|
||||
Reference in New Issue
Block a user