feat(lessons): the document shape is the stored record, and it travels (#3730)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / integration (push) Successful in 51s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m31s
CI & Build / Build & push image (push) Successful in 25s
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / integration (push) Successful in 51s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m31s
CI & Build / Build & push image (push) Successful in 25s
Milestone 385 step 3 — the step where the kind either works or is cosmetic.
THE DOCUMENT, and why there is no `lesson_document()` beside `rule_document()`
in embeddings. The step expected one. The difference is where the sharp shape
LIVES. A rule keeps its trigger in a column and its title is a plain name, so
`{title} — {trigger}` has to be synthesised at embed time and exists nowhere
else. A snippet — the only sharp record in the corpus by #2485's measurement,
0.153 top-to-second against 0.010–0.023 — gets there the other way: its STORED
title is already the join and its stored body already opens with the trigger,
so the ordinary `title\nbody` join IS the sharp document. Step 1 chose the
snippet route and step 2 built it, so `lessons.lesson_document` composes what
is STORED and the generic chunker does the rest.
The consequence the step asked about: `chunk_document` is untouched, so
CHUNKER_VERSION does not move and NOTHING re-embeds. The step's "Re-embed"
section describes a change this design does not make.
THE NARRATIVE stays in the body, departing from the step's instruction to keep
it out. `rule_document` excludes `why` because long dated narrative made
sixteen dev-logs land on the centroid of "development" — but that finding
predates chunking (#280). A body over budget is now split, and every chunk is
prefixed with the title, which for a lesson carries the trigger. The story
occupies its own vectors instead of averaging itself into the trigger's, and
each of those is still anchored to when the lesson applies. A guard asserts
exactly that. Holding the story out would cost the reader the only part that
explains the insight, to buy a sharpness the chunker already provides.
GLOBAL IN THE SEARCH is the real new code: `GLOBAL_NOTE_TYPES` and
`include_global_kinds` on `semantic_search_notes`, widening the PROJECT filter
alone. Off by default, because two callers depend on that filter holding — the
near-duplicate gate compares a record only against its own project on purpose,
and a globally visible kind there would let a lesson block an unrelated note's
create on a project its author never touched. It composes with `note_type`
rather than overriding it, so narrowing to snippets does not quietly acquire
lessons, and it changes nothing about the ACL: `notes_visibility_clause` still
gates every row.
Wired into the explicit MCP search only — the operator asked, and there is no
budget to spend. The unasked-for injection arms are step 5's subject (#3732)
and the legibility of a lesson appearing on a foreign project is step 7's
(#3734), so neither is turned on here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
@@ -136,3 +136,61 @@ def compose_title(what: str, when_to_apply: str = "") -> str:
|
||||
from scribe.services.embeddings import trigger_title
|
||||
|
||||
return trigger_title(what, when_to_apply)
|
||||
|
||||
|
||||
def compose_body(insight: str, when_to_apply: str = "") -> str:
|
||||
"""The lesson body — the trigger line first, the insight after.
|
||||
|
||||
The mirror of `compose_title` on the other half of the document, and the
|
||||
reason the pair is what makes a lesson findable: `chunk_document` joins
|
||||
them as `{title}\\n{body}`, so a lesson composed here states WHEN IT
|
||||
APPLIES in the title and again in the first line of the body. That is the
|
||||
twice-in-a-short-document shape note #2485 measured as the only sharp one
|
||||
in the corpus, reached the way a snippet reaches it — by being in the text
|
||||
— rather than by a second document builder at embed time.
|
||||
|
||||
`**When to apply:**` rather than plain text: the body is the READABLE
|
||||
form, `data` is the queryable mirror, and `_BODY_TRIGGER_RE` reads this
|
||||
line back when the mirror is missing. Its markdown must therefore match
|
||||
what that pattern expects, which is why neither is written by hand
|
||||
anywhere else.
|
||||
|
||||
The insight goes in the body rather than being held out of the document.
|
||||
`rule_document` excludes a rule's `why` because long dated narrative made
|
||||
sixteen dev-logs land on the centroid of "development" — but that finding
|
||||
predates chunking (#280). A body over the budget is now split into several
|
||||
chunks, EACH prefixed with the title, so a lesson's story no longer
|
||||
averages itself into its trigger: it occupies its own vectors, and every
|
||||
one of them still carries the trigger in its prefix. Holding it out would
|
||||
cost the reader the only part that explains the insight and would buy a
|
||||
sharpness the chunker already provides.
|
||||
"""
|
||||
lines = []
|
||||
trigger = (when_to_apply or "").strip()
|
||||
if trigger:
|
||||
lines.append(f"**When to apply:** {trigger}")
|
||||
insight = (insight or "").strip()
|
||||
if insight:
|
||||
lines.append(insight)
|
||||
return "\n\n".join(lines)
|
||||
|
||||
|
||||
def lesson_document(
|
||||
what: str, when_to_apply: str = "", insight: str = "",
|
||||
) -> tuple[str, str]:
|
||||
"""The (title, body) a lesson is STORED — and therefore embedded — as.
|
||||
|
||||
One call so the two halves cannot be composed apart. A lesson whose title
|
||||
carried the trigger and whose body did not would embed as an ordinary
|
||||
note wearing a label, and nothing would report it: the record would look
|
||||
right in every listing and simply never be retrieved at the moment it
|
||||
applies.
|
||||
|
||||
Deliberately returns what is STORED, not a separate embed-time shape.
|
||||
Rules need `rule_document` because a rule keeps its trigger in a column
|
||||
and its title is a plain name, so the sharp document has to be synthesised
|
||||
for the ranker and exists nowhere else. A lesson follows the snippet
|
||||
instead — the stored record IS the sharp document — which is why nothing
|
||||
re-embeds and `CHUNKER_VERSION` does not move.
|
||||
"""
|
||||
return compose_title(what, when_to_apply), compose_body(insight, when_to_apply)
|
||||
|
||||
Reference in New Issue
Block a user