feat(lessons): the document shape is the stored record, and it travels (#3730)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / integration (push) Successful in 51s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m31s
CI & Build / Build & push image (push) Successful in 25s

Milestone 385 step 3 — the step where the kind either works or is cosmetic.

THE DOCUMENT, and why there is no `lesson_document()` beside `rule_document()`
in embeddings. The step expected one. The difference is where the sharp shape
LIVES. A rule keeps its trigger in a column and its title is a plain name, so
`{title} — {trigger}` has to be synthesised at embed time and exists nowhere
else. A snippet — the only sharp record in the corpus by #2485's measurement,
0.153 top-to-second against 0.010–0.023 — gets there the other way: its STORED
title is already the join and its stored body already opens with the trigger,
so the ordinary `title\nbody` join IS the sharp document. Step 1 chose the
snippet route and step 2 built it, so `lessons.lesson_document` composes what
is STORED and the generic chunker does the rest.

The consequence the step asked about: `chunk_document` is untouched, so
CHUNKER_VERSION does not move and NOTHING re-embeds. The step's "Re-embed"
section describes a change this design does not make.

THE NARRATIVE stays in the body, departing from the step's instruction to keep
it out. `rule_document` excludes `why` because long dated narrative made
sixteen dev-logs land on the centroid of "development" — but that finding
predates chunking (#280). A body over budget is now split, and every chunk is
prefixed with the title, which for a lesson carries the trigger. The story
occupies its own vectors instead of averaging itself into the trigger's, and
each of those is still anchored to when the lesson applies. A guard asserts
exactly that. Holding the story out would cost the reader the only part that
explains the insight, to buy a sharpness the chunker already provides.

GLOBAL IN THE SEARCH is the real new code: `GLOBAL_NOTE_TYPES` and
`include_global_kinds` on `semantic_search_notes`, widening the PROJECT filter
alone. Off by default, because two callers depend on that filter holding — the
near-duplicate gate compares a record only against its own project on purpose,
and a globally visible kind there would let a lesson block an unrelated note's
create on a project its author never touched. It composes with `note_type`
rather than overriding it, so narrowing to snippets does not quietly acquire
lessons, and it changes nothing about the ACL: `notes_visibility_clause` still
gates every row.

Wired into the explicit MCP search only — the operator asked, and there is no
budget to spend. The unasked-for injection arms are step 5's subject (#3732)
and the legibility of a lesson appearing on a foreign project is step 7's
(#3734), so neither is turned on here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
2026-09-18 16:13:15 -04:00
co-authored by Claude Opus 5
parent d127d48c14
commit 1361ed7200
7 changed files with 424 additions and 1 deletions
+11
View File
@@ -140,6 +140,9 @@ async def search(
enter_project) — otherwise this searches across ALL projects and
bleeds unrelated work into the result set. 0 = search everything
(use only when you genuinely want a cross-project sweep).
A LESSON is the exception and arrives whatever the scope: the kind
records an insight that transfers, so it is reachable from a
project it was not written on.
system_id: Narrow to records tagged to one System (a named
subsystem/area — enter_project lists them). Use when investigating
a specific subsystem: it cuts the candidates to records someone
@@ -167,6 +170,14 @@ async def search(
uid, q, limit=limit, is_task=is_task,
project_id=project_id or None,
system_id=system_id or None,
# A LESSON is reachable from any project (milestone 385). The kind
# exists to carry an insight to the next project, so a project filter
# that hid it would hide it precisely where it is worth having. Only
# the project filter widens — everything else about the scoping holds,
# and a caller narrowing by `content_type` still gets what it asked
# for. This is the explicit search, where the operator asked; the
# unasked-for arms decide their own budget separately.
include_global_kinds=True,
# An explicit search reaches everything the operator may read, including
# records shared with them one-to-one.
scope="read",
+39 -1
View File
@@ -508,6 +508,25 @@ async def upsert_note_embedding(
logger.warning("Failed to persist embedding for note %d", note_id, exc_info=True)
# Kinds that belong to no single project, and are therefore reachable from a
# project-scoped search of a DIFFERENT project when a caller asks for them
# (milestone 385 step 3).
#
# A lesson is the whole reason this exists. "A better way to think about this
# problem" is not true only where it was learned, and a lesson confined to its
# origin project would be unreachable exactly where it is most useful — on the
# next project, which is the case the kind was created for (#3727).
# `semantic_search_rules` has always had this property: it scopes by OWNERSHIP
# rather than by what binds a given project, because "is there a rule about
# this" is a question asked across a whole rulebook. A lesson asks the same
# kind of question.
#
# Spelled as a literal rather than imported from `services.lessons`, which
# imports `trigger_title` from this module and would close a cycle. A guard
# pins the two equal instead.
GLOBAL_NOTE_TYPES: tuple[str, ...] = ("lesson",)
# Both searches rank WITHOUT the threshold and apply it in Python, so the best
# rejected score stays observable (#3670). The qualifying set is provably
# unchanged: rows arrive ordered by distance ascending, so every above-bar row
@@ -536,6 +555,7 @@ async def semantic_search_notes(
note_type: str | Sequence[str] | None = None,
task_kind: str | Sequence[str] | None = None,
orphan_only: bool = False,
include_global_kinds: bool = False,
scope: str = "own",
demote_superseded: bool = True,
system_id: int | None = None,
@@ -570,6 +590,19 @@ async def semantic_search_notes(
alone can express it. With `note_type="note", task_kind="issue"` a caller
gets fixed problems and durable notes without the open to-do list.
`include_global_kinds` lets a project-scoped search ALSO reach the kinds in
GLOBAL_NOTE_TYPES — records that belong to no single project — so a lesson
written on one project is found from another. It widens the PROJECT filter
only: a caller that also passes `note_type` still gets exactly the kinds it
asked for, so narrowing to snippets does not quietly acquire lessons.
Off by default, because two callers depend on the project filter holding.
The near-duplicate gate compares a record only against its own project on
purpose, and a globally-visible kind would let a lesson block an unrelated
note's create on a project its author never touched. Ordinary note recall
is project-scoped for the same reason — the point of the carve-out is that
ONE kind escapes, not that scoping is weaker.
`scope` ("own" | "browse" | "read", see access.notes_visibility_clause)
decides how far this may see. It exists because this one function serves
three different kinds of act: an explicit search, which should reach
@@ -631,7 +664,12 @@ async def semantic_search_notes(
if orphan_only:
stmt = stmt.where(Note.project_id.is_(None))
elif project_id is not None:
stmt = stmt.where(Note.project_id == project_id)
in_project = Note.project_id == project_id
if include_global_kinds:
in_project = or_(
in_project, Note.note_type.in_(GLOBAL_NOTE_TYPES)
)
stmt = stmt.where(in_project)
# Narrow to records tagged to one System (subsystem/area). An
# association filter, not a ranking signal — membership in the
# candidate set, decided before scoring, like project_id above.
+58
View File
@@ -136,3 +136,61 @@ def compose_title(what: str, when_to_apply: str = "") -> str:
from scribe.services.embeddings import trigger_title
return trigger_title(what, when_to_apply)
def compose_body(insight: str, when_to_apply: str = "") -> str:
"""The lesson body — the trigger line first, the insight after.
The mirror of `compose_title` on the other half of the document, and the
reason the pair is what makes a lesson findable: `chunk_document` joins
them as `{title}\\n{body}`, so a lesson composed here states WHEN IT
APPLIES in the title and again in the first line of the body. That is the
twice-in-a-short-document shape note #2485 measured as the only sharp one
in the corpus, reached the way a snippet reaches it — by being in the text
— rather than by a second document builder at embed time.
`**When to apply:**` rather than plain text: the body is the READABLE
form, `data` is the queryable mirror, and `_BODY_TRIGGER_RE` reads this
line back when the mirror is missing. Its markdown must therefore match
what that pattern expects, which is why neither is written by hand
anywhere else.
The insight goes in the body rather than being held out of the document.
`rule_document` excludes a rule's `why` because long dated narrative made
sixteen dev-logs land on the centroid of "development" — but that finding
predates chunking (#280). A body over the budget is now split into several
chunks, EACH prefixed with the title, so a lesson's story no longer
averages itself into its trigger: it occupies its own vectors, and every
one of them still carries the trigger in its prefix. Holding it out would
cost the reader the only part that explains the insight and would buy a
sharpness the chunker already provides.
"""
lines = []
trigger = (when_to_apply or "").strip()
if trigger:
lines.append(f"**When to apply:** {trigger}")
insight = (insight or "").strip()
if insight:
lines.append(insight)
return "\n\n".join(lines)
def lesson_document(
what: str, when_to_apply: str = "", insight: str = "",
) -> tuple[str, str]:
"""The (title, body) a lesson is STORED — and therefore embedded — as.
One call so the two halves cannot be composed apart. A lesson whose title
carried the trigger and whose body did not would embed as an ordinary
note wearing a label, and nothing would report it: the record would look
right in every listing and simply never be retrieved at the moment it
applies.
Deliberately returns what is STORED, not a separate embed-time shape.
Rules need `rule_document` because a rule keeps its trigger in a column
and its title is a plain name, so the sharp document has to be synthesised
for the ranker and exists nowhere else. A lesson follows the snippet
instead — the stored record IS the sharp document — which is why nothing
re-embeds and `CHUNKER_VERSION` does not move.
"""
return compose_title(what, when_to_apply), compose_body(insight, when_to_apply)