fix(lessons): a derived mirror survives the generic note door, by kind not by name (#3734)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / integration (push) Successful in 49s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m31s
CI & Build / Build & push image (push) Successful in 23s

Groundwork for step 7, and a data-integrity fix in its own right.

Two kinds keep a queryable mirror in `notes.data` derived from their body:
snippets and, since milestone 385, lessons. Every read prefers the mirror —
deliberately, because parsing markdown to answer what an index can answer is
how a hot path rots. So a write that moves the body must move the mirror.

#3128 found that hole for snippets and plugged it with a hard-coded
`if note.note_type == SNIPPET_NOTE_TYPE`. The plug was correct and did not
generalise: lessons arrived with the same design and none of the protection,
which is precisely the "don't add a fourth instance" defect #3734 was told to
avoid.

The cost is higher for a lesson. A stale snippet mirror reports the wrong
path. A stale lesson mirror reports the wrong TRIGGER, and the trigger is the
whole retrieval story — the lesson goes on firing for the situation it used
to name while displaying the one it now names. Silent, and confident.

So `update_note` now dispatches through `_mirror_recomposers()`, a
note_type -> recomposer table. A kind with a derived mirror is covered by
registering it, not by someone remembering to widen an if.

`lessons.recompose_data` is the lesson's entry. It recovers the subject with
`embeddings.untrigger_title` — new, and deliberately placed beside the join it
inverts rather than in the caller that wanted it, because a separator spelled
in two files is a separator that will one day be changed in one of them
(#3207). `TRIGGER_SEP` is now the one spelling, and `parse_snippet_fields`
uses it too; it had the third copy inline.

The two inverses stay distinct on purpose: a snippet partitions at the first
separator (its name is a symbol), a lesson strips an exact known suffix (its
subject may legitimately contain a dash). Different algorithms, one constant,
so they cannot disagree about where the seam is.

Provenance is DROPPED when the body drops it, which is the opposite call from
a snippet's `verification` — that is carried because it was never in the body
to delete. The body is the authority; carrying a value the reader just removed
is the failure the recompose exists to prevent.

Tests: test_snippet_mirror_generic_door.py becomes
test_derived_mirror_generic_door.py, since the concern is now plural. The
registry property is asserted directly (every kind with a mirror is in the
table; the dispatch names no kind inline), plus the lesson cases and the
join/inverse round-trip. `fake_lesson` moves to tests/helpers.py — it existed
in test_lesson_surfacing.py and a second copy was about to be written — and
gains the explicit `None`s `fake_snippet` carries, because update_note reads
`verify_with` and a MagicMock is truthy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
2026-09-19 13:58:54 -04:00
co-authored by Claude Opus 5
parent 1ade956cd5
commit 1252d0e305
8 changed files with 462 additions and 137 deletions
+34 -1
View File
@@ -203,6 +203,12 @@ def embedding_text(title: str | None, body: str | None) -> str:
return f"{title}\n{body}".strip() if body else title
# The join between a situation-keyed record's subject and its trigger. A
# CONSTANT because `untrigger_title` below has to spell the same thing to undo
# it, and two literals that must match are one edit away from not matching.
TRIGGER_SEP = ""
def trigger_title(subject: str | None, trigger: str | None) -> str:
"""`{subject}{trigger}` — the title half of a situation-keyed document.
@@ -226,10 +232,37 @@ def trigger_title(subject: str | None, trigger: str | None) -> str:
subject = (subject or "").strip()
trigger = (trigger or "").strip()
if subject and trigger:
return f"{subject}{trigger}"
return f"{subject}{TRIGGER_SEP}{trigger}"
return subject or trigger
def untrigger_title(title: str | None, trigger: str | None) -> str:
"""The subject back out of a `trigger_title` — the inverse of the join.
Kept HERE, beside the join, for the reason the join itself was
consolidated: a separator spelled in two files is a separator that will one
day be changed in one of them. #3207 records the shape — derive it before
the third copy — and an inverse written in a caller is that third copy
wearing a different name.
Needs the trigger passed in rather than guessing at the separator, because
a subject may legitimately contain an em dash. Given the trigger, the
suffix is exact and the split cannot be wrong.
Degrades to the whole title when the suffix is absent — a record written
before the join existed, or one with no trigger yet, still answers with
something a human recognises rather than with "".
"""
title = (title or "").strip()
trigger = (trigger or "").strip()
if not trigger:
return title
suffix = f"{TRIGGER_SEP}{trigger}"
if title.endswith(suffix):
return title[: -len(suffix)].strip()
return title
# --- chunking (#280): the document shape ------------------------------------
#
# bge-small reads at most 512 tokens and fastembed silently truncates the rest,
+40
View File
@@ -317,6 +317,46 @@ def compose_data(
return data
def recompose_data(note) -> dict:
"""Rebuild a lesson's `data` mirror from its own title and body.
For the GENERIC note door. `update_lesson` composes the mirror itself from
the merged field set and never needs this; a plain `update_note(body=...)`
has no idea the mirror exists and would leave it behind.
THE COST OF LEAVING IT BEHIND IS HIGHER HERE THAN FOR A SNIPPET. A stale
snippet mirror reports the wrong path. A stale lesson mirror reports the
wrong TRIGGER — and `lesson_trigger` prefers the mirror, so the lesson goes
on being retrieved for the situation it used to name while displaying the
one it now names. The trigger is the entire retrieval story (step 3), so
that is not a degraded record; it is a record that fires at the wrong
moment and looks right when it does.
The body is the authority and the mirror is derived — already this file's
rule. This is its enforcement on the path that bypasses `update_lesson`.
The subject comes back out of the title through `untrigger_title`, the
inverse of the join that composed it, rather than by splitting on a
separator spelled a second time here.
"""
from scribe.services.embeddings import untrigger_title
body = getattr(note, "body", None) or ""
trigger_match = _BODY_TRIGGER_RE.search(body)
trigger = trigger_match.group(1).strip() if trigger_match else ""
what = untrigger_title(getattr(note, "title", None), trigger)
# Sources through the normal read, which already falls back body →
# arose_from_id. A body edit that drops the provenance line should drop
# the mirror's copy too: the body is the authority, and carrying a value
# the reader just deleted is the failure this function exists to prevent.
sources_match = _BODY_SOURCES_RE.search(body)
sources = (
normalize_sources(_ID_RE.findall(sources_match.group(1)))
if sources_match else []
)
return compose_data(what, trigger, sources)
async def create_lesson(
user_id: int,
*,
+48 -22
View File
@@ -1,6 +1,6 @@
import logging
import re
from collections.abc import Iterable
from collections.abc import Callable, Iterable
from datetime import date, datetime, timezone
from sqlalchemy import func, or_, select, text
@@ -11,14 +11,44 @@ from scribe.models.base import iso
logger = logging.getLogger(__name__)
# The fields `snippets.parse_snippet_fields` reads. Writing any of them can
# change what a snippet's derived `data` mirror should say, so update_note
# recomposes the mirror when one moves. Kept here as a set of NAMES rather
# than imported, because it describes update_note's own `fields` dict, not the
# parser's signature.
# The fields a derived `data` mirror is parsed out of. Writing any of them can
# change what the mirror should say, so update_note recomposes it when one
# moves. Kept here as a set of NAMES rather than imported, because it
# describes update_note's own `fields` dict, not any parser's signature.
_PARSED_FROM_BODY = frozenset({"title", "body", "tags"})
def _mirror_recomposers() -> dict[str, Callable[[Note], dict]]:
"""note_type -> the function that rebuilds that kind's derived `data`.
A TABLE rather than a chain of `if note_type == ...`, because the previous
shape tested one constant and the next kind with a derived mirror was
silently not covered — which is exactly what happened: the snippet guard
(#3128) was hard-coded, and lessons arrived in milestone 385 with the same
body-is-authority/`data`-is-mirror design and none of the protection.
The failure is invisible from here. Nothing raises, nothing logs; the row
simply keeps answering queries from a mirror that no longer matches its
body, and every surface that prefers the mirror — which is all of them, by
design, because parsing markdown to answer what an index can answer is how
a hot path rots — reports the old value confidently.
Imported inside the function, not at module scope: both services call back
into this module (`update_snippet`/`update_lesson` -> `update_note`), so a
top-level import is a cycle.
"""
from scribe.services.lessons import (
LESSON_NOTE_TYPE, recompose_data as _lesson_mirror,
)
from scribe.services.snippets import (
SNIPPET_NOTE_TYPE, recompose_data as _snippet_mirror,
)
return {
SNIPPET_NOTE_TYPE: _snippet_mirror,
LESSON_NOTE_TYPE: _lesson_mirror,
}
# Text fields where EMPTY MEANS NULL (milestone 317). The sweep's whole signal
# is `verify_with IS NULL` = "this is a decision, there is nothing to go and
# check". An empty string that is not NULL makes a norm look like a constraint
@@ -604,23 +634,19 @@ async def update_note(
# costs exactly what the sweep exists to catch.
if note.verify_with != check_before:
note.verified_at = None
# A snippet's `data` is DERIVED from its body — so a write that moves
# the body through this generic door must move the mirror with it
# (#3128). Without this, PATCH /api/notes/<snippet_id> {body} left the
# mirror behind, and snippet_fields PREFERS the mirror: the row went on
# reporting its old repo/path/symbol to prior-art recall while showing
# its new body. `update_snippet` composes the mirror itself and passes
# it explicitly, so an explicit `data` always wins — the caller that
# knows the field set beats the one that can only re-read the body.
# Some kinds derive `data` from their body — so a write that moves the
# body through this generic door must move the mirror with it (#3128).
# Without this, PATCH /api/notes/<id> {body} left the mirror behind,
# and every read PREFERS the mirror: a snippet went on reporting its
# old repo/path/symbol to prior-art recall while showing its new body,
# and a lesson would go on being retrieved for the situation it used to
# name. The kind's own updater composes the mirror itself and passes it
# explicitly, so an explicit `data` always wins — the caller that knows
# the field set beats the one that can only re-read the body.
if "data" not in fields and not _PARSED_FROM_BODY.isdisjoint(fields):
# Imported here, not at module scope: services/snippets.py calls
# back into this module (update_snippet -> update_note), so a
# top-level import is a cycle.
from scribe.services.snippets import (
SNIPPET_NOTE_TYPE, recompose_data,
)
if note.note_type == SNIPPET_NOTE_TYPE:
note.data = recompose_data(note)
recompose = _mirror_recomposers().get(note.note_type or "")
if recompose is not None:
note.data = recompose(note)
# Auto-set lifecycle timestamps on status transitions
if "status" in fields:
_now = datetime.now(timezone.utc)
+13 -1
View File
@@ -275,9 +275,21 @@ def parse_snippet_fields(
``locations`` is a list of {repo,path,symbol}; ``repo``/``path``/``symbol``
mirror the FIRST location for back-compat with the single-location callers."""
from scribe.services.embeddings import TRIGGER_SEP
title = title or ""
body = body or ""
name, _, when_from_title = title.partition("")
# The inverse of `embeddings.trigger_title`, at the SEPARATOR it composed
# with — imported rather than spelled again, because a separator written
# in two files is a separator that will one day be changed in one of them.
#
# `partition` rather than the `untrigger_title` a lesson uses: that one is
# handed the trigger and strips an exact suffix, which a lesson needs
# because its subject may legitimately contain a dash. A snippet's name is
# a symbol, so the first separator is the right split and no trigger has
# to be known in advance. Two inverses, suited to their callers; one
# constant, so they cannot disagree about where the seam is.
name, _, when_from_title = title.partition(TRIGGER_SEP)
fields = {
"name": name.strip(),
"when_to_use": when_from_title.strip(),