feat(retrieval): every semantic search hands on the passage that matched
CI & Build / Python lint (push) Successful in 8s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Failing after 1m9s
CI & Build / Build & push image (push) Skipped

#4243 fixed one door. Scribe has three semantic searches over three chunk
tables, and all three collapsed chunk rows to the best one per record — each
of them KNEW which passage earned the hit, and each dropped it. Every surface
downstream then previewed the head of the document instead: a span the search
had already scored lower, with nothing saying so.

Mechanism, one place:
  - embeddings.record_best_chunk publishes {id: {index, text}} into `report`.
    Carried in `report`, NOT the return value: all three return
    list[tuple[float, Record]] and ~30 sites unpack that pair (lesson #4207).
  - semantic_search_rules and semantic_search_milestones now select
    chunk_index/chunk_text and publish the winner, as notes already did.
    semantic_search_milestones gains `report`, which it had no way to take.
  - services/text.matched_excerpt is the one choice of span, and
    excerpt_fields the one result block. Doors keep their own field names —
    the web renders `snippet`, MCP returns `excerpt` — because renaming a
    field a frontend reads is a different change from fixing what goes in it.

Surfaces:
  - knowledge.query_knowledge, whose own comment calls it "the human's MAIN
    search surface", was `(note.body or "")[:200]` on every row alike. Now the
    matched passage on a search, the opening on a browse, and `snippet_is`
    saying which. KnowledgeView renders that snippet, so this was live.
  - search(content_type='milestone') gains `matched` — the plan body stays
    out, but the passage that matched comes along, because recognising a plan
    means recognising the part you asked about and a description written at
    the start need not mention it.
  - The auto-inject menu and the write-path prior-art menu put the passage
    under their line. Both were title-only, which answers "does this apply?"
    for a lesson or snippet (the trigger is IN the title) and not at all for
    an issue or dev-log. No fallback to the body's opening: on a menu that is
    preamble dressed as a reason, and once indented it cannot be told apart.

Left alone deliberately: the rule arms. A rule hint already renders the rule's
TRIGGER, which is written to answer exactly "does this apply to me" and beats
a matched chunk at it; and that line's budget was measured at #3851. Adding a
passage there would duplicate the trigger and spend the budget twice.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
2026-09-21 09:44:10 -04:00
co-authored by Claude Opus 5
parent 6abedb0168
commit 253fb974f3
8 changed files with 527 additions and 51 deletions
+63
View File
@@ -36,3 +36,66 @@ def elide(text: str, budget: int) -> tuple[str, bool]:
head = text[:head_len].rstrip()
tail = text[-tail_len:].lstrip()
return f"{head}\n\n[… {omitted} characters omitted …]\n\n{tail}", True
# What a search result shows a reader, as one decision made in one place.
#
# Every door onto a semantic search faces the same question — which span of a
# record do I show someone deciding whether to open it? — and each used to
# answer it separately with a bare head cut of a different length: 240
# characters in the MCP search, 200 in the web's knowledge search. Both showed
# the document's OPENING, which is not the span that matched and not the span
# the ranking was built on.
#
# Doors keep their own field names (the web UI renders `snippet`, the MCP
# surface returns `excerpt`), because renaming a field a frontend consumes is
# a separate change from fixing which text goes in it. What they share is this
# function: the choice of span, and the obligation to say which span it is.
MATCHED_PASSAGE = "matched_passage"
BODY_OPENING = "body_opening"
def matched_excerpt(
body: str, chunk: dict | None, budget: int
) -> tuple[str, str, bool]:
"""Pick the span to show, and say which span it is.
`chunk` is a `report["best_chunk"]` entry from one of the semantic
searches — `{"index": int, "text": str}` — or None when the caller ran no
semantic search, passed no report, or the record has no embedding row.
Returns `(text, kind, was_cut)` where `kind` is MATCHED_PASSAGE or
BODY_OPENING. The kind is not decoration: a reader who cannot tell the
passage that earned the hit from the first paragraph of the document
cannot tell whether a thin-looking result is genuinely thin, and the whole
point of the excerpt is to support exactly that judgement.
A short record comes back whole either way — fragmenting a 200-character
note serves nobody, and its opening IS its content.
"""
passage = (chunk or {}).get("text") or ""
if passage.strip():
text, cut = elide(passage.strip(), budget)
return text, MATCHED_PASSAGE, cut
text, cut = elide(body or "", budget)
return text, BODY_OPENING, cut
def excerpt_fields(
body: str, chunk: dict | None, budget: int, *, key: str = "excerpt"
) -> dict:
"""`matched_excerpt` as the block a result row carries.
`key` names the text field so a door can keep the name its consumers
already read. The companion keys are derived from it, so a row never ends
up with an excerpt under one name and its label under another.
"""
text, kind, cut = matched_excerpt(body, chunk, budget)
out = {key: text, f"{key}_is": kind, "body_length": len(body or "")}
if cut or (kind == MATCHED_PASSAGE and len(text) < len(body or "")):
out["read_full"] = (
"This is one span of a longer record. Open it by id for the whole "
"thing rather than judging from what is here."
)
return out