feat(retrieval): every semantic search hands on the passage that matched
CI & Build / Python lint (push) Successful in 8s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Failing after 1m9s
CI & Build / Build & push image (push) Skipped
CI & Build / Python lint (push) Successful in 8s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Failing after 1m9s
CI & Build / Build & push image (push) Skipped
#4243 fixed one door. Scribe has three semantic searches over three chunk tables, and all three collapsed chunk rows to the best one per record — each of them KNEW which passage earned the hit, and each dropped it. Every surface downstream then previewed the head of the document instead: a span the search had already scored lower, with nothing saying so. Mechanism, one place: - embeddings.record_best_chunk publishes {id: {index, text}} into `report`. Carried in `report`, NOT the return value: all three return list[tuple[float, Record]] and ~30 sites unpack that pair (lesson #4207). - semantic_search_rules and semantic_search_milestones now select chunk_index/chunk_text and publish the winner, as notes already did. semantic_search_milestones gains `report`, which it had no way to take. - services/text.matched_excerpt is the one choice of span, and excerpt_fields the one result block. Doors keep their own field names — the web renders `snippet`, MCP returns `excerpt` — because renaming a field a frontend reads is a different change from fixing what goes in it. Surfaces: - knowledge.query_knowledge, whose own comment calls it "the human's MAIN search surface", was `(note.body or "")[:200]` on every row alike. Now the matched passage on a search, the opening on a browse, and `snippet_is` saying which. KnowledgeView renders that snippet, so this was live. - search(content_type='milestone') gains `matched` — the plan body stays out, but the passage that matched comes along, because recognising a plan means recognising the part you asked about and a description written at the start need not mention it. - The auto-inject menu and the write-path prior-art menu put the passage under their line. Both were title-only, which answers "does this apply?" for a lesson or snippet (the trigger is IN the title) and not at all for an issue or dev-log. No fallback to the body's opening: on a menu that is preamble dressed as a reason, and once indented it cannot be told apart. Left alone deliberately: the rule arms. A rule hint already renders the rule's TRIGGER, which is written to answer exactly "does this apply to me" and beats a matched chunk at it; and that line's budget was measured at #3851. Adding a passage there would duplicate the trigger and spend the budget twice. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
@@ -36,3 +36,66 @@ def elide(text: str, budget: int) -> tuple[str, bool]:
|
||||
head = text[:head_len].rstrip()
|
||||
tail = text[-tail_len:].lstrip()
|
||||
return f"{head}\n\n[… {omitted} characters omitted …]\n\n{tail}", True
|
||||
|
||||
|
||||
# What a search result shows a reader, as one decision made in one place.
|
||||
#
|
||||
# Every door onto a semantic search faces the same question — which span of a
|
||||
# record do I show someone deciding whether to open it? — and each used to
|
||||
# answer it separately with a bare head cut of a different length: 240
|
||||
# characters in the MCP search, 200 in the web's knowledge search. Both showed
|
||||
# the document's OPENING, which is not the span that matched and not the span
|
||||
# the ranking was built on.
|
||||
#
|
||||
# Doors keep their own field names (the web UI renders `snippet`, the MCP
|
||||
# surface returns `excerpt`), because renaming a field a frontend consumes is
|
||||
# a separate change from fixing which text goes in it. What they share is this
|
||||
# function: the choice of span, and the obligation to say which span it is.
|
||||
|
||||
MATCHED_PASSAGE = "matched_passage"
|
||||
BODY_OPENING = "body_opening"
|
||||
|
||||
|
||||
def matched_excerpt(
|
||||
body: str, chunk: dict | None, budget: int
|
||||
) -> tuple[str, str, bool]:
|
||||
"""Pick the span to show, and say which span it is.
|
||||
|
||||
`chunk` is a `report["best_chunk"]` entry from one of the semantic
|
||||
searches — `{"index": int, "text": str}` — or None when the caller ran no
|
||||
semantic search, passed no report, or the record has no embedding row.
|
||||
|
||||
Returns `(text, kind, was_cut)` where `kind` is MATCHED_PASSAGE or
|
||||
BODY_OPENING. The kind is not decoration: a reader who cannot tell the
|
||||
passage that earned the hit from the first paragraph of the document
|
||||
cannot tell whether a thin-looking result is genuinely thin, and the whole
|
||||
point of the excerpt is to support exactly that judgement.
|
||||
|
||||
A short record comes back whole either way — fragmenting a 200-character
|
||||
note serves nobody, and its opening IS its content.
|
||||
"""
|
||||
passage = (chunk or {}).get("text") or ""
|
||||
if passage.strip():
|
||||
text, cut = elide(passage.strip(), budget)
|
||||
return text, MATCHED_PASSAGE, cut
|
||||
text, cut = elide(body or "", budget)
|
||||
return text, BODY_OPENING, cut
|
||||
|
||||
|
||||
def excerpt_fields(
|
||||
body: str, chunk: dict | None, budget: int, *, key: str = "excerpt"
|
||||
) -> dict:
|
||||
"""`matched_excerpt` as the block a result row carries.
|
||||
|
||||
`key` names the text field so a door can keep the name its consumers
|
||||
already read. The companion keys are derived from it, so a row never ends
|
||||
up with an excerpt under one name and its label under another.
|
||||
"""
|
||||
text, kind, cut = matched_excerpt(body, chunk, budget)
|
||||
out = {key: text, f"{key}_is": kind, "body_length": len(body or "")}
|
||||
if cut or (kind == MATCHED_PASSAGE and len(text) < len(body or "")):
|
||||
out["read_full"] = (
|
||||
"This is one span of a longer record. Open it by id for the whole "
|
||||
"thing rather than judging from what is here."
|
||||
)
|
||||
return out
|
||||
|
||||
Reference in New Issue
Block a user