Files
FabledScribe/src/scribe/services/text.py
T
bvandeusenandClaude Opus 5 253fb974f3
CI & Build / Python lint (push) Successful in 8s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Failing after 1m9s
CI & Build / Build & push image (push) Skipped
feat(retrieval): every semantic search hands on the passage that matched
#4243 fixed one door. Scribe has three semantic searches over three chunk
tables, and all three collapsed chunk rows to the best one per record — each
of them KNEW which passage earned the hit, and each dropped it. Every surface
downstream then previewed the head of the document instead: a span the search
had already scored lower, with nothing saying so.

Mechanism, one place:
  - embeddings.record_best_chunk publishes {id: {index, text}} into `report`.
    Carried in `report`, NOT the return value: all three return
    list[tuple[float, Record]] and ~30 sites unpack that pair (lesson #4207).
  - semantic_search_rules and semantic_search_milestones now select
    chunk_index/chunk_text and publish the winner, as notes already did.
    semantic_search_milestones gains `report`, which it had no way to take.
  - services/text.matched_excerpt is the one choice of span, and
    excerpt_fields the one result block. Doors keep their own field names —
    the web renders `snippet`, MCP returns `excerpt` — because renaming a
    field a frontend reads is a different change from fixing what goes in it.

Surfaces:
  - knowledge.query_knowledge, whose own comment calls it "the human's MAIN
    search surface", was `(note.body or "")[:200]` on every row alike. Now the
    matched passage on a search, the opening on a browse, and `snippet_is`
    saying which. KnowledgeView renders that snippet, so this was live.
  - search(content_type='milestone') gains `matched` — the plan body stays
    out, but the passage that matched comes along, because recognising a plan
    means recognising the part you asked about and a description written at
    the start need not mention it.
  - The auto-inject menu and the write-path prior-art menu put the passage
    under their line. Both were title-only, which answers "does this apply?"
    for a lesson or snippet (the trigger is IN the title) and not at all for
    an issue or dev-log. No fallback to the body's opening: on a menu that is
    preamble dressed as a reason, and once indented it cannot be told apart.

Left alone deliberately: the rule arms. A rule hint already renders the rule's
TRIGGER, which is written to answer exactly "does this apply to me" and beats
a matched chunk at it; and that line's budget was measured at #3851. Adding a
passage there would duplicate the trigger and spend the budget twice.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
2026-09-21 09:44:10 -04:00

102 lines
4.6 KiB
Python

"""Text shortening for surfaces that cannot show a record whole.
One function, in one place, because both doors shorten and two copies would
drift — and because the reasoning below is the part that matters and should
not have to be re-derived at each call site.
"""
def elide(text: str, budget: int) -> tuple[str, bool]:
"""Cut to `budget` characters from the MIDDLE, keeping both ends.
A head-only cut — `text[:800]` — decides what a reader sees by character
position, which is uncorrelated with what matters. Prose does not put its
conclusion first: a passage that opens with what was attempted and closes
with "so this shipped in 04775c3" loses exactly the sentence that answers
the question. Worse, the reader cannot tell: a truncation marker says that
something was removed, never whether it mattered, so the decision "should
I look deeper?" gets made on evidence selected by length.
So keep the opening (what this is about) AND the closing (where it landed),
and state in between how much went. Two thirds to the head because that is
where the subject is established; a conclusion needs less room to carry.
Returns (text, was_cut). A `budget` of 0 or less means no cut.
This is the fallback, not the goal. Where the system knows WHICH span of a
record is the relevant one — a semantic search knows exactly that, and
stores it (#4243) — show that span and say so. Reach for this only when
nothing identifies a better part than "all of it".
"""
if budget <= 0 or len(text) <= budget:
return text, False
head_len = max(1, budget * 2 // 3)
tail_len = max(1, budget - head_len)
omitted = len(text) - head_len - tail_len
head = text[:head_len].rstrip()
tail = text[-tail_len:].lstrip()
return f"{head}\n\n[… {omitted} characters omitted …]\n\n{tail}", True
# What a search result shows a reader, as one decision made in one place.
#
# Every door onto a semantic search faces the same question — which span of a
# record do I show someone deciding whether to open it? — and each used to
# answer it separately with a bare head cut of a different length: 240
# characters in the MCP search, 200 in the web's knowledge search. Both showed
# the document's OPENING, which is not the span that matched and not the span
# the ranking was built on.
#
# Doors keep their own field names (the web UI renders `snippet`, the MCP
# surface returns `excerpt`), because renaming a field a frontend consumes is
# a separate change from fixing which text goes in it. What they share is this
# function: the choice of span, and the obligation to say which span it is.
MATCHED_PASSAGE = "matched_passage"
BODY_OPENING = "body_opening"
def matched_excerpt(
body: str, chunk: dict | None, budget: int
) -> tuple[str, str, bool]:
"""Pick the span to show, and say which span it is.
`chunk` is a `report["best_chunk"]` entry from one of the semantic
searches — `{"index": int, "text": str}` — or None when the caller ran no
semantic search, passed no report, or the record has no embedding row.
Returns `(text, kind, was_cut)` where `kind` is MATCHED_PASSAGE or
BODY_OPENING. The kind is not decoration: a reader who cannot tell the
passage that earned the hit from the first paragraph of the document
cannot tell whether a thin-looking result is genuinely thin, and the whole
point of the excerpt is to support exactly that judgement.
A short record comes back whole either way — fragmenting a 200-character
note serves nobody, and its opening IS its content.
"""
passage = (chunk or {}).get("text") or ""
if passage.strip():
text, cut = elide(passage.strip(), budget)
return text, MATCHED_PASSAGE, cut
text, cut = elide(body or "", budget)
return text, BODY_OPENING, cut
def excerpt_fields(
body: str, chunk: dict | None, budget: int, *, key: str = "excerpt"
) -> dict:
"""`matched_excerpt` as the block a result row carries.
`key` names the text field so a door can keep the name its consumers
already read. The companion keys are derived from it, so a row never ends
up with an excerpt under one name and its label under another.
"""
text, kind, cut = matched_excerpt(body, chunk, budget)
out = {key: text, f"{key}_is": kind, "body_length": len(body or "")}
if cut or (kind == MATCHED_PASSAGE and len(text) < len(body or "")):
out["read_full"] = (
"This is one span of a longer record. Open it by id for the whole "
"thing rather than judging from what is here."
)
return out