CI & Build / Python lint (push) Successful in 8s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Failing after 1m9s
CI & Build / Build & push image (push) Skipped
#4243 fixed one door. Scribe has three semantic searches over three chunk tables, and all three collapsed chunk rows to the best one per record — each of them KNEW which passage earned the hit, and each dropped it. Every surface downstream then previewed the head of the document instead: a span the search had already scored lower, with nothing saying so. Mechanism, one place: - embeddings.record_best_chunk publishes {id: {index, text}} into `report`. Carried in `report`, NOT the return value: all three return list[tuple[float, Record]] and ~30 sites unpack that pair (lesson #4207). - semantic_search_rules and semantic_search_milestones now select chunk_index/chunk_text and publish the winner, as notes already did. semantic_search_milestones gains `report`, which it had no way to take. - services/text.matched_excerpt is the one choice of span, and excerpt_fields the one result block. Doors keep their own field names — the web renders `snippet`, MCP returns `excerpt` — because renaming a field a frontend reads is a different change from fixing what goes in it. Surfaces: - knowledge.query_knowledge, whose own comment calls it "the human's MAIN search surface", was `(note.body or "")[:200]` on every row alike. Now the matched passage on a search, the opening on a browse, and `snippet_is` saying which. KnowledgeView renders that snippet, so this was live. - search(content_type='milestone') gains `matched` — the plan body stays out, but the passage that matched comes along, because recognising a plan means recognising the part you asked about and a description written at the start need not mention it. - The auto-inject menu and the write-path prior-art menu put the passage under their line. Both were title-only, which answers "does this apply?" for a lesson or snippet (the trigger is IN the title) and not at all for an issue or dev-log. No fallback to the body's opening: on a menu that is preamble dressed as a reason, and once indented it cannot be told apart. Left alone deliberately: the rule arms. A rule hint already renders the rule's TRIGGER, which is written to answer exactly "does this apply to me" and beats a matched chunk at it; and that line's budget was measured at #3851. Adding a passage there would duplicate the trigger and spend the budget twice. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
102 lines
4.6 KiB
Python
102 lines
4.6 KiB
Python
"""Text shortening for surfaces that cannot show a record whole.
|
|
|
|
One function, in one place, because both doors shorten and two copies would
|
|
drift — and because the reasoning below is the part that matters and should
|
|
not have to be re-derived at each call site.
|
|
"""
|
|
|
|
|
|
def elide(text: str, budget: int) -> tuple[str, bool]:
|
|
"""Cut to `budget` characters from the MIDDLE, keeping both ends.
|
|
|
|
A head-only cut — `text[:800]` — decides what a reader sees by character
|
|
position, which is uncorrelated with what matters. Prose does not put its
|
|
conclusion first: a passage that opens with what was attempted and closes
|
|
with "so this shipped in 04775c3" loses exactly the sentence that answers
|
|
the question. Worse, the reader cannot tell: a truncation marker says that
|
|
something was removed, never whether it mattered, so the decision "should
|
|
I look deeper?" gets made on evidence selected by length.
|
|
|
|
So keep the opening (what this is about) AND the closing (where it landed),
|
|
and state in between how much went. Two thirds to the head because that is
|
|
where the subject is established; a conclusion needs less room to carry.
|
|
|
|
Returns (text, was_cut). A `budget` of 0 or less means no cut.
|
|
|
|
This is the fallback, not the goal. Where the system knows WHICH span of a
|
|
record is the relevant one — a semantic search knows exactly that, and
|
|
stores it (#4243) — show that span and say so. Reach for this only when
|
|
nothing identifies a better part than "all of it".
|
|
"""
|
|
if budget <= 0 or len(text) <= budget:
|
|
return text, False
|
|
head_len = max(1, budget * 2 // 3)
|
|
tail_len = max(1, budget - head_len)
|
|
omitted = len(text) - head_len - tail_len
|
|
head = text[:head_len].rstrip()
|
|
tail = text[-tail_len:].lstrip()
|
|
return f"{head}\n\n[… {omitted} characters omitted …]\n\n{tail}", True
|
|
|
|
|
|
# What a search result shows a reader, as one decision made in one place.
|
|
#
|
|
# Every door onto a semantic search faces the same question — which span of a
|
|
# record do I show someone deciding whether to open it? — and each used to
|
|
# answer it separately with a bare head cut of a different length: 240
|
|
# characters in the MCP search, 200 in the web's knowledge search. Both showed
|
|
# the document's OPENING, which is not the span that matched and not the span
|
|
# the ranking was built on.
|
|
#
|
|
# Doors keep their own field names (the web UI renders `snippet`, the MCP
|
|
# surface returns `excerpt`), because renaming a field a frontend consumes is
|
|
# a separate change from fixing which text goes in it. What they share is this
|
|
# function: the choice of span, and the obligation to say which span it is.
|
|
|
|
MATCHED_PASSAGE = "matched_passage"
|
|
BODY_OPENING = "body_opening"
|
|
|
|
|
|
def matched_excerpt(
|
|
body: str, chunk: dict | None, budget: int
|
|
) -> tuple[str, str, bool]:
|
|
"""Pick the span to show, and say which span it is.
|
|
|
|
`chunk` is a `report["best_chunk"]` entry from one of the semantic
|
|
searches — `{"index": int, "text": str}` — or None when the caller ran no
|
|
semantic search, passed no report, or the record has no embedding row.
|
|
|
|
Returns `(text, kind, was_cut)` where `kind` is MATCHED_PASSAGE or
|
|
BODY_OPENING. The kind is not decoration: a reader who cannot tell the
|
|
passage that earned the hit from the first paragraph of the document
|
|
cannot tell whether a thin-looking result is genuinely thin, and the whole
|
|
point of the excerpt is to support exactly that judgement.
|
|
|
|
A short record comes back whole either way — fragmenting a 200-character
|
|
note serves nobody, and its opening IS its content.
|
|
"""
|
|
passage = (chunk or {}).get("text") or ""
|
|
if passage.strip():
|
|
text, cut = elide(passage.strip(), budget)
|
|
return text, MATCHED_PASSAGE, cut
|
|
text, cut = elide(body or "", budget)
|
|
return text, BODY_OPENING, cut
|
|
|
|
|
|
def excerpt_fields(
|
|
body: str, chunk: dict | None, budget: int, *, key: str = "excerpt"
|
|
) -> dict:
|
|
"""`matched_excerpt` as the block a result row carries.
|
|
|
|
`key` names the text field so a door can keep the name its consumers
|
|
already read. The companion keys are derived from it, so a row never ends
|
|
up with an excerpt under one name and its label under another.
|
|
"""
|
|
text, kind, cut = matched_excerpt(body, chunk, budget)
|
|
out = {key: text, f"{key}_is": kind, "body_length": len(body or "")}
|
|
if cut or (kind == MATCHED_PASSAGE and len(text) < len(body or "")):
|
|
out["read_full"] = (
|
|
"This is one span of a longer record. Open it by id for the whole "
|
|
"thing rather than judging from what is here."
|
|
)
|
|
return out
|