fix(search): show the passage that matched, not the opening of the body (#4243)
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Failing after 1m5s
CI & Build / Build & push image (push) Skipped
CI & Build / integration (push) Successful in 46s

Raised by the operator: are we limiting what comes back by character count,
and how do we verify the pertinent part is the part displayed?

We were not. mcp/tools/search.py sent (note.body or "")[:240] — a head cut,
with no marker that anything had been removed, so a 240-character preview of
a 4000-character record was indistinguishable from a complete short one.

The opening is the wrong span. The match is semantic and per chunk, and
semantic_search_notes collapses to best-chunk-per-note — its own comment at
the collapse says "the first appearance of a note is its best chunk". So the
system identified the passage that earned the hit and then discarded it:
select(Note, distance) kept no chunk column. A record could rank first on its
sixth paragraph, be previewed by its first, and be judged irrelevant on a
span the search had already scored lower. That biases against long records,
and it is self-concealing — the caller who does not open it never learns the
preview was misleading.

  - embeddings: chunk_index/chunk_text ride along in the select, and the
    collapse records the winner in report["best_chunk"]. Carried in `report`,
    NOT by widening the return tuple: ten callers unpack (score, note) at
    ~18 sites and nothing would catch the misses (lesson #4207). `report` is
    the side-channel this function already uses for best_available_score.
  - search(): excerpt / excerpt_is / body_length, and read_full when there is
    more. A caller that cannot tell a matched passage from a document opening
    cannot judge whether to look deeper, which is the only decision the field
    supports.

elide() moves to services/text.py so both callers share one copy, and it
keeps BOTH ends with a stated gap — it is the fallback for when nothing
identifies a better span than "all of it", not the goal.

Also fixes a guard that produced a false failure on the previous commit:
test_pull_telemetry checked `"project_id: int = 0" in body.split("\n")[0]`,
which sees only the first line, so wrapping get_task's signature over four
lines made it report a function that does take the project as one that does
not. Parsed with ast now, and proven to still reject an absent or
wrongly-typed parameter rather than being appeased by reflowing the code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
2026-09-21 08:50:16 -04:00
co-authored by Claude Opus 5
parent 4f2977b848
commit fdc07f2a2b
7 changed files with 245 additions and 50 deletions
+35 -4
View File
@@ -686,7 +686,16 @@ async def semantic_search_notes(
# to the note's owner, so filtering it would pin every scope to "own"
# and leave shared records unreachable by meaning.
stmt = (
select(Note, distance.label("distance"))
# chunk_index/chunk_text ride along so the collapse below can
# say WHICH passage matched. Without them the caller is left
# previewing the head of the body — a span this query has
# already determined is not why the record ranked (#4243).
select(
Note,
distance.label("distance"),
NoteEmbedding.chunk_index,
NoteEmbedding.chunk_text,
)
.select_from(NoteEmbedding)
.join(Note, NoteEmbedding.note_id == Note.id)
.where(
@@ -774,11 +783,23 @@ async def semantic_search_notes(
# Recover similarity (1 - distance); order stays highest-first.
scored: list[tuple[float, Note]] = []
seen: set[int] = set()
for note, dist in rows:
# The winning row IS the best chunk, by the ordering above — so this is the
# one place that knows which passage earned the hit. Kept beside the score
# rather than returned with it: the return type is list[tuple[float, Note]]
# and ten callers unpack it at ~18 sites, so widening the tuple would be an
# interface change to every one of them with nothing to catch the misses
# (lesson #4207). `report` is the side-channel this function already uses
# for best_available_score.
best_chunk: dict[int, dict] = {}
for note, dist, chunk_index, chunk_text in rows:
if int(note.id) in seen:
continue
seen.add(int(note.id))
scored.append((1.0 - float(dist), note))
best_chunk[int(note.id)] = {
"index": int(chunk_index),
"text": chunk_text or "",
}
# The best score anything reached, bar or no bar. Recorded BEFORE the
# filter because a call that returns nothing is exactly when it matters.
if report is not None:
@@ -797,8 +818,18 @@ async def semantic_search_notes(
report["best_available_id"] = int(best[1].id) if best else None
scored = [pair for pair in scored if pair[0] >= threshold]
if not demote_superseded:
return scored[:limit]
return await _apply_supersession_penalty(scored, limit)
final = scored[:limit]
else:
final = await _apply_supersession_penalty(scored, limit)
# Only for what actually came back, so a caller can key straight off the
# results without carrying chunks for records it never saw.
if report is not None:
report["best_chunk"] = {
int(n.id): best_chunk[int(n.id)]
for _s, n in final
if int(n.id) in best_chunk
}
return final
async def backfill_note_embeddings() -> None:
+38
View File
@@ -0,0 +1,38 @@
"""Text shortening for surfaces that cannot show a record whole.
One function, in one place, because both doors shorten and two copies would
drift — and because the reasoning below is the part that matters and should
not have to be re-derived at each call site.
"""
def elide(text: str, budget: int) -> tuple[str, bool]:
"""Cut to `budget` characters from the MIDDLE, keeping both ends.
A head-only cut — `text[:800]` — decides what a reader sees by character
position, which is uncorrelated with what matters. Prose does not put its
conclusion first: a passage that opens with what was attempted and closes
with "so this shipped in 04775c3" loses exactly the sentence that answers
the question. Worse, the reader cannot tell: a truncation marker says that
something was removed, never whether it mattered, so the decision "should
I look deeper?" gets made on evidence selected by length.
So keep the opening (what this is about) AND the closing (where it landed),
and state in between how much went. Two thirds to the head because that is
where the subject is established; a conclusion needs less room to carry.
Returns (text, was_cut). A `budget` of 0 or less means no cut.
This is the fallback, not the goal. Where the system knows WHICH span of a
record is the relevant one — a semantic search knows exactly that, and
stores it (#4243) — show that span and say so. Reach for this only when
nothing identifies a better part than "all of it".
"""
if budget <= 0 or len(text) <= budget:
return text, False
head_len = max(1, budget * 2 // 3)
tail_len = max(1, budget - head_len)
omitted = len(text) - head_len - tail_len
head = text[:head_len].rstrip()
tail = text[-tail_len:].lstrip()
return f"{head}\n\n[… {omitted} characters omitted …]\n\n{tail}", True