fix(search): show the passage that matched, not the opening of the body (#4243)
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Failing after 1m5s
CI & Build / Build & push image (push) Skipped
CI & Build / integration (push) Successful in 46s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Failing after 1m5s
CI & Build / Build & push image (push) Skipped
CI & Build / integration (push) Successful in 46s
Raised by the operator: are we limiting what comes back by character count,
and how do we verify the pertinent part is the part displayed?
We were not. mcp/tools/search.py sent (note.body or "")[:240] — a head cut,
with no marker that anything had been removed, so a 240-character preview of
a 4000-character record was indistinguishable from a complete short one.
The opening is the wrong span. The match is semantic and per chunk, and
semantic_search_notes collapses to best-chunk-per-note — its own comment at
the collapse says "the first appearance of a note is its best chunk". So the
system identified the passage that earned the hit and then discarded it:
select(Note, distance) kept no chunk column. A record could rank first on its
sixth paragraph, be previewed by its first, and be judged irrelevant on a
span the search had already scored lower. That biases against long records,
and it is self-concealing — the caller who does not open it never learns the
preview was misleading.
- embeddings: chunk_index/chunk_text ride along in the select, and the
collapse records the winner in report["best_chunk"]. Carried in `report`,
NOT by widening the return tuple: ten callers unpack (score, note) at
~18 sites and nothing would catch the misses (lesson #4207). `report` is
the side-channel this function already uses for best_available_score.
- search(): excerpt / excerpt_is / body_length, and read_full when there is
more. A caller that cannot tell a matched passage from a document opening
cannot judge whether to look deeper, which is the only decision the field
supports.
elide() moves to services/text.py so both callers share one copy, and it
keeps BOTH ends with a stated gap — it is the fallback for when nothing
identifies a better span than "all of it", not the goal.
Also fixes a guard that produced a false failure on the previous commit:
test_pull_telemetry checked `"project_id: int = 0" in body.split("\n")[0]`,
which sees only the first line, so wrapping get_task's signature over four
lines made it report a function that does take the project as one that does
not. Parsed with ast now, and proven to still reject an absent or
wrongly-typed parameter rather than being appeased by reflowing the code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
@@ -224,13 +224,34 @@ def _pulling_getters():
|
||||
yield module, name, body
|
||||
|
||||
|
||||
def _takes_reading_project(body: str) -> bool:
|
||||
"""Does this function declare `project_id: int = 0`?
|
||||
|
||||
Parsed, not string-matched against the first line. The first version read
|
||||
`body.split("\n")[0]`, which sees only as far as the first newline — so
|
||||
adding a parameter to `get_task` wrapped its signature over four lines and
|
||||
the guard reported a function that DOES take the project as one that does
|
||||
not. A guard that fails on formatting is a guard that gets appeased by
|
||||
reflowing the code it was meant to check.
|
||||
"""
|
||||
node = ast.parse(body).body[0]
|
||||
args = node.args
|
||||
params = list(args.posonlyargs) + list(args.args) + list(args.kwonlyargs)
|
||||
for arg in params:
|
||||
if arg.arg != "project_id":
|
||||
continue
|
||||
ann = getattr(arg, "annotation", None)
|
||||
return isinstance(ann, ast.Name) and ann.id == "int"
|
||||
return False
|
||||
|
||||
|
||||
def test_every_getter_that_pulls_takes_the_reading_project():
|
||||
"""Asserted on structure (rule 167): a behavioural test cannot see a
|
||||
parameter that was never threaded through."""
|
||||
missing = [
|
||||
f"{module}.{name}"
|
||||
for module, name, body in _pulling_getters()
|
||||
if "project_id: int = 0" not in body.split("\n")[0]
|
||||
if not _takes_reading_project(body)
|
||||
]
|
||||
assert not missing, (
|
||||
f"these getters record a pull but cannot say where the reader was: "
|
||||
|
||||
Reference in New Issue
Block a user