fix(search): show the passage that matched, not the opening of the body (#4243)
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Failing after 1m5s
CI & Build / Build & push image (push) Skipped
CI & Build / integration (push) Successful in 46s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Failing after 1m5s
CI & Build / Build & push image (push) Skipped
CI & Build / integration (push) Successful in 46s
Raised by the operator: are we limiting what comes back by character count,
and how do we verify the pertinent part is the part displayed?
We were not. mcp/tools/search.py sent (note.body or "")[:240] — a head cut,
with no marker that anything had been removed, so a 240-character preview of
a 4000-character record was indistinguishable from a complete short one.
The opening is the wrong span. The match is semantic and per chunk, and
semantic_search_notes collapses to best-chunk-per-note — its own comment at
the collapse says "the first appearance of a note is its best chunk". So the
system identified the passage that earned the hit and then discarded it:
select(Note, distance) kept no chunk column. A record could rank first on its
sixth paragraph, be previewed by its first, and be judged irrelevant on a
span the search had already scored lower. That biases against long records,
and it is self-concealing — the caller who does not open it never learns the
preview was misleading.
- embeddings: chunk_index/chunk_text ride along in the select, and the
collapse records the winner in report["best_chunk"]. Carried in `report`,
NOT by widening the return tuple: ten callers unpack (score, note) at
~18 sites and nothing would catch the misses (lesson #4207). `report` is
the side-channel this function already uses for best_available_score.
- search(): excerpt / excerpt_is / body_length, and read_full when there is
more. A caller that cannot tell a matched passage from a document opening
cannot judge whether to look deeper, which is the only decision the field
supports.
elide() moves to services/text.py so both callers share one copy, and it
keeps BOTH ends with a stated gap — it is the fallback for when nothing
identifies a better span than "all of it", not the goal.
Also fixes a guard that produced a false failure on the previous commit:
test_pull_telemetry checked `"project_id: int = 0" in body.split("\n")[0]`,
which sees only the first line, so wrapping get_task's signature over four
lines made it report a function that does take the project as one that does
not. Parsed with ast now, and proven to still reject an absent or
wrongly-typed parameter rather than being appeased by reflowing the code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
@@ -11,6 +11,7 @@ import time
|
||||
|
||||
from scribe.mcp._context import current_user_id
|
||||
from scribe.services.access import owner_names_for
|
||||
from scribe.services.text import elide
|
||||
from scribe.services.embeddings import (
|
||||
DEFAULT_SIMILARITY_THRESHOLD, semantic_search_milestones, semantic_search_notes,
|
||||
semantic_search_rules,
|
||||
@@ -102,6 +103,55 @@ async def _search_milestones(uid: int, q: str, limit: int, project_id: int) -> d
|
||||
}
|
||||
|
||||
|
||||
# A matched chunk is at most _CHUNK_CHAR_BUDGET (1400) characters, and it is
|
||||
# the evidence the ranking was built on — so it is worth more room than the 240
|
||||
# characters of document opening this used to send. Elision inside a chunk is
|
||||
# far less lossy than a head cut of a whole record: the region is already the
|
||||
# right one.
|
||||
_EXCERPT_CHARS = 1000
|
||||
|
||||
|
||||
def result_excerpt(note, chunk: dict | None) -> dict:
|
||||
"""The part of a record a caller judges "should I open this?" on.
|
||||
|
||||
This used to be `(note.body or "")[:240]` — the document's opening, with
|
||||
no marker that anything had been cut, so a 240-character preview of a
|
||||
4000-character record was indistinguishable from a complete short one.
|
||||
|
||||
The opening is the wrong span. The match was semantic and per-chunk, and
|
||||
`semantic_search_notes` collapses to best-chunk-per-note — so the system
|
||||
already knows which passage earned the hit and used to discard it. A
|
||||
record could rank first on its sixth paragraph and be previewed by its
|
||||
first, which the search had already judged less relevant, and the caller
|
||||
would decide from that and never know (#4243).
|
||||
|
||||
So: show the matched passage when there is one, the body when it fits, and
|
||||
in either case SAY which of the two this is. A caller that cannot tell an
|
||||
excerpt from a whole record cannot tell whether looking deeper is worth it,
|
||||
which is the only decision this field supports.
|
||||
"""
|
||||
body = note.body or ""
|
||||
matched = (chunk or {}).get("text") or ""
|
||||
out: dict = {"body_length": len(body)}
|
||||
|
||||
if matched.strip():
|
||||
text, cut = elide(matched.strip(), _EXCERPT_CHARS)
|
||||
out["excerpt"] = text
|
||||
out["excerpt_is"] = "matched_passage"
|
||||
if (chunk or {}).get("index") is not None:
|
||||
out["chunk_index"] = int(chunk["index"])
|
||||
else:
|
||||
# No stored chunk — an un-embedded record, or a caller that passed no
|
||||
# report. Fall back to the opening, and name it as the opening rather
|
||||
# than letting it pass for the relevant part.
|
||||
text, cut = elide(body, _EXCERPT_CHARS)
|
||||
out["excerpt"] = text
|
||||
out["excerpt_is"] = "body_opening"
|
||||
if cut or (out["excerpt_is"] == "matched_passage" and len(matched) < len(body)):
|
||||
out["read_full"] = "get_note / get_task by id for the whole record."
|
||||
return out
|
||||
|
||||
|
||||
async def search(
|
||||
q: str,
|
||||
content_type: str = "all",
|
||||
@@ -150,9 +200,21 @@ async def search(
|
||||
list_system_records gives the same slice unranked.
|
||||
|
||||
Returns:
|
||||
{"results": [{"id", "title", "body", "is_task", "tags", "similarity"}],
|
||||
{"results": [{"id", "title", "excerpt", "excerpt_is", "body_length",
|
||||
"is_task", "tags", "similarity"}],
|
||||
"total": int}
|
||||
|
||||
`excerpt` is a SPAN of the record, not the record. `excerpt_is` says
|
||||
which span: "matched_passage" is the chunk that actually earned the
|
||||
hit — the evidence the ranking was built on, and the right thing to
|
||||
judge relevance from. "body_opening" is a fallback for a record with
|
||||
no stored chunk, and is only the beginning of the text, which may say
|
||||
nothing about why it matched. `body_length` is the whole record's
|
||||
size, so a long record previewed by a short span is visible as one;
|
||||
`read_full` appears when there is more, and opening the id by
|
||||
get_note / get_task is how you get it. Judge from the passage, not
|
||||
from the fact that a preview looked thin.
|
||||
|
||||
A result marked `shared: true` with an `owner` belongs to another user —
|
||||
that person's suggestion, not the operator's own record or settled practice.
|
||||
Weigh it on its merits and say whose it is when you use it.
|
||||
@@ -195,12 +257,13 @@ async def search(
|
||||
owners = await owner_names_for(
|
||||
{int(note.user_id) for _s, note in raw if note.user_id != uid}
|
||||
)
|
||||
chunks = report.get("best_chunk") or {}
|
||||
return {
|
||||
"results": [
|
||||
{
|
||||
"id": note.id,
|
||||
"title": note.title,
|
||||
"body": (note.body or "")[:240],
|
||||
**result_excerpt(note, chunks.get(int(note.id))),
|
||||
"is_task": bool(note.is_task),
|
||||
"tags": list(note.tags or []),
|
||||
"similarity": float(score),
|
||||
|
||||
Reference in New Issue
Block a user