"""Text shortening for surfaces that cannot show a record whole. One function, in one place, because both doors shorten and two copies would drift — and because the reasoning below is the part that matters and should not have to be re-derived at each call site. """ def elide(text: str, budget: int) -> tuple[str, bool]: """Cut to `budget` characters from the MIDDLE, keeping both ends. A head-only cut — `text[:800]` — decides what a reader sees by character position, which is uncorrelated with what matters. Prose does not put its conclusion first: a passage that opens with what was attempted and closes with "so this shipped in 04775c3" loses exactly the sentence that answers the question. Worse, the reader cannot tell: a truncation marker says that something was removed, never whether it mattered, so the decision "should I look deeper?" gets made on evidence selected by length. So keep the opening (what this is about) AND the closing (where it landed), and state in between how much went. Two thirds to the head because that is where the subject is established; a conclusion needs less room to carry. Returns (text, was_cut). A `budget` of 0 or less means no cut. This is the fallback, not the goal. Where the system knows WHICH span of a record is the relevant one — a semantic search knows exactly that, and stores it (#4243) — show that span and say so. Reach for this only when nothing identifies a better part than "all of it". """ if budget <= 0 or len(text) <= budget: return text, False head_len = max(1, budget * 2 // 3) tail_len = max(1, budget - head_len) omitted = len(text) - head_len - tail_len head = text[:head_len].rstrip() tail = text[-tail_len:].lstrip() return f"{head}\n\n[… {omitted} characters omitted …]\n\n{tail}", True # What a search result shows a reader, as one decision made in one place. # # Every door onto a semantic search faces the same question — which span of a # record do I show someone deciding whether to open it? — and each used to # answer it separately with a bare head cut of a different length: 240 # characters in the MCP search, 200 in the web's knowledge search. Both showed # the document's OPENING, which is not the span that matched and not the span # the ranking was built on. # # Doors keep their own field names (the web UI renders `snippet`, the MCP # surface returns `excerpt`), because renaming a field a frontend consumes is # a separate change from fixing which text goes in it. What they share is this # function: the choice of span, and the obligation to say which span it is. MATCHED_PASSAGE = "matched_passage" BODY_OPENING = "body_opening" def matched_excerpt( body: str, chunk: dict | None, budget: int ) -> tuple[str, str, bool]: """Pick the span to show, and say which span it is. `chunk` is a `report["best_chunk"]` entry from one of the semantic searches — `{"index": int, "text": str}` — or None when the caller ran no semantic search, passed no report, or the record has no embedding row. Returns `(text, kind, was_cut)` where `kind` is MATCHED_PASSAGE or BODY_OPENING. The kind is not decoration: a reader who cannot tell the passage that earned the hit from the first paragraph of the document cannot tell whether a thin-looking result is genuinely thin, and the whole point of the excerpt is to support exactly that judgement. A short record comes back whole either way — fragmenting a 200-character note serves nobody, and its opening IS its content. """ passage = (chunk or {}).get("text") or "" if passage.strip(): text, cut = elide(passage.strip(), budget) return text, MATCHED_PASSAGE, cut text, cut = elide(body or "", budget) return text, BODY_OPENING, cut def excerpt_fields( body: str, chunk: dict | None, budget: int, *, key: str = "excerpt" ) -> dict: """`matched_excerpt` as the block a result row carries. `key` names the text field so a door can keep the name its consumers already read. The companion keys are derived from it, so a row never ends up with an excerpt under one name and its label under another. """ text, kind, cut = matched_excerpt(body, chunk, budget) out = {key: text, f"{key}_is": kind, "body_length": len(body or "")} if cut or (kind == MATCHED_PASSAGE and len(text) < len(body or "")): out["read_full"] = ( "This is one span of a longer record. Open it by id for the whole " "thing rather than judging from what is here." ) return out