feat(retrieval): every semantic search hands on the passage that matched
CI & Build / Python lint (push) Successful in 8s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Failing after 1m9s
CI & Build / Build & push image (push) Skipped

#4243 fixed one door. Scribe has three semantic searches over three chunk
tables, and all three collapsed chunk rows to the best one per record — each
of them KNEW which passage earned the hit, and each dropped it. Every surface
downstream then previewed the head of the document instead: a span the search
had already scored lower, with nothing saying so.

Mechanism, one place:
  - embeddings.record_best_chunk publishes {id: {index, text}} into `report`.
    Carried in `report`, NOT the return value: all three return
    list[tuple[float, Record]] and ~30 sites unpack that pair (lesson #4207).
  - semantic_search_rules and semantic_search_milestones now select
    chunk_index/chunk_text and publish the winner, as notes already did.
    semantic_search_milestones gains `report`, which it had no way to take.
  - services/text.matched_excerpt is the one choice of span, and
    excerpt_fields the one result block. Doors keep their own field names —
    the web renders `snippet`, MCP returns `excerpt` — because renaming a
    field a frontend reads is a different change from fixing what goes in it.

Surfaces:
  - knowledge.query_knowledge, whose own comment calls it "the human's MAIN
    search surface", was `(note.body or "")[:200]` on every row alike. Now the
    matched passage on a search, the opening on a browse, and `snippet_is`
    saying which. KnowledgeView renders that snippet, so this was live.
  - search(content_type='milestone') gains `matched` — the plan body stays
    out, but the passage that matched comes along, because recognising a plan
    means recognising the part you asked about and a description written at
    the start need not mention it.
  - The auto-inject menu and the write-path prior-art menu put the passage
    under their line. Both were title-only, which answers "does this apply?"
    for a lesson or snippet (the trigger is IN the title) and not at all for
    an issue or dev-log. No fallback to the body's opening: on a menu that is
    preamble dressed as a reason, and once indented it cannot be told apart.

Left alone deliberately: the rule arms. A rule hint already renders the rule's
TRIGGER, which is written to answer exactly "does this apply to me" and beats
a matched chunk at it; and that line's budget was measured at #3851. Adding a
passage there would duplicate the trigger and spend the budget twice.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
2026-09-21 09:44:10 -04:00
co-authored by Claude Opus 5
parent 6abedb0168
commit 253fb974f3
8 changed files with 527 additions and 51 deletions
+34 -1
View File
@@ -29,7 +29,17 @@ interface KnowledgeItem {
id: number;
note_type: "note" | "task" | "process" | "snippet" | "lesson";
title: string;
/** ONE SPAN of the record, never the whole thing. On a search this is the
* passage that actually matched; on a plain browse it is the opening,
* because nothing was matched and no span is better than another. */
snippet: string;
/** Which span `snippet` holds. A reader who cannot tell the matched passage
* from the document's first paragraph cannot tell whether a card that looks
* unrelated really is. */
snippet_is?: "matched_passage" | "body_opening";
/** Characters in the whole record, so a long record shown by a short span is
* visible as one. */
body_length?: number;
tags: string[];
project_id: number | null;
created_at: string;
@@ -181,6 +191,16 @@ const CONTENT_PAGE = 24; // items loaded per sentinel trigger
const REFILL_THRESHOLD = 48; // fetch more IDs when queue drops below this
const items = ref<KnowledgeItem[]>([]);
// True only when a search actually returned matched passages — not merely when
// the box has text in it. A keyword-only result set, or an embedder that is
// down, carries `body_opening` rows, and claiming otherwise would be the same
// species of lie this whole change is about (#4243).
const showsMatchedPassages = computed(
() =>
searchQuery.value.trim().length > 0 &&
items.value.some((i) => i.snippet_is === "matched_passage"),
);
const allTags = ref<string[]>([]);
const idQueue = ref<number[]>([]); // unloaded IDs ready to be content-fetched
const idOffset = ref(0); // next offset for ID batch requests
@@ -573,8 +593,16 @@ onUnmounted(() => {
<p v-else class="empty-narrator">Your story is unwritten. Create your first note to begin.</p>
</div>
<!-- Said once above the grid rather than per card: it is true of every
row at once, and a label repeated on each would cost more than it
tells. Only while a query is active — on a plain browse nothing
matched, so there is no matched passage to explain. -->
<p v-else-if="showsMatchedPassages" class="k-excerpt-note">
Excerpts below are the passage that matched your search, not the start of each record.
</p>
<!-- Card grid -->
<div v-else class="card-grid">
<div v-if="items.length" class="card-grid">
<div
v-for="item in items"
:key="item.id"
@@ -1004,6 +1032,11 @@ onUnmounted(() => {
line-height: 1.45;
margin: 0;
}
.k-excerpt-note {
font-size: 0.8rem;
color: var(--fs-text-tertiary);
margin: 0 0 var(--fs-space-3);
}
.k-card-footer {
display: flex;
align-items: center;