feat(embeddings): best-chunk-per-note on every retrieval surface (#280 step 4)
CI & Build / Plugin hooks (push) Failing after 1s
CI & Build / Python lint (push) Failing after 3s
CI & Build / integration (push) Successful in 17s
CI & Build / TypeScript typecheck (push) Successful in 32s
CI & Build / Python tests (push) Successful in 46s
CI & Build / Build & push image (push) Skipped

A note's relevance is now its best chunk's similarity, everywhere:

- semantic_search_notes keeps the indexed raw-distance top-k and over-fetches
  chunk rows (x4, composing with the x3 supersession over-fetch), then
  collapses to first-appearance-per-note — rows arrive distance-ordered, so
  first is best. Every ranked consumer (MCP/REST search, Browse, auto-inject,
  write-path, gate) inherits through the one function.
- list_notes semantic q swaps its join for a correlated MIN-distance
  subquery — the join would have repeated a long note once per matching chunk
  and made total count chunks.
- the duplicate report groups its self-join by note pair on MIN(distance):
  pair similarity = closest chunk pair, and the < join now also drops
  cross-chunk self-pairs that would flag every long note against itself.
- the write gate queries once per chunk of the candidate (capped at 8), so a
  note duplicating an existing record in ONE SECTION is caught — the
  whole-document query diluted exactly the section that mattered.

Integration test now seeds a two-chunk note and pins the collapse against
real pgvector.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UaYUaouG9jjhATyuxCKrQs
This commit is contained in:
2026-08-08 23:51:01 -04:00
co-authored by Claude Fable 5
parent 0e70a3896b
commit 041d8defbc
6 changed files with 164 additions and 51 deletions
+27
View File
@@ -68,6 +68,33 @@ async def test_semantic_match_when_body_substantial():
assert dup.similarity == 0.93
@pytest.mark.asyncio
async def test_gate_catches_a_duplicate_hiding_in_a_later_chunk():
"""The capability #280 adds to the gate: a long candidate that duplicates
an existing record in ONE SECTION is caught, where the whole-document
query this replaces diluted exactly the section that mattered. The gate
queries once per chunk and any chunk's hit blocks."""
para = ("This section restates an existing decision in enough words to be "
"a real paragraph of content for the chunker to keep. ") * 4
body = "\n\n".join(f"## Topic {i}\n\n{para} (t{i})" for i in range(8))
from scribe.services.embeddings import chunk_document
n_chunks = len(chunk_document("Title", body))
assert n_chunks > 1, "test body must actually chunk"
hit = _fake_note(id=30, title="The existing decision", note_type="note")
# Every chunk misses except the LAST one the gate will ask about.
sem = AsyncMock(side_effect=[[] for _ in range(n_chunks - 1)] + [[(0.94, hit)]])
with patch("scribe.services.dedup.async_session",
return_value=_session_returning(None)), \
patch("scribe.services.dedup.embeddings_svc.semantic_search_notes", sem):
dup = await find_duplicate_note(
7, "Title", body=body, project_id=2, is_task=False, note_type="note",
)
assert dup is not None and dup.id == 30
assert sem.await_count == n_chunks
@pytest.mark.asyncio
async def test_semantic_match_of_other_note_type_is_ignored():
other = _fake_note(id=21, title="X", note_type="process")