feat(embeddings): best-chunk-per-note on every retrieval surface (#280 step 4)
CI & Build / Plugin hooks (push) Failing after 1s
CI & Build / Python lint (push) Failing after 3s
CI & Build / integration (push) Successful in 17s
CI & Build / TypeScript typecheck (push) Successful in 32s
CI & Build / Python tests (push) Successful in 46s
CI & Build / Build & push image (push) Skipped

A note's relevance is now its best chunk's similarity, everywhere:

- semantic_search_notes keeps the indexed raw-distance top-k and over-fetches
  chunk rows (x4, composing with the x3 supersession over-fetch), then
  collapses to first-appearance-per-note — rows arrive distance-ordered, so
  first is best. Every ranked consumer (MCP/REST search, Browse, auto-inject,
  write-path, gate) inherits through the one function.
- list_notes semantic q swaps its join for a correlated MIN-distance
  subquery — the join would have repeated a long note once per matching chunk
  and made total count chunks.
- the duplicate report groups its self-join by note pair on MIN(distance):
  pair similarity = closest chunk pair, and the < join now also drops
  cross-chunk self-pairs that would flag every long note against itself.
- the write gate queries once per chunk of the candidate (capped at 8), so a
  note duplicating an existing record in ONE SECTION is caught — the
  whole-document query diluted exactly the section that mattered.

Integration test now seeds a two-chunk note and pins the collapse against
real pgvector.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UaYUaouG9jjhATyuxCKrQs
This commit is contained in:
2026-08-08 23:51:01 -04:00
co-authored by Claude Fable 5
parent 0e70a3896b
commit 041d8defbc
6 changed files with 164 additions and 51 deletions
@@ -68,7 +68,11 @@ async def seeded():
await s.flush()
# query vector will be [1,0,0,...]; near ~ identical (sim≈1.0),
# far is orthogonal (sim≈0.0 -> filtered by the default threshold).
# near gets a SECOND, weaker chunk (sim≈0.6) — the collapse to
# best-chunk-per-note (#280) is under test: near must come back once,
# at its best chunk's score, not twice.
s.add(_emb(near.id, user.id, 0, _vec(1.0)))
s.add(_emb(near.id, user.id, 1, _vec(0.6, 0.8)))
s.add(_emb(far.id, user.id, 0, _vec(0.0, 1.0)))
await s.commit()
ids = (user.id, near.id, far.id)
@@ -96,6 +100,9 @@ async def test_semantic_search_ranks_and_thresholds_via_pgvector(seeded):
assert near_id in ids
assert far_id not in ids
assert ids[0] == near_id
# Chunk collapse (#280): near has TWO chunk rows above the floor (sim≈1.0
# and ≈0.6) and must appear exactly once, at its best chunk's score.
assert ids.count(near_id) == 1
top_score = results[0][0]
assert top_score == pytest.approx(1.0, abs=1e-3)