A note's relevance is now its best chunk's similarity, everywhere:
- semantic_search_notes keeps the indexed raw-distance top-k and over-fetches
chunk rows (x4, composing with the x3 supersession over-fetch), then
collapses to first-appearance-per-note — rows arrive distance-ordered, so
first is best. Every ranked consumer (MCP/REST search, Browse, auto-inject,
write-path, gate) inherits through the one function.
- list_notes semantic q swaps its join for a correlated MIN-distance
subquery — the join would have repeated a long note once per matching chunk
and made total count chunks.
- the duplicate report groups its self-join by note pair on MIN(distance):
pair similarity = closest chunk pair, and the < join now also drops
cross-chunk self-pairs that would flag every long note against itself.
- the write gate queries once per chunk of the candidate (capped at 8), so a
note duplicating an existing record in ONE SECTION is caught — the
whole-document query diluted exactly the section that mattered.
Integration test now seeds a two-chunk note and pins the collapse against
real pgvector.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UaYUaouG9jjhATyuxCKrQs