feat(embeddings): best-chunk-per-note on every retrieval surface (#280 step 4)
CI & Build / Plugin hooks (push) Failing after 1s
CI & Build / Python lint (push) Failing after 3s
CI & Build / integration (push) Successful in 17s
CI & Build / TypeScript typecheck (push) Successful in 32s
CI & Build / Python tests (push) Successful in 46s
CI & Build / Build & push image (push) Skipped
CI & Build / Plugin hooks (push) Failing after 1s
CI & Build / Python lint (push) Failing after 3s
CI & Build / integration (push) Successful in 17s
CI & Build / TypeScript typecheck (push) Successful in 32s
CI & Build / Python tests (push) Successful in 46s
CI & Build / Build & push image (push) Skipped
A note's relevance is now its best chunk's similarity, everywhere: - semantic_search_notes keeps the indexed raw-distance top-k and over-fetches chunk rows (x4, composing with the x3 supersession over-fetch), then collapses to first-appearance-per-note — rows arrive distance-ordered, so first is best. Every ranked consumer (MCP/REST search, Browse, auto-inject, write-path, gate) inherits through the one function. - list_notes semantic q swaps its join for a correlated MIN-distance subquery — the join would have repeated a long note once per matching chunk and made total count chunks. - the duplicate report groups its self-join by note pair on MIN(distance): pair similarity = closest chunk pair, and the < join now also drops cross-chunk self-pairs that would flag every long note against itself. - the write gate queries once per chunk of the candidate (capped at 8), so a note duplicating an existing record in ONE SECTION is caught — the whole-document query diluted exactly the section that mattered. Integration test now seeds a two-chunk note and pins the collapse against real pgvector. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UaYUaouG9jjhATyuxCKrQs
This commit is contained in:
@@ -236,15 +236,25 @@ async def list_notes(
|
||||
if query_vec is not None:
|
||||
from scribe.models.embedding import NoteEmbedding
|
||||
from scribe.services.embeddings import INTERACTIVE_SEARCH_THRESHOLD
|
||||
distance = NoteEmbedding.embedding.cosine_distance(query_vec)
|
||||
sem_filter = distance <= (1.0 - INTERACTIVE_SEARCH_THRESHOLD)
|
||||
query = query.join(
|
||||
NoteEmbedding, NoteEmbedding.note_id == Note.id
|
||||
).where(sem_filter)
|
||||
count_query = count_query.join(
|
||||
NoteEmbedding, NoteEmbedding.note_id == Note.id
|
||||
).where(sem_filter)
|
||||
semantic_order = distance.asc()
|
||||
# Best-chunk-per-note as a correlated MIN, not a join (#280):
|
||||
# a note stores one embedding row PER CHUNK, so the plain join
|
||||
# this used to be would repeat a long note once per matching
|
||||
# chunk — duplicated list rows and a total that counts chunks.
|
||||
# This query is filter-heavy and paginated, never HNSW-bound,
|
||||
# so the scalar subquery costs what the join did.
|
||||
best_distance = (
|
||||
select(
|
||||
func.min(
|
||||
NoteEmbedding.embedding.cosine_distance(query_vec)
|
||||
)
|
||||
)
|
||||
.where(NoteEmbedding.note_id == Note.id)
|
||||
.scalar_subquery()
|
||||
)
|
||||
sem_filter = best_distance <= (1.0 - INTERACTIVE_SEARCH_THRESHOLD)
|
||||
query = query.where(sem_filter)
|
||||
count_query = count_query.where(sem_filter)
|
||||
semantic_order = best_distance.asc()
|
||||
else:
|
||||
terms = _strip_type_nouns(q)
|
||||
for term in terms:
|
||||
|
||||
Reference in New Issue
Block a user