feat(snippets): near-duplicate finder — surface the sets worth merging
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 14s
CI & Build / integration (push) Successful in 36s
CI & Build / Python tests (push) Successful in 55s
CI & Build / Build & push image (push) Successful in 44s

#231's premise was unifying reusable things already scattered as one-offs.
The create gate PREVENTS a new duplicate and merge_snippets CURES one you
point it at, but nothing FOUND the duplicates already in the record —
someone had to notice them by hand, which is the exact failure the Drafter
exists to remove.

One indexed self-join over note_embeddings, not an N² Python scan:
pgvector's cosine distance is the same operator semantic search uses, so a
similarity floor is a distance ceiling and the work stays in Postgres.
`left.note_id < right.note_id` yields each unordered pair once and drops
the self-pair that would otherwise dominate the ranking.

Pairs are collapsed into merge SETS by connected components. Transitive on
purpose: A~B plus B~C puts all three together even when A and C don't
directly clear the bar, which is what merge actually does (it folds every
source into one survivor). The cost is that a chain of mild resemblances
can rope in a member that isn't really alike — so the UI presents a set as
a proposal, shows the members, and never merges without a confirm.

Two scope decisions worth naming:

- OWN snippets only. merge_snippets requires one owner across the set, so
  surfacing someone else's would propose a merge that cannot be performed.
  The report is bounded by what the operator can act on, not what they can
  see.

- Threshold defaults to 0.82, LOOSER than the write gate's 0.90, and is a
  setting rather than a constant (rule #25). The gate blocks a create and
  has to be unforgiving of noise; this only suggests a merge under review,
  so it must reach further or it would never surface the pairs the gate
  already let through — which are precisely the ones that accumulated.

Fixes a real bug in the merge flow while wiring the UI: selectedList
filtered the selection against the CURRENT PAGE, and doMerge derives its
source ids from that list. A corpus-wide suggested group with off-page
members would have rendered incomplete and silently merged only the
visible subset. A group under review is now the authority for that list.

Refs #2088

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UaYUaouG9jjhATyuxCKrQs
This commit is contained in:
2026-07-28 18:32:40 -04:00
co-authored by Claude Opus 5
parent 84c5c0dc81
commit 6db791965f
10 changed files with 547 additions and 8 deletions
+3
View File
@@ -256,6 +256,9 @@ _READ_ONLY_TOOLS = frozenset({
"list_rules", "list_tags", "list_tasks", "list_topics", "list_trash",
"list_always_on_rules", "search",
"get_system", "list_systems", "list_system_records",
# Reports on the snippet corpus. Reads only — the merge it suggests is a
# separate, explicitly-called write.
"find_duplicate_snippets",
})
+39 -1
View File
@@ -190,6 +190,44 @@ async def get_snippet(snippet_id: int) -> dict:
return data
async def find_duplicate_snippets(threshold: float = 0.0) -> dict:
"""Find snippets already recorded that look like duplicates of each other.
The create gate PREVENTS a new duplicate and merge_snippets CURES one you
point it at — this is the missing third piece: it FINDS the ones already in
the record, so nobody has to notice them by hand.
Results are grouped into candidate merge SETS, not just pairs. Grouping is
transitive: if A resembles B and B resembles C, all three land in one set
even when A and C don't directly clear the bar. That mirrors what merge does
(it folds every source into one survivor), but it means a chain of mild
resemblances can rope in a member that isn't really alike — so read a set as
a proposal and check the members before acting.
Reports only YOUR snippets. merge_snippets requires one owner across the
whole set, so surfacing someone else's would propose a merge that can't be
performed.
Acting on a group: pick the best record as the canonical target, then
`merge_snippets(target_id, [other ids])`. Merge unions the fields and folds
every source's location in, so the survivor is findable at all their call
sites; the sources are trashed, recoverably. Prefer as target the one with
the clearest "when to reach for it" — merge keeps the target's title.
Args:
threshold: Similarity floor, 0-1. 0 (default) uses the configured
setting. Raise it if the report is noisy, lower it to catch more.
Returns {"groups": [{"note_ids", "snippets", "top_score"}], "pairs",
"threshold"}. An empty `groups` means nothing resembles anything else that
closely — the common and desirable case.
"""
uid = current_user_id()
return await dedup_svc.find_duplicate_snippets(
uid, threshold=threshold if threshold > 0 else None
)
async def verify_snippet(
snippet_id: int, status: str, detail: str = "", path: str = "",
) -> dict:
@@ -365,6 +403,6 @@ async def merge_snippets(target_id: int, source_ids: list[int]) -> dict:
def register(mcp) -> None:
for fn in (
list_snippets, create_snippet, get_snippet, update_snippet,
delete_snippet, merge_snippets, verify_snippet,
delete_snippet, merge_snippets, verify_snippet, find_duplicate_snippets,
):
mcp.tool(name=fn.__name__)(fn)