feat(snippets): near-duplicate finder — surface the sets worth merging
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 14s
CI & Build / integration (push) Successful in 36s
CI & Build / Python tests (push) Successful in 55s
CI & Build / Build & push image (push) Successful in 44s
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 14s
CI & Build / integration (push) Successful in 36s
CI & Build / Python tests (push) Successful in 55s
CI & Build / Build & push image (push) Successful in 44s
#231's premise was unifying reusable things already scattered as one-offs. The create gate PREVENTS a new duplicate and merge_snippets CURES one you point it at, but nothing FOUND the duplicates already in the record — someone had to notice them by hand, which is the exact failure the Drafter exists to remove. One indexed self-join over note_embeddings, not an N² Python scan: pgvector's cosine distance is the same operator semantic search uses, so a similarity floor is a distance ceiling and the work stays in Postgres. `left.note_id < right.note_id` yields each unordered pair once and drops the self-pair that would otherwise dominate the ranking. Pairs are collapsed into merge SETS by connected components. Transitive on purpose: A~B plus B~C puts all three together even when A and C don't directly clear the bar, which is what merge actually does (it folds every source into one survivor). The cost is that a chain of mild resemblances can rope in a member that isn't really alike — so the UI presents a set as a proposal, shows the members, and never merges without a confirm. Two scope decisions worth naming: - OWN snippets only. merge_snippets requires one owner across the set, so surfacing someone else's would propose a merge that cannot be performed. The report is bounded by what the operator can act on, not what they can see. - Threshold defaults to 0.82, LOOSER than the write gate's 0.90, and is a setting rather than a constant (rule #25). The gate blocks a create and has to be unforgiving of noise; this only suggests a merge under review, so it must reach further or it would never surface the pairs the gate already let through — which are precisely the ones that accumulated. Fixes a real bug in the merge flow while wiring the UI: selectedList filtered the selection against the CURRENT PAGE, and doMerge derives its source ids from that list. A corpus-wide suggested group with off-page members would have rendered incomplete and silently merged only the visible subset. A group under review is now the authority for that list. Refs #2088 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UaYUaouG9jjhATyuxCKrQs
This commit is contained in:
@@ -154,6 +154,25 @@ export async function deleteSnippet(id: number): Promise<void> {
|
|||||||
return apiDelete(`/api/snippets/${id}`);
|
return apiDelete(`/api/snippets/${id}`);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/** A set of snippets that resemble each other closely enough to be worth
|
||||||
|
* merging. Grouping is transitive, so a set can hold members that don't
|
||||||
|
* directly resemble each other — read it as a proposal, not a verdict. */
|
||||||
|
export interface DuplicateGroup {
|
||||||
|
note_ids: number[];
|
||||||
|
snippets: { id: number; title: string }[];
|
||||||
|
/** The strongest resemblance within the set — how confident the suggestion is. */
|
||||||
|
top_score: number;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Near-duplicates already in the record. The create gate prevents new ones and
|
||||||
|
* merge cures the ones you point it at; this is what finds them. */
|
||||||
|
export async function findDuplicateSnippets(
|
||||||
|
threshold?: number,
|
||||||
|
): Promise<{ groups: DuplicateGroup[]; threshold: number }> {
|
||||||
|
const qs = threshold ? `?threshold=${threshold}` : "";
|
||||||
|
return apiGet(`/api/snippets/duplicates${qs}`);
|
||||||
|
}
|
||||||
|
|
||||||
/** Record a drift-check verdict. The check itself runs where the code is — an
|
/** Record a drift-check verdict. The check itself runs where the code is — an
|
||||||
* agent with the working tree — since Scribe has no checkout. This stores what
|
* agent with the working tree — since Scribe has no checkout. This stores what
|
||||||
* was found, and is how the UI clears a stale marker after a manual fix. */
|
* was found, and is how the UI clears a stale marker after a manual fix. */
|
||||||
|
|||||||
@@ -23,6 +23,10 @@ const kbInjectEnabled = ref(true);
|
|||||||
const kbInjectThreshold = ref("0.55");
|
const kbInjectThreshold = ref("0.55");
|
||||||
const kbInjectTopK = ref("3");
|
const kbInjectTopK = ref("3");
|
||||||
const kbWritePathEnabled = ref(true);
|
const kbWritePathEnabled = ref(true);
|
||||||
|
// Near-duplicate report floor. Deliberately looser than the 0.90 write-time
|
||||||
|
// gate: that one BLOCKS a create and must be unforgiving of noise, this one only
|
||||||
|
// suggests a merge the operator reviews (services/dedup.py).
|
||||||
|
const kbDuplicateThreshold = ref("0.82");
|
||||||
const savingKbInject = ref(false);
|
const savingKbInject = ref(false);
|
||||||
const kbInjectSaved = ref(false);
|
const kbInjectSaved = ref(false);
|
||||||
|
|
||||||
@@ -68,8 +72,13 @@ async function saveRetention() {
|
|||||||
async function saveKbInject() {
|
async function saveKbInject() {
|
||||||
const t = Math.min(1, Math.max(0, Number(kbInjectThreshold.value) || 0));
|
const t = Math.min(1, Math.max(0, Number(kbInjectThreshold.value) || 0));
|
||||||
const k = Math.min(10, Math.max(1, Math.floor(Number(kbInjectTopK.value) || 1)));
|
const k = Math.min(10, Math.max(1, Math.floor(Number(kbInjectTopK.value) || 1)));
|
||||||
|
// `|| 0.82` not `|| 0`: an unparseable value here should fall back to the
|
||||||
|
// default, not to 0 — a 0 floor would report every snippet as a duplicate of
|
||||||
|
// every other one.
|
||||||
|
const dupT = Math.min(1, Math.max(0, Number(kbDuplicateThreshold.value) || 0.82));
|
||||||
kbInjectThreshold.value = String(t);
|
kbInjectThreshold.value = String(t);
|
||||||
kbInjectTopK.value = String(k);
|
kbInjectTopK.value = String(k);
|
||||||
|
kbDuplicateThreshold.value = String(dupT);
|
||||||
savingKbInject.value = true;
|
savingKbInject.value = true;
|
||||||
kbInjectSaved.value = false;
|
kbInjectSaved.value = false;
|
||||||
try {
|
try {
|
||||||
@@ -80,6 +89,7 @@ async function saveKbInject() {
|
|||||||
// Its own switch, but deliberately the same threshold/ceiling — see
|
// Its own switch, but deliberately the same threshold/ceiling — see
|
||||||
// WRITEPATH_ENABLED_KEY in services/plugin_context.py.
|
// WRITEPATH_ENABLED_KEY in services/plugin_context.py.
|
||||||
kb_writepath_enabled: kbWritePathEnabled.value ? 'true' : 'false',
|
kb_writepath_enabled: kbWritePathEnabled.value ? 'true' : 'false',
|
||||||
|
kb_duplicate_threshold: String(dupT),
|
||||||
});
|
});
|
||||||
kbInjectSaved.value = true;
|
kbInjectSaved.value = true;
|
||||||
setTimeout(() => (kbInjectSaved.value = false), 2000);
|
setTimeout(() => (kbInjectSaved.value = false), 2000);
|
||||||
@@ -464,6 +474,9 @@ onMounted(async () => {
|
|||||||
kbInjectTopK.value = allSettings.kb_autoinject_top_k;
|
kbInjectTopK.value = allSettings.kb_autoinject_top_k;
|
||||||
}
|
}
|
||||||
kbWritePathEnabled.value = allSettings.kb_writepath_enabled !== "false";
|
kbWritePathEnabled.value = allSettings.kb_writepath_enabled !== "false";
|
||||||
|
if (allSettings.kb_duplicate_threshold !== undefined) {
|
||||||
|
kbDuplicateThreshold.value = allSettings.kb_duplicate_threshold;
|
||||||
|
}
|
||||||
if (allSettings.notify_task_reminders !== undefined) {
|
if (allSettings.notify_task_reminders !== undefined) {
|
||||||
notifyTaskReminders.value = allSettings.notify_task_reminders !== "false";
|
notifyTaskReminders.value = allSettings.notify_task_reminders !== "false";
|
||||||
}
|
}
|
||||||
@@ -1211,6 +1224,25 @@ function formatUserDate(iso: string): string {
|
|||||||
edit. Off = prior art surfaces only on your own prompts.
|
edit. Off = prior art surfaces only on your own prompts.
|
||||||
</p>
|
</p>
|
||||||
</div>
|
</div>
|
||||||
|
<div class="field">
|
||||||
|
<label for="kb-duplicate-threshold">Near-duplicate report threshold</label>
|
||||||
|
<input
|
||||||
|
id="kb-duplicate-threshold"
|
||||||
|
v-model="kbDuplicateThreshold"
|
||||||
|
type="number"
|
||||||
|
min="0"
|
||||||
|
max="1"
|
||||||
|
step="0.01"
|
||||||
|
class="input"
|
||||||
|
style="max-width: 8rem"
|
||||||
|
/>
|
||||||
|
<p class="field-hint">
|
||||||
|
How alike two snippets must be before the Snippets page suggests merging
|
||||||
|
them. Lower = more suggestions, more false pairs. Looser than the 0.90
|
||||||
|
used to block a duplicate at creation, because this only proposes a merge
|
||||||
|
you review — it never acts on its own.
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
<div class="actions">
|
<div class="actions">
|
||||||
<button class="btn-save" @click="saveKbInject" :disabled="savingKbInject">
|
<button class="btn-save" @click="saveKbInject" :disabled="savingKbInject">
|
||||||
{{ savingKbInject ? 'Saving…' : 'Save' }}
|
{{ savingKbInject ? 'Saving…' : 'Save' }}
|
||||||
|
|||||||
@@ -1,7 +1,13 @@
|
|||||||
<script setup lang="ts">
|
<script setup lang="ts">
|
||||||
import { ref, computed, onMounted } from "vue";
|
import { ref, computed, onMounted } from "vue";
|
||||||
import { useRouter } from "vue-router";
|
import { useRouter } from "vue-router";
|
||||||
import { listSnippets, mergeSnippets, type SnippetListItem } from "@/api/snippets";
|
import {
|
||||||
|
findDuplicateSnippets,
|
||||||
|
listSnippets,
|
||||||
|
mergeSnippets,
|
||||||
|
type DuplicateGroup,
|
||||||
|
type SnippetListItem,
|
||||||
|
} from "@/api/snippets";
|
||||||
import { useToastStore } from "@/stores/toast";
|
import { useToastStore } from "@/stores/toast";
|
||||||
|
|
||||||
const router = useRouter();
|
const router = useRouter();
|
||||||
@@ -50,13 +56,54 @@ const showMergeModal = ref(false);
|
|||||||
const canonicalId = ref<number | null>(null);
|
const canonicalId = ref<number | null>(null);
|
||||||
const merging = ref(false);
|
const merging = ref(false);
|
||||||
|
|
||||||
const selectedList = computed(() =>
|
// Near-duplicate report (#2088). Loaded on demand, not with the list: it's a
|
||||||
snippets.value.filter((s) => selectedIds.value.has(s.id)),
|
// pairwise scan and most visits to this page aren't a tidy-up.
|
||||||
);
|
const duplicateGroups = ref<DuplicateGroup[]>([]);
|
||||||
|
const dupLoading = ref(false);
|
||||||
|
const dupChecked = ref(false);
|
||||||
|
// Set while merging a SUGGESTED group. The report reaches the whole corpus, so
|
||||||
|
// its members need not all be on the current page — see selectedList.
|
||||||
|
const reviewingGroup = ref<DuplicateGroup | null>(null);
|
||||||
|
|
||||||
|
/** The records the merge modal acts on.
|
||||||
|
*
|
||||||
|
* Normally that's the selection filtered against what's on screen. But a
|
||||||
|
* suggested group is corpus-wide: filtering it by the current page would render
|
||||||
|
* an incomplete set AND silently narrow what doMerge folds in, since it derives
|
||||||
|
* its source ids from this list. When a group is under review it is the
|
||||||
|
* authority. */
|
||||||
|
const selectedList = computed<{ id: number; title: string }[]>(() => {
|
||||||
|
if (reviewingGroup.value) return reviewingGroup.value.snippets;
|
||||||
|
return snippets.value.filter((s) => selectedIds.value.has(s.id));
|
||||||
|
});
|
||||||
|
|
||||||
|
async function loadDuplicates() {
|
||||||
|
dupLoading.value = true;
|
||||||
|
try {
|
||||||
|
const data = await findDuplicateSnippets();
|
||||||
|
duplicateGroups.value = data.groups;
|
||||||
|
dupChecked.value = true;
|
||||||
|
} catch {
|
||||||
|
toast.show("Couldn't check for duplicates", "error");
|
||||||
|
} finally {
|
||||||
|
dupLoading.value = false;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Hand a suggested group to the existing merge flow, pre-selected. The operator
|
||||||
|
* still picks which record survives and confirms — the report proposes, it
|
||||||
|
* never merges. */
|
||||||
|
function reviewGroup(group: DuplicateGroup) {
|
||||||
|
reviewingGroup.value = group;
|
||||||
|
selectedIds.value = new Set(group.note_ids);
|
||||||
|
canonicalId.value = group.note_ids[0] ?? null;
|
||||||
|
showMergeModal.value = true;
|
||||||
|
}
|
||||||
|
|
||||||
function exitSelectMode() {
|
function exitSelectMode() {
|
||||||
selectMode.value = false;
|
selectMode.value = false;
|
||||||
selectedIds.value = new Set();
|
selectedIds.value = new Set();
|
||||||
|
reviewingGroup.value = null;
|
||||||
}
|
}
|
||||||
function toggleSelectMode() {
|
function toggleSelectMode() {
|
||||||
if (selectMode.value) exitSelectMode();
|
if (selectMode.value) exitSelectMode();
|
||||||
@@ -77,6 +124,15 @@ function openMerge() {
|
|||||||
canonicalId.value = selectedList.value[0]?.id ?? null;
|
canonicalId.value = selectedList.value[0]?.id ?? null;
|
||||||
showMergeModal.value = true;
|
showMergeModal.value = true;
|
||||||
}
|
}
|
||||||
|
/** Dismiss the modal. Clears the reviewed group too — leaving it set would keep
|
||||||
|
* selectedList pinned to a corpus-wide set the operator has walked away from. */
|
||||||
|
function closeMerge() {
|
||||||
|
showMergeModal.value = false;
|
||||||
|
if (reviewingGroup.value) {
|
||||||
|
reviewingGroup.value = null;
|
||||||
|
selectedIds.value = new Set();
|
||||||
|
}
|
||||||
|
}
|
||||||
async function doMerge() {
|
async function doMerge() {
|
||||||
const target = canonicalId.value;
|
const target = canonicalId.value;
|
||||||
if (target == null) return;
|
if (target == null) return;
|
||||||
@@ -87,8 +143,13 @@ async function doMerge() {
|
|||||||
await mergeSnippets(target, sources);
|
await mergeSnippets(target, sources);
|
||||||
toast.show(`Merged ${sources.length} snippet${sources.length > 1 ? "s" : ""} in`);
|
toast.show(`Merged ${sources.length} snippet${sources.length > 1 ? "s" : ""} in`);
|
||||||
showMergeModal.value = false;
|
showMergeModal.value = false;
|
||||||
|
const wasSuggested = reviewingGroup.value !== null;
|
||||||
exitSelectMode();
|
exitSelectMode();
|
||||||
await loadSnippets();
|
await loadSnippets();
|
||||||
|
// The merged-away records are gone, so a stale report would keep offering
|
||||||
|
// them. Re-run it rather than clearing, so the operator can work through
|
||||||
|
// several groups without re-triggering the scan each time.
|
||||||
|
if (wasSuggested && dupChecked.value) await loadDuplicates();
|
||||||
} catch {
|
} catch {
|
||||||
toast.show("Failed to merge snippets", "error");
|
toast.show("Failed to merge snippets", "error");
|
||||||
} finally {
|
} finally {
|
||||||
@@ -259,6 +320,38 @@ function usageTitle(s: SnippetListItem): string {
|
|||||||
>
|
>
|
||||||
{{ needsAttentionOnly ? "Needs attention · filtering" : "Needs attention" }}
|
{{ needsAttentionOnly ? "Needs attention · filtering" : "Needs attention" }}
|
||||||
</button>
|
</button>
|
||||||
|
<button
|
||||||
|
class="btn-ghost"
|
||||||
|
:disabled="dupLoading"
|
||||||
|
title="Look for snippets already recorded that resemble each other closely enough to be worth merging"
|
||||||
|
@click="loadDuplicates"
|
||||||
|
>
|
||||||
|
{{ dupLoading ? "Checking…" : "Find duplicates" }}
|
||||||
|
</button>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<!-- Near-duplicate report. Only ever a proposal — merging is a separate,
|
||||||
|
confirmed act, and the operator chooses which record survives. -->
|
||||||
|
<div v-if="dupChecked && !dupLoading" class="dup-panel">
|
||||||
|
<p v-if="!duplicateGroups.length" class="dup-empty">
|
||||||
|
No near-duplicates found. Nothing recorded resembles anything else closely
|
||||||
|
enough to be worth merging.
|
||||||
|
</p>
|
||||||
|
<template v-else>
|
||||||
|
<p class="dup-head">
|
||||||
|
{{ duplicateGroups.length }} possible duplicate{{ duplicateGroups.length > 1 ? " sets" : " set" }}
|
||||||
|
— review each before merging; a set is a suggestion, not a verdict.
|
||||||
|
</p>
|
||||||
|
<div v-for="(g, i) in duplicateGroups" :key="i" class="dup-group">
|
||||||
|
<div class="dup-members">
|
||||||
|
<span v-for="s in g.snippets" :key="s.id" class="dup-member">
|
||||||
|
{{ splitTitle(s.title).name }}
|
||||||
|
</span>
|
||||||
|
</div>
|
||||||
|
<span class="dup-score">{{ Math.round(g.top_score * 100) }}% alike</span>
|
||||||
|
<button class="btn-ghost dup-action" @click="reviewGroup(g)">Review & merge</button>
|
||||||
|
</div>
|
||||||
|
</template>
|
||||||
</div>
|
</div>
|
||||||
|
|
||||||
<!-- Reverse lookup: what's already kept in this repo / file / symbol. -->
|
<!-- Reverse lookup: what's already kept in this repo / file / symbol. -->
|
||||||
@@ -389,7 +482,7 @@ function usageTitle(s: SnippetListItem): string {
|
|||||||
|
|
||||||
<!-- Merge modal -->
|
<!-- Merge modal -->
|
||||||
<teleport to="body">
|
<teleport to="body">
|
||||||
<div v-if="showMergeModal" class="modal-overlay" @click.self="showMergeModal = false">
|
<div v-if="showMergeModal" class="modal-overlay" @click.self="closeMerge">
|
||||||
<div class="modal-card" role="dialog" aria-modal="true" aria-label="Merge snippets">
|
<div class="modal-card" role="dialog" aria-modal="true" aria-label="Merge snippets">
|
||||||
<h3 class="modal-title">Merge snippets</h3>
|
<h3 class="modal-title">Merge snippets</h3>
|
||||||
<p class="modal-desc">
|
<p class="modal-desc">
|
||||||
@@ -409,7 +502,7 @@ function usageTitle(s: SnippetListItem): string {
|
|||||||
</label>
|
</label>
|
||||||
</div>
|
</div>
|
||||||
<div class="modal-actions">
|
<div class="modal-actions">
|
||||||
<button class="modal-btn" @click="showMergeModal = false">Cancel</button>
|
<button class="modal-btn" @click="closeMerge">Cancel</button>
|
||||||
<button
|
<button
|
||||||
class="modal-btn modal-btn-primary"
|
class="modal-btn modal-btn-primary"
|
||||||
:disabled="merging || canonicalId == null"
|
:disabled="merging || canonicalId == null"
|
||||||
@@ -686,6 +779,63 @@ function usageTitle(s: SnippetListItem): string {
|
|||||||
color: var(--color-text-muted);
|
color: var(--color-text-muted);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/* Near-duplicate report */
|
||||||
|
.dup-panel {
|
||||||
|
margin-bottom: 1.25rem;
|
||||||
|
padding: 0.85rem 1rem;
|
||||||
|
border: 1px solid var(--color-border);
|
||||||
|
border-radius: 8px;
|
||||||
|
background: var(--color-surface-alt, var(--color-surface));
|
||||||
|
}
|
||||||
|
|
||||||
|
.dup-empty,
|
||||||
|
.dup-head {
|
||||||
|
margin: 0 0 0.5rem;
|
||||||
|
font-size: 0.85rem;
|
||||||
|
color: var(--color-text-muted);
|
||||||
|
}
|
||||||
|
|
||||||
|
.dup-empty {
|
||||||
|
margin-bottom: 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
.dup-group {
|
||||||
|
display: flex;
|
||||||
|
align-items: center;
|
||||||
|
gap: 0.75rem;
|
||||||
|
flex-wrap: wrap;
|
||||||
|
padding: 0.5rem 0;
|
||||||
|
border-top: 1px solid var(--color-border);
|
||||||
|
}
|
||||||
|
|
||||||
|
.dup-members {
|
||||||
|
display: flex;
|
||||||
|
gap: 0.4rem;
|
||||||
|
flex-wrap: wrap;
|
||||||
|
flex: 1 1 20rem;
|
||||||
|
min-width: 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
.dup-member {
|
||||||
|
font-size: 0.8rem;
|
||||||
|
padding: 0.1rem 0.45rem;
|
||||||
|
border-radius: 4px;
|
||||||
|
background: color-mix(in srgb, var(--color-text-muted) 12%, transparent);
|
||||||
|
/* Long snippet names must not push the row into a horizontal scroll. */
|
||||||
|
overflow-wrap: anywhere;
|
||||||
|
}
|
||||||
|
|
||||||
|
.dup-score {
|
||||||
|
font-size: 0.75rem;
|
||||||
|
color: var(--color-text-muted);
|
||||||
|
font-variant-numeric: tabular-nums;
|
||||||
|
white-space: nowrap;
|
||||||
|
}
|
||||||
|
|
||||||
|
.dup-action {
|
||||||
|
white-space: nowrap;
|
||||||
|
}
|
||||||
|
|
||||||
/* Drift is a stronger signal than dead weight: the record may be actively
|
/* Drift is a stronger signal than dead weight: the record may be actively
|
||||||
misleading, not merely unused. Danger tone, and it sits first in the footer. */
|
misleading, not merely unused. Danger tone, and it sits first in the footer. */
|
||||||
.drift-tag {
|
.drift-tag {
|
||||||
|
|||||||
@@ -256,6 +256,9 @@ _READ_ONLY_TOOLS = frozenset({
|
|||||||
"list_rules", "list_tags", "list_tasks", "list_topics", "list_trash",
|
"list_rules", "list_tags", "list_tasks", "list_topics", "list_trash",
|
||||||
"list_always_on_rules", "search",
|
"list_always_on_rules", "search",
|
||||||
"get_system", "list_systems", "list_system_records",
|
"get_system", "list_systems", "list_system_records",
|
||||||
|
# Reports on the snippet corpus. Reads only — the merge it suggests is a
|
||||||
|
# separate, explicitly-called write.
|
||||||
|
"find_duplicate_snippets",
|
||||||
})
|
})
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
@@ -190,6 +190,44 @@ async def get_snippet(snippet_id: int) -> dict:
|
|||||||
return data
|
return data
|
||||||
|
|
||||||
|
|
||||||
|
async def find_duplicate_snippets(threshold: float = 0.0) -> dict:
|
||||||
|
"""Find snippets already recorded that look like duplicates of each other.
|
||||||
|
|
||||||
|
The create gate PREVENTS a new duplicate and merge_snippets CURES one you
|
||||||
|
point it at — this is the missing third piece: it FINDS the ones already in
|
||||||
|
the record, so nobody has to notice them by hand.
|
||||||
|
|
||||||
|
Results are grouped into candidate merge SETS, not just pairs. Grouping is
|
||||||
|
transitive: if A resembles B and B resembles C, all three land in one set
|
||||||
|
even when A and C don't directly clear the bar. That mirrors what merge does
|
||||||
|
(it folds every source into one survivor), but it means a chain of mild
|
||||||
|
resemblances can rope in a member that isn't really alike — so read a set as
|
||||||
|
a proposal and check the members before acting.
|
||||||
|
|
||||||
|
Reports only YOUR snippets. merge_snippets requires one owner across the
|
||||||
|
whole set, so surfacing someone else's would propose a merge that can't be
|
||||||
|
performed.
|
||||||
|
|
||||||
|
Acting on a group: pick the best record as the canonical target, then
|
||||||
|
`merge_snippets(target_id, [other ids])`. Merge unions the fields and folds
|
||||||
|
every source's location in, so the survivor is findable at all their call
|
||||||
|
sites; the sources are trashed, recoverably. Prefer as target the one with
|
||||||
|
the clearest "when to reach for it" — merge keeps the target's title.
|
||||||
|
|
||||||
|
Args:
|
||||||
|
threshold: Similarity floor, 0-1. 0 (default) uses the configured
|
||||||
|
setting. Raise it if the report is noisy, lower it to catch more.
|
||||||
|
|
||||||
|
Returns {"groups": [{"note_ids", "snippets", "top_score"}], "pairs",
|
||||||
|
"threshold"}. An empty `groups` means nothing resembles anything else that
|
||||||
|
closely — the common and desirable case.
|
||||||
|
"""
|
||||||
|
uid = current_user_id()
|
||||||
|
return await dedup_svc.find_duplicate_snippets(
|
||||||
|
uid, threshold=threshold if threshold > 0 else None
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
async def verify_snippet(
|
async def verify_snippet(
|
||||||
snippet_id: int, status: str, detail: str = "", path: str = "",
|
snippet_id: int, status: str, detail: str = "", path: str = "",
|
||||||
) -> dict:
|
) -> dict:
|
||||||
@@ -365,6 +403,6 @@ async def merge_snippets(target_id: int, source_ids: list[int]) -> dict:
|
|||||||
def register(mcp) -> None:
|
def register(mcp) -> None:
|
||||||
for fn in (
|
for fn in (
|
||||||
list_snippets, create_snippet, get_snippet, update_snippet,
|
list_snippets, create_snippet, get_snippet, update_snippet,
|
||||||
delete_snippet, merge_snippets, verify_snippet,
|
delete_snippet, merge_snippets, verify_snippet, find_duplicate_snippets,
|
||||||
):
|
):
|
||||||
mcp.tool(name=fn.__name__)(fn)
|
mcp.tool(name=fn.__name__)(fn)
|
||||||
|
|||||||
@@ -205,6 +205,24 @@ async def update_snippet_route(snippet_id: int):
|
|||||||
return jsonify(out)
|
return jsonify(out)
|
||||||
|
|
||||||
|
|
||||||
|
@snippets_bp.route("/duplicates", methods=["GET"])
|
||||||
|
@login_required
|
||||||
|
async def duplicate_snippets_route():
|
||||||
|
"""Near-duplicate snippets already recorded, grouped into merge candidates.
|
||||||
|
|
||||||
|
Registered ABOVE the `/<int:snippet_id>` routes on purpose — Quart matches
|
||||||
|
an int converter before a static segment either way, but keeping the literal
|
||||||
|
path first makes the precedence obvious to the next person reading this."""
|
||||||
|
uid = get_current_user_id()
|
||||||
|
try:
|
||||||
|
threshold = float(request.args.get("threshold", 0) or 0)
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
threshold = 0.0
|
||||||
|
return jsonify(await dedup_svc.find_duplicate_snippets(
|
||||||
|
uid, threshold=threshold if threshold > 0 else None
|
||||||
|
))
|
||||||
|
|
||||||
|
|
||||||
@snippets_bp.route("/<int:snippet_id>/verify", methods=["POST"])
|
@snippets_bp.route("/<int:snippet_id>/verify", methods=["POST"])
|
||||||
@login_required
|
@login_required
|
||||||
async def verify_snippet_route(snippet_id: int):
|
async def verify_snippet_route(snippet_id: int):
|
||||||
|
|||||||
@@ -26,11 +26,17 @@ import logging
|
|||||||
from dataclasses import dataclass
|
from dataclasses import dataclass
|
||||||
|
|
||||||
from sqlalchemy import func, select
|
from sqlalchemy import func, select
|
||||||
|
from sqlalchemy.orm import aliased
|
||||||
|
|
||||||
from scribe.models import async_session
|
from scribe.models import async_session
|
||||||
|
from scribe.models.embedding import NoteEmbedding
|
||||||
from scribe.models.note import Note
|
from scribe.models.note import Note
|
||||||
from scribe.models.rulebook import Rule
|
from scribe.models.rulebook import Rule
|
||||||
from scribe.services import embeddings as embeddings_svc
|
from scribe.services import embeddings as embeddings_svc
|
||||||
|
# Imported rather than redeclared: no service imports this module (the create
|
||||||
|
# gate is called from the routes/tools layer), so there is no cycle to dodge,
|
||||||
|
# and a second copy of the constant is a thing to drift.
|
||||||
|
from scribe.services.snippets import SNIPPET_NOTE_TYPE
|
||||||
|
|
||||||
logger = logging.getLogger(__name__)
|
logger = logging.getLogger(__name__)
|
||||||
|
|
||||||
@@ -139,6 +145,174 @@ async def find_duplicate_note(
|
|||||||
return None
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
# --- corpus-wide near-duplicate report (#2088) -------------------------------
|
||||||
|
# The gate above PREVENTS a new duplicate; merge_snippets CURES one you point it
|
||||||
|
# at. Neither FINDS the duplicates already sitting in the record — someone had to
|
||||||
|
# notice them by hand, which is the exact failure the Drafter exists to remove.
|
||||||
|
#
|
||||||
|
# WHY OWN SNIPPETS ONLY. merge_snippets requires every record to share the
|
||||||
|
# target's owner (cross-owner merge is out of scope), so a report that surfaced
|
||||||
|
# someone else's snippet would propose a merge that cannot be performed. The
|
||||||
|
# scope here is set by what the operator can actually act on, not by what they
|
||||||
|
# can see.
|
||||||
|
#
|
||||||
|
# WHY A LOWER THRESHOLD THAN THE GATE. The gate BLOCKS a write at 0.90 and has to
|
||||||
|
# be unforgiving of noise. This report only makes a suggestion the operator
|
||||||
|
# reviews, so it can afford to be looser and catch the pairs the gate lets
|
||||||
|
# through — which are precisely the ones that accumulated. It is a setting rather
|
||||||
|
# than a constant (rule #25) because the right value depends on how uniform a
|
||||||
|
# corpus is, and nobody can guess that from here.
|
||||||
|
|
||||||
|
DUPLICATE_THRESHOLD_KEY = "kb_duplicate_threshold"
|
||||||
|
DUPLICATE_DEFAULT_THRESHOLD = 0.82
|
||||||
|
# Hard cap on returned pairs. A pathologically uniform corpus is O(n²) pairs, and
|
||||||
|
# a report nobody can read is not a report.
|
||||||
|
_MAX_DUPLICATE_PAIRS = 200
|
||||||
|
|
||||||
|
|
||||||
|
async def get_duplicate_threshold(user_id: int) -> float:
|
||||||
|
"""The user's near-duplicate similarity floor, clamped to [0, 1]."""
|
||||||
|
from scribe.services.settings import get_setting
|
||||||
|
|
||||||
|
try:
|
||||||
|
value = float(await get_setting(
|
||||||
|
user_id, DUPLICATE_THRESHOLD_KEY, str(DUPLICATE_DEFAULT_THRESHOLD)
|
||||||
|
))
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
value = DUPLICATE_DEFAULT_THRESHOLD
|
||||||
|
return min(1.0, max(0.0, value))
|
||||||
|
|
||||||
|
|
||||||
|
def group_pairs(pairs: list[tuple[int, int, float]]) -> list[list[int]]:
|
||||||
|
"""Collapse similar-pairs into candidate merge SETS (connected components).
|
||||||
|
|
||||||
|
Pure and synchronous so the grouping rule is testable without a database.
|
||||||
|
|
||||||
|
Transitive on purpose: if A~B and B~C, all three land in one set even when
|
||||||
|
A and C fall below the threshold. That matches what merge does — it folds
|
||||||
|
every source into one survivor — and it avoids handing the operator three
|
||||||
|
overlapping pairs to reconcile by hand, which is the chore being removed.
|
||||||
|
The cost is that a chain of mild resemblances can rope in a pair that isn't
|
||||||
|
really alike; the operator sees the members and picks, so a set is a
|
||||||
|
proposal, never an action.
|
||||||
|
"""
|
||||||
|
parent: dict[int, int] = {}
|
||||||
|
|
||||||
|
def find(x: int) -> int:
|
||||||
|
parent.setdefault(x, x)
|
||||||
|
while parent[x] != x:
|
||||||
|
parent[x] = parent[parent[x]]
|
||||||
|
x = parent[x]
|
||||||
|
return x
|
||||||
|
|
||||||
|
def union(a: int, b: int) -> None:
|
||||||
|
ra, rb = find(a), find(b)
|
||||||
|
if ra != rb:
|
||||||
|
parent[rb] = ra
|
||||||
|
|
||||||
|
for left, right, _score in pairs:
|
||||||
|
union(left, right)
|
||||||
|
|
||||||
|
groups: dict[int, list[int]] = {}
|
||||||
|
for node in parent:
|
||||||
|
groups.setdefault(find(node), []).append(node)
|
||||||
|
# Biggest clusters first — the most tangled thing is the most worth fixing.
|
||||||
|
# Ids ascending within a set so the output is stable across runs.
|
||||||
|
return sorted((sorted(g) for g in groups.values() if len(g) > 1),
|
||||||
|
key=lambda g: (-len(g), g[0]))
|
||||||
|
|
||||||
|
|
||||||
|
async def find_duplicate_snippets(
|
||||||
|
user_id: int, *, threshold: float | None = None, limit: int = _MAX_DUPLICATE_PAIRS
|
||||||
|
) -> dict:
|
||||||
|
"""Near-duplicate snippets already in the record, grouped into merge sets.
|
||||||
|
|
||||||
|
One indexed self-join over `note_embeddings` rather than an N² Python scan:
|
||||||
|
pgvector's cosine distance is the same operator semantic search uses, so a
|
||||||
|
similarity floor is a distance ceiling and the work stays in Postgres.
|
||||||
|
|
||||||
|
Returns {"groups": [{"note_ids": [...], "snippets": [...], "top_score": f}],
|
||||||
|
"pairs": [...], "threshold": f}. Fail-open (an empty report) like the rest of
|
||||||
|
this module — a suggestion feature must not be able to break the page it
|
||||||
|
decorates.
|
||||||
|
"""
|
||||||
|
floor = await get_duplicate_threshold(user_id) if threshold is None else threshold
|
||||||
|
floor = min(1.0, max(0.0, floor))
|
||||||
|
max_distance = min(2.0, max(0.0, 1.0 - floor))
|
||||||
|
|
||||||
|
left = aliased(NoteEmbedding, name="left_emb")
|
||||||
|
right = aliased(NoteEmbedding, name="right_emb")
|
||||||
|
left_note = aliased(Note, name="left_note")
|
||||||
|
right_note = aliased(Note, name="right_note")
|
||||||
|
distance = left.embedding.cosine_distance(right.embedding)
|
||||||
|
|
||||||
|
pairs: list[tuple[int, int, float]] = []
|
||||||
|
try:
|
||||||
|
async with async_session() as session:
|
||||||
|
stmt = (
|
||||||
|
select(left.note_id, right.note_id, distance.label("distance"))
|
||||||
|
.select_from(left)
|
||||||
|
# `<` not `!=`: each unordered pair exactly once, and it drops
|
||||||
|
# the self-pair (distance 0) that would otherwise dominate.
|
||||||
|
.join(right, left.note_id < right.note_id)
|
||||||
|
.join(left_note, left_note.id == left.note_id)
|
||||||
|
.join(right_note, right_note.id == right.note_id)
|
||||||
|
.where(
|
||||||
|
left_note.note_type == SNIPPET_NOTE_TYPE,
|
||||||
|
right_note.note_type == SNIPPET_NOTE_TYPE,
|
||||||
|
left_note.deleted_at.is_(None),
|
||||||
|
right_note.deleted_at.is_(None),
|
||||||
|
# Owner-scoped on both sides — see the note above on why the
|
||||||
|
# report is bounded by what merge can actually act on.
|
||||||
|
left_note.user_id == user_id,
|
||||||
|
right_note.user_id == user_id,
|
||||||
|
distance <= max_distance,
|
||||||
|
)
|
||||||
|
.order_by(distance.asc())
|
||||||
|
.limit(max(1, limit))
|
||||||
|
)
|
||||||
|
rows = list((await session.execute(stmt)).all())
|
||||||
|
pairs = [(int(a), int(b), round(1.0 - float(d), 4)) for a, b, d in rows]
|
||||||
|
except Exception:
|
||||||
|
logger.warning("Near-duplicate snippet scan failed", exc_info=True)
|
||||||
|
return {"groups": [], "pairs": [], "threshold": floor}
|
||||||
|
|
||||||
|
if not pairs:
|
||||||
|
return {"groups": [], "pairs": [], "threshold": floor}
|
||||||
|
|
||||||
|
best: dict[tuple[int, int], float] = {(a, b): s for a, b, s in pairs}
|
||||||
|
grouped = group_pairs(pairs)
|
||||||
|
|
||||||
|
# Titles for presentation. One fetch for every id in the report.
|
||||||
|
ids = sorted({n for g in grouped for n in g})
|
||||||
|
titles: dict[int, str] = {}
|
||||||
|
try:
|
||||||
|
async with async_session() as session:
|
||||||
|
rows = (await session.execute(
|
||||||
|
select(Note.id, Note.title).where(Note.id.in_(ids))
|
||||||
|
)).all()
|
||||||
|
titles = {int(i): t for i, t in rows}
|
||||||
|
except Exception:
|
||||||
|
logger.debug("duplicate report titles unavailable", exc_info=True)
|
||||||
|
|
||||||
|
groups = []
|
||||||
|
for members in grouped:
|
||||||
|
scores = [
|
||||||
|
s for (a, b), s in best.items() if a in members and b in members
|
||||||
|
]
|
||||||
|
groups.append({
|
||||||
|
"note_ids": members,
|
||||||
|
"snippets": [
|
||||||
|
{"id": nid, "title": titles.get(nid, "")} for nid in members
|
||||||
|
],
|
||||||
|
# The strongest resemblance in the set — how confident the suggestion
|
||||||
|
# is, and what the list sorts on.
|
||||||
|
"top_score": max(scores) if scores else floor,
|
||||||
|
})
|
||||||
|
groups.sort(key=lambda g: (-g["top_score"], g["note_ids"][0]))
|
||||||
|
return {"groups": groups, "pairs": pairs, "threshold": floor}
|
||||||
|
|
||||||
|
|
||||||
async def find_duplicate_rule(
|
async def find_duplicate_rule(
|
||||||
title: str,
|
title: str,
|
||||||
topic_id: int | None = None,
|
topic_id: int | None = None,
|
||||||
|
|||||||
@@ -229,4 +229,5 @@ def test_register_attaches_all_tools():
|
|||||||
assert set(names) == {
|
assert set(names) == {
|
||||||
"list_snippets", "create_snippet", "get_snippet", "update_snippet",
|
"list_snippets", "create_snippet", "get_snippet", "update_snippet",
|
||||||
"delete_snippet", "merge_snippets", "verify_snippet",
|
"delete_snippet", "merge_snippets", "verify_snippet",
|
||||||
|
"find_duplicate_snippets",
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -20,7 +20,7 @@ def test_snippet_handlers_callable():
|
|||||||
for name in (
|
for name in (
|
||||||
"list_snippets_route", "create_snippet_route", "get_snippet_route",
|
"list_snippets_route", "create_snippet_route", "get_snippet_route",
|
||||||
"update_snippet_route", "delete_snippet_route", "merge_snippet_route",
|
"update_snippet_route", "delete_snippet_route", "merge_snippet_route",
|
||||||
"verify_snippet_route",
|
"verify_snippet_route", "duplicate_snippets_route",
|
||||||
):
|
):
|
||||||
assert callable(getattr(routes, name))
|
assert callable(getattr(routes, name))
|
||||||
|
|
||||||
@@ -68,6 +68,11 @@ def test_agent_and_web_surfaces_stay_at_parity():
|
|||||||
assert "verification" in inspect.signature(tools.list_snippets).parameters
|
assert "verification" in inspect.signature(tools.list_snippets).parameters
|
||||||
assert "verification" in inspect.getsource(routes.list_snippets_route)
|
assert "verification" in inspect.getsource(routes.list_snippets_route)
|
||||||
|
|
||||||
|
# The near-duplicate report (#2088) reaches both. The web side can only hand
|
||||||
|
# a group to the merge flow if it can get the groups in the first place.
|
||||||
|
assert callable(getattr(tools, "find_duplicate_snippets", None))
|
||||||
|
assert callable(getattr(routes, "duplicate_snippets_route", None))
|
||||||
|
|
||||||
|
|
||||||
def test_project_scoping_reaches_every_caller():
|
def test_project_scoping_reaches_every_caller():
|
||||||
"""A snippet search has to be narrowable to one project from both surfaces."""
|
"""A snippet search has to be narrowable to one project from both surfaces."""
|
||||||
|
|||||||
@@ -0,0 +1,99 @@
|
|||||||
|
"""Tests for the near-duplicate finder (#2088).
|
||||||
|
|
||||||
|
The SQL half needs a database and lives in the integration lane; what's covered
|
||||||
|
here is the grouping rule — which is where the interesting decisions are — plus
|
||||||
|
the fail-open contract and the threshold resolution.
|
||||||
|
"""
|
||||||
|
from unittest.mock import AsyncMock, patch
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from scribe.services import dedup as dedup_svc
|
||||||
|
from scribe.services.dedup import (
|
||||||
|
DUPLICATE_DEFAULT_THRESHOLD,
|
||||||
|
get_duplicate_threshold,
|
||||||
|
group_pairs,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
# --- grouping -------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_simple_pair_becomes_one_set():
|
||||||
|
assert group_pairs([(1, 2, 0.9)]) == [[1, 2]]
|
||||||
|
|
||||||
|
|
||||||
|
def test_grouping_is_transitive():
|
||||||
|
"""A~B and B~C puts all three in one set even though A and C never cleared
|
||||||
|
the bar together. That mirrors merge, which folds every source into one
|
||||||
|
survivor — three overlapping pairs would just be the same chore, unsorted."""
|
||||||
|
assert group_pairs([(1, 2, 0.9), (2, 3, 0.9)]) == [[1, 2, 3]]
|
||||||
|
|
||||||
|
|
||||||
|
def test_disjoint_clusters_stay_separate():
|
||||||
|
groups = group_pairs([(1, 2, 0.9), (3, 4, 0.9)])
|
||||||
|
assert groups == [[1, 2], [3, 4]]
|
||||||
|
|
||||||
|
|
||||||
|
def test_largest_cluster_comes_first():
|
||||||
|
"""The most tangled thing is the most worth fixing."""
|
||||||
|
groups = group_pairs([(5, 6, 0.9), (1, 2, 0.9), (2, 3, 0.9), (3, 4, 0.9)])
|
||||||
|
assert groups[0] == [1, 2, 3, 4]
|
||||||
|
|
||||||
|
|
||||||
|
def test_output_is_stable_across_input_order():
|
||||||
|
"""A report that reshuffles between runs is one nobody can work through."""
|
||||||
|
a = group_pairs([(1, 2, 0.9), (2, 3, 0.9), (7, 8, 0.9)])
|
||||||
|
b = group_pairs([(7, 8, 0.9), (2, 3, 0.9), (1, 2, 0.9)])
|
||||||
|
assert a == b
|
||||||
|
|
||||||
|
|
||||||
|
def test_no_pairs_means_no_groups():
|
||||||
|
assert group_pairs([]) == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_node_never_forms_a_group_with_itself():
|
||||||
|
"""The SQL uses `note_id <` so this shouldn't arrive, but a singleton set
|
||||||
|
would render as a "duplicate" of nothing and offer an impossible merge."""
|
||||||
|
assert group_pairs([(1, 1, 1.0)]) == []
|
||||||
|
|
||||||
|
|
||||||
|
# --- threshold ------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
async def test_threshold_falls_back_to_the_default_when_unset():
|
||||||
|
with patch("scribe.services.settings.get_setting",
|
||||||
|
AsyncMock(return_value=str(DUPLICATE_DEFAULT_THRESHOLD))):
|
||||||
|
assert await get_duplicate_threshold(1) == DUPLICATE_DEFAULT_THRESHOLD
|
||||||
|
|
||||||
|
|
||||||
|
async def test_a_garbage_setting_falls_back_rather_than_raising():
|
||||||
|
with patch("scribe.services.settings.get_setting",
|
||||||
|
AsyncMock(return_value="not-a-number")):
|
||||||
|
assert await get_duplicate_threshold(1) == DUPLICATE_DEFAULT_THRESHOLD
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("stored, expected", [("2.5", 1.0), ("-3", 0.0)])
|
||||||
|
async def test_threshold_is_clamped_to_the_valid_range(stored, expected):
|
||||||
|
with patch("scribe.services.settings.get_setting", AsyncMock(return_value=stored)):
|
||||||
|
assert await get_duplicate_threshold(1) == expected
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_report_threshold_is_looser_than_the_write_gate():
|
||||||
|
"""The gate BLOCKS a create and must be unforgiving of noise; this only
|
||||||
|
suggests a merge the operator reviews, so it has to reach further or it
|
||||||
|
would never surface the pairs the gate already let through."""
|
||||||
|
assert DUPLICATE_DEFAULT_THRESHOLD < dedup_svc._SEMANTIC_THRESHOLD
|
||||||
|
|
||||||
|
|
||||||
|
# --- fail-open ------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
async def test_a_failed_scan_returns_an_empty_report_not_an_error():
|
||||||
|
"""A suggestion feature must not be able to break the page it decorates."""
|
||||||
|
with (
|
||||||
|
patch.object(dedup_svc, "get_duplicate_threshold", AsyncMock(return_value=0.8)),
|
||||||
|
patch.object(dedup_svc, "async_session", side_effect=RuntimeError("boom")),
|
||||||
|
):
|
||||||
|
out = await dedup_svc.find_duplicate_snippets(1)
|
||||||
|
assert out == {"groups": [], "pairs": [], "threshold": 0.8}
|
||||||
Reference in New Issue
Block a user