From c8bfa6947c157dcd83c494b3cabe98a04a2ee082 Mon Sep 17 00:00:00 2001 From: Bryan Van Deusen Date: Thu, 17 Sep 2026 11:38:16 -0400 Subject: [PATCH] docs(retrieval): stop asking the operator to diagnose retrieval (#4102) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The last item in this step's done-when: no wording anywhere asks an operator to tune for correctness. Four Settings hints told them to do exactly that — "raise it if rules keep arriving unread", "lower it if a git push arrives with nothing", "lower this if genuine duplicates go unnoticed". Every one of those asks the operator to diagnose a ranker from symptoms, which is the job the model now does from the records: the telemetry says what each bar refused, and reading those records is what separates a real miss from a bar doing its job. The hints keep the explanation of WHAT each number is — that is worth reading — and drop the homework. `plugin_context.py` said the defaults "are meant to be tuned from retrieval_logs once data accrues", which was true and had no owner. It now names who does it and with what. `retrieval_telemetry`'s docstring gained the warning that belongs beside it: this readout has been measured pointing the wrong way, so the ids it returns are the point, not its percentiles. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy --- frontend/src/views/SettingsView.vue | 20 ++++++++++---------- src/scribe/mcp/tools/search.py | 12 +++++++----- src/scribe/services/plugin_context.py | 11 ++++++++--- 3 files changed, 25 insertions(+), 18 deletions(-) diff --git a/frontend/src/views/SettingsView.vue b/frontend/src/views/SettingsView.vue index 79920f2..8b1b452 100644 --- a/frontend/src/views/SettingsView.vue +++ b/frontend/src/views/SettingsView.vue @@ -1605,9 +1605,9 @@ async function deleteUser(userId: number) { Stricter than the prompt threshold above on purpose. Any two pieces of code look somewhat alike — shared keywords, indentation, structure — so resemblance scores start higher for code than for prose, and a bar tuned - for prompts flags unrelated code as prior art. Lower this if genuine - duplicates go unnoticed; raise it if you're being offered snippets that - have nothing to do with what's being written. Snippets recorded at the + for prompts flags unrelated code as prior art. Claude keeps this + one current from what the arm actually surfaced and refused; set it + yourself if you disagree with where it has landed. Snippets recorded at the exact file are always shown regardless — those are prior art by location, not by resemblance.

@@ -1644,8 +1644,9 @@ async function deleteUser(userId: number) { preloaded any more, so this is the only way a rule reaches a write. Stricter than the threshold above, because there are far fewer rules than snippets: with a small set something always ranks first, so the - bar has to carry more of the judgement. Raise it if rules keep - arriving unread; lower it if a rule you needed never showed up. + bar has to carry more of the judgement. If rules arrive unread, or one + you needed never showed up, that is Claude's to notice and correct + from the telemetry — and the budget below is usually the better lever.

@@ -1679,8 +1680,7 @@ async function deleteUser(userId: number) { query is the command text rather than code. Lower than the one above on purpose: a shell command is short, so it scores lower for the same relevance — at a shared bar this arm spoke on 2% of calls against the - write path's 37%. Raise it if commands attract rules that do not - apply; lower it if a git push arrives with nothing. + write path's 37%.

@@ -1748,9 +1748,9 @@ async function deleteUser(userId: number) { written, looked up when a task closes. Unlike every other bar here, the question this arm asks never changes — so its score is fixed by your preferences alone, and it will either always find one - or never find one. If you have written a preference for report shape - and it is not arriving, lower this; there is no run of calls that - will reveal the problem on its own. + or never find one, and no run of calls will reveal a dead one on its + own. That is why this arm is worth looking up in the panel below when + a report preference never seems to arrive.

diff --git a/src/scribe/mcp/tools/search.py b/src/scribe/mcp/tools/search.py index 57c2a22..17affb3 100644 --- a/src/scribe/mcp/tools/search.py +++ b/src/scribe/mcp/tools/search.py @@ -209,11 +209,13 @@ async def retrieval_telemetry( ) -> dict: """What the retrieval telemetry says about YOUR surfaces, over a window. - The read half of the loop the ranker's thresholds are meant to be tuned - from (#2975). Reach for it before changing a similarity threshold, a top-k, - or deciding whether a reranker is worth building — the alternative is - hand-probing the live instance, which is how the last such decision had to - be made. + The read half of the tuning loop, whose write half is `tune_retrieval` + (#2975, #4102). Reach for it before moving any floor or budget, and read + the records it names rather than its percentiles alone: this readout has + been measured pointing the WRONG WAY — 69 consecutive declines where every + percentile said "lower the bar" and the refused record was a false positive + — so `near_miss_samples=5` and opening the ids it returns is the step that + separates a real miss from a bar doing its job. Three readouts, from the three tables built for them: diff --git a/src/scribe/services/plugin_context.py b/src/scribe/services/plugin_context.py index 3a6b449..6d154f9 100644 --- a/src/scribe/services/plugin_context.py +++ b/src/scribe/services/plugin_context.py @@ -56,9 +56,14 @@ _GOAL_CHARS = 200 # Per-user settings (keys live in the generic settings table). The threshold is # deliberately STRICTER than the pull-search default (embeddings # DEFAULT_SIMILARITY_THRESHOLD = 0.45): an unsolicited per-turn inject must clear -# a higher bar than a search the agent chose to run. Defaults start conservative -# and are meant to be tuned from retrieval_logs (source='auto_inject') once data -# accrues — they're exposed in the Settings UI, no restart needed. +# a higher bar than a search the agent chose to run. +# +# The defaults below are STARTING POINTS, and correcting them is the model's +# job, not the operator's (#4102): `retrieval_surfaces` says what is in force, +# `retrieval_telemetry(near_miss_samples=N)` says what it refused, and +# `tune_retrieval` moves it with the reason attached. The operator can set any +# of them in Settings and their change is recorded the same way — but nobody +# has to read a log to get correct behaviour out of this. AUTOINJECT_ENABLED_KEY = "kb_autoinject_enabled" # The key and the value both live in the registry now (#4102); these names # survive because the comments above them are where each number's measurement