Every ledger is cleared on a compact, and a retrieval floor becomes something the model maintains #163

Merged
bvandeusen merged 12 commits from dev into main 2026-09-17 12:01:42 -04:00
3 changed files with 25 additions and 18 deletions
Showing only changes of commit c8bfa6947c - Show all commits
+10 -10
View File
@@ -1605,9 +1605,9 @@ async function deleteUser(userId: number) {
Stricter than the prompt threshold above on purpose. Any two pieces of
code look somewhat alike shared keywords, indentation, structure so
resemblance scores start higher for code than for prose, and a bar tuned
for prompts flags unrelated code as prior art. Lower this if genuine
duplicates go unnoticed; raise it if you're being offered snippets that
have nothing to do with what's being written. Snippets recorded at the
for prompts flags unrelated code as prior art. Claude keeps this
one current from what the arm actually surfaced and refused; set it
yourself if you disagree with where it has landed. Snippets recorded at the
exact file are always shown regardless those are prior art by
location, not by resemblance.
</p>
@@ -1644,8 +1644,9 @@ async function deleteUser(userId: number) {
preloaded any more, so this is the only way a rule reaches a write.
Stricter than the threshold above, because there are far fewer rules
than snippets: with a small set something always ranks first, so the
bar has to carry more of the judgement. Raise it if rules keep
arriving unread; lower it if a rule you needed never showed up.
bar has to carry more of the judgement. If rules arrive unread, or one
you needed never showed up, that is Claude's to notice and correct
from the telemetry — and the budget below is usually the better lever.
</p>
</div>
<div class="field">
@@ -1679,8 +1680,7 @@ async function deleteUser(userId: number) {
query is the command text rather than code. Lower than the one above
on purpose: a shell command is short, so it scores lower for the same
relevance — at a shared bar this arm spoke on 2% of calls against the
write path's 37%. Raise it if commands attract rules that do not
apply; lower it if a <code>git push</code> arrives with nothing.
write path's 37%.
</p>
</div>
<div class="field">
@@ -1748,9 +1748,9 @@ async function deleteUser(userId: number) {
written</em>, looked up when a task closes. Unlike every other bar
here, the question this arm asks never changes so its score is
fixed by your preferences alone, and it will either always find one
or never find one. If you have written a preference for report shape
and it is not arriving, lower this; there is no run of calls that
will reveal the problem on its own.
or never find one, and no run of calls will reveal a dead one on its
own. That is why this arm is worth looking up in the panel below when
a report preference never seems to arrive.
</p>
</div>
<div class="field">
+7 -5
View File
@@ -209,11 +209,13 @@ async def retrieval_telemetry(
) -> dict:
"""What the retrieval telemetry says about YOUR surfaces, over a window.
The read half of the loop the ranker's thresholds are meant to be tuned
from (#2975). Reach for it before changing a similarity threshold, a top-k,
or deciding whether a reranker is worth building — the alternative is
hand-probing the live instance, which is how the last such decision had to
be made.
The read half of the tuning loop, whose write half is `tune_retrieval`
(#2975, #4102). Reach for it before moving any floor or budget, and read
the records it names rather than its percentiles alone: this readout has
been measured pointing the WRONG WAY — 69 consecutive declines where every
percentile said "lower the bar" and the refused record was a false positive
— so `near_miss_samples=5` and opening the ids it returns is the step that
separates a real miss from a bar doing its job.
Three readouts, from the three tables built for them:
+8 -3
View File
@@ -56,9 +56,14 @@ _GOAL_CHARS = 200
# Per-user settings (keys live in the generic settings table). The threshold is
# deliberately STRICTER than the pull-search default (embeddings
# DEFAULT_SIMILARITY_THRESHOLD = 0.45): an unsolicited per-turn inject must clear
# a higher bar than a search the agent chose to run. Defaults start conservative
# and are meant to be tuned from retrieval_logs (source='auto_inject') once data
# accrues — they're exposed in the Settings UI, no restart needed.
# a higher bar than a search the agent chose to run.
#
# The defaults below are STARTING POINTS, and correcting them is the model's
# job, not the operator's (#4102): `retrieval_surfaces` says what is in force,
# `retrieval_telemetry(near_miss_samples=N)` says what it refused, and
# `tune_retrieval` moves it with the reason attached. The operator can set any
# of them in Settings and their change is recorded the same way — but nobody
# has to read a log to get correct behaviour out of this.
AUTOINJECT_ENABLED_KEY = "kb_autoinject_enabled"
# The key and the value both live in the registry now (#4102); these names
# survive because the comments above them are where each number's measurement