docs(retrieval): stop asking the operator to diagnose retrieval (#4102)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m30s
CI & Build / Build & push image (push) Successful in 33s
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m30s
CI & Build / Build & push image (push) Successful in 33s
The last item in this step's done-when: no wording anywhere asks an operator to tune for correctness. Four Settings hints told them to do exactly that — "raise it if rules keep arriving unread", "lower it if a git push arrives with nothing", "lower this if genuine duplicates go unnoticed". Every one of those asks the operator to diagnose a ranker from symptoms, which is the job the model now does from the records: the telemetry says what each bar refused, and reading those records is what separates a real miss from a bar doing its job. The hints keep the explanation of WHAT each number is — that is worth reading — and drop the homework. `plugin_context.py` said the defaults "are meant to be tuned from retrieval_logs once data accrues", which was true and had no owner. It now names who does it and with what. `retrieval_telemetry`'s docstring gained the warning that belongs beside it: this readout has been measured pointing the wrong way, so the ids it returns are the point, not its percentiles. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
@@ -1605,9 +1605,9 @@ async function deleteUser(userId: number) {
|
||||
Stricter than the prompt threshold above on purpose. Any two pieces of
|
||||
code look somewhat alike — shared keywords, indentation, structure — so
|
||||
resemblance scores start higher for code than for prose, and a bar tuned
|
||||
for prompts flags unrelated code as prior art. Lower this if genuine
|
||||
duplicates go unnoticed; raise it if you're being offered snippets that
|
||||
have nothing to do with what's being written. Snippets recorded at the
|
||||
for prompts flags unrelated code as prior art. Claude keeps this
|
||||
one current from what the arm actually surfaced and refused; set it
|
||||
yourself if you disagree with where it has landed. Snippets recorded at the
|
||||
exact file are always shown regardless — those are prior art by
|
||||
location, not by resemblance.
|
||||
</p>
|
||||
@@ -1644,8 +1644,9 @@ async function deleteUser(userId: number) {
|
||||
preloaded any more, so this is the only way a rule reaches a write.
|
||||
Stricter than the threshold above, because there are far fewer rules
|
||||
than snippets: with a small set something always ranks first, so the
|
||||
bar has to carry more of the judgement. Raise it if rules keep
|
||||
arriving unread; lower it if a rule you needed never showed up.
|
||||
bar has to carry more of the judgement. If rules arrive unread, or one
|
||||
you needed never showed up, that is Claude's to notice and correct
|
||||
from the telemetry — and the budget below is usually the better lever.
|
||||
</p>
|
||||
</div>
|
||||
<div class="field">
|
||||
@@ -1679,8 +1680,7 @@ async function deleteUser(userId: number) {
|
||||
query is the command text rather than code. Lower than the one above
|
||||
on purpose: a shell command is short, so it scores lower for the same
|
||||
relevance — at a shared bar this arm spoke on 2% of calls against the
|
||||
write path's 37%. Raise it if commands attract rules that do not
|
||||
apply; lower it if a <code>git push</code> arrives with nothing.
|
||||
write path's 37%.
|
||||
</p>
|
||||
</div>
|
||||
<div class="field">
|
||||
@@ -1748,9 +1748,9 @@ async function deleteUser(userId: number) {
|
||||
written</em>, looked up when a task closes. Unlike every other bar
|
||||
here, the question this arm asks never changes — so its score is
|
||||
fixed by your preferences alone, and it will either always find one
|
||||
or never find one. If you have written a preference for report shape
|
||||
and it is not arriving, lower this; there is no run of calls that
|
||||
will reveal the problem on its own.
|
||||
or never find one, and no run of calls will reveal a dead one on its
|
||||
own. That is why this arm is worth looking up in the panel below when
|
||||
a report preference never seems to arrive.
|
||||
</p>
|
||||
</div>
|
||||
<div class="field">
|
||||
|
||||
Reference in New Issue
Block a user