docs(retrieval): stop asking the operator to diagnose retrieval (#4102)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m30s
CI & Build / Build & push image (push) Successful in 33s

The last item in this step's done-when: no wording anywhere asks an
operator to tune for correctness.

Four Settings hints told them to do exactly that — "raise it if rules
keep arriving unread", "lower it if a git push arrives with nothing",
"lower this if genuine duplicates go unnoticed". Every one of those
asks the operator to diagnose a ranker from symptoms, which is the job
the model now does from the records: the telemetry says what each bar
refused, and reading those records is what separates a real miss from
a bar doing its job. The hints keep the explanation of WHAT each number
is — that is worth reading — and drop the homework.

`plugin_context.py` said the defaults "are meant to be tuned from
retrieval_logs once data accrues", which was true and had no owner.
It now names who does it and with what.

`retrieval_telemetry`'s docstring gained the warning that belongs
beside it: this readout has been measured pointing the wrong way, so
the ids it returns are the point, not its percentiles.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
2026-09-17 11:38:16 -04:00
co-authored by Claude Opus 5
parent 6240652dce
commit c8bfa6947c
3 changed files with 25 additions and 18 deletions
+7 -5
View File
@@ -209,11 +209,13 @@ async def retrieval_telemetry(
) -> dict:
"""What the retrieval telemetry says about YOUR surfaces, over a window.
The read half of the loop the ranker's thresholds are meant to be tuned
from (#2975). Reach for it before changing a similarity threshold, a top-k,
or deciding whether a reranker is worth building — the alternative is
hand-probing the live instance, which is how the last such decision had to
be made.
The read half of the tuning loop, whose write half is `tune_retrieval`
(#2975, #4102). Reach for it before moving any floor or budget, and read
the records it names rather than its percentiles alone: this readout has
been measured pointing the WRONG WAY — 69 consecutive declines where every
percentile said "lower the bar" and the refused record was a false positive
— so `near_miss_samples=5` and opening the ids it returns is the step that
separates a real miss from a bar doing its job.
Three readouts, from the three tables built for them: