feat(retrieval): a review pass judges whether injected lines related — menus_to_review, judge_menu and a judged readout (#4772)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 44s

An open rate cannot say whether a menu line related: every line carries its
matched passage (#4364), so "not opened" covers unrelated, enough as shown,
and already in context. #4772 "Injected notes are never judged".

- retrieval_judgments (0115): a reviewer verdict per line of a logged call,
  on_point / adjacent / unrelated, with its reason, rank, budget side and
  whether the agent opened it within the hour.
- menus_to_review re-runs a random sample of unjudged auto_inject calls with
  the arm's own parameters, past its budget, passage on every line.
  judge_menu records verdicts, re-deriving rank from a fresh re-run.
- retrieval_telemetry gains a judged block (by rank, within/beyond budget,
  on_point_unopened). surfaced_never_pulled stops blaming titles.
- missed-retrieval guidance names the review before a budget move.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-10-03 15:34:14 -04:00
co-authored by Claude Opus 5.5
parent 0a1bb68808
commit 6598c7fa85
13 changed files with 879 additions and 11 deletions
+7
View File
@@ -133,6 +133,10 @@ _READ_ONLY_TOOLS = frozenset({
# the prefixes the completeness test derives from, so nothing would have
# prompted this decision.
"notes_due_for_verification",
# The review pass's sample (#4772). Re-runs logged queries and writes no
# row of any kind — not even telemetry, since nothing is put in front of a
# working session. Spelled out for retrieval_telemetry's reason.
"menus_to_review",
# Its rule twin and a rule's edit history (milestones 312 and 323). Both
# pure reads, and both sat unlisted — so a read key was refused them — for
# the same reason: no read prefix, back when the completeness test only
@@ -198,6 +202,9 @@ _WRITE_TOOLS = frozenset({
# reads, and it appends the reason to the audit trail (#4102).
"tune_retrieval",
"migrate_retrieval_floor",
# A reviewer's verdicts on logged menu lines (#4772) — rows carrying free
# prose the agent authored, `rule_outcome`'s reason for being a write.
"judge_menu",
# trash
"restore", "purge_trash",
})