feat(retrieval): a review pass judges whether injected lines related — menus_to_review, judge_menu and a judged readout (#4772)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 44s
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 44s
An open rate cannot say whether a menu line related: every line carries its matched passage (#4364), so "not opened" covers unrelated, enough as shown, and already in context. #4772 "Injected notes are never judged". - retrieval_judgments (0115): a reviewer verdict per line of a logged call, on_point / adjacent / unrelated, with its reason, rank, budget side and whether the agent opened it within the hour. - menus_to_review re-runs a random sample of unjudged auto_inject calls with the arm's own parameters, past its budget, passage on every line. judge_menu records verdicts, re-deriving rank from a fresh re-run. - retrieval_telemetry gains a judged block (by rank, within/beyond budget, on_point_unopened). surfaced_never_pulled stops blaming titles. - missed-retrieval guidance names the review before a budget move. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -1,7 +1,7 @@
|
||||
{
|
||||
"name": "scribe",
|
||||
"description": "Scribe for Claude Code: connects the scribe MCP server, adds the hooks that deliver live project state and relevant records at the right moment, ships the shared client-neutral Scribe skills (using-scribe, writing-plans, reporting-back, systematic-debugging, verification, brainstorming, reusing-code, shape-accounting), and syncs your saved Scribe Processes as skills (/scribe:sync).",
|
||||
"version": "2026.10.03.0308",
|
||||
"version": "2026.10.03.1934",
|
||||
"author": {
|
||||
"name": "Bryan Van Deusen"
|
||||
},
|
||||
|
||||
@@ -36,6 +36,15 @@ everything else that was sitting in the same band.
|
||||
`reason` is required and has to say what you read, because it is what lets
|
||||
the operator disagree with a number they did not choose.
|
||||
|
||||
**A budget is judged by what sits past it, and an open rate cannot judge
|
||||
it.** Every menu line carries its matched passage, so a record left
|
||||
unopened may have been unrelated, enough as shown, or already in context.
|
||||
`menus_to_review` re-runs a sample of logged menus to a depth past the
|
||||
budget; judge each line from what is shown with `judge_menu`, then read
|
||||
`retrieval_telemetry`'s `judged` block. If the lines past the cut are
|
||||
mostly `on_point`, the budget is costing hits; mostly `unrelated`, it is
|
||||
doing its job.
|
||||
|
||||
Reaching for `tune_retrieval` before opening a single record is the wrong move,
|
||||
and it is the one that feels efficient. Worked example, measured on this
|
||||
install: a rule granting a routine push scored 0.6515 and ranked 5th for the
|
||||
|
||||
Reference in New Issue
Block a user