feat(retrieval): a review pass judges whether injected lines related — menus_to_review, judge_menu and a judged readout (#4772)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 44s

An open rate cannot say whether a menu line related: every line carries its
matched passage (#4364), so "not opened" covers unrelated, enough as shown,
and already in context. #4772 "Injected notes are never judged".

- retrieval_judgments (0115): a reviewer verdict per line of a logged call,
  on_point / adjacent / unrelated, with its reason, rank, budget side and
  whether the agent opened it within the hour.
- menus_to_review re-runs a random sample of unjudged auto_inject calls with
  the arm's own parameters, past its budget, passage on every line.
  judge_menu records verdicts, re-deriving rank from a fresh re-run.
- retrieval_telemetry gains a judged block (by rank, within/beyond budget,
  on_point_unopened). surfaced_never_pulled stops blaming titles.
- missed-retrieval guidance names the review before a budget move.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-10-03 15:34:14 -04:00
co-authored by Claude Opus 5.5
parent 0a1bb68808
commit 6598c7fa85
13 changed files with 879 additions and 11 deletions
+7
View File
@@ -133,6 +133,10 @@ _READ_ONLY_TOOLS = frozenset({
# the prefixes the completeness test derives from, so nothing would have
# prompted this decision.
"notes_due_for_verification",
# The review pass's sample (#4772). Re-runs logged queries and writes no
# row of any kind — not even telemetry, since nothing is put in front of a
# working session. Spelled out for retrieval_telemetry's reason.
"menus_to_review",
# Its rule twin and a rule's edit history (milestones 312 and 323). Both
# pure reads, and both sat unlisted — so a read key was refused them — for
# the same reason: no read prefix, back when the completeness test only
@@ -198,6 +202,9 @@ _WRITE_TOOLS = frozenset({
# reads, and it appends the reason to the audit trail (#4102).
"tune_retrieval",
"migrate_retrieval_floor",
# A reviewer's verdicts on logged menu lines (#4772) — rows carrying free
# prose the agent authored, `rule_outcome`'s reason for being a write.
"judge_menu",
# trash
"restore", "purge_trash",
})
+2 -1
View File
@@ -6,7 +6,7 @@ from `mcp.server.build_mcp_server`.
"""
from scribe.mcp.tools import (
design_systems, lessons, milestones, notes, processes, projects, recent, repos,
retrieval_tuning,
retrieval_review, retrieval_tuning,
wide_net,
rulebooks, search, shapes, snippets, systems, tags, tasks, trash,
)
@@ -16,6 +16,7 @@ def register_all(mcp) -> None:
"""Register every tool module's tools on the given MCPServer instance."""
search.register(mcp)
retrieval_tuning.register(mcp)
retrieval_review.register(mcp)
wide_net.register(mcp)
notes.register(mcp)
tasks.register(mcp)
+74
View File
@@ -0,0 +1,74 @@
"""Judging what a retrieval arm offered (#4772).
`retrieval_telemetry` says what the reader did with each line — opened it or
not. These two say whether the line RELATED, which is a reading only someone
looking at the query beside the line can do.
"""
from scribe.mcp._context import current_user_id
from scribe.services import retrieval_review as review_svc
async def menus_to_review(
source: str = "auto_inject", n: int = 5, days: int = 14, depth: int = 6,
) -> dict:
"""A sample of logged injection menus, re-run so you can judge each line.
Reach for this before moving a budget or a floor, and whenever an open rate
is about to be read as a verdict on relevance. It is not one: every menu
line carries the passage that matched, so a record left unopened may have
been unrelated, enough as shown, or already in context, and the usage
counters cannot tell those apart. A judged sample can.
Each menu is one logged call: the query the arm searched (enriched, as it
was asked), its `budget`, and `lines` — the query re-run with the arm's own
parameters to `depth`, each line with `rank`, `score`, `kind`, `name`, the
`passage` that matched, `within_budget` (was it inside the cut) and
`opened_after` (did your agent open it within an hour). Lines past the
budget are there on purpose: they are what a larger budget would add.
`missing` names ids the call logged that the re-run no longer finds — the
corpus has moved since. Judge what is in front of you.
JUDGE FROM THE LINE; DO NOT OPEN THE RECORD. Opening counts as a pull and
would credit the arm with an open it never earned. Then record verdicts
with `judge_menu`. `how_to_judge` in the response repeats the vocabulary.
Args:
source: the arm to review. Only arms whose search can be re-run
exactly are offered; today that is `auto_inject`.
n: how many calls to sample, 1–20. Random among the calls you have not
judged, so a review does not describe one afternoon.
days: how far back to sample from.
depth: how many candidates to re-run per call, 1–10. Past the budget,
so the cut itself can be judged.
"""
return await review_svc.menus_to_review(
current_user_id(), source=source, n=n, days=days, depth=depth,
)
async def judge_menu(log_id: int, verdicts: list[dict]) -> dict:
"""Record your verdict on lines of one menu from `menus_to_review`.
`verdicts` is a list of `{record_id, verdict, reason}`:
"on_point" — someone in the query's situation should read this.
"adjacent" — the same area, but it would not change what they do.
"unrelated" — close in words only.
`reason` is REQUIRED for every verdict: what in the query and the passage
decided it. A verdict nobody can re-read is a threshold, and this exists to
replace one. Rank, score and which side of the budget the line sat are
taken from a fresh re-run, not from you. Judging a line again replaces your
earlier verdict on it.
The verdicts are read back as `retrieval_telemetry`'s `judged` block.
"""
return await review_svc.judge_menu(
current_user_id(), log_id=log_id, verdicts=verdicts,
)
def register(mcp) -> None:
mcp.tool(name="menus_to_review")(menus_to_review)
mcp.tool(name="judge_menu")(judge_menu)
+17 -3
View File
@@ -380,7 +380,7 @@ async def retrieval_telemetry(
— so `near_miss_samples=5` and opening the ids it returns is the step that
separates a real miss from a bar doing its job.
Four readouts, from the four tables built for them:
Five readouts, from the five tables built for them:
`sources` — per retrieval surface (`auto_inject`, `write_path`,
`mcp_search`, …), from `retrieval_logs`: `calls`, `zero_result_calls`,
@@ -539,6 +539,19 @@ It is an UPPER BOUND per surface: a pull records the door it came
a ratio of opens would read near zero on an arm that is working.
`system_usage_failed: true` means its read broke and the zeros mean nothing.
`judged` — whether the lines RELATED (#4772), from `retrieval_judgments`:
per source, a reviewer's verdicts (`on_point` / `adjacent` / `unrelated`)
on a sample of logged menus, `within_budget` against `beyond_budget` and
`by_rank`. Every usage block above counts what the reader DID; a line
carries its matched passage, so "not opened" is no verdict on relevance,
and this is the block that holds one. `on_point_unopened` counts lines
that related and were left closed. Read it before moving a budget: if the
ranks past the cut are mostly `on_point`, the budget is costing hits; if
they are mostly `unrelated`, it is doing its job. Empty until someone
reviews — `menus_to_review` and `judge_menu` fill it. A sample is small by
nature, so the counts are printed, not rates. `judged_failed: true` means
the read broke.
EVERY COUNTER BLOCK CARRIES ITS OWN COVERAGE — `complete_from` and
`covers_window`. `complete_from` is when the number became trustworthy:
for one source, its first recorded row; for a section that sums several,
@@ -608,8 +621,9 @@ It is an UPPER BOUND per surface: a pull records the door it came
- `no_duration` — rows written without timings. A logging gap, not a slow
arm, and it devalues every other number from that source.
- `surfaced_never_pulled` — distinct records shown and never opened, per
corpus. Read their titles before touching a threshold: a record nobody
opens is usually one whose title does not say when it matters.
corpus. Not a verdict on them: a line carries its matched passage, so
an unopened record may have been unrelated, enough as shown, or already
in context. Judge a sample (`menus_to_review`) before acting on it.
- `read_and_unacted` — distinct rules OPENED in the window that recorded
no outcome, against the ones that did. The failure milestone 419 was
opened on, and the worse sibling of `surfaced_never_pulled` above: a