feat(retrieval): a review pass judges whether injected lines related — menus_to_review, judge_menu and a judged readout (#4772)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 44s
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 44s
An open rate cannot say whether a menu line related: every line carries its matched passage (#4364), so "not opened" covers unrelated, enough as shown, and already in context. #4772 "Injected notes are never judged". - retrieval_judgments (0115): a reviewer verdict per line of a logged call, on_point / adjacent / unrelated, with its reason, rank, budget side and whether the agent opened it within the hour. - menus_to_review re-runs a random sample of unjudged auto_inject calls with the arm's own parameters, past its budget, passage on every line. judge_menu records verdicts, re-deriving rank from a fresh re-run. - retrieval_telemetry gains a judged block (by rank, within/beyond budget, on_point_unopened). surfaced_never_pulled stops blaming titles. - missed-retrieval guidance names the review before a budget move. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -133,6 +133,10 @@ _READ_ONLY_TOOLS = frozenset({
|
||||
# the prefixes the completeness test derives from, so nothing would have
|
||||
# prompted this decision.
|
||||
"notes_due_for_verification",
|
||||
# The review pass's sample (#4772). Re-runs logged queries and writes no
|
||||
# row of any kind — not even telemetry, since nothing is put in front of a
|
||||
# working session. Spelled out for retrieval_telemetry's reason.
|
||||
"menus_to_review",
|
||||
# Its rule twin and a rule's edit history (milestones 312 and 323). Both
|
||||
# pure reads, and both sat unlisted — so a read key was refused them — for
|
||||
# the same reason: no read prefix, back when the completeness test only
|
||||
@@ -198,6 +202,9 @@ _WRITE_TOOLS = frozenset({
|
||||
# reads, and it appends the reason to the audit trail (#4102).
|
||||
"tune_retrieval",
|
||||
"migrate_retrieval_floor",
|
||||
# A reviewer's verdicts on logged menu lines (#4772) — rows carrying free
|
||||
# prose the agent authored, `rule_outcome`'s reason for being a write.
|
||||
"judge_menu",
|
||||
# trash
|
||||
"restore", "purge_trash",
|
||||
})
|
||||
|
||||
@@ -6,7 +6,7 @@ from `mcp.server.build_mcp_server`.
|
||||
"""
|
||||
from scribe.mcp.tools import (
|
||||
design_systems, lessons, milestones, notes, processes, projects, recent, repos,
|
||||
retrieval_tuning,
|
||||
retrieval_review, retrieval_tuning,
|
||||
wide_net,
|
||||
rulebooks, search, shapes, snippets, systems, tags, tasks, trash,
|
||||
)
|
||||
@@ -16,6 +16,7 @@ def register_all(mcp) -> None:
|
||||
"""Register every tool module's tools on the given MCPServer instance."""
|
||||
search.register(mcp)
|
||||
retrieval_tuning.register(mcp)
|
||||
retrieval_review.register(mcp)
|
||||
wide_net.register(mcp)
|
||||
notes.register(mcp)
|
||||
tasks.register(mcp)
|
||||
|
||||
@@ -0,0 +1,74 @@
|
||||
"""Judging what a retrieval arm offered (#4772).
|
||||
|
||||
`retrieval_telemetry` says what the reader did with each line — opened it or
|
||||
not. These two say whether the line RELATED, which is a reading only someone
|
||||
looking at the query beside the line can do.
|
||||
"""
|
||||
from scribe.mcp._context import current_user_id
|
||||
from scribe.services import retrieval_review as review_svc
|
||||
|
||||
|
||||
async def menus_to_review(
|
||||
source: str = "auto_inject", n: int = 5, days: int = 14, depth: int = 6,
|
||||
) -> dict:
|
||||
"""A sample of logged injection menus, re-run so you can judge each line.
|
||||
|
||||
Reach for this before moving a budget or a floor, and whenever an open rate
|
||||
is about to be read as a verdict on relevance. It is not one: every menu
|
||||
line carries the passage that matched, so a record left unopened may have
|
||||
been unrelated, enough as shown, or already in context, and the usage
|
||||
counters cannot tell those apart. A judged sample can.
|
||||
|
||||
Each menu is one logged call: the query the arm searched (enriched, as it
|
||||
was asked), its `budget`, and `lines` — the query re-run with the arm's own
|
||||
parameters to `depth`, each line with `rank`, `score`, `kind`, `name`, the
|
||||
`passage` that matched, `within_budget` (was it inside the cut) and
|
||||
`opened_after` (did your agent open it within an hour). Lines past the
|
||||
budget are there on purpose: they are what a larger budget would add.
|
||||
|
||||
`missing` names ids the call logged that the re-run no longer finds — the
|
||||
corpus has moved since. Judge what is in front of you.
|
||||
|
||||
JUDGE FROM THE LINE; DO NOT OPEN THE RECORD. Opening counts as a pull and
|
||||
would credit the arm with an open it never earned. Then record verdicts
|
||||
with `judge_menu`. `how_to_judge` in the response repeats the vocabulary.
|
||||
|
||||
Args:
|
||||
source: the arm to review. Only arms whose search can be re-run
|
||||
exactly are offered; today that is `auto_inject`.
|
||||
n: how many calls to sample, 1–20. Random among the calls you have not
|
||||
judged, so a review does not describe one afternoon.
|
||||
days: how far back to sample from.
|
||||
depth: how many candidates to re-run per call, 1–10. Past the budget,
|
||||
so the cut itself can be judged.
|
||||
"""
|
||||
return await review_svc.menus_to_review(
|
||||
current_user_id(), source=source, n=n, days=days, depth=depth,
|
||||
)
|
||||
|
||||
|
||||
async def judge_menu(log_id: int, verdicts: list[dict]) -> dict:
|
||||
"""Record your verdict on lines of one menu from `menus_to_review`.
|
||||
|
||||
`verdicts` is a list of `{record_id, verdict, reason}`:
|
||||
|
||||
"on_point" — someone in the query's situation should read this.
|
||||
"adjacent" — the same area, but it would not change what they do.
|
||||
"unrelated" — close in words only.
|
||||
|
||||
`reason` is REQUIRED for every verdict: what in the query and the passage
|
||||
decided it. A verdict nobody can re-read is a threshold, and this exists to
|
||||
replace one. Rank, score and which side of the budget the line sat are
|
||||
taken from a fresh re-run, not from you. Judging a line again replaces your
|
||||
earlier verdict on it.
|
||||
|
||||
The verdicts are read back as `retrieval_telemetry`'s `judged` block.
|
||||
"""
|
||||
return await review_svc.judge_menu(
|
||||
current_user_id(), log_id=log_id, verdicts=verdicts,
|
||||
)
|
||||
|
||||
|
||||
def register(mcp) -> None:
|
||||
mcp.tool(name="menus_to_review")(menus_to_review)
|
||||
mcp.tool(name="judge_menu")(judge_menu)
|
||||
@@ -380,7 +380,7 @@ async def retrieval_telemetry(
|
||||
— so `near_miss_samples=5` and opening the ids it returns is the step that
|
||||
separates a real miss from a bar doing its job.
|
||||
|
||||
Four readouts, from the four tables built for them:
|
||||
Five readouts, from the five tables built for them:
|
||||
|
||||
`sources` — per retrieval surface (`auto_inject`, `write_path`,
|
||||
`mcp_search`, …), from `retrieval_logs`: `calls`, `zero_result_calls`,
|
||||
@@ -539,6 +539,19 @@ It is an UPPER BOUND per surface: a pull records the door it came
|
||||
a ratio of opens would read near zero on an arm that is working.
|
||||
`system_usage_failed: true` means its read broke and the zeros mean nothing.
|
||||
|
||||
`judged` — whether the lines RELATED (#4772), from `retrieval_judgments`:
|
||||
per source, a reviewer's verdicts (`on_point` / `adjacent` / `unrelated`)
|
||||
on a sample of logged menus, `within_budget` against `beyond_budget` and
|
||||
`by_rank`. Every usage block above counts what the reader DID; a line
|
||||
carries its matched passage, so "not opened" is no verdict on relevance,
|
||||
and this is the block that holds one. `on_point_unopened` counts lines
|
||||
that related and were left closed. Read it before moving a budget: if the
|
||||
ranks past the cut are mostly `on_point`, the budget is costing hits; if
|
||||
they are mostly `unrelated`, it is doing its job. Empty until someone
|
||||
reviews — `menus_to_review` and `judge_menu` fill it. A sample is small by
|
||||
nature, so the counts are printed, not rates. `judged_failed: true` means
|
||||
the read broke.
|
||||
|
||||
EVERY COUNTER BLOCK CARRIES ITS OWN COVERAGE — `complete_from` and
|
||||
`covers_window`. `complete_from` is when the number became trustworthy:
|
||||
for one source, its first recorded row; for a section that sums several,
|
||||
@@ -608,8 +621,9 @@ It is an UPPER BOUND per surface: a pull records the door it came
|
||||
- `no_duration` — rows written without timings. A logging gap, not a slow
|
||||
arm, and it devalues every other number from that source.
|
||||
- `surfaced_never_pulled` — distinct records shown and never opened, per
|
||||
corpus. Read their titles before touching a threshold: a record nobody
|
||||
opens is usually one whose title does not say when it matters.
|
||||
corpus. Not a verdict on them: a line carries its matched passage, so
|
||||
an unopened record may have been unrelated, enough as shown, or already
|
||||
in context. Judge a sample (`menus_to_review`) before acting on it.
|
||||
- `read_and_unacted` — distinct rules OPENED in the window that recorded
|
||||
no outcome, against the ones that did. The failure milestone 419 was
|
||||
opened on, and the worse sibling of `surfaced_never_pulled` above: a
|
||||
|
||||
Reference in New Issue
Block a user