The first live sample ranked records the call could never have been offered:
the session that made a call writes the decision, often quoting the message,
and that record tops the re-run. Judged, it inflates on_point exactly where
the budget is decided.
- _rerun over-fetches by POSTDATED_SLACK, drops records created after the
call before ranking, and names them in `postdated`.
- each line carries `changed_since_call` (updated_at or a work log after
the call) and `logged` (shown fresh then).
- judge_menu shares the re-run, so a post-dated record cannot be judged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
An open rate cannot say whether a menu line related: every line carries its
matched passage (#4364), so "not opened" covers unrelated, enough as shown,
and already in context. #4772 "Injected notes are never judged".
- retrieval_judgments (0115): a reviewer verdict per line of a logged call,
on_point / adjacent / unrelated, with its reason, rank, budget side and
whether the agent opened it within the hour.
- menus_to_review re-runs a random sample of unjudged auto_inject calls with
the arm's own parameters, past its budget, passage on every line.
judge_menu records verdicts, re-deriving rank from a fresh re-run.
- retrieval_telemetry gains a judged block (by rank, within/beyond budget,
on_point_unopened). surfaced_never_pulled stops blaming titles.
- missed-retrieval guidance names the review before a budget move.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>