Files
FabledScribe/plugin/skills/using-scribe/missed-retrieval.md
T
bvandeusenandClaude Opus 5.5 6598c7fa85
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 44s
feat(retrieval): a review pass judges whether injected lines related — menus_to_review, judge_menu and a judged readout (#4772)
An open rate cannot say whether a menu line related: every line carries its
matched passage (#4364), so "not opened" covers unrelated, enough as shown,
and already in context. #4772 "Injected notes are never judged".

- retrieval_judgments (0115): a reviewer verdict per line of a logged call,
  on_point / adjacent / unrelated, with its reason, rank, budget side and
  whether the agent opened it within the hour.
- menus_to_review re-runs a random sample of unjudged auto_inject calls with
  the arm's own parameters, past its budget, passage on every line.
  judge_menu records verdicts, re-deriving rank from a fresh re-run.
- retrieval_telemetry gains a judged block (by rank, within/beyond budget,
  on_point_unopened). surfaced_never_pulled stops blaming titles.
- missed-retrieval guidance names the review before a budget move.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 15:34:14 -04:00

56 lines
3.4 KiB
Markdown

# When a record doesn't reach the moment it should
Part of the using-scribe skill. Read it when a rule should have governed a
moment and never arrived, when one arrives on every turn and never applies,
or before touching a retrieval floor.
Retrieval misjudging is ordinary, and it is fixable — but only by whoever
notices. **Either direction counts:** a rule that should have governed a moment
and never arrived, and a rule that arrives on every turn and never applies. So
does **either noticer**: the operator saying *"that should have fired"*, and you
noticing it yourself — you reached for a rule nobody offered you, or you were
handed the same rule five times and set it aside five times.
**Take it to the record first and the dial second.** A rule's `when_to_apply`
IS the text its similarity score is computed against, so when a rule misses a
moment it governs, the overwhelmingly likely cause is that its trigger does not
describe that moment in the words a session actually produces. Rewording one
trigger changes one rule's reach. Moving a floor changes what every record on
that surface does, and a floor cannot tell a badly-worded trigger from a
genuinely distant record — so one lowered to rescue a single rule admits
everything else that was sitting in the same band.
1. **Read the refused records.** `retrieval_telemetry(days=N,
near_miss_samples=5)` names by id what each surface refused and by how much.
Open them with `get_rule` / `get_note`. This is the step that carries the
answer: the statistic says a record was close, and only the record says
whether it was *right*.
2. **Fix the trigger.** `update_rule(when_to_apply=...)`, written as the
symptom — what the session was doing or saying at the moment it needed this
rule — not the situation the rule belongs to. Then check that it worked:
`what_might_apply("the moment, in the operator's own words")` and read where
the rule now ranks. The change is measurable, so measure it, and say the
before and after when you report it.
3. **Then consider the dial.** `retrieval_surfaces` shows what is in force per
arm and whether the number is still calibrated; `tune_retrieval` moves it.
`reason` is required and has to say what you read, because it is what lets
the operator disagree with a number they did not choose.
**A budget is judged by what sits past it, and an open rate cannot judge
it.** Every menu line carries its matched passage, so a record left
unopened may have been unrelated, enough as shown, or already in context.
`menus_to_review` re-runs a sample of logged menus to a depth past the
budget; judge each line from what is shown with `judge_menu`, then read
`retrieval_telemetry`'s `judged` block. If the lines past the cut are
mostly `on_point`, the budget is costing hits; mostly `unrelated`, it is
doing its job.
Reaching for `tune_retrieval` before opening a single record is the wrong move,
and it is the one that feels efficient. Worked example, measured on this
install: a rule granting a routine push scored 0.6515 and ranked 5th for the
moment it governed, behind three rules that *restrained* the same act. Every
percentile said "lower the floor" — and lowering it would have delivered those
three restraints and still not the rule. Rewriting the trigger to lead with the
symptom moved the same rule to 1st at 0.7130, ahead of all three. Only then was
the floor worth touching.