CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 44s
An open rate cannot say whether a menu line related: every line carries its matched passage (#4364), so "not opened" covers unrelated, enough as shown, and already in context. #4772 "Injected notes are never judged". - retrieval_judgments (0115): a reviewer verdict per line of a logged call, on_point / adjacent / unrelated, with its reason, rank, budget side and whether the agent opened it within the hour. - menus_to_review re-runs a random sample of unjudged auto_inject calls with the arm's own parameters, past its budget, passage on every line. judge_menu records verdicts, re-deriving rank from a fresh re-run. - retrieval_telemetry gains a judged block (by rank, within/beyond budget, on_point_unopened). surfaced_never_pulled stops blaming titles. - missed-retrieval guidance names the review before a budget move. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
56 lines
3.4 KiB
Markdown
56 lines
3.4 KiB
Markdown
# When a record doesn't reach the moment it should
|
|
|
|
Part of the using-scribe skill. Read it when a rule should have governed a
|
|
moment and never arrived, when one arrives on every turn and never applies,
|
|
or before touching a retrieval floor.
|
|
|
|
Retrieval misjudging is ordinary, and it is fixable — but only by whoever
|
|
notices. **Either direction counts:** a rule that should have governed a moment
|
|
and never arrived, and a rule that arrives on every turn and never applies. So
|
|
does **either noticer**: the operator saying *"that should have fired"*, and you
|
|
noticing it yourself — you reached for a rule nobody offered you, or you were
|
|
handed the same rule five times and set it aside five times.
|
|
|
|
**Take it to the record first and the dial second.** A rule's `when_to_apply`
|
|
IS the text its similarity score is computed against, so when a rule misses a
|
|
moment it governs, the overwhelmingly likely cause is that its trigger does not
|
|
describe that moment in the words a session actually produces. Rewording one
|
|
trigger changes one rule's reach. Moving a floor changes what every record on
|
|
that surface does, and a floor cannot tell a badly-worded trigger from a
|
|
genuinely distant record — so one lowered to rescue a single rule admits
|
|
everything else that was sitting in the same band.
|
|
|
|
1. **Read the refused records.** `retrieval_telemetry(days=N,
|
|
near_miss_samples=5)` names by id what each surface refused and by how much.
|
|
Open them with `get_rule` / `get_note`. This is the step that carries the
|
|
answer: the statistic says a record was close, and only the record says
|
|
whether it was *right*.
|
|
2. **Fix the trigger.** `update_rule(when_to_apply=...)`, written as the
|
|
symptom — what the session was doing or saying at the moment it needed this
|
|
rule — not the situation the rule belongs to. Then check that it worked:
|
|
`what_might_apply("the moment, in the operator's own words")` and read where
|
|
the rule now ranks. The change is measurable, so measure it, and say the
|
|
before and after when you report it.
|
|
3. **Then consider the dial.** `retrieval_surfaces` shows what is in force per
|
|
arm and whether the number is still calibrated; `tune_retrieval` moves it.
|
|
`reason` is required and has to say what you read, because it is what lets
|
|
the operator disagree with a number they did not choose.
|
|
|
|
**A budget is judged by what sits past it, and an open rate cannot judge
|
|
it.** Every menu line carries its matched passage, so a record left
|
|
unopened may have been unrelated, enough as shown, or already in context.
|
|
`menus_to_review` re-runs a sample of logged menus to a depth past the
|
|
budget; judge each line from what is shown with `judge_menu`, then read
|
|
`retrieval_telemetry`'s `judged` block. If the lines past the cut are
|
|
mostly `on_point`, the budget is costing hits; mostly `unrelated`, it is
|
|
doing its job.
|
|
|
|
Reaching for `tune_retrieval` before opening a single record is the wrong move,
|
|
and it is the one that feels efficient. Worked example, measured on this
|
|
install: a rule granting a routine push scored 0.6515 and ranked 5th for the
|
|
moment it governed, behind three rules that *restrained* the same act. Every
|
|
percentile said "lower the floor" — and lowering it would have delivered those
|
|
three restraints and still not the rule. Rewriting the trigger to lead with the
|
|
symptom moved the same rule to 1st at 0.7130, ahead of all three. Only then was
|
|
the floor worth touching.
|