fix(retrieval): the rules never-opened warning stops sending the reader to a review tool that refuses rules (#4798)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m7s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 26s

#4798 "The rules corpus's surfaced_never_pulled warning sends the reader to
menus_to_review, which can only review auto_inject". #4772 gave the warning
one remedy for both corpora: "judge a sample with menus_to_review". That is
right for notes, where a line carries its passage. For rules it is a dead
end: retrieval_review.REVIEWABLE holds only auto_inject, so the tool
refuses every rule arm.

- The reading is now per corpus (_NEVER_PULLED_READING). For rules, the
  text says:
  - the count includes rules that only arrived in a listing;
  - a rule can rightly be set aside on its trigger alone;
  - no judged sample exists for the rule arms;
  - a rule set aside again and again is a trigger to fix (update_rule
    when_to_apply, then what_might_apply), not a floor.
- The tool docstring says the same.
- The guard is tied to REVIEWABLE, so it can fail in both directions: if
  the rules text names menus_to_review while no rule arm is reviewable, or
  if a rule arm becomes reviewable and the text still says there is no
  sample.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-10-03 23:02:32 -04:00
co-authored by Claude Opus 5.5
parent f95ffae972
commit 9909cd2450
3 changed files with 54 additions and 5 deletions
+2
View File
@@ -624,6 +624,8 @@ It is an UPPER BOUND per surface: a pull records the door it came
corpus. Not a verdict on them: a line carries its matched passage, so
an unopened record may have been unrelated, enough as shown, or already
in context. Judge a sample (`menus_to_review`) before acting on it.
For RULES there is no judged sample: a rule set aside again and again
is a trigger to fix (`update_rule(when_to_apply=...)`), not a floor.
- `read_and_unacted` — distinct rules OPENED in the window that recorded
no outcome, against the ones that did. The failure milestone 419 was
opened on, and the worse sibling of `surfaced_never_pulled` above: a
+29 -5
View File
@@ -492,6 +492,34 @@ def _num(raw: str, fallback):
return fallback
# What "surfaced and never opened" can and cannot mean, per corpus, and where
# to look next. ONE SENTENCE PER CORPUS because the two lines are built
# differently (#4798): a note line carries its matched passage, so an unopened
# note may have done its job, and the judged sample is the instrument that can
# tell. A rule line carries only its trigger and asks to be opened, and no
# judged sample covers the rule arms (`retrieval_review.REVIEWABLE`), so
# sending that reader to `menus_to_review` sent them to a refusal.
_NEVER_PULLED_READING = {
"notes": (
"That is not a verdict on them: a line carries its matched passage, "
"so an unopened record may have been unrelated, enough as shown, or "
"already in context. Before acting on it — a title, a floor, a "
"budget — judge a sample with `menus_to_review` and read the "
"`judged` block."
),
"rules": (
"Not a verdict either, and read differently from notes: the count "
"includes rules that only arrived in a listing (the project "
"handshake, a planning read), and a rule line shows its trigger, so a "
"session can rightly set one aside on the trigger alone. There is no "
"judged sample for the rule arms. A rule set aside again and again "
"where it does not apply is a trigger to fix, not a floor to move — "
"`update_rule(when_to_apply=...)`, then `what_might_apply` with the "
"moment's own words to see where it ranks."
),
}
def _warn(code, detail, source=None, **numbers) -> dict:
"""One finding, carrying the numbers that produced it.
@@ -767,11 +795,7 @@ def _compute_warnings(sources: dict, usage: dict, rule_usage: dict,
out.append(_warn(
"surfaced_never_pulled",
f"{never} of {shown} distinct {label} were surfaced in this "
f"window and never opened. That is not a verdict on them: a "
f"line carries its matched passage, so an unopened record may "
f"have been unrelated, enough as shown, or already in context. "
f"Before acting on it — a title, a floor, a budget — judge a "
f"sample with `menus_to_review` and read the `judged` block.",
f"window and never opened. {_NEVER_PULLED_READING[label]}",
source=None, corpus=label,
surfaced=int(shown), pulled=int(pulled or 0), never_pulled=never,
))