fix(rules): a shortened rule line must not decide what it says about holding (#3851)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 8s
CI & Build / integration (push) Successful in 40s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m24s
CI & Build / Build & push image (push) Successful in 26s

CI run 6485 was red. Six failures, three causes, and only one of them was a
stale test.

THE REAL DEFECT. The compact branch dropped the `seen` TAIL along with the
trigger, so a rule the session had already been told rendered exactly like
one it had not. #3750's whole argument is that those are different claims —
a repeat is rendered precisely because the session may no longer HOLD what
it was told — and the tail is the entire difference a reader can act on.
test_a_rule_the_session_already_holds_is_referenced_not_re_offered caught it
within one commit, which is that guard working as intended.

Fixed by keeping the tail and dropping only the trigger, which is both the
cheaper and the safer cut: a trigger runs 300-400 characters after #3855, a
tail about 100. Re-measured on the real renderer — top-full-plus-references
is ~299 tokens against ~568 for five full lines, so about 2x the old single
line rather than the 1.4x claimed before, for four more rules and no lost
information. The comments carrying the old figure are corrected rather than
left to read as a decision nobody made.

THE FIXTURE THAT STRADDLED THE BAND. `_THREE_HITS` spanned 0.81-0.74 against
a 0.05 band, so the act arms dropped its lowest hit and four cases of
test_both_recorders_report_the_same_rules_for_one_call failed reporting a
count mismatch — under a message blaming the exclusion filter. A guard
pointing confidently at the wrong subsystem costs more than no guard,
because it is believed. Scores retightened to 0.81/0.80/0.79 and the
precondition is now asserted by a named test, so a future band change is
told where the problem is instead of through four confusing failures.

THE STALE CONSTANT GUARD. test_the_rule_arm_asks_for_one_rule_not_two pinned
RULEHINT_LIMIT == 1 — a real decision, correctly guarded, for a world with a
resident set. Rewritten to pin what replaced it, as relationships rather
than values (rule 115): the arm can return several, and rules are narrowed
HARDER than the notes menu because they measured flatter, not sharper.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cPyzNnegXHr5iRMzzy5KJ
This commit is contained in:
2026-09-11 14:07:32 -04:00
co-authored by Claude Opus 5
parent 10343a6019
commit 40189147d2
4 changed files with 121 additions and 27 deletions
+23 -5
View File
@@ -499,13 +499,31 @@ def test_the_rule_bar_defaults_above_the_code_bar():
assert pc.RULEHINT_DEFAULT_THRESHOLD > pc.WRITEPATH_DEFAULT_THRESHOLD
def test_the_rule_arm_asks_for_one_rule_not_two():
"""With a corpus this small, top-k does as much damage as the threshold:
k=2 over a few dozen candidates means the second line is almost always the
second-best noise, carrying the same confident framing as the first."""
def test_the_rule_arm_asks_for_a_set_and_lets_the_band_narrow_it():
"""WAS `..._asks_for_one_rule_not_two`, pinning `RULEHINT_LIMIT == 1`.
That guarded a real decision: with retrieval SUPPLEMENTING a 33-rule
resident set, k=2 over a few dozen candidates made the second line the
second-best noise wearing the first line's confident framing. Milestone
394 removes residency, so this arm becomes the whole delivery and a
`git push` governed by four rules cannot be served by one slot (#3851).
What replaces it is not simply a bigger k — that is the thing the old
test was right to fear, because a fixed k fills its slots whether or not
anything deserves them. The cap is a ceiling and the BAND is the control,
so both must exist for the arm to be shaped as intended.
Pinned as relationships rather than values, like the threshold test above
and for the same rule-115 reason: the band is a tuning number measured
against one corpus, and a test asserting 0.05 would fail on every retune
while proving nothing. What must not silently invert is that the arm can
return several, and that rules — which measured FLATTER than notes, not
sharper — are narrowed harder than the notes menu is.
"""
from scribe.services import plugin_context as pc
assert pc.RULEHINT_LIMIT == 1
assert pc.RULEHINT_LIMIT > 1
assert 0 < pc._RULEHINT_BAND < pc._AUTOINJECT_BAND
# --- the minimum-substance floor on the semantic arm (#2223) ------------------