Files
FabledScribe/tests/test_rule_hint_band.py
T
bvandeusenandClaude Opus 5 10343a6019
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / Python lint (push) Successful in 2s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Failing after 53s
CI & Build / Build & push image (push) Skipped
CI & Build / integration (push) Successful in 45s
feat(rules): an act surfaces a banded SET of rules, quieter after the first (#3851)
RULEHINT_LIMIT was 1. That was correct while retrieval SUPPLEMENTED a
33-rule resident set — one salient rule beside everything already loaded.
Milestone 394 removes residency, and then this arm is the whole delivery:
`git push origin dev` is governed by rules 1, 2, 9 and 140 simultaneously,
and each alone permits the mistake the others catch.

A cap plus a band, not a bigger cap. The old argument's real content is that
a fixed k invents lines — it fills slots whether or not anything deserves
them. A band keeps only what scored close to the top, so one clearly
relevant rule still shows one and four competing rules show four. The corpus
decides; the cap is a ceiling on the worst case, not the usual answer.

MEASURED, AND IT CORRECTED THE PREDICTION. The expectation was that rules
would rank sharply, since rule_document() shapes them like snippets and note
2485 measured snippets separating their top hit by 0.153 against 0.010-0.023
for every other kind. Three probes against real act queries say otherwise:

  `git push origin dev`      top 0.757, gap 0.022
  `docker compose up -d`     top 0.685, gap 0.016
  a bare-owner-filter query  top 0.656, gap 0.020

Dev-log territory, not snippet territory — rules arrive as a packed block,
so shaping alone did not buy separation. The band is therefore narrow: at
0.10 (the notes menu's value) every one of the top eight on the push probe
falls inside, including a CI-registry rule and another project's branch
policy. 0.05 admits about three ranks.

COST, MEASURED RATHER THAN ASSUMED. The old comment claimed a line costs
~40 tokens. It is ~143 once the trigger is rendered, and #3855 roughly
tripled trigger lengths, so five full lines are ~646 tokens before EVERY
Bash call. Hence rank decides volume: the top hit keeps the full rendering,
later hits are cited (~198 tokens total, 1.4x the old single line, for four
more rules). The old paragraph's instinct — a fourth voice at full volume is
where a reader stops reading — is answered by making later lines quieter
rather than by refusing to have them.

Band before dedup, deliberately. The band is a statement about scores;
letting the ledger reorder it would make "you were told this already" change
what counts as relevant. Same axis independence the renderer already keeps
between `kind` and `seen`, and `rule_ids` stays fresh-only (#3752) so
#3668's identity between logged results and surfacing rows survives.
`suppressed` now covers both causes and says so.

Both act arms take the same band: their score distributions are the same
shape, and only the query differs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cPyzNnegXHr5iRMzzy5KJ
2026-09-11 14:00:36 -04:00

170 lines
6.9 KiB
Python

"""An act surfaces a SET of rules, and rank decides how loudly (#3851).
WHY THIS EXISTS
`RULEHINT_LIMIT` was 1. That was right while retrieval merely SUPPLEMENTED a
33-rule resident set — one salient rule beside everything already loaded. It
stops being right the moment milestone 394 removes residency, because then
this arm is the whole delivery, and `git push origin dev` is governed by
rules 1, 2, 9 and 140 at once, each of which alone permits the mistake the
others catch.
Two instruments, and they answer different questions:
- `_rule_band` decides HOW MANY. A fixed k fills its slots whether or not
anything deserves them; a band keeps only what scored close to the top,
so one clearly-relevant rule still shows one.
- `compact` decides HOW LOUD. Measured at #3851: a full line is ~143 tokens
once the trigger is rendered, so five of them cost ~646 before every Bash
call. Top-full-plus-references costs ~198.
WHAT THIS PINS
Structure, never wording — the lines are prose and will be rewritten:
1. The band keeps the top hit and everything within `_RULEHINT_BAND`, and
drops what falls outside. Falsified below from both sides: a hit just
inside survives, a hit just outside does not.
2. Rank decides volume — the first line carries the trigger, later lines do
not, and every line names its rule's id so any of them can be pulled.
3. The band reads SCORES ONLY. A top hit the session has already seen still
anchors the band, and its score still sets the cutoff. This is the axis
independence the renderer already keeps between `kind` and `seen`, and
the regression it prevents is subtle: letting the ledger reorder the
band would make "you were told this" change what counts as relevant.
The band width itself is deliberately NOT pinned. It is a tuning value with
a comment recording the measurement behind it, and a test asserting 0.05
would fail on every future retune while proving nothing about behaviour —
so the cases below express their scores as offsets from the constant.
"""
import pytest
from scribe.services.plugin_context import (
_RULEHINT_BAND,
_rule_band,
_rule_hint_line,
)
from tests.helpers import fake_rule
_TRIGGER = "about to run git push with an earlier CI run still unread"
def _hit(score: float, rule_id: int):
return (score, fake_rule(id=rule_id, title=f"rule {rule_id}",
when_to_apply=_TRIGGER))
def test_an_empty_result_stays_empty():
"""No hits is not a crash and not a phantom line."""
assert _rule_band([]) == []
def test_the_band_keeps_a_hit_just_inside_it():
"""The whole point: a close second rule reaches the agent."""
top = 0.75
hits = [_hit(top, 1), _hit(top - _RULEHINT_BAND + 0.01, 2)]
assert [r.id for _s, r in _rule_band(hits)] == [1, 2]
def test_the_band_drops_a_hit_just_outside_it():
"""And the band must actually BIND, or it is a fixed k wearing a hat."""
top = 0.75
hits = [_hit(top, 1), _hit(top - _RULEHINT_BAND - 0.01, 2)]
assert [r.id for _s, r in _rule_band(hits)] == [1]
def test_one_clearly_better_rule_still_surfaces_alone():
"""The behaviour the old limit of 1 got right, which must not regress.
A moment with a single relevant rule shows one line, because the corpus
said so — not because a constant capped it.
"""
hits = [_hit(0.80, 1), _hit(0.55, 2), _hit(0.54, 3)]
assert [r.id for _s, r in _rule_band(hits)] == [1]
def test_a_flat_cluster_surfaces_together():
"""Measured shape of this corpus: adjacent rules sit ~0.02 apart.
Four rules governing one act is the `git push` case the step exists for,
and at the measured spacing they must arrive together rather than the
ranker picking one of four near-ties.
"""
hits = [_hit(0.757, 2), _hit(0.735, 7), _hit(0.726, 1), _hit(0.711, 9)]
assert [r.id for _s, r in _rule_band(hits)] == [2, 7, 1, 9]
def test_the_band_is_computed_from_scores_not_from_the_ledger():
"""A seen top hit still anchors the band (#3750 x #3851).
`_rule_band` never learns what the session has seen — dedup happens after
it, in the arms. Pinned here because the tempting "reorder so a fresh rule
leads" would silently change the cutoff, and the failure is invisible: the
arm would still emit lines, just the wrong set.
"""
hits = [_hit(0.80, 1), _hit(0.78, 2), _hit(0.60, 3)]
kept = _rule_band(hits)
# Independent of any `already` set, because it is not consulted.
assert [r.id for _s, r in kept] == [1, 2]
@pytest.mark.parametrize("seen", [False, True])
def test_the_leading_line_carries_the_trigger(seen):
"""Rank 0 gets the full rendering, on either tail."""
line = _rule_hint_line(
fake_rule(id=4, title="dev is home", when_to_apply=_TRIGGER),
where="to this Bash call", seen=seen, compact=False,
)
assert _TRIGGER in line
assert "get_rule(4)" in line
@pytest.mark.parametrize("seen", [False, True])
def test_a_later_line_cites_its_rule_without_quoting_the_trigger(seen):
"""Rank > 0 is a reference: identity and pointer, no trigger.
The trigger is the expensive half — it is what made a full line ~143
tokens — and the leading line has already demonstrated the shape. Both
assertions matter: dropping the trigger is the saving, and keeping
`get_rule(id)` is what makes the saving safe, because a cited rule the
reader cannot pull is just noise.
"""
line = _rule_hint_line(
fake_rule(id=4, title="dev is home", when_to_apply=_TRIGGER),
where="to this Bash call", seen=seen, compact=True,
)
assert _TRIGGER not in line
assert "dev is home" in line
assert "get_rule(4)" in line
def test_a_compact_line_is_materially_shorter_than_a_full_one():
"""The cost claim, asserted rather than asserted-in-a-comment.
Not a token count — that would pin the tokenizer. Half the characters is
the property that makes widening the arm affordable, and it is what fails
if a later edit reintroduces the trigger into the compact branch.
"""
rule = fake_rule(id=4, title="dev is home", when_to_apply=_TRIGGER)
full = _rule_hint_line(rule, where="here", seen=False, compact=False)
compact = _rule_hint_line(rule, where="here", seen=False, compact=True)
assert len(compact) * 2 < len(full)
def test_a_preference_keeps_its_noun_when_compact():
"""Force survives the shortening (#3849).
`kind` and `compact` are independent axes. A preference rendered as a
reference must still not read as a rule — the noun is the whole of the
visual difference, so losing it in the compact branch would make every
cited preference bind.
"""
line = _rule_hint_line(
fake_rule(id=5, title="pace debugging", kind="preference",
when_to_apply=_TRIGGER),
where="here", seen=False, compact=True,
)
assert "preference" in line.lower()
assert "standing rule" not in line.lower()