Files
FabledScribe/tests/test_retrieval_specs.py
T
bvandeusenandClaude Opus 5.5 f68b73f922
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m10s
CI & Build / Python tests (push) Successful in 1m54s
CI & Build / Build & push image (push) Successful in 31s
feat(500): retire the fixed-question preference arm - completion-report preferences ride the moments
report_preference searched one constant string when a task closed and handed
matches back as update_task's reply_preferences. Milestone 500 step 3 delivers
the reply shapes and the preferences mounted beside them on the moments, so
the arm goes, whole (rule 22):

- services/reply_preferences.py, the REPORT_PREFERENCE RuleArm, its re-scorer
  and corpus entries, and the REPORTPREF constants
- update_task's reply_preferences block and its cue; the docstring now points
  at moment_rules
- the fixed_query field and the fixed_query_never_clears warning: this arm was
  the only one that set it, so the concept and the cannot_decline exemption
  go with it
- the two Settings fields for its floor and budget
- migration 0121 deletes its rows (logs, judgments, rule-usage events, tuning
  history, settings keys), operator-approved 2026-10-09. Without it the
  readout would call the source unregistered and its surfacings would count
  as ambient. The unindexed rule_usage delete is bounded by created_at.

#5496 (step 4 of milestone 500).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-09 15:51:04 -04:00

71 lines
3.1 KiB
Python

"""The specs are the registry (milestone 456 step 5).
Three lists used to be kept by hand beside the arms — the tuning table, the
ranked rows of the telemetry registry, and the rule-usage denominator — with
tests to make them agree. They are now read off the pipeline's specs. These
pin the derivation itself: every arm is tunable, every ranked source is
measured with the declaration its spec carries, and the rule-usage
denominator counts exactly the rule surfaces a ranker chose.
"""
from __future__ import annotations
from scribe.services import retrieval_pipeline as rp
from scribe.services.retrieval_registry import POINTS, UNBIDDEN
from scribe.services.retrieval_surfaces import SURFACES
from scribe.services.rule_usage import RANKED_SOURCES, is_ambient
def test_every_arm_is_one_tunable_surface_in_the_order_settings_lists_them():
assert [arm.source for arm in rp.TUNED_ARMS] == list(SURFACES)
for arm in rp.TUNED_ARMS:
assert arm.tuning is not None, f"{arm.source} has no floor or budget"
# The join key: the surface a tuner moves and the rows it is judged by
# must name the same arm.
assert arm.tuning.name == arm.source
assert SURFACES[arm.source] is arm.tuning
def test_every_ranked_source_is_measured_as_its_spec_declares():
sources = [spec.source for spec in rp.RANKED]
assert len(sources) == len(set(sources)), f"a source is declared twice: {sources}"
for spec in rp.RANKED:
d = spec.declared
assert d is not None and d.what.strip(), f"{spec.source} declares nothing"
point = POINTS[spec.source]
assert point.kind == UNBIDDEN
assert point.what == d.what
# A quiet source must say why, and only a quiet source may (#2475).
assert point.expects_traffic == (not d.quiet_because)
assert point.quiet_because == d.quiet_because
def test_the_slots_are_measured_but_never_tuned():
for slot in (rp.PREFERENCE_SLOT, *rp.NOTE_SLOTS):
assert slot.source in POINTS
assert slot.source not in SURFACES
def test_the_rule_denominator_is_every_rule_surface_a_ranker_chose():
rule_arms = {arm.source for arm in rp.RULE_ARMS}
assert rule_arms <= set(RANKED_SOURCES)
assert set(RANKED_SOURCES) == (
rule_arms | {rp.PREFERENCE_SLOT_SOURCE, rp.VIA_LESSON_SOURCE,
rp.MOMENT_RULE_SOURCE}
)
# And a notes arm is not a RULE surface: its lines are not rules, so a
# rule pull cannot confirm them.
for arm in rp.NOTE_ARMS:
assert is_ambient(arm.source)
def test_an_arm_added_without_tuning_would_be_caught():
"""Rule 167: the first test has to be able to bite. SURFACES skips an arm
with no tuning rather than raising, so an arm added without one would
silently be untunable — and the order-equality assertion is what notices,
as this replays with one such arm appended."""
bare = rp.RuleArm("x_rule", band=False, compact_tail=False,
checkpoint=False, preference_slot=False)
arms = (*rp.TUNED_ARMS, bare)
derived = {a.tuning.name: a.tuning for a in arms if a.tuning is not None}
assert [a.source for a in arms] != list(derived)