Files
FabledScribe/tests/test_retrieval_tuning.py
T
bvandeusenandClaude Opus 5 512d0326a0
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / integration (push) Successful in 55s
CI & Build / TypeScript typecheck (push) Successful in 1m0s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Build & push image (push) Successful in 33s
fix(telemetry): a floor that moved inside the window makes the band check a comparison of two populations (#4225)
`retrieval_telemetry(days=30)` reported, for write_path_rule:

  "the weakest tenth of what this arm returns scores 0.6984, only -0.0216
   above its floor of 0.72"

A negative distance above something. The tenth percentile of what an arm
RETURNED cannot sit below the floor that gates what it may return — not
inside one population.

MEASURED CAUSE. write_path_rule's floor was 0.68 until 2026-09-02, when
2385100 (#3318) raised the shipped default to 0.72. The window opened
2026-08-22, so six days of it are calls made under the old bar; top_score.min
for the surface is exactly 0.68, the old bar still in the sample.

AND THE CHANGE LEFT NO TRACE THE READOUT COULD SEE. retrieval_tuning_events
records dial turns — a person or a model choosing a number. It was silent
about the other way a floor moves: somebody edits floor_default and ships it.
retrieval_tuning_history returned {"events": []} and retrieval_surfaces said
last_change: {}, source: "shipped". All true, and all of it silent about a
floor that had in fact moved.

THE RAISE ANNOUNCED ITSELF. A LOWERED FLOOR WOULD NOT: the gap comes out
comfortably positive and reads as a clean bill of health on a sample that
half predates the bar being judged. Both directions are now pinned.

So the check is SUSPENDED, not softened. band_hugs_floor asks whether the
scores are piled on the bar; that needs the scores and the bar to come from
the same regime. Where they do not, the honest answer is that this sample
cannot say, plus the date after which one can — floor_moved_mid_window
replaces band_hugs_floor for that arm and never accompanies it. A reader told
a number is unavailable goes and gets one; a reader handed a qualified number
uses it.

NO MIGRATION. `actor` is Text with no CHECK precisely so a new kind of actor
is not one — the model's own comment says so, and this is the case it
anticipated. "release" joins "model" and "human". user_id is already
nullable, which is right: no user did this, a release acts on every account
that has not overridden the dial, and a row per user would both multiply and
misattribute it. Both readers now take the newest of (this user's change, the
release's).

THE FIRST SIGHTING IS A BASELINE, written with old_value NULL. Nothing moved;
the row exists so the next release has a predecessor. That null is
load-bearing: floor_moves_since asks for old_value IS NOT NULL, so a fresh
install's baseline does not silently retire the check on every new install.

UI: the tuning history rendered actor as `human ? 'you' : 'Claude'`, so a
release row would have told the operator that Claude moved a floor it never
touched — the one failure the actor column exists to prevent. Three-way now,
with an unknown value printing itself rather than guessing.

Recorded at startup, inline and awaited. What #4181 cost three hours was
concurrency — a background task racing the hook for the same pool. Sequential
creates no contention, and this is twelve single-row reads. It must finish
before serving because a readout served before the change was recorded is the
exact answer this exists to stop giving.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
2026-09-21 01:17:37 -04:00

343 lines
15 KiB
Python

"""Moving a floor, and the reason that has to come with it (#4102).
WHY THIS EXISTS
The operator handed the dial to the model: *"the user should be able to touch it
but the model should be the thing handling it 9 times out of 10."* Everything
here guards the half of that sentence people skip — the operator still has to be
able to see what was done on their behalf and disagree with it.
WHAT THIS PINS
1. **A reason is required, and "" does not count.** The column is NOT NULL,
which an empty string satisfies; the service is where the requirement is
real. This is the whole guardrail: a caller who must write down why has to
have looked, and a caller who writes down something wrong leaves the
operator a sentence to argue with. A number that moved silently leaves
nothing.
2. **The change is recorded, with what it was before.** Without `old_value` a
history cannot answer "was this always like that", which is the first
question anyone asks of a surface behaving oddly.
3. **A clamp is reported, never swallowed.** A caller that believes it set 1.4
will read the next telemetry as evidence about a bar that was never in
force — the same class of error as #3739, one layer up.
4. **An unknown surface is refused before anything is written.** Settings keys
are free-form strings, so a typo'd surface would write a key nothing reads:
a change that reports success and alters nothing.
5. **The tool teaches the procedure that works**, not the one the numbers
suggest. Asserted on the docstring because the docstring IS the contract an
agent reads, and the failure it prevents is a model tuning from a
percentile — which has been measured pointing the wrong way.
"""
from unittest.mock import AsyncMock, MagicMock, patch
import pytest
from scribe.services import retrieval_tuning as rt
from tests.helpers import make_mock_session
def _patches(floor=0.72, budget=3):
"""Patch everything `set_dial` touches except the thing under test."""
session = make_mock_session()
return session, (
patch.object(rt, "async_session", MagicMock(return_value=session)),
patch.object(rt, "set_setting", AsyncMock()),
patch.object(rt, "floor_for", AsyncMock(return_value=floor)),
patch.object(rt, "budget_for", AsyncMock(return_value=budget)),
)
@pytest.mark.asyncio
@pytest.mark.parametrize("reason", ["", " ", "noisy", "too high"])
async def test_a_change_without_a_real_reason_is_refused(reason):
"""The guardrail. "too high" is a restatement of the change, not a basis."""
session, ctx = _patches()
with ctx[0], ctx[1], ctx[2], ctx[3], pytest.raises(ValueError) as e:
await rt.set_dial(1, "prompt_rule", "floor", 0.66, reason=reason)
# The message has to name the tool that produces a real basis, or the
# caller's next move is a longer sentence rather than a look at the records.
assert "near_miss_samples" in str(e.value)
session.add.assert_not_called()
@pytest.mark.asyncio
async def test_nothing_is_written_when_the_reason_is_refused():
"""Refused BEFORE the setting is touched, not after.
Writing the value and then raising would leave the number moved and the
history empty — the exact state this table exists to make impossible.
"""
session, ctx = _patches()
with ctx[0], patch.object(rt, "set_setting", AsyncMock()) as setter, \
ctx[2], ctx[3]:
with pytest.raises(ValueError):
await rt.set_dial(1, "prompt_rule", "floor", 0.66, reason="x")
setter.assert_not_called()
@pytest.mark.asyncio
async def test_a_good_change_writes_the_setting_and_the_event():
session, ctx = _patches(floor=0.72)
reason = ("read prompt_rule's 5 highest declines: 3 were project rules for "
"another repo, so the bar is doing its job here")
with ctx[0], patch.object(rt, "set_setting", AsyncMock()) as setter, \
ctx[2], ctx[3]:
out = await rt.set_dial(1, "prompt_rule", "floor", 0.66, reason=reason)
setter.assert_awaited_once()
_uid, key, stored = setter.await_args.args
assert key == "kb_promptrule_threshold" and stored == "0.66"
event = session.add.call_args.args[0]
assert event.surface == "prompt_rule" and event.dial == "floor"
# The before-value is what makes the history answerable.
assert event.old_value == 0.72 and event.new_value == 0.66
assert event.reason == reason and event.actor == "model"
assert out["previous"] == 0.72 and out["applied"] == 0.66
assert out["clamped"] is False
@pytest.mark.asyncio
@pytest.mark.parametrize("dial, sent, applied", [
("floor", 1.4, 1.0),
("floor", -0.2, 0.0),
("budget", 99, 10),
("budget", 0, 1),
])
async def test_a_clamp_is_reported_rather_than_swallowed(dial, sent, applied):
"""Said out loud, because silence here poisons the next reading.
A caller that believes it set 1.4 treats the following week's telemetry as
evidence about a bar that never existed, and then moves the dial again to
fix a problem it invented.
"""
session, ctx = _patches()
reason = "checked the refused records for this surface and they were fine"
with ctx[0], ctx[1], ctx[2], ctx[3]:
out = await rt.set_dial(1, "auto_inject", dial, sent, reason=reason)
assert out["applied"] == applied
assert out["clamped"] is True
@pytest.mark.asyncio
async def test_an_unknown_surface_is_refused_before_anything_is_written():
session, ctx = _patches()
with ctx[0], patch.object(rt, "set_setting", AsyncMock()) as setter, \
ctx[2], ctx[3]:
with pytest.raises(ValueError):
await rt.set_dial(1, "pretool_rule", "floor", 0.6,
reason="a perfectly good reason that is long enough")
setter.assert_not_called()
session.add.assert_not_called()
@pytest.mark.asyncio
@pytest.mark.parametrize("dial", ["threshold", "limit", "", "FLOOR"])
async def test_an_unknown_dial_is_refused(dial):
"""The near-misses are the old vocabulary — `threshold` and `limit` are what
these were called before this step, so they are exactly what a stale caller
will send, and a silent no-op there would be indistinguishable from a
change that did not take."""
session, ctx = _patches()
with ctx[0], ctx[1], ctx[2], ctx[3], pytest.raises(ValueError):
await rt.set_dial(1, "auto_inject", dial, 0.6,
reason="a perfectly good reason that is long enough")
session.add.assert_not_called()
@pytest.mark.asyncio
async def test_a_human_change_is_distinguishable_from_the_models():
"""Both act as the same user, so the id cannot tell them apart — and "did I
do this, or did the session?" is the first question the history is asked."""
session, ctx = _patches()
with ctx[0], ctx[1], ctx[2], ctx[3]:
await rt.set_dial(1, "auto_inject", "floor", 0.6, actor="human",
reason="operator set this themselves in Settings")
assert session.add.call_args.args[0].actor == "human"
@pytest.mark.asyncio
async def test_an_unknown_actor_is_refused():
"""Free text here would make the column unreadable within a month."""
session, ctx = _patches()
with ctx[0], ctx[1], ctx[2], ctx[3], pytest.raises(ValueError):
await rt.set_dial(1, "auto_inject", "floor", 0.6, actor="agent",
reason="a perfectly good reason that is long enough")
def test_the_tool_teaches_reading_the_records_not_the_percentile():
"""The contract an agent actually reads (rule 167).
The failure this prevents is a model moving a bar because `near_misses.p90`
sat close to it. That has been measured pointing the wrong way — 69 declines
where every percentile said "lower it" and the refused record was a false
positive — so the docstring has to carry the method, not just the warning.
"""
from scribe.mcp.tools import retrieval_tuning as tool
doc = tool.tune_retrieval.__doc__
assert "near_miss_samples" in doc, "the tool does not name how to get records"
assert "PERCENTILE ALONE" in doc.upper()
# And the worked example, because an abstract warning loses to a number.
assert "69" in doc
def test_every_tool_in_the_module_is_registered():
from scribe.mcp.tools import retrieval_tuning as tool
from tests.helpers import FakeMCP
mcp = FakeMCP()
tool.register(mcp)
# Order is the module's, and asserted rather than sorted: an unregistered
# tool is invisible to every caller, so the list is worth reading literally.
assert mcp.names == [
"retrieval_surfaces", "migrate_retrieval_floor",
"tune_retrieval", "retrieval_tuning_history",
]
# ── record_release_defaults / floor_moves_since (#4225) ───────────────────
#
# THE HOLE THESE FILL. This table records dial turns — a person or a model
# choosing a number. It was silent about the other way a floor moves: somebody
# edits `floor_default` in the registry and ships it. That change is invisible
# to every consumer of this table, and one consumer is `band_hugs_floor`, which
# compares a band against a floor and could not notice they came from different
# regimes.
#
# Measured on the instance that found it: `write_path_rule` went 0.68 -> 0.72
# on 2026-09-02 as a shipped default, so a 30-day window opening 2026-08-22
# held six days of calls made under the old bar. The readout reported the band
# sitting "-0.0216 above" its floor — which is what a two-population comparison
# looks like when it finally says so out loud.
def _release_rows(rows):
"""Patch `async_session` so the recorder sees `rows` as what is on record.
`rows` maps (surface, dial) -> new_value already recorded by a release.
"""
session = make_mock_session()
seen = []
async def execute(stmt):
seen.append(stmt)
result = MagicMock()
# The recorder asks one question at a time, in registry order, so the
# answers are handed back in the order it asks them.
key = seen_keys.pop(0) if seen_keys else None
row = None
if key is not None and key in rows:
row = MagicMock()
row.new_value = rows[key]
result.scalars.return_value.first.return_value = row
return result
seen_keys = [(n, d) for n in rt.surface_names() for d in ("floor", "budget")]
session.execute = AsyncMock(side_effect=execute)
return session
@pytest.mark.asyncio
async def test_a_first_boot_records_a_baseline_and_calls_it_one():
"""Nothing moved. The row exists so the NEXT release has a predecessor.
Load-bearing downstream: the band check suspends itself on a genuine move
and must NOT suspend itself on a fresh install's baseline. It tells them
apart by `old_value` being null, so a baseline that claimed a change would
silently retire the check on every new install.
"""
session = _release_rows({})
with patch.object(rt, "async_session", MagicMock(return_value=session)):
written = await rt.record_release_defaults()
assert written, "a first boot records every dial"
assert all(w["baseline"] for w in written)
assert all(w["old_value"] is None for w in written)
@pytest.mark.asyncio
async def test_a_default_that_did_not_move_writes_nothing():
"""Called on every boot, so it has to be idempotent — otherwise the
history fills with rows saying the release shipped the same number again,
and a history nobody can skim is one nobody reads."""
current = {(n, d): (rt.get_surface(n).floor_default if d == "floor"
else float(rt.get_surface(n).budget_default))
for n in rt.surface_names() for d in ("floor", "budget")}
session = _release_rows(current)
with patch.object(rt, "async_session", MagicMock(return_value=session)):
written = await rt.record_release_defaults()
assert written == []
session.add.assert_not_called()
session.commit.assert_not_called()
@pytest.mark.asyncio
async def test_a_moved_default_is_recorded_with_both_values():
"""The event the whole task is about, and it carries what it moved FROM —
without that a reader knows a change happened and nothing about whether
the old sample can be pooled with the new one."""
name = rt.surface_names()[0]
s = rt.get_surface(name)
current = {(n, d): (rt.get_surface(n).floor_default if d == "floor"
else float(rt.get_surface(n).budget_default))
for n in rt.surface_names() for d in ("floor", "budget")}
current[(name, "floor")] = s.floor_default - 0.04 # what the last release shipped
session = _release_rows(current)
with patch.object(rt, "async_session", MagicMock(return_value=session)):
written = await rt.record_release_defaults()
moved = [w for w in written if w["surface"] == name and w["dial"] == "floor"]
assert len(moved) == 1
assert moved[0]["baseline"] is False
assert moved[0]["old_value"] == pytest.approx(s.floor_default - 0.04)
assert moved[0]["new_value"] == pytest.approx(s.floor_default)
@pytest.mark.asyncio
async def test_a_release_row_belongs_to_no_account():
"""`user_id IS NULL`, because no user did this.
A release acts on every account that has not overridden the dial. Writing
one row per user would both multiply the row and misattribute it, and the
readers take the newest of (this user's change, the release's) — which only
works if the release's is distinguishable.
"""
session = _release_rows({})
with patch.object(rt, "async_session", MagicMock(return_value=session)):
await rt.record_release_defaults()
added = [c.args[0] for c in session.add.call_args_list]
assert added
assert all(e.user_id is None for e in added)
assert all(e.actor == rt.RELEASE_ACTOR for e in added)
@pytest.mark.asyncio
async def test_every_release_row_states_why_it_exists():
"""`reason` is the guardrail on this table, and a row written by machinery
is the one most likely to arrive blank."""
session = _release_rows({})
with patch.object(rt, "async_session", MagicMock(return_value=session)):
await rt.record_release_defaults()
added = [c.args[0] for c in session.add.call_args_list]
assert all(e.reason and e.reason.strip() for e in added)
assert all(e.surface in e.reason for e in added)
@pytest.mark.asyncio
async def test_a_baseline_is_not_reported_as_a_floor_move():
"""`floor_moves_since` is what suspends the band check, so it must ask for
a genuine move — `old_value IS NOT NULL` — rather than for any row."""
from datetime import datetime, timezone
session = make_mock_session()
result = MagicMock()
result.scalars.return_value.all.return_value = []
session.execute = AsyncMock(return_value=result)
with patch.object(rt, "async_session", MagicMock(return_value=session)):
out = await rt.floor_moves_since(datetime.now(timezone.utc))
assert out == {}
# The filter is the whole correctness argument; assert it is in the query.
stmt = str(session.execute.call_args.args[0])
assert "old_value IS NOT NULL" in stmt
assert "dial" in stmt