CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Failing after 1m4s
CI & Build / Build & push image (push) Skipped
Milestone 416 step 6. A retrieval floor is a cosine similarity, which only means something inside one embedding model's geometry over documents cut one particular way. Change either and every floor on the install keeps applying while describing nothing — and nothing anywhere says so, because the scores simply come out different and the bar goes on cutting. `CHUNKER_VERSION` already solved this for documents: stamped per row, so the backfill re-embeds precisely what is stale. The same idea, applied to the numbers: - `calibration_stamp()` — embedding model + document shape, one definition. TWO fields, never a fused string (rule 149): a mismatch has to say WHICH half moved, because they call for different responses. - `retrieval_tuning_events` gains `embedding_model` / `shape_version` (migration 0104), stamped on every write. Nullable and NOT backfilled — "unstamped" is the honest answer for a row written before this existed, and it reports as `stale: null`, never as fine. - `current_settings` reports calibration per dial: tuned rows from their event, untouched dials from the registry default's own stamp. - `retrieval_surfaces` and the Settings panel show the mismatch. The panel renders ONLY when something is stale, so seeing it at all is the signal. - `migrate_floor` / `migrate_retrieval_floor` answers "a path for thresholds to be inherited by the next model so that they don't have to recalibrate a lot": the raw cosine cannot cross models, but the PERCENTILE it represented can. Measure what fraction of a surface's logged calls the old floor admitted, re-score those queries under the current model, take the value admitting the same fraction. Dry run by default; applying writes an ordinary tuning event with the arithmetic in its reason. Nothing auto-retunes. A stale stamp says a number is no longer a measurement; it does not say what the number should be, and #4102 measured the one case where the statistic and the correct action pointed opposite ways. The load-bearing test is an ABSENCE: no chat-model identifier may appear anywhere in the calibration path. Claude produces none of these scores, so a Claude upgrade must trigger nothing — a false alarm here teaches the operator to ignore the real one on the day bge-small becomes bge-base. Backup v17 carries both columns, unfilled on the way out and on the way back: a round trip must not turn "we don't know" into a stated fact. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
192 lines
8.0 KiB
Python
192 lines
8.0 KiB
Python
"""A tuned number carries the space it was measured in — and only that (#4104).
|
|
|
|
WHY THIS EXISTS
|
|
|
|
Milestone 416 step 6 answers *"a path for thresholds to be inherited by the next
|
|
version or different model so that they don't have to recalibrate a lot."* The
|
|
mechanism is a stamp: every tuning event records the embedding model and
|
|
document shape the number was chosen under, and a mismatch is reported rather
|
|
than left to be noticed.
|
|
|
|
THE ONE THAT MATTERS MOST IS THE ABSENCE
|
|
|
|
`test_no_chat_model_identifier_reaches_the_calibration_path` is the load-bearing
|
|
guard here, and it is a guard against a plausible mistake rather than a
|
|
hypothetical one. These floors live in BAAI/bge-small-en-v1.5's vector space.
|
|
The CHAT model — Claude — produces none of these scores and changes none of
|
|
them, so a Claude upgrade must trigger nothing at all. If it ever did, the
|
|
operator would be asked to recalibrate six dials for no reason, learn that this
|
|
surface cries wolf, and ignore it on the day `bge-small` → `bge-base` actually
|
|
moves every score on the install. A false alarm here does not cost a
|
|
notification; it costs the real alarm.
|
|
|
|
The rest pins the honest-unknown behaviour: a row written before stamps existed
|
|
reports as `unstamped` with `stale: None`, never as fine and never as a model
|
|
name somebody guessed at.
|
|
"""
|
|
import re
|
|
from unittest.mock import AsyncMock, MagicMock, patch
|
|
|
|
import pytest
|
|
|
|
from scribe.services import embeddings as emb
|
|
from scribe.services import retrieval_tuning as rt
|
|
from scribe.services.retrieval_surfaces import SURFACES, get_surface
|
|
from tests.helpers import make_mock_session, tool_doc
|
|
|
|
# Names for the thing that talks, as opposed to the thing that embeds. Any of
|
|
# these appearing in the calibration path means a chat-model change could move a
|
|
# stamp — see the module docstring for why that is the expensive failure.
|
|
_CHAT_MODEL_WORDS = (
|
|
"claude", "opus", "sonnet", "haiku", "gpt", "llama", "mistral", "gemini",
|
|
"anthropic", "openai", "chat_model", "chatmodel", "model_version",
|
|
)
|
|
|
|
|
|
def test_calibration_stamp_names_the_embedder_and_the_shape():
|
|
stamp = emb.calibration_stamp()
|
|
assert stamp == {
|
|
"embedding_model": emb.EMBEDDING_MODEL,
|
|
"shape_version": emb.CHUNKER_VERSION,
|
|
}
|
|
# Two keys, never one fused string (rule 149): a mismatch has to be able to
|
|
# say WHICH half moved, because a new embedder and a re-cut document
|
|
# invalidate the same numbers for different reasons.
|
|
assert len(stamp) == 2
|
|
|
|
|
|
def test_the_stamp_tracks_the_embedding_model_constant_not_a_copy_of_it():
|
|
"""A second literal would drift, and drift here reads as a model change."""
|
|
with patch.object(emb, "EMBEDDING_MODEL", "some/other-model"):
|
|
assert emb.calibration_stamp()["embedding_model"] == "some/other-model"
|
|
with patch.object(emb, "CHUNKER_VERSION", 99):
|
|
assert emb.calibration_stamp()["shape_version"] == 99
|
|
|
|
|
|
def test_no_chat_model_identifier_reaches_the_calibration_path():
|
|
"""THE GUARD. A Claude upgrade must move nothing in this path.
|
|
|
|
Structural rather than behavioural on purpose: the failure is someone
|
|
*adding* a chat-model field in good faith ("surely the model matters"), and
|
|
no behavioural test catches a field that has not been written yet. Asserted
|
|
over the source of every module that produces or consumes a stamp.
|
|
"""
|
|
import inspect
|
|
|
|
from scribe.services import retrieval_migration as rm
|
|
|
|
for label, obj in (
|
|
("calibration_stamp", emb.calibration_stamp),
|
|
("_calibration", rt._calibration),
|
|
("migrate_floor", rm.migrate_floor),
|
|
):
|
|
src = inspect.getsource(obj).lower()
|
|
# Comments and docstrings here legitimately discuss the chat model in
|
|
# order to rule it out, so strip prose and assert on code only.
|
|
code = "\n".join(
|
|
line.split("#", 1)[0]
|
|
for line in re.sub(r'""".*?"""', "", src, flags=re.S).splitlines()
|
|
)
|
|
for word in _CHAT_MODEL_WORDS:
|
|
assert word not in code, f"{label} names the chat model: {word!r}"
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_a_moved_dial_is_written_with_the_live_stamp():
|
|
session = make_mock_session()
|
|
with patch.object(rt, "async_session", MagicMock(return_value=session)), \
|
|
patch.object(rt, "set_setting", AsyncMock()), \
|
|
patch.object(rt, "floor_for", AsyncMock(return_value=0.72)), \
|
|
patch.object(rt, "budget_for", AsyncMock(return_value=3)), \
|
|
patch.object(rt, "calibration_stamp",
|
|
MagicMock(return_value={"embedding_model": "m/x",
|
|
"shape_version": 7})):
|
|
await rt.set_dial(
|
|
1, "prompt_rule", "floor", 0.66,
|
|
reason="read the five refused records; four were genuine matches",
|
|
)
|
|
event = session.add.call_args[0][0]
|
|
assert event.embedding_model == "m/x"
|
|
assert event.shape_version == 7
|
|
|
|
|
|
def _row(model="BAAI/bge-small-en-v1.5", shape=1):
|
|
return MagicMock(embedding_model=model, shape_version=shape)
|
|
|
|
|
|
def test_an_untouched_dial_reports_the_shipped_default_it_is_still_on():
|
|
s = get_surface("prompt_rule")
|
|
live = {"embedding_model": s.measured_model, "shape_version": s.measured_shape}
|
|
cal = rt._calibration(None, s, live)
|
|
assert cal["source"] == "shipped"
|
|
assert cal["stale"] is False
|
|
|
|
|
|
def test_a_dial_moved_before_stamps_existed_is_unknown_not_fine():
|
|
"""`stale: None`, because "we don't know" and "it's fine" are not the same.
|
|
|
|
Collapsing them would hide the dials MOST likely to be wrong — the ones
|
|
somebody tuned longest ago.
|
|
"""
|
|
s = get_surface("prompt_rule")
|
|
cal = rt._calibration(_row(model=None, shape=None), s, emb.calibration_stamp())
|
|
assert cal["source"] == "unstamped"
|
|
assert cal["stale"] is None
|
|
assert cal["model_changed"] is None and cal["shape_changed"] is None
|
|
|
|
|
|
def test_a_model_change_and_a_shape_change_are_reported_apart():
|
|
"""Rule 149. Two conditions, two answers — they call for different work."""
|
|
s = get_surface("prompt_rule")
|
|
live = {"embedding_model": "new/model", "shape_version": 1}
|
|
cal = rt._calibration(_row(model="old/model", shape=1), s, live)
|
|
assert (cal["model_changed"], cal["shape_changed"], cal["stale"]) == (
|
|
True, False, True,
|
|
)
|
|
|
|
live = {"embedding_model": "old/model", "shape_version": 2}
|
|
cal = rt._calibration(_row(model="old/model", shape=1), s, live)
|
|
assert (cal["model_changed"], cal["shape_changed"], cal["stale"]) == (
|
|
False, True, True,
|
|
)
|
|
|
|
|
|
def test_a_dial_tuned_under_the_live_stamp_is_not_stale():
|
|
s = get_surface("prompt_rule")
|
|
live = emb.calibration_stamp()
|
|
cal = rt._calibration(
|
|
_row(model=live["embedding_model"], shape=live["shape_version"]), s, live,
|
|
)
|
|
assert cal["source"] == "tuned"
|
|
assert cal["stale"] is False
|
|
|
|
|
|
def test_every_surface_ships_a_stamp_for_its_defaults():
|
|
"""A default with no stamp cannot be told from one measured yesterday."""
|
|
for name, s in SURFACES.items():
|
|
assert s.measured_model, f"{name} default has no measured model"
|
|
assert isinstance(s.measured_shape, int), f"{name} shape is not a version"
|
|
|
|
|
|
def test_the_registry_stamp_is_a_literal_not_the_live_value():
|
|
"""It says what WAS true, so it must not follow the current constants.
|
|
|
|
A field that tracked `EMBEDDING_MODEL` would agree with it forever and could
|
|
never report the one thing it exists to report.
|
|
"""
|
|
with patch.object(emb, "EMBEDDING_MODEL", "some/other-model"):
|
|
assert get_surface("prompt_rule").measured_model != "some/other-model"
|
|
|
|
|
|
def test_the_tool_says_nothing_auto_retunes():
|
|
"""The contract an agent reads. A stale stamp is a prompt to MEASURE.
|
|
|
|
Without this, the obvious next move on seeing `stale: true` is to move the
|
|
dial — which is tuning from a statistic, the exact failure #4102 measured
|
|
pointing the wrong way.
|
|
"""
|
|
doc = tool_doc("scribe.mcp.tools.retrieval_tuning", "retrieval_surfaces")
|
|
assert "calibration" in doc
|
|
assert "NOTHING IS RETUNED AUTOMATICALLY" in doc
|
|
assert "unstamped" in doc
|