feat(retrieval): a tuned number carries the space it was measured in (#4104)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Failing after 1m4s
CI & Build / Build & push image (push) Skipped

Milestone 416 step 6. A retrieval floor is a cosine similarity, which only
means something inside one embedding model's geometry over documents cut one
particular way. Change either and every floor on the install keeps applying
while describing nothing — and nothing anywhere says so, because the scores
simply come out different and the bar goes on cutting.

`CHUNKER_VERSION` already solved this for documents: stamped per row, so the
backfill re-embeds precisely what is stale. The same idea, applied to the
numbers:

- `calibration_stamp()` — embedding model + document shape, one definition.
  TWO fields, never a fused string (rule 149): a mismatch has to say WHICH half
  moved, because they call for different responses.
- `retrieval_tuning_events` gains `embedding_model` / `shape_version`
  (migration 0104), stamped on every write. Nullable and NOT backfilled —
  "unstamped" is the honest answer for a row written before this existed, and
  it reports as `stale: null`, never as fine.
- `current_settings` reports calibration per dial: tuned rows from their event,
  untouched dials from the registry default's own stamp.
- `retrieval_surfaces` and the Settings panel show the mismatch. The panel
  renders ONLY when something is stale, so seeing it at all is the signal.
- `migrate_floor` / `migrate_retrieval_floor` answers "a path for thresholds to
  be inherited by the next model so that they don't have to recalibrate a lot":
  the raw cosine cannot cross models, but the PERCENTILE it represented can.
  Measure what fraction of a surface's logged calls the old floor admitted,
  re-score those queries under the current model, take the value admitting the
  same fraction. Dry run by default; applying writes an ordinary tuning event
  with the arithmetic in its reason.

Nothing auto-retunes. A stale stamp says a number is no longer a measurement;
it does not say what the number should be, and #4102 measured the one case
where the statistic and the correct action pointed opposite ways.

The load-bearing test is an ABSENCE: no chat-model identifier may appear
anywhere in the calibration path. Claude produces none of these scores, so a
Claude upgrade must trigger nothing — a false alarm here teaches the operator
to ignore the real one on the day bge-small becomes bge-base.

Backup v17 carries both columns, unfilled on the way out and on the way back:
a round trip must not turn "we don't know" into a stated fact.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
2026-09-17 12:45:48 -04:00
co-authored by Claude Opus 5
parent dcf800ed65
commit aee24c9c1c
13 changed files with 1002 additions and 2 deletions
+68
View File
@@ -47,6 +47,7 @@ from sqlalchemy import select
from scribe.models import async_session
from scribe.models.retrieval_tuning import RetrievalTuningEvent
from scribe.services.embeddings import calibration_stamp
from scribe.services.retrieval_surfaces import (
MAX_BUDGET,
budget_for,
@@ -86,6 +87,50 @@ def _clean_reason(reason: str) -> str:
return text
def _calibration(row, s, live: dict) -> dict:
"""What space one dial's number was chosen in, and whether that space moved.
Three sources, and they are not interchangeable:
- `tuned` — the dial was moved after #4104 and the event carries its
stamp. The only case where the answer is known.
- `shipped` — the dial has never been moved, so the number in force is
the registry default, and the registry records what THAT
was measured against (`Surface.measured_model`).
- `unstamped` — the dial was moved before this step existed. It was
measured under something; naming it would be inventing a
fact, so `stale` is None rather than True or False.
"Unknown" and "fine" must not render the same.
`model_changed` and `shape_changed` are reported apart (rule 149) because
they call for different responses: a new embedding model means every number
is a distance in a geometry that no longer exists, while a re-cut document
shape means the same records now embed different text. A caller that only
ever sees `stale: true` cannot tell those apart.
"""
if row is None:
model, shape, source = s.measured_model, s.measured_shape, "shipped"
elif row.embedding_model is None and row.shape_version is None:
return {
"source": "unstamped", "embedding_model": None,
"shape_version": None, "model_changed": None,
"shape_changed": None, "stale": None,
}
else:
model, shape, source = row.embedding_model, row.shape_version, "tuned"
model_changed = model != live["embedding_model"]
shape_changed = shape != live["shape_version"]
return {
"source": source,
"embedding_model": model,
"shape_version": shape,
"model_changed": model_changed,
"shape_changed": shape_changed,
"stale": model_changed or shape_changed,
}
async def current_settings(user_id: int) -> list[dict]:
"""Every tunable surface with its live pair and its last stated reason.
@@ -94,7 +139,16 @@ async def current_settings(user_id: int) -> list[dict]:
often — because a floor cannot be moved sensibly without those three — and
the reason last given, so the next change argues with the last one instead
of overwriting it blind.
From #4104 it also carries `calibration` per dial: the embedding model and
document shape the number in force was measured in, and whether either has
moved since. NOTHING AUTO-RETUNES on the strength of it. A stale stamp says
a number is a measurement of a space that no longer exists — it does not
say what the number should be now, and the one time a statistic was allowed
to answer that question it was wrong (see the module docstring). The stamp
is here so a reader knows which floors to go and re-measure.
"""
live = calibration_stamp()
out: list[dict] = []
async with async_session() as session:
for name in surface_names():
@@ -126,6 +180,13 @@ async def current_settings(user_id: int) -> list[dict]:
"last_change": {
dial: last[dial].to_dict() for dial in DIALS if dial in last
},
# Always present for BOTH dials, unlike `last_change`: an
# untouched dial still has a number in force, and that number
# was still measured in some space. Reported, never acted on —
# see the note below on why nothing auto-retunes.
"calibration": {
dial: _calibration(last.get(dial), s, live) for dial in DIALS
},
})
return out
@@ -179,9 +240,16 @@ async def set_dial(
await set_setting(user_id, key, stored)
async with async_session() as session:
# Stamped with the space this number was chosen in (#4104). Read at
# write time rather than passed in: the caller measuring a floor and
# the caller recording it are the same call, so there is no window in
# which they could disagree.
stamp = calibration_stamp()
session.add(RetrievalTuningEvent(
user_id=user_id, surface=surface, dial=dial,
old_value=old, new_value=applied, actor=actor, reason=text,
embedding_model=stamp["embedding_model"],
shape_version=stamp["shape_version"],
))
await session.commit()