Milestone 416 step 6. A retrieval floor is a cosine similarity, which only
means something inside one embedding model's geometry over documents cut one
particular way. Change either and every floor on the install keeps applying
while describing nothing — and nothing anywhere says so, because the scores
simply come out different and the bar goes on cutting.
`CHUNKER_VERSION` already solved this for documents: stamped per row, so the
backfill re-embeds precisely what is stale. The same idea, applied to the
numbers:
- `calibration_stamp()` — embedding model + document shape, one definition.
TWO fields, never a fused string (rule 149): a mismatch has to say WHICH half
moved, because they call for different responses.
- `retrieval_tuning_events` gains `embedding_model` / `shape_version`
(migration 0104), stamped on every write. Nullable and NOT backfilled —
"unstamped" is the honest answer for a row written before this existed, and
it reports as `stale: null`, never as fine.
- `current_settings` reports calibration per dial: tuned rows from their event,
untouched dials from the registry default's own stamp.
- `retrieval_surfaces` and the Settings panel show the mismatch. The panel
renders ONLY when something is stale, so seeing it at all is the signal.
- `migrate_floor` / `migrate_retrieval_floor` answers "a path for thresholds to
be inherited by the next model so that they don't have to recalibrate a lot":
the raw cosine cannot cross models, but the PERCENTILE it represented can.
Measure what fraction of a surface's logged calls the old floor admitted,
re-score those queries under the current model, take the value admitting the
same fraction. Dry run by default; applying writes an ordinary tuning event
with the arithmetic in its reason.
Nothing auto-retunes. A stale stamp says a number is no longer a measurement;
it does not say what the number should be, and #4102 measured the one case
where the statistic and the correct action pointed opposite ways.
The load-bearing test is an ABSENCE: no chat-model identifier may appear
anywhere in the calibration path. Claude produces none of these scores, so a
Claude upgrade must trigger nothing — a false alarm here teaches the operator
to ignore the real one on the day bge-small becomes bge-base.
Backup v17 carries both columns, unfilled on the way out and on the way back:
a round trip must not turn "we don't know" into a stated fact.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy