CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / integration (push) Successful in 43s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Failing after 1m2s
CI & Build / Build & push image (push) Skipped
CI on 09b4845 caught both, and the second is an improvement rather than a
repair.
`test_services_reply_preferences` patched `reply_preferences.get_setting`,
which the registry refactor removed — its floor and budget are resolved through
`retrieval_surfaces` now. Its key-independence test also asserted the arm asks
for exactly one key; it asks for two, because both of its numbers are its own
since this step, so the assertion names both and adds a check that the registry
and the module constant still agree about the floor key. They writing different
keys is the failure where the Settings form saves one string and the arm reads
another.
`test_settings_defaults_agree` parsed `plugin_context.py` for a bare
module-level float, and those constants now alias the registry. Rather than
teach the regex about aliases, the six retrieval floors are keyed on their
SURFACE NAME and read from the registry directly — which is strictly better for
this guard: a surface name is also its telemetry source, so a row names the same
arm the readout does, and a floor cannot be checked against a stale constant
that happened to keep its old value. `PLAN_MATCH_DEFAULT_THRESHOLD` is not a
push surface and keeps the older shape, with a note saying why.
Added while there: `test_every_tunable_surface_has_a_control`, derived from the
registry, so a seventh surface arrives as a failing test rather than as a number
only the model can reach (rules 25, 27).
Also lands the audit trail this step needs — `retrieval_tuning_events` (model +
migration 0103) and `services/retrieval_tuning.py`. The tool layer on top is
the next commit; the table is here because `reason` being REQUIRED is the whole
guardrail, and the schema is where that starts. A number moved silently leaves
nothing for the operator to review or disagree with, and the operator's decision
is that the model moves these "9 times out of 10".
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
76 lines
2.9 KiB
Python
76 lines
2.9 KiB
Python
"""retrieval_tuning_events — why a floor is where it is (#4102)
|
|
|
|
Revision ID: 0103
|
|
Revises: 0102
|
|
Create Date: 2026-09-17
|
|
|
|
Milestone 416 stops shipping similarity thresholds as values somebody has to
|
|
defend, and hands the adjustment to the model that reads the surface's own
|
|
telemetry. The operator's decision:
|
|
|
|
"the floor should be chosen and adjusted by the model using it… the user
|
|
should be able to touch it but the model should be the thing handling it 9
|
|
times out of 10."
|
|
|
|
The number itself already has a home — the generic settings table. What has no
|
|
home is the ARGUMENT, and once the values move on their own the argument is the
|
|
part an operator needs: what changed, from what to what, who moved it, and on
|
|
what evidence. This table is that trail, and it is what makes the delegation
|
|
reviewable rather than merely automatic.
|
|
|
|
Nothing is backfilled. A surface with no rows here is sitting on its shipped
|
|
starting point, which is a true and useful thing for the history to say.
|
|
"""
|
|
import sqlalchemy as sa
|
|
from alembic import op
|
|
|
|
revision = "0103"
|
|
down_revision = "0102"
|
|
branch_labels = None
|
|
depends_on = None
|
|
|
|
|
|
def upgrade() -> None:
|
|
op.create_table(
|
|
"retrieval_tuning_events",
|
|
sa.Column("id", sa.Integer(), primary_key=True),
|
|
sa.Column(
|
|
"created_at",
|
|
sa.DateTime(timezone=True),
|
|
server_default=sa.text("now()"),
|
|
nullable=False,
|
|
),
|
|
# FK-free, like retrieval_logs and app_logs: the record of why a number
|
|
# is where it is must outlive the account that moved it.
|
|
sa.Column("user_id", sa.Integer(), nullable=True),
|
|
# Also the surface's `retrieval_logs.source`, so a change can be read
|
|
# next to what the change did.
|
|
sa.Column("surface", sa.Text(), nullable=False),
|
|
sa.Column("dial", sa.Text(), nullable=False),
|
|
# Nullable: the first change to a surface has no stored predecessor. It
|
|
# moved off the shipped starting point, which is a different event from
|
|
# moving off a value somebody chose.
|
|
sa.Column("old_value", sa.Float(), nullable=True),
|
|
sa.Column("new_value", sa.Float(), nullable=False),
|
|
sa.Column(
|
|
"actor", sa.Text(), nullable=False, server_default=sa.text("'model'")
|
|
),
|
|
# Non-null here; non-BLANK is enforced at the service boundary, because
|
|
# a column that merely forbids NULL is satisfied by "" and a required
|
|
# field that accepts "" is a formality.
|
|
sa.Column("reason", sa.Text(), nullable=False),
|
|
)
|
|
# The only read this table has: one surface's history, newest first.
|
|
op.create_index(
|
|
"ix_retrieval_tuning_surface_created",
|
|
"retrieval_tuning_events",
|
|
["surface", sa.text("created_at DESC")],
|
|
)
|
|
|
|
|
|
def downgrade() -> None:
|
|
op.drop_index(
|
|
"ix_retrieval_tuning_surface_created", table_name="retrieval_tuning_events"
|
|
)
|
|
op.drop_table("retrieval_tuning_events")
|