feat(telemetry): record WHAT the bar turned away, not only how close it came (#3807)
CI & Build / Python lint (push) Successful in 8s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / integration (push) Successful in 40s
CI & Build / TypeScript typecheck (push) Successful in 43s
CI & Build / Python tests (push) Successful in 1m14s
CI & Build / Build & push image (push) Successful in 2m59s
CI & Build / Python lint (push) Successful in 8s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / integration (push) Successful in 40s
CI & Build / TypeScript typecheck (push) Successful in 43s
CI & Build / Python tests (push) Successful in 1m14s
CI & Build / Build & push image (push) Successful in 2m59s
#3670 added `best_available_score` so a threshold could be judged from its rejections. It records how CLOSE the bar came to firing and not WHAT it refused, and that is the half a decision actually needs. Live, pre_tool_rule sits at a ~0.72 bar with a near-miss p90 of 0.7071 — about 117 declines a day within 0.013 of firing. Dropping to 0.707 would take that arm from 22 hits a day to roughly 139: six-fold, on a surface that runs before every Bash call. The percentile says the mass is there. Nothing said whether it was worth showing. NEITHER OBVIOUS INSTRUMENT ANSWERS IT. Pull-through cannot: the injected rule line already carries title and trigger, so a session can comply without ever calling get_rule, and rule pull-through understates usefulness by construction. Reading the rejected records can — and `result_ids` holds only what was RETURNED, so on a zero-result call the near-missed record had no name at all. So the id, from the SAME ranked candidate as the score. Both searches unpack `best` once and read both fields off it, because splitting that into two expressions is exactly how a later edit pairs a score with its neighbour's id — and a score attached to the wrong record is worse than no id, since it invites judging the wrong one and concluding the bar is fine. write_path withholds the id on the same condition it withholds the score (#3739): a surviving id beside a null score names a record without saying what it scored, the pair disagreeing in the other direction. THE READ PATH IS A LISTING, NOT A STATISTIC — an id cannot be percentiled, and a reader tuning a bar needs to go and read the records. Opt-in via `near_miss_samples` (0-20, default 0) so the ordinary readout keeps its size, and deliberately NOT a window function: this module's one production outage was a grouped query Postgres rejected, swallowed by the broad except, every counter reading zero while the mocked tests passed (#2663). One flat ordered query, overfetched, bucketed in Python — the shape that lesson prescribes. Migration 0097, nullable and unbackfilled. Not a foreign key: the table spans record types and `source` says which, exactly as result_ids works. The integration guard pins the listing as PER SOURCE. A global LIMIT would let a noisy source eat the whole quota and leave the surface being tuned showing nothing — which reads as "nothing was close", the misreading this milestone has spent itself correcting. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011cPyzNnegXHr5iRMzzy5KJ
This commit is contained in:
@@ -0,0 +1,60 @@
|
||||
"""add retrieval_logs.best_available_id — WHICH record the bar refused (#3807)
|
||||
|
||||
Revision ID: 0097
|
||||
Revises: 0096
|
||||
Create Date: 2026-09-09
|
||||
|
||||
0096 added `best_available_score` so a threshold could be judged from what it
|
||||
rejected. It records how CLOSE the bar came to firing and not WHAT it turned
|
||||
away, and that turns out to be the half a decision actually needs.
|
||||
|
||||
Live, `pre_tool_rule` shows a bar of ~0.72 with a near-miss p90 of 0.7071 —
|
||||
about 117 declines a day sitting within 0.013 of firing. Dropping the bar to
|
||||
0.707 would take that arm from 22 hits a day to roughly 139, a six-fold change
|
||||
on a surface that runs before every Bash call. The percentile says the mass is
|
||||
there. Nothing says whether it is worth showing.
|
||||
|
||||
AND THE TWO OBVIOUS INSTRUMENTS DO NOT ANSWER IT. Pull-through cannot: the
|
||||
injected rule line already carries title and trigger, so a session can comply
|
||||
without ever calling `get_rule`, and rule pull-through therefore understates
|
||||
usefulness by construction. Reading the rejected records can — and `result_ids`
|
||||
holds only what was RETURNED, so on a zero-result call it is empty and the
|
||||
near-missed record has no name.
|
||||
|
||||
So: the id, beside the score, from the SAME ranked candidate. The two must
|
||||
never be able to describe different records — a score paired with its
|
||||
neighbour's id would be worse than no id at all, because it invites a reader to
|
||||
judge the wrong record and conclude the bar is fine.
|
||||
|
||||
NULLABLE AND UNBACKFILLED, for the reason 0095 and 0096 both spell out: a row
|
||||
written before this genuinely does not know, and inventing a value would put an
|
||||
artifact where a measurement belongs. Null here means "not measured", never
|
||||
"nothing was close".
|
||||
|
||||
NOT A FOREIGN KEY, deliberately. `retrieval_logs` spans record types — the
|
||||
rules arms store rule ids, the note arms store note ids — and `source` is what
|
||||
says which table an id belongs to, exactly as `result_ids` has always worked.
|
||||
A constraint would have to point at one table and would be wrong for the other.
|
||||
|
||||
Downgrade drops the column. Purely observational — nothing reads it for
|
||||
correctness.
|
||||
"""
|
||||
from alembic import op
|
||||
import sqlalchemy as sa
|
||||
|
||||
|
||||
revision = "0097"
|
||||
down_revision = "0096"
|
||||
branch_labels = None
|
||||
depends_on = None
|
||||
|
||||
|
||||
def upgrade() -> None:
|
||||
op.add_column(
|
||||
"retrieval_logs",
|
||||
sa.Column("best_available_id", sa.Integer(), nullable=True),
|
||||
)
|
||||
|
||||
|
||||
def downgrade() -> None:
|
||||
op.drop_column("retrieval_logs", "best_available_id")
|
||||
Reference in New Issue
Block a user