feat(retrieval): a review pass judges whether injected lines related — menus_to_review, judge_menu and a judged readout (#4772)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 44s
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 44s
An open rate cannot say whether a menu line related: every line carries its matched passage (#4364), so "not opened" covers unrelated, enough as shown, and already in context. #4772 "Injected notes are never judged". - retrieval_judgments (0115): a reviewer verdict per line of a logged call, on_point / adjacent / unrelated, with its reason, rank, budget side and whether the agent opened it within the hour. - menus_to_review re-runs a random sample of unjudged auto_inject calls with the arm's own parameters, past its budget, passage on every line. judge_menu records verdicts, re-deriving rank from a fresh re-run. - retrieval_telemetry gains a judged block (by rank, within/beyond budget, on_point_unopened). surfaced_never_pulled stops blaming titles. - missed-retrieval guidance names the review before a budget move. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,56 @@
|
||||
"""retrieval_judgments — a reviewer's verdict on each line of a logged menu
|
||||
(#4772)
|
||||
|
||||
Revision ID: 0115
|
||||
Revises: 0114
|
||||
Create Date: 2026-10-03
|
||||
|
||||
Whether an injected line RELATED, which the usage tables cannot say: since
|
||||
every line carries its matched passage, "not opened" no longer implies "not
|
||||
relevant". FK-free like the telemetry tables beside it; no CHECK on `verdict`,
|
||||
which the service validates.
|
||||
"""
|
||||
import sqlalchemy as sa
|
||||
from alembic import op
|
||||
|
||||
revision = "0115"
|
||||
down_revision = "0114"
|
||||
branch_labels = None
|
||||
depends_on = None
|
||||
|
||||
|
||||
def upgrade() -> None:
|
||||
op.create_table(
|
||||
"retrieval_judgments",
|
||||
sa.Column("id", sa.BigInteger(), primary_key=True),
|
||||
sa.Column(
|
||||
"created_at", sa.DateTime(timezone=True), nullable=False,
|
||||
server_default=sa.text("now()"),
|
||||
),
|
||||
sa.Column("user_id", sa.BigInteger(), nullable=True),
|
||||
sa.Column("retrieval_log_id", sa.BigInteger(), nullable=False),
|
||||
sa.Column("source", sa.Text(), nullable=False),
|
||||
sa.Column("query", sa.Text(), nullable=True),
|
||||
sa.Column("record_id", sa.BigInteger(), nullable=False),
|
||||
sa.Column("rank", sa.Integer(), nullable=False),
|
||||
sa.Column("score", sa.Float(), nullable=True),
|
||||
sa.Column("within_budget", sa.Boolean(), nullable=False),
|
||||
sa.Column("verdict", sa.Text(), nullable=False),
|
||||
sa.Column("reason", sa.Text(), nullable=False),
|
||||
sa.Column("opened_after", sa.Boolean(), nullable=True),
|
||||
sa.UniqueConstraint(
|
||||
"retrieval_log_id", "record_id", "user_id",
|
||||
name="uq_retrieval_judgment_line",
|
||||
),
|
||||
)
|
||||
op.create_index(
|
||||
"ix_retrieval_judgments_source_created", "retrieval_judgments",
|
||||
["source", "created_at"],
|
||||
)
|
||||
|
||||
|
||||
def downgrade() -> None:
|
||||
op.drop_index(
|
||||
"ix_retrieval_judgments_source_created", table_name="retrieval_judgments",
|
||||
)
|
||||
op.drop_table("retrieval_judgments")
|
||||
Reference in New Issue
Block a user