feat(retrieval): a review pass judges whether injected lines related — menus_to_review, judge_menu and a judged readout (#4772)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 44s

An open rate cannot say whether a menu line related: every line carries its
matched passage (#4364), so "not opened" covers unrelated, enough as shown,
and already in context. #4772 "Injected notes are never judged".

- retrieval_judgments (0115): a reviewer verdict per line of a logged call,
  on_point / adjacent / unrelated, with its reason, rank, budget side and
  whether the agent opened it within the hour.
- menus_to_review re-runs a random sample of unjudged auto_inject calls with
  the arm's own parameters, past its budget, passage on every line.
  judge_menu records verdicts, re-deriving rank from a fresh re-run.
- retrieval_telemetry gains a judged block (by rank, within/beyond budget,
  on_point_unopened). surfaced_never_pulled stops blaming titles.
- missed-retrieval guidance names the review before a budget move.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-10-03 15:34:14 -04:00
co-authored by Claude Opus 5.5
parent 0a1bb68808
commit 6598c7fa85
13 changed files with 879 additions and 11 deletions
+1
View File
@@ -61,6 +61,7 @@ from scribe.models.retrieval_tuning import RetrievalTuningEvent # noqa: E402, F
from scribe.models.note_usage import NoteUsageEvent # noqa: E402, F401
from scribe.models.rule_usage import RuleUsageEvent # noqa: E402, F401
from scribe.models.system_usage import SystemUsageEvent # noqa: E402, F401
from scribe.models.retrieval_judgment import RetrievalJudgment # noqa: E402, F401
from scribe.models.project import Project # noqa: E402, F401
from scribe.models.milestone import Milestone # noqa: E402, F401
from scribe.models.task_log import TaskLog # noqa: E402, F401
+82
View File
@@ -0,0 +1,82 @@
from sqlalchemy import BigInteger, Boolean, Float, Index, Integer, Text, UniqueConstraint
from sqlalchemy.orm import Mapped, mapped_column
from scribe.models import Base
from scribe.models.base import CreatedAtMixin, iso
ON_POINT = "on_point"
ADJACENT = "adjacent"
UNRELATED = "unrelated"
VERDICTS = (ON_POINT, ADJACENT, UNRELATED)
class RetrievalJudgment(Base, CreatedAtMixin):
"""One reviewer's verdict on one candidate of one logged retrieval call
(#4772).
The relevance half the usage tables cannot hold. `note_usage_events` says
whether a line was OPENED, and since every menu line carries its matched
passage (#4364) "not opened" covers an unrelated line, a line whose passage
was enough, and a line the reader already had. Only someone reading the
query beside the line can tell those apart, so this records that reading —
with its reason, because a verdict nobody can re-read is a threshold.
Keyed to the `retrieval_logs` row it judges, by id and FK-free like every
telemetry table here. `query` is COPIED rather than joined: the verdict is
only re-readable beside the words it was judged against.
`rank` and `score` are the candidate's place when the logged query was
re-run for review, which may run past the arm's budget on purpose — the
candidates just under the cut are the ones a budget change would add.
`within_budget` says which side of the cut the line was.
`opened_after` is whether the reviewer's own agent pulled the record within
an hour of the call, read from `note_usage_events`. A correlation, not a
session join — the server has no session identity (see NoteUsageEvent) —
and None when it was not measured.
"""
__tablename__ = "retrieval_judgments"
id: Mapped[int] = mapped_column(BigInteger, primary_key=True)
user_id: Mapped[int | None] = mapped_column(BigInteger, nullable=True)
retrieval_log_id: Mapped[int] = mapped_column(BigInteger, nullable=False)
# The judged call's surface, copied so the readout groups without a join.
source: Mapped[str] = mapped_column(Text, nullable=False)
query: Mapped[str | None] = mapped_column(Text, nullable=True)
record_id: Mapped[int] = mapped_column(BigInteger, nullable=False)
# 1-based, in the review's re-run.
rank: Mapped[int] = mapped_column(Integer, nullable=False)
score: Mapped[float | None] = mapped_column(Float, nullable=True)
within_budget: Mapped[bool] = mapped_column(Boolean, nullable=False)
# One of VERDICTS. Plain Text, no CHECK, like the usage tables; the
# service refuses anything else.
verdict: Mapped[str] = mapped_column(Text, nullable=False)
reason: Mapped[str] = mapped_column(Text, nullable=False)
opened_after: Mapped[bool | None] = mapped_column(Boolean, nullable=True)
__table_args__ = (
# One verdict per reviewer per line; judging it again replaces it.
UniqueConstraint(
"retrieval_log_id", "record_id", "user_id",
name="uq_retrieval_judgment_line",
),
Index("ix_retrieval_judgments_source_created", "source", "created_at"),
)
def to_dict(self) -> dict:
return {
"id": self.id,
"created_at": iso(self.created_at),
"user_id": self.user_id,
"retrieval_log_id": self.retrieval_log_id,
"source": self.source,
"query": self.query,
"record_id": self.record_id,
"rank": self.rank,
"score": self.score,
"within_budget": self.within_budget,
"verdict": self.verdict,
"reason": self.reason,
"opened_after": self.opened_after,
}