feat(retrieval): a review pass judges whether injected lines related — menus_to_review, judge_menu and a judged readout (#4772)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 44s
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 44s
An open rate cannot say whether a menu line related: every line carries its matched passage (#4364), so "not opened" covers unrelated, enough as shown, and already in context. #4772 "Injected notes are never judged". - retrieval_judgments (0115): a reviewer verdict per line of a logged call, on_point / adjacent / unrelated, with its reason, rank, budget side and whether the agent opened it within the hour. - menus_to_review re-runs a random sample of unjudged auto_inject calls with the arm's own parameters, past its budget, passage on every line. judge_menu records verdicts, re-deriving rank from a fresh re-run. - retrieval_telemetry gains a judged block (by rank, within/beyond budget, on_point_unopened). surfaced_never_pulled stops blaming titles. - missed-retrieval guidance names the review before a budget move. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -61,6 +61,7 @@ from scribe.models.retrieval_tuning import RetrievalTuningEvent # noqa: E402, F
|
||||
from scribe.models.note_usage import NoteUsageEvent # noqa: E402, F401
|
||||
from scribe.models.rule_usage import RuleUsageEvent # noqa: E402, F401
|
||||
from scribe.models.system_usage import SystemUsageEvent # noqa: E402, F401
|
||||
from scribe.models.retrieval_judgment import RetrievalJudgment # noqa: E402, F401
|
||||
from scribe.models.project import Project # noqa: E402, F401
|
||||
from scribe.models.milestone import Milestone # noqa: E402, F401
|
||||
from scribe.models.task_log import TaskLog # noqa: E402, F401
|
||||
|
||||
@@ -0,0 +1,82 @@
|
||||
from sqlalchemy import BigInteger, Boolean, Float, Index, Integer, Text, UniqueConstraint
|
||||
from sqlalchemy.orm import Mapped, mapped_column
|
||||
|
||||
from scribe.models import Base
|
||||
from scribe.models.base import CreatedAtMixin, iso
|
||||
|
||||
ON_POINT = "on_point"
|
||||
ADJACENT = "adjacent"
|
||||
UNRELATED = "unrelated"
|
||||
VERDICTS = (ON_POINT, ADJACENT, UNRELATED)
|
||||
|
||||
|
||||
class RetrievalJudgment(Base, CreatedAtMixin):
|
||||
"""One reviewer's verdict on one candidate of one logged retrieval call
|
||||
(#4772).
|
||||
|
||||
The relevance half the usage tables cannot hold. `note_usage_events` says
|
||||
whether a line was OPENED, and since every menu line carries its matched
|
||||
passage (#4364) "not opened" covers an unrelated line, a line whose passage
|
||||
was enough, and a line the reader already had. Only someone reading the
|
||||
query beside the line can tell those apart, so this records that reading —
|
||||
with its reason, because a verdict nobody can re-read is a threshold.
|
||||
|
||||
Keyed to the `retrieval_logs` row it judges, by id and FK-free like every
|
||||
telemetry table here. `query` is COPIED rather than joined: the verdict is
|
||||
only re-readable beside the words it was judged against.
|
||||
|
||||
`rank` and `score` are the candidate's place when the logged query was
|
||||
re-run for review, which may run past the arm's budget on purpose — the
|
||||
candidates just under the cut are the ones a budget change would add.
|
||||
`within_budget` says which side of the cut the line was.
|
||||
|
||||
`opened_after` is whether the reviewer's own agent pulled the record within
|
||||
an hour of the call, read from `note_usage_events`. A correlation, not a
|
||||
session join — the server has no session identity (see NoteUsageEvent) —
|
||||
and None when it was not measured.
|
||||
"""
|
||||
|
||||
__tablename__ = "retrieval_judgments"
|
||||
|
||||
id: Mapped[int] = mapped_column(BigInteger, primary_key=True)
|
||||
user_id: Mapped[int | None] = mapped_column(BigInteger, nullable=True)
|
||||
retrieval_log_id: Mapped[int] = mapped_column(BigInteger, nullable=False)
|
||||
# The judged call's surface, copied so the readout groups without a join.
|
||||
source: Mapped[str] = mapped_column(Text, nullable=False)
|
||||
query: Mapped[str | None] = mapped_column(Text, nullable=True)
|
||||
record_id: Mapped[int] = mapped_column(BigInteger, nullable=False)
|
||||
# 1-based, in the review's re-run.
|
||||
rank: Mapped[int] = mapped_column(Integer, nullable=False)
|
||||
score: Mapped[float | None] = mapped_column(Float, nullable=True)
|
||||
within_budget: Mapped[bool] = mapped_column(Boolean, nullable=False)
|
||||
# One of VERDICTS. Plain Text, no CHECK, like the usage tables; the
|
||||
# service refuses anything else.
|
||||
verdict: Mapped[str] = mapped_column(Text, nullable=False)
|
||||
reason: Mapped[str] = mapped_column(Text, nullable=False)
|
||||
opened_after: Mapped[bool | None] = mapped_column(Boolean, nullable=True)
|
||||
|
||||
__table_args__ = (
|
||||
# One verdict per reviewer per line; judging it again replaces it.
|
||||
UniqueConstraint(
|
||||
"retrieval_log_id", "record_id", "user_id",
|
||||
name="uq_retrieval_judgment_line",
|
||||
),
|
||||
Index("ix_retrieval_judgments_source_created", "source", "created_at"),
|
||||
)
|
||||
|
||||
def to_dict(self) -> dict:
|
||||
return {
|
||||
"id": self.id,
|
||||
"created_at": iso(self.created_at),
|
||||
"user_id": self.user_id,
|
||||
"retrieval_log_id": self.retrieval_log_id,
|
||||
"source": self.source,
|
||||
"query": self.query,
|
||||
"record_id": self.record_id,
|
||||
"rank": self.rank,
|
||||
"score": self.score,
|
||||
"within_budget": self.within_budget,
|
||||
"verdict": self.verdict,
|
||||
"reason": self.reason,
|
||||
"opened_after": self.opened_after,
|
||||
}
|
||||
Reference in New Issue
Block a user