fix(retrieval): the completion-report arm gets its own bar, not the prompt arm's (#3860)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m30s
CI & Build / Build & push image (push) Successful in 34s

`report_preference` shipped reading PROMPTRULE_THRESHOLD_KEY, so the two arms
were one dial: tuning the bar for an operator's prose silently retuned the
lookup that runs when a task closes.

That coupling is worse on this arm than it would be anywhere else. Every other
retrieval arm scores a query that varies per call, so a mis-set bar shows up as
a changed clear-rate. COMPLETION_QUERY is a fixed string, so this arm's best
score for a given corpus is a CONSTANT — and a constant sitting under the bar is
a dead arm rather than a quiet one. No volume of traffic reveals it.

Found by the first live read for milestone 394 step 9: 69 calls, 69 declines,
every one naming the same record at the same score (0.7194 against a 0.72 bar).
Reading `best_available_id` (#3807) showed the record was about interpreting a
REQUEST, not about report shape — so the declines were correct and the arm is
healthy. The percentile alone would have said "lower the bar", which would have
delivered a false positive on every completion report ever written.

The bar does not move; the key does. Both defaults stay 0.72, so this changes no
behaviour on any install — it makes "leave this one where it is" expressible,
which it was not before. Settings grows the control (rules 25, 27), and the
default-agreement check grows a row.

Deliberately NOT included: a change to PROMPTRULE_DEFAULT_THRESHOLD. The
evidence for moving it is this install's near-miss table, and rule 115 keeps a
shipped default from being justified by one instance's corpus. That bar is a
per-user setting and belongs in the operator's Settings, not in the product.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
2026-09-16 13:47:34 -04:00
co-authored by Claude Opus 5
parent dd08f9858c
commit 5ae60734bb
5 changed files with 117 additions and 11 deletions
+19
View File
@@ -471,6 +471,25 @@ PROMPTRULE_DEFAULT_THRESHOLD = 0.72
# at all, where the act arms never face it because a command is one thing.
PROMPTRULE_LIMIT = 3
# THE COMPLETION-REPORT ARM'S OWN BAR (services/reply_preferences.py).
#
# It borrowed PROMPTRULE_THRESHOLD_KEY when it shipped, which made the two
# arms one dial: an operator lowering the bar for their own prose moved this
# one with it, silently. That contradicts the rule every other bar here
# follows — one number cannot serve arms whose queries are different shapes —
# and this arm's query is the most different of all. The others score an
# operator's prose or a session's code, both of which vary per call; this one
# scores a FIXED string (`COMPLETION_QUERY`) against rule triggers, so its
# score for a given corpus is a constant. A constant that lands under the bar
# is not a quiet arm, it is a dead one, and nothing about the prose arm's
# traffic would ever reveal it.
#
# Kept at the prose arm's starting value rather than tuned: the split is what
# makes the two independently movable, and a default is a product decision
# that this install's corpus cannot settle (rule 115).
REPORTPREF_THRESHOLD_KEY = "kb_reportpref_threshold"
REPORTPREF_DEFAULT_THRESHOLD = 0.72
def _slugify(text: str) -> str:
"""kebab-case slug for a skill directory name (a-z0-9 + single hyphens)."""
+23 -11
View File
@@ -34,14 +34,26 @@ delivery entirely in retrieval, as 394 decided, and leaves an operator nothing
new to learn: a preference reaches the completion report the same way every
other record reaches its moment.
THE BAR IS THE PROMPT ARM'S, AND IT IS NOT YET EARNED HERE
THE BAR IS ITS OWN, AND THE FIRST READING EARNED IT (#3860)
A fixed query against triggers is a different score distribution from an
operator's message against the same documents. Starting at the prompt arm's
setting is the value with evidence behind it, and an operator's tuning of that
bar reaches this too. Every call logs under its own source,
`report_preference`, so step 6 can read this surface's near misses apart from
the prompt arm's before anyone moves the number.
This borrowed the prompt arm's key when it shipped, on the argument that a
fixed query against triggers is a different score distribution from an
operator's message against the same documents — true, and the reason the two
could not stay one dial. Five days of traffic settled it.
What the readout said: 69 calls, 69 declines, every one naming the SAME record
at the SAME score (rule 77 at 0.7194 against a 0.72 bar). That constancy is
the signature of this arm — `COMPLETION_QUERY` never varies, so for a given
corpus its best score is a constant, and a constant sitting under the bar is a
dead arm rather than a quiet one. The record it kept declining was about
reading a REQUEST, not about the shape of a report, so the decline was right
and the arm is healthy: this install simply has no completion-report
preference on file.
The bar stayed at 0.72, and the key moved out (REPORTPREF_THRESHOLD_KEY) so
that staying is a decision rather than a side effect of what the prose arm is
set to. A surface whose score cannot vary is the one surface where a borrowed
bar can be wrong forever without a single call looking unusual.
"""
from __future__ import annotations
@@ -50,8 +62,8 @@ import time
from scribe.services.embeddings import semantic_search_rules
from scribe.services.plugin_context import (
PROMPTRULE_DEFAULT_THRESHOLD,
PROMPTRULE_THRESHOLD_KEY,
REPORTPREF_DEFAULT_THRESHOLD,
REPORTPREF_THRESHOLD_KEY,
)
from scribe.services.retrieval_telemetry import record_retrieval
from scribe.services.rule_usage import record_rule_surfaced
@@ -79,9 +91,9 @@ LIMIT = 3
async def _threshold(user_id: int) -> float:
try:
value = float(await get_setting(
user_id, PROMPTRULE_THRESHOLD_KEY, str(PROMPTRULE_DEFAULT_THRESHOLD)))
user_id, REPORTPREF_THRESHOLD_KEY, str(REPORTPREF_DEFAULT_THRESHOLD)))
except (TypeError, ValueError):
value = PROMPTRULE_DEFAULT_THRESHOLD
value = REPORTPREF_DEFAULT_THRESHOLD
return min(1.0, max(0.0, value))