fix(retrieval): the completion-report arm gets its own bar, not the prompt arm's (#3860)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m30s
CI & Build / Build & push image (push) Successful in 34s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m30s
CI & Build / Build & push image (push) Successful in 34s
`report_preference` shipped reading PROMPTRULE_THRESHOLD_KEY, so the two arms were one dial: tuning the bar for an operator's prose silently retuned the lookup that runs when a task closes. That coupling is worse on this arm than it would be anywhere else. Every other retrieval arm scores a query that varies per call, so a mis-set bar shows up as a changed clear-rate. COMPLETION_QUERY is a fixed string, so this arm's best score for a given corpus is a CONSTANT — and a constant sitting under the bar is a dead arm rather than a quiet one. No volume of traffic reveals it. Found by the first live read for milestone 394 step 9: 69 calls, 69 declines, every one naming the same record at the same score (0.7194 against a 0.72 bar). Reading `best_available_id` (#3807) showed the record was about interpreting a REQUEST, not about report shape — so the declines were correct and the arm is healthy. The percentile alone would have said "lower the bar", which would have delivered a false positive on every completion report ever written. The bar does not move; the key does. Both defaults stay 0.72, so this changes no behaviour on any install — it makes "leave this one where it is" expressible, which it was not before. Settings grows the control (rules 25, 27), and the default-agreement check grows a row. Deliberately NOT included: a change to PROMPTRULE_DEFAULT_THRESHOLD. The evidence for moving it is this install's near-miss table, and rule 115 keeps a shipped default from being justified by one instance's corpus. That bar is a per-user setting and belongs in the operator's Settings, not in the product. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
@@ -471,6 +471,25 @@ PROMPTRULE_DEFAULT_THRESHOLD = 0.72
|
||||
# at all, where the act arms never face it because a command is one thing.
|
||||
PROMPTRULE_LIMIT = 3
|
||||
|
||||
# THE COMPLETION-REPORT ARM'S OWN BAR (services/reply_preferences.py).
|
||||
#
|
||||
# It borrowed PROMPTRULE_THRESHOLD_KEY when it shipped, which made the two
|
||||
# arms one dial: an operator lowering the bar for their own prose moved this
|
||||
# one with it, silently. That contradicts the rule every other bar here
|
||||
# follows — one number cannot serve arms whose queries are different shapes —
|
||||
# and this arm's query is the most different of all. The others score an
|
||||
# operator's prose or a session's code, both of which vary per call; this one
|
||||
# scores a FIXED string (`COMPLETION_QUERY`) against rule triggers, so its
|
||||
# score for a given corpus is a constant. A constant that lands under the bar
|
||||
# is not a quiet arm, it is a dead one, and nothing about the prose arm's
|
||||
# traffic would ever reveal it.
|
||||
#
|
||||
# Kept at the prose arm's starting value rather than tuned: the split is what
|
||||
# makes the two independently movable, and a default is a product decision
|
||||
# that this install's corpus cannot settle (rule 115).
|
||||
REPORTPREF_THRESHOLD_KEY = "kb_reportpref_threshold"
|
||||
REPORTPREF_DEFAULT_THRESHOLD = 0.72
|
||||
|
||||
|
||||
def _slugify(text: str) -> str:
|
||||
"""kebab-case slug for a skill directory name (a-z0-9 + single hyphens)."""
|
||||
|
||||
@@ -34,14 +34,26 @@ delivery entirely in retrieval, as 394 decided, and leaves an operator nothing
|
||||
new to learn: a preference reaches the completion report the same way every
|
||||
other record reaches its moment.
|
||||
|
||||
THE BAR IS THE PROMPT ARM'S, AND IT IS NOT YET EARNED HERE
|
||||
THE BAR IS ITS OWN, AND THE FIRST READING EARNED IT (#3860)
|
||||
|
||||
A fixed query against triggers is a different score distribution from an
|
||||
operator's message against the same documents. Starting at the prompt arm's
|
||||
setting is the value with evidence behind it, and an operator's tuning of that
|
||||
bar reaches this too. Every call logs under its own source,
|
||||
`report_preference`, so step 6 can read this surface's near misses apart from
|
||||
the prompt arm's before anyone moves the number.
|
||||
This borrowed the prompt arm's key when it shipped, on the argument that a
|
||||
fixed query against triggers is a different score distribution from an
|
||||
operator's message against the same documents — true, and the reason the two
|
||||
could not stay one dial. Five days of traffic settled it.
|
||||
|
||||
What the readout said: 69 calls, 69 declines, every one naming the SAME record
|
||||
at the SAME score (rule 77 at 0.7194 against a 0.72 bar). That constancy is
|
||||
the signature of this arm — `COMPLETION_QUERY` never varies, so for a given
|
||||
corpus its best score is a constant, and a constant sitting under the bar is a
|
||||
dead arm rather than a quiet one. The record it kept declining was about
|
||||
reading a REQUEST, not about the shape of a report, so the decline was right
|
||||
and the arm is healthy: this install simply has no completion-report
|
||||
preference on file.
|
||||
|
||||
The bar stayed at 0.72, and the key moved out (REPORTPREF_THRESHOLD_KEY) so
|
||||
that staying is a decision rather than a side effect of what the prose arm is
|
||||
set to. A surface whose score cannot vary is the one surface where a borrowed
|
||||
bar can be wrong forever without a single call looking unusual.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
@@ -50,8 +62,8 @@ import time
|
||||
|
||||
from scribe.services.embeddings import semantic_search_rules
|
||||
from scribe.services.plugin_context import (
|
||||
PROMPTRULE_DEFAULT_THRESHOLD,
|
||||
PROMPTRULE_THRESHOLD_KEY,
|
||||
REPORTPREF_DEFAULT_THRESHOLD,
|
||||
REPORTPREF_THRESHOLD_KEY,
|
||||
)
|
||||
from scribe.services.retrieval_telemetry import record_retrieval
|
||||
from scribe.services.rule_usage import record_rule_surfaced
|
||||
@@ -79,9 +91,9 @@ LIMIT = 3
|
||||
async def _threshold(user_id: int) -> float:
|
||||
try:
|
||||
value = float(await get_setting(
|
||||
user_id, PROMPTRULE_THRESHOLD_KEY, str(PROMPTRULE_DEFAULT_THRESHOLD)))
|
||||
user_id, REPORTPREF_THRESHOLD_KEY, str(REPORTPREF_DEFAULT_THRESHOLD)))
|
||||
except (TypeError, ValueError):
|
||||
value = PROMPTRULE_DEFAULT_THRESHOLD
|
||||
value = REPORTPREF_DEFAULT_THRESHOLD
|
||||
return min(1.0, max(0.0, value))
|
||||
|
||||
|
||||
|
||||
Reference in New Issue
Block a user