fix(retrieval): the completion-report arm gets its own bar, not the prompt arm's (#3860)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m30s
CI & Build / Build & push image (push) Successful in 34s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m30s
CI & Build / Build & push image (push) Successful in 34s
`report_preference` shipped reading PROMPTRULE_THRESHOLD_KEY, so the two arms were one dial: tuning the bar for an operator's prose silently retuned the lookup that runs when a task closes. That coupling is worse on this arm than it would be anywhere else. Every other retrieval arm scores a query that varies per call, so a mis-set bar shows up as a changed clear-rate. COMPLETION_QUERY is a fixed string, so this arm's best score for a given corpus is a CONSTANT — and a constant sitting under the bar is a dead arm rather than a quiet one. No volume of traffic reveals it. Found by the first live read for milestone 394 step 9: 69 calls, 69 declines, every one naming the same record at the same score (0.7194 against a 0.72 bar). Reading `best_available_id` (#3807) showed the record was about interpreting a REQUEST, not about report shape — so the declines were correct and the arm is healthy. The percentile alone would have said "lower the bar", which would have delivered a false positive on every completion report ever written. The bar does not move; the key does. Both defaults stay 0.72, so this changes no behaviour on any install — it makes "leave this one where it is" expressible, which it was not before. Settings grows the control (rules 25, 27), and the default-agreement check grows a row. Deliberately NOT included: a change to PROMPTRULE_DEFAULT_THRESHOLD. The evidence for moving it is this install's near-miss table, and rule 115 keeps a shipped default from being justified by one instance's corpus. That bar is a per-user setting and belongs in the operator's Settings, not in the product. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
@@ -96,6 +96,7 @@ const kbToolRuleThreshold = ref("0.68");
|
||||
// And the prompt boundary is a third query shape again — the operator's own
|
||||
// prose rather than anything a tool produced (#3852).
|
||||
const kbPromptRuleThreshold = ref("0.72");
|
||||
const kbReportPrefThreshold = ref("0.72");
|
||||
// Near-duplicate report floors, one per record kind (services/dedup.py).
|
||||
// Snippets are single-chunk, so their floor sits below the 0.90 write-time
|
||||
// gate and catches what it lets through. Notes/tasks are scored at chunk
|
||||
@@ -171,6 +172,7 @@ async function saveKbInject() {
|
||||
// Bash call, so a fallback of 0 would put a rule in front of every command.
|
||||
const trT = Math.min(1, Math.max(0, Number(kbToolRuleThreshold.value) || 0.68));
|
||||
const prT = Math.min(1, Math.max(0, Number(kbPromptRuleThreshold.value) || 0.72));
|
||||
const rpT = Math.min(1, Math.max(0, Number(kbReportPrefThreshold.value) || 0.72));
|
||||
kbInjectThreshold.value = String(t);
|
||||
kbInjectTopK.value = String(k);
|
||||
kbDupThresholdSnippet.value = String(dupSnip);
|
||||
@@ -181,6 +183,7 @@ async function saveKbInject() {
|
||||
kbRuleHintThreshold.value = String(rhT);
|
||||
kbToolRuleThreshold.value = String(trT);
|
||||
kbPromptRuleThreshold.value = String(prT);
|
||||
kbReportPrefThreshold.value = String(rpT);
|
||||
savingKbInject.value = true;
|
||||
kbInjectSaved.value = false;
|
||||
try {
|
||||
@@ -202,6 +205,12 @@ async function saveKbInject() {
|
||||
// queries are different shapes. Moving one must not move the others.
|
||||
kb_toolrule_threshold: String(trT),
|
||||
kb_promptrule_threshold: String(prT),
|
||||
// A SIXTH, and the one that most needed its own key: this arm's query
|
||||
// is a fixed string, so its score is a constant for a given corpus.
|
||||
// While it borrowed the prompt bar, tuning prose silently retuned it —
|
||||
// and a constant that lands under the bar is a dead arm, not a quiet
|
||||
// one (#3860).
|
||||
kb_reportpref_threshold: String(rpT),
|
||||
kb_duplicate_threshold_snippet: String(dupSnip),
|
||||
kb_duplicate_threshold_note: String(dupNote),
|
||||
kb_duplicate_threshold_task: String(dupTask),
|
||||
@@ -657,6 +666,9 @@ onMounted(async () => {
|
||||
if (allSettings.kb_promptrule_threshold !== undefined) {
|
||||
kbPromptRuleThreshold.value = allSettings.kb_promptrule_threshold;
|
||||
}
|
||||
if (allSettings.kb_reportpref_threshold !== undefined) {
|
||||
kbReportPrefThreshold.value = allSettings.kb_reportpref_threshold;
|
||||
}
|
||||
if (allSettings.kb_writepath_threshold !== undefined) {
|
||||
kbWritePathThreshold.value = allSettings.kb_writepath_threshold;
|
||||
}
|
||||
@@ -1568,6 +1580,28 @@ async function deleteUser(userId: number) {
|
||||
not a command or a file — which is why it carries its own number.
|
||||
</p>
|
||||
</div>
|
||||
<div class="field">
|
||||
<label for="kb-reportpref-threshold">Completion-report confidence threshold (0–1)</label>
|
||||
<input
|
||||
id="kb-reportpref-threshold"
|
||||
v-model="kbReportPrefThreshold"
|
||||
type="number"
|
||||
min="0"
|
||||
max="1"
|
||||
step="0.01"
|
||||
class="fs-input input"
|
||||
style="max-width: 8rem"
|
||||
/>
|
||||
<p class="field-hint">
|
||||
The bar for a <em>preference about how a completion report should be
|
||||
written</em>, looked up when a task closes. Unlike every other bar
|
||||
here, the question this arm asks never changes — so its score is
|
||||
fixed by your preferences alone, and it will either always find one
|
||||
or never find one. If you have written a preference for report shape
|
||||
and it is not arriving, lower this; there is no run of calls that
|
||||
will reveal the problem on its own.
|
||||
</p>
|
||||
</div>
|
||||
<!-- A design system belongs to a PROJECT, and the picker for it lives on
|
||||
the project. There was a setting here that designated the system
|
||||
this install's own interface was built from; it only ever described
|
||||
|
||||
@@ -471,6 +471,25 @@ PROMPTRULE_DEFAULT_THRESHOLD = 0.72
|
||||
# at all, where the act arms never face it because a command is one thing.
|
||||
PROMPTRULE_LIMIT = 3
|
||||
|
||||
# THE COMPLETION-REPORT ARM'S OWN BAR (services/reply_preferences.py).
|
||||
#
|
||||
# It borrowed PROMPTRULE_THRESHOLD_KEY when it shipped, which made the two
|
||||
# arms one dial: an operator lowering the bar for their own prose moved this
|
||||
# one with it, silently. That contradicts the rule every other bar here
|
||||
# follows — one number cannot serve arms whose queries are different shapes —
|
||||
# and this arm's query is the most different of all. The others score an
|
||||
# operator's prose or a session's code, both of which vary per call; this one
|
||||
# scores a FIXED string (`COMPLETION_QUERY`) against rule triggers, so its
|
||||
# score for a given corpus is a constant. A constant that lands under the bar
|
||||
# is not a quiet arm, it is a dead one, and nothing about the prose arm's
|
||||
# traffic would ever reveal it.
|
||||
#
|
||||
# Kept at the prose arm's starting value rather than tuned: the split is what
|
||||
# makes the two independently movable, and a default is a product decision
|
||||
# that this install's corpus cannot settle (rule 115).
|
||||
REPORTPREF_THRESHOLD_KEY = "kb_reportpref_threshold"
|
||||
REPORTPREF_DEFAULT_THRESHOLD = 0.72
|
||||
|
||||
|
||||
def _slugify(text: str) -> str:
|
||||
"""kebab-case slug for a skill directory name (a-z0-9 + single hyphens)."""
|
||||
|
||||
@@ -34,14 +34,26 @@ delivery entirely in retrieval, as 394 decided, and leaves an operator nothing
|
||||
new to learn: a preference reaches the completion report the same way every
|
||||
other record reaches its moment.
|
||||
|
||||
THE BAR IS THE PROMPT ARM'S, AND IT IS NOT YET EARNED HERE
|
||||
THE BAR IS ITS OWN, AND THE FIRST READING EARNED IT (#3860)
|
||||
|
||||
A fixed query against triggers is a different score distribution from an
|
||||
operator's message against the same documents. Starting at the prompt arm's
|
||||
setting is the value with evidence behind it, and an operator's tuning of that
|
||||
bar reaches this too. Every call logs under its own source,
|
||||
`report_preference`, so step 6 can read this surface's near misses apart from
|
||||
the prompt arm's before anyone moves the number.
|
||||
This borrowed the prompt arm's key when it shipped, on the argument that a
|
||||
fixed query against triggers is a different score distribution from an
|
||||
operator's message against the same documents — true, and the reason the two
|
||||
could not stay one dial. Five days of traffic settled it.
|
||||
|
||||
What the readout said: 69 calls, 69 declines, every one naming the SAME record
|
||||
at the SAME score (rule 77 at 0.7194 against a 0.72 bar). That constancy is
|
||||
the signature of this arm — `COMPLETION_QUERY` never varies, so for a given
|
||||
corpus its best score is a constant, and a constant sitting under the bar is a
|
||||
dead arm rather than a quiet one. The record it kept declining was about
|
||||
reading a REQUEST, not about the shape of a report, so the decline was right
|
||||
and the arm is healthy: this install simply has no completion-report
|
||||
preference on file.
|
||||
|
||||
The bar stayed at 0.72, and the key moved out (REPORTPREF_THRESHOLD_KEY) so
|
||||
that staying is a decision rather than a side effect of what the prose arm is
|
||||
set to. A surface whose score cannot vary is the one surface where a borrowed
|
||||
bar can be wrong forever without a single call looking unusual.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
@@ -50,8 +62,8 @@ import time
|
||||
|
||||
from scribe.services.embeddings import semantic_search_rules
|
||||
from scribe.services.plugin_context import (
|
||||
PROMPTRULE_DEFAULT_THRESHOLD,
|
||||
PROMPTRULE_THRESHOLD_KEY,
|
||||
REPORTPREF_DEFAULT_THRESHOLD,
|
||||
REPORTPREF_THRESHOLD_KEY,
|
||||
)
|
||||
from scribe.services.retrieval_telemetry import record_retrieval
|
||||
from scribe.services.rule_usage import record_rule_surfaced
|
||||
@@ -79,9 +91,9 @@ LIMIT = 3
|
||||
async def _threshold(user_id: int) -> float:
|
||||
try:
|
||||
value = float(await get_setting(
|
||||
user_id, PROMPTRULE_THRESHOLD_KEY, str(PROMPTRULE_DEFAULT_THRESHOLD)))
|
||||
user_id, REPORTPREF_THRESHOLD_KEY, str(REPORTPREF_DEFAULT_THRESHOLD)))
|
||||
except (TypeError, ValueError):
|
||||
value = PROMPTRULE_DEFAULT_THRESHOLD
|
||||
value = REPORTPREF_DEFAULT_THRESHOLD
|
||||
return min(1.0, max(0.0, value))
|
||||
|
||||
|
||||
|
||||
@@ -73,6 +73,46 @@ def test_it_is_a_ranked_source():
|
||||
assert not is_ambient("report_preference")
|
||||
|
||||
|
||||
def test_its_bar_is_its_own_key_not_the_prompt_arm_s(monkeypatch):
|
||||
"""Regression on the coupling #3860 found (and on the fix being real).
|
||||
|
||||
This arm shipped reading PROMPTRULE_THRESHOLD_KEY, so an operator tuning
|
||||
the bar for their own prose moved this one with it and was never told.
|
||||
That is worse here than anywhere else: every other arm scores a query that
|
||||
varies per call, while COMPLETION_QUERY is fixed — so this arm's score for
|
||||
a given corpus is a CONSTANT, and a constant that lands under the bar is a
|
||||
dead arm rather than a quiet one. No amount of traffic reveals it.
|
||||
|
||||
Asserted on the key the lookup actually asks for, which is the thing that
|
||||
broke, rather than on the constant being defined somewhere.
|
||||
"""
|
||||
import asyncio
|
||||
|
||||
from scribe.services import plugin_context
|
||||
from scribe.services.reply_preferences import completion_preferences
|
||||
|
||||
asked: list[str] = []
|
||||
|
||||
async def get_setting(user_id, key, default):
|
||||
asked.append(key)
|
||||
return default
|
||||
|
||||
async def search(user_id, query, **kw):
|
||||
kw["report"].update({"searched": True})
|
||||
return []
|
||||
|
||||
with patch(f"{M}.get_setting", AsyncMock(side_effect=get_setting)), \
|
||||
patch(f"{M}.semantic_search_rules", AsyncMock(side_effect=search)), \
|
||||
patch(f"{M}.record_retrieval"):
|
||||
asyncio.run(completion_preferences(7))
|
||||
|
||||
assert asked == [plugin_context.REPORTPREF_THRESHOLD_KEY]
|
||||
assert plugin_context.REPORTPREF_THRESHOLD_KEY != plugin_context.PROMPTRULE_THRESHOLD_KEY, (
|
||||
"the two keys are the same string again, so the settings form has one "
|
||||
"dial driving two arms — which is the defect, whatever the value is"
|
||||
)
|
||||
|
||||
|
||||
def test_the_query_assumes_no_particular_domain():
|
||||
from scribe.services.reply_preferences import COMPLETION_QUERY
|
||||
|
||||
|
||||
@@ -50,6 +50,7 @@ _PAIRS = (
|
||||
("plugin_context.py", "RULEHINT_DEFAULT_THRESHOLD", "kbRuleHintThreshold"),
|
||||
("plugin_context.py", "TOOLRULE_DEFAULT_THRESHOLD", "kbToolRuleThreshold"),
|
||||
("plugin_context.py", "PROMPTRULE_DEFAULT_THRESHOLD", "kbPromptRuleThreshold"),
|
||||
("plugin_context.py", "REPORTPREF_DEFAULT_THRESHOLD", "kbReportPrefThreshold"),
|
||||
# The plan gate (milestone 415): it blocks a create, so a form showing a
|
||||
# looser bar than the one in force would be the more misleading drift.
|
||||
("dedup.py", "PLAN_MATCH_DEFAULT_THRESHOLD", "kbPlanMatchThreshold"),
|
||||
|
||||
Reference in New Issue
Block a user