fix(retrieval): the completion-report arm gets its own bar, not the prompt arm's (#3860)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m30s
CI & Build / Build & push image (push) Successful in 34s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m30s
CI & Build / Build & push image (push) Successful in 34s
`report_preference` shipped reading PROMPTRULE_THRESHOLD_KEY, so the two arms were one dial: tuning the bar for an operator's prose silently retuned the lookup that runs when a task closes. That coupling is worse on this arm than it would be anywhere else. Every other retrieval arm scores a query that varies per call, so a mis-set bar shows up as a changed clear-rate. COMPLETION_QUERY is a fixed string, so this arm's best score for a given corpus is a CONSTANT — and a constant sitting under the bar is a dead arm rather than a quiet one. No volume of traffic reveals it. Found by the first live read for milestone 394 step 9: 69 calls, 69 declines, every one naming the same record at the same score (0.7194 against a 0.72 bar). Reading `best_available_id` (#3807) showed the record was about interpreting a REQUEST, not about report shape — so the declines were correct and the arm is healthy. The percentile alone would have said "lower the bar", which would have delivered a false positive on every completion report ever written. The bar does not move; the key does. Both defaults stay 0.72, so this changes no behaviour on any install — it makes "leave this one where it is" expressible, which it was not before. Settings grows the control (rules 25, 27), and the default-agreement check grows a row. Deliberately NOT included: a change to PROMPTRULE_DEFAULT_THRESHOLD. The evidence for moving it is this install's near-miss table, and rule 115 keeps a shipped default from being justified by one instance's corpus. That bar is a per-user setting and belongs in the operator's Settings, not in the product. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
@@ -96,6 +96,7 @@ const kbToolRuleThreshold = ref("0.68");
|
|||||||
// And the prompt boundary is a third query shape again — the operator's own
|
// And the prompt boundary is a third query shape again — the operator's own
|
||||||
// prose rather than anything a tool produced (#3852).
|
// prose rather than anything a tool produced (#3852).
|
||||||
const kbPromptRuleThreshold = ref("0.72");
|
const kbPromptRuleThreshold = ref("0.72");
|
||||||
|
const kbReportPrefThreshold = ref("0.72");
|
||||||
// Near-duplicate report floors, one per record kind (services/dedup.py).
|
// Near-duplicate report floors, one per record kind (services/dedup.py).
|
||||||
// Snippets are single-chunk, so their floor sits below the 0.90 write-time
|
// Snippets are single-chunk, so their floor sits below the 0.90 write-time
|
||||||
// gate and catches what it lets through. Notes/tasks are scored at chunk
|
// gate and catches what it lets through. Notes/tasks are scored at chunk
|
||||||
@@ -171,6 +172,7 @@ async function saveKbInject() {
|
|||||||
// Bash call, so a fallback of 0 would put a rule in front of every command.
|
// Bash call, so a fallback of 0 would put a rule in front of every command.
|
||||||
const trT = Math.min(1, Math.max(0, Number(kbToolRuleThreshold.value) || 0.68));
|
const trT = Math.min(1, Math.max(0, Number(kbToolRuleThreshold.value) || 0.68));
|
||||||
const prT = Math.min(1, Math.max(0, Number(kbPromptRuleThreshold.value) || 0.72));
|
const prT = Math.min(1, Math.max(0, Number(kbPromptRuleThreshold.value) || 0.72));
|
||||||
|
const rpT = Math.min(1, Math.max(0, Number(kbReportPrefThreshold.value) || 0.72));
|
||||||
kbInjectThreshold.value = String(t);
|
kbInjectThreshold.value = String(t);
|
||||||
kbInjectTopK.value = String(k);
|
kbInjectTopK.value = String(k);
|
||||||
kbDupThresholdSnippet.value = String(dupSnip);
|
kbDupThresholdSnippet.value = String(dupSnip);
|
||||||
@@ -181,6 +183,7 @@ async function saveKbInject() {
|
|||||||
kbRuleHintThreshold.value = String(rhT);
|
kbRuleHintThreshold.value = String(rhT);
|
||||||
kbToolRuleThreshold.value = String(trT);
|
kbToolRuleThreshold.value = String(trT);
|
||||||
kbPromptRuleThreshold.value = String(prT);
|
kbPromptRuleThreshold.value = String(prT);
|
||||||
|
kbReportPrefThreshold.value = String(rpT);
|
||||||
savingKbInject.value = true;
|
savingKbInject.value = true;
|
||||||
kbInjectSaved.value = false;
|
kbInjectSaved.value = false;
|
||||||
try {
|
try {
|
||||||
@@ -202,6 +205,12 @@ async function saveKbInject() {
|
|||||||
// queries are different shapes. Moving one must not move the others.
|
// queries are different shapes. Moving one must not move the others.
|
||||||
kb_toolrule_threshold: String(trT),
|
kb_toolrule_threshold: String(trT),
|
||||||
kb_promptrule_threshold: String(prT),
|
kb_promptrule_threshold: String(prT),
|
||||||
|
// A SIXTH, and the one that most needed its own key: this arm's query
|
||||||
|
// is a fixed string, so its score is a constant for a given corpus.
|
||||||
|
// While it borrowed the prompt bar, tuning prose silently retuned it —
|
||||||
|
// and a constant that lands under the bar is a dead arm, not a quiet
|
||||||
|
// one (#3860).
|
||||||
|
kb_reportpref_threshold: String(rpT),
|
||||||
kb_duplicate_threshold_snippet: String(dupSnip),
|
kb_duplicate_threshold_snippet: String(dupSnip),
|
||||||
kb_duplicate_threshold_note: String(dupNote),
|
kb_duplicate_threshold_note: String(dupNote),
|
||||||
kb_duplicate_threshold_task: String(dupTask),
|
kb_duplicate_threshold_task: String(dupTask),
|
||||||
@@ -657,6 +666,9 @@ onMounted(async () => {
|
|||||||
if (allSettings.kb_promptrule_threshold !== undefined) {
|
if (allSettings.kb_promptrule_threshold !== undefined) {
|
||||||
kbPromptRuleThreshold.value = allSettings.kb_promptrule_threshold;
|
kbPromptRuleThreshold.value = allSettings.kb_promptrule_threshold;
|
||||||
}
|
}
|
||||||
|
if (allSettings.kb_reportpref_threshold !== undefined) {
|
||||||
|
kbReportPrefThreshold.value = allSettings.kb_reportpref_threshold;
|
||||||
|
}
|
||||||
if (allSettings.kb_writepath_threshold !== undefined) {
|
if (allSettings.kb_writepath_threshold !== undefined) {
|
||||||
kbWritePathThreshold.value = allSettings.kb_writepath_threshold;
|
kbWritePathThreshold.value = allSettings.kb_writepath_threshold;
|
||||||
}
|
}
|
||||||
@@ -1568,6 +1580,28 @@ async function deleteUser(userId: number) {
|
|||||||
not a command or a file — which is why it carries its own number.
|
not a command or a file — which is why it carries its own number.
|
||||||
</p>
|
</p>
|
||||||
</div>
|
</div>
|
||||||
|
<div class="field">
|
||||||
|
<label for="kb-reportpref-threshold">Completion-report confidence threshold (0–1)</label>
|
||||||
|
<input
|
||||||
|
id="kb-reportpref-threshold"
|
||||||
|
v-model="kbReportPrefThreshold"
|
||||||
|
type="number"
|
||||||
|
min="0"
|
||||||
|
max="1"
|
||||||
|
step="0.01"
|
||||||
|
class="fs-input input"
|
||||||
|
style="max-width: 8rem"
|
||||||
|
/>
|
||||||
|
<p class="field-hint">
|
||||||
|
The bar for a <em>preference about how a completion report should be
|
||||||
|
written</em>, looked up when a task closes. Unlike every other bar
|
||||||
|
here, the question this arm asks never changes — so its score is
|
||||||
|
fixed by your preferences alone, and it will either always find one
|
||||||
|
or never find one. If you have written a preference for report shape
|
||||||
|
and it is not arriving, lower this; there is no run of calls that
|
||||||
|
will reveal the problem on its own.
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
<!-- A design system belongs to a PROJECT, and the picker for it lives on
|
<!-- A design system belongs to a PROJECT, and the picker for it lives on
|
||||||
the project. There was a setting here that designated the system
|
the project. There was a setting here that designated the system
|
||||||
this install's own interface was built from; it only ever described
|
this install's own interface was built from; it only ever described
|
||||||
|
|||||||
@@ -471,6 +471,25 @@ PROMPTRULE_DEFAULT_THRESHOLD = 0.72
|
|||||||
# at all, where the act arms never face it because a command is one thing.
|
# at all, where the act arms never face it because a command is one thing.
|
||||||
PROMPTRULE_LIMIT = 3
|
PROMPTRULE_LIMIT = 3
|
||||||
|
|
||||||
|
# THE COMPLETION-REPORT ARM'S OWN BAR (services/reply_preferences.py).
|
||||||
|
#
|
||||||
|
# It borrowed PROMPTRULE_THRESHOLD_KEY when it shipped, which made the two
|
||||||
|
# arms one dial: an operator lowering the bar for their own prose moved this
|
||||||
|
# one with it, silently. That contradicts the rule every other bar here
|
||||||
|
# follows — one number cannot serve arms whose queries are different shapes —
|
||||||
|
# and this arm's query is the most different of all. The others score an
|
||||||
|
# operator's prose or a session's code, both of which vary per call; this one
|
||||||
|
# scores a FIXED string (`COMPLETION_QUERY`) against rule triggers, so its
|
||||||
|
# score for a given corpus is a constant. A constant that lands under the bar
|
||||||
|
# is not a quiet arm, it is a dead one, and nothing about the prose arm's
|
||||||
|
# traffic would ever reveal it.
|
||||||
|
#
|
||||||
|
# Kept at the prose arm's starting value rather than tuned: the split is what
|
||||||
|
# makes the two independently movable, and a default is a product decision
|
||||||
|
# that this install's corpus cannot settle (rule 115).
|
||||||
|
REPORTPREF_THRESHOLD_KEY = "kb_reportpref_threshold"
|
||||||
|
REPORTPREF_DEFAULT_THRESHOLD = 0.72
|
||||||
|
|
||||||
|
|
||||||
def _slugify(text: str) -> str:
|
def _slugify(text: str) -> str:
|
||||||
"""kebab-case slug for a skill directory name (a-z0-9 + single hyphens)."""
|
"""kebab-case slug for a skill directory name (a-z0-9 + single hyphens)."""
|
||||||
|
|||||||
@@ -34,14 +34,26 @@ delivery entirely in retrieval, as 394 decided, and leaves an operator nothing
|
|||||||
new to learn: a preference reaches the completion report the same way every
|
new to learn: a preference reaches the completion report the same way every
|
||||||
other record reaches its moment.
|
other record reaches its moment.
|
||||||
|
|
||||||
THE BAR IS THE PROMPT ARM'S, AND IT IS NOT YET EARNED HERE
|
THE BAR IS ITS OWN, AND THE FIRST READING EARNED IT (#3860)
|
||||||
|
|
||||||
A fixed query against triggers is a different score distribution from an
|
This borrowed the prompt arm's key when it shipped, on the argument that a
|
||||||
operator's message against the same documents. Starting at the prompt arm's
|
fixed query against triggers is a different score distribution from an
|
||||||
setting is the value with evidence behind it, and an operator's tuning of that
|
operator's message against the same documents — true, and the reason the two
|
||||||
bar reaches this too. Every call logs under its own source,
|
could not stay one dial. Five days of traffic settled it.
|
||||||
`report_preference`, so step 6 can read this surface's near misses apart from
|
|
||||||
the prompt arm's before anyone moves the number.
|
What the readout said: 69 calls, 69 declines, every one naming the SAME record
|
||||||
|
at the SAME score (rule 77 at 0.7194 against a 0.72 bar). That constancy is
|
||||||
|
the signature of this arm — `COMPLETION_QUERY` never varies, so for a given
|
||||||
|
corpus its best score is a constant, and a constant sitting under the bar is a
|
||||||
|
dead arm rather than a quiet one. The record it kept declining was about
|
||||||
|
reading a REQUEST, not about the shape of a report, so the decline was right
|
||||||
|
and the arm is healthy: this install simply has no completion-report
|
||||||
|
preference on file.
|
||||||
|
|
||||||
|
The bar stayed at 0.72, and the key moved out (REPORTPREF_THRESHOLD_KEY) so
|
||||||
|
that staying is a decision rather than a side effect of what the prose arm is
|
||||||
|
set to. A surface whose score cannot vary is the one surface where a borrowed
|
||||||
|
bar can be wrong forever without a single call looking unusual.
|
||||||
"""
|
"""
|
||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
@@ -50,8 +62,8 @@ import time
|
|||||||
|
|
||||||
from scribe.services.embeddings import semantic_search_rules
|
from scribe.services.embeddings import semantic_search_rules
|
||||||
from scribe.services.plugin_context import (
|
from scribe.services.plugin_context import (
|
||||||
PROMPTRULE_DEFAULT_THRESHOLD,
|
REPORTPREF_DEFAULT_THRESHOLD,
|
||||||
PROMPTRULE_THRESHOLD_KEY,
|
REPORTPREF_THRESHOLD_KEY,
|
||||||
)
|
)
|
||||||
from scribe.services.retrieval_telemetry import record_retrieval
|
from scribe.services.retrieval_telemetry import record_retrieval
|
||||||
from scribe.services.rule_usage import record_rule_surfaced
|
from scribe.services.rule_usage import record_rule_surfaced
|
||||||
@@ -79,9 +91,9 @@ LIMIT = 3
|
|||||||
async def _threshold(user_id: int) -> float:
|
async def _threshold(user_id: int) -> float:
|
||||||
try:
|
try:
|
||||||
value = float(await get_setting(
|
value = float(await get_setting(
|
||||||
user_id, PROMPTRULE_THRESHOLD_KEY, str(PROMPTRULE_DEFAULT_THRESHOLD)))
|
user_id, REPORTPREF_THRESHOLD_KEY, str(REPORTPREF_DEFAULT_THRESHOLD)))
|
||||||
except (TypeError, ValueError):
|
except (TypeError, ValueError):
|
||||||
value = PROMPTRULE_DEFAULT_THRESHOLD
|
value = REPORTPREF_DEFAULT_THRESHOLD
|
||||||
return min(1.0, max(0.0, value))
|
return min(1.0, max(0.0, value))
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
@@ -73,6 +73,46 @@ def test_it_is_a_ranked_source():
|
|||||||
assert not is_ambient("report_preference")
|
assert not is_ambient("report_preference")
|
||||||
|
|
||||||
|
|
||||||
|
def test_its_bar_is_its_own_key_not_the_prompt_arm_s(monkeypatch):
|
||||||
|
"""Regression on the coupling #3860 found (and on the fix being real).
|
||||||
|
|
||||||
|
This arm shipped reading PROMPTRULE_THRESHOLD_KEY, so an operator tuning
|
||||||
|
the bar for their own prose moved this one with it and was never told.
|
||||||
|
That is worse here than anywhere else: every other arm scores a query that
|
||||||
|
varies per call, while COMPLETION_QUERY is fixed — so this arm's score for
|
||||||
|
a given corpus is a CONSTANT, and a constant that lands under the bar is a
|
||||||
|
dead arm rather than a quiet one. No amount of traffic reveals it.
|
||||||
|
|
||||||
|
Asserted on the key the lookup actually asks for, which is the thing that
|
||||||
|
broke, rather than on the constant being defined somewhere.
|
||||||
|
"""
|
||||||
|
import asyncio
|
||||||
|
|
||||||
|
from scribe.services import plugin_context
|
||||||
|
from scribe.services.reply_preferences import completion_preferences
|
||||||
|
|
||||||
|
asked: list[str] = []
|
||||||
|
|
||||||
|
async def get_setting(user_id, key, default):
|
||||||
|
asked.append(key)
|
||||||
|
return default
|
||||||
|
|
||||||
|
async def search(user_id, query, **kw):
|
||||||
|
kw["report"].update({"searched": True})
|
||||||
|
return []
|
||||||
|
|
||||||
|
with patch(f"{M}.get_setting", AsyncMock(side_effect=get_setting)), \
|
||||||
|
patch(f"{M}.semantic_search_rules", AsyncMock(side_effect=search)), \
|
||||||
|
patch(f"{M}.record_retrieval"):
|
||||||
|
asyncio.run(completion_preferences(7))
|
||||||
|
|
||||||
|
assert asked == [plugin_context.REPORTPREF_THRESHOLD_KEY]
|
||||||
|
assert plugin_context.REPORTPREF_THRESHOLD_KEY != plugin_context.PROMPTRULE_THRESHOLD_KEY, (
|
||||||
|
"the two keys are the same string again, so the settings form has one "
|
||||||
|
"dial driving two arms — which is the defect, whatever the value is"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
def test_the_query_assumes_no_particular_domain():
|
def test_the_query_assumes_no_particular_domain():
|
||||||
from scribe.services.reply_preferences import COMPLETION_QUERY
|
from scribe.services.reply_preferences import COMPLETION_QUERY
|
||||||
|
|
||||||
|
|||||||
@@ -50,6 +50,7 @@ _PAIRS = (
|
|||||||
("plugin_context.py", "RULEHINT_DEFAULT_THRESHOLD", "kbRuleHintThreshold"),
|
("plugin_context.py", "RULEHINT_DEFAULT_THRESHOLD", "kbRuleHintThreshold"),
|
||||||
("plugin_context.py", "TOOLRULE_DEFAULT_THRESHOLD", "kbToolRuleThreshold"),
|
("plugin_context.py", "TOOLRULE_DEFAULT_THRESHOLD", "kbToolRuleThreshold"),
|
||||||
("plugin_context.py", "PROMPTRULE_DEFAULT_THRESHOLD", "kbPromptRuleThreshold"),
|
("plugin_context.py", "PROMPTRULE_DEFAULT_THRESHOLD", "kbPromptRuleThreshold"),
|
||||||
|
("plugin_context.py", "REPORTPREF_DEFAULT_THRESHOLD", "kbReportPrefThreshold"),
|
||||||
# The plan gate (milestone 415): it blocks a create, so a form showing a
|
# The plan gate (milestone 415): it blocks a create, so a form showing a
|
||||||
# looser bar than the one in force would be the more misleading drift.
|
# looser bar than the one in force would be the more misleading drift.
|
||||||
("dedup.py", "PLAN_MATCH_DEFAULT_THRESHOLD", "kbPlanMatchThreshold"),
|
("dedup.py", "PLAN_MATCH_DEFAULT_THRESHOLD", "kbPlanMatchThreshold"),
|
||||||
|
|||||||
Reference in New Issue
Block a user