fix(retrieval): the completion-report arm gets its own bar, not the prompt arm's (#3860)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m30s
CI & Build / Build & push image (push) Successful in 34s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m30s
CI & Build / Build & push image (push) Successful in 34s
`report_preference` shipped reading PROMPTRULE_THRESHOLD_KEY, so the two arms were one dial: tuning the bar for an operator's prose silently retuned the lookup that runs when a task closes. That coupling is worse on this arm than it would be anywhere else. Every other retrieval arm scores a query that varies per call, so a mis-set bar shows up as a changed clear-rate. COMPLETION_QUERY is a fixed string, so this arm's best score for a given corpus is a CONSTANT — and a constant sitting under the bar is a dead arm rather than a quiet one. No volume of traffic reveals it. Found by the first live read for milestone 394 step 9: 69 calls, 69 declines, every one naming the same record at the same score (0.7194 against a 0.72 bar). Reading `best_available_id` (#3807) showed the record was about interpreting a REQUEST, not about report shape — so the declines were correct and the arm is healthy. The percentile alone would have said "lower the bar", which would have delivered a false positive on every completion report ever written. The bar does not move; the key does. Both defaults stay 0.72, so this changes no behaviour on any install — it makes "leave this one where it is" expressible, which it was not before. Settings grows the control (rules 25, 27), and the default-agreement check grows a row. Deliberately NOT included: a change to PROMPTRULE_DEFAULT_THRESHOLD. The evidence for moving it is this install's near-miss table, and rule 115 keeps a shipped default from being justified by one instance's corpus. That bar is a per-user setting and belongs in the operator's Settings, not in the product. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
@@ -73,6 +73,46 @@ def test_it_is_a_ranked_source():
|
||||
assert not is_ambient("report_preference")
|
||||
|
||||
|
||||
def test_its_bar_is_its_own_key_not_the_prompt_arm_s(monkeypatch):
|
||||
"""Regression on the coupling #3860 found (and on the fix being real).
|
||||
|
||||
This arm shipped reading PROMPTRULE_THRESHOLD_KEY, so an operator tuning
|
||||
the bar for their own prose moved this one with it and was never told.
|
||||
That is worse here than anywhere else: every other arm scores a query that
|
||||
varies per call, while COMPLETION_QUERY is fixed — so this arm's score for
|
||||
a given corpus is a CONSTANT, and a constant that lands under the bar is a
|
||||
dead arm rather than a quiet one. No amount of traffic reveals it.
|
||||
|
||||
Asserted on the key the lookup actually asks for, which is the thing that
|
||||
broke, rather than on the constant being defined somewhere.
|
||||
"""
|
||||
import asyncio
|
||||
|
||||
from scribe.services import plugin_context
|
||||
from scribe.services.reply_preferences import completion_preferences
|
||||
|
||||
asked: list[str] = []
|
||||
|
||||
async def get_setting(user_id, key, default):
|
||||
asked.append(key)
|
||||
return default
|
||||
|
||||
async def search(user_id, query, **kw):
|
||||
kw["report"].update({"searched": True})
|
||||
return []
|
||||
|
||||
with patch(f"{M}.get_setting", AsyncMock(side_effect=get_setting)), \
|
||||
patch(f"{M}.semantic_search_rules", AsyncMock(side_effect=search)), \
|
||||
patch(f"{M}.record_retrieval"):
|
||||
asyncio.run(completion_preferences(7))
|
||||
|
||||
assert asked == [plugin_context.REPORTPREF_THRESHOLD_KEY]
|
||||
assert plugin_context.REPORTPREF_THRESHOLD_KEY != plugin_context.PROMPTRULE_THRESHOLD_KEY, (
|
||||
"the two keys are the same string again, so the settings form has one "
|
||||
"dial driving two arms — which is the defect, whatever the value is"
|
||||
)
|
||||
|
||||
|
||||
def test_the_query_assumes_no_particular_domain():
|
||||
from scribe.services.reply_preferences import COMPLETION_QUERY
|
||||
|
||||
|
||||
Reference in New Issue
Block a user