feat(409): an operator's own reply shapes reach the reply they are about (#4013)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m24s
CI & Build / Build & push image (push) Successful in 23s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m24s
CI & Build / Build & push image (push) Successful in 23s
The reporting-back skill ships default shapes; an operator's adjustments to them are preference records. Prompt-time retrieval matches the operator's message, and a shape preference is about the reply, so those preferences were on file and never arrived. Operator's decision (logged on #4013): the server delivers them for a completion report, and the skill asks for every other kind. - Completion reports (option C): closing a task with update_task runs a kind-filtered preference search for the moment "writing the completion report after finishing a task" and returns matches as `reply_preferences` ({id, title, statement, kind}), with a sentence added to `report_back` naming the key. A preference says it is about completion reports through its own when_to_apply; no tag or column. Omitted when nothing matches, and the lookup fails open. - Telemetry: every call logs to retrieval_logs under `report_preference` (empty calls included; a search that never ran writes no row) and hits are recorded surfaced. The source is ranked, so it counts toward pull-through. The bar is the prompt arm's setting until step 6 reads this source's near misses. - Every other reply (option A): reporting-back gains "The operator's own shapes come first". Before a finding, decision, handoff or "where are we", search(content_type="rule") in the words of that moment and follow what comes back. Registered in the ownership guard with reporting-back as owner. - Loading reply shapes at session start (option B) was rejected: it would be a small copy of the preloading milestone 394 retired. Domain-neutral query (pinned); works on an install with no preferences. Plugin version minted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,81 @@
|
||||
"""The completion-report preference lookup (milestone 409 step 4).
|
||||
|
||||
What it pins: the lookup asks for PREFERENCES only, logs every call under its
|
||||
own source (the empty ones too), counts only what it showed as surfaced, and
|
||||
fails open. The query stays domain-neutral, because every kind of project
|
||||
closes tasks.
|
||||
"""
|
||||
import re
|
||||
from unittest.mock import AsyncMock, MagicMock, patch
|
||||
|
||||
M = "scribe.services.reply_preferences"
|
||||
|
||||
|
||||
def _rule(rid, kind="preference"):
|
||||
return MagicMock(id=rid, kind=kind, title=f"t{rid}", statement=f"s{rid}")
|
||||
|
||||
|
||||
async def _run(hits, *, searched=True, raises=None):
|
||||
from scribe.services.reply_preferences import completion_preferences
|
||||
|
||||
async def search(user_id, query, **kw):
|
||||
if raises:
|
||||
raise raises
|
||||
kw["report"].update({"searched": searched, "best_available_score": 0.7})
|
||||
return hits
|
||||
|
||||
search_mock = AsyncMock(side_effect=search)
|
||||
with patch(f"{M}.semantic_search_rules", search_mock), \
|
||||
patch(f"{M}.get_setting", AsyncMock(return_value="0.72")), \
|
||||
patch(f"{M}.record_retrieval") as logged, \
|
||||
patch(f"{M}.record_rule_surfaced") as surfaced:
|
||||
out = await completion_preferences(7, project_id=3)
|
||||
return out, search_mock, logged, surfaced
|
||||
|
||||
|
||||
async def test_asks_for_preferences_only_and_returns_them_best_first():
|
||||
out, search, logged, surfaced = await _run([(0.9, _rule(1)), (0.8, _rule(2))])
|
||||
assert search.await_args.kwargs["kind"] == "preference"
|
||||
assert [p["id"] for p in out] == [1, 2]
|
||||
assert all(p["kind"] == "preference" for p in out)
|
||||
assert logged.call_args.kwargs["source"] == "report_preference"
|
||||
assert surfaced.call_args.kwargs == {"user_id": 7, "rule_ids": [1, 2], "source": "report_preference"}
|
||||
|
||||
|
||||
async def test_a_rule_that_slips_through_is_not_handed_back_as_a_preference():
|
||||
out, _, _, surfaced = await _run([(0.9, _rule(1, kind="rule")), (0.8, _rule(2))])
|
||||
assert [p["id"] for p in out] == [2]
|
||||
assert surfaced.call_args.kwargs["rule_ids"] == [2]
|
||||
|
||||
|
||||
async def test_an_empty_call_is_still_logged_and_surfaces_nothing():
|
||||
out, _, logged, surfaced = await _run([])
|
||||
assert out == []
|
||||
logged.assert_called_once()
|
||||
assert logged.call_args.kwargs["results"] == []
|
||||
surfaced.assert_not_called()
|
||||
|
||||
|
||||
async def test_a_search_that_never_ran_says_so_to_the_log():
|
||||
_, _, logged, _ = await _run([], searched=False)
|
||||
assert logged.call_args.kwargs["searched"] is False
|
||||
|
||||
|
||||
async def test_fails_open():
|
||||
out, _, _, surfaced = await _run([], raises=RuntimeError("embedder down"))
|
||||
assert out == []
|
||||
surfaced.assert_not_called()
|
||||
|
||||
|
||||
def test_it_is_a_ranked_source():
|
||||
from scribe.services.rule_usage import is_ambient
|
||||
|
||||
assert not is_ambient("report_preference")
|
||||
|
||||
|
||||
def test_the_query_assumes_no_particular_domain():
|
||||
from scribe.services.reply_preferences import COMPLETION_QUERY
|
||||
|
||||
dev_only = [w for w in (r"\bCI\b", r"\bcommit", r"\bpull request", r"\bcode\b", r"\btest")
|
||||
if re.search(w, COMPLETION_QUERY, re.IGNORECASE)]
|
||||
assert not dev_only, f"software-only vocabulary in a query every project runs: {dev_only}"
|
||||
Reference in New Issue
Block a user