feat(409): an operator's own reply shapes reach the reply they are about (#4013)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m24s
CI & Build / Build & push image (push) Successful in 23s

The reporting-back skill ships default shapes; an operator's adjustments to
them are preference records. Prompt-time retrieval matches the operator's
message, and a shape preference is about the reply, so those preferences were
on file and never arrived. Operator's decision (logged on #4013): the server
delivers them for a completion report, and the skill asks for every other kind.

- Completion reports (option C): closing a task with update_task runs a
  kind-filtered preference search for the moment "writing the completion
  report after finishing a task" and returns matches as `reply_preferences`
  ({id, title, statement, kind}), with a sentence added to `report_back`
  naming the key. A preference says it is about completion reports through
  its own when_to_apply; no tag or column. Omitted when nothing matches, and
  the lookup fails open.
- Telemetry: every call logs to retrieval_logs under `report_preference`
  (empty calls included; a search that never ran writes no row) and hits are
  recorded surfaced. The source is ranked, so it counts toward pull-through.
  The bar is the prompt arm's setting until step 6 reads this source's near
  misses.
- Every other reply (option A): reporting-back gains "The operator's own
  shapes come first". Before a finding, decision, handoff or "where are we",
  search(content_type="rule") in the words of that moment and follow what
  comes back. Registered in the ownership guard with reporting-back as owner.
- Loading reply shapes at session start (option B) was rejected: it would be
  a small copy of the preloading milestone 394 retired.

Domain-neutral query (pinned); works on an install with no preferences.
Plugin version minted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-09-14 17:43:16 -04:00
co-authored by Claude Opus 5
parent 9071cb05da
commit 921565696c
8 changed files with 294 additions and 8 deletions
+81
View File
@@ -0,0 +1,81 @@
"""The completion-report preference lookup (milestone 409 step 4).
What it pins: the lookup asks for PREFERENCES only, logs every call under its
own source (the empty ones too), counts only what it showed as surfaced, and
fails open. The query stays domain-neutral, because every kind of project
closes tasks.
"""
import re
from unittest.mock import AsyncMock, MagicMock, patch
M = "scribe.services.reply_preferences"
def _rule(rid, kind="preference"):
return MagicMock(id=rid, kind=kind, title=f"t{rid}", statement=f"s{rid}")
async def _run(hits, *, searched=True, raises=None):
from scribe.services.reply_preferences import completion_preferences
async def search(user_id, query, **kw):
if raises:
raise raises
kw["report"].update({"searched": searched, "best_available_score": 0.7})
return hits
search_mock = AsyncMock(side_effect=search)
with patch(f"{M}.semantic_search_rules", search_mock), \
patch(f"{M}.get_setting", AsyncMock(return_value="0.72")), \
patch(f"{M}.record_retrieval") as logged, \
patch(f"{M}.record_rule_surfaced") as surfaced:
out = await completion_preferences(7, project_id=3)
return out, search_mock, logged, surfaced
async def test_asks_for_preferences_only_and_returns_them_best_first():
out, search, logged, surfaced = await _run([(0.9, _rule(1)), (0.8, _rule(2))])
assert search.await_args.kwargs["kind"] == "preference"
assert [p["id"] for p in out] == [1, 2]
assert all(p["kind"] == "preference" for p in out)
assert logged.call_args.kwargs["source"] == "report_preference"
assert surfaced.call_args.kwargs == {"user_id": 7, "rule_ids": [1, 2], "source": "report_preference"}
async def test_a_rule_that_slips_through_is_not_handed_back_as_a_preference():
out, _, _, surfaced = await _run([(0.9, _rule(1, kind="rule")), (0.8, _rule(2))])
assert [p["id"] for p in out] == [2]
assert surfaced.call_args.kwargs["rule_ids"] == [2]
async def test_an_empty_call_is_still_logged_and_surfaces_nothing():
out, _, logged, surfaced = await _run([])
assert out == []
logged.assert_called_once()
assert logged.call_args.kwargs["results"] == []
surfaced.assert_not_called()
async def test_a_search_that_never_ran_says_so_to_the_log():
_, _, logged, _ = await _run([], searched=False)
assert logged.call_args.kwargs["searched"] is False
async def test_fails_open():
out, _, _, surfaced = await _run([], raises=RuntimeError("embedder down"))
assert out == []
surfaced.assert_not_called()
def test_it_is_a_ranked_source():
from scribe.services.rule_usage import is_ambient
assert not is_ambient("report_preference")
def test_the_query_assumes_no_particular_domain():
from scribe.services.reply_preferences import COMPLETION_QUERY
dev_only = [w for w in (r"\bCI\b", r"\bcommit", r"\bpull request", r"\bcode\b", r"\btest")
if re.search(w, COMPLETION_QUERY, re.IGNORECASE)]
assert not dev_only, f"software-only vocabulary in a query every project runs: {dev_only}"