feat(500): retire the fixed-question preference arm - completion-report preferences ride the moments
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m10s
CI & Build / Python tests (push) Successful in 1m54s
CI & Build / Build & push image (push) Successful in 31s

report_preference searched one constant string when a task closed and handed
matches back as update_task's reply_preferences. Milestone 500 step 3 delivers
the reply shapes and the preferences mounted beside them on the moments, so
the arm goes, whole (rule 22):

- services/reply_preferences.py, the REPORT_PREFERENCE RuleArm, its re-scorer
  and corpus entries, and the REPORTPREF constants
- update_task's reply_preferences block and its cue; the docstring now points
  at moment_rules
- the fixed_query field and the fixed_query_never_clears warning: this arm was
  the only one that set it, so the concept and the cannot_decline exemption
  go with it
- the two Settings fields for its floor and budget
- migration 0121 deletes its rows (logs, judgments, rule-usage events, tuning
  history, settings keys), operator-approved 2026-10-09. Without it the
  readout would call the source unregistered and its surfacings would count
  as ambient. The unindexed rule_usage delete is bounded by created_at.

#5496 (step 4 of milestone 500).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-10-09 15:51:04 -04:00
co-authored by Claude Opus 5.5
parent cee1eb1328
commit f68b73f922
19 changed files with 94 additions and 619 deletions
+1 -13
View File
@@ -596,19 +596,7 @@ It is an UPPER BOUND per surface: a pull records the door it came
Only ever raised for unbidden arms known to log unconditionally: a
search returning a list every time is doing its job, and an arm whose
zeros were never written would flag a LOGGING bug while pointing you at
a threshold, which is #3497 exactly. Nor for an arm whose query never
changes — see the next entry.
- `fixed_query_never_clears` — an arm that always searches the SAME query
returned nothing on every call. Its score is one constant, so this is
not a quiet window: the bar sits above that constant and no amount of
further traffic will produce a different result. The arm is off rather
than silent, and nothing else here would say so. The same property is
why `cannot_decline` is not raised for these arms: with a constant
score the decline rate is 0% or 100% by construction, so "never
declined" is arithmetic and not evidence about the floor. Read the
refused record (`near_miss_samples`) BEFORE moving the dial — the last
time an arm sat here, every percentile said lower it and the refused
record showed the refusal was right.
a threshold, which is #3497 exactly.
- `band_hugs_floor` — the weakest tenth of what an arm returns sits on
its floor. The bar is doing the selecting and the score is not, so
moving that floor changes how MUCH you get, not how good it is.
+5 -22
View File
@@ -40,7 +40,6 @@ from scribe.services.notes import minted_kind
from scribe.services import placement as placement_svc
from scribe.services import planning as planning_svc
from scribe.services import record_batch as batch_svc
from scribe.services import reply_preferences as reply_prefs_svc
from scribe.services import rulebooks as rulebooks_svc
from scribe.services import systems as systems_svc
from scribe.services import task_logs as task_logs_svc
@@ -418,13 +417,11 @@ async def update_task(
Closing a task (done or cancelled) also returns `reply_shape`, the
default shape of the completion report you are about to write, and
`report_back`: a one-line reminder of what that reply should cover. When the
operator has preferences for how a completion report is written, they
come back as `reply_preferences` ({id, title, statement, kind}) — found
by their `when_to_apply`, so a preference whose trigger is writing the
report after finishing a task is the one that arrives here. Where one
differs from the default shape, the preference is what the operator
asked for. When family adoptions were answered `owed` while the task was
`report_back`: a one-line reminder of what that reply should cover. The
operator's own preferences for that report arrive beside it in
`moment_rules` — the ones mounted on finishing work. Where one differs
from the default shape, the preference is what the operator asked for.
When family adoptions were answered `owed` while the task was
open, they come back as `family_owed` ({idea_id, idea_title, project_id,
project_title, owed_task_id}): name each in the report — it is work
filed into that project.
@@ -470,14 +467,6 @@ async def update_task(
await family_svc.attach_family_hint(uid, data, note, created=False)
if status in _CLOSING_STATUSES:
data["report_back"] = REPORT_BACK_CUE
# The operator's own adjustments to the completion report, retrieved
# at the one moment a server can see that report coming (milestone
# 409 step 4). Omitted rather than sent empty, like every decoration.
prefs = await reply_prefs_svc.completion_preferences(
uid, project_id=getattr(note, "project_id", None))
if prefs:
data["reply_preferences"] = prefs
data["report_back"] = REPORT_BACK_CUE + " " + REPLY_PREFERENCES_CUE
# Owed family adoptions filed while the task was open (milestone 463
# step 6) — work now waiting in some project, which the report names.
await family_adoption_svc.attach_owed_adoptions(uid, data, note)
@@ -541,12 +530,6 @@ REPORT_BACK_CUE = (
"Reporting this to the operator? Say where it sits (from `placement`), "
"what now works, what needs them, and what comes next."
)
# Appended only when `reply_preferences` is present, so the key never arrives
# unexplained and a session with no preferences reads exactly what it did.
REPLY_PREFERENCES_CUE = (
"The operator has preferences for how this report is written — "
"follow `reply_preferences` over the default shape where they differ."
)
_ITEM_KEYS = {"title", "body", "type", "status", "priority", "kind", "tags", "system_ids"}
+2 -20
View File
@@ -601,25 +601,6 @@ PROMPTRULE_DEFAULT_THRESHOLD = SURFACES["prompt_rule"].floor_default
# at all, where the act arms never face it because a command is one thing.
PROMPTRULE_LIMIT = SURFACES["prompt_rule"].budget_default
# THE COMPLETION-REPORT ARM'S OWN BAR (services/reply_preferences.py).
#
# It borrowed PROMPTRULE_THRESHOLD_KEY when it shipped, which made the two
# arms one dial: an operator lowering the bar for their own prose moved this
# one with it, silently. That contradicts the rule every other bar here
# follows — one number cannot serve arms whose queries are different shapes —
# and this arm's query is the most different of all. The others score an
# operator's prose or a session's code, both of which vary per call; this one
# scores a FIXED string (`COMPLETION_QUERY`) against rule triggers, so its
# score for a given corpus is a constant. A constant that lands under the bar
# is not a quiet arm, it is a dead one, and nothing about the prose arm's
# traffic would ever reveal it.
#
# Kept at the prose arm's starting value rather than tuned: the split is what
# makes the two independently movable, and a default is a product decision
# that this install's corpus cannot settle (rule 115).
REPORTPREF_THRESHOLD_KEY = SURFACES["report_preference"].floor_key
REPORTPREF_DEFAULT_THRESHOLD = SURFACES["report_preference"].floor_default
def _slugify(text: str) -> str:
"""kebab-case slug for a skill directory name (a-z0-9 + single hyphens)."""
@@ -726,7 +707,8 @@ async def get_autoinject_config(user_id: int) -> dict:
THE TWO NUMBERS COME FROM THE REGISTRY NOW (#4102). They used to be read and
clamped here, and identically again in `get_writepath_config`, and again in
three rule arms, and once more in `reply_preferences`. That was
three rule arms, and once more in the completion-report arm (since
retired, milestone 500). That was
tolerable while the values were shipped constants. It stops being tolerable
once a tool is expected to MOVE them, because a tuning surface cannot be
consistent across arms that each spell their configuration differently.
-141
View File
@@ -1,141 +0,0 @@
"""The operator's own preferences for a completion report, at the moment it is written.
WHY THIS EXISTS (milestone 409 step 4)
The reporting-back skill ships DEFAULT shapes. An operator will want some of
them different ("my completion reports also say how it was tested", "decisions
as a numbered list"), and those adjustments are `preference` records. The gap
is the query: prompt-time retrieval matches the OPERATOR'S MESSAGE, and a shape
preference is about the REPLY. "Fix the flaky test" never retrieves "completion
reports should say how it was tested", so the preference is on file and never
arrives.
THE DECISION (operator, 2026-09-14, logged on the step): two deliveries, split
by reply kind.
- A COMPLETION REPORT has a moment the server can see — a task closing — so
the server retrieves for it and hands the matches back beside
`report_back`. That is this module.
- EVERY OTHER REPLY KIND (a finding, a decision, a handoff…) has no tool call
in front of it, so the reporting-back skill asks: it tells the agent to
`search(content_type="rule")` for that kind before writing.
Loading reply-shape preferences at session start was the rejected third option:
it is a small copy of the preloading milestone 394 retired, and just as
unmeasurable.
HOW A PREFERENCE SAYS IT IS ABOUT COMPLETION REPORTS
By its trigger, which is what it already has — no tag, no new column. A
preference's `when_to_apply` dominates its embedded document, so one written
for this moment ("writing the report after finishing a task") resembles
COMPLETION_QUERY below, and one about anything else does not. That keeps
delivery entirely in retrieval, as 394 decided, and leaves an operator nothing
new to learn: a preference reaches the completion report the same way every
other record reaches its moment.
THE BAR IS ITS OWN, AND THE FIRST READING EARNED IT (#3860)
This borrowed the prompt arm's key when it shipped, on the argument that a
fixed query against triggers is a different score distribution from an
operator's message against the same documents — true, and the reason the two
could not stay one dial. Five days of traffic settled it.
What the readout said: 69 calls, 69 declines, every one naming the SAME record
at the SAME score (rule 77 at 0.7194 against a 0.72 bar). That constancy is
the signature of this arm — `COMPLETION_QUERY` never varies, so for a given
corpus its best score is a constant, and a constant sitting under the bar is a
dead arm rather than a quiet one. The record it kept declining was about
reading a REQUEST, not about the shape of a report, so the decline was right
and the arm is healthy: this install simply has no completion-report
preference on file.
The bar stayed at 0.72, and the key moved out (REPORTPREF_THRESHOLD_KEY) so
that staying is a decision rather than a side effect of what the prose arm is
set to. A surface whose score cannot vary is the one surface where a borrowed
bar can be wrong forever without a single call looking unusual.
"""
from __future__ import annotations
import logging
from scribe.services import retrieval_pipeline as rp
from scribe.services.embeddings import semantic_search_rules
from scribe.services.retrieval_surfaces import SURFACES, budget_for, floor_for
from scribe.services.retrieval_telemetry import record_retrieval
from scribe.services.rule_usage import record_rule_surfaced
logger = logging.getLogger(__name__)
SOURCE = rp.REPORT_PREFERENCE.source
# Written in the vocabulary of the MOMENT, because that is what a trigger is
# written in and what this query is scored against. Domain-neutral on purpose
# (rule #115): a writing project or a home-infrastructure project closes tasks
# too, and its operator's preferences must match as well as a developer's.
COMPLETION_QUERY = (
"writing the completion report to the operator after finishing a task — "
"how that reply should be laid out and what it should include"
)
# A handful, not a menu. More than a few shape preferences for ONE kind of
# reply would contradict each other before they helped; the limit is here to
# keep one noisy corpus from turning a status change into a wall of text.
# The STARTING budget, not the budget (#4102). Both numbers this arm runs on
# now come from the surface registry, so the model that reads this arm's
# telemetry can move either — which matters more here than anywhere else,
# because a fixed query makes this arm's score a constant and a floor a hair
# above it produces a dead arm no amount of traffic will ever reveal.
LIMIT = SURFACES["report_preference"].budget_default
async def _threshold(user_id: int) -> float:
return await floor_for(user_id, SOURCE)
async def _limit(user_id: int) -> int:
return await budget_for(user_id, SOURCE)
async def completion_preferences(user_id: int, *, project_id: int | None = None) -> list[dict]:
"""The operator's preferences for a completion report, best match first.
KIND-FILTERED, for the reason `_reserve_slot_for_preference` gives: a
binding rule that happened to resemble the query would otherwise ride out
under a key that says "how the operator likes this written", which is a
claim about force the record does not make.
Every call is logged, the empty ones included: a surface that records only
the calls it liked reports a flawless clear-rate however badly its bar is
set. An install with no preferences at all logs nothing, because no search
ran (`searched` stays False) — that is not a decline.
Fails open to an empty list: this decorates a write that has already
happened, and a lookup that errors must not turn it into a failure.
"""
try:
# The stages — the kind filter, the unconditional call row, the
# fresh-only surfacing rows — are the one pipeline's (milestone 456).
# RANKED: this surface chose what it showed, so the name is in
# rule_usage.RANKED_SOURCES and its hits count toward pull-through.
# There is no session ledger here: a completion report is written
# once, so nothing it could repeat has been shown before.
result = await rp.run_rule_arm(
rp.REPORT_PREFERENCE,
rp.RuleMoment(
user_id=user_id, query=COMPLETION_QUERY, project_id=project_id,
),
floor=await _threshold(user_id), budget=await _limit(user_id),
io=rp.RuleIO(
search=semantic_search_rules,
record_retrieval=record_retrieval,
record_rule_surfaced=record_rule_surfaced,
),
)
return [
{"id": rule.id, "title": rule.title, "statement": rule.statement, "kind": "preference"}
for _score, rule in result.shown
]
except Exception: # noqa: BLE001 - a decoration never breaks its payload
logger.warning("completion preference lookup failed", exc_info=True)
return []
@@ -117,7 +117,6 @@ _RESCORERS = {
"pre_tool_rule": lambda u, q, p: _rescore_rules(u, q, p, None),
"reply_rule": lambda u, q, p: _rescore_rules(u, q, p, None),
"prompt_rule": lambda u, q, p: _rescore_rules(u, q, p, None),
"report_preference": lambda u, q, p: _rescore_rules(u, q, p, "preference"),
}
# The embedding table each surface's re-scorer reads. A migration from a corpus
@@ -131,7 +130,6 @@ _CORPUS = {
"pre_tool_rule": RuleEmbedding,
"reply_rule": RuleEmbedding,
"prompt_rule": RuleEmbedding,
"report_preference": RuleEmbedding,
}
+1 -36
View File
@@ -421,10 +421,6 @@ class Declared:
quiet_because: str = ""
"""Set when silence over an active window is correct, saying why (#2475)."""
fixed_query: bool = False
"""The arm always searches the same string, so its decline rate is 0% or
100% and `cannot_decline` says nothing about it."""
@dataclass(frozen=True)
class RankedSource:
@@ -523,37 +519,6 @@ WRITE_PATH_RULE = RuleArm(
),
declared=Declared("rules that may govern the file being written"),
)
# The completion report's preferences (milestone 409 step 4): a FIXED query,
# preferences only, read by update_task as records rather than as lines. Its
# query is `reply_preferences.COMPLETION_QUERY`.
REPORT_PREFERENCE = RuleArm(
"report_preference", band=False, compact_tail=False, checkpoint=False,
preference_slot=False, kind="preference",
tuning=Surface(
name="report_preference",
floor_key="kb_reportpref_threshold",
floor_default=0.72,
budget_key="kb_reportpref_top_k",
budget_default=3,
# THE ONE FIXED QUERY, and the reason this arm behaves unlike the rest.
# The others score something that varies per call; this one scores a
# constant string, so its top score for a given corpus is also a
# constant. A floor a hair above that constant is not a quiet arm, it
# is a dead one, and no amount of traffic will ever reveal it — which
# is precisely how this arm spent 69 calls declining the same record.
asks="a fixed question about how to lay out a completion report",
over="preferences",
fires="when a task finishes",
),
# `fixed_query`: COMPLETION_QUERY is a module constant, so this arm's top
# score is the same number on every call — measured at 0.791 across 45
# consecutive calls, with p10, p50, p90, min and max all identical. Five
# equal percentiles is the tell.
declared=Declared(
"the fixed question asked when a task finishes: how should this report read",
fixed_query=True,
),
)
# The backstop for every arm that ran earlier in the turn and missed: the
# finished reply against every rule's trigger. Its floor IS its stop bar
# (the reply_rule surface), so it is passed as both.
@@ -581,7 +546,7 @@ REPLY_RULE = RuleArm(
),
)
RULE_ARMS: tuple[RuleArm, ...] = (
WRITE_PATH_RULE, PRE_TOOL_RULE, PROMPT_RULE, REPORT_PREFERENCE, REPLY_RULE,
WRITE_PATH_RULE, PRE_TOOL_RULE, PROMPT_RULE, REPLY_RULE,
)
PREFERENCE_SLOT_SOURCE = "preference_slot"
+1 -25
View File
@@ -94,29 +94,6 @@ class Point:
quiet_because: str = ""
fixed_query: bool = False
"""Whether this arm always searches the SAME query string.
THE #3497 GUARD, ONE STEP OVER. `logs_unconditionally` below exists
because a warning computed over a LOGGING property read as a ranking
problem. This field exists because a warning computed over a QUERY-SHAPE
property does the same thing.
An arm with a fixed query scores against one constant. Its decline rate is
therefore 0% or 100% and nothing in between — which of the two depends
only on whether the bar sits below or above that single number. So "never
returned nothing" says nothing at all about whether a floor is applied,
and `cannot_decline` — whose whole remedy is "check that it applies its
floor" — is uninformative here and skips these arms.
What IS informative for them is the mirror image, and `reply_preferences`
names it in its own docstring: every call returning nothing means the bar
sits above the constant, no traffic will ever move it, and the arm is
dead. That has happened — 69 consecutive declines at 0.0006 under the bar
(see `retrieval_surfaces`) — so it gets its own warning rather than
inheriting one written for arms whose score can vary.
"""
logs_unconditionally: bool = True
"""Whether this arm writes a row even when it returns NOTHING.
@@ -144,8 +121,7 @@ POINTS: dict[str, Point] = dict([
# ambient or pulled.
*(_p(spec.source, UNBIDDEN, spec.declared.what,
expects_traffic=not spec.declared.quiet_because,
quiet_because=spec.declared.quiet_because,
fixed_query=spec.declared.fixed_query)
quiet_because=spec.declared.quiet_because)
for spec in RANKED),
# A LOOKUP, not a ranker (#4796): a record the operator named by number
# in the message, fetched by id. No score and no bar, so it writes no
+3 -3
View File
@@ -5,9 +5,9 @@ WHY THIS EXISTS
Six push arms each carried their own loose copy of the same shape: a settings
key, a default, and a limit that was usually a module constant nobody could
change. The read-and-clamp was written out separately in `plugin_context` (twice
over, for auto-inject and the write path), in three rule arms, and again as
`reply_preferences._threshold`. That was survivable while the numbers were
shipped constants an operator occasionally edited.
over, for auto-inject and the write path), in three rule arms, and again in
the completion-report arm (since retired, milestone 500). That was survivable
while the numbers were shipped constants an operator occasionally edited.
It stops being survivable once the numbers are meant to MOVE. The operator's
decision for this step:
@@ -591,22 +591,12 @@ def _compute_warnings(sources: dict, usage: dict, rule_usage: dict,
# arms once recorded only their hits, so their decline count was
# structurally zero and this warning would have fired on a LOGGING
# defect while pointing the reader at the threshold.
#
# FOUR NOW. A FIXED-QUERY arm is exempt for the same reason one step
# over (#4232): it scores against one constant, so its decline rate is
# 0% or 100% and never in between, and which one it is depends only on
# where the bar sits relative to that single number. "Never returned
# nothing" is then not evidence about the floor — it is arithmetic —
# and this warning's own remedy, "check that it applies its floor",
# cannot be answered from it. Those arms get `fixed_query_never_clears`
# below, which asks the question that IS answerable for them.
if (
calls >= min_calls
and (b.get("zero_result_calls") or 0) == 0
and point is not None
and point.kind == UNBIDDEN
and point.logs_unconditionally
and not point.fixed_query
):
out.append(_warn(
"cannot_decline",
@@ -617,46 +607,6 @@ def _compute_warnings(sources: dict, usage: dict, rule_usage: dict,
source=name, calls=calls, zero_result_calls=0,
))
# ── A fixed-query arm that never clears its bar ──────────────────
#
# The mirror image of `cannot_decline`, and the state that actually
# threatens these arms. `reply_preferences` names it in its own
# docstring: "a fixed query makes this arm's score a constant and a
# floor a hair above it produces a dead arm no amount of traffic will
# ever reveal."
#
# For an arm whose score can vary, a window of all-empty calls is
# ordinary — it means nothing matched, which is an answer. For one
# whose score is a constant it means the bar is above that constant,
# and no volume of further calls will ever produce a different result.
# The arm is not quiet; it is switched off, and nothing else in this
# readout would say so.
#
# It has happened: `report_preference` logged 69 consecutive declines
# at 0.0006 under the bar. Note what that incident also proves — the
# fix is NOT automatically to lower the floor. Reading the refused
# record showed the refusal was correct, so this warning sends the
# reader to `near_miss_samples` rather than to the dial.
if (
calls >= min_calls
and point is not None
and point.fixed_query
and point.logs_unconditionally
and (b.get("zero_result_calls") or 0) == calls
):
out.append(_warn(
"fixed_query_never_clears",
f"{calls} calls, every one of them empty — and this arm always "
f"searches the same query, so its score is a constant. That "
f"means the bar sits above it and no amount of further traffic "
f"will change the result: the arm is off, not quiet. Read the "
f"record it refused (`near_miss_samples`) before touching the "
f"floor — the last time this arm sat here, every percentile "
f"said lower it and the refused record showed the refusal was "
f"right.",
source=name, calls=calls, zero_result_calls=calls,
))
# ── Band hugs its floor ──────────────────────────────────────────
#
# Read on p10, the WEAKEST tenth of what the arm returned. If even