feat(moments): the reply moment holds a finished reply for one read (milestone 458 step 4b, #4922)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Failing after 1m25s
CI & Build / Build & push image (push) Skipped

The reply is the one act no tool call marks, and it is where "let me
know if it works" gets said. A new Stop hook (scribe_reply_check.sh)
sends the finished reply to POST /api/plugin/reply-rules, which checks
it twice:

- mounted: every unopened RULE on reply.report, plus reply.ask when the
  reply asks a question. Deterministic.
- semantic: the reply's head and tail against every rule's trigger, on a
  new ranked surface, reply_rule. It is the backstop for whatever the
  earlier arms missed. Its floor is its stop bar (default 0.80, budget
  1), with its own Settings dials. The new stop_only stage records
  surfacing for the rule that holds and nothing else, because nothing
  else reached anyone.

Following the operator's ruling from 456 step 8, a rule that holds blocks
once, in the server's words. The hook blocks only on a reason it was
given, so an unreachable instance never stops a session, and it never
holds the rewrite. The ledger is the act checkpoint's own, so a rule
holds a session once across both doors and the per-session cap counts
both.

The turn reader moved from the report check into scribe_defs.sh
(scribe_turn_facts / scribe_turn_fact), so the two Stop hooks read a
turn the same way. The output was checked identical on a real transcript.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-10-05 12:30:07 -04:00
co-authored by Claude Opus 5.5
parent b3b616b20a
commit c6cdfc2172
17 changed files with 740 additions and 59 deletions
+36
View File
@@ -294,6 +294,42 @@ async def moment():
})
@plugin_bp.post("/reply-rules")
@login_required
@memoized_query_embeddings
async def reply_rules():
"""Whether the reply that ends a turn is held for one read (milestone 458).
Called by the plugin's Stop hook with the finished reply. Two halves: the
rules MOUNTED on the reply moments (`reply.report`, and `reply.ask` when
the reply asks something), and the reply text against every rule's
trigger — the backstop for whatever the earlier arms missed.
Body: `{"reply": "<text>"}`. Query: `repo` / `project_id` for the scope;
`held_rule_ids` (opened this session — exempt, as at the act checkpoint);
`exclude_rule_ids` (named this session — not re-counted as surfaced);
`stopped_rule_ids` (already held a reply or an act this session — a rule
holds once, and the per-session cap counts these).
Returns `reason` — the words the hook blocks with, empty when nothing
holds — plus `rule_ids` for the hook's ledger and `moments` reached.
"""
data = await request.get_json(silent=True) or {}
reply = str(data.get("reply") or "")
project_id, _repo, _unbound = await _project_scope()
held = await moment_delivery_svc.reply_hold(
g.user.id, reply, project_id=project_id,
exclude=frozenset(_int_list(request.args.get("exclude_rule_ids"))),
held=frozenset(_int_list(request.args.get("held_rule_ids"))),
stopped=frozenset(_int_list(request.args.get("stopped_rule_ids"))),
)
return jsonify({
"reason": held.get("reason", ""),
"rule_ids": held.get("rule_ids", []),
"moments": held.get("moments", []),
})
@plugin_bp.get("/prior-art")
@login_required
@memoized_query_embeddings
+150
View File
@@ -117,3 +117,153 @@ async def attach_moment_rules(
except Exception: # noqa: BLE001 - a decoration never breaks the payload
logger.debug("moment rules for %s could not be attached", tool, exc_info=True)
return data
# ── The reply moment: the backstop at the end of a turn ─────────────────
#
# The Stop hook's door, folded in from milestone 456 step 8. A reply is the
# one act no tool call marks, and the moment where "please verify this on
# your end" gets said — so it is checked twice over: the rules MOUNTED on the
# reply moments (deterministic, somebody said they belong here), and the
# reply text against every rule's trigger (the backstop for whatever the
# earlier arms missed). An unopened rule from either half HOLDS the reply for
# one read, in these words; the hook never holds the rewrite.
# The reply's head and tail. The embedder reads ~512 tokens, and the part of
# a report that asks something of the reader — "let me know if it works" —
# is at the END, where a head-only query would cut it off.
_REPLY_HEAD_CHARS = 600
_REPLY_TAIL_CHARS = 1200
_REACHED_BY_REPLY = "the reply that ends this turn"
def reply_moments(reply: str) -> list[dict]:
"""The moments a finished reply reaches: always `reply.report`, and
`reply.ask` when a line of its prose ends in a question."""
reached = [{"moment": "reply.report", "tool": "", "match": _REACHED_BY_REPLY,
"via": moment_actions.DEFAULT}]
in_code = False
for line in (reply or "").splitlines():
stripped = line.strip()
if stripped.startswith("```"):
in_code = not in_code
continue
if not in_code and stripped.endswith("?"):
reached.append({"moment": "reply.ask", "tool": "", "match": "a question in the reply",
"via": moment_actions.DEFAULT})
break
return reached
def _reply_query(reply: str) -> str:
text = " ".join((reply or "").split())
if len(text) <= _REPLY_HEAD_CHARS + _REPLY_TAIL_CHARS:
return text
return text[:_REPLY_HEAD_CHARS] + " … " + text[-_REPLY_TAIL_CHARS:]
def reply_hold_reason(held: list[dict]) -> str:
"""The text the agent reads INSTEAD of its reply going out.
A practice, not a prohibition (rule 165), on the act checkpoint's model:
nothing here knows the reply is wrong. It names each rule with why it was
raised, gives the remedy as one call per rule, and says the reply may go
out unchanged — so a reader who finds the rules beside the point is out in
as many calls as there are rules, and is never held twice.
"""
if not held:
return ""
parts = []
for item in held:
why = (f"mounted on {item['moment']}" if item.get("moment")
else f"scores {item['score']} against this reply")
trigger = f"; it applies when {item['trigger']}" if item.get("trigger") else ""
parts.append(f"“{item['title']}” ({why}{trigger}) — get_rule({item['rule_id']})")
noun = "a standing rule" if len(held) == 1 else "standing rules"
return (
f"Held for one read before this reply goes out: {noun} this session "
f"has not opened: " + "; ".join(parts) + ". Read "
+ ("it" if len(held) == 1 else "them")
+ ", then send the reply — unchanged if it already does what the rule "
"asks, which is a judgement only you can make. This check runs once "
"per rule; the rewrite is never held."
)
async def reply_hold(
user_id: int, reply: str, *, project_id: int | None = None,
exclude: frozenset[int] = frozenset(), held: frozenset[int] = frozenset(),
stopped: frozenset[int] = frozenset(),
) -> dict:
"""Whether this finished reply is held, and in what words.
`held` is what the session OPENED (the act checkpoint's exemption);
`stopped` is what an earlier hold already put in front of it, so a rule
holds a session once. Past the act checkpoint's per-session cap nothing
holds — a mis-set bar degrades to a quiet session, never a stuck one.
Returns {} or {"reason", "rule_ids", "moments"}. Fails open to {}.
"""
from scribe.services import plugin_context as pc
from scribe.services import rule_usage, rulebooks
from scribe.services.retrieval_surfaces import budget_for, floor_for
try:
if not (reply or "").strip() or len(stopped) >= pc.CHECKPOINT_SESSION_CAP:
return {}
reached = reply_moments(reply)
names = [hit["moment"] for hit in reached]
skip = set(held) | set(stopped)
room = pc.CHECKPOINT_SESSION_CAP - len(stopped)
out: list[dict] = []
# The mounted half: deterministic, so every unopened RULE on these
# moments holds. A preference claims no such force.
for rule, at in await rulebooks.rules_on_moments(user_id, names, project_id or None):
if rule.id in skip or rule.kind == "preference":
continue
out.append({"rule_id": rule.id, "title": rule.title, "moment": at,
"trigger": (rule.when_to_apply or "").strip()})
# Trimmed BEFORE recording: a rule the cap cut was shown to nobody.
out = out[:room]
mounted = [item["rule_id"] for item in out]
fresh = [rid for rid in mounted if rid not in exclude]
if fresh:
try:
rule_usage.record_rule_surfaced(
user_id=user_id, rule_ids=fresh, source=rp.MOMENT_RULE_SOURCE,
detail={item["rule_id"]: item["moment"] for item in out
if item["rule_id"] in fresh},
)
except Exception: # noqa: BLE001 - observation never breaks the observed
logger.debug("reply mount surfacing not recorded", exc_info=True)
# The semantic half: the backstop. A mounted rule already holding is
# treated as held here, so one rule is never named twice.
floor = await floor_for(user_id, rp.REPLY_RULE.source)
result = await rp.run_rule_arm(
rp.REPLY_RULE,
rp.RuleMoment(
user_id=user_id, query=_reply_query(reply), project_id=project_id,
where="to this reply", checkpoint_where="this reply",
exclude=exclude, held=frozenset(skip | set(mounted)),
),
floor=floor, budget=await budget_for(user_id, rp.REPLY_RULE.source),
io=pc._rule_io(), checkpoint_floor=floor,
)
if result.checkpoint and len(out) < room:
cp = result.checkpoint
out.append({"rule_id": cp["rule_id"], "title": cp["title"],
"score": cp["score"], "trigger": cp.get("trigger", "")})
if not out:
return {}
return {
"reason": reply_hold_reason(out),
"rule_ids": [item["rule_id"] for item in out],
"moments": names,
}
except Exception: # noqa: BLE001 - a recall aid never stops a session by failing
logger.debug("reply hold failed", exc_info=True)
return {}
@@ -115,6 +115,7 @@ _RESCORERS = {
),
"write_path_rule": lambda u, q, p: _rescore_rules(u, q, p, None),
"pre_tool_rule": lambda u, q, p: _rescore_rules(u, q, p, None),
"reply_rule": lambda u, q, p: _rescore_rules(u, q, p, None),
"prompt_rule": lambda u, q, p: _rescore_rules(u, q, p, None),
"report_preference": lambda u, q, p: _rescore_rules(u, q, p, "preference"),
}
@@ -128,6 +129,7 @@ _CORPUS = {
"write_path": NoteEmbedding,
"write_path_rule": RuleEmbedding,
"pre_tool_rule": RuleEmbedding,
"reply_rule": RuleEmbedding,
"prompt_rule": RuleEmbedding,
"report_preference": RuleEmbedding,
}
+25 -8
View File
@@ -354,6 +354,14 @@ class RuleArm:
kind: str | None = None
"""Ask the ranker for one record kind only ("preference"), or every kind."""
stop_only: bool = False
"""The door shows nothing but the stop (milestone 458's reply moment).
At the end of a turn there is no context to annotate: a hit either holds
the reply for one read or reaches nobody. So only the checkpoint's rule is
recorded as surfaced — the rest were ranked, and the call row counts them,
but nobody was shown them."""
# Today's differences, reproduced exactly (milestone 456 step 2). Whether the
# prompt arm should band, and whether the act arms should reserve a
@@ -377,8 +385,15 @@ REPORT_PREFERENCE = RuleArm(
"report_preference", band=False, compact_tail=False, checkpoint=False,
preference_slot=False, kind="preference",
)
# The backstop for every arm that ran earlier in the turn and missed: the
# finished reply against every rule's trigger. Its floor IS its stop bar
# (the reply_rule surface), so it is passed as both.
REPLY_RULE = RuleArm(
"reply_rule", band=False, compact_tail=False, checkpoint=True,
preference_slot=False, stop_only=True,
)
RULE_ARMS: tuple[RuleArm, ...] = (
WRITE_PATH_RULE, PRE_TOOL_RULE, PROMPT_RULE, REPORT_PREFERENCE,
WRITE_PATH_RULE, PRE_TOOL_RULE, PROMPT_RULE, REPORT_PREFERENCE, REPLY_RULE,
)
PREFERENCE_SLOT_SOURCE = "preference_slot"
@@ -639,13 +654,6 @@ async def run_rule_arm(
# ambient: the source is in `rule_usage.RANKED_SOURCES`.
rule_ids = [rule.id for _score, rule in fresh]
source = arm.source
if rule_ids:
try:
io.record_rule_surfaced(
user_id=moment.user_id, rule_ids=rule_ids, source=source,
)
except Exception: # noqa: BLE001 - observation never breaks the observed
_telemetry_failed(source)
checkpoint = (
checkpoint_for(
kept, held=set(moment.held), floor=checkpoint_floor,
@@ -653,6 +661,15 @@ async def run_rule_arm(
)
if arm.checkpoint else {}
)
if arm.stop_only:
rule_ids = [rid for rid in rule_ids if rid == checkpoint.get("rule_id")]
if rule_ids:
try:
io.record_rule_surfaced(
user_id=moment.user_id, rule_ids=rule_ids, source=source,
)
except Exception: # noqa: BLE001 - observation never breaks the observed
_telemetry_failed(source)
return RuleResult(
lines=lines, rule_ids=rule_ids,
shown_rule_ids=[rule.id for _score, rule in shown],
@@ -135,6 +135,9 @@ POINTS: dict[str, Point] = dict([
_p("write_path_rule", UNBIDDEN, "rules that may govern the file being written"),
_p("pre_tool_rule", UNBIDDEN, "rules that may govern a command about to run"),
_p("prompt_rule", UNBIDDEN, "rules that may govern what the operator just asked"),
_p("reply_rule", UNBIDDEN,
"a rule that holds the finished reply for one read — the backstop "
"for whatever the earlier arms missed"),
_p("preference_slot", UNBIDDEN,
"the one line reserved for a preference at the prompt boundary"),
_p("reuse_slot", UNBIDDEN, "the one line reserved for a reusable snippet"),
+15
View File
@@ -207,6 +207,21 @@ SURFACES: dict[str, Surface] = {
over="preferences",
fires="when a task finishes",
),
# THE REPLY BACKSTOP (milestone 458, folded in from 456 step 8). Its floor
# is a STOP bar, not a hint bar: at the end of a turn nothing can be shown
# beside the reply, so a hit either holds the reply for one read or says
# nothing. Hence a default at the checkpoint's level and a budget of one —
# the call row's results are then exactly the rule that would hold.
"reply_rule": Surface(
name="reply_rule",
floor_key="kb_replyrule_threshold",
floor_default=0.80,
budget_key="kb_replyrule_top_k",
budget_default=1,
asks="the reply that ends a turn, against rule triggers",
over="global rules plus the bound project's own",
fires="once per turn, when the reply is finished",
),
}
# Reserved slots are deliberately absent. `preference_slot`, `reuse_slot` and
+4
View File
@@ -101,6 +101,10 @@ logger = logging.getLogger(__name__)
# Add a source here only when a ranker picked it.
RANKED_SOURCES = (
"write_path_rule", "pre_tool_rule", "prompt_rule",
# The reply backstop records only the rule that HELD the reply — a claim
# put in front of the reader as plainly as any line, and one a pull
# confirms or refutes the same way.
"reply_rule",
# A rule mounted on a moment (milestone 458). Not a ranker's pick, but
# not bulk either: somebody decided this rule applies at this moment, and
# that is exactly the claim a pull can confirm or refute. Its