feat(moments): the reply moment holds a finished reply for one read (milestone 458 step 4b, #4922)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Failing after 1m25s
CI & Build / Build & push image (push) Skipped

The reply is the one act no tool call marks, and it is where "let me
know if it works" gets said. A new Stop hook (scribe_reply_check.sh)
sends the finished reply to POST /api/plugin/reply-rules, which checks
it twice:

- mounted: every unopened RULE on reply.report, plus reply.ask when the
  reply asks a question. Deterministic.
- semantic: the reply's head and tail against every rule's trigger, on a
  new ranked surface, reply_rule. It is the backstop for whatever the
  earlier arms missed. Its floor is its stop bar (default 0.80, budget
  1), with its own Settings dials. The new stop_only stage records
  surfacing for the rule that holds and nothing else, because nothing
  else reached anyone.

Following the operator's ruling from 456 step 8, a rule that holds blocks
once, in the server's words. The hook blocks only on a reason it was
given, so an unreachable instance never stops a session, and it never
holds the rewrite. The ledger is the act checkpoint's own, so a rule
holds a session once across both doors and the per-session cap counts
both.

The turn reader moved from the report check into scribe_defs.sh
(scribe_turn_facts / scribe_turn_fact), so the two Stop hooks read a
turn the same way. The output was checked identical on a real transcript.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-10-05 12:30:07 -04:00
co-authored by Claude Opus 5.5
parent b3b616b20a
commit c6cdfc2172
17 changed files with 740 additions and 59 deletions
+25 -8
View File
@@ -354,6 +354,14 @@ class RuleArm:
kind: str | None = None
"""Ask the ranker for one record kind only ("preference"), or every kind."""
stop_only: bool = False
"""The door shows nothing but the stop (milestone 458's reply moment).
At the end of a turn there is no context to annotate: a hit either holds
the reply for one read or reaches nobody. So only the checkpoint's rule is
recorded as surfaced — the rest were ranked, and the call row counts them,
but nobody was shown them."""
# Today's differences, reproduced exactly (milestone 456 step 2). Whether the
# prompt arm should band, and whether the act arms should reserve a
@@ -377,8 +385,15 @@ REPORT_PREFERENCE = RuleArm(
"report_preference", band=False, compact_tail=False, checkpoint=False,
preference_slot=False, kind="preference",
)
# The backstop for every arm that ran earlier in the turn and missed: the
# finished reply against every rule's trigger. Its floor IS its stop bar
# (the reply_rule surface), so it is passed as both.
REPLY_RULE = RuleArm(
"reply_rule", band=False, compact_tail=False, checkpoint=True,
preference_slot=False, stop_only=True,
)
RULE_ARMS: tuple[RuleArm, ...] = (
WRITE_PATH_RULE, PRE_TOOL_RULE, PROMPT_RULE, REPORT_PREFERENCE,
WRITE_PATH_RULE, PRE_TOOL_RULE, PROMPT_RULE, REPORT_PREFERENCE, REPLY_RULE,
)
PREFERENCE_SLOT_SOURCE = "preference_slot"
@@ -639,13 +654,6 @@ async def run_rule_arm(
# ambient: the source is in `rule_usage.RANKED_SOURCES`.
rule_ids = [rule.id for _score, rule in fresh]
source = arm.source
if rule_ids:
try:
io.record_rule_surfaced(
user_id=moment.user_id, rule_ids=rule_ids, source=source,
)
except Exception: # noqa: BLE001 - observation never breaks the observed
_telemetry_failed(source)
checkpoint = (
checkpoint_for(
kept, held=set(moment.held), floor=checkpoint_floor,
@@ -653,6 +661,15 @@ async def run_rule_arm(
)
if arm.checkpoint else {}
)
if arm.stop_only:
rule_ids = [rid for rid in rule_ids if rid == checkpoint.get("rule_id")]
if rule_ids:
try:
io.record_rule_surfaced(
user_id=moment.user_id, rule_ids=rule_ids, source=source,
)
except Exception: # noqa: BLE001 - observation never breaks the observed
_telemetry_failed(source)
return RuleResult(
lines=lines, rule_ids=rule_ids,
shown_rule_ids=[rule.id for _score, rule in shown],