feat(moments): the reply moment holds a finished reply for one read (milestone 458 step 4b, #4922)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Failing after 1m25s
CI & Build / Build & push image (push) Skipped
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Failing after 1m25s
CI & Build / Build & push image (push) Skipped
The reply is the one act no tool call marks, and it is where "let me know if it works" gets said. A new Stop hook (scribe_reply_check.sh) sends the finished reply to POST /api/plugin/reply-rules, which checks it twice: - mounted: every unopened RULE on reply.report, plus reply.ask when the reply asks a question. Deterministic. - semantic: the reply's head and tail against every rule's trigger, on a new ranked surface, reply_rule. It is the backstop for whatever the earlier arms missed. Its floor is its stop bar (default 0.80, budget 1), with its own Settings dials. The new stop_only stage records surfacing for the rule that holds and nothing else, because nothing else reached anyone. Following the operator's ruling from 456 step 8, a rule that holds blocks once, in the server's words. The hook blocks only on a reason it was given, so an unreachable instance never stops a session, and it never holds the rewrite. The ledger is the act checkpoint's own, so a rule holds a session once across both doors and the per-session cap counts both. The turn reader moved from the report check into scribe_defs.sh (scribe_turn_facts / scribe_turn_fact), so the two Stop hooks read a turn the same way. The output was checked identical on a real transcript. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -294,6 +294,42 @@ async def moment():
|
||||
})
|
||||
|
||||
|
||||
@plugin_bp.post("/reply-rules")
|
||||
@login_required
|
||||
@memoized_query_embeddings
|
||||
async def reply_rules():
|
||||
"""Whether the reply that ends a turn is held for one read (milestone 458).
|
||||
|
||||
Called by the plugin's Stop hook with the finished reply. Two halves: the
|
||||
rules MOUNTED on the reply moments (`reply.report`, and `reply.ask` when
|
||||
the reply asks something), and the reply text against every rule's
|
||||
trigger — the backstop for whatever the earlier arms missed.
|
||||
|
||||
Body: `{"reply": "<text>"}`. Query: `repo` / `project_id` for the scope;
|
||||
`held_rule_ids` (opened this session — exempt, as at the act checkpoint);
|
||||
`exclude_rule_ids` (named this session — not re-counted as surfaced);
|
||||
`stopped_rule_ids` (already held a reply or an act this session — a rule
|
||||
holds once, and the per-session cap counts these).
|
||||
|
||||
Returns `reason` — the words the hook blocks with, empty when nothing
|
||||
holds — plus `rule_ids` for the hook's ledger and `moments` reached.
|
||||
"""
|
||||
data = await request.get_json(silent=True) or {}
|
||||
reply = str(data.get("reply") or "")
|
||||
project_id, _repo, _unbound = await _project_scope()
|
||||
held = await moment_delivery_svc.reply_hold(
|
||||
g.user.id, reply, project_id=project_id,
|
||||
exclude=frozenset(_int_list(request.args.get("exclude_rule_ids"))),
|
||||
held=frozenset(_int_list(request.args.get("held_rule_ids"))),
|
||||
stopped=frozenset(_int_list(request.args.get("stopped_rule_ids"))),
|
||||
)
|
||||
return jsonify({
|
||||
"reason": held.get("reason", ""),
|
||||
"rule_ids": held.get("rule_ids", []),
|
||||
"moments": held.get("moments", []),
|
||||
})
|
||||
|
||||
|
||||
@plugin_bp.get("/prior-art")
|
||||
@login_required
|
||||
@memoized_query_embeddings
|
||||
|
||||
@@ -117,3 +117,153 @@ async def attach_moment_rules(
|
||||
except Exception: # noqa: BLE001 - a decoration never breaks the payload
|
||||
logger.debug("moment rules for %s could not be attached", tool, exc_info=True)
|
||||
return data
|
||||
|
||||
|
||||
# ── The reply moment: the backstop at the end of a turn ─────────────────
|
||||
#
|
||||
# The Stop hook's door, folded in from milestone 456 step 8. A reply is the
|
||||
# one act no tool call marks, and the moment where "please verify this on
|
||||
# your end" gets said — so it is checked twice over: the rules MOUNTED on the
|
||||
# reply moments (deterministic, somebody said they belong here), and the
|
||||
# reply text against every rule's trigger (the backstop for whatever the
|
||||
# earlier arms missed). An unopened rule from either half HOLDS the reply for
|
||||
# one read, in these words; the hook never holds the rewrite.
|
||||
|
||||
# The reply's head and tail. The embedder reads ~512 tokens, and the part of
|
||||
# a report that asks something of the reader — "let me know if it works" —
|
||||
# is at the END, where a head-only query would cut it off.
|
||||
_REPLY_HEAD_CHARS = 600
|
||||
_REPLY_TAIL_CHARS = 1200
|
||||
|
||||
_REACHED_BY_REPLY = "the reply that ends this turn"
|
||||
|
||||
|
||||
def reply_moments(reply: str) -> list[dict]:
|
||||
"""The moments a finished reply reaches: always `reply.report`, and
|
||||
`reply.ask` when a line of its prose ends in a question."""
|
||||
reached = [{"moment": "reply.report", "tool": "", "match": _REACHED_BY_REPLY,
|
||||
"via": moment_actions.DEFAULT}]
|
||||
in_code = False
|
||||
for line in (reply or "").splitlines():
|
||||
stripped = line.strip()
|
||||
if stripped.startswith("```"):
|
||||
in_code = not in_code
|
||||
continue
|
||||
if not in_code and stripped.endswith("?"):
|
||||
reached.append({"moment": "reply.ask", "tool": "", "match": "a question in the reply",
|
||||
"via": moment_actions.DEFAULT})
|
||||
break
|
||||
return reached
|
||||
|
||||
|
||||
def _reply_query(reply: str) -> str:
|
||||
text = " ".join((reply or "").split())
|
||||
if len(text) <= _REPLY_HEAD_CHARS + _REPLY_TAIL_CHARS:
|
||||
return text
|
||||
return text[:_REPLY_HEAD_CHARS] + " … " + text[-_REPLY_TAIL_CHARS:]
|
||||
|
||||
|
||||
def reply_hold_reason(held: list[dict]) -> str:
|
||||
"""The text the agent reads INSTEAD of its reply going out.
|
||||
|
||||
A practice, not a prohibition (rule 165), on the act checkpoint's model:
|
||||
nothing here knows the reply is wrong. It names each rule with why it was
|
||||
raised, gives the remedy as one call per rule, and says the reply may go
|
||||
out unchanged — so a reader who finds the rules beside the point is out in
|
||||
as many calls as there are rules, and is never held twice.
|
||||
"""
|
||||
if not held:
|
||||
return ""
|
||||
parts = []
|
||||
for item in held:
|
||||
why = (f"mounted on {item['moment']}" if item.get("moment")
|
||||
else f"scores {item['score']} against this reply")
|
||||
trigger = f"; it applies when {item['trigger']}" if item.get("trigger") else ""
|
||||
parts.append(f"“{item['title']}” ({why}{trigger}) — get_rule({item['rule_id']})")
|
||||
noun = "a standing rule" if len(held) == 1 else "standing rules"
|
||||
return (
|
||||
f"Held for one read before this reply goes out: {noun} this session "
|
||||
f"has not opened: " + "; ".join(parts) + ". Read "
|
||||
+ ("it" if len(held) == 1 else "them")
|
||||
+ ", then send the reply — unchanged if it already does what the rule "
|
||||
"asks, which is a judgement only you can make. This check runs once "
|
||||
"per rule; the rewrite is never held."
|
||||
)
|
||||
|
||||
|
||||
async def reply_hold(
|
||||
user_id: int, reply: str, *, project_id: int | None = None,
|
||||
exclude: frozenset[int] = frozenset(), held: frozenset[int] = frozenset(),
|
||||
stopped: frozenset[int] = frozenset(),
|
||||
) -> dict:
|
||||
"""Whether this finished reply is held, and in what words.
|
||||
|
||||
`held` is what the session OPENED (the act checkpoint's exemption);
|
||||
`stopped` is what an earlier hold already put in front of it, so a rule
|
||||
holds a session once. Past the act checkpoint's per-session cap nothing
|
||||
holds — a mis-set bar degrades to a quiet session, never a stuck one.
|
||||
|
||||
Returns {} or {"reason", "rule_ids", "moments"}. Fails open to {}.
|
||||
"""
|
||||
from scribe.services import plugin_context as pc
|
||||
from scribe.services import rule_usage, rulebooks
|
||||
from scribe.services.retrieval_surfaces import budget_for, floor_for
|
||||
|
||||
try:
|
||||
if not (reply or "").strip() or len(stopped) >= pc.CHECKPOINT_SESSION_CAP:
|
||||
return {}
|
||||
reached = reply_moments(reply)
|
||||
names = [hit["moment"] for hit in reached]
|
||||
skip = set(held) | set(stopped)
|
||||
room = pc.CHECKPOINT_SESSION_CAP - len(stopped)
|
||||
out: list[dict] = []
|
||||
|
||||
# The mounted half: deterministic, so every unopened RULE on these
|
||||
# moments holds. A preference claims no such force.
|
||||
for rule, at in await rulebooks.rules_on_moments(user_id, names, project_id or None):
|
||||
if rule.id in skip or rule.kind == "preference":
|
||||
continue
|
||||
out.append({"rule_id": rule.id, "title": rule.title, "moment": at,
|
||||
"trigger": (rule.when_to_apply or "").strip()})
|
||||
# Trimmed BEFORE recording: a rule the cap cut was shown to nobody.
|
||||
out = out[:room]
|
||||
mounted = [item["rule_id"] for item in out]
|
||||
fresh = [rid for rid in mounted if rid not in exclude]
|
||||
if fresh:
|
||||
try:
|
||||
rule_usage.record_rule_surfaced(
|
||||
user_id=user_id, rule_ids=fresh, source=rp.MOMENT_RULE_SOURCE,
|
||||
detail={item["rule_id"]: item["moment"] for item in out
|
||||
if item["rule_id"] in fresh},
|
||||
)
|
||||
except Exception: # noqa: BLE001 - observation never breaks the observed
|
||||
logger.debug("reply mount surfacing not recorded", exc_info=True)
|
||||
|
||||
# The semantic half: the backstop. A mounted rule already holding is
|
||||
# treated as held here, so one rule is never named twice.
|
||||
floor = await floor_for(user_id, rp.REPLY_RULE.source)
|
||||
result = await rp.run_rule_arm(
|
||||
rp.REPLY_RULE,
|
||||
rp.RuleMoment(
|
||||
user_id=user_id, query=_reply_query(reply), project_id=project_id,
|
||||
where="to this reply", checkpoint_where="this reply",
|
||||
exclude=exclude, held=frozenset(skip | set(mounted)),
|
||||
),
|
||||
floor=floor, budget=await budget_for(user_id, rp.REPLY_RULE.source),
|
||||
io=pc._rule_io(), checkpoint_floor=floor,
|
||||
)
|
||||
if result.checkpoint and len(out) < room:
|
||||
cp = result.checkpoint
|
||||
out.append({"rule_id": cp["rule_id"], "title": cp["title"],
|
||||
"score": cp["score"], "trigger": cp.get("trigger", "")})
|
||||
|
||||
if not out:
|
||||
return {}
|
||||
return {
|
||||
"reason": reply_hold_reason(out),
|
||||
"rule_ids": [item["rule_id"] for item in out],
|
||||
"moments": names,
|
||||
}
|
||||
except Exception: # noqa: BLE001 - a recall aid never stops a session by failing
|
||||
logger.debug("reply hold failed", exc_info=True)
|
||||
return {}
|
||||
|
||||
@@ -115,6 +115,7 @@ _RESCORERS = {
|
||||
),
|
||||
"write_path_rule": lambda u, q, p: _rescore_rules(u, q, p, None),
|
||||
"pre_tool_rule": lambda u, q, p: _rescore_rules(u, q, p, None),
|
||||
"reply_rule": lambda u, q, p: _rescore_rules(u, q, p, None),
|
||||
"prompt_rule": lambda u, q, p: _rescore_rules(u, q, p, None),
|
||||
"report_preference": lambda u, q, p: _rescore_rules(u, q, p, "preference"),
|
||||
}
|
||||
@@ -128,6 +129,7 @@ _CORPUS = {
|
||||
"write_path": NoteEmbedding,
|
||||
"write_path_rule": RuleEmbedding,
|
||||
"pre_tool_rule": RuleEmbedding,
|
||||
"reply_rule": RuleEmbedding,
|
||||
"prompt_rule": RuleEmbedding,
|
||||
"report_preference": RuleEmbedding,
|
||||
}
|
||||
|
||||
@@ -354,6 +354,14 @@ class RuleArm:
|
||||
kind: str | None = None
|
||||
"""Ask the ranker for one record kind only ("preference"), or every kind."""
|
||||
|
||||
stop_only: bool = False
|
||||
"""The door shows nothing but the stop (milestone 458's reply moment).
|
||||
|
||||
At the end of a turn there is no context to annotate: a hit either holds
|
||||
the reply for one read or reaches nobody. So only the checkpoint's rule is
|
||||
recorded as surfaced — the rest were ranked, and the call row counts them,
|
||||
but nobody was shown them."""
|
||||
|
||||
|
||||
# Today's differences, reproduced exactly (milestone 456 step 2). Whether the
|
||||
# prompt arm should band, and whether the act arms should reserve a
|
||||
@@ -377,8 +385,15 @@ REPORT_PREFERENCE = RuleArm(
|
||||
"report_preference", band=False, compact_tail=False, checkpoint=False,
|
||||
preference_slot=False, kind="preference",
|
||||
)
|
||||
# The backstop for every arm that ran earlier in the turn and missed: the
|
||||
# finished reply against every rule's trigger. Its floor IS its stop bar
|
||||
# (the reply_rule surface), so it is passed as both.
|
||||
REPLY_RULE = RuleArm(
|
||||
"reply_rule", band=False, compact_tail=False, checkpoint=True,
|
||||
preference_slot=False, stop_only=True,
|
||||
)
|
||||
RULE_ARMS: tuple[RuleArm, ...] = (
|
||||
WRITE_PATH_RULE, PRE_TOOL_RULE, PROMPT_RULE, REPORT_PREFERENCE,
|
||||
WRITE_PATH_RULE, PRE_TOOL_RULE, PROMPT_RULE, REPORT_PREFERENCE, REPLY_RULE,
|
||||
)
|
||||
|
||||
PREFERENCE_SLOT_SOURCE = "preference_slot"
|
||||
@@ -639,13 +654,6 @@ async def run_rule_arm(
|
||||
# ambient: the source is in `rule_usage.RANKED_SOURCES`.
|
||||
rule_ids = [rule.id for _score, rule in fresh]
|
||||
source = arm.source
|
||||
if rule_ids:
|
||||
try:
|
||||
io.record_rule_surfaced(
|
||||
user_id=moment.user_id, rule_ids=rule_ids, source=source,
|
||||
)
|
||||
except Exception: # noqa: BLE001 - observation never breaks the observed
|
||||
_telemetry_failed(source)
|
||||
checkpoint = (
|
||||
checkpoint_for(
|
||||
kept, held=set(moment.held), floor=checkpoint_floor,
|
||||
@@ -653,6 +661,15 @@ async def run_rule_arm(
|
||||
)
|
||||
if arm.checkpoint else {}
|
||||
)
|
||||
if arm.stop_only:
|
||||
rule_ids = [rid for rid in rule_ids if rid == checkpoint.get("rule_id")]
|
||||
if rule_ids:
|
||||
try:
|
||||
io.record_rule_surfaced(
|
||||
user_id=moment.user_id, rule_ids=rule_ids, source=source,
|
||||
)
|
||||
except Exception: # noqa: BLE001 - observation never breaks the observed
|
||||
_telemetry_failed(source)
|
||||
return RuleResult(
|
||||
lines=lines, rule_ids=rule_ids,
|
||||
shown_rule_ids=[rule.id for _score, rule in shown],
|
||||
|
||||
@@ -135,6 +135,9 @@ POINTS: dict[str, Point] = dict([
|
||||
_p("write_path_rule", UNBIDDEN, "rules that may govern the file being written"),
|
||||
_p("pre_tool_rule", UNBIDDEN, "rules that may govern a command about to run"),
|
||||
_p("prompt_rule", UNBIDDEN, "rules that may govern what the operator just asked"),
|
||||
_p("reply_rule", UNBIDDEN,
|
||||
"a rule that holds the finished reply for one read — the backstop "
|
||||
"for whatever the earlier arms missed"),
|
||||
_p("preference_slot", UNBIDDEN,
|
||||
"the one line reserved for a preference at the prompt boundary"),
|
||||
_p("reuse_slot", UNBIDDEN, "the one line reserved for a reusable snippet"),
|
||||
|
||||
@@ -207,6 +207,21 @@ SURFACES: dict[str, Surface] = {
|
||||
over="preferences",
|
||||
fires="when a task finishes",
|
||||
),
|
||||
# THE REPLY BACKSTOP (milestone 458, folded in from 456 step 8). Its floor
|
||||
# is a STOP bar, not a hint bar: at the end of a turn nothing can be shown
|
||||
# beside the reply, so a hit either holds the reply for one read or says
|
||||
# nothing. Hence a default at the checkpoint's level and a budget of one —
|
||||
# the call row's results are then exactly the rule that would hold.
|
||||
"reply_rule": Surface(
|
||||
name="reply_rule",
|
||||
floor_key="kb_replyrule_threshold",
|
||||
floor_default=0.80,
|
||||
budget_key="kb_replyrule_top_k",
|
||||
budget_default=1,
|
||||
asks="the reply that ends a turn, against rule triggers",
|
||||
over="global rules plus the bound project's own",
|
||||
fires="once per turn, when the reply is finished",
|
||||
),
|
||||
}
|
||||
|
||||
# Reserved slots are deliberately absent. `preference_slot`, `reuse_slot` and
|
||||
|
||||
@@ -101,6 +101,10 @@ logger = logging.getLogger(__name__)
|
||||
# Add a source here only when a ranker picked it.
|
||||
RANKED_SOURCES = (
|
||||
"write_path_rule", "pre_tool_rule", "prompt_rule",
|
||||
# The reply backstop records only the rule that HELD the reply — a claim
|
||||
# put in front of the reader as plainly as any line, and one a pull
|
||||
# confirms or refutes the same way.
|
||||
"reply_rule",
|
||||
# A rule mounted on a moment (milestone 458). Not a ranker's pick, but
|
||||
# not bulk either: somebody decided this rule applies at this moment, and
|
||||
# that is exactly the claim a pull can confirm or refute. Its
|
||||
|
||||
Reference in New Issue
Block a user