feat(plugin): the seam that erases the evidence is where the unresolved rules get named (#4216)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m46s
CI & Build / Build & push image (push) Successful in 14s

Step 5 of milestone 419. The milestone's subject is that a rule read and
ignored is arithmetically identical to a rule read and followed, and the
compaction is where that identity becomes permanent — the turns holding the
evidence are summarised away, and the unjudged thing survives as nothing.

WHY THIS IS ASSEMBLED IN THE HOOK. `rule_usage_events` has no session column;
it is per user over a window. A session-scoped answer therefore cannot be
asked of the server, and has to be built where a session is a thing that
exists. Four ledgers four hooks already write:

  .rules.ids      an arm NAMED the rule
  .opened.ids     the session called get_rule      (#4100)
  .acted.ids      the session called rule_outcome  (new here)
  .checkpoint.ids the rule HELD an act             (#4214)

Every one is an observed tool call. Nothing asks the model what it followed —
milestone 386 ruled that out, because a model asked "did you apply rule 156?"
says yes. Two subtractions: named-minus-opened is the arm talking to nobody,
opened-minus-acted is the milestone's whole subject.

PreCompact stdout is the compaction's custom instructions (#3680), not a
message to the model, so the readout does not say "you slipped" — it says
which ids must be carried through, which is the one thing a summary can do
about an unjudged finding.

SILENT WHEN NOTHING HAPPENED, and the accusations are conditional on having
members. "0 rules unresolved" on every compaction is how a readout teaches
its reader to skip it. Traffic is still reported, because the static
instructions already ask for it in prose; these lines are the measured
version.

scribe_record_outcome.sh is the third ledger's writer, matched on
mcp__.*__rule_outcome and mirroring scribe_record_opened.sh: TMPDIR only,
silent, exit 0 on every path. A PostToolUse hook that spoke would put a line
after every rule_outcome call and give recording an outcome a cost.

Also: check_plugin.py skipped the new hook for want of a smoke event, which
would have left the newest of the three ledgers as the only one the plugin
lane never runs. Added, mirroring its sibling.

tests/test_precompact_hook.py now isolates TMPDIR — the hook reads session
ledgers from there, so without isolation a test would see whatever this real
session had accumulated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
2026-09-21 00:53:29 -04:00
co-authored by Claude Opus 5
parent fdfb2d94ac
commit ae773740b4
8 changed files with 484 additions and 9 deletions
+92
View File
@@ -525,6 +525,98 @@ scribe_contract_block() {
return 0
}
# ── The session slippage readout (#4216, milestone 419) ───────────────────
#
# WHAT IT ANSWERS, AND WHY IT CANNOT BE ASKED OF THE SERVER. Which rules fired
# this session, which changed an action, which did not. `rule_usage_events`
# carries no session column — it is per user over a window — so a
# session-scoped answer has to be assembled where a session is a thing that
# exists. That is here, from ledgers four hooks already write:
#
# .rules.ids an arm NAMED the rule (a teaser was shown)
# .opened.ids the session called get_rule (#4100, an observed event)
# .acted.ids the session called rule_outcome (#4216)
# .checkpoint.ids the rule HELD an act (#4214, the strongest)
#
# EVERY LINE IS AN OBSERVED TOOL CALL. Nothing here asks the model what it
# followed — milestone 386 ruled that out, and rightly: a model asked "did you
# apply rule 156?" will say yes. These four files record what HAPPENED.
#
# THE SUBTRACTIONS ARE THE POINT. Named-minus-opened is the arm talking to
# nobody; opened-minus-acted is the milestone's whole subject, a rule read and
# then indistinguishable from one that worked. Neither is an accusation — a
# rule may be read and correctly judged not to apply — which is why the lines
# below ask for the leftovers to be CARRIED, not explained.
scribe_ledger_ids() {
# Live ids from one ledger as space-separated words, for set arithmetic.
# Aged through scribe_rules_live so a rule named two hours ago does not read
# as something this context still holds; a bare-id ledger (the checkpoint
# one) simply has no stamps and survives the ageing unchanged.
local f="$1"
[ -n "$f" ] && [ -f "$f" ] || return 0
scribe_rules_live "$f" | tr ',' ' '
}
scribe_slippage_lines() {
# $1 state dir, $2 sanitised session id.
#
# SILENT ONLY WHEN NO RULE TOUCHED THE SESSION AT ALL. Traffic that did
# happen is always reported, because "which rules governed this work" is
# what the static instructions above already ask the summariser to preserve
# in prose — these lines are the measured version of that, and they are
# three short lines.
#
# What is conditional is the ACCUSATION. Each subtraction prints only when
# it has members, so a session that opened everything it was shown and
# resolved everything it opened gets the traffic and no more. "0 rules
# unresolved" on every compaction is how a readout teaches its reader to
# skip it.
local dir="$1" sid="$2" named opened acted held
[ -n "$dir" ] && [ -n "$sid" ] || return 0
named=$(scribe_ledger_ids "$dir/${sid}.rules.ids")
opened=$(scribe_ledger_ids "$dir/${sid}.opened.ids")
acted=$(scribe_ledger_ids "$dir/${sid}.acted.ids")
held=$(scribe_ledger_ids "$dir/${sid}.checkpoint.ids")
[ -n "$named$opened" ] || return 0
local unread unresolved
unread=$(scribe_ids_minus "$named" "$opened")
unresolved=$(scribe_ids_minus "$opened" "$acted")
printf -- '- This session'"'"'s rule traffic, from what actually happened rather than from recollection:\n'
[ -n "$opened" ] && printf -- ' read: %s\n' "$opened"
[ -n "$held" ] && printf -- ' held an act before it ran: %s\n' "$held"
[ -n "$unread" ] && printf -- ' named by an arm and never opened: %s\n' "$unread"
if [ -n "$unresolved" ]; then
printf -- ' READ WITH NO OUTCOME RECORDED: %s. Carry these over as outstanding. A rule read and left unresolved looks exactly like one that worked, and the summary is where that difference is lost for good — say `rule_outcome(id, "applied")`, or `rule_outcome(id, "departed", why=...)` where you deliberately went another way.\n' "$unresolved"
fi
return 0
}
scribe_ids_minus() {
# Set difference over two space-separated id lists, order preserved. Written
# as one awk pass rather than a nested shell loop because the ledgers can
# hold a few dozen ids by the end of a long session and this runs inside the
# compaction path, where a slow hook delays the thing it is decorating.
local a="$1" b="$2"
[ -n "$a" ] || return 0
awk -v a="$a" -v b="$b" '
BEGIN {
n = split(b, drop, " ")
for (i = 1; i <= n; i++) if (drop[i] != "") skip[drop[i]] = 1
m = split(a, keep, " ")
out = ""
for (i = 1; i <= m; i++) {
id = keep[i]
if (id == "" || (id in skip) || (id in done)) continue
done[id] = 1
out = out (out == "" ? "" : " ") id
}
if (out != "") print out
}
' </dev/null 2>/dev/null
}
scribe_local_dups() {
local root="$1" rel="$2" kind name pat hits count label files
while IFS=$'\t' read -r kind name; do