feat(plugin): the seam that erases the evidence is where the unresolved rules get named (#4216)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m46s
CI & Build / Build & push image (push) Successful in 14s

Step 5 of milestone 419. The milestone's subject is that a rule read and
ignored is arithmetically identical to a rule read and followed, and the
compaction is where that identity becomes permanent — the turns holding the
evidence are summarised away, and the unjudged thing survives as nothing.

WHY THIS IS ASSEMBLED IN THE HOOK. `rule_usage_events` has no session column;
it is per user over a window. A session-scoped answer therefore cannot be
asked of the server, and has to be built where a session is a thing that
exists. Four ledgers four hooks already write:

  .rules.ids      an arm NAMED the rule
  .opened.ids     the session called get_rule      (#4100)
  .acted.ids      the session called rule_outcome  (new here)
  .checkpoint.ids the rule HELD an act             (#4214)

Every one is an observed tool call. Nothing asks the model what it followed —
milestone 386 ruled that out, because a model asked "did you apply rule 156?"
says yes. Two subtractions: named-minus-opened is the arm talking to nobody,
opened-minus-acted is the milestone's whole subject.

PreCompact stdout is the compaction's custom instructions (#3680), not a
message to the model, so the readout does not say "you slipped" — it says
which ids must be carried through, which is the one thing a summary can do
about an unjudged finding.

SILENT WHEN NOTHING HAPPENED, and the accusations are conditional on having
members. "0 rules unresolved" on every compaction is how a readout teaches
its reader to skip it. Traffic is still reported, because the static
instructions already ask for it in prose; these lines are the measured
version.

scribe_record_outcome.sh is the third ledger's writer, matched on
mcp__.*__rule_outcome and mirroring scribe_record_opened.sh: TMPDIR only,
silent, exit 0 on every path. A PostToolUse hook that spoke would put a line
after every rule_outcome call and give recording an outcome a cost.

Also: check_plugin.py skipped the new hook for want of a smoke event, which
would have left the newest of the three ledgers as the only one the plugin
lane never runs. Added, mirroring its sibling.

tests/test_precompact_hook.py now isolates TMPDIR — the hook reads session
ledgers from there, so without isolation a test would see whatever this real
session had accumulated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
2026-09-21 00:53:29 -04:00
co-authored by Claude Opus 5
parent fdfb2d94ac
commit ae773740b4
8 changed files with 484 additions and 9 deletions
+57
View File
@@ -0,0 +1,57 @@
#!/usr/bin/env bash
# Scribe — record that this session said what a rule DID, not merely that it
# read one (#4216, milestone 419).
#
# THE THIRD LEDGER, AND WHY THE TWO THAT EXIST ARE NOT ENOUGH.
#
# `.rules.ids` says a rule was NAMED. `.opened.ids` says it was READ (#4100 —
# a PostToolUse hook watches the `get_rule` call, so it is a recorded event
# rather than a model's claim about its own context). Neither can say what
# happened next, and that is the whole of what milestone 419 is about: a rule
# read and followed and a rule read and forgotten leave identical traces.
#
# `rule_outcome` (#4212) is the call that closes that gap, and the server
# records it. But the server's row carries no session — `rule_usage_events` is
# per user over a window — so a SESSION-scoped readout cannot be asked of it.
# The question "which rules changed something in THIS session" has to be
# answered where a session is a thing that exists, which is here.
#
# SAME EVIDENCE CLASS AS `.opened.ids`, deliberately. A tool call happened or
# it did not, and the harness reports it either way; nothing here asks the
# model whether it followed anything. That is the distinction milestone 386
# drew when it ruled out self-report, and this stays on the right side of it.
#
# WHAT IT CANNOT SAY: that the rule was followed WELL, or that `applied` was
# honest. It records that an outcome was declared. The value of that is not the
# claim itself — it is that the absence of one becomes visible, which is the
# state nothing could previously name.
#
# EXIT 0, ALWAYS. This decorates a ledger; a bookkeeping failure must never
# turn a successful tool call into a hook error.
set -uo pipefail
# shellcheck source=plugin/hooks/scribe_defs.sh
. "$(dirname "${BASH_SOURCE[0]}")/scribe_defs.sh"
event=$(cat 2>/dev/null || true)
[ -n "$event" ] || exit 0
event_flat=$(printf '%s' "$event" | scribe_json_flat)
session_id=$(scribe_json_pick "$event_flat" '.session_id')
[ -n "$session_id" ] || exit 0
# The matcher in hooks.json narrows to the rule_outcome tools, but the server
# segment of an MCP tool name varies with how the plugin was installed, so the
# id is read from the field rather than from an assumed tool name — the same
# reasoning scribe_record_opened.sh gives.
rule_id=$(scribe_json_pick "$event_flat" '.tool_input.rule_id')
rule_id=$(printf '%s' "$rule_id" | tr -cd '0-9')
[ -n "$rule_id" ] || exit 0
state_dir="${TMPDIR:-/tmp}/scribe-priorart"
mkdir -p "$state_dir" 2>/dev/null || true
safe_sid=$(printf '%s' "$session_id" | tr -c 'A-Za-z0-9._-' '_')
# Stamped and append-only like its two siblings, so one reader ages them all.
printf '%s\n' "$rule_id" | scribe_rules_append "$state_dir/${safe_sid}.acted.ids"
exit 0