feat(500): one end-of-turn request - the completion-section check folds into the reply check
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m7s
CI & Build / Python tests (push) Failing after 1m25s
CI & Build / Build & push image (push) Skipped

The Stop hook sent the finished reply twice: scribe_report_check.sh checked a
task-closing reply for the completion sections in shell and reported to
/report-check, and scribe_reply_check.sh sent the same reply to /reply-rules
for the rule hold. Now the reply goes once. When the turn closed a task the
hook adds the close count and ids, and the server runs the section check
(services/report_check, the same three patterns) beside the reply hold,
folding both into one reason.

The report_check adherence log is still written for every checked reply
(milestone 409's number). A section hold marks the session, so its rewrite is
sent back once with rewrite:true to record how it came out, and is never held.
The block reason carries the completion shape's one line and points at
list_reply_shapes rather than at the skill. #5496 (step 4 of milestone 500).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-10-09 15:38:17 -04:00
co-authored by Claude Opus 5.5
parent ae66d83213
commit 6071007fe2
15 changed files with 348 additions and 433 deletions
+1 -1
View File
@@ -1,7 +1,7 @@
{ {
"name": "scribe", "name": "scribe",
"description": "Scribe for Claude Code: connects the scribe MCP server, adds the hooks that deliver live project state and relevant records at the right moment, ships the shared client-neutral Scribe skills (using-scribe, writing-plans, reporting-back, systematic-debugging, verification, brainstorming, reusing-code, shape-accounting, family-canon), and syncs your saved Scribe Processes as skills (/scribe:sync).", "description": "Scribe for Claude Code: connects the scribe MCP server, adds the hooks that deliver live project state and relevant records at the right moment, ships the shared client-neutral Scribe skills (using-scribe, writing-plans, reporting-back, systematic-debugging, verification, brainstorming, reusing-code, shape-accounting, family-canon), and syncs your saved Scribe Processes as skills (/scribe:sync).",
"version": "2026.10.09.1850", "version": "2026.10.09.1938",
"author": { "author": {
"name": "Bryan Van Deusen" "name": "Bryan Van Deusen"
}, },
+2 -3
View File
@@ -15,7 +15,7 @@ another one means adding files, not moving or rewriting any.
|---|---|---| |---|---|---|
| **The skills** | `plugin/skills/*/SKILL.md` | Agent Skills (the open SKILL.md format). They state every Scribe reflex in full and name no client. `tests/test_guidance_ownership.py` fails if a skill names a particular client, or references anything outside its own folder. Every client package ships this folder verbatim. | | **The skills** | `plugin/skills/*/SKILL.md` | Agent Skills (the open SKILL.md format). They state every Scribe reflex in full and name no client. `tests/test_guidance_ownership.py` fails if a skill names a particular client, or references anything outside its own folder. Every client package ships this folder verbatim. |
| **The MCP server** | `<base URL>/mcp` | HTTP, `Authorization: Bearer <fmcp_ key>`. Its `_INSTRUCTIONS` is a client-neutral orientation — the workflow across tools (≤1,600 chars, #4389); each tool's description carries its contract; in-band responses (`placement`, `report_back`, `systems_hint`, the duplicate gate, the guessed-id refusal) fire in every client. | | **The MCP server** | `<base URL>/mcp` | HTTP, `Authorization: Bearer <fmcp_ key>`. Its `_INSTRUCTIONS` is a client-neutral orientation — the workflow across tools (≤1,600 chars, #4389); each tool's description carries its contract; in-band responses (`placement`, `report_back`, `systems_hint`, the duplicate gate, the guessed-id refusal) fire in every client. |
| **The adapter API** | `<base URL>/api/plugin/*` | Plain `GET` endpoints any client's hooks can call with the same key (read scope is enough): `context` (live session state), `retrieve` (rules, preferences and notes for a message), `prior-art` (records and shape-ledger hints for code being written), `tool-rules` (rules for a command about to run), `report-check` (records a completion-report check and returns the reason for a block), `processes` (stored Processes to expose as skills). | | **The adapter API** | `<base URL>/api/plugin/*` | Plain `GET` endpoints any client's hooks can call with the same key (read scope is enough): `context` (live session state), `retrieve` (rules, preferences and notes for a message), `prior-art` (records and shape-ledger hints for code being written), `tool-rules` (rules for a command about to run), `reply-rules` (a `POST`: the finished reply, held for an unopened rule or — when the turn closed a task — for missing completion sections), `processes` (stored Processes to expose as skills). |
| **The API key** | Scribe → Settings → API Keys | One `fmcp_` key per install. Read scope for hooks; write scope for the MCP tools. | | **The API key** | Scribe → Settings → API Keys | One `fmcp_` key per install. Read scope for hooks; write scope for the MCP tools. |
## Added by each client ## Added by each client
@@ -40,8 +40,7 @@ another one means adding files, not moving or rewriting any.
| `hooks/scribe_after_write.sh` | PostToolUse on shell commands: the same check for code written through the shell. | | `hooks/scribe_after_write.sh` | PostToolUse on shell commands: the same check for code written through the shell. |
| `hooks/scribe_tool_rules.sh` | PreToolUse on shell commands: `GET /api/plugin/tool-rules`. | | `hooks/scribe_tool_rules.sh` | PreToolUse on shell commands: `GET /api/plugin/tool-rules`. |
| `hooks/scribe_moment.sh` | PreToolUse on every tool: the rules mounted on the moments the call reaches, `POST /api/plugin/moment`; skips tools `GET /api/plugin/moment-tools` says reach nothing mounted, and Scribe's own (their responses carry `moment_rules`). | | `hooks/scribe_moment.sh` | PreToolUse on every tool: the rules mounted on the moments the call reaches, `POST /api/plugin/moment`; skips tools `GET /api/plugin/moment-tools` says reach nothing mounted, and Scribe's own (their responses carry `moment_rules`). |
| `hooks/scribe_report_check.sh` | Stop: when the turn closed a task, checks the reply for the completion sections and reports to `GET /api/plugin/report-check`; blocks once, with the reason the server returns. | | `hooks/scribe_reply_check.sh` | Stop: sends the finished reply to `POST /api/plugin/reply-rules` — the rules mounted on the reply moments, and the reply against every rule's trigger, and — when the turn closed a task — the completion-section check; holds once per rule, with the reason the server returns, and never holds the rewrite. |
| `hooks/scribe_reply_check.sh` | Stop: sends the finished reply to `POST /api/plugin/reply-rules` — the rules mounted on the reply moments, and the reply against every rule's trigger; holds once per rule, with the reason the server returns, and never holds the rewrite. |
| `hooks/scribe_shape_check.sh` | Stop: sends the definitions the turn wrote (the write hooks' `<sid>.written.ids` ledger) to `GET /api/plugin/shape-check`; blocks once, with the reason the server returns, so the agent judges what it built. | | `hooks/scribe_shape_check.sh` | Stop: sends the definitions the turn wrote (the write hooks' `<sid>.written.ids` ledger) to `GET /api/plugin/shape-check`; blocks once, with the reason the server returns, so the agent judges what it built. |
| `hooks/scribe_sync_processes.sh` + `commands/sync.md` | `GET /api/plugin/processes` → `~/.claude/skills/scribe-proc-*` stubs; `/scribe:sync` on demand. | | `hooks/scribe_sync_processes.sh` + `commands/sync.md` | `GET /api/plugin/processes` → `~/.claude/skills/scribe-proc-*` stubs; `/scribe:sync` on demand. |
| `hooks/scribe_defs.sh` | Shared shell helpers: config, dedup ledgers, outage line. | | `hooks/scribe_defs.sh` | Shared shell helpers: config, dedup ledgers, outage line. |
+10 -9
View File
@@ -82,15 +82,16 @@ On install you'll be asked for:
answer" line (8 s budget here — it runs after the tool, so it gates answer" line (8 s budget here — it runs after the tool, so it gates
nothing). The extractor, the prose/data skip list, the local by-name nothing). The extractor, the prose/data skip list, the local by-name
duplicate arm and the outage line are shared in `hooks/scribe_defs.sh`. duplicate arm and the outage line are shared in `hooks/scribe_defs.sh`.
- `hooks/hooks.json` → Stop hook (`hooks/scribe_report_check.sh`): when the - `hooks/hooks.json` → Stop hook (`hooks/scribe_reply_check.sh`): sends the
turn closed a Scribe task (`update_task`/`create_task` with status done), reply that ends the turn to `POST /api/plugin/reply-rules`, which holds it
checks the reply that ends it for the completion sections — where the work once for an unopened rule mounted on the reply moments or matching the reply.
sits, what needs you, what comes next — and reports the outcome to When the turn closed a Scribe task (`update_task`/`create_task` with status
`GET /api/plugin/report-check`. If sections are missing it blocks once with done), the same request has the server check the reply for the completion
the reason the server returns, and records how the rewrite came out; it sections — where the work sits, what needs you, what comes next. If they are
never blocks twice, and never blocks when the instance did not record the missing it blocks once with the reason the server returns, and records how
check (unconfigured or unreachable). Outcomes land in the admin logs under the rewrite came out; it never blocks twice, and never blocks when the
category `plugin`, action `report_check`. instance did not record the check (unconfigured or unreachable). Outcomes
land in the admin logs under category `plugin`, action `report_check`.
- `hooks/hooks.json` → a second Stop hook (`hooks/scribe_shape_check.sh`): the - `hooks/hooks.json` → a second Stop hook (`hooks/scribe_shape_check.sh`): the
write hooks note every definition a write names (and every new file) in a write hooks note every definition a write names (and every new file) in a
session ledger; at the end of the turn this sends them to session ledger; at the end of the turn this sends them to
-4
View File
@@ -104,10 +104,6 @@
"Stop": [ "Stop": [
{ {
"hooks": [ "hooks": [
{
"type": "command",
"command": "bash \"${CLAUDE_PLUGIN_ROOT}/hooks/scribe_report_check.sh\""
},
{ {
"type": "command", "type": "command",
"command": "bash \"${CLAUDE_PLUGIN_ROOT}/hooks/scribe_reply_check.sh\"" "command": "bash \"${CLAUDE_PLUGIN_ROOT}/hooks/scribe_reply_check.sh\""
+2 -2
View File
@@ -135,8 +135,8 @@ scribe_json_flat_lines() {
# Cheap by construction, because it runs on every prompt: the last 512 KB only, # Cheap by construction, because it runs on every prompt: the last 512 KB only,
# and grep narrows to assistant records carrying a text block BEFORE anything is # and grep narrows to assistant records carrying a text block BEFORE anything is
# parsed — a transcript is mostly tool results, and parsing those in awk is the # parsed — a transcript is mostly tool results, and parsing those in awk is the
# multi-second cost scribe_report_check.sh already measured. Fixed strings with # multi-second cost the completion-report check measured (#4107). Fixed strings
# UNESCAPED quotes can only match at a record's own top level (see that hook). # with UNESCAPED quotes can only match at a record's own top level.
# A sidechain is a subagent talking, not this session, so it is skipped. # A sidechain is a subagent talking, not this session, so it is skipped.
scribe_recent_context() { scribe_recent_context() {
[ -n "${1:-}" ] && [ -f "$1" ] || return 0 [ -n "${1:-}" ] && [ -f "$1" ] || return 0
+1 -1
View File
@@ -39,7 +39,7 @@
# #
# mode=lines the input is JSONL — one JSON value per line — and a line that # mode=lines the input is JSONL — one JSON value per line — and a line that
# does not parse is DROPPED, the rest still read. This is the # does not parse is DROPPED, the rest still read. This is the
# transcript in scribe_report_check.sh, where the window starts # transcript the Stop hooks read, where the window starts
# mid-record by construction: `tail -n 3000` cuts wherever it # mid-record by construction: `tail -n 3000` cuts wherever it
# cuts, and the first line is routinely half a record. It is # cuts, and the first line is routinely half a record. It is
# exactly the `map(try fromjson catch empty)` the jq program it # exactly the `map(try fromjson catch empty)` the jq program it
+46 -10
View File
@@ -9,7 +9,15 @@
# text against every rule's trigger, as the backstop for whatever the earlier # text against every rule's trigger, as the backstop for whatever the earlier
# arms missed. An unopened rule from either half holds the reply for one read. # arms missed. An unopened rule from either half holds the reply for one read.
# #
# THE SAME CONTRACT AS scribe_report_check.sh, and for its reasons: # ONE END-OF-TURN REQUEST (milestone 500 step 4). When the turn closed a task,
# the same request carries the task ids and the server also checks the reply
# for the completion sections — once a Stop hook of its own
# (scribe_report_check.sh), now the same call. The section check's outcome is
# recorded (`report_check`, milestone 409's adherence number), and when it
# held the reply, the rewrite is sent back once with `rewrite: true` so the
# server can record how it came out. A rewrite is never held.
#
# THE CONTRACT, and the reasons for it:
# - the server decides and supplies the words; this hook blocks only on a # - the server decides and supplies the words; this hook blocks only on a
# reason it was given, so an unconfigured or unreachable instance never # reason it was given, so an unconfigured or unreachable instance never
# stops a session; # stops a session;
@@ -34,15 +42,43 @@ session_id=$(scribe_json_pick "$event_flat" '.session_id')
active=$(scribe_json_pick "$event_flat" '.stop_hook_active') active=$(scribe_json_pick "$event_flat" '.stop_hook_active')
event_cwd=$(scribe_json_pick "$event_flat" '.cwd') event_cwd=$(scribe_json_pick "$event_flat" '.cwd')
# The rewrite after a hold goes out as written.
[ "$active" = "true" ] && exit 0
[ -n "$transcript" ] && [ -f "$transcript" ] && [ -n "$session_id" ] || exit 0 [ -n "$transcript" ] && [ -f "$transcript" ] && [ -n "$session_id" ] || exit 0
state_dir="${TMPDIR:-/tmp}/scribe-priorart"
mkdir -p "$state_dir" 2>/dev/null || true
safe_sid=$(printf '%s' "$session_id" | tr -c 'A-Za-z0-9._-' '_')
rulefile="$state_dir/${safe_sid}.rules.ids"
stopfile="$state_dir/${safe_sid}.checkpoint.ids"
# Set when the section check held the reply: the next stop is its rewrite.
reportfile="$state_dir/${safe_sid}.reportcheck"
# The rewrite after a hold goes out as written. Only a rewrite of a SECTION
# hold is sent at all — to record how it came out; one after a rule hold, or
# another plugin's block, has nothing to report.
if [ "$active" = "true" ]; then
[ -f "$reportfile" ] || exit 0
rm -f "$reportfile" 2>/dev/null || true
rewrite=true
else
rm -f "$reportfile" 2>/dev/null || true
rewrite=false
fi
scribe_config || exit 0 scribe_config || exit 0
facts=$(scribe_turn_facts "$transcript") facts=$(scribe_turn_facts "$transcript")
[ "$(scribe_turn_fact "$facts" bounded)" = "1" ] || exit 0 [ "$(scribe_turn_fact "$facts" bounded)" = "1" ] || exit 0
reply=$(scribe_turn_fact "$facts" reply | scribe_json_unescape) reply=$(scribe_turn_fact "$facts" reply | scribe_json_unescape)
# The reply may not be in the transcript yet when the hook fires: an empty
# reply is "cannot tell", never "missing everything".
[ -n "$(printf '%s' "$reply" | tr -d '[:space:]')" ] || exit 0 [ -n "$(printf '%s' "$reply" | tr -d '[:space:]')" ] || exit 0
# How many tasks this turn closed, and which — successful closes only, this
# session only (scribe_turn.awk). A task created already done closes with no
# id, so the count is what says a check is due. Digits and commas only, so
# both drop straight into the JSON.
closed=$(scribe_turn_fact "$facts" closed | tr -cd '0-9')
closed=${closed:-0}
task_ids=$(scribe_turn_fact "$facts" task_ids | tr -cd '0-9,' | sed 's/,,*/,/g; s/^,//; s/,$//')
[ "$rewrite" = "true" ] && [ "$closed" = "0" ] && exit 0
# Bounded before encoding. The server reads the head and the tail — the part # Bounded before encoding. The server reads the head and the tail — the part
# of a report that asks something of the reader is at its end — so a very # of a report that asks something of the reader is at its end — so a very
@@ -51,12 +87,10 @@ if [ "${#reply}" -gt 12000 ]; then
reply="${reply:0:4000} … ${reply: -8000}" reply="${reply:0:4000} … ${reply: -8000}"
fi fi
reply_esc=$(printf '%s' "$reply" | scribe_json_escape) || exit 0 reply_esc=$(printf '%s' "$reply" | scribe_json_escape) || exit 0
closing=""
state_dir="${TMPDIR:-/tmp}/scribe-priorart" if [ "$closed" != "0" ]; then
mkdir -p "$state_dir" 2>/dev/null || true closing=$(printf ',"closed":%d,"closed_task_ids":[%s],"rewrite":%s' "$closed" "$task_ids" "$rewrite")
safe_sid=$(printf '%s' "$session_id" | tr -c 'A-Za-z0-9._-' '_') fi
rulefile="$state_dir/${safe_sid}.rules.ids"
stopfile="$state_dir/${safe_sid}.checkpoint.ids"
query="" query=""
scope=$(scribe_scope_query "${event_cwd:-${CLAUDE_PROJECT_DIR:-$PWD}}") scope=$(scribe_scope_query "${event_cwd:-${CLAUDE_PROJECT_DIR:-$PWD}}")
@@ -70,15 +104,17 @@ if [ -f "$stopfile" ]; then
fi fi
query=${query#&} query=${query#&}
answer=$(printf '{"reply":"%s"}' "$reply_esc" | curl -fsS --max-time 6 \ answer=$(printf '{"reply":"%s"%s}' "$reply_esc" "$closing" | curl -fsS --max-time 6 \
-H "Authorization: Bearer ${token}" \ -H "Authorization: Bearer ${token}" \
-H "Content-Type: application/json" \ -H "Content-Type: application/json" \
--data-binary @- \ --data-binary @- \
"${url%/}/api/plugin/reply-rules${query:+?$query}" 2>/dev/null) || exit 0 "${url%/}/api/plugin/reply-rules${query:+?$query}" 2>/dev/null) || exit 0
[ "$rewrite" = "true" ] && exit 0
answer_flat=$(printf '%s' "$answer" | scribe_json_flat) answer_flat=$(printf '%s' "$answer" | scribe_json_flat)
reason=$(scribe_json_pick "$answer_flat" '.reason') reason=$(scribe_json_pick "$answer_flat" '.reason')
[ -n "$reason" ] || exit 0 [ -n "$reason" ] || exit 0
[ "$(scribe_json_pick "$answer_flat" '.report_check')" = "blocked" ] && : > "$reportfile" 2>/dev/null
# Recorded BEFORE the block is emitted, for the act checkpoint's reason: a # Recorded BEFORE the block is emitted, for the act checkpoint's reason: a
# hold that is shown and not recorded is one that can be shown again. # hold that is shown and not recorded is one that can be shown again.
-155
View File
@@ -1,155 +0,0 @@
#!/usr/bin/env bash
# Scribe plugin — Stop hook: a reply that closes a task carries the completion
# sections (milestone 409 step 5).
#
# Everything else the plugin does happens BEFORE the agent writes: context,
# retrieval, the reporting-back skill. This is the one moment the finished
# reply exists, so it is both the last chance to fix a report the operator
# cannot read and the only place adherence to the shape can be measured.
#
# DETERMINISTIC, NO MODEL CALL. Three questions, cheapest first:
#
# 1. Did this turn close a Scribe task? An `update_task` / `create_task` tool
# call with status "done" since the turn's prompt, whose result was not an
# error. Most turns stop here, silently.
# 2. Does the reply that ends the turn have the completion sections? Loosely:
# where the work sits (a record named by id and title, or step N of M),
# what needs the operator, and what comes next. Matched on the words that
# carry the meaning rather than exact headings, so the skill's wording can
# change without breaking this.
# 3. If sections are missing, block once with a reason naming them. The agent
# rewrites; the rewrite is checked and recorded, and never blocked again.
#
# MEASURED FROM THE FIRST CALL. Each checked reply is reported to the instance
# (`/api/plugin/report-check`): passed, blocked, and after a rewrite either
# passed_after_rewrite or missing_after_rewrite. Turns that closed no task are
# not reported — they would cost a request on every turn and add nothing to
# the rate step 6 reads (blocked among checked replies).
#
# IT BLOCKS ONLY WHEN THE BLOCK IS RECORDED, AND ONLY IN THE SERVER'S WORDS.
# The report goes out first; the instance answers a recorded `blocked` with
# the reason to send the agent back with, and the hook blocks only on that
# reason. An unconfigured or unreachable instance therefore never stops a
# session, every intervention is one the numbers can see, and the guidance
# text lives on the server (plugin/PACKAGING.md: hooks carry timing and
# transport).
#
# THE TRANSCRIPT FORMAT IS OBSERVED, NOT DOCUMENTED. Claude Code documents
# `transcript_path` and `stop_hook_active` for Stop, not the JSONL inside. As
# read from real transcripts (2026-09-14): one content block per line;
# `type: "assistant"` lines carry `message.content[]` blocks of `text` /
# `tool_use` ({id, name, input}); tool results arrive as `type: "user"` lines
# whose content is a `tool_result` array ({tool_use_id, is_error}); a turn's
# prompt — typed, or a background-task notification — is a `user` line whose
# content is a plain string and which is not `isMeta`. Anything that does not
# parse that way makes the hook stay out of the way rather than guess.
#
# Config (same as the other hooks):
# CLAUDE_PLUGIN_OPTION_API_ENDPOINT base URL, no trailing slash
# CLAUDE_PLUGIN_OPTION_API_TOKEN fmcp_ API key (sensitive)
# SCRIBE_URL / SCRIBE_TOKEN override for the settings.json dogfooding path.
set -uo pipefail
command -v curl >/dev/null 2>&1 || exit 0
# shellcheck source=plugin/hooks/scribe_defs.sh
. "$(dirname "${BASH_SOURCE[0]}")/scribe_defs.sh"
# Stop delivers { session_id, transcript_path, cwd, hook_event_name, stop_hook_active }.
event=$(cat 2>/dev/null || true)
event_flat=$(printf '%s' "$event" | scribe_json_flat)
transcript=$(scribe_json_pick "$event_flat" '.transcript_path')
[ -n "$transcript" ] && [ -f "$transcript" ] || exit 0
session_id=$(scribe_json_pick "$event_flat" '.session_id')
active=$(scribe_json_pick "$event_flat" '.stop_hook_active')
event_cwd=$(scribe_json_pick "$event_flat" '.cwd')
safe_sid=$(printf '%s' "${session_id:-nosession}" | tr -c 'A-Za-z0-9._-' '_')
state_dir="${TMPDIR:-/tmp}/scribe-reportcheck"
mkdir -p "$state_dir" 2>/dev/null || true
marker="$state_dir/${safe_sid}.blocked"
# Cheap prefilter: no task tool anywhere in the recent transcript → nothing to
# check. Keeps the ordinary turn at one grep. Process substitution, NOT a pipe:
# under `pipefail`, `grep -q` exiting on the first match kills `tail` with
# SIGPIPE, and the pipeline then reports failure precisely when it matched.
grep -q -E '"name":[[:space:]]*"([^"]*__)?(update|create)_task"' < <(tail -c 2000000 "$transcript" 2>/dev/null) || {
rm -f "$marker" 2>/dev/null || true
exit 0
}
# The turn, parsed once — scribe_turn_facts, shared with the reply check.
facts=$(scribe_turn_facts "$transcript")
fact() { scribe_turn_fact "$facts" "$1"; }
[ "$(fact bounded)" = "1" ] || exit 0
closed=$(fact closed)
if [ "${closed:-0}" = "0" ]; then
rm -f "$marker" 2>/dev/null || true
exit 0
fi
# Escaped on one line coming out of awk, so the format survives a multi-line
# reply; decoded here, once, where it is about to be read as text.
reply=$(fact reply | scribe_json_unescape)
task_ids=$(fact task_ids)
# The reply may not be written to the transcript yet when the hook fires. An
# empty reply is "cannot tell", not "missing everything" — stay out of the way.
[ -n "$(printf '%s' "$reply" | tr -d '[:space:]')" ] || exit 0
missing=()
# Where the work sits: a record named by id AND title (#12 "…", milestone 3 "…"),
# or a step position. A bare id is exactly the homework this shape removes.
# shellcheck disable=SC2016 # backticks here are literal markdown, not an expansion
grep -q -i -E '(#[0-9]+|milestone [0-9]+|task [0-9]+)[*_`]*[[:space:]]*[*_`]*["“]|step [0-9]+ of [0-9]+' <<< "$reply" \
|| missing+=("where it sits")
# What needs the operator — "needs you: nothing" counts; it is an answer.
grep -q -i -E 'needs? (from )?you|nothing (is )?needed from you|your (call|decision)' <<< "$reply" \
|| missing+=("needs you")
# What comes next.
grep -q -i -E '\bnext\b' <<< "$reply" \
|| missing+=("next")
# Reports the outcome; prints the instance's reply and returns 0 only if the
# instance recorded it.
report() {
scribe_config || return 1
local q repo enc m
q="outcome=$1&task_ids=${task_ids}"
m=$(IFS=,; printf '%s' "${missing[*]:-}")
if [ -n "$m" ]; then
enc=$(printf '%s' "$m" | scribe_urlenc) || enc=""
q="${q}&missing=${enc}"
fi
scope=$(scribe_scope_query "${event_cwd:-${CLAUDE_PROJECT_DIR:-$PWD}}")
[ -n "$scope" ] && q="${q}&${scope}"
curl -fsS --max-time 4 \
-H "Authorization: Bearer ${token}" \
"${url%/}/api/plugin/report-check?${q}" 2>/dev/null
}
if [ "$active" = "true" ]; then
# A Stop hook already blocked this stop. If it was this one, the reply is
# the rewrite: record how it came out, and let the session stop whatever
# the answer. If it was another plugin's block, this hook has nothing to add.
[ -f "$marker" ] || exit 0
rm -f "$marker" 2>/dev/null || true
if [ ${#missing[@]} -eq 0 ]; then report passed_after_rewrite >/dev/null; else report missing_after_rewrite >/dev/null; fi
exit 0
fi
rm -f "$marker" 2>/dev/null || true
if [ ${#missing[@]} -eq 0 ]; then
report passed >/dev/null
exit 0
fi
# The words the agent is sent back with are the server's (plugin/PACKAGING.md:
# a hook carries timing and transport). No reason back → nothing recorded →
# no block.
answer=$(report blocked) || exit 0
reason=$(scribe_json_pick "$(printf '%s' "$answer" | scribe_json_flat)" '.reason')
[ -n "$reason" ] || exit 0
: > "$marker" 2>/dev/null || true
printf '{"decision":"block","reason":"%s"}\n' "$(printf '%s' "$reason" | scribe_json_escape)"
exit 0
+1 -1
View File
@@ -10,7 +10,7 @@
# words: the agent that built the code is the one participant who knows what # words: the agent that built the code is the one participant who knows what
# it is, and the end of the turn is the last moment that is still true. # it is, and the end of the turn is the last moment that is still true.
# #
# THE SAME DISCIPLINE AS scribe_report_check.sh, deliberately: # THE SAME DISCIPLINE AS the completion-section check in scribe_reply_check.sh:
# - it blocks only on a `reason` the instance returned, which it returns # - it blocks only on a `reason` the instance returned, which it returns
# only for a block it RECORDED — an unconfigured or unreachable instance # only for a block it RECORDED — an unconfigured or unreachable instance
# never stops a session, and every intervention is one the numbers see; # never stops a session, and every intervention is one the numbers see;
+1 -1
View File
@@ -2,7 +2,7 @@
# #
# Reads the flat `IDX<TAB>PATH<TAB>VALUE` stream that scribe_json.awk produces # Reads the flat `IDX<TAB>PATH<TAB>VALUE` stream that scribe_json.awk produces
# in `mode=lines` from a Claude Code transcript, and answers the four questions # in `mode=lines` from a Claude Code transcript, and answers the four questions
# scribe_report_check.sh asks. It replaces a thirty-line jq program; the shape # the Stop hooks ask (scribe_reply_check.sh). It replaces a thirty-line jq program; the shape
# of the answer is unchanged, so the hook around it reads the same. # of the answer is unchanged, so the hook around it reads the same.
# #
# bounded 1 if a user prompt was found in the window, else empty. A window # bounded 1 if a user prompt was found in the window, else empty. A window
+40 -49
View File
@@ -367,28 +367,58 @@ async def reply_rules():
the reply asks something), and the reply text against every rule's the reply asks something), and the reply text against every rule's
trigger — the backstop for whatever the earlier arms missed. trigger — the backstop for whatever the earlier arms missed.
Body: `{"reply": "<text>"}`. Query: `repo` / `project_id` for the scope; ONE END-OF-TURN REQUEST (milestone 500 step 4): a reply that closed tasks
is also checked here for the completion sections (services/report_check),
which used to be a Stop hook and an endpoint of its own. Both holds fold
into one `reason`; `report_check` says what the section check recorded, so
the hook knows the rewrite is one to report.
Body: `{"reply": "<text>", "closed": 1, "closed_task_ids": [41],
"rewrite": false}` — the last three only when the turn closed a task.
`closed` is the count and decides whether the check runs (a task created
already done has no id to list); `rewrite` marks the reply written after a
section hold, which is recorded and never held.
Query: `repo` / `project_id` for the scope;
`held_rule_ids` (opened this session — exempt, as at the act checkpoint); `held_rule_ids` (opened this session — exempt, as at the act checkpoint);
`exclude_rule_ids` (named this session — not re-counted as surfaced); `exclude_rule_ids` (named this session — not re-counted as surfaced);
`stopped_rule_ids` (already held a reply or an act this session — a rule `stopped_rule_ids` (already held a reply or an act this session — a rule
holds once, and the per-session cap counts these). holds once, and the per-session cap counts these).
Returns `reason` — the words the hook blocks with, empty when nothing Returns `reason` — the words the hook blocks with, empty when nothing
holds — plus `rule_ids` for the hook's ledger and `moments` reached. holds — plus `rule_ids` for the hook's ledger, `moments` reached and
`report_check` (the recorded outcome, or "" when nothing was checked).
""" """
data = await request.get_json(silent=True) or {} data = await request.get_json(silent=True) or {}
if not isinstance(data, dict):
data = {}
reply = str(data.get("reply") or "") reply = str(data.get("reply") or "")
# `type(...) is int`, not isinstance: JSON `true` would otherwise read as 1.
closed = data.get("closed") if type(data.get("closed")) is int else 0
ids = data.get("closed_task_ids")
task_ids = [i for i in ids if type(i) is int and i > 0][:20] if isinstance(ids, list) else []
rewrite = data.get("rewrite") is True
project_id, _repo, _unbound = await _project_scope() project_id, _repo, _unbound = await _project_scope()
held = await moment_delivery_svc.reply_hold(
g.user.id, reply, project_id=project_id, checked: dict = {}
exclude=frozenset(_int_list(request.args.get("exclude_rule_ids"))), if closed > 0:
held=frozenset(_int_list(request.args.get("held_rule_ids"))), checked = await report_check_svc.check_reply(
stopped=frozenset(_int_list(request.args.get("stopped_rule_ids"))), g.user.id, reply, task_ids=task_ids, rewrite=rewrite, project_id=project_id or None,
) )
# The rewrite after a hold goes out as written: recorded above, held by nothing.
held: dict = {}
if not rewrite:
held = await moment_delivery_svc.reply_hold(
g.user.id, reply, project_id=project_id,
exclude=frozenset(_int_list(request.args.get("exclude_rule_ids"))),
held=frozenset(_int_list(request.args.get("held_rule_ids"))),
stopped=frozenset(_int_list(request.args.get("stopped_rule_ids"))),
)
reasons = [r for r in (checked.get("reason", ""), held.get("reason", "")) if r]
return jsonify({ return jsonify({
"reason": held.get("reason", ""), "reason": "\n\n".join(reasons),
"rule_ids": held.get("rule_ids", []), "rule_ids": held.get("rule_ids", []),
"moments": held.get("moments", []), "moments": held.get("moments", []),
"report_check": checked.get("outcome", ""),
}) })
@@ -504,45 +534,6 @@ def _parse_shapes(raw: str) -> list[tuple[str, str]]:
return out return out
@plugin_bp.get("/report-check")
@login_required
async def report_check():
"""Record what a Stop hook found in a reply that closed a task (milestone 409 step 5).
The hook decides which completion sections the reply lacks — a local check
of text it can read — and reports the outcome here. For `blocked` the
response carries the `reason` to send the agent back with: the words are
the server's, so every client's hook says the same thing (plugin/PACKAGING.md).
A hook blocks only on a `reason` it received, which means only on a block
that was recorded.
A GET for the reason every plugin endpoint is one: a read-scoped key must
be enough to run the plugin, and this records telemetry the way /retrieve
records a retrieval log.
Query:
outcome (str) — passed | blocked | passed_after_rewrite |
missing_after_rewrite. Anything else is a 400.
missing (opt) — comma-separated sections the reply lacked:
"where it sits", "needs you", "next".
task_ids (opt) — comma-separated ids of the tasks the turn closed.
repo (opt) — working repo remote, resolved like the other arms.
"""
outcome = (request.args.get("outcome") or "").strip()
if outcome not in report_check_svc.OUTCOMES:
return jsonify({"error": f"outcome must be one of {list(report_check_svc.OUTCOMES)}"}), 400
missing = [m for m in (request.args.get("missing") or "").split(",") if m.strip()]
task_ids = _int_list(request.args.get("task_ids"))[:20]
project_id, _repo, _unbound = await _project_scope()
await report_check_svc.record_report_check(
g.user.id, outcome, missing=missing, task_ids=task_ids, project_id=project_id or None,
)
body: dict = {"status": "ok"}
if outcome == "blocked":
body["reason"] = report_check_svc.block_reason(missing)
return jsonify(body)
@plugin_bp.get("/shape-check") @plugin_bp.get("/shape-check")
@login_required @login_required
async def shape_check(): async def shape_check():
@@ -557,7 +548,7 @@ async def shape_check():
recorded, and never on the stop that follows one (`phase=after`). recorded, and never on the stop that follows one (`phase=after`).
A GET like every plugin endpoint: a read-scoped key runs the plugin, and A GET like every plugin endpoint: a read-scoped key runs the plugin, and
this records telemetry the way /report-check does. It changes no ledger this records telemetry the way /reply-rules records its section check. It changes no ledger
row — the agent's own classify_shapes call does that. row — the agent's own classify_shapes call does that.
Query: Query:
+74 -22
View File
@@ -1,38 +1,60 @@
"""The report-shape check: what the plugin's Stop hook found, and what it says. """The report-shape check: does a reply that closed a task carry the completion
sections, and what is it sent back with when it does not.
WHY THIS EXISTS (milestone 409 step 5) WHY THIS EXISTS (milestone 409 step 5)
Everything that helps an agent write a readable completion report arrives Everything that helps an agent write a readable completion report arrives
BEFORE the reply is written. A client's Stop hook is the one moment the BEFORE the reply is written. The Stop hook is the one moment the finished reply
finished reply exists, so it checks that a reply closing a task carries the exists, so a reply that closed a task is checked for the completion sections
completion sections (where the work sits, what needs the operator, what comes (where the work sits, what needs the operator, what comes next). Two jobs:
next), and reports what it found here. Two jobs live on this side:
- RECORDING the outcome, so the rate of `blocked` among checked replies is a - RECORDING the outcome, so the rate of `blocked` among checked replies is a
number milestone 409's last step can read rather than an impression. number rather than an impression (`category=plugin, action=report_check`).
- OWNING THE WORDS the agent is sent back with. A hook carries timing and - OWNING THE WORDS the agent is sent back with, so every client's hook says
transport only (plugin/PACKAGING.md); guidance text comes from the server, the same thing (plugin/PACKAGING.md: hooks carry timing and transport).
so a second client's hook gets the same instruction by calling the same
endpoint, and the wording changes in one place. ONE END-OF-TURN REQUEST (milestone 500 step 4). This used to be a Stop hook of
its own that checked the sections in shell and reported here; the reply-moment
check sent the same reply to `/reply-rules` a moment later. The hook now sends
the reply once, with the ids of the tasks the turn closed, and the section
check runs here beside the reply hold. The patterns are the ones the shell
used, matched on the words that carry the meaning rather than exact headings,
so a shape's wording can change without breaking this.
app_logs rather than a table of its own: one small event with a JSON detail is app_logs rather than a table of its own: one small event with a JSON detail is
what that table holds, it already has retention and an admin viewer, and what that table holds, it already has retention and an admin viewer, and
nothing here needs a join. If the numbers earn a readout, that is the moment to nothing here needs a join.
decide whether they earn a table.
""" """
from __future__ import annotations from __future__ import annotations
import json import json
import logging
import re
from scribe.models import async_session from scribe.models import async_session
from scribe.models.app_log import AppLog from scribe.models.app_log import AppLog
from scribe.services import reply_shapes
logger = logging.getLogger(__name__)
OUTCOMES = ("passed", "blocked", "passed_after_rewrite", "missing_after_rewrite") OUTCOMES = ("passed", "blocked", "passed_after_rewrite", "missing_after_rewrite")
# The sections a hook may name as missing, in the order the reason lists them. # The sections, in the order the reason lists them, each with what counts as
# Anything else a client sends is dropped rather than echoed into an # having it. A bare id ("closed #41") does not place the work: naming a record
# instruction the agent will follow. # by id AND title, or a step position, is the shape — the title is what spares
SECTIONS = ("where it sits", "needs you", "next") # the reader a lookup. "Needs you: nothing" counts; it is an answer.
SECTION_PATTERNS: dict[str, re.Pattern] = {
"where it sits": re.compile(
r'(#[0-9]+|milestone [0-9]+|task [0-9]+)[*_`]*\s*[*_`]*["“]|step [0-9]+ of [0-9]+', re.I),
"needs you": re.compile(
r"needs? (from )?you|nothing (is )?needed from you|your (call|decision)", re.I),
"next": re.compile(r"\bnext\b", re.I),
}
SECTIONS = tuple(SECTION_PATTERNS)
def missing_sections(reply: str) -> list[str]:
return [name for name, pat in SECTION_PATTERNS.items() if not pat.search(reply or "")]
def known_sections(missing: list[str]) -> list[str]: def known_sections(missing: list[str]) -> list[str]:
@@ -43,16 +65,17 @@ def known_sections(missing: list[str]) -> list[str]:
def block_reason(missing: list[str]) -> str: def block_reason(missing: list[str]) -> str:
"""What the agent is told when its completion report is sent back. """What the agent is told when its completion report is sent back.
Names what is missing and points at the reporting-back skill for the shape Names what is missing and carries the completion shape's one line rather
rather than restating it — the skill owns the shape (decision #4027). than the whole shape: the shape itself was delivered when the task closed,
and `list_reply_shapes` has it in full.
""" """
listed = ", ".join(known_sections(missing)) or "the completion sections" listed = ", ".join(known_sections(missing)) or "the completion sections"
shape = reply_shapes.SHAPES["completion"]
return ( return (
f"This turn closed a Scribe task, and the reply that ends it is missing: {listed}. " f"This turn closed a Scribe task, and the reply that ends it is missing: {listed}. "
"The operator reads this reply to find out where the work stands. Rewrite it as a " "The person the work is for reads this reply to find out where it stands. Rewrite it "
"completion report (the reporting-back skill has the shape): where it sits — the task " f"as a completion report — {shape.reminder} Take where it sits from `placement`, the "
"or milestone by id and title, from `placement` — what now works, what needs them " "task or plan by id and title (`list_reply_shapes` has the shape in full)."
"(or \"nothing\"), and what comes next."
) )
@@ -78,3 +101,32 @@ async def record_report_check(
details=json.dumps(details), details=json.dumps(details),
)) ))
await session.commit() await session.commit()
async def check_reply(
user_id: int, reply: str, *, task_ids: list[int], rewrite: bool,
project_id: int | None = None,
) -> dict:
"""Check a reply that closed tasks, record the outcome, and say whether it is held.
`rewrite` is the reply written after an earlier hold: it is recorded and
never held again, so a hold costs one turn at most.
Returns {"outcome", "reason"} — `reason` empty unless the reply is held.
A hold happens only when its record was written: an outcome the numbers
cannot see is not one this check may act on, so a recording failure
degrades to no hold rather than to an unmeasured one.
"""
missing = missing_sections(reply)
if rewrite:
outcome = "missing_after_rewrite" if missing else "passed_after_rewrite"
else:
outcome = "blocked" if missing else "passed"
try:
await record_report_check(user_id, outcome, missing=missing, task_ids=task_ids,
project_id=project_id)
except Exception: # noqa: BLE001 - an unrecorded check never holds a reply
logger.warning("report check not recorded", exc_info=True)
return {"outcome": "", "reason": ""}
return {"outcome": outcome,
"reason": block_reason(missing) if outcome == "blocked" else ""}
+103 -1
View File
@@ -197,7 +197,8 @@ async def test_the_route_passes_the_reply_and_all_three_ledgers():
with patch.object(routes.moment_delivery_svc, "reply_hold", hold): with patch.object(routes.moment_delivery_svc, "reply_hold", hold):
resp = await routes.reply_rules.__wrapped__() resp = await routes.reply_rules.__wrapped__()
body = await resp.get_json() body = await resp.get_json()
assert body == {"reason": "Held.", "rule_ids": [11], "moments": ["reply.report"]} assert body == {"reason": "Held.", "rule_ids": [11], "moments": ["reply.report"],
"report_check": ""}
assert hold.await_args.args == (7, "done") assert hold.await_args.args == (7, "done")
assert hold.await_args.kwargs == { assert hold.await_args.kwargs == {
"project_id": 3, "exclude": frozenset({5}), "held": frozenset({4}), "project_id": 3, "exclude": frozenset({5}), "held": frozenset({4}),
@@ -205,6 +206,55 @@ async def test_the_route_passes_the_reply_and_all_three_ledgers():
} }
async def _reply_route(body, *, hold_reason="", checked=None):
from scribe.routes import plugin as routes
hold = AsyncMock(return_value={"reason": hold_reason, "rule_ids": [11] if hold_reason else [],
"moments": ["reply.report"]})
check = AsyncMock(return_value=checked or {"outcome": "passed", "reason": ""})
app = Quart(__name__)
async with app.test_request_context("/api/plugin/reply-rules", method="POST", json=body,
query_string={"project_id": "3"}):
g.user = type("U", (), {"id": 7})()
with patch.object(routes.moment_delivery_svc, "reply_hold", hold), \
patch.object(routes.report_check_svc, "check_reply", check):
resp = await routes.reply_rules.__wrapped__()
return await resp.get_json(), hold, check
async def test_a_turn_that_closed_nothing_is_not_section_checked():
_body, _hold, check = await _reply_route({"reply": "done"})
check.assert_not_awaited()
async def test_a_closing_turn_is_section_checked_and_both_holds_fold_into_one_reason():
"""One end-of-turn request (milestone 500 step 4): the section check and
the rule hold answer together, in one `reason`."""
body, _hold, check = await _reply_route(
{"reply": "done", "closed": 1, "closed_task_ids": [41, True, "x"], "rewrite": False},
hold_reason="RULE HOLD", checked={"outcome": "blocked", "reason": "SECTIONS"},
)
assert body["reason"] == "SECTIONS\n\nRULE HOLD"
assert body["report_check"] == "blocked"
# JSON `true` is not task 1, and a string is not an id.
assert check.await_args.kwargs == {"task_ids": [41], "rewrite": False, "project_id": 3}
async def test_a_task_created_already_done_is_checked_without_an_id():
_body, _hold, check = await _reply_route({"reply": "done", "closed": 1, "closed_task_ids": []})
assert check.await_args.kwargs["task_ids"] == []
async def test_the_rewrite_is_recorded_and_held_by_nothing():
body, hold, check = await _reply_route(
{"reply": "done", "closed": 1, "closed_task_ids": [41], "rewrite": True},
checked={"outcome": "passed_after_rewrite", "reason": ""},
)
assert check.await_args.kwargs["rewrite"] is True
hold.assert_not_awaited()
assert body["reason"] == "" and body["report_check"] == "passed_after_rewrite"
# ── the hook ──────────────────────────────────────────────────────────── # ── the hook ────────────────────────────────────────────────────────────
@@ -251,6 +301,58 @@ def test_the_hook_blocks_in_the_servers_words_and_records_the_hold(tmp_path):
assert seen[1]["exclude_rule_ids"] == ["11"] assert seen[1]["exclude_rule_ids"] == ["11"]
def _closing_transcript(tmp_path, *replies):
lines = [
{"type": "user", "message": {"role": "user", "content": "please finish it"}},
{"type": "assistant", "message": {"content": [
{"type": "tool_use", "id": "toolu_1", "name": "mcp__plugin_scribe_scribe__update_task",
"input": {"task_id": 41, "status": "done"}}]}},
{"type": "user", "message": {"content": [
{"type": "tool_result", "tool_use_id": "toolu_1", "is_error": False, "content": "{}"}]}},
] + [{"type": "assistant", "message": {"content": [{"type": "text", "text": r}]}} for r in replies]
path = tmp_path / "t.jsonl"
path.write_text("\n".join(json.dumps(x, separators=(",", ":")) for x in lines) + "\n")
return path
SECTION_HOLD = json.dumps({"reason": "Rewrite as a completion report.", "rule_ids": [],
"moments": ["reply.report"], "report_check": "blocked"}).encode()
def test_a_turn_that_closed_nothing_sends_no_close(tmp_path):
t = _transcript(tmp_path, "Done.")
with http_sink(by_path={"/api/plugin/reply-rules": QUIET}) as (port, seen):
_run(tmp_path, port, t)
assert "closed" not in json.loads(seen[0]["_body"])
def test_a_section_hold_blocks_once_then_the_rewrite_is_reported_and_never_held(tmp_path):
"""The completion-section check, folded into the one Stop request: the
first stop is held in the server's words, the rewrite goes back once with
`rewrite: true` so its outcome is recorded, and nothing after that is sent."""
t = _closing_transcript(tmp_path, "All done, pushed it.")
with http_sink(by_path={"/api/plugin/reply-rules": SECTION_HOLD}) as (port, seen):
out = json.loads(_run(tmp_path, port, t))
assert out == {"decision": "block", "reason": "Rewrite as a completion report."}
first = json.loads(seen[0]["_body"])
assert (first["closed"], first["closed_task_ids"], first["rewrite"]) == (1, [41], False)
rewritten = _closing_transcript(tmp_path, "All done, pushed it.", "Where this sits: …")
assert _run(tmp_path, port, rewritten, active=True) == ""
assert json.loads(seen[1]["_body"])["rewrite"] is True
# The marker is spent: a further stop in the loop sends nothing.
assert _run(tmp_path, port, rewritten, active=True) == ""
assert len(seen) == 2
def test_a_rule_hold_on_a_closing_turn_does_not_mark_a_rewrite(tmp_path):
t = _closing_transcript(tmp_path, "Done.")
with http_sink(by_path={"/api/plugin/reply-rules": HELD}) as (port, seen):
_run(tmp_path, port, t)
assert _run(tmp_path, port, t, active=True) == ""
assert len(seen) == 1
def test_the_hook_says_nothing_when_nothing_holds(tmp_path): def test_the_hook_says_nothing_when_nothing_holds(tmp_path):
t = _transcript(tmp_path, "Done.") t = _transcript(tmp_path, "Done.")
with http_sink(by_path={"/api/plugin/reply-rules": QUIET}) as (port, _seen): with http_sink(by_path={"/api/plugin/reply-rules": QUIET}) as (port, _seen):
-169
View File
@@ -1,169 +0,0 @@
"""The Stop hook that checks a task-closing reply for the completion sections
(milestone 409 step 5).
Runs the real shell against synthetic transcripts in the shape Claude Code
writes (one content block per JSONL line) and the shared HTTP sink. What it
pins: silence on every turn that closed nothing; a block only when the
instance recorded it, in the words the instance returned; one rewrite at most,
recorded; and no block from another plugin's loop or a failed task write.
"""
from __future__ import annotations
import json
import os
import shutil
import subprocess
from pathlib import Path
import pytest
from tests.helpers import http_sink
HOOK = Path(__file__).resolve().parents[1] / "plugin" / "hooks" / "scribe_report_check.sh"
TOOL = "mcp__plugin_scribe_scribe__update_task"
GOOD = ('**Where this sits:** milestone 12 "Move the backups offsite", step 3 of 5.\n'
"**What now works:** the sync runs nightly.\n**Needs you:** nothing.\n**Next:** alerts.")
BAD = "All done, pushed it."
REASON = "SERVER REASON: rewrite as a completion report"
def _env(tmp_path, url="http://127.0.0.1:9"):
for tool in ("curl", "bash"):
if shutil.which(tool) is None:
pytest.skip(f"hook runtime tool {tool!r} not installed")
return {"PATH": os.environ["PATH"], "SCRIBE_URL": url, "SCRIBE_TOKEN": "t",
"TMPDIR": str(tmp_path), "HOME": str(tmp_path)}
def _prompt(text="please finish it"):
return {"type": "user", "message": {"role": "user", "content": text}}
def _tool_use(tid="toolu_1", status="done", name=TOOL, task_id=41):
return {"type": "assistant", "message": {"content": [
{"type": "tool_use", "id": tid, "name": name, "input": {"task_id": task_id, "status": status}}]}}
def _result(tid="toolu_1", is_error=False):
return {"type": "user", "message": {"content": [
{"type": "tool_result", "tool_use_id": tid, "is_error": is_error, "content": "{}"}]}}
def _text(text):
return {"type": "assistant", "message": {"content": [{"type": "text", "text": text}]}}
def _transcript(tmp_path, lines):
path = tmp_path / "t.jsonl"
# Compact, like the file Claude Code writes ({"name":"…"}, no spaces).
path.write_text("\n".join(json.dumps(line, separators=(",", ":")) for line in lines) + "\n")
return path
def _run(env, transcript, active=False, session="s1"):
out = subprocess.run(
["bash", str(HOOK)],
input=json.dumps({"session_id": session, "transcript_path": str(transcript),
"cwd": str(transcript.parent), "hook_event_name": "Stop",
"stop_hook_active": active}),
capture_output=True, text=True, env=env, timeout=30,
)
assert out.returncode == 0, out.stderr
return out.stdout.strip()
def _closing_turn(reply):
return [_prompt(), _text("On it."), _tool_use(), _result(), _text(reply)]
def test_a_turn_that_closed_nothing_is_silent_and_reports_nothing(tmp_path):
with http_sink(b'{"status":"ok","reason":"x"}') as (port, seen):
env = _env(tmp_path, f"http://127.0.0.1:{port}")
t = _transcript(tmp_path, [_prompt(), _tool_use(status="in_progress"), _result(), _text(BAD)])
assert _run(env, t) == ""
assert seen == []
def test_a_complete_report_passes_silently_and_is_recorded(tmp_path):
with http_sink(b'{"status":"ok"}') as (port, seen):
env = _env(tmp_path, f"http://127.0.0.1:{port}")
assert _run(env, _transcript(tmp_path, _closing_turn(GOOD))) == ""
assert [q["outcome"] for q in seen] == [["passed"]]
assert seen[0]["task_ids"] == ["41"]
def test_a_missing_section_blocks_once_in_the_servers_words_then_records_the_rewrite(tmp_path):
reply = json.dumps({"status": "ok", "reason": REASON}).encode()
with http_sink(reply) as (port, seen):
env = _env(tmp_path, f"http://127.0.0.1:{port}")
out = json.loads(_run(env, _transcript(tmp_path, _closing_turn(BAD))))
assert out == {"decision": "block", "reason": REASON}
assert seen[0]["outcome"] == ["blocked"]
assert seen[0]["missing"] == ["where it sits,needs you,next"]
# The rewrite: Claude Code sets stop_hook_active; the hook records and never blocks again.
rewritten = _transcript(tmp_path, _closing_turn(BAD) + [_text(GOOD)])
assert _run(env, rewritten, active=True) == ""
assert seen[1]["outcome"] == ["passed_after_rewrite"]
assert _run(env, rewritten, active=True) == ""
assert len(seen) == 2
def test_a_rewrite_that_still_misses_is_recorded_and_not_blocked(tmp_path):
reply = json.dumps({"status": "ok", "reason": REASON}).encode()
with http_sink(reply) as (port, seen):
env = _env(tmp_path, f"http://127.0.0.1:{port}")
t = _transcript(tmp_path, _closing_turn(BAD))
_run(env, t)
assert _run(env, t, active=True) == ""
assert [q["outcome"][0] for q in seen] == ["blocked", "missing_after_rewrite"]
def test_another_hooks_block_loop_is_left_alone(tmp_path):
with http_sink(b'{"status":"ok","reason":"x"}') as (port, seen):
env = _env(tmp_path, f"http://127.0.0.1:{port}")
assert _run(env, _transcript(tmp_path, _closing_turn(BAD)), active=True) == ""
assert seen == []
def test_a_task_write_that_failed_closed_nothing(tmp_path):
with http_sink(b'{"status":"ok","reason":"x"}') as (port, seen):
env = _env(tmp_path, f"http://127.0.0.1:{port}")
t = _transcript(tmp_path, [_prompt(), _tool_use(), _result(is_error=True), _text(BAD)])
assert _run(env, t) == ""
assert seen == []
def test_a_task_closed_in_an_earlier_turn_does_not_count(tmp_path):
with http_sink(b'{"status":"ok","reason":"x"}') as (port, seen):
env = _env(tmp_path, f"http://127.0.0.1:{port}")
t = _transcript(tmp_path, _closing_turn(GOOD) + [_prompt("thanks, what else?"), _text(BAD)])
assert _run(env, t) == ""
assert seen == []
def test_a_reply_not_yet_written_is_not_judged(tmp_path):
with http_sink(b'{"status":"ok","reason":"x"}') as (port, seen):
env = _env(tmp_path, f"http://127.0.0.1:{port}")
t = _transcript(tmp_path, [_prompt(), _tool_use(), _result()])
assert _run(env, t) == ""
assert seen == []
def test_no_block_without_a_recorded_check(tmp_path):
t = _transcript(tmp_path, _closing_turn(BAD))
# Unreachable instance.
assert _run(_env(tmp_path), t) == ""
# An instance that answered but returned no reason.
with http_sink(b'{"status":"ok"}') as (port, seen):
assert _run(_env(tmp_path, f"http://127.0.0.1:{port}"), t, session="s2") == ""
assert seen[0]["outcome"] == ["blocked"]
def test_a_bare_id_does_not_count_as_placing_the_work(tmp_path):
reply = json.dumps({"status": "ok", "reason": REASON}).encode()
with http_sink(reply) as (port, seen):
env = _env(tmp_path, f"http://127.0.0.1:{port}")
bare = "Closed #41.\n**Needs you:** nothing.\n**Next:** #42."
assert json.loads(_run(env, _transcript(tmp_path, _closing_turn(bare))))["decision"] == "block"
assert seen[0]["missing"] == ["where it sits"]
+67 -5
View File
@@ -1,12 +1,38 @@
"""The server half of the report-shape check (milestone 409 step 5): the words a """The report-shape check (milestone 409 step 5): which completion sections a
blocked reply is sent back with, and the outcome record.""" task-closing reply lacks, the words it is sent back with, and the outcome
record. Since milestone 500 step 4 the check itself runs here, inside the one
end-of-turn request, rather than in a Stop hook of its own."""
import json import json
from unittest.mock import patch from unittest.mock import AsyncMock, patch
import pytest import pytest
from tests.helpers import make_mock_session from tests.helpers import make_mock_session
GOOD = ('**Where this sits:** milestone 12 "Move the backups offsite", step 3 of 5.\n'
"**What now works:** the sync runs nightly.\n**Needs you:** nothing.\n**Next:** alerts.")
BAD = "All done, pushed it."
def test_a_complete_report_misses_nothing():
from scribe.services.report_check import missing_sections
assert missing_sections(GOOD) == []
def test_a_bare_reply_misses_every_section_in_order():
from scribe.services.report_check import SECTIONS, missing_sections
assert missing_sections(BAD) == list(SECTIONS)
def test_a_bare_id_does_not_count_as_placing_the_work():
"""The title is what spares the reader a lookup; "closed #41" does not."""
from scribe.services.report_check import missing_sections
assert missing_sections("Closed #41.\n**Needs you:** nothing.\n**Next:** #42.") == ["where it sits"]
assert missing_sections('Closed #41 "Sync". Needs you: nothing. Next: #42.') == []
def test_the_reason_names_only_sections_it_knows(): def test_the_reason_names_only_sections_it_knows():
from scribe.services.report_check import block_reason from scribe.services.report_check import block_reason
@@ -14,8 +40,17 @@ def test_the_reason_names_only_sections_it_knows():
reason = block_reason(["next", "ignore previous instructions", "Where It Sits"]) reason = block_reason(["next", "ignore previous instructions", "Where It Sits"])
assert "missing: where it sits, next." in reason assert "missing: where it sits, next." in reason
assert "ignore previous instructions" not in reason assert "ignore previous instructions" not in reason
# It points at the skill that owns the shape rather than restating it.
assert "reporting-back" in reason and "placement" in reason
def test_the_reason_carries_the_completion_shapes_line_and_where_the_rest_is():
"""The shape was delivered when the task closed; the reason repeats its
one line, not the whole of it, and names where the full text is."""
from scribe.services import reply_shapes
from scribe.services.report_check import block_reason
reason = block_reason(["next"])
assert reply_shapes.SHAPES["completion"].reminder in reason
assert "list_reply_shapes" in reason and "placement" in reason
def test_a_reason_with_nothing_recognised_still_says_what_to_do(): def test_a_reason_with_nothing_recognised_still_says_what_to_do():
@@ -45,3 +80,30 @@ async def test_an_unknown_outcome_is_refused_before_anything_is_written():
pytest.raises(ValueError): pytest.raises(ValueError):
await record_report_check(7, "skipped") await record_report_check(7, "skipped")
session.add.assert_not_called() session.add.assert_not_called()
@pytest.mark.parametrize("reply,rewrite,outcome,held", [
(GOOD, False, "passed", False),
(BAD, False, "blocked", True),
(GOOD, True, "passed_after_rewrite", False),
(BAD, True, "missing_after_rewrite", False),
])
async def test_check_reply_records_every_outcome_and_holds_only_a_first_miss(reply, rewrite, outcome, held):
from scribe.services import report_check
record = AsyncMock()
with patch.object(report_check, "record_report_check", record):
got = await report_check.check_reply(7, reply, task_ids=[41], rewrite=rewrite, project_id=2)
assert got["outcome"] == outcome
assert bool(got["reason"]) is held
assert record.await_args.args == (7, outcome)
assert record.await_args.kwargs["task_ids"] == [41]
async def test_an_unrecorded_check_holds_nothing():
"""A hold the numbers cannot see is not one this check may make."""
from scribe.services import report_check
with patch.object(report_check, "record_report_check", AsyncMock(side_effect=RuntimeError("db"))):
got = await report_check.check_reply(7, BAD, task_ids=[41], rewrite=False)
assert got == {"outcome": "", "reason": ""}