feat(500): one end-of-turn request - the completion-section check folds into the reply check
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m7s
CI & Build / Python tests (push) Failing after 1m25s
CI & Build / Build & push image (push) Skipped
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m7s
CI & Build / Python tests (push) Failing after 1m25s
CI & Build / Build & push image (push) Skipped
The Stop hook sent the finished reply twice: scribe_report_check.sh checked a task-closing reply for the completion sections in shell and reported to /report-check, and scribe_reply_check.sh sent the same reply to /reply-rules for the rule hold. Now the reply goes once. When the turn closed a task the hook adds the close count and ids, and the server runs the section check (services/report_check, the same three patterns) beside the reply hold, folding both into one reason. The report_check adherence log is still written for every checked reply (milestone 409's number). A section hold marks the session, so its rewrite is sent back once with rewrite:true to record how it came out, and is never held. The block reason carries the completion shape's one line and points at list_reply_shapes rather than at the skill. #5496 (step 4 of milestone 500). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -1,7 +1,7 @@
|
||||
{
|
||||
"name": "scribe",
|
||||
"description": "Scribe for Claude Code: connects the scribe MCP server, adds the hooks that deliver live project state and relevant records at the right moment, ships the shared client-neutral Scribe skills (using-scribe, writing-plans, reporting-back, systematic-debugging, verification, brainstorming, reusing-code, shape-accounting, family-canon), and syncs your saved Scribe Processes as skills (/scribe:sync).",
|
||||
"version": "2026.10.09.1850",
|
||||
"version": "2026.10.09.1938",
|
||||
"author": {
|
||||
"name": "Bryan Van Deusen"
|
||||
},
|
||||
|
||||
+2
-3
@@ -15,7 +15,7 @@ another one means adding files, not moving or rewriting any.
|
||||
|---|---|---|
|
||||
| **The skills** | `plugin/skills/*/SKILL.md` | Agent Skills (the open SKILL.md format). They state every Scribe reflex in full and name no client. `tests/test_guidance_ownership.py` fails if a skill names a particular client, or references anything outside its own folder. Every client package ships this folder verbatim. |
|
||||
| **The MCP server** | `<base URL>/mcp` | HTTP, `Authorization: Bearer <fmcp_ key>`. Its `_INSTRUCTIONS` is a client-neutral orientation — the workflow across tools (≤1,600 chars, #4389); each tool's description carries its contract; in-band responses (`placement`, `report_back`, `systems_hint`, the duplicate gate, the guessed-id refusal) fire in every client. |
|
||||
| **The adapter API** | `<base URL>/api/plugin/*` | Plain `GET` endpoints any client's hooks can call with the same key (read scope is enough): `context` (live session state), `retrieve` (rules, preferences and notes for a message), `prior-art` (records and shape-ledger hints for code being written), `tool-rules` (rules for a command about to run), `report-check` (records a completion-report check and returns the reason for a block), `processes` (stored Processes to expose as skills). |
|
||||
| **The adapter API** | `<base URL>/api/plugin/*` | Plain `GET` endpoints any client's hooks can call with the same key (read scope is enough): `context` (live session state), `retrieve` (rules, preferences and notes for a message), `prior-art` (records and shape-ledger hints for code being written), `tool-rules` (rules for a command about to run), `reply-rules` (a `POST`: the finished reply, held for an unopened rule or — when the turn closed a task — for missing completion sections), `processes` (stored Processes to expose as skills). |
|
||||
| **The API key** | Scribe → Settings → API Keys | One `fmcp_` key per install. Read scope for hooks; write scope for the MCP tools. |
|
||||
|
||||
## Added by each client
|
||||
@@ -40,8 +40,7 @@ another one means adding files, not moving or rewriting any.
|
||||
| `hooks/scribe_after_write.sh` | PostToolUse on shell commands: the same check for code written through the shell. |
|
||||
| `hooks/scribe_tool_rules.sh` | PreToolUse on shell commands: `GET /api/plugin/tool-rules`. |
|
||||
| `hooks/scribe_moment.sh` | PreToolUse on every tool: the rules mounted on the moments the call reaches, `POST /api/plugin/moment`; skips tools `GET /api/plugin/moment-tools` says reach nothing mounted, and Scribe's own (their responses carry `moment_rules`). |
|
||||
| `hooks/scribe_report_check.sh` | Stop: when the turn closed a task, checks the reply for the completion sections and reports to `GET /api/plugin/report-check`; blocks once, with the reason the server returns. |
|
||||
| `hooks/scribe_reply_check.sh` | Stop: sends the finished reply to `POST /api/plugin/reply-rules` — the rules mounted on the reply moments, and the reply against every rule's trigger; holds once per rule, with the reason the server returns, and never holds the rewrite. |
|
||||
| `hooks/scribe_reply_check.sh` | Stop: sends the finished reply to `POST /api/plugin/reply-rules` — the rules mounted on the reply moments, and the reply against every rule's trigger, and — when the turn closed a task — the completion-section check; holds once per rule, with the reason the server returns, and never holds the rewrite. |
|
||||
| `hooks/scribe_shape_check.sh` | Stop: sends the definitions the turn wrote (the write hooks' `<sid>.written.ids` ledger) to `GET /api/plugin/shape-check`; blocks once, with the reason the server returns, so the agent judges what it built. |
|
||||
| `hooks/scribe_sync_processes.sh` + `commands/sync.md` | `GET /api/plugin/processes` → `~/.claude/skills/scribe-proc-*` stubs; `/scribe:sync` on demand. |
|
||||
| `hooks/scribe_defs.sh` | Shared shell helpers: config, dedup ledgers, outage line. |
|
||||
|
||||
+10
-9
@@ -82,15 +82,16 @@ On install you'll be asked for:
|
||||
answer" line (8 s budget here — it runs after the tool, so it gates
|
||||
nothing). The extractor, the prose/data skip list, the local by-name
|
||||
duplicate arm and the outage line are shared in `hooks/scribe_defs.sh`.
|
||||
- `hooks/hooks.json` → Stop hook (`hooks/scribe_report_check.sh`): when the
|
||||
turn closed a Scribe task (`update_task`/`create_task` with status done),
|
||||
checks the reply that ends it for the completion sections — where the work
|
||||
sits, what needs you, what comes next — and reports the outcome to
|
||||
`GET /api/plugin/report-check`. If sections are missing it blocks once with
|
||||
the reason the server returns, and records how the rewrite came out; it
|
||||
never blocks twice, and never blocks when the instance did not record the
|
||||
check (unconfigured or unreachable). Outcomes land in the admin logs under
|
||||
category `plugin`, action `report_check`.
|
||||
- `hooks/hooks.json` → Stop hook (`hooks/scribe_reply_check.sh`): sends the
|
||||
reply that ends the turn to `POST /api/plugin/reply-rules`, which holds it
|
||||
once for an unopened rule mounted on the reply moments or matching the reply.
|
||||
When the turn closed a Scribe task (`update_task`/`create_task` with status
|
||||
done), the same request has the server check the reply for the completion
|
||||
sections — where the work sits, what needs you, what comes next. If they are
|
||||
missing it blocks once with the reason the server returns, and records how
|
||||
the rewrite came out; it never blocks twice, and never blocks when the
|
||||
instance did not record the check (unconfigured or unreachable). Outcomes
|
||||
land in the admin logs under category `plugin`, action `report_check`.
|
||||
- `hooks/hooks.json` → a second Stop hook (`hooks/scribe_shape_check.sh`): the
|
||||
write hooks note every definition a write names (and every new file) in a
|
||||
session ledger; at the end of the turn this sends them to
|
||||
|
||||
@@ -104,10 +104,6 @@
|
||||
"Stop": [
|
||||
{
|
||||
"hooks": [
|
||||
{
|
||||
"type": "command",
|
||||
"command": "bash \"${CLAUDE_PLUGIN_ROOT}/hooks/scribe_report_check.sh\""
|
||||
},
|
||||
{
|
||||
"type": "command",
|
||||
"command": "bash \"${CLAUDE_PLUGIN_ROOT}/hooks/scribe_reply_check.sh\""
|
||||
|
||||
@@ -135,8 +135,8 @@ scribe_json_flat_lines() {
|
||||
# Cheap by construction, because it runs on every prompt: the last 512 KB only,
|
||||
# and grep narrows to assistant records carrying a text block BEFORE anything is
|
||||
# parsed — a transcript is mostly tool results, and parsing those in awk is the
|
||||
# multi-second cost scribe_report_check.sh already measured. Fixed strings with
|
||||
# UNESCAPED quotes can only match at a record's own top level (see that hook).
|
||||
# multi-second cost the completion-report check measured (#4107). Fixed strings
|
||||
# with UNESCAPED quotes can only match at a record's own top level.
|
||||
# A sidechain is a subagent talking, not this session, so it is skipped.
|
||||
scribe_recent_context() {
|
||||
[ -n "${1:-}" ] && [ -f "$1" ] || return 0
|
||||
|
||||
@@ -39,7 +39,7 @@
|
||||
#
|
||||
# mode=lines the input is JSONL — one JSON value per line — and a line that
|
||||
# does not parse is DROPPED, the rest still read. This is the
|
||||
# transcript in scribe_report_check.sh, where the window starts
|
||||
# transcript the Stop hooks read, where the window starts
|
||||
# mid-record by construction: `tail -n 3000` cuts wherever it
|
||||
# cuts, and the first line is routinely half a record. It is
|
||||
# exactly the `map(try fromjson catch empty)` the jq program it
|
||||
|
||||
@@ -9,7 +9,15 @@
|
||||
# text against every rule's trigger, as the backstop for whatever the earlier
|
||||
# arms missed. An unopened rule from either half holds the reply for one read.
|
||||
#
|
||||
# THE SAME CONTRACT AS scribe_report_check.sh, and for its reasons:
|
||||
# ONE END-OF-TURN REQUEST (milestone 500 step 4). When the turn closed a task,
|
||||
# the same request carries the task ids and the server also checks the reply
|
||||
# for the completion sections — once a Stop hook of its own
|
||||
# (scribe_report_check.sh), now the same call. The section check's outcome is
|
||||
# recorded (`report_check`, milestone 409's adherence number), and when it
|
||||
# held the reply, the rewrite is sent back once with `rewrite: true` so the
|
||||
# server can record how it came out. A rewrite is never held.
|
||||
#
|
||||
# THE CONTRACT, and the reasons for it:
|
||||
# - the server decides and supplies the words; this hook blocks only on a
|
||||
# reason it was given, so an unconfigured or unreachable instance never
|
||||
# stops a session;
|
||||
@@ -34,15 +42,43 @@ session_id=$(scribe_json_pick "$event_flat" '.session_id')
|
||||
active=$(scribe_json_pick "$event_flat" '.stop_hook_active')
|
||||
event_cwd=$(scribe_json_pick "$event_flat" '.cwd')
|
||||
|
||||
# The rewrite after a hold goes out as written.
|
||||
[ "$active" = "true" ] && exit 0
|
||||
[ -n "$transcript" ] && [ -f "$transcript" ] && [ -n "$session_id" ] || exit 0
|
||||
|
||||
state_dir="${TMPDIR:-/tmp}/scribe-priorart"
|
||||
mkdir -p "$state_dir" 2>/dev/null || true
|
||||
safe_sid=$(printf '%s' "$session_id" | tr -c 'A-Za-z0-9._-' '_')
|
||||
rulefile="$state_dir/${safe_sid}.rules.ids"
|
||||
stopfile="$state_dir/${safe_sid}.checkpoint.ids"
|
||||
# Set when the section check held the reply: the next stop is its rewrite.
|
||||
reportfile="$state_dir/${safe_sid}.reportcheck"
|
||||
|
||||
# The rewrite after a hold goes out as written. Only a rewrite of a SECTION
|
||||
# hold is sent at all — to record how it came out; one after a rule hold, or
|
||||
# another plugin's block, has nothing to report.
|
||||
if [ "$active" = "true" ]; then
|
||||
[ -f "$reportfile" ] || exit 0
|
||||
rm -f "$reportfile" 2>/dev/null || true
|
||||
rewrite=true
|
||||
else
|
||||
rm -f "$reportfile" 2>/dev/null || true
|
||||
rewrite=false
|
||||
fi
|
||||
scribe_config || exit 0
|
||||
|
||||
facts=$(scribe_turn_facts "$transcript")
|
||||
[ "$(scribe_turn_fact "$facts" bounded)" = "1" ] || exit 0
|
||||
reply=$(scribe_turn_fact "$facts" reply | scribe_json_unescape)
|
||||
# The reply may not be in the transcript yet when the hook fires: an empty
|
||||
# reply is "cannot tell", never "missing everything".
|
||||
[ -n "$(printf '%s' "$reply" | tr -d '[:space:]')" ] || exit 0
|
||||
# How many tasks this turn closed, and which — successful closes only, this
|
||||
# session only (scribe_turn.awk). A task created already done closes with no
|
||||
# id, so the count is what says a check is due. Digits and commas only, so
|
||||
# both drop straight into the JSON.
|
||||
closed=$(scribe_turn_fact "$facts" closed | tr -cd '0-9')
|
||||
closed=${closed:-0}
|
||||
task_ids=$(scribe_turn_fact "$facts" task_ids | tr -cd '0-9,' | sed 's/,,*/,/g; s/^,//; s/,$//')
|
||||
[ "$rewrite" = "true" ] && [ "$closed" = "0" ] && exit 0
|
||||
|
||||
# Bounded before encoding. The server reads the head and the tail — the part
|
||||
# of a report that asks something of the reader is at its end — so a very
|
||||
@@ -51,12 +87,10 @@ if [ "${#reply}" -gt 12000 ]; then
|
||||
reply="${reply:0:4000} … ${reply: -8000}"
|
||||
fi
|
||||
reply_esc=$(printf '%s' "$reply" | scribe_json_escape) || exit 0
|
||||
|
||||
state_dir="${TMPDIR:-/tmp}/scribe-priorart"
|
||||
mkdir -p "$state_dir" 2>/dev/null || true
|
||||
safe_sid=$(printf '%s' "$session_id" | tr -c 'A-Za-z0-9._-' '_')
|
||||
rulefile="$state_dir/${safe_sid}.rules.ids"
|
||||
stopfile="$state_dir/${safe_sid}.checkpoint.ids"
|
||||
closing=""
|
||||
if [ "$closed" != "0" ]; then
|
||||
closing=$(printf ',"closed":%d,"closed_task_ids":[%s],"rewrite":%s' "$closed" "$task_ids" "$rewrite")
|
||||
fi
|
||||
|
||||
query=""
|
||||
scope=$(scribe_scope_query "${event_cwd:-${CLAUDE_PROJECT_DIR:-$PWD}}")
|
||||
@@ -70,15 +104,17 @@ if [ -f "$stopfile" ]; then
|
||||
fi
|
||||
query=${query#&}
|
||||
|
||||
answer=$(printf '{"reply":"%s"}' "$reply_esc" | curl -fsS --max-time 6 \
|
||||
answer=$(printf '{"reply":"%s"%s}' "$reply_esc" "$closing" | curl -fsS --max-time 6 \
|
||||
-H "Authorization: Bearer ${token}" \
|
||||
-H "Content-Type: application/json" \
|
||||
--data-binary @- \
|
||||
"${url%/}/api/plugin/reply-rules${query:+?$query}" 2>/dev/null) || exit 0
|
||||
|
||||
[ "$rewrite" = "true" ] && exit 0
|
||||
answer_flat=$(printf '%s' "$answer" | scribe_json_flat)
|
||||
reason=$(scribe_json_pick "$answer_flat" '.reason')
|
||||
[ -n "$reason" ] || exit 0
|
||||
[ "$(scribe_json_pick "$answer_flat" '.report_check')" = "blocked" ] && : > "$reportfile" 2>/dev/null
|
||||
|
||||
# Recorded BEFORE the block is emitted, for the act checkpoint's reason: a
|
||||
# hold that is shown and not recorded is one that can be shown again.
|
||||
|
||||
@@ -1,155 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Scribe plugin — Stop hook: a reply that closes a task carries the completion
|
||||
# sections (milestone 409 step 5).
|
||||
#
|
||||
# Everything else the plugin does happens BEFORE the agent writes: context,
|
||||
# retrieval, the reporting-back skill. This is the one moment the finished
|
||||
# reply exists, so it is both the last chance to fix a report the operator
|
||||
# cannot read and the only place adherence to the shape can be measured.
|
||||
#
|
||||
# DETERMINISTIC, NO MODEL CALL. Three questions, cheapest first:
|
||||
#
|
||||
# 1. Did this turn close a Scribe task? An `update_task` / `create_task` tool
|
||||
# call with status "done" since the turn's prompt, whose result was not an
|
||||
# error. Most turns stop here, silently.
|
||||
# 2. Does the reply that ends the turn have the completion sections? Loosely:
|
||||
# where the work sits (a record named by id and title, or step N of M),
|
||||
# what needs the operator, and what comes next. Matched on the words that
|
||||
# carry the meaning rather than exact headings, so the skill's wording can
|
||||
# change without breaking this.
|
||||
# 3. If sections are missing, block once with a reason naming them. The agent
|
||||
# rewrites; the rewrite is checked and recorded, and never blocked again.
|
||||
#
|
||||
# MEASURED FROM THE FIRST CALL. Each checked reply is reported to the instance
|
||||
# (`/api/plugin/report-check`): passed, blocked, and after a rewrite either
|
||||
# passed_after_rewrite or missing_after_rewrite. Turns that closed no task are
|
||||
# not reported — they would cost a request on every turn and add nothing to
|
||||
# the rate step 6 reads (blocked among checked replies).
|
||||
#
|
||||
# IT BLOCKS ONLY WHEN THE BLOCK IS RECORDED, AND ONLY IN THE SERVER'S WORDS.
|
||||
# The report goes out first; the instance answers a recorded `blocked` with
|
||||
# the reason to send the agent back with, and the hook blocks only on that
|
||||
# reason. An unconfigured or unreachable instance therefore never stops a
|
||||
# session, every intervention is one the numbers can see, and the guidance
|
||||
# text lives on the server (plugin/PACKAGING.md: hooks carry timing and
|
||||
# transport).
|
||||
#
|
||||
# THE TRANSCRIPT FORMAT IS OBSERVED, NOT DOCUMENTED. Claude Code documents
|
||||
# `transcript_path` and `stop_hook_active` for Stop, not the JSONL inside. As
|
||||
# read from real transcripts (2026-09-14): one content block per line;
|
||||
# `type: "assistant"` lines carry `message.content[]` blocks of `text` /
|
||||
# `tool_use` ({id, name, input}); tool results arrive as `type: "user"` lines
|
||||
# whose content is a `tool_result` array ({tool_use_id, is_error}); a turn's
|
||||
# prompt — typed, or a background-task notification — is a `user` line whose
|
||||
# content is a plain string and which is not `isMeta`. Anything that does not
|
||||
# parse that way makes the hook stay out of the way rather than guess.
|
||||
#
|
||||
# Config (same as the other hooks):
|
||||
# CLAUDE_PLUGIN_OPTION_API_ENDPOINT base URL, no trailing slash
|
||||
# CLAUDE_PLUGIN_OPTION_API_TOKEN fmcp_ API key (sensitive)
|
||||
# SCRIBE_URL / SCRIBE_TOKEN override for the settings.json dogfooding path.
|
||||
set -uo pipefail
|
||||
|
||||
command -v curl >/dev/null 2>&1 || exit 0
|
||||
|
||||
# shellcheck source=plugin/hooks/scribe_defs.sh
|
||||
. "$(dirname "${BASH_SOURCE[0]}")/scribe_defs.sh"
|
||||
|
||||
# Stop delivers { session_id, transcript_path, cwd, hook_event_name, stop_hook_active }.
|
||||
event=$(cat 2>/dev/null || true)
|
||||
event_flat=$(printf '%s' "$event" | scribe_json_flat)
|
||||
transcript=$(scribe_json_pick "$event_flat" '.transcript_path')
|
||||
[ -n "$transcript" ] && [ -f "$transcript" ] || exit 0
|
||||
session_id=$(scribe_json_pick "$event_flat" '.session_id')
|
||||
active=$(scribe_json_pick "$event_flat" '.stop_hook_active')
|
||||
event_cwd=$(scribe_json_pick "$event_flat" '.cwd')
|
||||
|
||||
safe_sid=$(printf '%s' "${session_id:-nosession}" | tr -c 'A-Za-z0-9._-' '_')
|
||||
state_dir="${TMPDIR:-/tmp}/scribe-reportcheck"
|
||||
mkdir -p "$state_dir" 2>/dev/null || true
|
||||
marker="$state_dir/${safe_sid}.blocked"
|
||||
|
||||
# Cheap prefilter: no task tool anywhere in the recent transcript → nothing to
|
||||
# check. Keeps the ordinary turn at one grep. Process substitution, NOT a pipe:
|
||||
# under `pipefail`, `grep -q` exiting on the first match kills `tail` with
|
||||
# SIGPIPE, and the pipeline then reports failure precisely when it matched.
|
||||
grep -q -E '"name":[[:space:]]*"([^"]*__)?(update|create)_task"' < <(tail -c 2000000 "$transcript" 2>/dev/null) || {
|
||||
rm -f "$marker" 2>/dev/null || true
|
||||
exit 0
|
||||
}
|
||||
|
||||
# The turn, parsed once — scribe_turn_facts, shared with the reply check.
|
||||
facts=$(scribe_turn_facts "$transcript")
|
||||
fact() { scribe_turn_fact "$facts" "$1"; }
|
||||
|
||||
[ "$(fact bounded)" = "1" ] || exit 0
|
||||
closed=$(fact closed)
|
||||
if [ "${closed:-0}" = "0" ]; then
|
||||
rm -f "$marker" 2>/dev/null || true
|
||||
exit 0
|
||||
fi
|
||||
# Escaped on one line coming out of awk, so the format survives a multi-line
|
||||
# reply; decoded here, once, where it is about to be read as text.
|
||||
reply=$(fact reply | scribe_json_unescape)
|
||||
task_ids=$(fact task_ids)
|
||||
|
||||
# The reply may not be written to the transcript yet when the hook fires. An
|
||||
# empty reply is "cannot tell", not "missing everything" — stay out of the way.
|
||||
[ -n "$(printf '%s' "$reply" | tr -d '[:space:]')" ] || exit 0
|
||||
|
||||
missing=()
|
||||
# Where the work sits: a record named by id AND title (#12 "…", milestone 3 "…"),
|
||||
# or a step position. A bare id is exactly the homework this shape removes.
|
||||
# shellcheck disable=SC2016 # backticks here are literal markdown, not an expansion
|
||||
grep -q -i -E '(#[0-9]+|milestone [0-9]+|task [0-9]+)[*_`]*[[:space:]]*[*_`]*["“]|step [0-9]+ of [0-9]+' <<< "$reply" \
|
||||
|| missing+=("where it sits")
|
||||
# What needs the operator — "needs you: nothing" counts; it is an answer.
|
||||
grep -q -i -E 'needs? (from )?you|nothing (is )?needed from you|your (call|decision)' <<< "$reply" \
|
||||
|| missing+=("needs you")
|
||||
# What comes next.
|
||||
grep -q -i -E '\bnext\b' <<< "$reply" \
|
||||
|| missing+=("next")
|
||||
|
||||
# Reports the outcome; prints the instance's reply and returns 0 only if the
|
||||
# instance recorded it.
|
||||
report() {
|
||||
scribe_config || return 1
|
||||
local q repo enc m
|
||||
q="outcome=$1&task_ids=${task_ids}"
|
||||
m=$(IFS=,; printf '%s' "${missing[*]:-}")
|
||||
if [ -n "$m" ]; then
|
||||
enc=$(printf '%s' "$m" | scribe_urlenc) || enc=""
|
||||
q="${q}&missing=${enc}"
|
||||
fi
|
||||
scope=$(scribe_scope_query "${event_cwd:-${CLAUDE_PROJECT_DIR:-$PWD}}")
|
||||
[ -n "$scope" ] && q="${q}&${scope}"
|
||||
curl -fsS --max-time 4 \
|
||||
-H "Authorization: Bearer ${token}" \
|
||||
"${url%/}/api/plugin/report-check?${q}" 2>/dev/null
|
||||
}
|
||||
|
||||
if [ "$active" = "true" ]; then
|
||||
# A Stop hook already blocked this stop. If it was this one, the reply is
|
||||
# the rewrite: record how it came out, and let the session stop whatever
|
||||
# the answer. If it was another plugin's block, this hook has nothing to add.
|
||||
[ -f "$marker" ] || exit 0
|
||||
rm -f "$marker" 2>/dev/null || true
|
||||
if [ ${#missing[@]} -eq 0 ]; then report passed_after_rewrite >/dev/null; else report missing_after_rewrite >/dev/null; fi
|
||||
exit 0
|
||||
fi
|
||||
rm -f "$marker" 2>/dev/null || true
|
||||
|
||||
if [ ${#missing[@]} -eq 0 ]; then
|
||||
report passed >/dev/null
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# The words the agent is sent back with are the server's (plugin/PACKAGING.md:
|
||||
# a hook carries timing and transport). No reason back → nothing recorded →
|
||||
# no block.
|
||||
answer=$(report blocked) || exit 0
|
||||
reason=$(scribe_json_pick "$(printf '%s' "$answer" | scribe_json_flat)" '.reason')
|
||||
[ -n "$reason" ] || exit 0
|
||||
: > "$marker" 2>/dev/null || true
|
||||
printf '{"decision":"block","reason":"%s"}\n' "$(printf '%s' "$reason" | scribe_json_escape)"
|
||||
exit 0
|
||||
@@ -10,7 +10,7 @@
|
||||
# words: the agent that built the code is the one participant who knows what
|
||||
# it is, and the end of the turn is the last moment that is still true.
|
||||
#
|
||||
# THE SAME DISCIPLINE AS scribe_report_check.sh, deliberately:
|
||||
# THE SAME DISCIPLINE AS the completion-section check in scribe_reply_check.sh:
|
||||
# - it blocks only on a `reason` the instance returned, which it returns
|
||||
# only for a block it RECORDED — an unconfigured or unreachable instance
|
||||
# never stops a session, and every intervention is one the numbers see;
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
#
|
||||
# Reads the flat `IDX<TAB>PATH<TAB>VALUE` stream that scribe_json.awk produces
|
||||
# in `mode=lines` from a Claude Code transcript, and answers the four questions
|
||||
# scribe_report_check.sh asks. It replaces a thirty-line jq program; the shape
|
||||
# the Stop hooks ask (scribe_reply_check.sh). It replaces a thirty-line jq program; the shape
|
||||
# of the answer is unchanged, so the hook around it reads the same.
|
||||
#
|
||||
# bounded 1 if a user prompt was found in the window, else empty. A window
|
||||
|
||||
+40
-49
@@ -367,28 +367,58 @@ async def reply_rules():
|
||||
the reply asks something), and the reply text against every rule's
|
||||
trigger — the backstop for whatever the earlier arms missed.
|
||||
|
||||
Body: `{"reply": "<text>"}`. Query: `repo` / `project_id` for the scope;
|
||||
ONE END-OF-TURN REQUEST (milestone 500 step 4): a reply that closed tasks
|
||||
is also checked here for the completion sections (services/report_check),
|
||||
which used to be a Stop hook and an endpoint of its own. Both holds fold
|
||||
into one `reason`; `report_check` says what the section check recorded, so
|
||||
the hook knows the rewrite is one to report.
|
||||
|
||||
Body: `{"reply": "<text>", "closed": 1, "closed_task_ids": [41],
|
||||
"rewrite": false}` — the last three only when the turn closed a task.
|
||||
`closed` is the count and decides whether the check runs (a task created
|
||||
already done has no id to list); `rewrite` marks the reply written after a
|
||||
section hold, which is recorded and never held.
|
||||
Query: `repo` / `project_id` for the scope;
|
||||
`held_rule_ids` (opened this session — exempt, as at the act checkpoint);
|
||||
`exclude_rule_ids` (named this session — not re-counted as surfaced);
|
||||
`stopped_rule_ids` (already held a reply or an act this session — a rule
|
||||
holds once, and the per-session cap counts these).
|
||||
|
||||
Returns `reason` — the words the hook blocks with, empty when nothing
|
||||
holds — plus `rule_ids` for the hook's ledger and `moments` reached.
|
||||
holds — plus `rule_ids` for the hook's ledger, `moments` reached and
|
||||
`report_check` (the recorded outcome, or "" when nothing was checked).
|
||||
"""
|
||||
data = await request.get_json(silent=True) or {}
|
||||
if not isinstance(data, dict):
|
||||
data = {}
|
||||
reply = str(data.get("reply") or "")
|
||||
# `type(...) is int`, not isinstance: JSON `true` would otherwise read as 1.
|
||||
closed = data.get("closed") if type(data.get("closed")) is int else 0
|
||||
ids = data.get("closed_task_ids")
|
||||
task_ids = [i for i in ids if type(i) is int and i > 0][:20] if isinstance(ids, list) else []
|
||||
rewrite = data.get("rewrite") is True
|
||||
project_id, _repo, _unbound = await _project_scope()
|
||||
held = await moment_delivery_svc.reply_hold(
|
||||
g.user.id, reply, project_id=project_id,
|
||||
exclude=frozenset(_int_list(request.args.get("exclude_rule_ids"))),
|
||||
held=frozenset(_int_list(request.args.get("held_rule_ids"))),
|
||||
stopped=frozenset(_int_list(request.args.get("stopped_rule_ids"))),
|
||||
)
|
||||
|
||||
checked: dict = {}
|
||||
if closed > 0:
|
||||
checked = await report_check_svc.check_reply(
|
||||
g.user.id, reply, task_ids=task_ids, rewrite=rewrite, project_id=project_id or None,
|
||||
)
|
||||
# The rewrite after a hold goes out as written: recorded above, held by nothing.
|
||||
held: dict = {}
|
||||
if not rewrite:
|
||||
held = await moment_delivery_svc.reply_hold(
|
||||
g.user.id, reply, project_id=project_id,
|
||||
exclude=frozenset(_int_list(request.args.get("exclude_rule_ids"))),
|
||||
held=frozenset(_int_list(request.args.get("held_rule_ids"))),
|
||||
stopped=frozenset(_int_list(request.args.get("stopped_rule_ids"))),
|
||||
)
|
||||
reasons = [r for r in (checked.get("reason", ""), held.get("reason", "")) if r]
|
||||
return jsonify({
|
||||
"reason": held.get("reason", ""),
|
||||
"reason": "\n\n".join(reasons),
|
||||
"rule_ids": held.get("rule_ids", []),
|
||||
"moments": held.get("moments", []),
|
||||
"report_check": checked.get("outcome", ""),
|
||||
})
|
||||
|
||||
|
||||
@@ -504,45 +534,6 @@ def _parse_shapes(raw: str) -> list[tuple[str, str]]:
|
||||
return out
|
||||
|
||||
|
||||
@plugin_bp.get("/report-check")
|
||||
@login_required
|
||||
async def report_check():
|
||||
"""Record what a Stop hook found in a reply that closed a task (milestone 409 step 5).
|
||||
|
||||
The hook decides which completion sections the reply lacks — a local check
|
||||
of text it can read — and reports the outcome here. For `blocked` the
|
||||
response carries the `reason` to send the agent back with: the words are
|
||||
the server's, so every client's hook says the same thing (plugin/PACKAGING.md).
|
||||
A hook blocks only on a `reason` it received, which means only on a block
|
||||
that was recorded.
|
||||
|
||||
A GET for the reason every plugin endpoint is one: a read-scoped key must
|
||||
be enough to run the plugin, and this records telemetry the way /retrieve
|
||||
records a retrieval log.
|
||||
|
||||
Query:
|
||||
outcome (str) — passed | blocked | passed_after_rewrite |
|
||||
missing_after_rewrite. Anything else is a 400.
|
||||
missing (opt) — comma-separated sections the reply lacked:
|
||||
"where it sits", "needs you", "next".
|
||||
task_ids (opt) — comma-separated ids of the tasks the turn closed.
|
||||
repo (opt) — working repo remote, resolved like the other arms.
|
||||
"""
|
||||
outcome = (request.args.get("outcome") or "").strip()
|
||||
if outcome not in report_check_svc.OUTCOMES:
|
||||
return jsonify({"error": f"outcome must be one of {list(report_check_svc.OUTCOMES)}"}), 400
|
||||
missing = [m for m in (request.args.get("missing") or "").split(",") if m.strip()]
|
||||
task_ids = _int_list(request.args.get("task_ids"))[:20]
|
||||
project_id, _repo, _unbound = await _project_scope()
|
||||
await report_check_svc.record_report_check(
|
||||
g.user.id, outcome, missing=missing, task_ids=task_ids, project_id=project_id or None,
|
||||
)
|
||||
body: dict = {"status": "ok"}
|
||||
if outcome == "blocked":
|
||||
body["reason"] = report_check_svc.block_reason(missing)
|
||||
return jsonify(body)
|
||||
|
||||
|
||||
@plugin_bp.get("/shape-check")
|
||||
@login_required
|
||||
async def shape_check():
|
||||
@@ -557,7 +548,7 @@ async def shape_check():
|
||||
recorded, and never on the stop that follows one (`phase=after`).
|
||||
|
||||
A GET like every plugin endpoint: a read-scoped key runs the plugin, and
|
||||
this records telemetry the way /report-check does. It changes no ledger
|
||||
this records telemetry the way /reply-rules records its section check. It changes no ledger
|
||||
row — the agent's own classify_shapes call does that.
|
||||
|
||||
Query:
|
||||
|
||||
@@ -1,38 +1,60 @@
|
||||
"""The report-shape check: what the plugin's Stop hook found, and what it says.
|
||||
"""The report-shape check: does a reply that closed a task carry the completion
|
||||
sections, and what is it sent back with when it does not.
|
||||
|
||||
WHY THIS EXISTS (milestone 409 step 5)
|
||||
|
||||
Everything that helps an agent write a readable completion report arrives
|
||||
BEFORE the reply is written. A client's Stop hook is the one moment the
|
||||
finished reply exists, so it checks that a reply closing a task carries the
|
||||
completion sections (where the work sits, what needs the operator, what comes
|
||||
next), and reports what it found here. Two jobs live on this side:
|
||||
BEFORE the reply is written. The Stop hook is the one moment the finished reply
|
||||
exists, so a reply that closed a task is checked for the completion sections
|
||||
(where the work sits, what needs the operator, what comes next). Two jobs:
|
||||
|
||||
- RECORDING the outcome, so the rate of `blocked` among checked replies is a
|
||||
number milestone 409's last step can read rather than an impression.
|
||||
- OWNING THE WORDS the agent is sent back with. A hook carries timing and
|
||||
transport only (plugin/PACKAGING.md); guidance text comes from the server,
|
||||
so a second client's hook gets the same instruction by calling the same
|
||||
endpoint, and the wording changes in one place.
|
||||
number rather than an impression (`category=plugin, action=report_check`).
|
||||
- OWNING THE WORDS the agent is sent back with, so every client's hook says
|
||||
the same thing (plugin/PACKAGING.md: hooks carry timing and transport).
|
||||
|
||||
ONE END-OF-TURN REQUEST (milestone 500 step 4). This used to be a Stop hook of
|
||||
its own that checked the sections in shell and reported here; the reply-moment
|
||||
check sent the same reply to `/reply-rules` a moment later. The hook now sends
|
||||
the reply once, with the ids of the tasks the turn closed, and the section
|
||||
check runs here beside the reply hold. The patterns are the ones the shell
|
||||
used, matched on the words that carry the meaning rather than exact headings,
|
||||
so a shape's wording can change without breaking this.
|
||||
|
||||
app_logs rather than a table of its own: one small event with a JSON detail is
|
||||
what that table holds, it already has retention and an admin viewer, and
|
||||
nothing here needs a join. If the numbers earn a readout, that is the moment to
|
||||
decide whether they earn a table.
|
||||
nothing here needs a join.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import logging
|
||||
import re
|
||||
|
||||
from scribe.models import async_session
|
||||
from scribe.models.app_log import AppLog
|
||||
from scribe.services import reply_shapes
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
OUTCOMES = ("passed", "blocked", "passed_after_rewrite", "missing_after_rewrite")
|
||||
|
||||
# The sections a hook may name as missing, in the order the reason lists them.
|
||||
# Anything else a client sends is dropped rather than echoed into an
|
||||
# instruction the agent will follow.
|
||||
SECTIONS = ("where it sits", "needs you", "next")
|
||||
# The sections, in the order the reason lists them, each with what counts as
|
||||
# having it. A bare id ("closed #41") does not place the work: naming a record
|
||||
# by id AND title, or a step position, is the shape — the title is what spares
|
||||
# the reader a lookup. "Needs you: nothing" counts; it is an answer.
|
||||
SECTION_PATTERNS: dict[str, re.Pattern] = {
|
||||
"where it sits": re.compile(
|
||||
r'(#[0-9]+|milestone [0-9]+|task [0-9]+)[*_`]*\s*[*_`]*["“]|step [0-9]+ of [0-9]+', re.I),
|
||||
"needs you": re.compile(
|
||||
r"needs? (from )?you|nothing (is )?needed from you|your (call|decision)", re.I),
|
||||
"next": re.compile(r"\bnext\b", re.I),
|
||||
}
|
||||
SECTIONS = tuple(SECTION_PATTERNS)
|
||||
|
||||
|
||||
def missing_sections(reply: str) -> list[str]:
|
||||
return [name for name, pat in SECTION_PATTERNS.items() if not pat.search(reply or "")]
|
||||
|
||||
|
||||
def known_sections(missing: list[str]) -> list[str]:
|
||||
@@ -43,16 +65,17 @@ def known_sections(missing: list[str]) -> list[str]:
|
||||
def block_reason(missing: list[str]) -> str:
|
||||
"""What the agent is told when its completion report is sent back.
|
||||
|
||||
Names what is missing and points at the reporting-back skill for the shape
|
||||
rather than restating it — the skill owns the shape (decision #4027).
|
||||
Names what is missing and carries the completion shape's one line rather
|
||||
than the whole shape: the shape itself was delivered when the task closed,
|
||||
and `list_reply_shapes` has it in full.
|
||||
"""
|
||||
listed = ", ".join(known_sections(missing)) or "the completion sections"
|
||||
shape = reply_shapes.SHAPES["completion"]
|
||||
return (
|
||||
f"This turn closed a Scribe task, and the reply that ends it is missing: {listed}. "
|
||||
"The operator reads this reply to find out where the work stands. Rewrite it as a "
|
||||
"completion report (the reporting-back skill has the shape): where it sits — the task "
|
||||
"or milestone by id and title, from `placement` — what now works, what needs them "
|
||||
"(or \"nothing\"), and what comes next."
|
||||
"The person the work is for reads this reply to find out where it stands. Rewrite it "
|
||||
f"as a completion report — {shape.reminder} Take where it sits from `placement`, the "
|
||||
"task or plan by id and title (`list_reply_shapes` has the shape in full)."
|
||||
)
|
||||
|
||||
|
||||
@@ -78,3 +101,32 @@ async def record_report_check(
|
||||
details=json.dumps(details),
|
||||
))
|
||||
await session.commit()
|
||||
|
||||
|
||||
async def check_reply(
|
||||
user_id: int, reply: str, *, task_ids: list[int], rewrite: bool,
|
||||
project_id: int | None = None,
|
||||
) -> dict:
|
||||
"""Check a reply that closed tasks, record the outcome, and say whether it is held.
|
||||
|
||||
`rewrite` is the reply written after an earlier hold: it is recorded and
|
||||
never held again, so a hold costs one turn at most.
|
||||
|
||||
Returns {"outcome", "reason"} — `reason` empty unless the reply is held.
|
||||
A hold happens only when its record was written: an outcome the numbers
|
||||
cannot see is not one this check may act on, so a recording failure
|
||||
degrades to no hold rather than to an unmeasured one.
|
||||
"""
|
||||
missing = missing_sections(reply)
|
||||
if rewrite:
|
||||
outcome = "missing_after_rewrite" if missing else "passed_after_rewrite"
|
||||
else:
|
||||
outcome = "blocked" if missing else "passed"
|
||||
try:
|
||||
await record_report_check(user_id, outcome, missing=missing, task_ids=task_ids,
|
||||
project_id=project_id)
|
||||
except Exception: # noqa: BLE001 - an unrecorded check never holds a reply
|
||||
logger.warning("report check not recorded", exc_info=True)
|
||||
return {"outcome": "", "reason": ""}
|
||||
return {"outcome": outcome,
|
||||
"reason": block_reason(missing) if outcome == "blocked" else ""}
|
||||
|
||||
+103
-1
@@ -197,7 +197,8 @@ async def test_the_route_passes_the_reply_and_all_three_ledgers():
|
||||
with patch.object(routes.moment_delivery_svc, "reply_hold", hold):
|
||||
resp = await routes.reply_rules.__wrapped__()
|
||||
body = await resp.get_json()
|
||||
assert body == {"reason": "Held.", "rule_ids": [11], "moments": ["reply.report"]}
|
||||
assert body == {"reason": "Held.", "rule_ids": [11], "moments": ["reply.report"],
|
||||
"report_check": ""}
|
||||
assert hold.await_args.args == (7, "done")
|
||||
assert hold.await_args.kwargs == {
|
||||
"project_id": 3, "exclude": frozenset({5}), "held": frozenset({4}),
|
||||
@@ -205,6 +206,55 @@ async def test_the_route_passes_the_reply_and_all_three_ledgers():
|
||||
}
|
||||
|
||||
|
||||
async def _reply_route(body, *, hold_reason="", checked=None):
|
||||
from scribe.routes import plugin as routes
|
||||
|
||||
hold = AsyncMock(return_value={"reason": hold_reason, "rule_ids": [11] if hold_reason else [],
|
||||
"moments": ["reply.report"]})
|
||||
check = AsyncMock(return_value=checked or {"outcome": "passed", "reason": ""})
|
||||
app = Quart(__name__)
|
||||
async with app.test_request_context("/api/plugin/reply-rules", method="POST", json=body,
|
||||
query_string={"project_id": "3"}):
|
||||
g.user = type("U", (), {"id": 7})()
|
||||
with patch.object(routes.moment_delivery_svc, "reply_hold", hold), \
|
||||
patch.object(routes.report_check_svc, "check_reply", check):
|
||||
resp = await routes.reply_rules.__wrapped__()
|
||||
return await resp.get_json(), hold, check
|
||||
|
||||
|
||||
async def test_a_turn_that_closed_nothing_is_not_section_checked():
|
||||
_body, _hold, check = await _reply_route({"reply": "done"})
|
||||
check.assert_not_awaited()
|
||||
|
||||
|
||||
async def test_a_closing_turn_is_section_checked_and_both_holds_fold_into_one_reason():
|
||||
"""One end-of-turn request (milestone 500 step 4): the section check and
|
||||
the rule hold answer together, in one `reason`."""
|
||||
body, _hold, check = await _reply_route(
|
||||
{"reply": "done", "closed": 1, "closed_task_ids": [41, True, "x"], "rewrite": False},
|
||||
hold_reason="RULE HOLD", checked={"outcome": "blocked", "reason": "SECTIONS"},
|
||||
)
|
||||
assert body["reason"] == "SECTIONS\n\nRULE HOLD"
|
||||
assert body["report_check"] == "blocked"
|
||||
# JSON `true` is not task 1, and a string is not an id.
|
||||
assert check.await_args.kwargs == {"task_ids": [41], "rewrite": False, "project_id": 3}
|
||||
|
||||
|
||||
async def test_a_task_created_already_done_is_checked_without_an_id():
|
||||
_body, _hold, check = await _reply_route({"reply": "done", "closed": 1, "closed_task_ids": []})
|
||||
assert check.await_args.kwargs["task_ids"] == []
|
||||
|
||||
|
||||
async def test_the_rewrite_is_recorded_and_held_by_nothing():
|
||||
body, hold, check = await _reply_route(
|
||||
{"reply": "done", "closed": 1, "closed_task_ids": [41], "rewrite": True},
|
||||
checked={"outcome": "passed_after_rewrite", "reason": ""},
|
||||
)
|
||||
assert check.await_args.kwargs["rewrite"] is True
|
||||
hold.assert_not_awaited()
|
||||
assert body["reason"] == "" and body["report_check"] == "passed_after_rewrite"
|
||||
|
||||
|
||||
# ── the hook ────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
@@ -251,6 +301,58 @@ def test_the_hook_blocks_in_the_servers_words_and_records_the_hold(tmp_path):
|
||||
assert seen[1]["exclude_rule_ids"] == ["11"]
|
||||
|
||||
|
||||
def _closing_transcript(tmp_path, *replies):
|
||||
lines = [
|
||||
{"type": "user", "message": {"role": "user", "content": "please finish it"}},
|
||||
{"type": "assistant", "message": {"content": [
|
||||
{"type": "tool_use", "id": "toolu_1", "name": "mcp__plugin_scribe_scribe__update_task",
|
||||
"input": {"task_id": 41, "status": "done"}}]}},
|
||||
{"type": "user", "message": {"content": [
|
||||
{"type": "tool_result", "tool_use_id": "toolu_1", "is_error": False, "content": "{}"}]}},
|
||||
] + [{"type": "assistant", "message": {"content": [{"type": "text", "text": r}]}} for r in replies]
|
||||
path = tmp_path / "t.jsonl"
|
||||
path.write_text("\n".join(json.dumps(x, separators=(",", ":")) for x in lines) + "\n")
|
||||
return path
|
||||
|
||||
|
||||
SECTION_HOLD = json.dumps({"reason": "Rewrite as a completion report.", "rule_ids": [],
|
||||
"moments": ["reply.report"], "report_check": "blocked"}).encode()
|
||||
|
||||
|
||||
def test_a_turn_that_closed_nothing_sends_no_close(tmp_path):
|
||||
t = _transcript(tmp_path, "Done.")
|
||||
with http_sink(by_path={"/api/plugin/reply-rules": QUIET}) as (port, seen):
|
||||
_run(tmp_path, port, t)
|
||||
assert "closed" not in json.loads(seen[0]["_body"])
|
||||
|
||||
|
||||
def test_a_section_hold_blocks_once_then_the_rewrite_is_reported_and_never_held(tmp_path):
|
||||
"""The completion-section check, folded into the one Stop request: the
|
||||
first stop is held in the server's words, the rewrite goes back once with
|
||||
`rewrite: true` so its outcome is recorded, and nothing after that is sent."""
|
||||
t = _closing_transcript(tmp_path, "All done, pushed it.")
|
||||
with http_sink(by_path={"/api/plugin/reply-rules": SECTION_HOLD}) as (port, seen):
|
||||
out = json.loads(_run(tmp_path, port, t))
|
||||
assert out == {"decision": "block", "reason": "Rewrite as a completion report."}
|
||||
first = json.loads(seen[0]["_body"])
|
||||
assert (first["closed"], first["closed_task_ids"], first["rewrite"]) == (1, [41], False)
|
||||
|
||||
rewritten = _closing_transcript(tmp_path, "All done, pushed it.", "Where this sits: …")
|
||||
assert _run(tmp_path, port, rewritten, active=True) == ""
|
||||
assert json.loads(seen[1]["_body"])["rewrite"] is True
|
||||
# The marker is spent: a further stop in the loop sends nothing.
|
||||
assert _run(tmp_path, port, rewritten, active=True) == ""
|
||||
assert len(seen) == 2
|
||||
|
||||
|
||||
def test_a_rule_hold_on_a_closing_turn_does_not_mark_a_rewrite(tmp_path):
|
||||
t = _closing_transcript(tmp_path, "Done.")
|
||||
with http_sink(by_path={"/api/plugin/reply-rules": HELD}) as (port, seen):
|
||||
_run(tmp_path, port, t)
|
||||
assert _run(tmp_path, port, t, active=True) == ""
|
||||
assert len(seen) == 1
|
||||
|
||||
|
||||
def test_the_hook_says_nothing_when_nothing_holds(tmp_path):
|
||||
t = _transcript(tmp_path, "Done.")
|
||||
with http_sink(by_path={"/api/plugin/reply-rules": QUIET}) as (port, _seen):
|
||||
|
||||
@@ -1,169 +0,0 @@
|
||||
"""The Stop hook that checks a task-closing reply for the completion sections
|
||||
(milestone 409 step 5).
|
||||
|
||||
Runs the real shell against synthetic transcripts in the shape Claude Code
|
||||
writes (one content block per JSONL line) and the shared HTTP sink. What it
|
||||
pins: silence on every turn that closed nothing; a block only when the
|
||||
instance recorded it, in the words the instance returned; one rewrite at most,
|
||||
recorded; and no block from another plugin's loop or a failed task write.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import shutil
|
||||
import subprocess
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from tests.helpers import http_sink
|
||||
|
||||
HOOK = Path(__file__).resolve().parents[1] / "plugin" / "hooks" / "scribe_report_check.sh"
|
||||
TOOL = "mcp__plugin_scribe_scribe__update_task"
|
||||
GOOD = ('**Where this sits:** milestone 12 "Move the backups offsite", step 3 of 5.\n'
|
||||
"**What now works:** the sync runs nightly.\n**Needs you:** nothing.\n**Next:** alerts.")
|
||||
BAD = "All done, pushed it."
|
||||
REASON = "SERVER REASON: rewrite as a completion report"
|
||||
|
||||
|
||||
def _env(tmp_path, url="http://127.0.0.1:9"):
|
||||
for tool in ("curl", "bash"):
|
||||
if shutil.which(tool) is None:
|
||||
pytest.skip(f"hook runtime tool {tool!r} not installed")
|
||||
return {"PATH": os.environ["PATH"], "SCRIBE_URL": url, "SCRIBE_TOKEN": "t",
|
||||
"TMPDIR": str(tmp_path), "HOME": str(tmp_path)}
|
||||
|
||||
|
||||
def _prompt(text="please finish it"):
|
||||
return {"type": "user", "message": {"role": "user", "content": text}}
|
||||
|
||||
|
||||
def _tool_use(tid="toolu_1", status="done", name=TOOL, task_id=41):
|
||||
return {"type": "assistant", "message": {"content": [
|
||||
{"type": "tool_use", "id": tid, "name": name, "input": {"task_id": task_id, "status": status}}]}}
|
||||
|
||||
|
||||
def _result(tid="toolu_1", is_error=False):
|
||||
return {"type": "user", "message": {"content": [
|
||||
{"type": "tool_result", "tool_use_id": tid, "is_error": is_error, "content": "{}"}]}}
|
||||
|
||||
|
||||
def _text(text):
|
||||
return {"type": "assistant", "message": {"content": [{"type": "text", "text": text}]}}
|
||||
|
||||
|
||||
def _transcript(tmp_path, lines):
|
||||
path = tmp_path / "t.jsonl"
|
||||
# Compact, like the file Claude Code writes ({"name":"…"}, no spaces).
|
||||
path.write_text("\n".join(json.dumps(line, separators=(",", ":")) for line in lines) + "\n")
|
||||
return path
|
||||
|
||||
|
||||
def _run(env, transcript, active=False, session="s1"):
|
||||
out = subprocess.run(
|
||||
["bash", str(HOOK)],
|
||||
input=json.dumps({"session_id": session, "transcript_path": str(transcript),
|
||||
"cwd": str(transcript.parent), "hook_event_name": "Stop",
|
||||
"stop_hook_active": active}),
|
||||
capture_output=True, text=True, env=env, timeout=30,
|
||||
)
|
||||
assert out.returncode == 0, out.stderr
|
||||
return out.stdout.strip()
|
||||
|
||||
|
||||
def _closing_turn(reply):
|
||||
return [_prompt(), _text("On it."), _tool_use(), _result(), _text(reply)]
|
||||
|
||||
|
||||
def test_a_turn_that_closed_nothing_is_silent_and_reports_nothing(tmp_path):
|
||||
with http_sink(b'{"status":"ok","reason":"x"}') as (port, seen):
|
||||
env = _env(tmp_path, f"http://127.0.0.1:{port}")
|
||||
t = _transcript(tmp_path, [_prompt(), _tool_use(status="in_progress"), _result(), _text(BAD)])
|
||||
assert _run(env, t) == ""
|
||||
assert seen == []
|
||||
|
||||
|
||||
def test_a_complete_report_passes_silently_and_is_recorded(tmp_path):
|
||||
with http_sink(b'{"status":"ok"}') as (port, seen):
|
||||
env = _env(tmp_path, f"http://127.0.0.1:{port}")
|
||||
assert _run(env, _transcript(tmp_path, _closing_turn(GOOD))) == ""
|
||||
assert [q["outcome"] for q in seen] == [["passed"]]
|
||||
assert seen[0]["task_ids"] == ["41"]
|
||||
|
||||
|
||||
def test_a_missing_section_blocks_once_in_the_servers_words_then_records_the_rewrite(tmp_path):
|
||||
reply = json.dumps({"status": "ok", "reason": REASON}).encode()
|
||||
with http_sink(reply) as (port, seen):
|
||||
env = _env(tmp_path, f"http://127.0.0.1:{port}")
|
||||
out = json.loads(_run(env, _transcript(tmp_path, _closing_turn(BAD))))
|
||||
assert out == {"decision": "block", "reason": REASON}
|
||||
assert seen[0]["outcome"] == ["blocked"]
|
||||
assert seen[0]["missing"] == ["where it sits,needs you,next"]
|
||||
|
||||
# The rewrite: Claude Code sets stop_hook_active; the hook records and never blocks again.
|
||||
rewritten = _transcript(tmp_path, _closing_turn(BAD) + [_text(GOOD)])
|
||||
assert _run(env, rewritten, active=True) == ""
|
||||
assert seen[1]["outcome"] == ["passed_after_rewrite"]
|
||||
assert _run(env, rewritten, active=True) == ""
|
||||
assert len(seen) == 2
|
||||
|
||||
|
||||
def test_a_rewrite_that_still_misses_is_recorded_and_not_blocked(tmp_path):
|
||||
reply = json.dumps({"status": "ok", "reason": REASON}).encode()
|
||||
with http_sink(reply) as (port, seen):
|
||||
env = _env(tmp_path, f"http://127.0.0.1:{port}")
|
||||
t = _transcript(tmp_path, _closing_turn(BAD))
|
||||
_run(env, t)
|
||||
assert _run(env, t, active=True) == ""
|
||||
assert [q["outcome"][0] for q in seen] == ["blocked", "missing_after_rewrite"]
|
||||
|
||||
|
||||
def test_another_hooks_block_loop_is_left_alone(tmp_path):
|
||||
with http_sink(b'{"status":"ok","reason":"x"}') as (port, seen):
|
||||
env = _env(tmp_path, f"http://127.0.0.1:{port}")
|
||||
assert _run(env, _transcript(tmp_path, _closing_turn(BAD)), active=True) == ""
|
||||
assert seen == []
|
||||
|
||||
|
||||
def test_a_task_write_that_failed_closed_nothing(tmp_path):
|
||||
with http_sink(b'{"status":"ok","reason":"x"}') as (port, seen):
|
||||
env = _env(tmp_path, f"http://127.0.0.1:{port}")
|
||||
t = _transcript(tmp_path, [_prompt(), _tool_use(), _result(is_error=True), _text(BAD)])
|
||||
assert _run(env, t) == ""
|
||||
assert seen == []
|
||||
|
||||
|
||||
def test_a_task_closed_in_an_earlier_turn_does_not_count(tmp_path):
|
||||
with http_sink(b'{"status":"ok","reason":"x"}') as (port, seen):
|
||||
env = _env(tmp_path, f"http://127.0.0.1:{port}")
|
||||
t = _transcript(tmp_path, _closing_turn(GOOD) + [_prompt("thanks, what else?"), _text(BAD)])
|
||||
assert _run(env, t) == ""
|
||||
assert seen == []
|
||||
|
||||
|
||||
def test_a_reply_not_yet_written_is_not_judged(tmp_path):
|
||||
with http_sink(b'{"status":"ok","reason":"x"}') as (port, seen):
|
||||
env = _env(tmp_path, f"http://127.0.0.1:{port}")
|
||||
t = _transcript(tmp_path, [_prompt(), _tool_use(), _result()])
|
||||
assert _run(env, t) == ""
|
||||
assert seen == []
|
||||
|
||||
|
||||
def test_no_block_without_a_recorded_check(tmp_path):
|
||||
t = _transcript(tmp_path, _closing_turn(BAD))
|
||||
# Unreachable instance.
|
||||
assert _run(_env(tmp_path), t) == ""
|
||||
# An instance that answered but returned no reason.
|
||||
with http_sink(b'{"status":"ok"}') as (port, seen):
|
||||
assert _run(_env(tmp_path, f"http://127.0.0.1:{port}"), t, session="s2") == ""
|
||||
assert seen[0]["outcome"] == ["blocked"]
|
||||
|
||||
|
||||
def test_a_bare_id_does_not_count_as_placing_the_work(tmp_path):
|
||||
reply = json.dumps({"status": "ok", "reason": REASON}).encode()
|
||||
with http_sink(reply) as (port, seen):
|
||||
env = _env(tmp_path, f"http://127.0.0.1:{port}")
|
||||
bare = "Closed #41.\n**Needs you:** nothing.\n**Next:** #42."
|
||||
assert json.loads(_run(env, _transcript(tmp_path, _closing_turn(bare))))["decision"] == "block"
|
||||
assert seen[0]["missing"] == ["where it sits"]
|
||||
@@ -1,12 +1,38 @@
|
||||
"""The server half of the report-shape check (milestone 409 step 5): the words a
|
||||
blocked reply is sent back with, and the outcome record."""
|
||||
"""The report-shape check (milestone 409 step 5): which completion sections a
|
||||
task-closing reply lacks, the words it is sent back with, and the outcome
|
||||
record. Since milestone 500 step 4 the check itself runs here, inside the one
|
||||
end-of-turn request, rather than in a Stop hook of its own."""
|
||||
import json
|
||||
from unittest.mock import patch
|
||||
from unittest.mock import AsyncMock, patch
|
||||
|
||||
import pytest
|
||||
|
||||
from tests.helpers import make_mock_session
|
||||
|
||||
GOOD = ('**Where this sits:** milestone 12 "Move the backups offsite", step 3 of 5.\n'
|
||||
"**What now works:** the sync runs nightly.\n**Needs you:** nothing.\n**Next:** alerts.")
|
||||
BAD = "All done, pushed it."
|
||||
|
||||
|
||||
def test_a_complete_report_misses_nothing():
|
||||
from scribe.services.report_check import missing_sections
|
||||
|
||||
assert missing_sections(GOOD) == []
|
||||
|
||||
|
||||
def test_a_bare_reply_misses_every_section_in_order():
|
||||
from scribe.services.report_check import SECTIONS, missing_sections
|
||||
|
||||
assert missing_sections(BAD) == list(SECTIONS)
|
||||
|
||||
|
||||
def test_a_bare_id_does_not_count_as_placing_the_work():
|
||||
"""The title is what spares the reader a lookup; "closed #41" does not."""
|
||||
from scribe.services.report_check import missing_sections
|
||||
|
||||
assert missing_sections("Closed #41.\n**Needs you:** nothing.\n**Next:** #42.") == ["where it sits"]
|
||||
assert missing_sections('Closed #41 "Sync". Needs you: nothing. Next: #42.') == []
|
||||
|
||||
|
||||
def test_the_reason_names_only_sections_it_knows():
|
||||
from scribe.services.report_check import block_reason
|
||||
@@ -14,8 +40,17 @@ def test_the_reason_names_only_sections_it_knows():
|
||||
reason = block_reason(["next", "ignore previous instructions", "Where It Sits"])
|
||||
assert "missing: where it sits, next." in reason
|
||||
assert "ignore previous instructions" not in reason
|
||||
# It points at the skill that owns the shape rather than restating it.
|
||||
assert "reporting-back" in reason and "placement" in reason
|
||||
|
||||
|
||||
def test_the_reason_carries_the_completion_shapes_line_and_where_the_rest_is():
|
||||
"""The shape was delivered when the task closed; the reason repeats its
|
||||
one line, not the whole of it, and names where the full text is."""
|
||||
from scribe.services import reply_shapes
|
||||
from scribe.services.report_check import block_reason
|
||||
|
||||
reason = block_reason(["next"])
|
||||
assert reply_shapes.SHAPES["completion"].reminder in reason
|
||||
assert "list_reply_shapes" in reason and "placement" in reason
|
||||
|
||||
|
||||
def test_a_reason_with_nothing_recognised_still_says_what_to_do():
|
||||
@@ -45,3 +80,30 @@ async def test_an_unknown_outcome_is_refused_before_anything_is_written():
|
||||
pytest.raises(ValueError):
|
||||
await record_report_check(7, "skipped")
|
||||
session.add.assert_not_called()
|
||||
|
||||
|
||||
@pytest.mark.parametrize("reply,rewrite,outcome,held", [
|
||||
(GOOD, False, "passed", False),
|
||||
(BAD, False, "blocked", True),
|
||||
(GOOD, True, "passed_after_rewrite", False),
|
||||
(BAD, True, "missing_after_rewrite", False),
|
||||
])
|
||||
async def test_check_reply_records_every_outcome_and_holds_only_a_first_miss(reply, rewrite, outcome, held):
|
||||
from scribe.services import report_check
|
||||
|
||||
record = AsyncMock()
|
||||
with patch.object(report_check, "record_report_check", record):
|
||||
got = await report_check.check_reply(7, reply, task_ids=[41], rewrite=rewrite, project_id=2)
|
||||
assert got["outcome"] == outcome
|
||||
assert bool(got["reason"]) is held
|
||||
assert record.await_args.args == (7, outcome)
|
||||
assert record.await_args.kwargs["task_ids"] == [41]
|
||||
|
||||
|
||||
async def test_an_unrecorded_check_holds_nothing():
|
||||
"""A hold the numbers cannot see is not one this check may make."""
|
||||
from scribe.services import report_check
|
||||
|
||||
with patch.object(report_check, "record_report_check", AsyncMock(side_effect=RuntimeError("db"))):
|
||||
got = await report_check.check_reply(7, BAD, task_ids=[41], rewrite=False)
|
||||
assert got == {"outcome": "", "reason": ""}
|
||||
|
||||
Reference in New Issue
Block a user