diff --git a/plugin/.claude-plugin/plugin.json b/plugin/.claude-plugin/plugin.json index aeb17858..ea4e2cb5 100644 --- a/plugin/.claude-plugin/plugin.json +++ b/plugin/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "name": "scribe", "description": "Scribe for Claude Code: connects the scribe MCP server, adds the hooks that deliver live project state and relevant records at the right moment, ships the shared client-neutral Scribe skills (using-scribe, writing-plans, reporting-back, systematic-debugging, verification, brainstorming, reusing-code, shape-accounting, family-canon), and syncs your saved Scribe Processes as skills (/scribe:sync).", - "version": "2026.10.09.1850", + "version": "2026.10.09.1938", "author": { "name": "Bryan Van Deusen" }, diff --git a/plugin/PACKAGING.md b/plugin/PACKAGING.md index 14037f6d..e44c79bc 100644 --- a/plugin/PACKAGING.md +++ b/plugin/PACKAGING.md @@ -15,7 +15,7 @@ another one means adding files, not moving or rewriting any. |---|---|---| | **The skills** | `plugin/skills/*/SKILL.md` | Agent Skills (the open SKILL.md format). They state every Scribe reflex in full and name no client. `tests/test_guidance_ownership.py` fails if a skill names a particular client, or references anything outside its own folder. Every client package ships this folder verbatim. | | **The MCP server** | `/mcp` | HTTP, `Authorization: Bearer `. Its `_INSTRUCTIONS` is a client-neutral orientation — the workflow across tools (≤1,600 chars, #4389); each tool's description carries its contract; in-band responses (`placement`, `report_back`, `systems_hint`, the duplicate gate, the guessed-id refusal) fire in every client. | -| **The adapter API** | `/api/plugin/*` | Plain `GET` endpoints any client's hooks can call with the same key (read scope is enough): `context` (live session state), `retrieve` (rules, preferences and notes for a message), `prior-art` (records and shape-ledger hints for code being written), `tool-rules` (rules for a command about to run), `report-check` (records a completion-report check and returns the reason for a block), `processes` (stored Processes to expose as skills). | +| **The adapter API** | `/api/plugin/*` | Plain `GET` endpoints any client's hooks can call with the same key (read scope is enough): `context` (live session state), `retrieve` (rules, preferences and notes for a message), `prior-art` (records and shape-ledger hints for code being written), `tool-rules` (rules for a command about to run), `reply-rules` (a `POST`: the finished reply, held for an unopened rule or — when the turn closed a task — for missing completion sections), `processes` (stored Processes to expose as skills). | | **The API key** | Scribe → Settings → API Keys | One `fmcp_` key per install. Read scope for hooks; write scope for the MCP tools. | ## Added by each client @@ -40,8 +40,7 @@ another one means adding files, not moving or rewriting any. | `hooks/scribe_after_write.sh` | PostToolUse on shell commands: the same check for code written through the shell. | | `hooks/scribe_tool_rules.sh` | PreToolUse on shell commands: `GET /api/plugin/tool-rules`. | | `hooks/scribe_moment.sh` | PreToolUse on every tool: the rules mounted on the moments the call reaches, `POST /api/plugin/moment`; skips tools `GET /api/plugin/moment-tools` says reach nothing mounted, and Scribe's own (their responses carry `moment_rules`). | -| `hooks/scribe_report_check.sh` | Stop: when the turn closed a task, checks the reply for the completion sections and reports to `GET /api/plugin/report-check`; blocks once, with the reason the server returns. | -| `hooks/scribe_reply_check.sh` | Stop: sends the finished reply to `POST /api/plugin/reply-rules` — the rules mounted on the reply moments, and the reply against every rule's trigger; holds once per rule, with the reason the server returns, and never holds the rewrite. | +| `hooks/scribe_reply_check.sh` | Stop: sends the finished reply to `POST /api/plugin/reply-rules` — the rules mounted on the reply moments, and the reply against every rule's trigger, and — when the turn closed a task — the completion-section check; holds once per rule, with the reason the server returns, and never holds the rewrite. | | `hooks/scribe_shape_check.sh` | Stop: sends the definitions the turn wrote (the write hooks' `.written.ids` ledger) to `GET /api/plugin/shape-check`; blocks once, with the reason the server returns, so the agent judges what it built. | | `hooks/scribe_sync_processes.sh` + `commands/sync.md` | `GET /api/plugin/processes` → `~/.claude/skills/scribe-proc-*` stubs; `/scribe:sync` on demand. | | `hooks/scribe_defs.sh` | Shared shell helpers: config, dedup ledgers, outage line. | diff --git a/plugin/README.md b/plugin/README.md index 859e1518..1c3de12c 100644 --- a/plugin/README.md +++ b/plugin/README.md @@ -82,15 +82,16 @@ On install you'll be asked for: answer" line (8 s budget here — it runs after the tool, so it gates nothing). The extractor, the prose/data skip list, the local by-name duplicate arm and the outage line are shared in `hooks/scribe_defs.sh`. -- `hooks/hooks.json` → Stop hook (`hooks/scribe_report_check.sh`): when the - turn closed a Scribe task (`update_task`/`create_task` with status done), - checks the reply that ends it for the completion sections — where the work - sits, what needs you, what comes next — and reports the outcome to - `GET /api/plugin/report-check`. If sections are missing it blocks once with - the reason the server returns, and records how the rewrite came out; it - never blocks twice, and never blocks when the instance did not record the - check (unconfigured or unreachable). Outcomes land in the admin logs under - category `plugin`, action `report_check`. +- `hooks/hooks.json` → Stop hook (`hooks/scribe_reply_check.sh`): sends the + reply that ends the turn to `POST /api/plugin/reply-rules`, which holds it + once for an unopened rule mounted on the reply moments or matching the reply. + When the turn closed a Scribe task (`update_task`/`create_task` with status + done), the same request has the server check the reply for the completion + sections — where the work sits, what needs you, what comes next. If they are + missing it blocks once with the reason the server returns, and records how + the rewrite came out; it never blocks twice, and never blocks when the + instance did not record the check (unconfigured or unreachable). Outcomes + land in the admin logs under category `plugin`, action `report_check`. - `hooks/hooks.json` → a second Stop hook (`hooks/scribe_shape_check.sh`): the write hooks note every definition a write names (and every new file) in a session ledger; at the end of the turn this sends them to diff --git a/plugin/hooks/hooks.json b/plugin/hooks/hooks.json index 20793ef6..213aaf7c 100644 --- a/plugin/hooks/hooks.json +++ b/plugin/hooks/hooks.json @@ -104,10 +104,6 @@ "Stop": [ { "hooks": [ - { - "type": "command", - "command": "bash \"${CLAUDE_PLUGIN_ROOT}/hooks/scribe_report_check.sh\"" - }, { "type": "command", "command": "bash \"${CLAUDE_PLUGIN_ROOT}/hooks/scribe_reply_check.sh\"" diff --git a/plugin/hooks/scribe_defs.sh b/plugin/hooks/scribe_defs.sh index af43fb5e..61d7c838 100644 --- a/plugin/hooks/scribe_defs.sh +++ b/plugin/hooks/scribe_defs.sh @@ -135,8 +135,8 @@ scribe_json_flat_lines() { # Cheap by construction, because it runs on every prompt: the last 512 KB only, # and grep narrows to assistant records carrying a text block BEFORE anything is # parsed — a transcript is mostly tool results, and parsing those in awk is the -# multi-second cost scribe_report_check.sh already measured. Fixed strings with -# UNESCAPED quotes can only match at a record's own top level (see that hook). +# multi-second cost the completion-report check measured (#4107). Fixed strings +# with UNESCAPED quotes can only match at a record's own top level. # A sidechain is a subagent talking, not this session, so it is skipped. scribe_recent_context() { [ -n "${1:-}" ] && [ -f "$1" ] || return 0 diff --git a/plugin/hooks/scribe_json.awk b/plugin/hooks/scribe_json.awk index 2b61ab8f..6c6c8965 100644 --- a/plugin/hooks/scribe_json.awk +++ b/plugin/hooks/scribe_json.awk @@ -39,7 +39,7 @@ # # mode=lines the input is JSONL — one JSON value per line — and a line that # does not parse is DROPPED, the rest still read. This is the -# transcript in scribe_report_check.sh, where the window starts +# transcript the Stop hooks read, where the window starts # mid-record by construction: `tail -n 3000` cuts wherever it # cuts, and the first line is routinely half a record. It is # exactly the `map(try fromjson catch empty)` the jq program it diff --git a/plugin/hooks/scribe_reply_check.sh b/plugin/hooks/scribe_reply_check.sh index 38206434..a68dcd80 100644 --- a/plugin/hooks/scribe_reply_check.sh +++ b/plugin/hooks/scribe_reply_check.sh @@ -9,7 +9,15 @@ # text against every rule's trigger, as the backstop for whatever the earlier # arms missed. An unopened rule from either half holds the reply for one read. # -# THE SAME CONTRACT AS scribe_report_check.sh, and for its reasons: +# ONE END-OF-TURN REQUEST (milestone 500 step 4). When the turn closed a task, +# the same request carries the task ids and the server also checks the reply +# for the completion sections — once a Stop hook of its own +# (scribe_report_check.sh), now the same call. The section check's outcome is +# recorded (`report_check`, milestone 409's adherence number), and when it +# held the reply, the rewrite is sent back once with `rewrite: true` so the +# server can record how it came out. A rewrite is never held. +# +# THE CONTRACT, and the reasons for it: # - the server decides and supplies the words; this hook blocks only on a # reason it was given, so an unconfigured or unreachable instance never # stops a session; @@ -34,15 +42,43 @@ session_id=$(scribe_json_pick "$event_flat" '.session_id') active=$(scribe_json_pick "$event_flat" '.stop_hook_active') event_cwd=$(scribe_json_pick "$event_flat" '.cwd') -# The rewrite after a hold goes out as written. -[ "$active" = "true" ] && exit 0 [ -n "$transcript" ] && [ -f "$transcript" ] && [ -n "$session_id" ] || exit 0 + +state_dir="${TMPDIR:-/tmp}/scribe-priorart" +mkdir -p "$state_dir" 2>/dev/null || true +safe_sid=$(printf '%s' "$session_id" | tr -c 'A-Za-z0-9._-' '_') +rulefile="$state_dir/${safe_sid}.rules.ids" +stopfile="$state_dir/${safe_sid}.checkpoint.ids" +# Set when the section check held the reply: the next stop is its rewrite. +reportfile="$state_dir/${safe_sid}.reportcheck" + +# The rewrite after a hold goes out as written. Only a rewrite of a SECTION +# hold is sent at all — to record how it came out; one after a rule hold, or +# another plugin's block, has nothing to report. +if [ "$active" = "true" ]; then + [ -f "$reportfile" ] || exit 0 + rm -f "$reportfile" 2>/dev/null || true + rewrite=true +else + rm -f "$reportfile" 2>/dev/null || true + rewrite=false +fi scribe_config || exit 0 facts=$(scribe_turn_facts "$transcript") [ "$(scribe_turn_fact "$facts" bounded)" = "1" ] || exit 0 reply=$(scribe_turn_fact "$facts" reply | scribe_json_unescape) +# The reply may not be in the transcript yet when the hook fires: an empty +# reply is "cannot tell", never "missing everything". [ -n "$(printf '%s' "$reply" | tr -d '[:space:]')" ] || exit 0 +# How many tasks this turn closed, and which — successful closes only, this +# session only (scribe_turn.awk). A task created already done closes with no +# id, so the count is what says a check is due. Digits and commas only, so +# both drop straight into the JSON. +closed=$(scribe_turn_fact "$facts" closed | tr -cd '0-9') +closed=${closed:-0} +task_ids=$(scribe_turn_fact "$facts" task_ids | tr -cd '0-9,' | sed 's/,,*/,/g; s/^,//; s/,$//') +[ "$rewrite" = "true" ] && [ "$closed" = "0" ] && exit 0 # Bounded before encoding. The server reads the head and the tail — the part # of a report that asks something of the reader is at its end — so a very @@ -51,12 +87,10 @@ if [ "${#reply}" -gt 12000 ]; then reply="${reply:0:4000} … ${reply: -8000}" fi reply_esc=$(printf '%s' "$reply" | scribe_json_escape) || exit 0 - -state_dir="${TMPDIR:-/tmp}/scribe-priorart" -mkdir -p "$state_dir" 2>/dev/null || true -safe_sid=$(printf '%s' "$session_id" | tr -c 'A-Za-z0-9._-' '_') -rulefile="$state_dir/${safe_sid}.rules.ids" -stopfile="$state_dir/${safe_sid}.checkpoint.ids" +closing="" +if [ "$closed" != "0" ]; then + closing=$(printf ',"closed":%d,"closed_task_ids":[%s],"rewrite":%s' "$closed" "$task_ids" "$rewrite") +fi query="" scope=$(scribe_scope_query "${event_cwd:-${CLAUDE_PROJECT_DIR:-$PWD}}") @@ -70,15 +104,17 @@ if [ -f "$stopfile" ]; then fi query=${query#&} -answer=$(printf '{"reply":"%s"}' "$reply_esc" | curl -fsS --max-time 6 \ +answer=$(printf '{"reply":"%s"%s}' "$reply_esc" "$closing" | curl -fsS --max-time 6 \ -H "Authorization: Bearer ${token}" \ -H "Content-Type: application/json" \ --data-binary @- \ "${url%/}/api/plugin/reply-rules${query:+?$query}" 2>/dev/null) || exit 0 +[ "$rewrite" = "true" ] && exit 0 answer_flat=$(printf '%s' "$answer" | scribe_json_flat) reason=$(scribe_json_pick "$answer_flat" '.reason') [ -n "$reason" ] || exit 0 +[ "$(scribe_json_pick "$answer_flat" '.report_check')" = "blocked" ] && : > "$reportfile" 2>/dev/null # Recorded BEFORE the block is emitted, for the act checkpoint's reason: a # hold that is shown and not recorded is one that can be shown again. diff --git a/plugin/hooks/scribe_report_check.sh b/plugin/hooks/scribe_report_check.sh deleted file mode 100644 index b8742851..00000000 --- a/plugin/hooks/scribe_report_check.sh +++ /dev/null @@ -1,155 +0,0 @@ -#!/usr/bin/env bash -# Scribe plugin — Stop hook: a reply that closes a task carries the completion -# sections (milestone 409 step 5). -# -# Everything else the plugin does happens BEFORE the agent writes: context, -# retrieval, the reporting-back skill. This is the one moment the finished -# reply exists, so it is both the last chance to fix a report the operator -# cannot read and the only place adherence to the shape can be measured. -# -# DETERMINISTIC, NO MODEL CALL. Three questions, cheapest first: -# -# 1. Did this turn close a Scribe task? An `update_task` / `create_task` tool -# call with status "done" since the turn's prompt, whose result was not an -# error. Most turns stop here, silently. -# 2. Does the reply that ends the turn have the completion sections? Loosely: -# where the work sits (a record named by id and title, or step N of M), -# what needs the operator, and what comes next. Matched on the words that -# carry the meaning rather than exact headings, so the skill's wording can -# change without breaking this. -# 3. If sections are missing, block once with a reason naming them. The agent -# rewrites; the rewrite is checked and recorded, and never blocked again. -# -# MEASURED FROM THE FIRST CALL. Each checked reply is reported to the instance -# (`/api/plugin/report-check`): passed, blocked, and after a rewrite either -# passed_after_rewrite or missing_after_rewrite. Turns that closed no task are -# not reported — they would cost a request on every turn and add nothing to -# the rate step 6 reads (blocked among checked replies). -# -# IT BLOCKS ONLY WHEN THE BLOCK IS RECORDED, AND ONLY IN THE SERVER'S WORDS. -# The report goes out first; the instance answers a recorded `blocked` with -# the reason to send the agent back with, and the hook blocks only on that -# reason. An unconfigured or unreachable instance therefore never stops a -# session, every intervention is one the numbers can see, and the guidance -# text lives on the server (plugin/PACKAGING.md: hooks carry timing and -# transport). -# -# THE TRANSCRIPT FORMAT IS OBSERVED, NOT DOCUMENTED. Claude Code documents -# `transcript_path` and `stop_hook_active` for Stop, not the JSONL inside. As -# read from real transcripts (2026-09-14): one content block per line; -# `type: "assistant"` lines carry `message.content[]` blocks of `text` / -# `tool_use` ({id, name, input}); tool results arrive as `type: "user"` lines -# whose content is a `tool_result` array ({tool_use_id, is_error}); a turn's -# prompt — typed, or a background-task notification — is a `user` line whose -# content is a plain string and which is not `isMeta`. Anything that does not -# parse that way makes the hook stay out of the way rather than guess. -# -# Config (same as the other hooks): -# CLAUDE_PLUGIN_OPTION_API_ENDPOINT base URL, no trailing slash -# CLAUDE_PLUGIN_OPTION_API_TOKEN fmcp_ API key (sensitive) -# SCRIBE_URL / SCRIBE_TOKEN override for the settings.json dogfooding path. -set -uo pipefail - -command -v curl >/dev/null 2>&1 || exit 0 - -# shellcheck source=plugin/hooks/scribe_defs.sh -. "$(dirname "${BASH_SOURCE[0]}")/scribe_defs.sh" - -# Stop delivers { session_id, transcript_path, cwd, hook_event_name, stop_hook_active }. -event=$(cat 2>/dev/null || true) -event_flat=$(printf '%s' "$event" | scribe_json_flat) -transcript=$(scribe_json_pick "$event_flat" '.transcript_path') -[ -n "$transcript" ] && [ -f "$transcript" ] || exit 0 -session_id=$(scribe_json_pick "$event_flat" '.session_id') -active=$(scribe_json_pick "$event_flat" '.stop_hook_active') -event_cwd=$(scribe_json_pick "$event_flat" '.cwd') - -safe_sid=$(printf '%s' "${session_id:-nosession}" | tr -c 'A-Za-z0-9._-' '_') -state_dir="${TMPDIR:-/tmp}/scribe-reportcheck" -mkdir -p "$state_dir" 2>/dev/null || true -marker="$state_dir/${safe_sid}.blocked" - -# Cheap prefilter: no task tool anywhere in the recent transcript → nothing to -# check. Keeps the ordinary turn at one grep. Process substitution, NOT a pipe: -# under `pipefail`, `grep -q` exiting on the first match kills `tail` with -# SIGPIPE, and the pipeline then reports failure precisely when it matched. -grep -q -E '"name":[[:space:]]*"([^"]*__)?(update|create)_task"' < <(tail -c 2000000 "$transcript" 2>/dev/null) || { - rm -f "$marker" 2>/dev/null || true - exit 0 -} - -# The turn, parsed once — scribe_turn_facts, shared with the reply check. -facts=$(scribe_turn_facts "$transcript") -fact() { scribe_turn_fact "$facts" "$1"; } - -[ "$(fact bounded)" = "1" ] || exit 0 -closed=$(fact closed) -if [ "${closed:-0}" = "0" ]; then - rm -f "$marker" 2>/dev/null || true - exit 0 -fi -# Escaped on one line coming out of awk, so the format survives a multi-line -# reply; decoded here, once, where it is about to be read as text. -reply=$(fact reply | scribe_json_unescape) -task_ids=$(fact task_ids) - -# The reply may not be written to the transcript yet when the hook fires. An -# empty reply is "cannot tell", not "missing everything" — stay out of the way. -[ -n "$(printf '%s' "$reply" | tr -d '[:space:]')" ] || exit 0 - -missing=() -# Where the work sits: a record named by id AND title (#12 "…", milestone 3 "…"), -# or a step position. A bare id is exactly the homework this shape removes. -# shellcheck disable=SC2016 # backticks here are literal markdown, not an expansion -grep -q -i -E '(#[0-9]+|milestone [0-9]+|task [0-9]+)[*_`]*[[:space:]]*[*_`]*["“]|step [0-9]+ of [0-9]+' <<< "$reply" \ - || missing+=("where it sits") -# What needs the operator — "needs you: nothing" counts; it is an answer. -grep -q -i -E 'needs? (from )?you|nothing (is )?needed from you|your (call|decision)' <<< "$reply" \ - || missing+=("needs you") -# What comes next. -grep -q -i -E '\bnext\b' <<< "$reply" \ - || missing+=("next") - -# Reports the outcome; prints the instance's reply and returns 0 only if the -# instance recorded it. -report() { - scribe_config || return 1 - local q repo enc m - q="outcome=$1&task_ids=${task_ids}" - m=$(IFS=,; printf '%s' "${missing[*]:-}") - if [ -n "$m" ]; then - enc=$(printf '%s' "$m" | scribe_urlenc) || enc="" - q="${q}&missing=${enc}" - fi - scope=$(scribe_scope_query "${event_cwd:-${CLAUDE_PROJECT_DIR:-$PWD}}") - [ -n "$scope" ] && q="${q}&${scope}" - curl -fsS --max-time 4 \ - -H "Authorization: Bearer ${token}" \ - "${url%/}/api/plugin/report-check?${q}" 2>/dev/null -} - -if [ "$active" = "true" ]; then - # A Stop hook already blocked this stop. If it was this one, the reply is - # the rewrite: record how it came out, and let the session stop whatever - # the answer. If it was another plugin's block, this hook has nothing to add. - [ -f "$marker" ] || exit 0 - rm -f "$marker" 2>/dev/null || true - if [ ${#missing[@]} -eq 0 ]; then report passed_after_rewrite >/dev/null; else report missing_after_rewrite >/dev/null; fi - exit 0 -fi -rm -f "$marker" 2>/dev/null || true - -if [ ${#missing[@]} -eq 0 ]; then - report passed >/dev/null - exit 0 -fi - -# The words the agent is sent back with are the server's (plugin/PACKAGING.md: -# a hook carries timing and transport). No reason back → nothing recorded → -# no block. -answer=$(report blocked) || exit 0 -reason=$(scribe_json_pick "$(printf '%s' "$answer" | scribe_json_flat)" '.reason') -[ -n "$reason" ] || exit 0 -: > "$marker" 2>/dev/null || true -printf '{"decision":"block","reason":"%s"}\n' "$(printf '%s' "$reason" | scribe_json_escape)" -exit 0 diff --git a/plugin/hooks/scribe_shape_check.sh b/plugin/hooks/scribe_shape_check.sh index a63097a9..c745c6c8 100644 --- a/plugin/hooks/scribe_shape_check.sh +++ b/plugin/hooks/scribe_shape_check.sh @@ -10,7 +10,7 @@ # words: the agent that built the code is the one participant who knows what # it is, and the end of the turn is the last moment that is still true. # -# THE SAME DISCIPLINE AS scribe_report_check.sh, deliberately: +# THE SAME DISCIPLINE AS the completion-section check in scribe_reply_check.sh: # - it blocks only on a `reason` the instance returned, which it returns # only for a block it RECORDED — an unconfigured or unreachable instance # never stops a session, and every intervention is one the numbers see; diff --git a/plugin/hooks/scribe_turn.awk b/plugin/hooks/scribe_turn.awk index 44f07a9d..a43a901d 100644 --- a/plugin/hooks/scribe_turn.awk +++ b/plugin/hooks/scribe_turn.awk @@ -2,7 +2,7 @@ # # Reads the flat `IDXPATHVALUE` stream that scribe_json.awk produces # in `mode=lines` from a Claude Code transcript, and answers the four questions -# scribe_report_check.sh asks. It replaces a thirty-line jq program; the shape +# the Stop hooks ask (scribe_reply_check.sh). It replaces a thirty-line jq program; the shape # of the answer is unchanged, so the hook around it reads the same. # # bounded 1 if a user prompt was found in the window, else empty. A window diff --git a/src/scribe/routes/plugin.py b/src/scribe/routes/plugin.py index 5777f097..6ba1b68f 100644 --- a/src/scribe/routes/plugin.py +++ b/src/scribe/routes/plugin.py @@ -367,28 +367,58 @@ async def reply_rules(): the reply asks something), and the reply text against every rule's trigger — the backstop for whatever the earlier arms missed. - Body: `{"reply": ""}`. Query: `repo` / `project_id` for the scope; + ONE END-OF-TURN REQUEST (milestone 500 step 4): a reply that closed tasks + is also checked here for the completion sections (services/report_check), + which used to be a Stop hook and an endpoint of its own. Both holds fold + into one `reason`; `report_check` says what the section check recorded, so + the hook knows the rewrite is one to report. + + Body: `{"reply": "", "closed": 1, "closed_task_ids": [41], + "rewrite": false}` — the last three only when the turn closed a task. + `closed` is the count and decides whether the check runs (a task created + already done has no id to list); `rewrite` marks the reply written after a + section hold, which is recorded and never held. + Query: `repo` / `project_id` for the scope; `held_rule_ids` (opened this session — exempt, as at the act checkpoint); `exclude_rule_ids` (named this session — not re-counted as surfaced); `stopped_rule_ids` (already held a reply or an act this session — a rule holds once, and the per-session cap counts these). Returns `reason` — the words the hook blocks with, empty when nothing - holds — plus `rule_ids` for the hook's ledger and `moments` reached. + holds — plus `rule_ids` for the hook's ledger, `moments` reached and + `report_check` (the recorded outcome, or "" when nothing was checked). """ data = await request.get_json(silent=True) or {} + if not isinstance(data, dict): + data = {} reply = str(data.get("reply") or "") + # `type(...) is int`, not isinstance: JSON `true` would otherwise read as 1. + closed = data.get("closed") if type(data.get("closed")) is int else 0 + ids = data.get("closed_task_ids") + task_ids = [i for i in ids if type(i) is int and i > 0][:20] if isinstance(ids, list) else [] + rewrite = data.get("rewrite") is True project_id, _repo, _unbound = await _project_scope() - held = await moment_delivery_svc.reply_hold( - g.user.id, reply, project_id=project_id, - exclude=frozenset(_int_list(request.args.get("exclude_rule_ids"))), - held=frozenset(_int_list(request.args.get("held_rule_ids"))), - stopped=frozenset(_int_list(request.args.get("stopped_rule_ids"))), - ) + + checked: dict = {} + if closed > 0: + checked = await report_check_svc.check_reply( + g.user.id, reply, task_ids=task_ids, rewrite=rewrite, project_id=project_id or None, + ) + # The rewrite after a hold goes out as written: recorded above, held by nothing. + held: dict = {} + if not rewrite: + held = await moment_delivery_svc.reply_hold( + g.user.id, reply, project_id=project_id, + exclude=frozenset(_int_list(request.args.get("exclude_rule_ids"))), + held=frozenset(_int_list(request.args.get("held_rule_ids"))), + stopped=frozenset(_int_list(request.args.get("stopped_rule_ids"))), + ) + reasons = [r for r in (checked.get("reason", ""), held.get("reason", "")) if r] return jsonify({ - "reason": held.get("reason", ""), + "reason": "\n\n".join(reasons), "rule_ids": held.get("rule_ids", []), "moments": held.get("moments", []), + "report_check": checked.get("outcome", ""), }) @@ -504,45 +534,6 @@ def _parse_shapes(raw: str) -> list[tuple[str, str]]: return out -@plugin_bp.get("/report-check") -@login_required -async def report_check(): - """Record what a Stop hook found in a reply that closed a task (milestone 409 step 5). - - The hook decides which completion sections the reply lacks — a local check - of text it can read — and reports the outcome here. For `blocked` the - response carries the `reason` to send the agent back with: the words are - the server's, so every client's hook says the same thing (plugin/PACKAGING.md). - A hook blocks only on a `reason` it received, which means only on a block - that was recorded. - - A GET for the reason every plugin endpoint is one: a read-scoped key must - be enough to run the plugin, and this records telemetry the way /retrieve - records a retrieval log. - - Query: - outcome (str) — passed | blocked | passed_after_rewrite | - missing_after_rewrite. Anything else is a 400. - missing (opt) — comma-separated sections the reply lacked: - "where it sits", "needs you", "next". - task_ids (opt) — comma-separated ids of the tasks the turn closed. - repo (opt) — working repo remote, resolved like the other arms. - """ - outcome = (request.args.get("outcome") or "").strip() - if outcome not in report_check_svc.OUTCOMES: - return jsonify({"error": f"outcome must be one of {list(report_check_svc.OUTCOMES)}"}), 400 - missing = [m for m in (request.args.get("missing") or "").split(",") if m.strip()] - task_ids = _int_list(request.args.get("task_ids"))[:20] - project_id, _repo, _unbound = await _project_scope() - await report_check_svc.record_report_check( - g.user.id, outcome, missing=missing, task_ids=task_ids, project_id=project_id or None, - ) - body: dict = {"status": "ok"} - if outcome == "blocked": - body["reason"] = report_check_svc.block_reason(missing) - return jsonify(body) - - @plugin_bp.get("/shape-check") @login_required async def shape_check(): @@ -557,7 +548,7 @@ async def shape_check(): recorded, and never on the stop that follows one (`phase=after`). A GET like every plugin endpoint: a read-scoped key runs the plugin, and - this records telemetry the way /report-check does. It changes no ledger + this records telemetry the way /reply-rules records its section check. It changes no ledger row — the agent's own classify_shapes call does that. Query: diff --git a/src/scribe/services/report_check.py b/src/scribe/services/report_check.py index ba881232..2c1dbed1 100644 --- a/src/scribe/services/report_check.py +++ b/src/scribe/services/report_check.py @@ -1,38 +1,60 @@ -"""The report-shape check: what the plugin's Stop hook found, and what it says. +"""The report-shape check: does a reply that closed a task carry the completion +sections, and what is it sent back with when it does not. WHY THIS EXISTS (milestone 409 step 5) Everything that helps an agent write a readable completion report arrives -BEFORE the reply is written. A client's Stop hook is the one moment the -finished reply exists, so it checks that a reply closing a task carries the -completion sections (where the work sits, what needs the operator, what comes -next), and reports what it found here. Two jobs live on this side: +BEFORE the reply is written. The Stop hook is the one moment the finished reply +exists, so a reply that closed a task is checked for the completion sections +(where the work sits, what needs the operator, what comes next). Two jobs: - RECORDING the outcome, so the rate of `blocked` among checked replies is a - number milestone 409's last step can read rather than an impression. - - OWNING THE WORDS the agent is sent back with. A hook carries timing and - transport only (plugin/PACKAGING.md); guidance text comes from the server, - so a second client's hook gets the same instruction by calling the same - endpoint, and the wording changes in one place. + number rather than an impression (`category=plugin, action=report_check`). + - OWNING THE WORDS the agent is sent back with, so every client's hook says + the same thing (plugin/PACKAGING.md: hooks carry timing and transport). + +ONE END-OF-TURN REQUEST (milestone 500 step 4). This used to be a Stop hook of +its own that checked the sections in shell and reported here; the reply-moment +check sent the same reply to `/reply-rules` a moment later. The hook now sends +the reply once, with the ids of the tasks the turn closed, and the section +check runs here beside the reply hold. The patterns are the ones the shell +used, matched on the words that carry the meaning rather than exact headings, +so a shape's wording can change without breaking this. app_logs rather than a table of its own: one small event with a JSON detail is what that table holds, it already has retention and an admin viewer, and -nothing here needs a join. If the numbers earn a readout, that is the moment to -decide whether they earn a table. +nothing here needs a join. """ from __future__ import annotations import json +import logging +import re from scribe.models import async_session from scribe.models.app_log import AppLog +from scribe.services import reply_shapes + +logger = logging.getLogger(__name__) OUTCOMES = ("passed", "blocked", "passed_after_rewrite", "missing_after_rewrite") -# The sections a hook may name as missing, in the order the reason lists them. -# Anything else a client sends is dropped rather than echoed into an -# instruction the agent will follow. -SECTIONS = ("where it sits", "needs you", "next") +# The sections, in the order the reason lists them, each with what counts as +# having it. A bare id ("closed #41") does not place the work: naming a record +# by id AND title, or a step position, is the shape — the title is what spares +# the reader a lookup. "Needs you: nothing" counts; it is an answer. +SECTION_PATTERNS: dict[str, re.Pattern] = { + "where it sits": re.compile( + r'(#[0-9]+|milestone [0-9]+|task [0-9]+)[*_`]*\s*[*_`]*["“]|step [0-9]+ of [0-9]+', re.I), + "needs you": re.compile( + r"needs? (from )?you|nothing (is )?needed from you|your (call|decision)", re.I), + "next": re.compile(r"\bnext\b", re.I), +} +SECTIONS = tuple(SECTION_PATTERNS) + + +def missing_sections(reply: str) -> list[str]: + return [name for name, pat in SECTION_PATTERNS.items() if not pat.search(reply or "")] def known_sections(missing: list[str]) -> list[str]: @@ -43,16 +65,17 @@ def known_sections(missing: list[str]) -> list[str]: def block_reason(missing: list[str]) -> str: """What the agent is told when its completion report is sent back. - Names what is missing and points at the reporting-back skill for the shape - rather than restating it — the skill owns the shape (decision #4027). + Names what is missing and carries the completion shape's one line rather + than the whole shape: the shape itself was delivered when the task closed, + and `list_reply_shapes` has it in full. """ listed = ", ".join(known_sections(missing)) or "the completion sections" + shape = reply_shapes.SHAPES["completion"] return ( f"This turn closed a Scribe task, and the reply that ends it is missing: {listed}. " - "The operator reads this reply to find out where the work stands. Rewrite it as a " - "completion report (the reporting-back skill has the shape): where it sits — the task " - "or milestone by id and title, from `placement` — what now works, what needs them " - "(or \"nothing\"), and what comes next." + "The person the work is for reads this reply to find out where it stands. Rewrite it " + f"as a completion report — {shape.reminder} Take where it sits from `placement`, the " + "task or plan by id and title (`list_reply_shapes` has the shape in full)." ) @@ -78,3 +101,32 @@ async def record_report_check( details=json.dumps(details), )) await session.commit() + + +async def check_reply( + user_id: int, reply: str, *, task_ids: list[int], rewrite: bool, + project_id: int | None = None, +) -> dict: + """Check a reply that closed tasks, record the outcome, and say whether it is held. + + `rewrite` is the reply written after an earlier hold: it is recorded and + never held again, so a hold costs one turn at most. + + Returns {"outcome", "reason"} — `reason` empty unless the reply is held. + A hold happens only when its record was written: an outcome the numbers + cannot see is not one this check may act on, so a recording failure + degrades to no hold rather than to an unmeasured one. + """ + missing = missing_sections(reply) + if rewrite: + outcome = "missing_after_rewrite" if missing else "passed_after_rewrite" + else: + outcome = "blocked" if missing else "passed" + try: + await record_report_check(user_id, outcome, missing=missing, task_ids=task_ids, + project_id=project_id) + except Exception: # noqa: BLE001 - an unrecorded check never holds a reply + logger.warning("report check not recorded", exc_info=True) + return {"outcome": "", "reason": ""} + return {"outcome": outcome, + "reason": block_reason(missing) if outcome == "blocked" else ""} diff --git a/tests/test_reply_moment.py b/tests/test_reply_moment.py index 9593f9ca..008ce820 100644 --- a/tests/test_reply_moment.py +++ b/tests/test_reply_moment.py @@ -197,7 +197,8 @@ async def test_the_route_passes_the_reply_and_all_three_ledgers(): with patch.object(routes.moment_delivery_svc, "reply_hold", hold): resp = await routes.reply_rules.__wrapped__() body = await resp.get_json() - assert body == {"reason": "Held.", "rule_ids": [11], "moments": ["reply.report"]} + assert body == {"reason": "Held.", "rule_ids": [11], "moments": ["reply.report"], + "report_check": ""} assert hold.await_args.args == (7, "done") assert hold.await_args.kwargs == { "project_id": 3, "exclude": frozenset({5}), "held": frozenset({4}), @@ -205,6 +206,55 @@ async def test_the_route_passes_the_reply_and_all_three_ledgers(): } +async def _reply_route(body, *, hold_reason="", checked=None): + from scribe.routes import plugin as routes + + hold = AsyncMock(return_value={"reason": hold_reason, "rule_ids": [11] if hold_reason else [], + "moments": ["reply.report"]}) + check = AsyncMock(return_value=checked or {"outcome": "passed", "reason": ""}) + app = Quart(__name__) + async with app.test_request_context("/api/plugin/reply-rules", method="POST", json=body, + query_string={"project_id": "3"}): + g.user = type("U", (), {"id": 7})() + with patch.object(routes.moment_delivery_svc, "reply_hold", hold), \ + patch.object(routes.report_check_svc, "check_reply", check): + resp = await routes.reply_rules.__wrapped__() + return await resp.get_json(), hold, check + + +async def test_a_turn_that_closed_nothing_is_not_section_checked(): + _body, _hold, check = await _reply_route({"reply": "done"}) + check.assert_not_awaited() + + +async def test_a_closing_turn_is_section_checked_and_both_holds_fold_into_one_reason(): + """One end-of-turn request (milestone 500 step 4): the section check and + the rule hold answer together, in one `reason`.""" + body, _hold, check = await _reply_route( + {"reply": "done", "closed": 1, "closed_task_ids": [41, True, "x"], "rewrite": False}, + hold_reason="RULE HOLD", checked={"outcome": "blocked", "reason": "SECTIONS"}, + ) + assert body["reason"] == "SECTIONS\n\nRULE HOLD" + assert body["report_check"] == "blocked" + # JSON `true` is not task 1, and a string is not an id. + assert check.await_args.kwargs == {"task_ids": [41], "rewrite": False, "project_id": 3} + + +async def test_a_task_created_already_done_is_checked_without_an_id(): + _body, _hold, check = await _reply_route({"reply": "done", "closed": 1, "closed_task_ids": []}) + assert check.await_args.kwargs["task_ids"] == [] + + +async def test_the_rewrite_is_recorded_and_held_by_nothing(): + body, hold, check = await _reply_route( + {"reply": "done", "closed": 1, "closed_task_ids": [41], "rewrite": True}, + checked={"outcome": "passed_after_rewrite", "reason": ""}, + ) + assert check.await_args.kwargs["rewrite"] is True + hold.assert_not_awaited() + assert body["reason"] == "" and body["report_check"] == "passed_after_rewrite" + + # ── the hook ──────────────────────────────────────────────────────────── @@ -251,6 +301,58 @@ def test_the_hook_blocks_in_the_servers_words_and_records_the_hold(tmp_path): assert seen[1]["exclude_rule_ids"] == ["11"] +def _closing_transcript(tmp_path, *replies): + lines = [ + {"type": "user", "message": {"role": "user", "content": "please finish it"}}, + {"type": "assistant", "message": {"content": [ + {"type": "tool_use", "id": "toolu_1", "name": "mcp__plugin_scribe_scribe__update_task", + "input": {"task_id": 41, "status": "done"}}]}}, + {"type": "user", "message": {"content": [ + {"type": "tool_result", "tool_use_id": "toolu_1", "is_error": False, "content": "{}"}]}}, + ] + [{"type": "assistant", "message": {"content": [{"type": "text", "text": r}]}} for r in replies] + path = tmp_path / "t.jsonl" + path.write_text("\n".join(json.dumps(x, separators=(",", ":")) for x in lines) + "\n") + return path + + +SECTION_HOLD = json.dumps({"reason": "Rewrite as a completion report.", "rule_ids": [], + "moments": ["reply.report"], "report_check": "blocked"}).encode() + + +def test_a_turn_that_closed_nothing_sends_no_close(tmp_path): + t = _transcript(tmp_path, "Done.") + with http_sink(by_path={"/api/plugin/reply-rules": QUIET}) as (port, seen): + _run(tmp_path, port, t) + assert "closed" not in json.loads(seen[0]["_body"]) + + +def test_a_section_hold_blocks_once_then_the_rewrite_is_reported_and_never_held(tmp_path): + """The completion-section check, folded into the one Stop request: the + first stop is held in the server's words, the rewrite goes back once with + `rewrite: true` so its outcome is recorded, and nothing after that is sent.""" + t = _closing_transcript(tmp_path, "All done, pushed it.") + with http_sink(by_path={"/api/plugin/reply-rules": SECTION_HOLD}) as (port, seen): + out = json.loads(_run(tmp_path, port, t)) + assert out == {"decision": "block", "reason": "Rewrite as a completion report."} + first = json.loads(seen[0]["_body"]) + assert (first["closed"], first["closed_task_ids"], first["rewrite"]) == (1, [41], False) + + rewritten = _closing_transcript(tmp_path, "All done, pushed it.", "Where this sits: …") + assert _run(tmp_path, port, rewritten, active=True) == "" + assert json.loads(seen[1]["_body"])["rewrite"] is True + # The marker is spent: a further stop in the loop sends nothing. + assert _run(tmp_path, port, rewritten, active=True) == "" + assert len(seen) == 2 + + +def test_a_rule_hold_on_a_closing_turn_does_not_mark_a_rewrite(tmp_path): + t = _closing_transcript(tmp_path, "Done.") + with http_sink(by_path={"/api/plugin/reply-rules": HELD}) as (port, seen): + _run(tmp_path, port, t) + assert _run(tmp_path, port, t, active=True) == "" + assert len(seen) == 1 + + def test_the_hook_says_nothing_when_nothing_holds(tmp_path): t = _transcript(tmp_path, "Done.") with http_sink(by_path={"/api/plugin/reply-rules": QUIET}) as (port, _seen): diff --git a/tests/test_report_check_hook.py b/tests/test_report_check_hook.py deleted file mode 100644 index 293ac98f..00000000 --- a/tests/test_report_check_hook.py +++ /dev/null @@ -1,169 +0,0 @@ -"""The Stop hook that checks a task-closing reply for the completion sections -(milestone 409 step 5). - -Runs the real shell against synthetic transcripts in the shape Claude Code -writes (one content block per JSONL line) and the shared HTTP sink. What it -pins: silence on every turn that closed nothing; a block only when the -instance recorded it, in the words the instance returned; one rewrite at most, -recorded; and no block from another plugin's loop or a failed task write. -""" -from __future__ import annotations - -import json -import os -import shutil -import subprocess -from pathlib import Path - -import pytest - -from tests.helpers import http_sink - -HOOK = Path(__file__).resolve().parents[1] / "plugin" / "hooks" / "scribe_report_check.sh" -TOOL = "mcp__plugin_scribe_scribe__update_task" -GOOD = ('**Where this sits:** milestone 12 "Move the backups offsite", step 3 of 5.\n' - "**What now works:** the sync runs nightly.\n**Needs you:** nothing.\n**Next:** alerts.") -BAD = "All done, pushed it." -REASON = "SERVER REASON: rewrite as a completion report" - - -def _env(tmp_path, url="http://127.0.0.1:9"): - for tool in ("curl", "bash"): - if shutil.which(tool) is None: - pytest.skip(f"hook runtime tool {tool!r} not installed") - return {"PATH": os.environ["PATH"], "SCRIBE_URL": url, "SCRIBE_TOKEN": "t", - "TMPDIR": str(tmp_path), "HOME": str(tmp_path)} - - -def _prompt(text="please finish it"): - return {"type": "user", "message": {"role": "user", "content": text}} - - -def _tool_use(tid="toolu_1", status="done", name=TOOL, task_id=41): - return {"type": "assistant", "message": {"content": [ - {"type": "tool_use", "id": tid, "name": name, "input": {"task_id": task_id, "status": status}}]}} - - -def _result(tid="toolu_1", is_error=False): - return {"type": "user", "message": {"content": [ - {"type": "tool_result", "tool_use_id": tid, "is_error": is_error, "content": "{}"}]}} - - -def _text(text): - return {"type": "assistant", "message": {"content": [{"type": "text", "text": text}]}} - - -def _transcript(tmp_path, lines): - path = tmp_path / "t.jsonl" - # Compact, like the file Claude Code writes ({"name":"…"}, no spaces). - path.write_text("\n".join(json.dumps(line, separators=(",", ":")) for line in lines) + "\n") - return path - - -def _run(env, transcript, active=False, session="s1"): - out = subprocess.run( - ["bash", str(HOOK)], - input=json.dumps({"session_id": session, "transcript_path": str(transcript), - "cwd": str(transcript.parent), "hook_event_name": "Stop", - "stop_hook_active": active}), - capture_output=True, text=True, env=env, timeout=30, - ) - assert out.returncode == 0, out.stderr - return out.stdout.strip() - - -def _closing_turn(reply): - return [_prompt(), _text("On it."), _tool_use(), _result(), _text(reply)] - - -def test_a_turn_that_closed_nothing_is_silent_and_reports_nothing(tmp_path): - with http_sink(b'{"status":"ok","reason":"x"}') as (port, seen): - env = _env(tmp_path, f"http://127.0.0.1:{port}") - t = _transcript(tmp_path, [_prompt(), _tool_use(status="in_progress"), _result(), _text(BAD)]) - assert _run(env, t) == "" - assert seen == [] - - -def test_a_complete_report_passes_silently_and_is_recorded(tmp_path): - with http_sink(b'{"status":"ok"}') as (port, seen): - env = _env(tmp_path, f"http://127.0.0.1:{port}") - assert _run(env, _transcript(tmp_path, _closing_turn(GOOD))) == "" - assert [q["outcome"] for q in seen] == [["passed"]] - assert seen[0]["task_ids"] == ["41"] - - -def test_a_missing_section_blocks_once_in_the_servers_words_then_records_the_rewrite(tmp_path): - reply = json.dumps({"status": "ok", "reason": REASON}).encode() - with http_sink(reply) as (port, seen): - env = _env(tmp_path, f"http://127.0.0.1:{port}") - out = json.loads(_run(env, _transcript(tmp_path, _closing_turn(BAD)))) - assert out == {"decision": "block", "reason": REASON} - assert seen[0]["outcome"] == ["blocked"] - assert seen[0]["missing"] == ["where it sits,needs you,next"] - - # The rewrite: Claude Code sets stop_hook_active; the hook records and never blocks again. - rewritten = _transcript(tmp_path, _closing_turn(BAD) + [_text(GOOD)]) - assert _run(env, rewritten, active=True) == "" - assert seen[1]["outcome"] == ["passed_after_rewrite"] - assert _run(env, rewritten, active=True) == "" - assert len(seen) == 2 - - -def test_a_rewrite_that_still_misses_is_recorded_and_not_blocked(tmp_path): - reply = json.dumps({"status": "ok", "reason": REASON}).encode() - with http_sink(reply) as (port, seen): - env = _env(tmp_path, f"http://127.0.0.1:{port}") - t = _transcript(tmp_path, _closing_turn(BAD)) - _run(env, t) - assert _run(env, t, active=True) == "" - assert [q["outcome"][0] for q in seen] == ["blocked", "missing_after_rewrite"] - - -def test_another_hooks_block_loop_is_left_alone(tmp_path): - with http_sink(b'{"status":"ok","reason":"x"}') as (port, seen): - env = _env(tmp_path, f"http://127.0.0.1:{port}") - assert _run(env, _transcript(tmp_path, _closing_turn(BAD)), active=True) == "" - assert seen == [] - - -def test_a_task_write_that_failed_closed_nothing(tmp_path): - with http_sink(b'{"status":"ok","reason":"x"}') as (port, seen): - env = _env(tmp_path, f"http://127.0.0.1:{port}") - t = _transcript(tmp_path, [_prompt(), _tool_use(), _result(is_error=True), _text(BAD)]) - assert _run(env, t) == "" - assert seen == [] - - -def test_a_task_closed_in_an_earlier_turn_does_not_count(tmp_path): - with http_sink(b'{"status":"ok","reason":"x"}') as (port, seen): - env = _env(tmp_path, f"http://127.0.0.1:{port}") - t = _transcript(tmp_path, _closing_turn(GOOD) + [_prompt("thanks, what else?"), _text(BAD)]) - assert _run(env, t) == "" - assert seen == [] - - -def test_a_reply_not_yet_written_is_not_judged(tmp_path): - with http_sink(b'{"status":"ok","reason":"x"}') as (port, seen): - env = _env(tmp_path, f"http://127.0.0.1:{port}") - t = _transcript(tmp_path, [_prompt(), _tool_use(), _result()]) - assert _run(env, t) == "" - assert seen == [] - - -def test_no_block_without_a_recorded_check(tmp_path): - t = _transcript(tmp_path, _closing_turn(BAD)) - # Unreachable instance. - assert _run(_env(tmp_path), t) == "" - # An instance that answered but returned no reason. - with http_sink(b'{"status":"ok"}') as (port, seen): - assert _run(_env(tmp_path, f"http://127.0.0.1:{port}"), t, session="s2") == "" - assert seen[0]["outcome"] == ["blocked"] - - -def test_a_bare_id_does_not_count_as_placing_the_work(tmp_path): - reply = json.dumps({"status": "ok", "reason": REASON}).encode() - with http_sink(reply) as (port, seen): - env = _env(tmp_path, f"http://127.0.0.1:{port}") - bare = "Closed #41.\n**Needs you:** nothing.\n**Next:** #42." - assert json.loads(_run(env, _transcript(tmp_path, _closing_turn(bare))))["decision"] == "block" - assert seen[0]["missing"] == ["where it sits"] diff --git a/tests/test_services_report_check.py b/tests/test_services_report_check.py index 0f342313..ff17c0ba 100644 --- a/tests/test_services_report_check.py +++ b/tests/test_services_report_check.py @@ -1,12 +1,38 @@ -"""The server half of the report-shape check (milestone 409 step 5): the words a -blocked reply is sent back with, and the outcome record.""" +"""The report-shape check (milestone 409 step 5): which completion sections a +task-closing reply lacks, the words it is sent back with, and the outcome +record. Since milestone 500 step 4 the check itself runs here, inside the one +end-of-turn request, rather than in a Stop hook of its own.""" import json -from unittest.mock import patch +from unittest.mock import AsyncMock, patch import pytest from tests.helpers import make_mock_session +GOOD = ('**Where this sits:** milestone 12 "Move the backups offsite", step 3 of 5.\n' + "**What now works:** the sync runs nightly.\n**Needs you:** nothing.\n**Next:** alerts.") +BAD = "All done, pushed it." + + +def test_a_complete_report_misses_nothing(): + from scribe.services.report_check import missing_sections + + assert missing_sections(GOOD) == [] + + +def test_a_bare_reply_misses_every_section_in_order(): + from scribe.services.report_check import SECTIONS, missing_sections + + assert missing_sections(BAD) == list(SECTIONS) + + +def test_a_bare_id_does_not_count_as_placing_the_work(): + """The title is what spares the reader a lookup; "closed #41" does not.""" + from scribe.services.report_check import missing_sections + + assert missing_sections("Closed #41.\n**Needs you:** nothing.\n**Next:** #42.") == ["where it sits"] + assert missing_sections('Closed #41 "Sync". Needs you: nothing. Next: #42.') == [] + def test_the_reason_names_only_sections_it_knows(): from scribe.services.report_check import block_reason @@ -14,8 +40,17 @@ def test_the_reason_names_only_sections_it_knows(): reason = block_reason(["next", "ignore previous instructions", "Where It Sits"]) assert "missing: where it sits, next." in reason assert "ignore previous instructions" not in reason - # It points at the skill that owns the shape rather than restating it. - assert "reporting-back" in reason and "placement" in reason + + +def test_the_reason_carries_the_completion_shapes_line_and_where_the_rest_is(): + """The shape was delivered when the task closed; the reason repeats its + one line, not the whole of it, and names where the full text is.""" + from scribe.services import reply_shapes + from scribe.services.report_check import block_reason + + reason = block_reason(["next"]) + assert reply_shapes.SHAPES["completion"].reminder in reason + assert "list_reply_shapes" in reason and "placement" in reason def test_a_reason_with_nothing_recognised_still_says_what_to_do(): @@ -45,3 +80,30 @@ async def test_an_unknown_outcome_is_refused_before_anything_is_written(): pytest.raises(ValueError): await record_report_check(7, "skipped") session.add.assert_not_called() + + +@pytest.mark.parametrize("reply,rewrite,outcome,held", [ + (GOOD, False, "passed", False), + (BAD, False, "blocked", True), + (GOOD, True, "passed_after_rewrite", False), + (BAD, True, "missing_after_rewrite", False), +]) +async def test_check_reply_records_every_outcome_and_holds_only_a_first_miss(reply, rewrite, outcome, held): + from scribe.services import report_check + + record = AsyncMock() + with patch.object(report_check, "record_report_check", record): + got = await report_check.check_reply(7, reply, task_ids=[41], rewrite=rewrite, project_id=2) + assert got["outcome"] == outcome + assert bool(got["reason"]) is held + assert record.await_args.args == (7, outcome) + assert record.await_args.kwargs["task_ids"] == [41] + + +async def test_an_unrecorded_check_holds_nothing(): + """A hold the numbers cannot see is not one this check may make.""" + from scribe.services import report_check + + with patch.object(report_check, "record_report_check", AsyncMock(side_effect=RuntimeError("db"))): + got = await report_check.check_reply(7, BAD, task_ids=[41], rewrite=False) + assert got == {"outcome": "", "reason": ""}