CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / integration (push) Successful in 49s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / Python tests (push) Failing after 1m7s
CI & Build / Build & push image (push) Skipped
Every hook opened `command -v jq >/dev/null 2>&1 || exit 0`, so on a machine without jq the operator got no session context, no rules, no prior art and no process sync — and not one word saying why, because `exit 0` is indistinguishable from "ran fine, nothing to say". jq is absent by default on macOS, on the Debian/Ubuntu slim images, on Alpine and in most CI containers. That is not a prerequisite to document; it is the plugin handing its own packaging problem to whoever installs it. `tac` was worse: GNU-only, so the prior-art hook's enclosing-definition arm did nothing at all on every Mac, silently, from the day it shipped. It is not replaced but removed — scribe_defs judges each line independently, so extracting forward and taking `tail -1` is the same answer as reversing and taking the head, and it drops the early-exit `head` that #4042 was filed for. No server contract changed, so a lagging plugin cache keeps working. scribe_json.awk JSON -> IDX<TAB>PATH<TAB>VALUE. Two modes: `whole` for an event or a response body, `lines` for a transcript, where an unparseable record is dropped and the rest still read — the `map(try fromjson catch empty)` the jq program opened with. Arrays also report their LENGTH at `[#]`, which is what keeps "zero notes" distinct from "no answer" (#2932). scribe_turn.awk the turn-bounding program, replacing the thirty lines of jq in the Stop hook. scribe_defs.sh scribe_json_flat / _pick / _list / _len / _list_minus read, scribe_json_out writes the envelope (five copies of one shape, gone), scribe_urlenc replaces `jq -sRr '@uri'`. Percent-encoding goes through `od -tu1` rather than an awk character loop on purpose: awk's idea of a character follows the locale, so gawk reads an accented letter as one and mawk as two, and an encoder built on substr() would emit a different URL depending on which awk is installed. Encoding is defined on bytes. Verified byte-identical to `jq -sRr '@uri'`. Measured, not assumed. The per-event path costs 8ms against jq's 3ms. The transcript path was 70x slower until two fixes: the Stop hook now finds where the turn starts with a fixed-string grep before parsing (a needle carrying unescaped quotes cannot occur inside a JSON string, so it matches only at a record's top level — checked against a full JSON parse of a 27MB transcript: 152 prompt records, 152 matches, no misses, no extras), and the parser reads each token out of a 1024-byte window instead of copying the rest of the buffer per token, which was quadratic in line length on the 400KB tool results a transcript carries. Differential-tested against the jq program it replaces over 724 windows cut from three real transcripts — 724 identical, 0 mismatched, 45 of them exercising a real task close and a real reply. That sweep is what caught `scribe_turn.awk` never setting FS, which truncated every multi-word reply at its first space and was invisible to a test whose replies were all empty. check_plugin.py's `jq -R` lint becomes a guard against either binary coming back, and three smoke checks lose their `shutil.which("jq")` skip. jq is not in `ci-python` either, so those three announced a skip on every CI run and had never once run there: removing the dependency from the product also closed a permanent hole in its verification. They pass now across all ten hooks. tests/test_hook_json_reader.py is a differential against Python's `json` over nested objects, arrays, unicode, escapes, control characters, empty cases and a value longer than the token window, plus the envelope, the encoder and the turn analyzer. 139 cases. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
109 lines
4.5 KiB
Awk
109 lines
4.5 KiB
Awk
# Scribe plugin — what happened in the LAST TURN of a transcript (#4107).
|
|
#
|
|
# Reads the flat `IDX<TAB>PATH<TAB>VALUE` stream that scribe_json.awk produces
|
|
# in `mode=lines` from a Claude Code transcript, and answers the four questions
|
|
# scribe_report_check.sh asks. It replaces a thirty-line jq program; the shape
|
|
# of the answer is unchanged, so the hook around it reads the same.
|
|
#
|
|
# bounded 1 if a user prompt was found in the window, else empty. A window
|
|
# with no prompt cannot be cut into a turn, and the hook then says
|
|
# nothing rather than guessing — `tail -n 3000` cuts wherever it
|
|
# cuts, and a turn that began before the cut is not this hook's.
|
|
# closed how many task-closing tool calls SUCCEEDED in that turn.
|
|
# task_ids their task ids, comma-joined, for the server to record.
|
|
# reply the assistant text after the last action, STILL JSON-ESCAPED and
|
|
# on one line. The caller decodes it with scribe_json_unescape —
|
|
# decoding here would put newlines into a line-oriented format.
|
|
#
|
|
# A SIDECHAIN IS NOT THIS SESSION. Subagent records interleave into the same
|
|
# file, and a subagent closing a task is not the operator's session closing
|
|
# one — counting those made the check fire on turns that closed nothing.
|
|
#
|
|
# A CLOSE THAT ERRORED IS NOT A CLOSE, which is why the error pass runs first:
|
|
# the tool_result carrying `is_error` arrives in a LATER record than the
|
|
# tool_use it refutes, so a single forward pass would have already counted it.
|
|
|
|
# TAB-SEPARATED, and stated rather than assumed. Under awk's default splitting
|
|
# a value is cut at its first SPACE, so `$3` of a reply line was the reply's
|
|
# first word — the kind of defect that hides completely behind a test whose
|
|
# replies are all empty. Found by differential-testing this against the jq
|
|
# program it replaces, over real transcript windows (#4107).
|
|
BEGIN { FS = "\t" }
|
|
|
|
{
|
|
i = $1 + 0
|
|
if (i > maxidx) maxidx = i
|
|
p = $2
|
|
if (p == ".type") { f[i, "type"] = $3; next }
|
|
if (p == ".isSidechain") { f[i, "side"] = $3; next }
|
|
if (p == ".isMeta") { f[i, "meta"] = $3; next }
|
|
# `.message.content` as a SCALAR is what marks a real user prompt; a tool
|
|
# result carries an array at the same path, and counting one as a prompt
|
|
# would cut the turn at the wrong place.
|
|
if (p == ".message.content") { if ($3 != "null") f[i, "str"] = 1; next }
|
|
if (substr(p, 1, 17) != ".message.content[") next
|
|
rest = substr(p, 18)
|
|
if (rest == "#]") { nb[i] = $3 + 0; next }
|
|
c = index(rest, "]")
|
|
if (c < 2) next
|
|
b[i, substr(rest, 1, c - 1) + 0, substr(rest, c + 1)] = $3
|
|
}
|
|
|
|
function is_mine(i) {
|
|
return (f[i, "side"] != "true")
|
|
}
|
|
|
|
function has_tool_use(i, j) {
|
|
for (j = 0; j < nb[i]; j++) if (b[i, j, ".type"] == "tool_use") return 1
|
|
return 0
|
|
}
|
|
|
|
END {
|
|
for (i = 1; i <= maxidx; i++)
|
|
if (is_mine(i) && f[i, "type"] == "user" && f[i, "meta"] != "true" && f[i, "str"] == 1)
|
|
prompt = i
|
|
if (!prompt) { print "bounded\t"; exit 0 }
|
|
|
|
for (i = prompt + 1; i <= maxidx; i++) {
|
|
if (!is_mine(i) || f[i, "type"] != "user") continue
|
|
for (j = 0; j < nb[i]; j++)
|
|
if (b[i, j, ".type"] == "tool_result" && b[i, j, ".is_error"] == "true")
|
|
errored[b[i, j, ".tool_use_id"]] = 1
|
|
}
|
|
|
|
for (i = prompt + 1; i <= maxidx; i++) {
|
|
if (!is_mine(i) || f[i, "type"] != "assistant") continue
|
|
for (j = 0; j < nb[i]; j++) {
|
|
if (b[i, j, ".type"] != "tool_use") continue
|
|
if (b[i, j, ".name"] !~ /(^|__)(update|create)_task$/) continue
|
|
if (b[i, j, ".input.status"] != "done") continue
|
|
if (b[i, j, ".id"] in errored) continue
|
|
closed++
|
|
t = b[i, j, ".input.task_id"]
|
|
if (t != "") ids = ids (ids == "" ? "" : ",") t
|
|
}
|
|
}
|
|
|
|
# Where the WORK stopped and the report began. Everything after the last
|
|
# action is the reply being checked; text emitted between two tool calls is
|
|
# narration mid-work, not a report, and holding it to the report shape would
|
|
# block turns that did report properly at the end.
|
|
act = prompt
|
|
for (i = prompt + 1; i <= maxidx; i++) {
|
|
if (!is_mine(i)) continue
|
|
if (f[i, "type"] == "user" || (f[i, "type"] == "assistant" && has_tool_use(i))) act = i
|
|
}
|
|
|
|
for (i = act + 1; i <= maxidx; i++) {
|
|
if (!is_mine(i) || f[i, "type"] != "assistant") continue
|
|
for (j = 0; j < nb[i]; j++)
|
|
if (b[i, j, ".type"] == "text")
|
|
reply = reply (reply == "" ? "" : "\\n") b[i, j, ".text"]
|
|
}
|
|
|
|
printf "bounded\t1\n"
|
|
printf "closed\t%d\n", closed + 0
|
|
printf "task_ids\t%s\n", ids
|
|
printf "reply\t%s\n", reply
|
|
}
|