Commit Graph
6 Commits
Author SHA1 Message Date
bvandeusenandClaude Opus 5.5 c6cdfc2172 feat(moments): the reply moment holds a finished reply for one read (milestone 458 step 4b, #4922)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / integration (push) Successful in 1m1s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Python tests (push) Failing after 1m25s
CI & Build / Build & push image (push) Skipped
The reply is the one act no tool call marks, and it is where "let me
know if it works" gets said. A new Stop hook (scribe_reply_check.sh)
sends the finished reply to POST /api/plugin/reply-rules, which checks
it twice:

- mounted: every unopened RULE on reply.report, plus reply.ask when the
  reply asks a question. Deterministic.
- semantic: the reply's head and tail against every rule's trigger, on a
  new ranked surface, reply_rule. It is the backstop for whatever the
  earlier arms missed. Its floor is its stop bar (default 0.80, budget
  1), with its own Settings dials. The new stop_only stage records
  surfacing for the rule that holds and nothing else, because nothing
  else reached anyone.

Following the operator's ruling from 456 step 8, a rule that holds blocks
once, in the server's words. The hook blocks only on a reason it was
given, so an unreachable instance never stops a session, and it never
holds the rewrite. The ledger is the act checkpoint's own, so a rule
holds a session once across both doors and the per-session cap counts
both.

The turn reader moved from the report check into scribe_defs.sh
(scribe_turn_facts / scribe_turn_fact), so the two Stop hooks read a
turn the same way. The output was checked identical on a real transcript.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 12:30:07 -04:00
bvandeusenandClaude Opus 5.5 78653130d6 feat(moments): mounted rules arrive when their moment happens, through every door (milestone 458 step 4a, #4922)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 57s
CI & Build / Python tests (push) Failing after 1m19s
CI & Build / Build & push image (push) Skipped
A rule mounted on a moment now reaches the session when an act reaches
that moment, with no semantic match involved:

- run_moment_arm on the pipeline: a lookup, not a ranked search. Each
  line names the moment and the act that reached it ("at work.deliver,
  reached by `git push`"), so a misfire is visible where it lands and
  can be unmapped in-session. A repeat is cited, not quoted; fresh
  rules are recorded surfaced under source moment_rule with the moment
  in detail. No retrieval_logs row, as for the other lookups, so no
  latency is persisted for this arm.
- rule_scope: a rule's home clause, moved out of semantic_search_rules
  so the moment lookup scopes by the same one.
- rulebooks.rules_on_moments / mounted_moments.
- The plugin door: a catch-all PreToolUse hook (scribe_moment.sh). It
  keeps /moment-tools' answer on disk for five minutes, so a call to a
  tool that cannot reach a mounted rule sends nothing, and an install
  that has mounted nothing sends one request per window. It shares the
  rules ledger with the other arms and fails open silently.
- The MCP door: Scribe's own tools named by the shipped mappings carry
  moment_rules in their response, so a client without the plugin gets
  them too. The hook skips those tools. A guard pins the attach on
  every one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 12:21:45 -04:00
bvandeusenandClaude Opus 5.5 f1fbdf746a feat(shapes): the agent judges what it wrote, at the end of the turn (milestone 439 steps 1-3)
CI & Build / Python lint (push) Successful in 5s
CI & Build / Plugin hooks (push) Successful in 18s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m38s
CI & Build / Build & push image (push) Successful in 33s
Recording used to be decided by machinery — the only "record it" prompt
fired when a same-named copy already existed (#2664), so a first instance of
a reusable piece was never asked about, and judgment arrived only through
audits. Now the question is asked where the knowledge is: the end of the
turn that wrote the code, of the agent that wrote it.

- Write hooks keep `<sid>.written.ids` (path, kind, name) for every
  definition a write names; a new file adds a `file` line for its stem — a
  candidate in any language without a framework rule (scribe_written_append).
- Stop hook scribe_shape_check.sh sends the ledger to GET
  /api/plugin/shape-check and blocks once, in the server's words, when
  anything is unjudged. Same discipline as the report check: never twice,
  never without a recorded check, another hook's loop left alone; the ledger
  is kept when the instance cannot be reached.
- shape_ledger.unjudged_shapes: no row, unclassified, scoped and hook stamps
  are unjudged; an agent/audit/import verdict is not. A snippet recorded at
  the shape answers for it until the refresh stamps it canonical.
- services/shape_check owns the reason text and records every outcome in
  app_logs (passed / blocked / judged_after_block / left_after_block).
- classify_shapes(repo=…) judges a shape the ledger has not synced yet via a
  provisional row under a bound repo; the sync confirms it, or vanishes and
  revives it with the verdict intact. An unbound repo is refused.

Plugin version minted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 08:39:36 -04:00
bvandeusenandClaude Opus 5.5 4fb53b844d refactor(mcp): _INSTRUCTIONS orients the workflow, not a rulebook (#4389)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Failing after 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m35s
CI & Build / Build & push image (push) Successful in 23s
Spike #4389 read the spec, Claude's docs and a dozen servers: the field is
for how the tools fit together, and the field runs ~600-1,600 characters.
Ours sat at the 2,048 cap as a keyword index that also carried stance.

- JUDGE, REPORT and MISSED leave the index. They fire mid-work, not at
  session start; using-scribe and reporting-back state them in full, and
  `placement`/`report_back` cue reporting in-band. No skill text changes.
- The rest is rewritten as plain practices (1,503 chars) and keeps every
  session-start marker the ownership registry pins.
- INSTRUCTIONS_BUDGET 2000 -> 1600; the three index markers are dropped
  from the registry; the miss-route index test now checks that the index
  keeps what_might_apply and stays off the route.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 07:50:34 -04:00
bvandeusenandClaude Opus 5 dd80e2bc86 feat(409): a Stop hook checks that a reply closing a task has the completion sections (#4014)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 8s
CI & Build / integration (push) Successful in 47s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Python tests (push) Failing after 59s
CI & Build / Build & push image (push) Skipped
Everything else Scribe gives an agent arrives before the reply is written. A
Stop hook is the one moment the finished reply exists, so it is the last
chance to fix a report the operator can't read, and the only place adherence
to the shape can be measured.

- plugin/hooks/scribe_report_check.sh (Stop): deterministic, no model call.
  1. Did this turn close a task? That means an update_task/create_task call
     with status "done" since the turn's prompt, whose tool_result is not an
     error. Otherwise it stays silent, which covers most turns (one grep).
  2. Does the reply that ends the turn say where the work sits (a record by
     id and title, or step N of M), what needs the operator, and what comes
     next? Matched on those words, not on exact headings.
  3. If sections are missing, it blocks once. With stop_hook_active set, a
     rewrite is recorded (passed_after_rewrite / missing_after_rewrite) and
     never blocked again. A block loop started by another plugin (no marker
     from this hook) is left alone.
- Measured: every checked reply is reported to GET /api/plugin/report-check
  (passed / blocked / after rewrite). Turns that close nothing are not
  reported; they would cost a request per turn and add nothing to the rate.
  Outcomes go to app_logs as category "plugin", action "report_check".
- It blocks only when the block was recorded, and only in the server's words.
  The endpoint returns the block reason, so the hook carries timing and
  transport only (PACKAGING.md), and an unconfigured or unreachable instance
  never stops a session.
- The transcript format is read from real transcripts and marked in the hook
  as observed rather than documented. The Stop contract (transcript_path,
  stop_hook_active, decision/reason, no matcher, SubagentStop separate) was
  checked against the Claude Code hooks docs. A prompt-type hook was not
  needed: the deterministic check passed a real completion report from this
  session and blocked a stripped one.
- A pipefail trap was caught while exercising the hook: `tail | grep -q`
  reports failure exactly when grep matches, because tail dies of SIGPIPE.
  The prefilter reads through process substitution; the section checks use
  here-strings.
- Tests: an end-to-end hook suite over synthetic transcripts and the shared
  HTTP sink (silence, pass, server-worded block, rewrite recorded, foreign
  loop, errored write, earlier turn, unwritten reply, no recorded check, bare
  id), and service tests for the reason wording and the outcome record. Smoke
  event added to check_plugin; README and PACKAGING list the hook and
  endpoint. Plugin version minted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 18:34:46 -04:00
bvandeusenandClaude Opus 5 15621fa873 docs(410): a packaging contract, so a second client is a manifest and an adapter (#4032)
CI & Build / Python lint (push) Successful in 7s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 56s
CI & Build / Python tests (push) Successful in 1m26s
CI & Build / Build & push image (push) Successful in 16s
Step 5 of milestone 410 "One owner per piece of guidance". The operator
wants any attempt to package Scribe for another agent client to find the
repo already in the right shape.

plugin/PACKAGING.md (linked from the README) states:
- what every client package shares: plugin/skills/ (verbatim), the /mcp
  endpoint and its in-band responses, the /api/plugin/* adapter endpoints
  (context, retrieve, prior-art, tool-rules, processes), and one fmcp_ key
- what each client adds: a manifest; hooks limited to timing and transport;
  optional commands; adapter static text that never copies a skill or the
  index
- the Claude Code adapter file by file, as the worked example
- how Agent Plugins 1.0 clients (Codex, Cursor, Copilot/VS Code, Kiro,
  ChatGPT) and Gemini CLI would map, marked researched-not-tested (#4023)
- four open questions for the second package: passing the key to the MCP
  server, whether two manifests can share one folder, hook parity, and
  where process skills go

Hook audit: every hook prints only server-provided text, status or outage
lines, the running version, or the compaction reload pointer. No guidance
copies, so nothing moved.

Guard: test_the_skills_reference_nothing_outside_their_folder fails on a
relative path upward or a reference to plugin/, hooks/, commands/ or a
manifest from inside a skill, with a companion test showing it can fail.
Plugin version minted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 13:23:25 -04:00