Reply shapes ride the moments (milestone 500, steps 1–5) #209

Merged
bvandeusen merged 13 commits from dev into main 2026-10-09 16:40:04 -04:00
13 Commits
Author SHA1 Message Date
bvandeusenandClaude Opus 5.5 69dddb7551 feat(500): Settings shows the shipped reply shapes, what adjusts each, and how often each arrives
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / integration (push) Successful in 1m19s
CI & Build / Python tests (push) Successful in 1m58s
CI & Build / Build & push image (push) Successful in 39s
Settings > General > Reply shapes. One menu row per shape (core, completion,
asks, plan): its title and when it arrives, how many of the operator's
preferences adjust it, and its deliveries over 30 days (in full / as a
reminder, "—" when the counts cannot be read). A verdict line leads. Opening a
row shows the shipped text, read-only because it is product, and the
preferences mounted on its moment; each opens in the rule editor, and "Adjust
this" opens a new preference already mounted on the shape's moment with a
starting trigger.

The rule editor takes an optional preset (kind, moments, trigger) and, opened
outside the Rules view, asks where a new record lives instead of dropping it.
GET /api/retrieval/reply-shapes now returns reply_shapes.overview: the
catalog plus mounted preferences (the same lookup that delivers them) and
counts from the reply_shape delivery rows. #5497 (step 5 of milestone 500).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-09 16:19:45 -04:00
bvandeusenandClaude Opus 5.5 b00dc7c4c2 feat(500): reporting-back becomes the long-form reference behind the delivered shapes
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / integration (push) Successful in 1m17s
CI & Build / Python tests (push) Successful in 1m59s
CI & Build / Build & push image (push) Successful in 23s
The reply shapes are server product now, delivered at their moments, so the
skill stops restating them. Gone from it: the "Every reply" list, the
per-kind tables, and "The operator's own shapes come first" (preferences
arrive beside the shape on the same moments). It opens by saying where the
shapes come from (list_reply_shapes, the delivered core's header) and keeps
the reasoning: sections chosen not filled, a settled decision acted on,
placement from the record, who decides what, the assumptions an option
carries, the completion report worked in full, and the second pass. 13.5k to
10.7k characters.

The core gains the Finding kind the skill's table carried (2,150 of 2,200).
Tests follow the content: the kind and Approval-row pins move to the shapes,
a new test holds the worked example and the completion shape to the same
sections, and the guidance-ownership registry reads the delivered shapes as a
surface, owning the preference-wins and length topics there. using-scribe
points at the delivery and list_reply_shapes. #5496.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-09 15:55:02 -04:00
bvandeusenandClaude Opus 5.5 f68b73f922 feat(500): retire the fixed-question preference arm - completion-report preferences ride the moments
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m10s
CI & Build / Python tests (push) Successful in 1m54s
CI & Build / Build & push image (push) Successful in 31s
report_preference searched one constant string when a task closed and handed
matches back as update_task's reply_preferences. Milestone 500 step 3 delivers
the reply shapes and the preferences mounted beside them on the moments, so
the arm goes, whole (rule 22):

- services/reply_preferences.py, the REPORT_PREFERENCE RuleArm, its re-scorer
  and corpus entries, and the REPORTPREF constants
- update_task's reply_preferences block and its cue; the docstring now points
  at moment_rules
- the fixed_query field and the fixed_query_never_clears warning: this arm was
  the only one that set it, so the concept and the cannot_decline exemption
  go with it
- the two Settings fields for its floor and budget
- migration 0121 deletes its rows (logs, judgments, rule-usage events, tuning
  history, settings keys), operator-approved 2026-10-09. Without it the
  readout would call the source unregistered and its surfacings would count
  as ambient. The unindexed rule_usage delete is bounded by created_at.

#5496 (step 4 of milestone 500).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-09 15:51:04 -04:00
bvandeusenandClaude Opus 5.5 cee1eb1328 fix(500): the section-hold marker is a ledger by name, so a compaction sweeps it
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 1m3s
CI & Build / Python tests (push) Successful in 1m54s
CI & Build / Build & push image (push) Successful in 22s
The marker scribe_reply_check.sh writes after a completion-section hold sat in
the swept ledger directory as `<sid>.reportcheck`, and the sweep matches
`.ids` - test_every_session_file_in_a_swept_directory_follows_the_convention
caught it (run 799). Renamed `<sid>.reportcheck.ids` and added to the roster.
#5496.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-09 15:42:25 -04:00
bvandeusenandClaude Opus 5.5 6071007fe2 feat(500): one end-of-turn request - the completion-section check folds into the reply check
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m7s
CI & Build / Python tests (push) Failing after 1m25s
CI & Build / Build & push image (push) Skipped
The Stop hook sent the finished reply twice: scribe_report_check.sh checked a
task-closing reply for the completion sections in shell and reported to
/report-check, and scribe_reply_check.sh sent the same reply to /reply-rules
for the rule hold. Now the reply goes once. When the turn closed a task the
hook adds the close count and ids, and the server runs the section check
(services/report_check, the same three patterns) beside the reply hold,
folding both into one reason.

The report_check adherence log is still written for every checked reply
(milestone 409's number). A section hold marks the session, so its rewrite is
sent back once with rewrite:true to record how it came out, and is never held.
The block reason carries the completion shape's one line and points at
list_reply_shapes rather than at the skill. #5496 (step 4 of milestone 500).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-09 15:38:17 -04:00
bvandeusenandClaude Opus 5.5 ae66d83213 fix(500): the hooks decode a \u escape to UTF-8 in any locale and any awk
CI & Build / Python lint (push) Successful in 2s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m57s
CI & Build / Build & push image (push) Successful in 25s
scribe_json_unescape wrote each escaped code point with sprintf("%c", n),
which writes the character only under gawk in a UTF-8 locale; gawk in the C
locale and mawk in any locale write the single byte n % 256. The server
escapes every non-ASCII character, so "·" reached the session as 0xB7, "—"
as 0x14 and an emoji as a NUL wherever a hook ran without a UTF-8 locale.
The decoder now encodes UTF-8 itself under LC_ALL=C. A new test sends what
the server sends from a bare environment, in three locales (#5495).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-09 14:50:29 -04:00
bvandeusenandClaude Opus 5.5 eadb08c347 feat(500): reply shapes are delivered - the core every turn through the ledger, each slice at its moment, reply mounts before the reply (#5495)
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 1m12s
CI & Build / Python tests (push) Failing after 1m31s
CI & Build / Build & push image (push) Skipped
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 56s
- Every turn (/api/plugin/retrieve, UserPromptSubmit): the core reply shape
  leads the payload - in full the first time, as its one-line reminder
  after that - followed by whatever is mounted on reply.report, under the
  shared rule ledger. Fresh keys come back as shape_keys.
- The ledger is <sid>.shapes.ids in scribe-priorart (scribe_shapes_file /
  _seen / _append), so the compaction sweep that clears every .ids ledger
  is what brings the full core back after one.
- At a moment (/api/plugin/moment): the slice for that reply - completion
  on work.finish, asks on reply.ask, plan on work.plan - ahead of the
  mounted rules. reachable_tools now lists tools reaching a shaped moment
  even on an install with nothing mounted.
- Scribe's own tools (attach_moment_rules): reply_shape in the response,
  in full, since that door has no ledger. enter_project carries the core
  for clients with no prompt hook.
- Telemetry: one AppLog row per delivery (plugin / reply_shape), each
  shape with full or pointer and the door (turn, hook, mcp).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-09 14:45:25 -04:00
bvandeusenandClaude Opus 5.5 9fedcbea3d feat(500): who decides what - the agent settles what only it can see, the operator decides direction (#5494)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m7s
CI & Build / Python tests (push) Successful in 1m58s
CI & Build / Build & push image (push) Successful in 27s
The operator's line (2026-10-09), reconciling their 2026-09-21 ruling
(#4218, "you are the judge") with "not that Claude should make all
decisions": judging covers the work and its records - facts, findings,
measurements, whether work is done - and direction stays theirs:
priorities, what the thing should be, trade-offs only they can price,
anything they live with, and every act hard to undo or facing outward.

- reply_shapes: a one-line form in the core (every turn), the full text
  in the asks slice (where a question gets handed back). Slice budget
  1400 -> 1600 for it.
- reporting-back "You are the judge": "the operator reads the decision -
  they do not make it" replaced; the line is drawn by who can see the
  evidence, not by how hard the call is.
- plugin version minted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-09 14:35:55 -04:00
bvandeusenandClaude Opus 5.5 3648c78286 fix(500): the completion slice says "two questions", not "two tests" - the domain guard reads the word as software (#5493)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m3s
CI & Build / Python tests (push) Successful in 1m58s
CI & Build / Build & push image (push) Successful in 1m5s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-09 14:32:05 -04:00
bvandeusenandClaude Opus 5.5 a0390d9da7 feat(500): the reply shapes become server product content - a compact core and per-kind slices, each tied to its moment (#5493)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m12s
CI & Build / Python tests (push) Failing after 1m29s
CI & Build / Build & push image (push) Skipped
services/reply_shapes.py is the single source of the default reply shapes:
the core (every reply, ~1,800 chars against a 2,200 budget) and three
slices - completion on work.finish, asks on reply.ask, plan on work.plan.
Read through list_reply_shapes (MCP, read-only) and GET
/api/retrieval/reply-shapes, one service behind both doors.

Nothing delivers them yet; that is step 3. The skill still carries its
copy until step 4 shrinks it to the long-form reference.

The software-only vocabulary guard moves into tests/helpers.py
(DEV_ONLY, dev_only_hits) rather than becoming a fourth copy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-09 14:29:44 -04:00
bvandeusenandClaude Opus 5.5 acda59fc73 ci: integration finds only its own service containers
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m12s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 25s
The runner's docker daemon is shared, so `--filter name=<job> | head -n1` can
pick another repo's Postgres when two jobs run side by side (Steward run 8358,
Inkwell run 8653). Scope the lookup to this job's own task prefix and require
exactly one match. Done from Inkwell #5313 with the operator's permission; the
recipe is rule 79.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-07 23:17:10 -04:00
bvandeusenandClaude Opus 5.5 37ca2ee958 fix(lessons): convergence is measured against each lesson's own register, and its members must resemble each other - a flat band of no-rule lessons names nothing (#5193)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / integration (push) Successful in 2m1s
CI & Build / Python tests (push) Successful in 2m23s
CI & Build / Build & push image (push) Successful in 52s
Lessons share one register, so any two score well above unrelated text and
a fixed similarity bar sat inside that band: every no-rule lesson
"resembled" most of the others and the named group grew with the pool.

A neighbour now counts only when it stands above the lesson's own
background (median and MAD of its similarity to every lesson it can reach,
in that register's spread), and a group is a clique: every pair clears the
bar from both sides, so one broad lesson cannot join unrelated ones. The
bar, the minimum background and the fetch sizes are stated as defaults.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 22:58:39 -04:00
bvandeusenandClaude Opus 5.5 890c93e126 feat(scripts): measure_duplication - the DRY close-out measure: before/after duplicate share, and the copies that exist only after (#4745)
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / integration (push) Successful in 2m0s
CI & Build / Python tests (push) Successful in 2m40s
CI & Build / Build & push image (push) Successful in 25s
A multi-pass DRY audit could not answer whether it created duplication:
a pass's own shortening can leave two statements identical to a third,
and no per-pass scan sees a copy that did not exist when it ran. The
Librarian retrospective improvised this measure in a scratchpad; this is
it as a tool the DRY Pass process can name.

6-line windows over significant lines (comments, blanks and bare
punctuation dropped, strings folded), grouped by glob or extension.
Revisions are read through git archive, so nothing is checked out.
Stdlib only, so any project can run a scratch copy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 21:24:05 -04:00