Commit Graph
100 Commits
Author SHA1 Message Date
bvandeusen 2b52afcd72 Merge pull request 'Mount the corpus by proposal: a pass and an open-after-moment signal (milestone 458 step 7)' (#203) from dev into main
CI & Build / Python lint (push) Successful in 5s
CI & Build / Plugin hooks (push) Successful in 17s
CI & Build / Python tests (push) Successful in 1m59s
CI & Build / TypeScript typecheck (push) Successful in 56s
CI & Build / integration (push) Successful in 1m12s
CI & Build / Build & push image (push) Successful in 22s
2026-10-05 16:31:42 -04:00
bvandeusenandClaude Opus 5.5 c404a3127a test(retrieval): the proposal endpoints are routed on the app (milestone 458 step 7, #4925)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 19s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m2s
CI & Build / Python tests (push) Successful in 2m17s
CI & Build / Build & push image (push) Successful in 56s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 16:16:59 -04:00
bvandeusenandClaude Opus 5.5 570d1b6d5a test(backup): register the rule_moment_judgments builder in the import column guard (milestone 458 step 7, #4925)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 56s
CI & Build / Python tests (push) Failing after 1m44s
CI & Build / Build & push image (push) Skipped
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 16:12:30 -04:00
bvandeusenandClaude Opus 5.5 dfcf4df2e9 feat(moments): mount the corpus by proposal - a pass and an open-after-moment signal, both stopping at the operator (milestone 458 step 7, #4925)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 19s
CI & Build / integration (push) Failing after 43s
CI & Build / Python tests (push) Failing after 46s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Build & push image (push) Skipped
A rule written before moments existed is mounted on nothing. Step 7 records,
per (rule, moment), whether it belongs there and who said so:

- rule_moment_judgments (migration 0118, backup v23): suggested / confirmed /
  rejected, from a pass, the signal, or an edit. Moment "" is "no moment fits".
- The pass: rules_to_mount lists unjudged rules; propose_rule_moments records
  suggestions that mount nothing; rule_moment_proposals and
  judge_rule_moments put them to the operator. A confirm mounts, a reject is
  kept so the pair is never proposed again. Same service behind REST and a
  "Waiting on you" panel in Settings > Moments.
- Edits are judgments: set_rule_moments, the one mount write path, confirms
  what was added and rejects what was removed in the same transaction.
- The signal: scribe_moment.sh keeps a per-session acts ledger; when a rule
  is opened, scribe_record_opened.sh sends the last three minutes of it to
  /api/plugin/rule-opened. The acts resolve through the install's mappings;
  work.run and work.change are not evidence. Counted per distinct session
  with lesson_rules' evidence model, and once due the open returns one line
  asking the reader to offer the mount.
- scribe_session_end.sh removes the session's scribe-moment files.

Plugin 2026.10.05.2003.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 16:03:53 -04:00
bvandeusen 2db9e0c0c1 Merge pull request 'Skills and processes declare their moments, and the UI for moments (milestone 458 steps 5–6)' (#202) from dev into main
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m9s
CI & Build / Python tests (push) Successful in 1m59s
CI & Build / Build & push image (push) Successful in 19s
2026-10-05 15:41:56 -04:00
bvandeusenandClaude Opus 5.5 5bcc603310 fix(moments): the store loads rather than fetches, so the deadline guard does not read it as a bare fetch (milestone 458 step 6, #4924)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 52s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 2m2s
CI & Build / Build & push image (push) Successful in 45s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 14:47:12 -04:00
bvandeusenandClaude Opus 5.5 b73a689849 feat(moments): the human door onto mounts, mappings and per-moment telemetry (milestone 458 step 6, #4924)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 18s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Python tests (push) Failing after 1m34s
CI & Build / Build & push image (push) Skipped
CI & Build / integration (push) Successful in 2m7s
Everything an agent can do with moments, a person can now see and change
in the app.

- Rule editor: a moment picker beside the trigger. Catalog moments are
  ticked; a named procedure's `skill.<name>` is typed and checked as the
  server checks it. `moments` is always sent, so unticking the last moment
  unmounts the rule.
- Settings, Moments section (General tab): for each moment, what it
  means, the actions that reach it on this install (shipped ones can be
  switched off, the install's own removed), how many rules are mounted on
  it, and deliveries and agent opens over the window. Below that: named
  procedures with mounts, switched-off defaults with Restore, and a form to
  add an action.
- retrieval_telemetry.moment_usage: per moment, `delivered`, `rules`,
  `opened` (agent pulls after the first delivery there; an upper bound, as
  by_source is) and `last_delivered_at`. No ratio, because a mount is a
  person's statement, not a ranker's guess. Guarded on its own, and also
  reported in retrieval_summary as `moment_usage`.
- rulebooks.mount_counts; mounted_moments now derives from it.
- GET /api/retrieval/moments carries `mounted` and `usage` (?days=).
  DELETE /moments/mappings also reads the mapping from query parameters,
  since the browser's DELETE sends no body.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 14:42:30 -04:00
bvandeusenandClaude Opus 5.5 c489452a2d test(moments): an install removal still drops its tool while the skill loader stays listed (milestone 458 step 5, #4923)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 19s
CI & Build / TypeScript typecheck (push) Successful in 1m2s
CI & Build / integration (push) Successful in 1m20s
CI & Build / Python tests (push) Successful in 2m12s
CI & Build / Build & push image (push) Successful in 33s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 14:28:50 -04:00
bvandeusenandClaude Opus 5.5 f1fdc4a951 feat(moments): skills and stored processes declare the moments they are for (milestone 458 step 5, #4923)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m3s
CI & Build / Python tests (push) Failing after 1m25s
CI & Build / Build & push image (push) Skipped
Loading a procedure now also reaches the moment it is for. Loading the
reporting procedure is a report; loading the release procedure is a delivery.

- Bundled skills: each SKILL.md declares `metadata: moments:`. The same
  declaration ships as Skill defaults (BUNDLED_SKILL_MOMENTS), because the
  server never sees the plugin's files. test_skill_moments holds the two
  together and pins the plugin name that qualifies the skill.
- Stored processes: `moments` on create_process and update_process, stored
  in the note's data and returned by get_process. A `scribe-proc-<slug>`
  load resolves its process through the sync manifest at load time. The
  moments are not copied into the stub, which would go stale mid-session.
- reachable_tools lists the skill loader whenever anything is mounted, since
  a process's moments are known only when it loads.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 14:23:24 -04:00
bvandeusen 9fc2df080d Merge pull request 'Rules mount on moments, and one retrieval pipeline (milestones 458 steps 1–4, 456 steps 1–3)' (#201) from dev into main
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 1m0s
CI & Build / Python tests (push) Successful in 1m54s
CI & Build / Build & push image (push) Successful in 19s
2026-10-05 12:41:56 -04:00
bvandeusenandClaude Opus 5.5 b29689d4de fix(retrieval): declare the reply moment's surfacing site in FAN_OUT_SITES (milestone 458 step 4b, #4922)
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / Python tests (push) Successful in 1m50s
CI & Build / Build & push image (push) Successful in 40s
moment_delivery records the mounted half's surfacing under
source=rp.MOMENT_RULE_SOURCE. That is an attribute the registry's
extractor cannot resolve, so the site is now declared with the value it
emits, read from the pipeline's own constant.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 12:32:07 -04:00
bvandeusenandClaude Opus 5.5 c6cdfc2172 feat(moments): the reply moment holds a finished reply for one read (milestone 458 step 4b, #4922)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / integration (push) Successful in 1m1s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Python tests (push) Failing after 1m25s
CI & Build / Build & push image (push) Skipped
The reply is the one act no tool call marks, and it is where "let me
know if it works" gets said. A new Stop hook (scribe_reply_check.sh)
sends the finished reply to POST /api/plugin/reply-rules, which checks
it twice:

- mounted: every unopened RULE on reply.report, plus reply.ask when the
  reply asks a question. Deterministic.
- semantic: the reply's head and tail against every rule's trigger, on a
  new ranked surface, reply_rule. It is the backstop for whatever the
  earlier arms missed. Its floor is its stop bar (default 0.80, budget
  1), with its own Settings dials. The new stop_only stage records
  surfacing for the rule that holds and nothing else, because nothing
  else reached anyone.

Following the operator's ruling from 456 step 8, a rule that holds blocks
once, in the server's words. The hook blocks only on a reason it was
given, so an unreachable instance never stops a session, and it never
holds the rewrite. The ledger is the act checkpoint's own, so a rule
holds a session once across both doors and the per-session cap counts
both.

The turn reader moved from the report check into scribe_defs.sh
(scribe_turn_facts / scribe_turn_fact), so the two Stop hooks read a
turn the same way. The output was checked identical on a real transcript.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 12:30:07 -04:00
bvandeusenandClaude Opus 5.5 b3b616b20a fix(moments): the tools import the delivery module, and the tool-list cache leaves the swept directory (milestone 458 step 4a, #4922)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / integration (push) Successful in 49s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m48s
CI & Build / Build & push image (push) Successful in 32s
Two guards caught 7865313:

- test_mcp_tool_processes reads every coroutine in a tool module's
  namespace as a tool, and a name-imported attach_moment_rules looked
  like one. All seven modules now call moment_delivery.attach_moment_rules,
  and the parity guard accepts the attribute form.
- The session-ledger convention: a file in a swept directory must be a
  .ids ledger. The tool-list cache describes the install, not the
  context, so it moves to its own directory, scribe-moment, where a
  compaction does not sweep it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 12:24:30 -04:00
bvandeusenandClaude Opus 5.5 78653130d6 feat(moments): mounted rules arrive when their moment happens, through every door (milestone 458 step 4a, #4922)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 57s
CI & Build / Python tests (push) Failing after 1m19s
CI & Build / Build & push image (push) Skipped
A rule mounted on a moment now reaches the session when an act reaches
that moment, with no semantic match involved:

- run_moment_arm on the pipeline: a lookup, not a ranked search. Each
  line names the moment and the act that reached it ("at work.deliver,
  reached by `git push`"), so a misfire is visible where it lands and
  can be unmapped in-session. A repeat is cited, not quoted; fresh
  rules are recorded surfaced under source moment_rule with the moment
  in detail. No retrieval_logs row, as for the other lookups, so no
  latency is persisted for this arm.
- rule_scope: a rule's home clause, moved out of semantic_search_rules
  so the moment lookup scopes by the same one.
- rulebooks.rules_on_moments / mounted_moments.
- The plugin door: a catch-all PreToolUse hook (scribe_moment.sh). It
  keeps /moment-tools' answer on disk for five minutes, so a call to a
  tool that cannot reach a mounted rule sends nothing, and an install
  that has mounted nothing sends one request per window. It shares the
  rules ledger with the other arms and fails open silently.
- The MCP door: Scribe's own tools named by the shipped mappings carry
  moment_rules in their response, so a client without the plugin gets
  them too. The hook skips those tools. A guard pins the attach on
  every one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 12:21:45 -04:00
bvandeusenandClaude Opus 5.5 cc26054437 feat(moments): rules mount on moments, through every rule door (milestone 458 step 3, #4921)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m0s
CI & Build / Python tests (push) Successful in 1m51s
CI & Build / Build & push image (push) Successful in 29s
rule_moments (migration 0117) records which moments a rule arrives at, by
catalog name, cascading with the rule. rule_detail, the one seam every
rule door already returns through, gains moments beside system_ids: None
leaves the mounts alone, a list replaces them. get_rule and both
list_rules doors read them back, batched per page.

All five MCP rule/preference writes and the three REST ones take moments
and validate them before their create or update. An unknown name is
refused with the catalog listed and leaves no half-made rule behind; a
parity test pins that ordering on every door.

Backup v22 carries the mounts as a join table remapped through the rule
map; a real-Postgres round trip checks they land on the restored rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 11:03:51 -04:00
bvandeusenandClaude Opus 5.5 cf3de5bae1 feat(moments): actions map onto moments, with in-session corrections (milestone 458 step 2, #4920)
CI & Build / Plugin hooks (push) Successful in 18s
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m23s
CI & Build / Python tests (push) Successful in 2m0s
CI & Build / Build & push image (push) Successful in 36s
moment_actions.resolve(tool, input) names every moment a call reaches and
the action that reached it. One call can reach several: kubectl apply is
a run, a deliver and a reach outside the workspace. Command tools match by
how each segment of the line starts, with a word boundary; other tools by
field=value arguments. The MCP server prefix and case are ignored.

56 shipped defaults cover the harness tools, Scribe tools and common
command shapes. moment_mappings (migration 0116) holds what an install
adds and the defaults it switches off. A removal is a stored row, so an
upgrade does not switch the default back on.

Per the operator ruling, corrections happen in the session: map_action
and unmap_action (write tools) return now_reaches so the fix can be
confirmed in the same reply. list_moments now shows each moment's
actions on this install. REST mirrors both doors, recorded as human.
Backup v21 carries the mappings.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 10:58:00 -04:00
bvandeusenandClaude Opus 5.5 2ff7f2f34f feat(moments): the moment catalog rules will mount on, readable in-session (milestone 458 step 1, #4919)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 58s
CI & Build / Python tests (push) Successful in 1m53s
CI & Build / Build & push image (push) Successful in 32s
Fourteen generic moments of work (session.start, work.start … reply.ask)
plus the skill.<name> family, each with what it means and the kinds of
action that reach it, written for any kind of work rather than software
alone. The catalog is code because every install needs the same mount
points; which actions reach a moment is per-install data (step 2).

require_moment refuses an unknown name with the catalog listed, so a
typo cannot become a mount that never fires. list_moments (read-only)
and GET /api/retrieval/moments hand out the same catalog.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 10:50:45 -04:00
bvandeusenandClaude Opus 5.5 0720ab6dcf refactor(retrieval): the completion-report preferences run through the pipeline (milestone 456 step 3, #4905)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Build & push image (push) Successful in 28s
reply_preferences.completion_preferences was the last hand-written
rule search + record_retrieval + record_rule_surfaced triple outside the
pipeline. It is now REPORT_PREFERENCE, a RuleArm with kind="preference"
over COMPLETION_QUERY, read back as records rather than as lines.

- RuleArm gains `kind`. _ranked asks the ranker for that kind and drops
  anything else on the way out, before the row is logged.
- RuleResult gains `shown`, the (score, rule) pairs behind the lines, for
  callers that return records.
- RuleMoment.project_id may be None: logged as given, searched as
  `project_id or None`. report_preference rows keep their NULL project.
- Guards: reply_preferences now has 0 direct rule searches. The registry
  constant-source example moves from report_preference to preference_slot,
  the pipeline slot that records through a module constant.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 08:38:05 -04:00
bvandeusenandClaude Opus 5.5 b72c9a92e2 test(retrieval): two source guards read the pipeline where the rule arms now live (milestone 456 step 2, #4904)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 49s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Build & push image (push) Successful in 30s
CI run 8137 had two failures, both source-scanning guards whose property
moved to the pipeline. The property itself still held:

- test_every_surface_name_is_a_real_telemetry_source looked for
  source="<name>" literals. The rule arms now record under their spec field,
  so the spec is the join key, read from RULE_ARMS.
- test_every_hook_rule_search_says_which_project_it_is_for counted 4 direct
  searches in plugin_context. It now expects 0 there, so a copied arm fails
  the count. It also asserts that the pipeline has exactly one search and
  that its keyword literal carries project_id and never everywhere.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 08:07:26 -04:00
bvandeusenandClaude Opus 5.5 f6b824b214 refactor(retrieval): the three rule arms run through one pipeline (milestone 456 step 2, #4904)
CI & Build / Plugin hooks (push) Successful in 19s
CI & Build / Python lint (push) Successful in 2s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / Python tests (push) Failing after 1m28s
CI & Build / Build & push image (push) Skipped
CI & Build / integration (push) Successful in 1m15s
prompt_rule, pre_tool_rule and the rule half of the write path each wrote
the same steps out by hand: search, band, split fresh from repeats, log the
call before any early return, reserve a slot, render, record surfacings.
The copies drifted, and #3497, #3750 and #3752 were fixed one copy at a time.

- New services/retrieval_pipeline.py. run_rule_arm runs the stages in one
  order. RuleArm is the spec (source, band, compact_tail, checkpoint,
  preference_slot). RuleMoment is the query and the session ledger.
- The I/O (ranker and recorders) is passed in as RuleIO. plugin_context
  resolves it from its own names at call time, so existing patch points
  still apply.
- _rule_band, _rule_hint_line, checkpoint_for, checkpoint_reason, the band
  constant and the preference slot moved into the pipeline unchanged.
  plugin_context re-exports them.
- A recorder that raises now costs only its row, never a rendered line.
- The flags reproduce today exactly. Whether prompt_rule should band, and
  whether the act arms should reserve a preference, are step 7.
- Registry: the pipeline call sites are declared in FAN_OUT_SITES, with
  values read from the specs.
- The #3497 structural guard now checks the one implementation, and that
  plugin_context writes no rule-source row of its own.

plugin_context.py: 3,558 -> 2,980 lines.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 08:01:23 -04:00
bvandeusenandClaude Opus 5.5 be62abf142 perf(retrieval): one embedding per query per plugin request (milestone 456 step 1, #4903)
CI & Build / Plugin hooks (push) Successful in 17s
CI & Build / Python lint (push) Successful in 2s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m25s
CI & Build / Python tests (push) Successful in 2m6s
CI & Build / Build & push image (push) Successful in 1m25s
A single plugin request fans one query out to several ranked arms. The
operator message is searched by auto_inject, its reuse and lesson slots,
prompt_rule, the preference slot and rule_via_lesson, and each arm called
get_embedding on its own: up to six model calls for one vector.

- embeddings.query_embedding_memo(): a request-scoped ContextVar memo that
  get_embedding consults. Outside a scope the default is None, so every
  other caller is unchanged. A failed embedding is not remembered.
- memoized_query_embeddings decorates /retrieve, /tool-rules and
  /prior-art.
- get_embeddings (document chunks) is untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 07:55:06 -04:00
bvandeusen 802ead748b Merge pull request 'Rules never-opened advice per corpus (#4798); a lesson's name is one line (#4797)' (#200) from dev into main
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 53s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Build & push image (push) Successful in 22s
2026-10-03 23:11:18 -04:00
bvandeusenandClaude Opus 5.5 d2ac7bf220 fix(lessons): a lesson's name is one claim on one line (#4797)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m48s
CI & Build / Build & push image (push) Successful in 28s
#4797 "A lesson's what takes a whole narrative, and the narrative becomes
its name". `what` is the title every listing and menu prints. Nothing
enforced its documented "one line", so lessons written with the incident
in `what` printed up to ~1,500 characters as a menu line, buried the
claim, and diluted the trigger in the embedded title. That was 14 of the
39 lessons on this install.

- lessons.require_claim refuses a `what` over WHAT_MAX_CHARS (240) or
  running over several lines. The refusal says the story goes in
  `insight`.
- Both doors call it before writing:
  - MCP create_lesson / update_lesson;
  - REST create / update.
  An update checks only a NEW name, so a lesson stored with a long one can
  still take the edit that repairs it.
- lessons.claim_line is the display half. _menu_name shows an over-long
  stored name as its first sentence marked " …", because a door guard
  does not undo rows already stored and an unrepaired install would keep
  printing them.
- The tool docstrings state the bound.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 23:04:53 -04:00
bvandeusenandClaude Opus 5.5 9909cd2450 fix(retrieval): the rules never-opened warning stops sending the reader to a review tool that refuses rules (#4798)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m7s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 26s
#4798 "The rules corpus's surfaced_never_pulled warning sends the reader to
menus_to_review, which can only review auto_inject". #4772 gave the warning
one remedy for both corpora: "judge a sample with menus_to_review". That is
right for notes, where a line carries its passage. For rules it is a dead
end: retrieval_review.REVIEWABLE holds only auto_inject, so the tool
refuses every rule arm.

- The reading is now per corpus (_NEVER_PULLED_READING). For rules, the
  text says:
  - the count includes rules that only arrived in a listing;
  - a rule can rightly be set aside on its trigger alone;
  - no judged sample exists for the rule arms;
  - a rule set aside again and again is a trigger to fix (update_rule
    when_to_apply, then what_might_apply), not a floor.
- The tool docstring says the same.
- The guard is tied to REVIEWABLE, so it can fail in both directions: if
  the rules text names menus_to_review while no rule arm is reviewable, or
  if a rule arm becomes reviewable and the text still says there is no
  sample.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 23:02:32 -04:00
bvandeusen 8b9ff0c5c1 Merge pull request 'Auto-inject looks up records named by number (#4796)' (#199) from dev into main
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 1m47s
CI & Build / Build & push image (push) Successful in 18s
2026-10-03 22:16:01 -04:00
bvandeusenandClaude Opus 5.5 f95ffae972 feat(retrieval): a record the operator names by number reaches the menu (#4796)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m47s
CI & Build / Build & push image (push) Successful in 37s
#4796 "A record the operator names by number reaches auto-inject only if
its wording happens to match". "yes go ahead with 4448" says which record
is meant, but a number means nothing to an embedding, so the prompt menu
filled with records resembling the words around it.

- record_refs.named_record_ids reads the operator's raw prompt, never the
  reply-enriched query:
  - `#N`, unless the word before it marks another numbering (PR, CI, rule,
    milestone, system, log...);
  - a bare number of 3 or more digits that opens the message, follows a
    reference word or continues a list one started;
  - never a quantity ("300 seconds"), a date, version, path or fenced code.
- Each id is resolved through the ACL check; trashed or inaccessible ids
  are dropped.
- A "Named in your message" block leads the menu: the kind, the System,
  the name and the opening of the body, or the seen pointer if the record
  is already on the ledger. It takes no share of top_k, and the semantic
  lines leave those ids out.
- The block is booked under the new `named_ref` source, registered as an
  unbidden lookup that is allowed to be quiet. A named id that also ranked
  counts as suppressed in the auto_inject row, so #3668's identity holds.
- Named records now arrive even when the search finds nothing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 22:10:14 -04:00
bvandeusen a19c344770 Merge pull request 'No heading-only chunks, agent-message skip, refusals reach the agent (#4784, #4785, #4794)' (#198) from dev into main
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m14s
CI & Build / Python tests (push) Successful in 2m2s
CI & Build / Build & push image (push) Successful in 18s
2026-10-03 21:47:03 -04:00
bvandeusenandClaude Opus 5.5 5c6d9a7b55 fix(mcp): a tool's refusal reaches the agent with its reason (#4794)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 1m52s
CI & Build / Build & push image (push) Successful in 27s
SDK 2.x passes on only a ToolError's text. Any other exception becomes
UnexpectedToolError("Error executing tool X") and its message stays on the
server. ValueError is how every Scribe tool refuses — what was refused, why,
what to do instead — so every refusal reached the agent bare, and it could
only retry blind. Seen live: tune_retrieval(actor="operator") and
judge_menu(verdicts=[]) both answered "Error executing tool …" and nothing
else.

StrictArgsMCPServer.call_tool re-raises an UnexpectedToolError caused by a
ValueError as a ToolError carrying the message, in the SDK's own
"Error executing tool X: <reason>" shape. Other exceptions are crashes and
stay masked. The stale comment claiming the SDK returns ValueError text is
corrected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 21:31:10 -04:00
bvandeusenandClaude Opus 5.5 df4b37673f fix(plugin): a subagent hand-back is not an operator prompt (#4785)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 54s
CI & Build / TypeScript typecheck (push) Successful in 56s
CI & Build / Python tests (push) Successful in 1m49s
CI & Build / Build & push image (push) Successful in 29s
`<agent-message …>` turns were scored by auto-inject and the prompt rule
arm as if the operator had typed them, spending a retrieval and writing a
retrieval_logs row that reads as a real message (log 54355, found by the
#4772 review). Same class as #4142. Matched up to the tag name, since the
tag carries attributes; `<agent-messages>` and the tag named mid-sentence
are still retrieved against.

Plugin 2026.10.04.0125.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 21:25:37 -04:00
bvandeusenandClaude Opus 5.5 78fa01b025 fix(embeddings): no chunk is a heading alone (#4784)
Judging live auto-inject menus (#4772) found "## Work log — 2026-08-21",
"## Findings worth carrying forward" and "Two dev→main PRs this session."
each embedded as a whole chunk. Text about nothing embeds close to
everything, so they took ranks 1–3 on vague queries and pushed real records
down.

Three ways the chunker made them, each closed:
- _split_paragraphs flushed a heading on its own when the paragraph under it
  was long. A thin lead is now carried into the paragraph that follows, and
  a hard cut never lands in the first quarter of the budget, where the
  newline it finds closes a heading.
- A heading with no body (above its ### parts) became a section of its own
  when the section before it was full. A thin section now takes the next.
- A short lead-in or a short last section stood alone. The lead-in joins
  what follows; a thin tail joins what precedes it.

A continuation piece now repeats the LAST heading before it, since one
section can now hold several. CHUNKER_VERSION 2 → 3, so the startup backfill
re-embeds; tuned dials will report their shape as changed, which is true.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 21:25:37 -04:00
bvandeusen c389ed61d2 Merge pull request 'The review drops records written after the call it re-runs (#4773)' (#197) from dev into main
CI & Build / Python lint (push) Successful in 5s
CI & Build / Plugin hooks (push) Successful in 18s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m59s
CI & Build / integration (push) Successful in 1m5s
CI & Build / Build & push image (push) Successful in 23s
2026-10-03 20:37:08 -04:00
bvandeusenandClaude Opus 5.5 5d0e97576b fix(retrieval): the review drops records written after the call it re-runs (#4773)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m8s
CI & Build / Python tests (push) Successful in 1m53s
CI & Build / Build & push image (push) Successful in 36s
The first live sample ranked records the call could never have been offered:
the session that made a call writes the decision, often quoting the message,
and that record tops the re-run. Judged, it inflates on_point exactly where
the budget is decided.

- _rerun over-fetches by POSTDATED_SLACK, drops records created after the
  call before ranking, and names them in `postdated`.
- each line carries `changed_since_call` (updated_at or a work log after
  the call) and `logged` (shown fresh then).
- judge_menu shares the re-run, so a post-dated record cannot be judged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 19:42:20 -04:00
bvandeusen c50c1beb83 Merge pull request 'A review pass judges whether injected lines related (#4772)' (#196) from dev into main
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / integration (push) Successful in 1m19s
CI & Build / Python tests (push) Successful in 2m6s
CI & Build / Build & push image (push) Successful in 19s
2026-10-03 18:17:49 -04:00
bvandeusenandClaude Opus 5.5 6598c7fa85 feat(retrieval): a review pass judges whether injected lines related — menus_to_review, judge_menu and a judged readout (#4772)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 44s
An open rate cannot say whether a menu line related: every line carries its
matched passage (#4364), so "not opened" covers unrelated, enough as shown,
and already in context. #4772 "Injected notes are never judged".

- retrieval_judgments (0115): a reviewer verdict per line of a logged call,
  on_point / adjacent / unrelated, with its reason, rank, budget side and
  whether the agent opened it within the hour.
- menus_to_review re-runs a random sample of unjudged auto_inject calls with
  the arm's own parameters, past its budget, passage on every line.
  judge_menu records verdicts, re-deriving rank from a fresh re-run.
- retrieval_telemetry gains a judged block (by rank, within/beyond budget,
  on_point_unopened). surfaced_never_pulled stops blaming titles.
- missed-retrieval guidance names the review before a budget move.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 15:34:14 -04:00
bvandeusen b4c637a7df Merge pull request 'Rulings readout and the prior-art 500 fix' (#195) from dev into main
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / Python lint (push) Successful in 2s
CI & Build / TypeScript typecheck (push) Successful in 59s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m43s
CI & Build / Build & push image (push) Successful in 20s
2026-10-03 08:57:36 -04:00
bvandeusenandClaude Opus 5.5 0a1bb68808 feat(rulings): system_usage_events is read back — per-System counts and a telemetry block (#4769)
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Python tests (push) Successful in 1m44s
CI & Build / Build & push image (push) Successful in 36s
#4769 "Rulings are counted where someone will read them": milestone 444
step 4 wrote system_usage_events and nothing read it.

- retrieval_telemetry gains a `system_usage` block: surfacings and opens by
  source, distinct counts, and `by_system` naming the areas most shown.
  There is deliberately no pull-through ratio, because rulings travel in full
  in the line and opens are the exception.
- usage_for_systems (one GROUP BY) adds `usage` to the REST Systems list and
  detail, and to MCP get_system. MCP list_systems is unchanged.
- The Systems UI shows a "rulings shown N×" chip.
- rulings_pre_tool, rulings_write_path and mcp_get_system are now declared
  registry points; the registry guard covers their recorders.
- The Systems store merges a PATCH reply instead of replacing the row.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 08:49:20 -04:00
bvandeusenandClaude Opus 5.5 1b973ebd13 fix(write-path): a search hit's excerpt under snippet no longer 500s the prior-art route (#4768)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m5s
CI & Build / Python tests (push) Successful in 1m50s
CI & Build / Build & push image (push) Successful in 39s
#4768 "Write-path prior-art route 500s when a nameless search hit carries
its excerpt under `snippet`": _prior_art_line read item["snippet"] as a
snippet record's field dict, but a search hit carries its matched passage
there as a string. Read the name only from a dict.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 08:39:11 -04:00
bvandeusen a24f6b6b41 Merge pull request 'Rulings reach the work they govern — milestone 444 steps 3-4' (#194) from dev into main
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m11s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 20s
2026-10-03 08:29:25 -04:00
bvandeusenandClaude Opus 5.5 891f375715 test(rulings): the hook/route parity guards pin the rulings params (#4757)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 49s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m49s
CI & Build / Build & push image (push) Successful in 33s
CI 7939 went red on test_the_hook_and_the_route_agree_on_every_parameter_name:
the Bash hook now sends root and cwd, and the guard pins the exact set. Both
guards (tool-rules and prior-art) now also pin seen_ruling_systems, spelled
once in scribe_rulings_query, and check the routes read all three.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 23:12:40 -04:00
bvandeusenandClaude Opus 5.5 556872c039 feat(rulings): a command or edit touching an area's files shows its rulings, once per session (milestone 444 step 4, #4757)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m6s
CI & Build / Python tests (push) Failing after 1m22s
CI & Build / Build & push image (push) Skipped
A System's rulings (the Rulings section of its description) now reach the
work by path, not by similarity. Both PreToolUse arms resolve the files a
command or edit names to the Systems whose path_patterns cover them, and the
first touch in a session shows each area's rulings in one line; a repeat is
a one-line reference. A lookup, so no floor, no budget, no retrieval_logs row.

- services/system_rulings: parse_rulings, command_paths (reads and writes,
  relative to the repo root from any cwd; flags, URLs, globs skipped),
  rulings_for_paths
- /tool-rules takes root, cwd and seen_ruling_systems; /prior-art takes
  seen_ruling_systems; both return ruling_system_ids
- hooks share <sid>.rulings.ids (cleared on compaction by the ledger naming
  convention); the Bash hook sends the repo root and cwd
- system_usage_events (migration 0114): surfacings by source, pulls from
  get_system; carried by backup (v20) through the system map
- writing-records: rulings also arrive when the area's files are touched

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 23:08:32 -04:00
bvandeusenandClaude Opus 5.5 a113c72b4f feat(systems): a System names its files — path patterns stored, validated and matched (milestone 444 step 3, #4756)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m6s
CI & Build / Python tests (push) Successful in 1m50s
CI & Build / Build & push image (push) Successful in 43s
A System gains path_patterns: globs relative to the repo root (* within one
directory, ** across any depth, a plain directory covering everything under
it). One service validates them for every door, so the web UI and MCP refuse
the same bad pattern with the same message. systems_for_paths resolves paths
to every active System that covers them, which step 4 (#4757) uses to deliver
an area's rulings when its files are touched.

- schema: systems.path_patterns JSONB NOT NULL default [] (migration 0113)
- service: normalize_path_patterns, path_matches, systems_for_paths
- routes + MCP create_system/update_system accept it; [] clears
- web UI: a Files field in the create and edit forms, patterns on the card
- backup carries it through export and restore
- using-scribe reflex 7: tagging work keeps a System's files current

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 22:55:26 -04:00
bvandeusen 6477c20664 Merge pull request 'Milestone 444 steps 1-2: an operator's ruling lives on the System it governs, and code is read as behaviour, not intent' (#193) from dev into main
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m46s
CI & Build / Build & push image (push) Successful in 18s
2026-10-02 19:46:50 -04:00
bvandeusenandClaude Opus 5.5 01d8a0b9f1 feat(guidance): an operator's ruling lives on the System it governs, and code is read as behaviour, not intent (milestone 444 steps 1-2, #4754 #4755)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m46s
CI & Build / Build & push image (push) Successful in 56s
A Librarian session contradicted a decision the operator had made 13 days
earlier. The ruling ("retry, then replace, never give up on a book") was
kept only as a quote in a work log, beside a session's reading of it that
capped replacements at 3. Three later sessions built on the reading, and one
carried the cap into an option as a "known cost", which the operator then
approved without being asked about it.

- writing-records.md: "A ruling goes on the System it governs". What a
  ruling is (the operator decided it; a later change could undo it), how it
  differs from a rule, and where it goes: a Rulings section at the end of
  the System description, one line each with who, when and the source
  record. Written the turn the operator decides; holds what is in force,
  not history; a charter line that contradicts a ruling is fixed in the
  same edit.
- using-scribe SKILL.md: reflex 1 says code tells you what a thing does,
  not what was wanted, and a limit read from code is unconfirmed until a
  System's Rulings says otherwise. Reflex 9 points to the ruling section.
- reporting-back: an option that carries existing behaviour says whose call
  it was (the operator's ruling, or a past session's never confirmed); one
  that contradicts a ruling is a Conflict.
- create_system / update_system docstrings: the Rulings section, and that
  description replaces the whole text.
- test_guidance_ownership: three topics pinned to their owners.
- Plugin version minted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 18:59:36 -04:00
bvandeusen f94ff9e8fc Merge pull request 'The web lesson editor records which rule a lesson is an instance of (milestone 440, #4658)' (#192) from dev into main
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m43s
CI & Build / Build & push image (push) Successful in 17s
2026-10-01 16:54:34 -04:00
bvandeusenandClaude Opus 5.5 cf4206469b feat(lessons): the web editor records which rule a lesson is an instance of (milestone 440, #4658)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 37s
The lesson editor now asks the question the create tool asks, while the
writer still has the situation in mind. It offers three answers:
- an instance of a rule, ticked from the rules the lesson resembles;
- "No rule fits", with a required reason;
- leave it open.

The rules are fetched before the save from a new GET
/api/lessons/rule-candidates. It runs the same rule_candidates search the
create door replies with, and passes "search unavailable" through as null,
apart from "nothing resembles it".

An edit sends an answer only when it changed. Re-sending the same rules
would re-stamp their judgments. Worse, a shared editor who cannot see the
owner's rule would reject it by sending a list without it. Leaving a linked
lesson open sends rule_ids=[], which is what clears the links.

"Leave it open" is not offered once "no rule fits" is on file: the server
has no way to take that answer back except by naming a rule. When "no rule
fits" completes a convergence group, the editor says so in a toast, since
the page it goes to does not recompute it.

Guards check that the payload uses the names both routes read, that the
three answers are offered, that the candidates route sits above the id route
and keeps null apart from [], and that the editor reuses .rule-chip.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 16:36:43 -04:00
bvandeusen 6024230fa9 Merge pull request 'Milestone 440 steps 4, 5, 7: rules surface through their lessons, convergence named at the write, both records show the link' (#191) from dev into main
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m50s
CI & Build / Build & push image (push) Successful in 19s
2026-10-01 16:06:03 -04:00
bvandeusenandClaude Opus 5.5 75cefe60e4 feat(lessons): both records show the link in the web UI (milestone 440 step 7, #4635)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / integration (push) Successful in 1m9s
CI & Build / Python tests (push) Successful in 1m53s
CI & Build / Build & push image (push) Successful in 41s
The lesson page gains an "Instance of" panel that holds one of three answers.
Unjudged is the fall-through, so it is stated rather than left blank:
- the rule(s) it was judged an instance of;
- "No rule — <why>";
- "Not yet judged".

Suggested links show what they rest on (distinct situations, and projects
when more than one), with Confirm / Not an instance for a reader who can
write. Rejected links stay listed with their reason.

The rule slide-over lists the lessons that are instances of it, plus the
suggestions waiting on a judgment. Each entry links through to the other
record, and kind and state wear the existing .rule-chip.

The write check moves to utils/permission.ts. The copy on the snippet page
looked for "edit", which the server never sends, so shared editors saw a
read-only page (#4640 "The snippet page hid its edit controls from shared
editors"). Guards pin the client's unions and write levels to the
service's, the model's and access.py's own values.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 15:46:27 -04:00
bvandeusenandClaude Opus 5.5 6d3dca0af5 feat(lessons): convergence is named at the write — no-rule lessons that keep landing in one situation suggest a rule (milestone 440 step 5, #4634)
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / Python lint (push) Successful in 2s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Build & push image (push) Successful in 27s
When a lesson is answered "no rule fits" (create_lesson / update_lesson on
both doors), the response looks for other no-rule lessons it resembles and,
once there are CONVERGENCE_LESSONS (3) of them, carries `convergence`: the
members, their incidents and projects, and a hint to draft the missing rule
with create_rule (operator approval as always) and point each lesson at it —
or to leave them as lessons when no single choice is right every time.

- convergence_group is the pure bar: distinct LESSONS count, incidents never
  stand in for them (one broad lesson cannot trigger it), and a group whose
  sources all point at one incident is one event written up several times.
- convergence_for searches lessons by the new one's claim + trigger
  (trigger_title) at CONVERGENCE_THRESHOLD 0.65 — above the menu's "worth
  showing", below the duplicate gate's "same record" — then keeps the ones
  with a lesson_no_rule answer. Fail-open. No sweep, no timer (#4183).
- Defaults stated as defaults (rules 32, 115).
- Tests: the bar (pure), the search with stubs, the door, and the no-rule
  filter against Postgres; conftest stubs convergence_for for unit tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 15:31:15 -04:00
bvandeusenandClaude Opus 5.5 f34249a2d8 test(rules): locate the pre-tool arm by what it does, not by its name (#4633)
CI & Build / integration (push) Successful in 52s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / Python tests (push) Successful in 1m48s
CI & Build / Build & push image (push) Successful in 26s
CI run 712 failed test_neither_rule_arm_logs_its_call_behind_a_results_guard
with "min() iterable argument is empty": the guard walked the function named
build_tool_rule_hint, which since fd2ebf4 is a wrapper that searches nothing
— the arm's body moved to _tool_rule_hint. The property it guards (the call
is logged before any early return) still held; the guard had pinned a name.

It now finds the async function that calls semantic_search_rules under the
pre_tool_rule source, so a later rename moves the guard with the arm instead
of emptying it (rule 167).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 15:11:28 -04:00
bvandeusenandClaude Opus 5.5 fd2ebf4c13 feat(rules): a rule surfaces through the lessons confirmed as its instances (milestone 440 step 4, #4633)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 18s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 55s
CI & Build / Python tests (push) Failing after 1m16s
CI & Build / Build & push image (push) Skipped
On all three rule arms (prompt_rule, pre_tool_rule, write_path_rule), after
the direct match, a lesson matching the moment at the notes menu's own bar
brings the rule(s) it is CONFIRMED to be an instance of, rendered in rule
voice with "Reached through lesson #N “…”, a recorded instance of it."

- Confirmed links only (lesson_rules.confirmed_lessons /
  confirmed_rules_in_scope). A suggested link carrying its rule would
  manufacture the co-arrival #4637 counts and prove itself.
- Scope kept: a lesson surfaces everywhere, but the rule it brings must be
  global or this project's own — milestone 414's boundary, not reopened
  through a side door.
- Suppression is by RULE: anything the direct band named or the session
  ledger holds is skipped, whichever lesson reached it.
- Its own slot (VIA_LESSON_LIMIT = 1), not a rule slot — a stated default,
  since the plan's "decide by measurement" has nothing to measure until
  links are confirmed (#4632). Its own source, rule_via_lesson: registered,
  RANKED for pull-through, logged whenever it searches.
- Nothing searched at all while no lesson carries a confirmed link, which is
  every install until one is judged.
- build_prompt_rule_hint / build_tool_rule_hint are now thin wrappers over
  the direct arms (_prompt_rule_hint / _tool_rule_hint); the arm leaves its
  query in `_via_query` only when it ran, so a disabled arm or a blank prompt
  brings no rule in either, and the key never leaves the server.
- Via-lesson rules are not fed to the #4637 co-surfacing recorder: only
  direct matches are evidence.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 15:03:41 -04:00
bvandeusen c08655914f Merge pull request 'Soft links: a lesson and a rule arriving together are proposed as a link (milestone 440 step 3, #4637)' (#190) from dev into main
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 56s
CI & Build / Python tests (push) Successful in 1m44s
CI & Build / Build & push image (push) Successful in 21s
2026-10-01 14:59:30 -04:00
bvandeusenandClaude Opus 5.5 dbbab859ce feat(lessons): soft links — a lesson and a rule arriving together in distinct situations are proposed as a link (milestone 440 step 3, #4637)
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Build & push image (push) Successful in 30s
When one hook response puts a lesson and a rule in front of the reader, the
pair is recorded as evidence on a SUGGESTED lesson_rule_links row; once the
pair has arrived together in PROPOSE_SITUATIONS (3) distinct situations, the
next co-arrival carries one line asking the reader to judge it with
judge_lesson_link. Nothing about surfacing changes: a suggested link carries
no rule anywhere (that is #4633, confirmed links only).

- lesson_rules: co_surfaced (fail-open; only pairs the reader could confirm —
  a lesson they may write, a rule they own; judged pairs gather nothing; one
  proposal per response; PROPOSE_COOLDOWN 6h between asks), plus the pure
  counting rules: situation_key, add_evidence, proposal_due, evidence_summary.
  A situation is the prompt on /retrieve (word tokens, sorted and
  de-duplicated, so trivial rewordings count once) and the FILE on
  /prior-art (every edit to one file is one situation).
- rules_for_lessons shows a suggested link's evidence counts.
- plugin_context: build_autoinject_hint returns lesson_ids,
  build_prompt_rule_hint returns shown_rule_ids, build_write_path_hint records
  its own pair; routes/plugin /retrieve records the prompt pair. Shown lines,
  repeats included: relevance makes a co-arrival, not the session ledger.
- Evidence lives in the existing evidence column — no migration, and backup
  already carries it.
- Tests: counting rules (unit), the recorder against Postgres (bar, repeat,
  judged pairs, ownership, cooldown, evidence kept on confirm); conftest stubs
  co_surfaced for unit tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 14:30:58 -04:00
bvandeusen 6ba40c9e10 Merge pull request 'Lessons point at rules (milestone 440 steps 1–2) + the away-from-home usage readout' (#189) from dev into main
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / TypeScript typecheck (push) Successful in 58s
CI & Build / integration (push) Successful in 57s
CI & Build / Python tests (push) Successful in 1m40s
CI & Build / Build & push image (push) Successful in 21s
2026-10-01 13:48:01 -04:00
bvandeusenandClaude Opus 5.5 95be2512ab fix(lessons): build the rule-candidate query with trigger_title (#4631)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 1m0s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m53s
CI & Build / Build & push image (push) Successful in 27s
CI run 707 failed test_the_trigger_separator_is_spelled_in_exactly_one_place:
rule_candidates joined the lesson's claim and trigger with an inline " — ",
a second spelling of embeddings.TRIGGER_SEP (#3207). It now calls
trigger_title — the same join a rule's own document title is built with,
which was the point of the query shape in the first place.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 13:34:08 -04:00
bvandeusenandClaude Opus 5.5 f7d8dc2e55 feat(lessons): judged when written — a new lesson is offered its rules, and "no rule fits" is an answer (milestone 440 step 2, #4631)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 53s
CI & Build / Python tests (push) Failing after 1m15s
CI & Build / Build & push image (push) Skipped
"Which rule is this lesson an instance of?" now has three recorded answers:
a rule named (a confirmed link, #4630), no rule fits (new), or unjudged.

- Model + migration 0112: lesson_no_rule (lesson_id PK, CASCADE from the
  note; why; judged_at). A table rather than a key in notes.data, because
  that mirror is re-composed from the body on every edit and would erase it.
- Service (lesson_rules): set_no_rule rejects any confirmed link with the
  reason; a confirmation (set_lesson_rules or judge_link) deletes the answer;
  require_one_answer refuses both answers in one call before any write;
  judgments_for_lessons + attach_lesson_rules add rule_judgment (and no_rule)
  to every lesson payload; list_unjudged lists the open ones; rule_candidates
  searches rules with the lesson's claim + trigger at the explicit-search bar,
  None when the search could not run.
- MCP: create_lesson/update_lesson take no_rule; an unanswered create returns
  rule_candidates, rule_judgment and a rule_hint; list_lessons(unjudged=true).
- REST: the same on POST/PATCH /api/lessons and GET ?unjudged=1; create
  returns rule_candidates.
- Backup v19: a lesson_no_rule section, export (full and per-user) and import.
- Guidance: create_lesson docstring, writing-records.md in using-scribe (owner,
  pinned in test_guidance_ownership), create_rule docstring on linking the
  lessons a new rule governs. Plugin version minted.
- Tests: door units, integration for the three states, the rejection reason,
  scoping, cascade; backup registries.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 13:25:27 -04:00
bvandeusenandClaude Opus 5.5 c8393975c3 fix(mcp): classify judge_lesson_link as a write tool (#4630)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m42s
CI & Build / Build & push image (push) Successful in 31s
The tool-classification guard (test_mcp_auth) failed CI run 705: the new
judge_lesson_link tool sat in no set, so a read key would have been silently
denied it. It changes a link's state, so it belongs in _WRITE_TOOLS.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 12:44:06 -04:00
bvandeusenandClaude Opus 5.5 41e4fbaba1 feat(lessons): a lesson names the rule it is an instance of — lesson_rule_links (milestone 440 step 1, #4630)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Failing after 1m13s
CI & Build / Build & push image (push) Skipped
The link between a lesson (one concrete situation) and the rule that governs
it, with the operator's soft-then-hard design built into its state:
suggested while evidence accumulates, confirmed or rejected once judged. Only
confirmed will carry a rule in retrieval (#4633); rejected is kept so the pair
is never proposed again.

- models/lesson_rule_link.py + migration 0111: one row per (lesson, rule),
  CASCADE on both ends, indexed both ways, CHECK on state (rule 36), evidence
  JSONB and judged_at.
- services/lesson_rules.py: require_rules (validated before any write, so
  a bad id leaves nothing half-linked), set_lesson_rules (set-semantics;
  a dropped rule becomes rejected, not forgotten), judge_link, and the two
  reads. ACL: write on the lesson (share-aware), ownership of the rule; a
  reader sees only rules they own. Decorations are fail-open (#4286).
- MCP: create_lesson / update_lesson take rule_ids; get/create/update return
  `rules`; new judge_lesson_link tool. REST: the same on /api/lessons plus
  PUT /api/lessons/<id>/rules/<rule_id>. Rules: rule_detail carries `lessons`.
- Backup v18: export (full and user-scoped, both ends in scope), builder,
  importer; both column guards register the table.
- Tests: integration (states, set-semantics, judge, ACL all-or-nothing,
  cascade both ways, CHECK, one row per pair); unit (door wiring, judge
  registered, migration/model state agreement, backup skip and unjudged
  stays unjudged). conftest stubs the decorations for unit tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 12:38:02 -04:00
bvandeusenandClaude Opus 5.5 2e4c2d9493 feat(usage): count what happened away from a record's own project — the readout #3735 needs
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m6s
CI & Build / Python tests (push) Successful in 1m50s
CI & Build / Build & push image (push) Successful in 38s
Milestone 385 step 8 (#3735 "a lesson is recalled on a project it was not
written on") is defined by opened-on-another-project. The reader's project has
been recorded on every usage event since 0c8e109, but nothing compared it with
the record's own, so the criterion was still unreadable.

- note_usage.usage_for_notes: a second aggregate in the same session joins
  notes and counts surfaced_away_count / pulled_away_count. Counted only where
  both projects are known and differ; ranked surfacings only (#2477). The
  first aggregate is untouched, so events on deleted notes still count.
- empty_usage carries both keys zero-filled; every door that attaches usage
  (list_lessons, get_lesson, snippets, knowledge) gets them through
  attach_usage.
- UsageBadge tooltip says "On other projects: surfaced N×, opened M×" when it
  happened, and nothing when it did not.
- Tests: the mocked split test feeds both aggregates; a unit test for the
  away counters; a real-Postgres test that home, unreported and ambient
  events are all left out.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 10:47:03 -04:00
bvandeusen d7641e57a8 Merge pull request 'Shapes are judged when they are written (milestone 439 steps 1–6) + #4608 fixes' (#188) from dev into main
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 56s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Build & push image (push) Successful in 17s
2026-10-01 08:55:10 -04:00
bvandeusenandClaude Opus 5.5 582a5a4f48 feat(shapes): the practice is written where it is read, and the coverage line measures the slip (milestone 439 step 6)
CI & Build / integration (push) Successful in 47s
CI & Build / Python tests (push) Successful in 1m39s
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Build & push image (push) Successful in 29s
- reusing-code: "Before the turn ends — say what you built" — the four
  verdicts, the one classify_shapes(repo=…) call, and the component file as a
  shape. Description names the end-of-turn moment.
- shape-accounting: the writer judges; audits are the check that it held.
  The write path SUGGESTS (no more hook instances); component file rows and
  whole-file canon described; scoped covers Svelte too.
- _INSTRUCTIONS reuse line: "before the turn ends, say what you built
  (create_snippet the reusable, classify_shapes the rest)" — 1570/1600.
- Coverage line: "written-shape check (7d): N turns checked, M asked, K left
  unjudged", from the Stop hook's recorded outcomes; silent until the
  question has been put.
- test_guidance_ownership pins the new topic on reusing-code.
- prior-art hook header no longer says it stamps instance rows.

Plugin version minted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 08:46:58 -04:00
bvandeusenandClaude Opus 5.5 8b1567cf2c feat(shapes): a component file is a candidate shape in its own right (milestone 439 step 5)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 54s
CI & Build / Build & push image (push) Canceled after 0s
CI & Build / Python tests (push) Canceled after 1m42s
A single-file component defines `.card`, `.title` and a `Props`, never
anything named after itself — so the ledger could not hold "StatusChip is
canon" and could not say a new card was built where one already existed.

- coverage.is_file_unit: a file that RENDERS (its markup names classes, or it
  opens a <script>/<template>/<style> block) and defines nothing named after
  its stem gets one `file` row. Structural, not a framework list: Svelte and
  Vue components qualify, a TSX component already has its function's row, a
  module is accounted for by its definitions. Emitting every file would have
  put every module of every project in the todo at once.
- No body fingerprint on file rows, so an edit never re-asks a judgment.
- `file` is its own form and family; derive grouping skips it (`index`,
  `+page` repeat by convention). Divergence buckets it on its own, so a new
  component where a component canon dominates is a fair question.
- mark_canonicals: a snippet recorded at a path with no symbol makes that
  file row canonical; it still covers no definition inside.
- The write hooks apply the same test before noting a new file for the
  end-of-turn question.

Plugin version minted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 08:45:14 -04:00
bvandeusenandClaude Opus 5.5 4cf1c6f042 feat(shapes): the write-path hook suggests and no longer stamps (milestone 439 step 4)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m49s
CI & Build / Build & push image (push) Successful in 27s
The hook used to land its evidence as `instance` rows, classified_by=hook —
a permanent verdict nobody read, and the source of the weak stamps that
poisoned two ledgers (#4608). Now the agent that wrote the code judges it at
the end of the turn, and the hook's evidence is what it is shown.

- shape_ledger.suggest_write_path_instances (was stamp_write_path_instances):
  same evidence and form gate, but it writes a PROPOSAL — proposed_snippet_id,
  basis "reference" (named) or "semantic" (resembles), score — and never a
  status. A judged row is left alone; a new shape gets an unclassified
  provisional row so the suggestion reaches the end-of-turn question as
  "looks like #N". By-name uses edges stay: a fact, not a verdict.
- The prior-art line says "looks like #N … when you judge this turn's
  shapes, say whether it is", and the result key is `suggested`.
- write_time_divergence takes `suggested`; _RESEMBLE_MIN re-documented as the
  suggestion bar and the weak-stamp line.
- reusing-code skill no longer says the pull stamps an instance.
- Tests moved to the new contract, plus a guard that nothing the hook does
  writes classified_by="hook".

Plugin version minted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 08:42:10 -04:00
bvandeusenandClaude Opus 5.5 f1fbdf746a feat(shapes): the agent judges what it wrote, at the end of the turn (milestone 439 steps 1-3)
CI & Build / Python lint (push) Successful in 5s
CI & Build / Plugin hooks (push) Successful in 18s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m38s
CI & Build / Build & push image (push) Successful in 33s
Recording used to be decided by machinery — the only "record it" prompt
fired when a same-named copy already existed (#2664), so a first instance of
a reusable piece was never asked about, and judgment arrived only through
audits. Now the question is asked where the knowledge is: the end of the
turn that wrote the code, of the agent that wrote it.

- Write hooks keep `<sid>.written.ids` (path, kind, name) for every
  definition a write names; a new file adds a `file` line for its stem — a
  candidate in any language without a framework rule (scribe_written_append).
- Stop hook scribe_shape_check.sh sends the ledger to GET
  /api/plugin/shape-check and blocks once, in the server's words, when
  anything is unjudged. Same discipline as the report check: never twice,
  never without a recorded check, another hook's loop left alone; the ledger
  is kept when the instance cannot be reached.
- shape_ledger.unjudged_shapes: no row, unclassified, scoped and hook stamps
  are unjudged; an agent/audit/import verdict is not. A snippet recorded at
  the shape answers for it until the refresh stamps it canonical.
- services/shape_check owns the reason text and records every outcome in
  app_logs (passed / blocked / judged_after_block / left_after_block).
- classify_shapes(repo=…) judges a shape the ledger has not synced yet via a
  provisional row under a bound repo; the sync confirms it, or vanishes and
  revives it with the verdict intact. An unbound repo is refused.

Plugin version minted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 08:39:36 -04:00
bvandeusenandClaude Opus 5.5 eb5cc6d3a7 fix(shapes): weak stamps elect no canon; the arrival line names the review; Svelte scopes by default (#4608)
CI & Build / Python lint (push) Successful in 8s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m3s
CI & Build / Python tests (push) Successful in 1m44s
CI & Build / Build & push image (push) Successful in 1m5s
Librarian and Stash had 132 and 488 write-path stamps written at 0.68-0.72,
before the 0.80 floor (#4204). Nothing re-judged them, and dominant_canon
counted them, so a YAML CI snippet (#3410) "dominated" Librarian's web
directory and every write there was told it diverged from it — 1408 flags
on one project, 902 on the other, all skimmed past.

- is_weak_stamp: one predicate for "hook-stamped below today's floor".
  dominant_canon and canon_form skip such rows; stamps_to_review lists them
  by the same predicate. No stored row changes — they stay for judgment.
- flag_divergence withdraws a standing flag whose canon no longer dominates
  its directory. A flag is a mechanical prompt, not a judgment.
- The coverage line ends "N weak stamps · M incoherent canons to judge —
  stamps_to_review" when either is nonzero, so a project's own session sees
  the queue on arrival. The shape-accounting skill says what to do with it.
- scoped_definitions handles .svelte: every <style> is component-scoped
  unless <style global> or :global(...); instance-script syms are scoped,
  module-script ones stay ordinary.

Plugin version minted for the skill change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 23:27:07 -04:00
bvandeusenandClaude Opus 5.5 3270fe90c1 refactor(settings): one bounded_float for every numeric bar read from settings
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / integration (push) Successful in 52s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m36s
CI & Build / Build & push image (push) Successful in 32s
The shape ledger's divergence readout flagged _gate_setting (#4385).
Reading it turned up a family with no canon: floor_for,
get_duplicate_threshold, get_plan_match_threshold and _gate_setting each
parsed a stored string, fell back to the default (never 0) and clamped
into [lo, 1].

- services/settings.bounded_float(raw, default, lo=0, hi=1) is the pure
  parse, fallback and clamp. Each caller keeps its own get_setting read,
  so tests patching get_setting per module still take effect, and
  _gate_setting still fails open on an unreadable setting.
- SettingsView.saveKbInject: eleven inline Math.min/Math.max clamps and
  the local gateAt become one asBar(v, d, lo = 0), mirroring the server.

No behaviour change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 12:30:58 -04:00
bvandeusen 33c0f2c95d Merge pull request 'Server instructions orient the workflow; using-scribe splits into reference files (#4389, #4398)' (#187) from dev into main
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m46s
CI & Build / Build & push image (push) Successful in 14s
2026-09-24 11:30:02 -04:00
bvandeusenandClaude Opus 5.5 d5dad587f1 refactor(skills): using-scribe keeps the every-turn practices; moment-specific depth moves to reference files (#4398)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m8s
CI & Build / Python tests (push) Successful in 1m44s
CI & Build / Build & push image (push) Successful in 21s
Anthropic's skill guidance: keep SKILL.md under 500 lines, split into
reference files linked one level deep as it nears that. using-scribe was
478 and every new practice lands there.

- SKILL.md 478 -> 317 lines. It keeps orientation, one copy, the reflexes,
  scope, the judge section, UI and the process-skill index, plus a "Read
  these when the moment comes" list naming each file with its moment.
- projects.md: binding a non-git directory (.scribe) and project inception.
- writing-records.md: where a new rule goes, lesson growth, and notes that
  carry their own check (reflex 10 keeps a pointer).
- missed-retrieval.md: the record-before-dial route, verbatim.
- Text moved, not rewritten, except for the seams and one cross-reference.

Tests:
- tests.helpers.skill_text reads SKILL.md plus its reference files. The
  ownership registry, the miss-route and the verification tests use it, so
  a topic stays owned by its skill whichever file holds it.
- The force test scans every skill .md on its own, since each file is read
  on its own.
- New test_skill_structure: SKILL.md <= 350 lines, every reference file is
  linked from SKILL.md, none links another, and one over 100 lines opens
  with Contents. Each guard is shown to fail.

The plugin version is minted. That also clears 4fb53b8's red Plugin hooks
lane, which failed only because PACKAGING.md changed without a mint.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 09:58:41 -04:00
bvandeusenandClaude Opus 5.5 4fb53b844d refactor(mcp): _INSTRUCTIONS orients the workflow, not a rulebook (#4389)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Failing after 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m35s
CI & Build / Build & push image (push) Successful in 23s
Spike #4389 read the spec, Claude's docs and a dozen servers: the field is
for how the tools fit together, and the field runs ~600-1,600 characters.
Ours sat at the 2,048 cap as a keyword index that also carried stance.

- JUDGE, REPORT and MISSED leave the index. They fire mid-work, not at
  session start; using-scribe and reporting-back state them in full, and
  `placement`/`report_back` cue reporting in-band. No skill text changes.
- The rest is rewritten as plain practices (1,503 chars) and keeps every
  session-start marker the ownership registry pins.
- INSTRUCTIONS_BUDGET 2000 -> 1600; the three index markers are dropped
  from the registry; the miss-route index test now checks that the index
  keeps what_might_apply and stays off the route.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 07:50:34 -04:00
bvandeusen 779efb2e74 Merge pull request 'dev → main: MCP SDK v2 port, create-gate thresholds in Settings' (#186) from dev into main
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 52s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m38s
CI & Build / Build & push image (push) Successful in 16s
2026-09-24 07:31:51 -04:00
bvandeusenandClaude Opus 5.5 c26b7f248e feat(mcp): port the server to MCP Python SDK v2 (FastMCP → MCPServer) (#2196)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m37s
CI & Build / Build & push image (push) Successful in 31s
- server.py: `mcp.server.mcpserver.MCPServer`; StrictArgsFastMCP becomes
  StrictArgsMCPServer, whose call_tool takes and forwards v2's `context`
  and raises ToolError (the SDK logs anything else as an unexpected crash;
  the message reaches the caller either way).
- stateless_http and transport_security moved from the constructor to
  `streamable_http_app(...)` in mount_mcp, with their reasons.
- Per-request user identity is unchanged: the contextvar set around the
  ASGI call reaches the handler on both v2 paths (legacy stateless spawns
  from the request task; the 2026-07-28 modern path opens its task group
  inside the request).
- pyproject: mcp[cli]>=2.2, no ceiling (installs are --locked). uv.lock
  regenerated with --upgrade-package mcp in the ci-python image: only mcp
  and its own dependencies moved (106 → 110 packages).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 07:21:27 -04:00
bvandeusenandClaude Opus 5.5 4502f0a1ae feat(dedup): the create gate's similarity bars are settings (#4385, rule 25)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / integration (push) Successful in 56s
CI & Build / Python tests (push) Successful in 1m38s
CI & Build / Build & push image (push) Successful in 38s
gate_bars(user_id, note_type) resolves the block bar and, for notes and
tasks, the overlap floor from kb_gate_* settings, with the old constants
as defaults. Fail-open on an unreadable value; a block bar clamps at 0.80
and the overlap floor at 0.70 and never above the bar. Five fields in
Settings beside the duplicate-report floors.

CLAIM_LEASE stays a constant, with the reason written at it: a per-user
lease would make one shared task live to one reader and dead to another.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 07:11:10 -04:00
bvandeusen f385d9075b Merge pull request 'dev → main: task claims (milestone 381), embedding model stamp, dedup copy band, sweep-shared.css' (#185) from dev into main
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / integration (push) Successful in 51s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m33s
CI & Build / Build & push image (push) Successful in 13s
2026-09-24 06:57:12 -04:00
bvandeusenandClaude Opus 5.5 baf22179ef refactor(frontend): the sweep pane's column wrapper joins sweep-shared.css (#3207)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Build & push image (push) Successful in 34s
.sweep was the one rule still written identically in both sweep panes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 06:53:11 -04:00
bvandeusenandClaude Opus 5.5 c787957ddf fix(frontend): load sweep-shared.css once from main.ts — a shared <style src> breaks vite build (#3207)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Python tests (push) Successful in 1m36s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 52s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Build & push image (push) Canceled after 35s
plugin-vue throws "Cannot read properties of undefined (reading 'scoped')"
when several SFCs share one src stylesheet. vue-tsc passed, so only the
image job caught it; :dev has not rebuilt since 22b7a92.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 06:50:59 -04:00
bvandeusenandClaude Opus 5.5 4a93c8b262 fix(task-logs): a collaborator who may write a task may log on it (rule 78)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / integration (push) Successful in 52s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m37s
CI & Build / Build & push image (push) Failing after 24s
create_log filtered on Note.user_id == user_id, a bare owner check, so a
collaborator with write access to a shared task was told it did not exist
— and, since a log now stamps the claim, could never be seen working it.
It now asks can_write_note. Editing and deleting a log still require its
author, which is authorship rather than access.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 06:48:01 -04:00
bvandeusenandClaude Opus 5.5 22b7a928da refactor(frontend): derive the sweep-row visual language into sweep-shared.css (#3207)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 52s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m32s
CI & Build / Build & push image (push) Failing after 22s
The note sweep, rule sweep, preference drift and rule history panes each
restated the same row recipe in a scoped block. It now lives once in
assets/sweep-shared.css under `sweep-` names (unprefixed globals would
leak into the unrelated scoped .row-title/.lede/.age/.actions elsewhere),
imported unscoped beside each pane's own scoped remainder.

Moved only what was shared: RuleHistoryPanel takes .sweep-state and keeps
its button row head; PreferenceDrift keeps its tiny-type footnote and
overrides the action row's layout. Lede and state now use the body-sm
token instead of 0.85rem/0.9rem (13px vs 13.6/14.4px) — one value, from
the design system.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 06:45:43 -04:00
bvandeusenandClaude Opus 5.5 952e56ee75 feat(tasks): the hand-off — SessionEnd releases a session's claims; the practice is written down (milestone 381 step 4)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 1m40s
CI & Build / Build & push image (push) Successful in 24s
Release is the mechanical half: scribe_session_end.sh sends the ending
session's id to /api/plugin/release-session, which releases the claims it
held. A tidy-up, not the guarantee (no SessionEnd on a crash; the lease
covers that), and skipped on /clear so SessionStart(clear) can still
hand the claimed work back.

Saying what happened is the half only the model can do. It is stated as a
practice where it is read: the using-scribe skill owns it ("Hand off before
this session's context stops existing", pinned in test_guidance_ownership),
the static context points at it for the wrap-up moment, and add_task_log's
docstring says a log claims the task. _INSTRUCTIONS is untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 06:43:14 -04:00
bvandeusenandClaude Opus 5.5 e46eea3b52 feat(tasks): SessionStart reads the claim — a compaction gets its work back (milestone 381 step 3)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m35s
CI & Build / Build & push image (push) Canceled after 45s
The claim now pays rent to the session that set it. The SessionStart hook
sends the host's `source` and the session id; the server renders a claim
section by source (task_claims.render_claims):

- compact / clear: the work this session had claimed, each with its two
  latest log entries, so a compacted session resumes from the record
  rather than from a count of open tasks.
- startup / clear: other sessions' live claims (may still be running) and
  in-progress tasks whose claim went quiet (abandoned mid-task).
- fork: the same, framed as "two sessions may now hold this".
- resume, or no source sent: nothing.

The task sidebar shows a live claim as "Being worked" and a dead one on
open work as "Went quiet", so the operator sees a session die mid-task.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 06:41:19 -04:00
bvandeusenandClaude Opus 5.5 a585fe0be0 fix(backup): a task claim does not travel in a backup (milestone 381)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m38s
CI & Build / Build & push image (push) Successful in 29s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 19:11:48 -04:00
bvandeusenandClaude Opus 5.5 b06b3a1d8a feat(tasks): a task can say a session is working it — the claim (milestone 381 step 2)
CI & Build / Python lint (push) Successful in 3s
CI & Build / integration (push) Successful in 53s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Python tests (push) Failing after 1m18s
CI & Build / Build & push image (push) Skipped
CI & Build / Plugin hooks (push) Successful in 14s
status is durable and nothing clears it, so in_progress cannot also mean
"someone is on this now". The claim is that second meaning, stored as who
and when so it dies on read rather than needing to be cleared.

- notes.claimed_by / claimed_at / claim_touched_at / claim_session (0110).
- The server stamps it on the write that is the work: reaching in_progress
  (update or create, via apply_status_transition) and a work log on an open
  task. done/cancelled/todo release it. Live while touched within
  CLAIM_LEASE (2h); dead on read past it, with no sweep.
- The plugin's PostToolUse hook on update_task/add_task_log binds the
  harness's session_id (GET /api/plugin/claim-session). It binds only to a
  live claim the caller holds; it cannot create one.
- to_dict carries `claim`. Plugin version minted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 19:10:06 -04:00
bvandeusenandClaude Opus 5.5 b8f543f45a fix(tasks): a create that names a status is stamped like an update to it (#3683)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / integration (push) Successful in 55s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Build & push image (push) Successful in 27s
build_note (create_note and create_records) set the status and nothing it
implies, so create_task(status="in_progress") wrote a started task with no
started_at, and done/cancelled had no completed_at or next recurrence.
The transition is now one function, apply_status_transition, called by
both paths; tests pin the invariant for every status.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 19:04:27 -04:00
bvandeusenandClaude Opus 5.5 f4e9cd429b feat(embeddings): every vector records the model whose space it lives in (#4132)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m40s
CI & Build / Build & push image (push) Canceled after 27s
The four embedding tables stamped chunker_version but not the model, and
vector(384) is a width, not an identity: a same-width model swap would
write a second geometry beside the first with no error.

- embedding_model on note/rule/milestone/system embeddings (0109; existing
  rows stamped with the only model any install has ever run).
- Every write stamps EMBEDDING_MODEL; every backfill's "current" test is
  is_current_stamp(), both halves of calibration_stamp().
- migrate_floor refuses while any row its surface searches is off the live
  model, before sampling: re-embed, then migrate.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 19:02:30 -04:00
bvandeusenandClaude Opus 5.5 11b286d786 test(dedup): a note is gated at its copy band, not the general floor (#4306)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m41s
CI & Build / Build & push image (push) Successful in 30s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 19:00:03 -04:00
bvandeusenandClaude Opus 5.5 fc1c463641 feat(dedup): a note or task blocks only as a copy; a close match is surfaced for judgement (#4306)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / TypeScript typecheck (push) Successful in 58s
CI & Build / integration (push) Successful in 1m12s
CI & Build / Python tests (push) Failing after 1m29s
CI & Build / Build & push image (push) Skipped
Measured on the live corpus, the 74 note pairs at or above the old 0.90 bar
were almost all distinct siblings — consecutive dev-logs, sub-notes of one
design, research parts — and the one clear copy sat at 0.997. The block
refused the next dev-log and taught force=true, as #4134 found for rules.

- The semantic arm blocks notes and tasks only at >= 0.98. The title block
  stays; processes keep 0.90 (not measured).
- 0.87 to 0.98 comes back as `overlaps` on the create reply, from the same
  per-chunk searches, with a note that leaves the call to the session:
  fold in and delete if it is the same record, keep both if a sibling.
- create_note, create_task, create_records and start_planning's steps all
  carry it; a batch names the record each overlap belongs to.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 18:53:22 -04:00
bvandeusen eb6912c722 Merge pull request 'Snippets gain notes; when_to_use is the situation it is ranked on (#4378)' (#184) from dev into main
CI & Build / Python lint (push) Successful in 3s
CI & Build / integration (push) Successful in 54s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m43s
CI & Build / Build & push image (push) Successful in 17s
CI & Build / Plugin hooks (push) Successful in 12s
2026-09-23 18:00:35 -04:00
bvandeusenandClaude Opus 5.5 4ae18a9dd9 feat(snippets): a snippet has notes; when_to_use is the situation it is ranked on (#4378)
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / Python tests (push) Successful in 1m37s
CI & Build / Python lint (push) Successful in 3s
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Build & push image (push) Successful in 34s
A snippet had no field for prose, so what a session learned about one went
into when_to_use — the trigger joined onto every chunk it is embedded as. A
sweep found write-ups of up to 3 KB there, headings and all.

- notes: stored after the code under `## Notes`, parsed back from the body,
  carried by every path that rebuilds it (update, merge, un-merge). A
  snippet with no notes composes the body it always did.
- create/update_snippet (MCP) take notes and return trigger_advice when
  when_to_use is long, headed or multi-paragraph. Advice, not a refusal.
- Tool docs, the reusing-code skill and the editor hint describe the trigger
  as the situation and point the explanation at notes.
- Editor gains a Notes field; the detail view renders it as markdown.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 17:27:22 -04:00
bvandeusen 14734ebe46 Merge pull request 'dev → main: auto-inject reads the conversation, startup names its failure, menu lines carry metadata, titles are names' (#183) from dev into main
CI & Build / integration (push) Successful in 50s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Build & push image (push) Successful in 17s
2026-09-23 16:57:42 -04:00
bvandeusenandClaude Opus 5.5 66e21a6c60 refactor(notes): a snippet's and lesson's stored title is its name; the trigger joins it only in the embedded document (milestone 427)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / integration (push) Successful in 52s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m35s
CI & Build / Build & push image (push) Successful in 32s
The title was `subject — trigger` because the stored title WAS the
embedded one, and the join is what makes these kinds rank on the
situation they apply to (#2485). Every surface that shows a title then
showed the trigger too -- menus, lists and search rows ran to kilobytes.

- embeddings.document_title(title, note_type, data, body) joins the
  trigger from `data` (body fallback) at embed time. Idempotent: an
  un-migrated composed title comes out the same, never doubled. The
  embed path, the startup backfill and the dedup gate's semantic signal
  all use it, so the embedded text -- and every vector -- is unchanged.
- Writers store the subject: snippet create/update (service, REST, MCP)
  and lesson_document. Both compose_title helpers are removed.
- Readers: dedup takes `data`; the menus strip the embedded title from a
  passage; list rows project `when_to_use`, which SnippetListView reads.
- 0108 rewrites existing rows on an exact `' — ' || <own trigger>`
  suffix with raw SQL, leaving updated_at alone so the backfill does not
  re-embed the corpus for identical vectors. Downgrade recomposes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 16:48:28 -04:00
bvandeusenandClaude Opus 5.5 bb632c4196 feat(retrieval): a menu line is a name, its kind and System, and the whole passage that matched (#4364)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 58s
CI & Build / Python tests (push) Successful in 1m36s
CI & Build / Build & push image (push) Successful in 26s
Injected lines rendered a snippet's or lesson's title, which carries its
whole trigger by construction (the embedding shape) and ran past 1,500
characters -- again on every `seen` repeat. The passage under a line was
cut to 200 chars from the middle, keeping its head (the title again) and
losing where the match was.

Now, on both the prompt menu and the write-path prior-art menu:
- the line shows the record's NAME (snippet data.name / lesson subject),
  with its kind and System (`[issue (done) · Plugin & hooks]`);
- the passage is the whole matched chunk, title prefix stripped, on one
  line so the blockquote holds; a title-only match hands over the trigger;
- a `seen` record is a one-line pointer to what is already in context.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 16:29:48 -04:00
bvandeusenandClaude Opus 5.5 20227ebb5d fix(plugin): recent-context cap counts text, and an empty transcript is not a failure (#4364)
CI & Build / Python lint (push) Successful in 3s
CI & Build / integration (push) Successful in 45s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Build & push image (push) Successful in 51s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / Python tests (push) Successful in 1m35s
The trailing newline became a space before the 600-char cut, so the cap
kept 599 characters of text; and under pipefail a grep that matched no
reply made the helper exit 1. Trim before cutting; return 0 explicitly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 16:08:01 -04:00
bvandeusenandClaude Opus 5.5 eb00554976 fix(plugin): a SessionStart that loads no project says why, and names the first move (#4366)
CI & Build / Python tests (push) Canceled after 23s
CI & Build / integration (push) Canceled after 27s
CI & Build / TypeScript typecheck (push) Canceled after 28s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / Build & push image (push) Canceled after 0s
The context fetch folded a timeout, an HTTP error and a refused key into
one sentence that ended "enter_project() as needed" -- read as optional,
so a session could start with no recent milestones or open tasks and no
way to know what prior work existed. The fetch now names the cause and
the elapsed time, retries once (6s) only where a retry can change the
answer, and the fallback states enter_project as the first step, with
the marker's project id when there is one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 16:07:33 -04:00
bvandeusenandClaude Opus 5.5 574b27ae74 feat(plugin): auto-inject retrieves on the conversation, not the typed words alone (#4364)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / integration (push) Successful in 49s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Failing after 1m7s
CI & Build / Build & push image (push) Skipped
The prompt hook sent only the operator's message, so a mid-session
follow-up ("yes do that") named nothing a note, lesson or rule could
match. The hook now reads the tail of the last assistant reply from the
transcript and sends it as `ctx`; the notes and rule arms append it to a
short prompt (<= 280 chars), prompt first, capped at 600 chars. No model
tokens: it is embedding input, and the injected menu's budget is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 16:05:21 -04:00
bvandeusen bf2871ffb4 Merge pull request 'Shape ledger: a semantic miss is not evidence; floors sized for the reader (#4208, #4306)' (#182) from dev into main
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / Python lint (push) Successful in 2s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Build & push image (push) Successful in 18s
2026-09-22 09:50:33 -04:00
bvandeusenandClaude Opus 5 a01deeb851 fix(shapes): the signature basis admits the band true instances score in (#4306)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 57s
CI & Build / Python tests (push) Successful in 1m44s
CI & Build / Build & push image (push) Successful in 27s
Measured over 459 judged instance rows against every same-family canon:
true instances score 1.0 or ~0.795 (a model whose mixin list differs from
the canon's), and 0.8 cut the second band entirely. The proposal is read
before it is confirmed, so the floor admits it: 0.75 takes every measured
true row, at 35 wrong-canon pairings instead of 14.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
2026-09-22 08:44:58 -04:00
bvandeusenandClaude Opus 5 312dc9f6f7 fix(shapes): the semantic arm proposes at the write-path floor, not a private 0.8 (#4208)
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 56s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Build & push image (push) Successful in 22s
0.8 was sized for proposals nobody reads; every semantic proposal is read
by a judge before it is confirmed. Measured, it proposed 0 of 150 while
judged instances of a canon score 0.68-0.71 - the write-path hint's own
floor asks the same question of the same documents at 0.68. The arm now
uses that floor, scans 8 hits instead of 3 (a true instance ranked 4th
behind snippets of other language families), and the proposer version
bumps to 5 so rows examined under the old floor are read again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
2026-09-22 08:33:01 -04:00
bvandeusenandClaude Opus 5 a875a1b2ee fix(shapes): a semantic miss is not evidence — the divergence check stops reading it (#4208)
CI & Build / Python lint (push) Successful in 2s
CI & Build / integration (push) Successful in 53s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / Python tests (push) Successful in 1m32s
CI & Build / Build & push image (push) Successful in 24s
Measured live: true instances of the service-unit canon score 0.68-0.71
against it at best, helpers 0.66-0.75 against unrelated snippets, and
nothing reaches the 0.8 floor. The "conclusive miss" fired for nearly every
body and silenced real divergences exactly as it silenced helpers. The
stored miss basis, its flag withdrawal and the report plumbing are removed;
the floor stays at 0.8 and the new-shapes-first ordering stays. False
prompts are answered by judgment (exempt with a reason).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
2026-09-22 08:27:15 -04:00
bvandeusen 3530ed5e87 Merge pull request 'dev → main: the meaning check reads new shapes first, and its miss withdraws an early flag (#4208)' (#181) from dev into main
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Build & push image (push) Successful in 14s
2026-09-22 08:14:50 -04:00
bvandeusenandClaude Opus 5 e2197ac799 fix(tests): assert the withdrawn flag on its row, not on a count the skew margin inflates (#4208)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Build & push image (push) Successful in 14s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
2026-09-22 08:09:52 -04:00
bvandeusenandClaude Opus 5 19438cc894 fix(shapes): the capped meaning pass reads new shapes first, and its miss withdraws an early flag (#4208)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Failing after 51s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m33s
CI & Build / Build & push image (push) Successful in 21s
Measured on the first live refresh after 91cde6c deployed: the version bump
queued 1,533 rows for the semantic arm, the cap read 150 of them in row
order, and all six shapes new since the previous refresh - the only rows
flag_divergence acts on, and the highest ids - were flagged before the arm
reached them. The gate silenced nothing because it never got to look, and
since a flag persists until judged, a later conclusive miss could not take
it back.

- propose_for_repo sorts the semantic todo by _semantic_priority:
  unclassified before scoped, newest first.
- flag_divergence withdraws a standing flag when the row now carries a
  conclusive miss; that evidence alone withdraws one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
2026-09-22 08:05:49 -04:00
bvandeusen cc791de0e5 Merge pull request 'dev → main: rule overlap check, design-guidance write arm, usage chip seam, divergence meaning gate' (#180) from dev into main
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / Python lint (push) Successful in 2s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 51s
CI & Build / Python tests (push) Successful in 1m33s
CI & Build / Build & push image (push) Successful in 16s
2026-09-22 07:48:53 -04:00