Commit Graph
1645 Commits
Author SHA1 Message Date
bvandeusenandClaude Opus 5.5 17f4348711 ci: the image build waits on the integration lane (rule 177, #4958)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 56s
CI & Build / Python tests (push) Successful in 1m53s
CI & Build / Build & push image (push) Successful in 18s
The integration job was added after the build gate (c6211e5) and never
joined its needs, so main run 8212 published :latest with integration red.
Both jobs carry the same ref condition, so the gate cannot be satisfied by
a skipped integration run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 19:34:16 -04:00
bvandeusenandClaude Opus 5.5 983fd2c4d1 fix(retrieval): scope the rule search before ranking it - an in-scope rule is no longer lost behind other owners' nearer rules (#4958)
CI & Build / Python lint (push) Successful in 5s
CI & Build / Plugin hooks (push) Successful in 20s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m11s
CI & Build / Python tests (push) Successful in 2m0s
CI & Build / Build & push image (push) Successful in 29s
Ordered straight off rule_embeddings, the planner walked the HNSW index,
which returns about hnsw.ef_search (40) nearest chunks across every owner
and project and filters by home only afterwards. A reader's own rule ranked
past the 40th chunk overall was silently dropped - on a shared install,
other users' rules fill those 40. This is the likely cause of the
intermittent test_integration_rule_scope failures, whose axis-vector
fixtures sit far from every real embedding in the graph.

The in-scope chunks now go through a MATERIALIZED CTE, which the index
cannot order, so the ranking over them is exact. A rulebook is hundreds of
chunks; exact is cheap. New test: 80 nearer rules belonging to someone else
no longer hide the reader's one rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 19:30:05 -04:00
bvandeusenandClaude Opus 5.5 79cd2341a6 test(rules): the scope test says why a rule search returned nothing - an error or an empty scan (#4958)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / integration (push) Successful in 1m6s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 18s
semantic_search_rules fails open, so both failures on run 8209 and main
run 8212 read as set() with no trace. Assert report["searched"] and carry
the swallowed traceback into the failure.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 18:34:16 -04:00
bvandeusenandClaude Opus 5.5 4e1320120d feat(moments): a mount that keeps arriving where it does not apply proposes its own removal (milestone 458 step 7b, #4955)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m19s
CI & Build / Python tests (push) Successful in 2m3s
CI & Build / Build & push image (push) Successful in 18s
The open-after-moment signal proposes a mount; nothing proposed taking one
off, so a wrong mount was noise at every occurrence until someone happened
to notice. rule_misfired(rule_id, moment, why, reached_by) records a report
against a MOUNTED pair, counted per distinct day (the MCP door carries no
session id) on a new rule_moment_judgments.misfire column (migration 0119,
backup v24). At three days the response carries a line asking the agent to
offer the operator the fix - reject takes the rule off, unmap_action stops
the action reaching the moment, confirm keeps the mount and stops the
asking - and Settings > Moments lists it as an unmount proposal with the
reasons and the actions that reached it. A re-mount clears the count.

Taught in moments.md, missed-retrieval.md and the reply hold's wording.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 18:19:59 -04:00
bvandeusenandClaude Opus 5.5 8afd6da8af docs(plugin): reference files point at each other in words, not links - one level deep (milestone 458 step 8, #4926)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 56s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 2m11s
CI & Build / Build & push image (push) Successful in 41s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 16:44:22 -04:00
bvandeusenandClaude Opus 5.5 a72a422534 docs(plugin): the instruction surfaces teach moments - reading a line that arrived at one, correcting a misfire, and giving a new rule its moments (milestone 458 step 8, #4926)
CI & Build / Python lint (push) Successful in 13s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / TypeScript typecheck (push) Successful in 56s
CI & Build / integration (push) Successful in 1m38s
CI & Build / Python tests (push) Failing after 1m59s
CI & Build / Build & push image (push) Skipped
Until now only the tool arguments knew moments existed. The guidance
surfaces described rules as reached by resemblance alone:

- using-scribe: a short reflex paragraph and a new reference file,
  moments.md. It covers reading "at <moment>, reached by <action>", the
  reply held once at reply.report, map_action / unmap_action offered in
  one line, and the step 7 proposal line answered with judge_rule_moments.
- writing-records: asks WHEN a rule applies as well as what it is about.
  A rule, preference or process about a point in the work gets
  moments=[...] as it is written, and the trigger stays as the net.
- missed-retrieval: a missed WHEN is mounted or mapped, not reworded. A
  misfire is unmounted or unmapped.
- _INSTRUCTIONS: one clause (list_moments; mount rules about WHEN),
  1594 of 1600 chars.
- static context: injected lines include the rules mounted on a moment
  that was reached.
- test_guidance_ownership: three owned topics, so the text cannot quietly
  drop out.

Plugin minted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 16:39:47 -04:00
bvandeusenandClaude Opus 5.5 c404a3127a test(retrieval): the proposal endpoints are routed on the app (milestone 458 step 7, #4925)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 19s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m2s
CI & Build / Python tests (push) Successful in 2m17s
CI & Build / Build & push image (push) Successful in 56s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 16:16:59 -04:00
bvandeusenandClaude Opus 5.5 570d1b6d5a test(backup): register the rule_moment_judgments builder in the import column guard (milestone 458 step 7, #4925)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 56s
CI & Build / Python tests (push) Failing after 1m44s
CI & Build / Build & push image (push) Skipped
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 16:12:30 -04:00
bvandeusenandClaude Opus 5.5 dfcf4df2e9 feat(moments): mount the corpus by proposal - a pass and an open-after-moment signal, both stopping at the operator (milestone 458 step 7, #4925)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 19s
CI & Build / integration (push) Failing after 43s
CI & Build / Python tests (push) Failing after 46s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Build & push image (push) Skipped
A rule written before moments existed is mounted on nothing. Step 7 records,
per (rule, moment), whether it belongs there and who said so:

- rule_moment_judgments (migration 0118, backup v23): suggested / confirmed /
  rejected, from a pass, the signal, or an edit. Moment "" is "no moment fits".
- The pass: rules_to_mount lists unjudged rules; propose_rule_moments records
  suggestions that mount nothing; rule_moment_proposals and
  judge_rule_moments put them to the operator. A confirm mounts, a reject is
  kept so the pair is never proposed again. Same service behind REST and a
  "Waiting on you" panel in Settings > Moments.
- Edits are judgments: set_rule_moments, the one mount write path, confirms
  what was added and rejects what was removed in the same transaction.
- The signal: scribe_moment.sh keeps a per-session acts ledger; when a rule
  is opened, scribe_record_opened.sh sends the last three minutes of it to
  /api/plugin/rule-opened. The acts resolve through the install's mappings;
  work.run and work.change are not evidence. Counted per distinct session
  with lesson_rules' evidence model, and once due the open returns one line
  asking the reader to offer the mount.
- scribe_session_end.sh removes the session's scribe-moment files.

Plugin 2026.10.05.2003.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 16:03:53 -04:00
bvandeusenandClaude Opus 5.5 5bcc603310 fix(moments): the store loads rather than fetches, so the deadline guard does not read it as a bare fetch (milestone 458 step 6, #4924)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 52s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 2m2s
CI & Build / Build & push image (push) Successful in 45s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 14:47:12 -04:00
bvandeusenandClaude Opus 5.5 b73a689849 feat(moments): the human door onto mounts, mappings and per-moment telemetry (milestone 458 step 6, #4924)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 18s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Python tests (push) Failing after 1m34s
CI & Build / Build & push image (push) Skipped
CI & Build / integration (push) Successful in 2m7s
Everything an agent can do with moments, a person can now see and change
in the app.

- Rule editor: a moment picker beside the trigger. Catalog moments are
  ticked; a named procedure's `skill.<name>` is typed and checked as the
  server checks it. `moments` is always sent, so unticking the last moment
  unmounts the rule.
- Settings, Moments section (General tab): for each moment, what it
  means, the actions that reach it on this install (shipped ones can be
  switched off, the install's own removed), how many rules are mounted on
  it, and deliveries and agent opens over the window. Below that: named
  procedures with mounts, switched-off defaults with Restore, and a form to
  add an action.
- retrieval_telemetry.moment_usage: per moment, `delivered`, `rules`,
  `opened` (agent pulls after the first delivery there; an upper bound, as
  by_source is) and `last_delivered_at`. No ratio, because a mount is a
  person's statement, not a ranker's guess. Guarded on its own, and also
  reported in retrieval_summary as `moment_usage`.
- rulebooks.mount_counts; mounted_moments now derives from it.
- GET /api/retrieval/moments carries `mounted` and `usage` (?days=).
  DELETE /moments/mappings also reads the mapping from query parameters,
  since the browser's DELETE sends no body.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 14:42:30 -04:00
bvandeusenandClaude Opus 5.5 c489452a2d test(moments): an install removal still drops its tool while the skill loader stays listed (milestone 458 step 5, #4923)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 19s
CI & Build / TypeScript typecheck (push) Successful in 1m2s
CI & Build / integration (push) Successful in 1m20s
CI & Build / Python tests (push) Successful in 2m12s
CI & Build / Build & push image (push) Successful in 33s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 14:28:50 -04:00
bvandeusenandClaude Opus 5.5 f1fdc4a951 feat(moments): skills and stored processes declare the moments they are for (milestone 458 step 5, #4923)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m3s
CI & Build / Python tests (push) Failing after 1m25s
CI & Build / Build & push image (push) Skipped
Loading a procedure now also reaches the moment it is for. Loading the
reporting procedure is a report; loading the release procedure is a delivery.

- Bundled skills: each SKILL.md declares `metadata: moments:`. The same
  declaration ships as Skill defaults (BUNDLED_SKILL_MOMENTS), because the
  server never sees the plugin's files. test_skill_moments holds the two
  together and pins the plugin name that qualifies the skill.
- Stored processes: `moments` on create_process and update_process, stored
  in the note's data and returned by get_process. A `scribe-proc-<slug>`
  load resolves its process through the sync manifest at load time. The
  moments are not copied into the stub, which would go stale mid-session.
- reachable_tools lists the skill loader whenever anything is mounted, since
  a process's moments are known only when it loads.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 14:23:24 -04:00
bvandeusenandClaude Opus 5.5 b29689d4de fix(retrieval): declare the reply moment's surfacing site in FAN_OUT_SITES (milestone 458 step 4b, #4922)
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / Python tests (push) Successful in 1m50s
CI & Build / Build & push image (push) Successful in 40s
moment_delivery records the mounted half's surfacing under
source=rp.MOMENT_RULE_SOURCE. That is an attribute the registry's
extractor cannot resolve, so the site is now declared with the value it
emits, read from the pipeline's own constant.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 12:32:07 -04:00
bvandeusenandClaude Opus 5.5 c6cdfc2172 feat(moments): the reply moment holds a finished reply for one read (milestone 458 step 4b, #4922)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / integration (push) Successful in 1m1s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Python tests (push) Failing after 1m25s
CI & Build / Build & push image (push) Skipped
The reply is the one act no tool call marks, and it is where "let me
know if it works" gets said. A new Stop hook (scribe_reply_check.sh)
sends the finished reply to POST /api/plugin/reply-rules, which checks
it twice:

- mounted: every unopened RULE on reply.report, plus reply.ask when the
  reply asks a question. Deterministic.
- semantic: the reply's head and tail against every rule's trigger, on a
  new ranked surface, reply_rule. It is the backstop for whatever the
  earlier arms missed. Its floor is its stop bar (default 0.80, budget
  1), with its own Settings dials. The new stop_only stage records
  surfacing for the rule that holds and nothing else, because nothing
  else reached anyone.

Following the operator's ruling from 456 step 8, a rule that holds blocks
once, in the server's words. The hook blocks only on a reason it was
given, so an unreachable instance never stops a session, and it never
holds the rewrite. The ledger is the act checkpoint's own, so a rule
holds a session once across both doors and the per-session cap counts
both.

The turn reader moved from the report check into scribe_defs.sh
(scribe_turn_facts / scribe_turn_fact), so the two Stop hooks read a
turn the same way. The output was checked identical on a real transcript.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 12:30:07 -04:00
bvandeusenandClaude Opus 5.5 b3b616b20a fix(moments): the tools import the delivery module, and the tool-list cache leaves the swept directory (milestone 458 step 4a, #4922)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / integration (push) Successful in 49s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m48s
CI & Build / Build & push image (push) Successful in 32s
Two guards caught 7865313:

- test_mcp_tool_processes reads every coroutine in a tool module's
  namespace as a tool, and a name-imported attach_moment_rules looked
  like one. All seven modules now call moment_delivery.attach_moment_rules,
  and the parity guard accepts the attribute form.
- The session-ledger convention: a file in a swept directory must be a
  .ids ledger. The tool-list cache describes the install, not the
  context, so it moves to its own directory, scribe-moment, where a
  compaction does not sweep it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 12:24:30 -04:00
bvandeusenandClaude Opus 5.5 78653130d6 feat(moments): mounted rules arrive when their moment happens, through every door (milestone 458 step 4a, #4922)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 57s
CI & Build / Python tests (push) Failing after 1m19s
CI & Build / Build & push image (push) Skipped
A rule mounted on a moment now reaches the session when an act reaches
that moment, with no semantic match involved:

- run_moment_arm on the pipeline: a lookup, not a ranked search. Each
  line names the moment and the act that reached it ("at work.deliver,
  reached by `git push`"), so a misfire is visible where it lands and
  can be unmapped in-session. A repeat is cited, not quoted; fresh
  rules are recorded surfaced under source moment_rule with the moment
  in detail. No retrieval_logs row, as for the other lookups, so no
  latency is persisted for this arm.
- rule_scope: a rule's home clause, moved out of semantic_search_rules
  so the moment lookup scopes by the same one.
- rulebooks.rules_on_moments / mounted_moments.
- The plugin door: a catch-all PreToolUse hook (scribe_moment.sh). It
  keeps /moment-tools' answer on disk for five minutes, so a call to a
  tool that cannot reach a mounted rule sends nothing, and an install
  that has mounted nothing sends one request per window. It shares the
  rules ledger with the other arms and fails open silently.
- The MCP door: Scribe's own tools named by the shipped mappings carry
  moment_rules in their response, so a client without the plugin gets
  them too. The hook skips those tools. A guard pins the attach on
  every one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 12:21:45 -04:00
bvandeusenandClaude Opus 5.5 cc26054437 feat(moments): rules mount on moments, through every rule door (milestone 458 step 3, #4921)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m0s
CI & Build / Python tests (push) Successful in 1m51s
CI & Build / Build & push image (push) Successful in 29s
rule_moments (migration 0117) records which moments a rule arrives at, by
catalog name, cascading with the rule. rule_detail, the one seam every
rule door already returns through, gains moments beside system_ids: None
leaves the mounts alone, a list replaces them. get_rule and both
list_rules doors read them back, batched per page.

All five MCP rule/preference writes and the three REST ones take moments
and validate them before their create or update. An unknown name is
refused with the catalog listed and leaves no half-made rule behind; a
parity test pins that ordering on every door.

Backup v22 carries the mounts as a join table remapped through the rule
map; a real-Postgres round trip checks they land on the restored rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 11:03:51 -04:00
bvandeusenandClaude Opus 5.5 cf3de5bae1 feat(moments): actions map onto moments, with in-session corrections (milestone 458 step 2, #4920)
CI & Build / Plugin hooks (push) Successful in 18s
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m23s
CI & Build / Python tests (push) Successful in 2m0s
CI & Build / Build & push image (push) Successful in 36s
moment_actions.resolve(tool, input) names every moment a call reaches and
the action that reached it. One call can reach several: kubectl apply is
a run, a deliver and a reach outside the workspace. Command tools match by
how each segment of the line starts, with a word boundary; other tools by
field=value arguments. The MCP server prefix and case are ignored.

56 shipped defaults cover the harness tools, Scribe tools and common
command shapes. moment_mappings (migration 0116) holds what an install
adds and the defaults it switches off. A removal is a stored row, so an
upgrade does not switch the default back on.

Per the operator ruling, corrections happen in the session: map_action
and unmap_action (write tools) return now_reaches so the fix can be
confirmed in the same reply. list_moments now shows each moment's
actions on this install. REST mirrors both doors, recorded as human.
Backup v21 carries the mappings.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 10:58:00 -04:00
bvandeusenandClaude Opus 5.5 2ff7f2f34f feat(moments): the moment catalog rules will mount on, readable in-session (milestone 458 step 1, #4919)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 58s
CI & Build / Python tests (push) Successful in 1m53s
CI & Build / Build & push image (push) Successful in 32s
Fourteen generic moments of work (session.start, work.start … reply.ask)
plus the skill.<name> family, each with what it means and the kinds of
action that reach it, written for any kind of work rather than software
alone. The catalog is code because every install needs the same mount
points; which actions reach a moment is per-install data (step 2).

require_moment refuses an unknown name with the catalog listed, so a
typo cannot become a mount that never fires. list_moments (read-only)
and GET /api/retrieval/moments hand out the same catalog.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 10:50:45 -04:00
bvandeusenandClaude Opus 5.5 0720ab6dcf refactor(retrieval): the completion-report preferences run through the pipeline (milestone 456 step 3, #4905)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Build & push image (push) Successful in 28s
reply_preferences.completion_preferences was the last hand-written
rule search + record_retrieval + record_rule_surfaced triple outside the
pipeline. It is now REPORT_PREFERENCE, a RuleArm with kind="preference"
over COMPLETION_QUERY, read back as records rather than as lines.

- RuleArm gains `kind`. _ranked asks the ranker for that kind and drops
  anything else on the way out, before the row is logged.
- RuleResult gains `shown`, the (score, rule) pairs behind the lines, for
  callers that return records.
- RuleMoment.project_id may be None: logged as given, searched as
  `project_id or None`. report_preference rows keep their NULL project.
- Guards: reply_preferences now has 0 direct rule searches. The registry
  constant-source example moves from report_preference to preference_slot,
  the pipeline slot that records through a module constant.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 08:38:05 -04:00
bvandeusenandClaude Opus 5.5 b72c9a92e2 test(retrieval): two source guards read the pipeline where the rule arms now live (milestone 456 step 2, #4904)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 49s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Build & push image (push) Successful in 30s
CI run 8137 had two failures, both source-scanning guards whose property
moved to the pipeline. The property itself still held:

- test_every_surface_name_is_a_real_telemetry_source looked for
  source="<name>" literals. The rule arms now record under their spec field,
  so the spec is the join key, read from RULE_ARMS.
- test_every_hook_rule_search_says_which_project_it_is_for counted 4 direct
  searches in plugin_context. It now expects 0 there, so a copied arm fails
  the count. It also asserts that the pipeline has exactly one search and
  that its keyword literal carries project_id and never everywhere.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 08:07:26 -04:00
bvandeusenandClaude Opus 5.5 f6b824b214 refactor(retrieval): the three rule arms run through one pipeline (milestone 456 step 2, #4904)
CI & Build / Plugin hooks (push) Successful in 19s
CI & Build / Python lint (push) Successful in 2s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / Python tests (push) Failing after 1m28s
CI & Build / Build & push image (push) Skipped
CI & Build / integration (push) Successful in 1m15s
prompt_rule, pre_tool_rule and the rule half of the write path each wrote
the same steps out by hand: search, band, split fresh from repeats, log the
call before any early return, reserve a slot, render, record surfacings.
The copies drifted, and #3497, #3750 and #3752 were fixed one copy at a time.

- New services/retrieval_pipeline.py. run_rule_arm runs the stages in one
  order. RuleArm is the spec (source, band, compact_tail, checkpoint,
  preference_slot). RuleMoment is the query and the session ledger.
- The I/O (ranker and recorders) is passed in as RuleIO. plugin_context
  resolves it from its own names at call time, so existing patch points
  still apply.
- _rule_band, _rule_hint_line, checkpoint_for, checkpoint_reason, the band
  constant and the preference slot moved into the pipeline unchanged.
  plugin_context re-exports them.
- A recorder that raises now costs only its row, never a rendered line.
- The flags reproduce today exactly. Whether prompt_rule should band, and
  whether the act arms should reserve a preference, are step 7.
- Registry: the pipeline call sites are declared in FAN_OUT_SITES, with
  values read from the specs.
- The #3497 structural guard now checks the one implementation, and that
  plugin_context writes no rule-source row of its own.

plugin_context.py: 3,558 -> 2,980 lines.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 08:01:23 -04:00
bvandeusenandClaude Opus 5.5 be62abf142 perf(retrieval): one embedding per query per plugin request (milestone 456 step 1, #4903)
CI & Build / Plugin hooks (push) Successful in 17s
CI & Build / Python lint (push) Successful in 2s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m25s
CI & Build / Python tests (push) Successful in 2m6s
CI & Build / Build & push image (push) Successful in 1m25s
A single plugin request fans one query out to several ranked arms. The
operator message is searched by auto_inject, its reuse and lesson slots,
prompt_rule, the preference slot and rule_via_lesson, and each arm called
get_embedding on its own: up to six model calls for one vector.

- embeddings.query_embedding_memo(): a request-scoped ContextVar memo that
  get_embedding consults. Outside a scope the default is None, so every
  other caller is unchanged. A failed embedding is not remembered.
- memoized_query_embeddings decorates /retrieve, /tool-rules and
  /prior-art.
- get_embeddings (document chunks) is untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 07:55:06 -04:00
bvandeusenandClaude Opus 5.5 d2ac7bf220 fix(lessons): a lesson's name is one claim on one line (#4797)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m48s
CI & Build / Build & push image (push) Successful in 28s
#4797 "A lesson's what takes a whole narrative, and the narrative becomes
its name". `what` is the title every listing and menu prints. Nothing
enforced its documented "one line", so lessons written with the incident
in `what` printed up to ~1,500 characters as a menu line, buried the
claim, and diluted the trigger in the embedded title. That was 14 of the
39 lessons on this install.

- lessons.require_claim refuses a `what` over WHAT_MAX_CHARS (240) or
  running over several lines. The refusal says the story goes in
  `insight`.
- Both doors call it before writing:
  - MCP create_lesson / update_lesson;
  - REST create / update.
  An update checks only a NEW name, so a lesson stored with a long one can
  still take the edit that repairs it.
- lessons.claim_line is the display half. _menu_name shows an over-long
  stored name as its first sentence marked " …", because a door guard
  does not undo rows already stored and an unrepaired install would keep
  printing them.
- The tool docstrings state the bound.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 23:04:53 -04:00
bvandeusenandClaude Opus 5.5 9909cd2450 fix(retrieval): the rules never-opened warning stops sending the reader to a review tool that refuses rules (#4798)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m7s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 26s
#4798 "The rules corpus's surfaced_never_pulled warning sends the reader to
menus_to_review, which can only review auto_inject". #4772 gave the warning
one remedy for both corpora: "judge a sample with menus_to_review". That is
right for notes, where a line carries its passage. For rules it is a dead
end: retrieval_review.REVIEWABLE holds only auto_inject, so the tool
refuses every rule arm.

- The reading is now per corpus (_NEVER_PULLED_READING). For rules, the
  text says:
  - the count includes rules that only arrived in a listing;
  - a rule can rightly be set aside on its trigger alone;
  - no judged sample exists for the rule arms;
  - a rule set aside again and again is a trigger to fix (update_rule
    when_to_apply, then what_might_apply), not a floor.
- The tool docstring says the same.
- The guard is tied to REVIEWABLE, so it can fail in both directions: if
  the rules text names menus_to_review while no rule arm is reviewable, or
  if a rule arm becomes reviewable and the text still says there is no
  sample.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 23:02:32 -04:00
bvandeusenandClaude Opus 5.5 f95ffae972 feat(retrieval): a record the operator names by number reaches the menu (#4796)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m47s
CI & Build / Build & push image (push) Successful in 37s
#4796 "A record the operator names by number reaches auto-inject only if
its wording happens to match". "yes go ahead with 4448" says which record
is meant, but a number means nothing to an embedding, so the prompt menu
filled with records resembling the words around it.

- record_refs.named_record_ids reads the operator's raw prompt, never the
  reply-enriched query:
  - `#N`, unless the word before it marks another numbering (PR, CI, rule,
    milestone, system, log...);
  - a bare number of 3 or more digits that opens the message, follows a
    reference word or continues a list one started;
  - never a quantity ("300 seconds"), a date, version, path or fenced code.
- Each id is resolved through the ACL check; trashed or inaccessible ids
  are dropped.
- A "Named in your message" block leads the menu: the kind, the System,
  the name and the opening of the body, or the seen pointer if the record
  is already on the ledger. It takes no share of top_k, and the semantic
  lines leave those ids out.
- The block is booked under the new `named_ref` source, registered as an
  unbidden lookup that is allowed to be quiet. A named id that also ranked
  counts as suppressed in the auto_inject row, so #3668's identity holds.
- Named records now arrive even when the search finds nothing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 22:10:14 -04:00
bvandeusenandClaude Opus 5.5 5c6d9a7b55 fix(mcp): a tool's refusal reaches the agent with its reason (#4794)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 1m52s
CI & Build / Build & push image (push) Successful in 27s
SDK 2.x passes on only a ToolError's text. Any other exception becomes
UnexpectedToolError("Error executing tool X") and its message stays on the
server. ValueError is how every Scribe tool refuses — what was refused, why,
what to do instead — so every refusal reached the agent bare, and it could
only retry blind. Seen live: tune_retrieval(actor="operator") and
judge_menu(verdicts=[]) both answered "Error executing tool …" and nothing
else.

StrictArgsMCPServer.call_tool re-raises an UnexpectedToolError caused by a
ValueError as a ToolError carrying the message, in the SDK's own
"Error executing tool X: <reason>" shape. Other exceptions are crashes and
stay masked. The stale comment claiming the SDK returns ValueError text is
corrected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 21:31:10 -04:00
bvandeusenandClaude Opus 5.5 df4b37673f fix(plugin): a subagent hand-back is not an operator prompt (#4785)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 54s
CI & Build / TypeScript typecheck (push) Successful in 56s
CI & Build / Python tests (push) Successful in 1m49s
CI & Build / Build & push image (push) Successful in 29s
`<agent-message …>` turns were scored by auto-inject and the prompt rule
arm as if the operator had typed them, spending a retrieval and writing a
retrieval_logs row that reads as a real message (log 54355, found by the
#4772 review). Same class as #4142. Matched up to the tag name, since the
tag carries attributes; `<agent-messages>` and the tag named mid-sentence
are still retrieved against.

Plugin 2026.10.04.0125.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 21:25:37 -04:00
bvandeusenandClaude Opus 5.5 78fa01b025 fix(embeddings): no chunk is a heading alone (#4784)
Judging live auto-inject menus (#4772) found "## Work log — 2026-08-21",
"## Findings worth carrying forward" and "Two dev→main PRs this session."
each embedded as a whole chunk. Text about nothing embeds close to
everything, so they took ranks 1–3 on vague queries and pushed real records
down.

Three ways the chunker made them, each closed:
- _split_paragraphs flushed a heading on its own when the paragraph under it
  was long. A thin lead is now carried into the paragraph that follows, and
  a hard cut never lands in the first quarter of the budget, where the
  newline it finds closes a heading.
- A heading with no body (above its ### parts) became a section of its own
  when the section before it was full. A thin section now takes the next.
- A short lead-in or a short last section stood alone. The lead-in joins
  what follows; a thin tail joins what precedes it.

A continuation piece now repeats the LAST heading before it, since one
section can now hold several. CHUNKER_VERSION 2 → 3, so the startup backfill
re-embeds; tuned dials will report their shape as changed, which is true.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 21:25:37 -04:00
bvandeusenandClaude Opus 5.5 5d0e97576b fix(retrieval): the review drops records written after the call it re-runs (#4773)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m8s
CI & Build / Python tests (push) Successful in 1m53s
CI & Build / Build & push image (push) Successful in 36s
The first live sample ranked records the call could never have been offered:
the session that made a call writes the decision, often quoting the message,
and that record tops the re-run. Judged, it inflates on_point exactly where
the budget is decided.

- _rerun over-fetches by POSTDATED_SLACK, drops records created after the
  call before ranking, and names them in `postdated`.
- each line carries `changed_since_call` (updated_at or a work log after
  the call) and `logged` (shown fresh then).
- judge_menu shares the re-run, so a post-dated record cannot be judged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 19:42:20 -04:00
bvandeusenandClaude Opus 5.5 6598c7fa85 feat(retrieval): a review pass judges whether injected lines related — menus_to_review, judge_menu and a judged readout (#4772)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 44s
An open rate cannot say whether a menu line related: every line carries its
matched passage (#4364), so "not opened" covers unrelated, enough as shown,
and already in context. #4772 "Injected notes are never judged".

- retrieval_judgments (0115): a reviewer verdict per line of a logged call,
  on_point / adjacent / unrelated, with its reason, rank, budget side and
  whether the agent opened it within the hour.
- menus_to_review re-runs a random sample of unjudged auto_inject calls with
  the arm's own parameters, past its budget, passage on every line.
  judge_menu records verdicts, re-deriving rank from a fresh re-run.
- retrieval_telemetry gains a judged block (by rank, within/beyond budget,
  on_point_unopened). surfaced_never_pulled stops blaming titles.
- missed-retrieval guidance names the review before a budget move.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 15:34:14 -04:00
bvandeusenandClaude Opus 5.5 0a1bb68808 feat(rulings): system_usage_events is read back — per-System counts and a telemetry block (#4769)
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Python tests (push) Successful in 1m44s
CI & Build / Build & push image (push) Successful in 36s
#4769 "Rulings are counted where someone will read them": milestone 444
step 4 wrote system_usage_events and nothing read it.

- retrieval_telemetry gains a `system_usage` block: surfacings and opens by
  source, distinct counts, and `by_system` naming the areas most shown.
  There is deliberately no pull-through ratio, because rulings travel in full
  in the line and opens are the exception.
- usage_for_systems (one GROUP BY) adds `usage` to the REST Systems list and
  detail, and to MCP get_system. MCP list_systems is unchanged.
- The Systems UI shows a "rulings shown N×" chip.
- rulings_pre_tool, rulings_write_path and mcp_get_system are now declared
  registry points; the registry guard covers their recorders.
- The Systems store merges a PATCH reply instead of replacing the row.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 08:49:20 -04:00
bvandeusenandClaude Opus 5.5 1b973ebd13 fix(write-path): a search hit's excerpt under snippet no longer 500s the prior-art route (#4768)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m5s
CI & Build / Python tests (push) Successful in 1m50s
CI & Build / Build & push image (push) Successful in 39s
#4768 "Write-path prior-art route 500s when a nameless search hit carries
its excerpt under `snippet`": _prior_art_line read item["snippet"] as a
snippet record's field dict, but a search hit carries its matched passage
there as a string. Read the name only from a dict.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 08:39:11 -04:00
bvandeusenandClaude Opus 5.5 891f375715 test(rulings): the hook/route parity guards pin the rulings params (#4757)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 49s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m49s
CI & Build / Build & push image (push) Successful in 33s
CI 7939 went red on test_the_hook_and_the_route_agree_on_every_parameter_name:
the Bash hook now sends root and cwd, and the guard pins the exact set. Both
guards (tool-rules and prior-art) now also pin seen_ruling_systems, spelled
once in scribe_rulings_query, and check the routes read all three.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 23:12:40 -04:00
bvandeusenandClaude Opus 5.5 556872c039 feat(rulings): a command or edit touching an area's files shows its rulings, once per session (milestone 444 step 4, #4757)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m6s
CI & Build / Python tests (push) Failing after 1m22s
CI & Build / Build & push image (push) Skipped
A System's rulings (the Rulings section of its description) now reach the
work by path, not by similarity. Both PreToolUse arms resolve the files a
command or edit names to the Systems whose path_patterns cover them, and the
first touch in a session shows each area's rulings in one line; a repeat is
a one-line reference. A lookup, so no floor, no budget, no retrieval_logs row.

- services/system_rulings: parse_rulings, command_paths (reads and writes,
  relative to the repo root from any cwd; flags, URLs, globs skipped),
  rulings_for_paths
- /tool-rules takes root, cwd and seen_ruling_systems; /prior-art takes
  seen_ruling_systems; both return ruling_system_ids
- hooks share <sid>.rulings.ids (cleared on compaction by the ledger naming
  convention); the Bash hook sends the repo root and cwd
- system_usage_events (migration 0114): surfacings by source, pulls from
  get_system; carried by backup (v20) through the system map
- writing-records: rulings also arrive when the area's files are touched

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 23:08:32 -04:00
bvandeusenandClaude Opus 5.5 a113c72b4f feat(systems): a System names its files — path patterns stored, validated and matched (milestone 444 step 3, #4756)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m6s
CI & Build / Python tests (push) Successful in 1m50s
CI & Build / Build & push image (push) Successful in 43s
A System gains path_patterns: globs relative to the repo root (* within one
directory, ** across any depth, a plain directory covering everything under
it). One service validates them for every door, so the web UI and MCP refuse
the same bad pattern with the same message. systems_for_paths resolves paths
to every active System that covers them, which step 4 (#4757) uses to deliver
an area's rulings when its files are touched.

- schema: systems.path_patterns JSONB NOT NULL default [] (migration 0113)
- service: normalize_path_patterns, path_matches, systems_for_paths
- routes + MCP create_system/update_system accept it; [] clears
- web UI: a Files field in the create and edit forms, patterns on the card
- backup carries it through export and restore
- using-scribe reflex 7: tagging work keeps a System's files current

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 22:55:26 -04:00
bvandeusenandClaude Opus 5.5 01d8a0b9f1 feat(guidance): an operator's ruling lives on the System it governs, and code is read as behaviour, not intent (milestone 444 steps 1-2, #4754 #4755)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m46s
CI & Build / Build & push image (push) Successful in 56s
A Librarian session contradicted a decision the operator had made 13 days
earlier. The ruling ("retry, then replace, never give up on a book") was
kept only as a quote in a work log, beside a session's reading of it that
capped replacements at 3. Three later sessions built on the reading, and one
carried the cap into an option as a "known cost", which the operator then
approved without being asked about it.

- writing-records.md: "A ruling goes on the System it governs". What a
  ruling is (the operator decided it; a later change could undo it), how it
  differs from a rule, and where it goes: a Rulings section at the end of
  the System description, one line each with who, when and the source
  record. Written the turn the operator decides; holds what is in force,
  not history; a charter line that contradicts a ruling is fixed in the
  same edit.
- using-scribe SKILL.md: reflex 1 says code tells you what a thing does,
  not what was wanted, and a limit read from code is unconfirmed until a
  System's Rulings says otherwise. Reflex 9 points to the ruling section.
- reporting-back: an option that carries existing behaviour says whose call
  it was (the operator's ruling, or a past session's never confirmed); one
  that contradicts a ruling is a Conflict.
- create_system / update_system docstrings: the Rulings section, and that
  description replaces the whole text.
- test_guidance_ownership: three topics pinned to their owners.
- Plugin version minted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 18:59:36 -04:00
bvandeusenandClaude Opus 5.5 cf4206469b feat(lessons): the web editor records which rule a lesson is an instance of (milestone 440, #4658)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 37s
The lesson editor now asks the question the create tool asks, while the
writer still has the situation in mind. It offers three answers:
- an instance of a rule, ticked from the rules the lesson resembles;
- "No rule fits", with a required reason;
- leave it open.

The rules are fetched before the save from a new GET
/api/lessons/rule-candidates. It runs the same rule_candidates search the
create door replies with, and passes "search unavailable" through as null,
apart from "nothing resembles it".

An edit sends an answer only when it changed. Re-sending the same rules
would re-stamp their judgments. Worse, a shared editor who cannot see the
owner's rule would reject it by sending a list without it. Leaving a linked
lesson open sends rule_ids=[], which is what clears the links.

"Leave it open" is not offered once "no rule fits" is on file: the server
has no way to take that answer back except by naming a rule. When "no rule
fits" completes a convergence group, the editor says so in a toast, since
the page it goes to does not recompute it.

Guards check that the payload uses the names both routes read, that the
three answers are offered, that the candidates route sits above the id route
and keeps null apart from [], and that the editor reuses .rule-chip.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 16:36:43 -04:00
bvandeusenandClaude Opus 5.5 75cefe60e4 feat(lessons): both records show the link in the web UI (milestone 440 step 7, #4635)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / integration (push) Successful in 1m9s
CI & Build / Python tests (push) Successful in 1m53s
CI & Build / Build & push image (push) Successful in 41s
The lesson page gains an "Instance of" panel that holds one of three answers.
Unjudged is the fall-through, so it is stated rather than left blank:
- the rule(s) it was judged an instance of;
- "No rule — <why>";
- "Not yet judged".

Suggested links show what they rest on (distinct situations, and projects
when more than one), with Confirm / Not an instance for a reader who can
write. Rejected links stay listed with their reason.

The rule slide-over lists the lessons that are instances of it, plus the
suggestions waiting on a judgment. Each entry links through to the other
record, and kind and state wear the existing .rule-chip.

The write check moves to utils/permission.ts. The copy on the snippet page
looked for "edit", which the server never sends, so shared editors saw a
read-only page (#4640 "The snippet page hid its edit controls from shared
editors"). Guards pin the client's unions and write levels to the
service's, the model's and access.py's own values.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 15:46:27 -04:00
bvandeusenandClaude Opus 5.5 6d3dca0af5 feat(lessons): convergence is named at the write — no-rule lessons that keep landing in one situation suggest a rule (milestone 440 step 5, #4634)
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / Python lint (push) Successful in 2s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Build & push image (push) Successful in 27s
When a lesson is answered "no rule fits" (create_lesson / update_lesson on
both doors), the response looks for other no-rule lessons it resembles and,
once there are CONVERGENCE_LESSONS (3) of them, carries `convergence`: the
members, their incidents and projects, and a hint to draft the missing rule
with create_rule (operator approval as always) and point each lesson at it —
or to leave them as lessons when no single choice is right every time.

- convergence_group is the pure bar: distinct LESSONS count, incidents never
  stand in for them (one broad lesson cannot trigger it), and a group whose
  sources all point at one incident is one event written up several times.
- convergence_for searches lessons by the new one's claim + trigger
  (trigger_title) at CONVERGENCE_THRESHOLD 0.65 — above the menu's "worth
  showing", below the duplicate gate's "same record" — then keeps the ones
  with a lesson_no_rule answer. Fail-open. No sweep, no timer (#4183).
- Defaults stated as defaults (rules 32, 115).
- Tests: the bar (pure), the search with stubs, the door, and the no-rule
  filter against Postgres; conftest stubs convergence_for for unit tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 15:31:15 -04:00
bvandeusenandClaude Opus 5.5 f34249a2d8 test(rules): locate the pre-tool arm by what it does, not by its name (#4633)
CI & Build / integration (push) Successful in 52s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / Python tests (push) Successful in 1m48s
CI & Build / Build & push image (push) Successful in 26s
CI run 712 failed test_neither_rule_arm_logs_its_call_behind_a_results_guard
with "min() iterable argument is empty": the guard walked the function named
build_tool_rule_hint, which since fd2ebf4 is a wrapper that searches nothing
— the arm's body moved to _tool_rule_hint. The property it guards (the call
is logged before any early return) still held; the guard had pinned a name.

It now finds the async function that calls semantic_search_rules under the
pre_tool_rule source, so a later rename moves the guard with the arm instead
of emptying it (rule 167).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 15:11:28 -04:00
bvandeusenandClaude Opus 5.5 fd2ebf4c13 feat(rules): a rule surfaces through the lessons confirmed as its instances (milestone 440 step 4, #4633)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 18s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 55s
CI & Build / Python tests (push) Failing after 1m16s
CI & Build / Build & push image (push) Skipped
On all three rule arms (prompt_rule, pre_tool_rule, write_path_rule), after
the direct match, a lesson matching the moment at the notes menu's own bar
brings the rule(s) it is CONFIRMED to be an instance of, rendered in rule
voice with "Reached through lesson #N “…”, a recorded instance of it."

- Confirmed links only (lesson_rules.confirmed_lessons /
  confirmed_rules_in_scope). A suggested link carrying its rule would
  manufacture the co-arrival #4637 counts and prove itself.
- Scope kept: a lesson surfaces everywhere, but the rule it brings must be
  global or this project's own — milestone 414's boundary, not reopened
  through a side door.
- Suppression is by RULE: anything the direct band named or the session
  ledger holds is skipped, whichever lesson reached it.
- Its own slot (VIA_LESSON_LIMIT = 1), not a rule slot — a stated default,
  since the plan's "decide by measurement" has nothing to measure until
  links are confirmed (#4632). Its own source, rule_via_lesson: registered,
  RANKED for pull-through, logged whenever it searches.
- Nothing searched at all while no lesson carries a confirmed link, which is
  every install until one is judged.
- build_prompt_rule_hint / build_tool_rule_hint are now thin wrappers over
  the direct arms (_prompt_rule_hint / _tool_rule_hint); the arm leaves its
  query in `_via_query` only when it ran, so a disabled arm or a blank prompt
  brings no rule in either, and the key never leaves the server.
- Via-lesson rules are not fed to the #4637 co-surfacing recorder: only
  direct matches are evidence.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 15:03:41 -04:00
bvandeusenandClaude Opus 5.5 dbbab859ce feat(lessons): soft links — a lesson and a rule arriving together in distinct situations are proposed as a link (milestone 440 step 3, #4637)
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Build & push image (push) Successful in 30s
When one hook response puts a lesson and a rule in front of the reader, the
pair is recorded as evidence on a SUGGESTED lesson_rule_links row; once the
pair has arrived together in PROPOSE_SITUATIONS (3) distinct situations, the
next co-arrival carries one line asking the reader to judge it with
judge_lesson_link. Nothing about surfacing changes: a suggested link carries
no rule anywhere (that is #4633, confirmed links only).

- lesson_rules: co_surfaced (fail-open; only pairs the reader could confirm —
  a lesson they may write, a rule they own; judged pairs gather nothing; one
  proposal per response; PROPOSE_COOLDOWN 6h between asks), plus the pure
  counting rules: situation_key, add_evidence, proposal_due, evidence_summary.
  A situation is the prompt on /retrieve (word tokens, sorted and
  de-duplicated, so trivial rewordings count once) and the FILE on
  /prior-art (every edit to one file is one situation).
- rules_for_lessons shows a suggested link's evidence counts.
- plugin_context: build_autoinject_hint returns lesson_ids,
  build_prompt_rule_hint returns shown_rule_ids, build_write_path_hint records
  its own pair; routes/plugin /retrieve records the prompt pair. Shown lines,
  repeats included: relevance makes a co-arrival, not the session ledger.
- Evidence lives in the existing evidence column — no migration, and backup
  already carries it.
- Tests: counting rules (unit), the recorder against Postgres (bar, repeat,
  judged pairs, ownership, cooldown, evidence kept on confirm); conftest stubs
  co_surfaced for unit tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 14:30:58 -04:00
bvandeusenandClaude Opus 5.5 95be2512ab fix(lessons): build the rule-candidate query with trigger_title (#4631)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 1m0s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m53s
CI & Build / Build & push image (push) Successful in 27s
CI run 707 failed test_the_trigger_separator_is_spelled_in_exactly_one_place:
rule_candidates joined the lesson's claim and trigger with an inline " — ",
a second spelling of embeddings.TRIGGER_SEP (#3207). It now calls
trigger_title — the same join a rule's own document title is built with,
which was the point of the query shape in the first place.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 13:34:08 -04:00
bvandeusenandClaude Opus 5.5 f7d8dc2e55 feat(lessons): judged when written — a new lesson is offered its rules, and "no rule fits" is an answer (milestone 440 step 2, #4631)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 53s
CI & Build / Python tests (push) Failing after 1m15s
CI & Build / Build & push image (push) Skipped
"Which rule is this lesson an instance of?" now has three recorded answers:
a rule named (a confirmed link, #4630), no rule fits (new), or unjudged.

- Model + migration 0112: lesson_no_rule (lesson_id PK, CASCADE from the
  note; why; judged_at). A table rather than a key in notes.data, because
  that mirror is re-composed from the body on every edit and would erase it.
- Service (lesson_rules): set_no_rule rejects any confirmed link with the
  reason; a confirmation (set_lesson_rules or judge_link) deletes the answer;
  require_one_answer refuses both answers in one call before any write;
  judgments_for_lessons + attach_lesson_rules add rule_judgment (and no_rule)
  to every lesson payload; list_unjudged lists the open ones; rule_candidates
  searches rules with the lesson's claim + trigger at the explicit-search bar,
  None when the search could not run.
- MCP: create_lesson/update_lesson take no_rule; an unanswered create returns
  rule_candidates, rule_judgment and a rule_hint; list_lessons(unjudged=true).
- REST: the same on POST/PATCH /api/lessons and GET ?unjudged=1; create
  returns rule_candidates.
- Backup v19: a lesson_no_rule section, export (full and per-user) and import.
- Guidance: create_lesson docstring, writing-records.md in using-scribe (owner,
  pinned in test_guidance_ownership), create_rule docstring on linking the
  lessons a new rule governs. Plugin version minted.
- Tests: door units, integration for the three states, the rejection reason,
  scoping, cascade; backup registries.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 13:25:27 -04:00
bvandeusenandClaude Opus 5.5 c8393975c3 fix(mcp): classify judge_lesson_link as a write tool (#4630)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m42s
CI & Build / Build & push image (push) Successful in 31s
The tool-classification guard (test_mcp_auth) failed CI run 705: the new
judge_lesson_link tool sat in no set, so a read key would have been silently
denied it. It changes a link's state, so it belongs in _WRITE_TOOLS.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 12:44:06 -04:00
bvandeusenandClaude Opus 5.5 41e4fbaba1 feat(lessons): a lesson names the rule it is an instance of — lesson_rule_links (milestone 440 step 1, #4630)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Failing after 1m13s
CI & Build / Build & push image (push) Skipped
The link between a lesson (one concrete situation) and the rule that governs
it, with the operator's soft-then-hard design built into its state:
suggested while evidence accumulates, confirmed or rejected once judged. Only
confirmed will carry a rule in retrieval (#4633); rejected is kept so the pair
is never proposed again.

- models/lesson_rule_link.py + migration 0111: one row per (lesson, rule),
  CASCADE on both ends, indexed both ways, CHECK on state (rule 36), evidence
  JSONB and judged_at.
- services/lesson_rules.py: require_rules (validated before any write, so
  a bad id leaves nothing half-linked), set_lesson_rules (set-semantics;
  a dropped rule becomes rejected, not forgotten), judge_link, and the two
  reads. ACL: write on the lesson (share-aware), ownership of the rule; a
  reader sees only rules they own. Decorations are fail-open (#4286).
- MCP: create_lesson / update_lesson take rule_ids; get/create/update return
  `rules`; new judge_lesson_link tool. REST: the same on /api/lessons plus
  PUT /api/lessons/<id>/rules/<rule_id>. Rules: rule_detail carries `lessons`.
- Backup v18: export (full and user-scoped, both ends in scope), builder,
  importer; both column guards register the table.
- Tests: integration (states, set-semantics, judge, ACL all-or-nothing,
  cascade both ways, CHECK, one row per pair); unit (door wiring, judge
  registered, migration/model state agreement, backup skip and unjudged
  stays unjudged). conftest stubs the decorations for unit tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 12:38:02 -04:00
bvandeusenandClaude Opus 5.5 2e4c2d9493 feat(usage): count what happened away from a record's own project — the readout #3735 needs
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m6s
CI & Build / Python tests (push) Successful in 1m50s
CI & Build / Build & push image (push) Successful in 38s
Milestone 385 step 8 (#3735 "a lesson is recalled on a project it was not
written on") is defined by opened-on-another-project. The reader's project has
been recorded on every usage event since 0c8e109, but nothing compared it with
the record's own, so the criterion was still unreadable.

- note_usage.usage_for_notes: a second aggregate in the same session joins
  notes and counts surfaced_away_count / pulled_away_count. Counted only where
  both projects are known and differ; ranked surfacings only (#2477). The
  first aggregate is untouched, so events on deleted notes still count.
- empty_usage carries both keys zero-filled; every door that attaches usage
  (list_lessons, get_lesson, snippets, knowledge) gets them through
  attach_usage.
- UsageBadge tooltip says "On other projects: surfaced N×, opened M×" when it
  happened, and nothing when it did not.
- Tests: the mocked split test feeds both aggregates; a unit test for the
  away counters; a real-Postgres test that home, unreported and ambient
  events are all left out.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 10:47:03 -04:00
bvandeusenandClaude Opus 5.5 582a5a4f48 feat(shapes): the practice is written where it is read, and the coverage line measures the slip (milestone 439 step 6)
CI & Build / integration (push) Successful in 47s
CI & Build / Python tests (push) Successful in 1m39s
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Build & push image (push) Successful in 29s
- reusing-code: "Before the turn ends — say what you built" — the four
  verdicts, the one classify_shapes(repo=…) call, and the component file as a
  shape. Description names the end-of-turn moment.
- shape-accounting: the writer judges; audits are the check that it held.
  The write path SUGGESTS (no more hook instances); component file rows and
  whole-file canon described; scoped covers Svelte too.
- _INSTRUCTIONS reuse line: "before the turn ends, say what you built
  (create_snippet the reusable, classify_shapes the rest)" — 1570/1600.
- Coverage line: "written-shape check (7d): N turns checked, M asked, K left
  unjudged", from the Stop hook's recorded outcomes; silent until the
  question has been put.
- test_guidance_ownership pins the new topic on reusing-code.
- prior-art hook header no longer says it stamps instance rows.

Plugin version minted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 08:46:58 -04:00