Commit Graph
1660 Commits
Author SHA1 Message Date
bvandeusen c08655914f Merge pull request 'Soft links: a lesson and a rule arriving together are proposed as a link (milestone 440 step 3, #4637)' (#190) from dev into main
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 56s
CI & Build / Python tests (push) Successful in 1m44s
CI & Build / Build & push image (push) Successful in 21s
2026-10-01 14:59:30 -04:00
bvandeusenandClaude Opus 5.5 dbbab859ce feat(lessons): soft links — a lesson and a rule arriving together in distinct situations are proposed as a link (milestone 440 step 3, #4637)
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Build & push image (push) Successful in 30s
When one hook response puts a lesson and a rule in front of the reader, the
pair is recorded as evidence on a SUGGESTED lesson_rule_links row; once the
pair has arrived together in PROPOSE_SITUATIONS (3) distinct situations, the
next co-arrival carries one line asking the reader to judge it with
judge_lesson_link. Nothing about surfacing changes: a suggested link carries
no rule anywhere (that is #4633, confirmed links only).

- lesson_rules: co_surfaced (fail-open; only pairs the reader could confirm —
  a lesson they may write, a rule they own; judged pairs gather nothing; one
  proposal per response; PROPOSE_COOLDOWN 6h between asks), plus the pure
  counting rules: situation_key, add_evidence, proposal_due, evidence_summary.
  A situation is the prompt on /retrieve (word tokens, sorted and
  de-duplicated, so trivial rewordings count once) and the FILE on
  /prior-art (every edit to one file is one situation).
- rules_for_lessons shows a suggested link's evidence counts.
- plugin_context: build_autoinject_hint returns lesson_ids,
  build_prompt_rule_hint returns shown_rule_ids, build_write_path_hint records
  its own pair; routes/plugin /retrieve records the prompt pair. Shown lines,
  repeats included: relevance makes a co-arrival, not the session ledger.
- Evidence lives in the existing evidence column — no migration, and backup
  already carries it.
- Tests: counting rules (unit), the recorder against Postgres (bar, repeat,
  judged pairs, ownership, cooldown, evidence kept on confirm); conftest stubs
  co_surfaced for unit tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 14:30:58 -04:00
bvandeusen 6ba40c9e10 Merge pull request 'Lessons point at rules (milestone 440 steps 1–2) + the away-from-home usage readout' (#189) from dev into main
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / TypeScript typecheck (push) Successful in 58s
CI & Build / integration (push) Successful in 57s
CI & Build / Python tests (push) Successful in 1m40s
CI & Build / Build & push image (push) Successful in 21s
2026-10-01 13:48:01 -04:00
bvandeusenandClaude Opus 5.5 95be2512ab fix(lessons): build the rule-candidate query with trigger_title (#4631)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 1m0s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m53s
CI & Build / Build & push image (push) Successful in 27s
CI run 707 failed test_the_trigger_separator_is_spelled_in_exactly_one_place:
rule_candidates joined the lesson's claim and trigger with an inline " — ",
a second spelling of embeddings.TRIGGER_SEP (#3207). It now calls
trigger_title — the same join a rule's own document title is built with,
which was the point of the query shape in the first place.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 13:34:08 -04:00
bvandeusenandClaude Opus 5.5 f7d8dc2e55 feat(lessons): judged when written — a new lesson is offered its rules, and "no rule fits" is an answer (milestone 440 step 2, #4631)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 53s
CI & Build / Python tests (push) Failing after 1m15s
CI & Build / Build & push image (push) Skipped
"Which rule is this lesson an instance of?" now has three recorded answers:
a rule named (a confirmed link, #4630), no rule fits (new), or unjudged.

- Model + migration 0112: lesson_no_rule (lesson_id PK, CASCADE from the
  note; why; judged_at). A table rather than a key in notes.data, because
  that mirror is re-composed from the body on every edit and would erase it.
- Service (lesson_rules): set_no_rule rejects any confirmed link with the
  reason; a confirmation (set_lesson_rules or judge_link) deletes the answer;
  require_one_answer refuses both answers in one call before any write;
  judgments_for_lessons + attach_lesson_rules add rule_judgment (and no_rule)
  to every lesson payload; list_unjudged lists the open ones; rule_candidates
  searches rules with the lesson's claim + trigger at the explicit-search bar,
  None when the search could not run.
- MCP: create_lesson/update_lesson take no_rule; an unanswered create returns
  rule_candidates, rule_judgment and a rule_hint; list_lessons(unjudged=true).
- REST: the same on POST/PATCH /api/lessons and GET ?unjudged=1; create
  returns rule_candidates.
- Backup v19: a lesson_no_rule section, export (full and per-user) and import.
- Guidance: create_lesson docstring, writing-records.md in using-scribe (owner,
  pinned in test_guidance_ownership), create_rule docstring on linking the
  lessons a new rule governs. Plugin version minted.
- Tests: door units, integration for the three states, the rejection reason,
  scoping, cascade; backup registries.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 13:25:27 -04:00
bvandeusenandClaude Opus 5.5 c8393975c3 fix(mcp): classify judge_lesson_link as a write tool (#4630)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m42s
CI & Build / Build & push image (push) Successful in 31s
The tool-classification guard (test_mcp_auth) failed CI run 705: the new
judge_lesson_link tool sat in no set, so a read key would have been silently
denied it. It changes a link's state, so it belongs in _WRITE_TOOLS.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 12:44:06 -04:00
bvandeusenandClaude Opus 5.5 41e4fbaba1 feat(lessons): a lesson names the rule it is an instance of — lesson_rule_links (milestone 440 step 1, #4630)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Failing after 1m13s
CI & Build / Build & push image (push) Skipped
The link between a lesson (one concrete situation) and the rule that governs
it, with the operator's soft-then-hard design built into its state:
suggested while evidence accumulates, confirmed or rejected once judged. Only
confirmed will carry a rule in retrieval (#4633); rejected is kept so the pair
is never proposed again.

- models/lesson_rule_link.py + migration 0111: one row per (lesson, rule),
  CASCADE on both ends, indexed both ways, CHECK on state (rule 36), evidence
  JSONB and judged_at.
- services/lesson_rules.py: require_rules (validated before any write, so
  a bad id leaves nothing half-linked), set_lesson_rules (set-semantics;
  a dropped rule becomes rejected, not forgotten), judge_link, and the two
  reads. ACL: write on the lesson (share-aware), ownership of the rule; a
  reader sees only rules they own. Decorations are fail-open (#4286).
- MCP: create_lesson / update_lesson take rule_ids; get/create/update return
  `rules`; new judge_lesson_link tool. REST: the same on /api/lessons plus
  PUT /api/lessons/<id>/rules/<rule_id>. Rules: rule_detail carries `lessons`.
- Backup v18: export (full and user-scoped, both ends in scope), builder,
  importer; both column guards register the table.
- Tests: integration (states, set-semantics, judge, ACL all-or-nothing,
  cascade both ways, CHECK, one row per pair); unit (door wiring, judge
  registered, migration/model state agreement, backup skip and unjudged
  stays unjudged). conftest stubs the decorations for unit tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 12:38:02 -04:00
bvandeusenandClaude Opus 5.5 2e4c2d9493 feat(usage): count what happened away from a record's own project — the readout #3735 needs
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m6s
CI & Build / Python tests (push) Successful in 1m50s
CI & Build / Build & push image (push) Successful in 38s
Milestone 385 step 8 (#3735 "a lesson is recalled on a project it was not
written on") is defined by opened-on-another-project. The reader's project has
been recorded on every usage event since 0c8e109, but nothing compared it with
the record's own, so the criterion was still unreadable.

- note_usage.usage_for_notes: a second aggregate in the same session joins
  notes and counts surfaced_away_count / pulled_away_count. Counted only where
  both projects are known and differ; ranked surfacings only (#2477). The
  first aggregate is untouched, so events on deleted notes still count.
- empty_usage carries both keys zero-filled; every door that attaches usage
  (list_lessons, get_lesson, snippets, knowledge) gets them through
  attach_usage.
- UsageBadge tooltip says "On other projects: surfaced N×, opened M×" when it
  happened, and nothing when it did not.
- Tests: the mocked split test feeds both aggregates; a unit test for the
  away counters; a real-Postgres test that home, unreported and ambient
  events are all left out.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 10:47:03 -04:00
bvandeusen d7641e57a8 Merge pull request 'Shapes are judged when they are written (milestone 439 steps 1–6) + #4608 fixes' (#188) from dev into main
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 56s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Build & push image (push) Successful in 17s
2026-10-01 08:55:10 -04:00
bvandeusenandClaude Opus 5.5 582a5a4f48 feat(shapes): the practice is written where it is read, and the coverage line measures the slip (milestone 439 step 6)
CI & Build / integration (push) Successful in 47s
CI & Build / Python tests (push) Successful in 1m39s
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Build & push image (push) Successful in 29s
- reusing-code: "Before the turn ends — say what you built" — the four
  verdicts, the one classify_shapes(repo=…) call, and the component file as a
  shape. Description names the end-of-turn moment.
- shape-accounting: the writer judges; audits are the check that it held.
  The write path SUGGESTS (no more hook instances); component file rows and
  whole-file canon described; scoped covers Svelte too.
- _INSTRUCTIONS reuse line: "before the turn ends, say what you built
  (create_snippet the reusable, classify_shapes the rest)" — 1570/1600.
- Coverage line: "written-shape check (7d): N turns checked, M asked, K left
  unjudged", from the Stop hook's recorded outcomes; silent until the
  question has been put.
- test_guidance_ownership pins the new topic on reusing-code.
- prior-art hook header no longer says it stamps instance rows.

Plugin version minted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 08:46:58 -04:00
bvandeusenandClaude Opus 5.5 8b1567cf2c feat(shapes): a component file is a candidate shape in its own right (milestone 439 step 5)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 54s
CI & Build / Build & push image (push) Canceled after 0s
CI & Build / Python tests (push) Canceled after 1m42s
A single-file component defines `.card`, `.title` and a `Props`, never
anything named after itself — so the ledger could not hold "StatusChip is
canon" and could not say a new card was built where one already existed.

- coverage.is_file_unit: a file that RENDERS (its markup names classes, or it
  opens a <script>/<template>/<style> block) and defines nothing named after
  its stem gets one `file` row. Structural, not a framework list: Svelte and
  Vue components qualify, a TSX component already has its function's row, a
  module is accounted for by its definitions. Emitting every file would have
  put every module of every project in the todo at once.
- No body fingerprint on file rows, so an edit never re-asks a judgment.
- `file` is its own form and family; derive grouping skips it (`index`,
  `+page` repeat by convention). Divergence buckets it on its own, so a new
  component where a component canon dominates is a fair question.
- mark_canonicals: a snippet recorded at a path with no symbol makes that
  file row canonical; it still covers no definition inside.
- The write hooks apply the same test before noting a new file for the
  end-of-turn question.

Plugin version minted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 08:45:14 -04:00
bvandeusenandClaude Opus 5.5 4cf1c6f042 feat(shapes): the write-path hook suggests and no longer stamps (milestone 439 step 4)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m49s
CI & Build / Build & push image (push) Successful in 27s
The hook used to land its evidence as `instance` rows, classified_by=hook —
a permanent verdict nobody read, and the source of the weak stamps that
poisoned two ledgers (#4608). Now the agent that wrote the code judges it at
the end of the turn, and the hook's evidence is what it is shown.

- shape_ledger.suggest_write_path_instances (was stamp_write_path_instances):
  same evidence and form gate, but it writes a PROPOSAL — proposed_snippet_id,
  basis "reference" (named) or "semantic" (resembles), score — and never a
  status. A judged row is left alone; a new shape gets an unclassified
  provisional row so the suggestion reaches the end-of-turn question as
  "looks like #N". By-name uses edges stay: a fact, not a verdict.
- The prior-art line says "looks like #N … when you judge this turn's
  shapes, say whether it is", and the result key is `suggested`.
- write_time_divergence takes `suggested`; _RESEMBLE_MIN re-documented as the
  suggestion bar and the weak-stamp line.
- reusing-code skill no longer says the pull stamps an instance.
- Tests moved to the new contract, plus a guard that nothing the hook does
  writes classified_by="hook".

Plugin version minted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 08:42:10 -04:00
bvandeusenandClaude Opus 5.5 f1fbdf746a feat(shapes): the agent judges what it wrote, at the end of the turn (milestone 439 steps 1-3)
CI & Build / Python lint (push) Successful in 5s
CI & Build / Plugin hooks (push) Successful in 18s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m38s
CI & Build / Build & push image (push) Successful in 33s
Recording used to be decided by machinery — the only "record it" prompt
fired when a same-named copy already existed (#2664), so a first instance of
a reusable piece was never asked about, and judgment arrived only through
audits. Now the question is asked where the knowledge is: the end of the
turn that wrote the code, of the agent that wrote it.

- Write hooks keep `<sid>.written.ids` (path, kind, name) for every
  definition a write names; a new file adds a `file` line for its stem — a
  candidate in any language without a framework rule (scribe_written_append).
- Stop hook scribe_shape_check.sh sends the ledger to GET
  /api/plugin/shape-check and blocks once, in the server's words, when
  anything is unjudged. Same discipline as the report check: never twice,
  never without a recorded check, another hook's loop left alone; the ledger
  is kept when the instance cannot be reached.
- shape_ledger.unjudged_shapes: no row, unclassified, scoped and hook stamps
  are unjudged; an agent/audit/import verdict is not. A snippet recorded at
  the shape answers for it until the refresh stamps it canonical.
- services/shape_check owns the reason text and records every outcome in
  app_logs (passed / blocked / judged_after_block / left_after_block).
- classify_shapes(repo=…) judges a shape the ledger has not synced yet via a
  provisional row under a bound repo; the sync confirms it, or vanishes and
  revives it with the verdict intact. An unbound repo is refused.

Plugin version minted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 08:39:36 -04:00
bvandeusenandClaude Opus 5.5 eb5cc6d3a7 fix(shapes): weak stamps elect no canon; the arrival line names the review; Svelte scopes by default (#4608)
CI & Build / Python lint (push) Successful in 8s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m3s
CI & Build / Python tests (push) Successful in 1m44s
CI & Build / Build & push image (push) Successful in 1m5s
Librarian and Stash had 132 and 488 write-path stamps written at 0.68-0.72,
before the 0.80 floor (#4204). Nothing re-judged them, and dominant_canon
counted them, so a YAML CI snippet (#3410) "dominated" Librarian's web
directory and every write there was told it diverged from it — 1408 flags
on one project, 902 on the other, all skimmed past.

- is_weak_stamp: one predicate for "hook-stamped below today's floor".
  dominant_canon and canon_form skip such rows; stamps_to_review lists them
  by the same predicate. No stored row changes — they stay for judgment.
- flag_divergence withdraws a standing flag whose canon no longer dominates
  its directory. A flag is a mechanical prompt, not a judgment.
- The coverage line ends "N weak stamps · M incoherent canons to judge —
  stamps_to_review" when either is nonzero, so a project's own session sees
  the queue on arrival. The shape-accounting skill says what to do with it.
- scoped_definitions handles .svelte: every <style> is component-scoped
  unless <style global> or :global(...); instance-script syms are scoped,
  module-script ones stay ordinary.

Plugin version minted for the skill change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 23:27:07 -04:00
bvandeusenandClaude Opus 5.5 3270fe90c1 refactor(settings): one bounded_float for every numeric bar read from settings
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / integration (push) Successful in 52s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m36s
CI & Build / Build & push image (push) Successful in 32s
The shape ledger's divergence readout flagged _gate_setting (#4385).
Reading it turned up a family with no canon: floor_for,
get_duplicate_threshold, get_plan_match_threshold and _gate_setting each
parsed a stored string, fell back to the default (never 0) and clamped
into [lo, 1].

- services/settings.bounded_float(raw, default, lo=0, hi=1) is the pure
  parse, fallback and clamp. Each caller keeps its own get_setting read,
  so tests patching get_setting per module still take effect, and
  _gate_setting still fails open on an unreadable setting.
- SettingsView.saveKbInject: eleven inline Math.min/Math.max clamps and
  the local gateAt become one asBar(v, d, lo = 0), mirroring the server.

No behaviour change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 12:30:58 -04:00
bvandeusen 33c0f2c95d Merge pull request 'Server instructions orient the workflow; using-scribe splits into reference files (#4389, #4398)' (#187) from dev into main
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m46s
CI & Build / Build & push image (push) Successful in 14s
2026-09-24 11:30:02 -04:00
bvandeusenandClaude Opus 5.5 d5dad587f1 refactor(skills): using-scribe keeps the every-turn practices; moment-specific depth moves to reference files (#4398)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m8s
CI & Build / Python tests (push) Successful in 1m44s
CI & Build / Build & push image (push) Successful in 21s
Anthropic's skill guidance: keep SKILL.md under 500 lines, split into
reference files linked one level deep as it nears that. using-scribe was
478 and every new practice lands there.

- SKILL.md 478 -> 317 lines. It keeps orientation, one copy, the reflexes,
  scope, the judge section, UI and the process-skill index, plus a "Read
  these when the moment comes" list naming each file with its moment.
- projects.md: binding a non-git directory (.scribe) and project inception.
- writing-records.md: where a new rule goes, lesson growth, and notes that
  carry their own check (reflex 10 keeps a pointer).
- missed-retrieval.md: the record-before-dial route, verbatim.
- Text moved, not rewritten, except for the seams and one cross-reference.

Tests:
- tests.helpers.skill_text reads SKILL.md plus its reference files. The
  ownership registry, the miss-route and the verification tests use it, so
  a topic stays owned by its skill whichever file holds it.
- The force test scans every skill .md on its own, since each file is read
  on its own.
- New test_skill_structure: SKILL.md <= 350 lines, every reference file is
  linked from SKILL.md, none links another, and one over 100 lines opens
  with Contents. Each guard is shown to fail.

The plugin version is minted. That also clears 4fb53b8's red Plugin hooks
lane, which failed only because PACKAGING.md changed without a mint.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 09:58:41 -04:00
bvandeusenandClaude Opus 5.5 4fb53b844d refactor(mcp): _INSTRUCTIONS orients the workflow, not a rulebook (#4389)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Failing after 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m35s
CI & Build / Build & push image (push) Successful in 23s
Spike #4389 read the spec, Claude's docs and a dozen servers: the field is
for how the tools fit together, and the field runs ~600-1,600 characters.
Ours sat at the 2,048 cap as a keyword index that also carried stance.

- JUDGE, REPORT and MISSED leave the index. They fire mid-work, not at
  session start; using-scribe and reporting-back state them in full, and
  `placement`/`report_back` cue reporting in-band. No skill text changes.
- The rest is rewritten as plain practices (1,503 chars) and keeps every
  session-start marker the ownership registry pins.
- INSTRUCTIONS_BUDGET 2000 -> 1600; the three index markers are dropped
  from the registry; the miss-route index test now checks that the index
  keeps what_might_apply and stays off the route.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 07:50:34 -04:00
bvandeusen 779efb2e74 Merge pull request 'dev → main: MCP SDK v2 port, create-gate thresholds in Settings' (#186) from dev into main
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 52s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m38s
CI & Build / Build & push image (push) Successful in 16s
2026-09-24 07:31:51 -04:00
bvandeusenandClaude Opus 5.5 c26b7f248e feat(mcp): port the server to MCP Python SDK v2 (FastMCP → MCPServer) (#2196)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m37s
CI & Build / Build & push image (push) Successful in 31s
- server.py: `mcp.server.mcpserver.MCPServer`; StrictArgsFastMCP becomes
  StrictArgsMCPServer, whose call_tool takes and forwards v2's `context`
  and raises ToolError (the SDK logs anything else as an unexpected crash;
  the message reaches the caller either way).
- stateless_http and transport_security moved from the constructor to
  `streamable_http_app(...)` in mount_mcp, with their reasons.
- Per-request user identity is unchanged: the contextvar set around the
  ASGI call reaches the handler on both v2 paths (legacy stateless spawns
  from the request task; the 2026-07-28 modern path opens its task group
  inside the request).
- pyproject: mcp[cli]>=2.2, no ceiling (installs are --locked). uv.lock
  regenerated with --upgrade-package mcp in the ci-python image: only mcp
  and its own dependencies moved (106 → 110 packages).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 07:21:27 -04:00
bvandeusenandClaude Opus 5.5 4502f0a1ae feat(dedup): the create gate's similarity bars are settings (#4385, rule 25)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / integration (push) Successful in 56s
CI & Build / Python tests (push) Successful in 1m38s
CI & Build / Build & push image (push) Successful in 38s
gate_bars(user_id, note_type) resolves the block bar and, for notes and
tasks, the overlap floor from kb_gate_* settings, with the old constants
as defaults. Fail-open on an unreadable value; a block bar clamps at 0.80
and the overlap floor at 0.70 and never above the bar. Five fields in
Settings beside the duplicate-report floors.

CLAIM_LEASE stays a constant, with the reason written at it: a per-user
lease would make one shared task live to one reader and dead to another.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 07:11:10 -04:00
bvandeusen f385d9075b Merge pull request 'dev → main: task claims (milestone 381), embedding model stamp, dedup copy band, sweep-shared.css' (#185) from dev into main
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / integration (push) Successful in 51s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m33s
CI & Build / Build & push image (push) Successful in 13s
2026-09-24 06:57:12 -04:00
bvandeusenandClaude Opus 5.5 baf22179ef refactor(frontend): the sweep pane's column wrapper joins sweep-shared.css (#3207)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Build & push image (push) Successful in 34s
.sweep was the one rule still written identically in both sweep panes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 06:53:11 -04:00
bvandeusenandClaude Opus 5.5 c787957ddf fix(frontend): load sweep-shared.css once from main.ts — a shared <style src> breaks vite build (#3207)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Python tests (push) Successful in 1m36s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 52s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Build & push image (push) Canceled after 35s
plugin-vue throws "Cannot read properties of undefined (reading 'scoped')"
when several SFCs share one src stylesheet. vue-tsc passed, so only the
image job caught it; :dev has not rebuilt since 22b7a92.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 06:50:59 -04:00
bvandeusenandClaude Opus 5.5 4a93c8b262 fix(task-logs): a collaborator who may write a task may log on it (rule 78)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / integration (push) Successful in 52s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m37s
CI & Build / Build & push image (push) Failing after 24s
create_log filtered on Note.user_id == user_id, a bare owner check, so a
collaborator with write access to a shared task was told it did not exist
— and, since a log now stamps the claim, could never be seen working it.
It now asks can_write_note. Editing and deleting a log still require its
author, which is authorship rather than access.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 06:48:01 -04:00
bvandeusenandClaude Opus 5.5 22b7a928da refactor(frontend): derive the sweep-row visual language into sweep-shared.css (#3207)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 52s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m32s
CI & Build / Build & push image (push) Failing after 22s
The note sweep, rule sweep, preference drift and rule history panes each
restated the same row recipe in a scoped block. It now lives once in
assets/sweep-shared.css under `sweep-` names (unprefixed globals would
leak into the unrelated scoped .row-title/.lede/.age/.actions elsewhere),
imported unscoped beside each pane's own scoped remainder.

Moved only what was shared: RuleHistoryPanel takes .sweep-state and keeps
its button row head; PreferenceDrift keeps its tiny-type footnote and
overrides the action row's layout. Lede and state now use the body-sm
token instead of 0.85rem/0.9rem (13px vs 13.6/14.4px) — one value, from
the design system.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 06:45:43 -04:00
bvandeusenandClaude Opus 5.5 952e56ee75 feat(tasks): the hand-off — SessionEnd releases a session's claims; the practice is written down (milestone 381 step 4)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 1m40s
CI & Build / Build & push image (push) Successful in 24s
Release is the mechanical half: scribe_session_end.sh sends the ending
session's id to /api/plugin/release-session, which releases the claims it
held. A tidy-up, not the guarantee (no SessionEnd on a crash; the lease
covers that), and skipped on /clear so SessionStart(clear) can still
hand the claimed work back.

Saying what happened is the half only the model can do. It is stated as a
practice where it is read: the using-scribe skill owns it ("Hand off before
this session's context stops existing", pinned in test_guidance_ownership),
the static context points at it for the wrap-up moment, and add_task_log's
docstring says a log claims the task. _INSTRUCTIONS is untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 06:43:14 -04:00
bvandeusenandClaude Opus 5.5 e46eea3b52 feat(tasks): SessionStart reads the claim — a compaction gets its work back (milestone 381 step 3)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m35s
CI & Build / Build & push image (push) Canceled after 45s
The claim now pays rent to the session that set it. The SessionStart hook
sends the host's `source` and the session id; the server renders a claim
section by source (task_claims.render_claims):

- compact / clear: the work this session had claimed, each with its two
  latest log entries, so a compacted session resumes from the record
  rather than from a count of open tasks.
- startup / clear: other sessions' live claims (may still be running) and
  in-progress tasks whose claim went quiet (abandoned mid-task).
- fork: the same, framed as "two sessions may now hold this".
- resume, or no source sent: nothing.

The task sidebar shows a live claim as "Being worked" and a dead one on
open work as "Went quiet", so the operator sees a session die mid-task.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 06:41:19 -04:00
bvandeusenandClaude Opus 5.5 a585fe0be0 fix(backup): a task claim does not travel in a backup (milestone 381)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m38s
CI & Build / Build & push image (push) Successful in 29s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 19:11:48 -04:00
bvandeusenandClaude Opus 5.5 b06b3a1d8a feat(tasks): a task can say a session is working it — the claim (milestone 381 step 2)
CI & Build / Python lint (push) Successful in 3s
CI & Build / integration (push) Successful in 53s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Python tests (push) Failing after 1m18s
CI & Build / Build & push image (push) Skipped
CI & Build / Plugin hooks (push) Successful in 14s
status is durable and nothing clears it, so in_progress cannot also mean
"someone is on this now". The claim is that second meaning, stored as who
and when so it dies on read rather than needing to be cleared.

- notes.claimed_by / claimed_at / claim_touched_at / claim_session (0110).
- The server stamps it on the write that is the work: reaching in_progress
  (update or create, via apply_status_transition) and a work log on an open
  task. done/cancelled/todo release it. Live while touched within
  CLAIM_LEASE (2h); dead on read past it, with no sweep.
- The plugin's PostToolUse hook on update_task/add_task_log binds the
  harness's session_id (GET /api/plugin/claim-session). It binds only to a
  live claim the caller holds; it cannot create one.
- to_dict carries `claim`. Plugin version minted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 19:10:06 -04:00
bvandeusenandClaude Opus 5.5 b8f543f45a fix(tasks): a create that names a status is stamped like an update to it (#3683)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / integration (push) Successful in 55s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Build & push image (push) Successful in 27s
build_note (create_note and create_records) set the status and nothing it
implies, so create_task(status="in_progress") wrote a started task with no
started_at, and done/cancelled had no completed_at or next recurrence.
The transition is now one function, apply_status_transition, called by
both paths; tests pin the invariant for every status.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 19:04:27 -04:00
bvandeusenandClaude Opus 5.5 f4e9cd429b feat(embeddings): every vector records the model whose space it lives in (#4132)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m40s
CI & Build / Build & push image (push) Canceled after 27s
The four embedding tables stamped chunker_version but not the model, and
vector(384) is a width, not an identity: a same-width model swap would
write a second geometry beside the first with no error.

- embedding_model on note/rule/milestone/system embeddings (0109; existing
  rows stamped with the only model any install has ever run).
- Every write stamps EMBEDDING_MODEL; every backfill's "current" test is
  is_current_stamp(), both halves of calibration_stamp().
- migrate_floor refuses while any row its surface searches is off the live
  model, before sampling: re-embed, then migrate.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 19:02:30 -04:00
bvandeusenandClaude Opus 5.5 11b286d786 test(dedup): a note is gated at its copy band, not the general floor (#4306)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m41s
CI & Build / Build & push image (push) Successful in 30s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 19:00:03 -04:00
bvandeusenandClaude Opus 5.5 fc1c463641 feat(dedup): a note or task blocks only as a copy; a close match is surfaced for judgement (#4306)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / TypeScript typecheck (push) Successful in 58s
CI & Build / integration (push) Successful in 1m12s
CI & Build / Python tests (push) Failing after 1m29s
CI & Build / Build & push image (push) Skipped
Measured on the live corpus, the 74 note pairs at or above the old 0.90 bar
were almost all distinct siblings — consecutive dev-logs, sub-notes of one
design, research parts — and the one clear copy sat at 0.997. The block
refused the next dev-log and taught force=true, as #4134 found for rules.

- The semantic arm blocks notes and tasks only at >= 0.98. The title block
  stays; processes keep 0.90 (not measured).
- 0.87 to 0.98 comes back as `overlaps` on the create reply, from the same
  per-chunk searches, with a note that leaves the call to the session:
  fold in and delete if it is the same record, keep both if a sibling.
- create_note, create_task, create_records and start_planning's steps all
  carry it; a batch names the record each overlap belongs to.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 18:53:22 -04:00
bvandeusen eb6912c722 Merge pull request 'Snippets gain notes; when_to_use is the situation it is ranked on (#4378)' (#184) from dev into main
CI & Build / Python lint (push) Successful in 3s
CI & Build / integration (push) Successful in 54s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m43s
CI & Build / Build & push image (push) Successful in 17s
CI & Build / Plugin hooks (push) Successful in 12s
2026-09-23 18:00:35 -04:00
bvandeusenandClaude Opus 5.5 4ae18a9dd9 feat(snippets): a snippet has notes; when_to_use is the situation it is ranked on (#4378)
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / Python tests (push) Successful in 1m37s
CI & Build / Python lint (push) Successful in 3s
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Build & push image (push) Successful in 34s
A snippet had no field for prose, so what a session learned about one went
into when_to_use — the trigger joined onto every chunk it is embedded as. A
sweep found write-ups of up to 3 KB there, headings and all.

- notes: stored after the code under `## Notes`, parsed back from the body,
  carried by every path that rebuilds it (update, merge, un-merge). A
  snippet with no notes composes the body it always did.
- create/update_snippet (MCP) take notes and return trigger_advice when
  when_to_use is long, headed or multi-paragraph. Advice, not a refusal.
- Tool docs, the reusing-code skill and the editor hint describe the trigger
  as the situation and point the explanation at notes.
- Editor gains a Notes field; the detail view renders it as markdown.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 17:27:22 -04:00
bvandeusen 14734ebe46 Merge pull request 'dev → main: auto-inject reads the conversation, startup names its failure, menu lines carry metadata, titles are names' (#183) from dev into main
CI & Build / integration (push) Successful in 50s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Build & push image (push) Successful in 17s
2026-09-23 16:57:42 -04:00
bvandeusenandClaude Opus 5.5 66e21a6c60 refactor(notes): a snippet's and lesson's stored title is its name; the trigger joins it only in the embedded document (milestone 427)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / integration (push) Successful in 52s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m35s
CI & Build / Build & push image (push) Successful in 32s
The title was `subject — trigger` because the stored title WAS the
embedded one, and the join is what makes these kinds rank on the
situation they apply to (#2485). Every surface that shows a title then
showed the trigger too -- menus, lists and search rows ran to kilobytes.

- embeddings.document_title(title, note_type, data, body) joins the
  trigger from `data` (body fallback) at embed time. Idempotent: an
  un-migrated composed title comes out the same, never doubled. The
  embed path, the startup backfill and the dedup gate's semantic signal
  all use it, so the embedded text -- and every vector -- is unchanged.
- Writers store the subject: snippet create/update (service, REST, MCP)
  and lesson_document. Both compose_title helpers are removed.
- Readers: dedup takes `data`; the menus strip the embedded title from a
  passage; list rows project `when_to_use`, which SnippetListView reads.
- 0108 rewrites existing rows on an exact `' — ' || <own trigger>`
  suffix with raw SQL, leaving updated_at alone so the backfill does not
  re-embed the corpus for identical vectors. Downgrade recomposes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 16:48:28 -04:00
bvandeusenandClaude Opus 5.5 bb632c4196 feat(retrieval): a menu line is a name, its kind and System, and the whole passage that matched (#4364)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 58s
CI & Build / Python tests (push) Successful in 1m36s
CI & Build / Build & push image (push) Successful in 26s
Injected lines rendered a snippet's or lesson's title, which carries its
whole trigger by construction (the embedding shape) and ran past 1,500
characters -- again on every `seen` repeat. The passage under a line was
cut to 200 chars from the middle, keeping its head (the title again) and
losing where the match was.

Now, on both the prompt menu and the write-path prior-art menu:
- the line shows the record's NAME (snippet data.name / lesson subject),
  with its kind and System (`[issue (done) · Plugin & hooks]`);
- the passage is the whole matched chunk, title prefix stripped, on one
  line so the blockquote holds; a title-only match hands over the trigger;
- a `seen` record is a one-line pointer to what is already in context.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 16:29:48 -04:00
bvandeusenandClaude Opus 5.5 20227ebb5d fix(plugin): recent-context cap counts text, and an empty transcript is not a failure (#4364)
CI & Build / Python lint (push) Successful in 3s
CI & Build / integration (push) Successful in 45s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Build & push image (push) Successful in 51s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / Python tests (push) Successful in 1m35s
The trailing newline became a space before the 600-char cut, so the cap
kept 599 characters of text; and under pipefail a grep that matched no
reply made the helper exit 1. Trim before cutting; return 0 explicitly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 16:08:01 -04:00
bvandeusenandClaude Opus 5.5 eb00554976 fix(plugin): a SessionStart that loads no project says why, and names the first move (#4366)
CI & Build / Python tests (push) Canceled after 23s
CI & Build / integration (push) Canceled after 27s
CI & Build / TypeScript typecheck (push) Canceled after 28s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / Build & push image (push) Canceled after 0s
The context fetch folded a timeout, an HTTP error and a refused key into
one sentence that ended "enter_project() as needed" -- read as optional,
so a session could start with no recent milestones or open tasks and no
way to know what prior work existed. The fetch now names the cause and
the elapsed time, retries once (6s) only where a retry can change the
answer, and the fallback states enter_project as the first step, with
the marker's project id when there is one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 16:07:33 -04:00
bvandeusenandClaude Opus 5.5 574b27ae74 feat(plugin): auto-inject retrieves on the conversation, not the typed words alone (#4364)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / integration (push) Successful in 49s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Failing after 1m7s
CI & Build / Build & push image (push) Skipped
The prompt hook sent only the operator's message, so a mid-session
follow-up ("yes do that") named nothing a note, lesson or rule could
match. The hook now reads the tail of the last assistant reply from the
transcript and sends it as `ctx`; the notes and rule arms append it to a
short prompt (<= 280 chars), prompt first, capped at 600 chars. No model
tokens: it is embedding input, and the injected menu's budget is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 16:05:21 -04:00
bvandeusen bf2871ffb4 Merge pull request 'Shape ledger: a semantic miss is not evidence; floors sized for the reader (#4208, #4306)' (#182) from dev into main
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / Python lint (push) Successful in 2s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Build & push image (push) Successful in 18s
2026-09-22 09:50:33 -04:00
bvandeusenandClaude Opus 5 a01deeb851 fix(shapes): the signature basis admits the band true instances score in (#4306)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 57s
CI & Build / Python tests (push) Successful in 1m44s
CI & Build / Build & push image (push) Successful in 27s
Measured over 459 judged instance rows against every same-family canon:
true instances score 1.0 or ~0.795 (a model whose mixin list differs from
the canon's), and 0.8 cut the second band entirely. The proposal is read
before it is confirmed, so the floor admits it: 0.75 takes every measured
true row, at 35 wrong-canon pairings instead of 14.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
2026-09-22 08:44:58 -04:00
bvandeusenandClaude Opus 5 312dc9f6f7 fix(shapes): the semantic arm proposes at the write-path floor, not a private 0.8 (#4208)
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 56s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Build & push image (push) Successful in 22s
0.8 was sized for proposals nobody reads; every semantic proposal is read
by a judge before it is confirmed. Measured, it proposed 0 of 150 while
judged instances of a canon score 0.68-0.71 - the write-path hint's own
floor asks the same question of the same documents at 0.68. The arm now
uses that floor, scans 8 hits instead of 3 (a true instance ranked 4th
behind snippets of other language families), and the proposer version
bumps to 5 so rows examined under the old floor are read again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
2026-09-22 08:33:01 -04:00
bvandeusenandClaude Opus 5 a875a1b2ee fix(shapes): a semantic miss is not evidence — the divergence check stops reading it (#4208)
CI & Build / Python lint (push) Successful in 2s
CI & Build / integration (push) Successful in 53s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / Python tests (push) Successful in 1m32s
CI & Build / Build & push image (push) Successful in 24s
Measured live: true instances of the service-unit canon score 0.68-0.71
against it at best, helpers 0.66-0.75 against unrelated snippets, and
nothing reaches the 0.8 floor. The "conclusive miss" fired for nearly every
body and silenced real divergences exactly as it silenced helpers. The
stored miss basis, its flag withdrawal and the report plumbing are removed;
the floor stays at 0.8 and the new-shapes-first ordering stays. False
prompts are answered by judgment (exempt with a reason).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
2026-09-22 08:27:15 -04:00
bvandeusen 3530ed5e87 Merge pull request 'dev → main: the meaning check reads new shapes first, and its miss withdraws an early flag (#4208)' (#181) from dev into main
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Build & push image (push) Successful in 14s
2026-09-22 08:14:50 -04:00
bvandeusenandClaude Opus 5 e2197ac799 fix(tests): assert the withdrawn flag on its row, not on a count the skew margin inflates (#4208)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Build & push image (push) Successful in 14s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
2026-09-22 08:09:52 -04:00
bvandeusenandClaude Opus 5 19438cc894 fix(shapes): the capped meaning pass reads new shapes first, and its miss withdraws an early flag (#4208)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Failing after 51s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m33s
CI & Build / Build & push image (push) Successful in 21s
Measured on the first live refresh after 91cde6c deployed: the version bump
queued 1,533 rows for the semantic arm, the cap read 150 of them in row
order, and all six shapes new since the previous refresh - the only rows
flag_divergence acts on, and the highest ids - were flagged before the arm
reached them. The gate silenced nothing because it never got to look, and
since a flag persists until judged, a later conclusive miss could not take
it back.

- propose_for_repo sorts the semantic todo by _semantic_priority:
  unclassified before scoped, newest first.
- flag_divergence withdraws a standing flag when the row now carries a
  conclusive miss; that evidence alone withdraws one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
2026-09-22 08:05:49 -04:00
bvandeusen cc791de0e5 Merge pull request 'dev → main: rule overlap check, design-guidance write arm, usage chip seam, divergence meaning gate' (#180) from dev into main
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / Python lint (push) Successful in 2s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 51s
CI & Build / Python tests (push) Successful in 1m33s
CI & Build / Build & push image (push) Successful in 16s
2026-09-22 07:48:53 -04:00