c404a3127a6e38e5b77aa403a61431e4a1a900f4
590
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c404a3127a |
test(retrieval): the proposal endpoints are routed on the app (milestone 458 step 7, #4925)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 19s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m2s
CI & Build / Python tests (push) Successful in 2m17s
CI & Build / Build & push image (push) Successful in 56s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
570d1b6d5a |
test(backup): register the rule_moment_judgments builder in the import column guard (milestone 458 step 7, #4925)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 56s
CI & Build / Python tests (push) Failing after 1m44s
CI & Build / Build & push image (push) Skipped
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
dfcf4df2e9 |
feat(moments): mount the corpus by proposal - a pass and an open-after-moment signal, both stopping at the operator (milestone 458 step 7, #4925)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 19s
CI & Build / integration (push) Failing after 43s
CI & Build / Python tests (push) Failing after 46s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Build & push image (push) Skipped
A rule written before moments existed is mounted on nothing. Step 7 records, per (rule, moment), whether it belongs there and who said so: - rule_moment_judgments (migration 0118, backup v23): suggested / confirmed / rejected, from a pass, the signal, or an edit. Moment "" is "no moment fits". - The pass: rules_to_mount lists unjudged rules; propose_rule_moments records suggestions that mount nothing; rule_moment_proposals and judge_rule_moments put them to the operator. A confirm mounts, a reject is kept so the pair is never proposed again. Same service behind REST and a "Waiting on you" panel in Settings > Moments. - Edits are judgments: set_rule_moments, the one mount write path, confirms what was added and rejects what was removed in the same transaction. - The signal: scribe_moment.sh keeps a per-session acts ledger; when a rule is opened, scribe_record_opened.sh sends the last three minutes of it to /api/plugin/rule-opened. The acts resolve through the install's mappings; work.run and work.change are not evidence. Counted per distinct session with lesson_rules' evidence model, and once due the open returns one line asking the reader to offer the mount. - scribe_session_end.sh removes the session's scribe-moment files. Plugin 2026.10.05.2003. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
b73a689849 |
feat(moments): the human door onto mounts, mappings and per-moment telemetry (milestone 458 step 6, #4924)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 18s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Python tests (push) Failing after 1m34s
CI & Build / Build & push image (push) Skipped
CI & Build / integration (push) Successful in 2m7s
Everything an agent can do with moments, a person can now see and change in the app. - Rule editor: a moment picker beside the trigger. Catalog moments are ticked; a named procedure's `skill.<name>` is typed and checked as the server checks it. `moments` is always sent, so unticking the last moment unmounts the rule. - Settings, Moments section (General tab): for each moment, what it means, the actions that reach it on this install (shipped ones can be switched off, the install's own removed), how many rules are mounted on it, and deliveries and agent opens over the window. Below that: named procedures with mounts, switched-off defaults with Restore, and a form to add an action. - retrieval_telemetry.moment_usage: per moment, `delivered`, `rules`, `opened` (agent pulls after the first delivery there; an upper bound, as by_source is) and `last_delivered_at`. No ratio, because a mount is a person's statement, not a ranker's guess. Guarded on its own, and also reported in retrieval_summary as `moment_usage`. - rulebooks.mount_counts; mounted_moments now derives from it. - GET /api/retrieval/moments carries `mounted` and `usage` (?days=). DELETE /moments/mappings also reads the mapping from query parameters, since the browser's DELETE sends no body. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
c489452a2d |
test(moments): an install removal still drops its tool while the skill loader stays listed (milestone 458 step 5, #4923)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 19s
CI & Build / TypeScript typecheck (push) Successful in 1m2s
CI & Build / integration (push) Successful in 1m20s
CI & Build / Python tests (push) Successful in 2m12s
CI & Build / Build & push image (push) Successful in 33s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f1fdc4a951 |
feat(moments): skills and stored processes declare the moments they are for (milestone 458 step 5, #4923)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m3s
CI & Build / Python tests (push) Failing after 1m25s
CI & Build / Build & push image (push) Skipped
Loading a procedure now also reaches the moment it is for. Loading the reporting procedure is a report; loading the release procedure is a delivery. - Bundled skills: each SKILL.md declares `metadata: moments:`. The same declaration ships as Skill defaults (BUNDLED_SKILL_MOMENTS), because the server never sees the plugin's files. test_skill_moments holds the two together and pins the plugin name that qualifies the skill. - Stored processes: `moments` on create_process and update_process, stored in the note's data and returned by get_process. A `scribe-proc-<slug>` load resolves its process through the sync manifest at load time. The moments are not copied into the stub, which would go stale mid-session. - reachable_tools lists the skill loader whenever anything is mounted, since a process's moments are known only when it loads. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
c6cdfc2172 |
feat(moments): the reply moment holds a finished reply for one read (milestone 458 step 4b, #4922)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / integration (push) Successful in 1m1s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Python tests (push) Failing after 1m25s
CI & Build / Build & push image (push) Skipped
The reply is the one act no tool call marks, and it is where "let me know if it works" gets said. A new Stop hook (scribe_reply_check.sh) sends the finished reply to POST /api/plugin/reply-rules, which checks it twice: - mounted: every unopened RULE on reply.report, plus reply.ask when the reply asks a question. Deterministic. - semantic: the reply's head and tail against every rule's trigger, on a new ranked surface, reply_rule. It is the backstop for whatever the earlier arms missed. Its floor is its stop bar (default 0.80, budget 1), with its own Settings dials. The new stop_only stage records surfacing for the rule that holds and nothing else, because nothing else reached anyone. Following the operator's ruling from 456 step 8, a rule that holds blocks once, in the server's words. The hook blocks only on a reason it was given, so an unreachable instance never stops a session, and it never holds the rewrite. The ledger is the act checkpoint's own, so a rule holds a session once across both doors and the per-session cap counts both. The turn reader moved from the report check into scribe_defs.sh (scribe_turn_facts / scribe_turn_fact), so the two Stop hooks read a turn the same way. The output was checked identical on a real transcript. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
b3b616b20a |
fix(moments): the tools import the delivery module, and the tool-list cache leaves the swept directory (milestone 458 step 4a, #4922)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / integration (push) Successful in 49s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m48s
CI & Build / Build & push image (push) Successful in 32s
Two guards caught
|
||
|
|
78653130d6 |
feat(moments): mounted rules arrive when their moment happens, through every door (milestone 458 step 4a, #4922)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 57s
CI & Build / Python tests (push) Failing after 1m19s
CI & Build / Build & push image (push) Skipped
A rule mounted on a moment now reaches the session when an act reaches
that moment, with no semantic match involved:
- run_moment_arm on the pipeline: a lookup, not a ranked search. Each
line names the moment and the act that reached it ("at work.deliver,
reached by `git push`"), so a misfire is visible where it lands and
can be unmapped in-session. A repeat is cited, not quoted; fresh
rules are recorded surfaced under source moment_rule with the moment
in detail. No retrieval_logs row, as for the other lookups, so no
latency is persisted for this arm.
- rule_scope: a rule's home clause, moved out of semantic_search_rules
so the moment lookup scopes by the same one.
- rulebooks.rules_on_moments / mounted_moments.
- The plugin door: a catch-all PreToolUse hook (scribe_moment.sh). It
keeps /moment-tools' answer on disk for five minutes, so a call to a
tool that cannot reach a mounted rule sends nothing, and an install
that has mounted nothing sends one request per window. It shares the
rules ledger with the other arms and fails open silently.
- The MCP door: Scribe's own tools named by the shipped mappings carry
moment_rules in their response, so a client without the plugin gets
them too. The hook skips those tools. A guard pins the attach on
every one.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
cc26054437 |
feat(moments): rules mount on moments, through every rule door (milestone 458 step 3, #4921)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m0s
CI & Build / Python tests (push) Successful in 1m51s
CI & Build / Build & push image (push) Successful in 29s
rule_moments (migration 0117) records which moments a rule arrives at, by catalog name, cascading with the rule. rule_detail, the one seam every rule door already returns through, gains moments beside system_ids: None leaves the mounts alone, a list replaces them. get_rule and both list_rules doors read them back, batched per page. All five MCP rule/preference writes and the three REST ones take moments and validate them before their create or update. An unknown name is refused with the catalog listed and leaves no half-made rule behind; a parity test pins that ordering on every door. Backup v22 carries the mounts as a join table remapped through the rule map; a real-Postgres round trip checks they land on the restored rule. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
cf3de5bae1 |
feat(moments): actions map onto moments, with in-session corrections (milestone 458 step 2, #4920)
CI & Build / Plugin hooks (push) Successful in 18s
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m23s
CI & Build / Python tests (push) Successful in 2m0s
CI & Build / Build & push image (push) Successful in 36s
moment_actions.resolve(tool, input) names every moment a call reaches and the action that reached it. One call can reach several: kubectl apply is a run, a deliver and a reach outside the workspace. Command tools match by how each segment of the line starts, with a word boundary; other tools by field=value arguments. The MCP server prefix and case are ignored. 56 shipped defaults cover the harness tools, Scribe tools and common command shapes. moment_mappings (migration 0116) holds what an install adds and the defaults it switches off. A removal is a stored row, so an upgrade does not switch the default back on. Per the operator ruling, corrections happen in the session: map_action and unmap_action (write tools) return now_reaches so the fix can be confirmed in the same reply. list_moments now shows each moment's actions on this install. REST mirrors both doors, recorded as human. Backup v21 carries the mappings. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
2ff7f2f34f |
feat(moments): the moment catalog rules will mount on, readable in-session (milestone 458 step 1, #4919)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 58s
CI & Build / Python tests (push) Successful in 1m53s
CI & Build / Build & push image (push) Successful in 32s
Fourteen generic moments of work (session.start, work.start … reply.ask) plus the skill.<name> family, each with what it means and the kinds of action that reach it, written for any kind of work rather than software alone. The catalog is code because every install needs the same mount points; which actions reach a moment is per-install data (step 2). require_moment refuses an unknown name with the catalog listed, so a typo cannot become a mount that never fires. list_moments (read-only) and GET /api/retrieval/moments hand out the same catalog. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
0720ab6dcf |
refactor(retrieval): the completion-report preferences run through the pipeline (milestone 456 step 3, #4905)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Build & push image (push) Successful in 28s
reply_preferences.completion_preferences was the last hand-written rule search + record_retrieval + record_rule_surfaced triple outside the pipeline. It is now REPORT_PREFERENCE, a RuleArm with kind="preference" over COMPLETION_QUERY, read back as records rather than as lines. - RuleArm gains `kind`. _ranked asks the ranker for that kind and drops anything else on the way out, before the row is logged. - RuleResult gains `shown`, the (score, rule) pairs behind the lines, for callers that return records. - RuleMoment.project_id may be None: logged as given, searched as `project_id or None`. report_preference rows keep their NULL project. - Guards: reply_preferences now has 0 direct rule searches. The registry constant-source example moves from report_preference to preference_slot, the pipeline slot that records through a module constant. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
b72c9a92e2 |
test(retrieval): two source guards read the pipeline where the rule arms now live (milestone 456 step 2, #4904)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 49s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Build & push image (push) Successful in 30s
CI run 8137 had two failures, both source-scanning guards whose property moved to the pipeline. The property itself still held: - test_every_surface_name_is_a_real_telemetry_source looked for source="<name>" literals. The rule arms now record under their spec field, so the spec is the join key, read from RULE_ARMS. - test_every_hook_rule_search_says_which_project_it_is_for counted 4 direct searches in plugin_context. It now expects 0 there, so a copied arm fails the count. It also asserts that the pipeline has exactly one search and that its keyword literal carries project_id and never everywhere. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f6b824b214 |
refactor(retrieval): the three rule arms run through one pipeline (milestone 456 step 2, #4904)
CI & Build / Plugin hooks (push) Successful in 19s
CI & Build / Python lint (push) Successful in 2s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / Python tests (push) Failing after 1m28s
CI & Build / Build & push image (push) Skipped
CI & Build / integration (push) Successful in 1m15s
prompt_rule, pre_tool_rule and the rule half of the write path each wrote the same steps out by hand: search, band, split fresh from repeats, log the call before any early return, reserve a slot, render, record surfacings. The copies drifted, and #3497, #3750 and #3752 were fixed one copy at a time. - New services/retrieval_pipeline.py. run_rule_arm runs the stages in one order. RuleArm is the spec (source, band, compact_tail, checkpoint, preference_slot). RuleMoment is the query and the session ledger. - The I/O (ranker and recorders) is passed in as RuleIO. plugin_context resolves it from its own names at call time, so existing patch points still apply. - _rule_band, _rule_hint_line, checkpoint_for, checkpoint_reason, the band constant and the preference slot moved into the pipeline unchanged. plugin_context re-exports them. - A recorder that raises now costs only its row, never a rendered line. - The flags reproduce today exactly. Whether prompt_rule should band, and whether the act arms should reserve a preference, are step 7. - Registry: the pipeline call sites are declared in FAN_OUT_SITES, with values read from the specs. - The #3497 structural guard now checks the one implementation, and that plugin_context writes no rule-source row of its own. plugin_context.py: 3,558 -> 2,980 lines. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
be62abf142 |
perf(retrieval): one embedding per query per plugin request (milestone 456 step 1, #4903)
CI & Build / Plugin hooks (push) Successful in 17s
CI & Build / Python lint (push) Successful in 2s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m25s
CI & Build / Python tests (push) Successful in 2m6s
CI & Build / Build & push image (push) Successful in 1m25s
A single plugin request fans one query out to several ranked arms. The operator message is searched by auto_inject, its reuse and lesson slots, prompt_rule, the preference slot and rule_via_lesson, and each arm called get_embedding on its own: up to six model calls for one vector. - embeddings.query_embedding_memo(): a request-scoped ContextVar memo that get_embedding consults. Outside a scope the default is None, so every other caller is unchanged. A failed embedding is not remembered. - memoized_query_embeddings decorates /retrieve, /tool-rules and /prior-art. - get_embeddings (document chunks) is untouched. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
d2ac7bf220 |
fix(lessons): a lesson's name is one claim on one line (#4797)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m48s
CI & Build / Build & push image (push) Successful in 28s
#4797 "A lesson's what takes a whole narrative, and the narrative becomes its name". `what` is the title every listing and menu prints. Nothing enforced its documented "one line", so lessons written with the incident in `what` printed up to ~1,500 characters as a menu line, buried the claim, and diluted the trigger in the embedded title. That was 14 of the 39 lessons on this install. - lessons.require_claim refuses a `what` over WHAT_MAX_CHARS (240) or running over several lines. The refusal says the story goes in `insight`. - Both doors call it before writing: - MCP create_lesson / update_lesson; - REST create / update. An update checks only a NEW name, so a lesson stored with a long one can still take the edit that repairs it. - lessons.claim_line is the display half. _menu_name shows an over-long stored name as its first sentence marked " …", because a door guard does not undo rows already stored and an unrepaired install would keep printing them. - The tool docstrings state the bound. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
9909cd2450 |
fix(retrieval): the rules never-opened warning stops sending the reader to a review tool that refuses rules (#4798)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m7s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 26s
#4798 "The rules corpus's surfaced_never_pulled warning sends the reader to menus_to_review, which can only review auto_inject". #4772 gave the warning one remedy for both corpora: "judge a sample with menus_to_review". That is right for notes, where a line carries its passage. For rules it is a dead end: retrieval_review.REVIEWABLE holds only auto_inject, so the tool refuses every rule arm. - The reading is now per corpus (_NEVER_PULLED_READING). For rules, the text says: - the count includes rules that only arrived in a listing; - a rule can rightly be set aside on its trigger alone; - no judged sample exists for the rule arms; - a rule set aside again and again is a trigger to fix (update_rule when_to_apply, then what_might_apply), not a floor. - The tool docstring says the same. - The guard is tied to REVIEWABLE, so it can fail in both directions: if the rules text names menus_to_review while no rule arm is reviewable, or if a rule arm becomes reviewable and the text still says there is no sample. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f95ffae972 |
feat(retrieval): a record the operator names by number reaches the menu (#4796)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m47s
CI & Build / Build & push image (push) Successful in 37s
#4796 "A record the operator names by number reaches auto-inject only if its wording happens to match". "yes go ahead with 4448" says which record is meant, but a number means nothing to an embedding, so the prompt menu filled with records resembling the words around it. - record_refs.named_record_ids reads the operator's raw prompt, never the reply-enriched query: - `#N`, unless the word before it marks another numbering (PR, CI, rule, milestone, system, log...); - a bare number of 3 or more digits that opens the message, follows a reference word or continues a list one started; - never a quantity ("300 seconds"), a date, version, path or fenced code. - Each id is resolved through the ACL check; trashed or inaccessible ids are dropped. - A "Named in your message" block leads the menu: the kind, the System, the name and the opening of the body, or the seen pointer if the record is already on the ledger. It takes no share of top_k, and the semantic lines leave those ids out. - The block is booked under the new `named_ref` source, registered as an unbidden lookup that is allowed to be quiet. A named id that also ranked counts as suppressed in the auto_inject row, so #3668's identity holds. - Named records now arrive even when the search finds nothing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
5c6d9a7b55 |
fix(mcp): a tool's refusal reaches the agent with its reason (#4794)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 1m52s
CI & Build / Build & push image (push) Successful in 27s
SDK 2.x passes on only a ToolError's text. Any other exception becomes
UnexpectedToolError("Error executing tool X") and its message stays on the
server. ValueError is how every Scribe tool refuses — what was refused, why,
what to do instead — so every refusal reached the agent bare, and it could
only retry blind. Seen live: tune_retrieval(actor="operator") and
judge_menu(verdicts=[]) both answered "Error executing tool …" and nothing
else.
StrictArgsMCPServer.call_tool re-raises an UnexpectedToolError caused by a
ValueError as a ToolError carrying the message, in the SDK's own
"Error executing tool X: <reason>" shape. Other exceptions are crashes and
stay masked. The stale comment claiming the SDK returns ValueError text is
corrected.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
df4b37673f |
fix(plugin): a subagent hand-back is not an operator prompt (#4785)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 54s
CI & Build / TypeScript typecheck (push) Successful in 56s
CI & Build / Python tests (push) Successful in 1m49s
CI & Build / Build & push image (push) Successful in 29s
`<agent-message …>` turns were scored by auto-inject and the prompt rule arm as if the operator had typed them, spending a retrieval and writing a retrieval_logs row that reads as a real message (log 54355, found by the #4772 review). Same class as #4142. Matched up to the tag name, since the tag carries attributes; `<agent-messages>` and the tag named mid-sentence are still retrieved against. Plugin 2026.10.04.0125. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
78fa01b025 |
fix(embeddings): no chunk is a heading alone (#4784)
Judging live auto-inject menus (#4772) found "## Work log — 2026-08-21", "## Findings worth carrying forward" and "Two dev→main PRs this session." each embedded as a whole chunk. Text about nothing embeds close to everything, so they took ranks 1–3 on vague queries and pushed real records down. Three ways the chunker made them, each closed: - _split_paragraphs flushed a heading on its own when the paragraph under it was long. A thin lead is now carried into the paragraph that follows, and a hard cut never lands in the first quarter of the budget, where the newline it finds closes a heading. - A heading with no body (above its ### parts) became a section of its own when the section before it was full. A thin section now takes the next. - A short lead-in or a short last section stood alone. The lead-in joins what follows; a thin tail joins what precedes it. A continuation piece now repeats the LAST heading before it, since one section can now hold several. CHUNKER_VERSION 2 → 3, so the startup backfill re-embeds; tuned dials will report their shape as changed, which is true. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
5d0e97576b |
fix(retrieval): the review drops records written after the call it re-runs (#4773)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m8s
CI & Build / Python tests (push) Successful in 1m53s
CI & Build / Build & push image (push) Successful in 36s
The first live sample ranked records the call could never have been offered: the session that made a call writes the decision, often quoting the message, and that record tops the re-run. Judged, it inflates on_point exactly where the budget is decided. - _rerun over-fetches by POSTDATED_SLACK, drops records created after the call before ranking, and names them in `postdated`. - each line carries `changed_since_call` (updated_at or a work log after the call) and `logged` (shown fresh then). - judge_menu shares the re-run, so a post-dated record cannot be judged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
6598c7fa85 |
feat(retrieval): a review pass judges whether injected lines related — menus_to_review, judge_menu and a judged readout (#4772)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 44s
An open rate cannot say whether a menu line related: every line carries its matched passage (#4364), so "not opened" covers unrelated, enough as shown, and already in context. #4772 "Injected notes are never judged". - retrieval_judgments (0115): a reviewer verdict per line of a logged call, on_point / adjacent / unrelated, with its reason, rank, budget side and whether the agent opened it within the hour. - menus_to_review re-runs a random sample of unjudged auto_inject calls with the arm's own parameters, past its budget, passage on every line. judge_menu records verdicts, re-deriving rank from a fresh re-run. - retrieval_telemetry gains a judged block (by rank, within/beyond budget, on_point_unopened). surfaced_never_pulled stops blaming titles. - missed-retrieval guidance names the review before a budget move. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
0a1bb68808 |
feat(rulings): system_usage_events is read back — per-System counts and a telemetry block (#4769)
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Python tests (push) Successful in 1m44s
CI & Build / Build & push image (push) Successful in 36s
#4769 "Rulings are counted where someone will read them": milestone 444 step 4 wrote system_usage_events and nothing read it. - retrieval_telemetry gains a `system_usage` block: surfacings and opens by source, distinct counts, and `by_system` naming the areas most shown. There is deliberately no pull-through ratio, because rulings travel in full in the line and opens are the exception. - usage_for_systems (one GROUP BY) adds `usage` to the REST Systems list and detail, and to MCP get_system. MCP list_systems is unchanged. - The Systems UI shows a "rulings shown N×" chip. - rulings_pre_tool, rulings_write_path and mcp_get_system are now declared registry points; the registry guard covers their recorders. - The Systems store merges a PATCH reply instead of replacing the row. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
1b973ebd13 |
fix(write-path): a search hit's excerpt under snippet no longer 500s the prior-art route (#4768)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m5s
CI & Build / Python tests (push) Successful in 1m50s
CI & Build / Build & push image (push) Successful in 39s
#4768 "Write-path prior-art route 500s when a nameless search hit carries its excerpt under `snippet`": _prior_art_line read item["snippet"] as a snippet record's field dict, but a search hit carries its matched passage there as a string. Read the name only from a dict. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
891f375715 |
test(rulings): the hook/route parity guards pin the rulings params (#4757)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 49s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m49s
CI & Build / Build & push image (push) Successful in 33s
CI 7939 went red on test_the_hook_and_the_route_agree_on_every_parameter_name: the Bash hook now sends root and cwd, and the guard pins the exact set. Both guards (tool-rules and prior-art) now also pin seen_ruling_systems, spelled once in scribe_rulings_query, and check the routes read all three. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
556872c039 |
feat(rulings): a command or edit touching an area's files shows its rulings, once per session (milestone 444 step 4, #4757)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m6s
CI & Build / Python tests (push) Failing after 1m22s
CI & Build / Build & push image (push) Skipped
A System's rulings (the Rulings section of its description) now reach the work by path, not by similarity. Both PreToolUse arms resolve the files a command or edit names to the Systems whose path_patterns cover them, and the first touch in a session shows each area's rulings in one line; a repeat is a one-line reference. A lookup, so no floor, no budget, no retrieval_logs row. - services/system_rulings: parse_rulings, command_paths (reads and writes, relative to the repo root from any cwd; flags, URLs, globs skipped), rulings_for_paths - /tool-rules takes root, cwd and seen_ruling_systems; /prior-art takes seen_ruling_systems; both return ruling_system_ids - hooks share <sid>.rulings.ids (cleared on compaction by the ledger naming convention); the Bash hook sends the repo root and cwd - system_usage_events (migration 0114): surfacings by source, pulls from get_system; carried by backup (v20) through the system map - writing-records: rulings also arrive when the area's files are touched Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
a113c72b4f |
feat(systems): a System names its files — path patterns stored, validated and matched (milestone 444 step 3, #4756)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m6s
CI & Build / Python tests (push) Successful in 1m50s
CI & Build / Build & push image (push) Successful in 43s
A System gains path_patterns: globs relative to the repo root (* within one directory, ** across any depth, a plain directory covering everything under it). One service validates them for every door, so the web UI and MCP refuse the same bad pattern with the same message. systems_for_paths resolves paths to every active System that covers them, which step 4 (#4757) uses to deliver an area's rulings when its files are touched. - schema: systems.path_patterns JSONB NOT NULL default [] (migration 0113) - service: normalize_path_patterns, path_matches, systems_for_paths - routes + MCP create_system/update_system accept it; [] clears - web UI: a Files field in the create and edit forms, patterns on the card - backup carries it through export and restore - using-scribe reflex 7: tagging work keeps a System's files current Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
01d8a0b9f1 |
feat(guidance): an operator's ruling lives on the System it governs, and code is read as behaviour, not intent (milestone 444 steps 1-2, #4754 #4755)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m46s
CI & Build / Build & push image (push) Successful in 56s
A Librarian session contradicted a decision the operator had made 13 days
earlier. The ruling ("retry, then replace, never give up on a book") was
kept only as a quote in a work log, beside a session's reading of it that
capped replacements at 3. Three later sessions built on the reading, and one
carried the cap into an option as a "known cost", which the operator then
approved without being asked about it.
- writing-records.md: "A ruling goes on the System it governs". What a
ruling is (the operator decided it; a later change could undo it), how it
differs from a rule, and where it goes: a Rulings section at the end of
the System description, one line each with who, when and the source
record. Written the turn the operator decides; holds what is in force,
not history; a charter line that contradicts a ruling is fixed in the
same edit.
- using-scribe SKILL.md: reflex 1 says code tells you what a thing does,
not what was wanted, and a limit read from code is unconfirmed until a
System's Rulings says otherwise. Reflex 9 points to the ruling section.
- reporting-back: an option that carries existing behaviour says whose call
it was (the operator's ruling, or a past session's never confirmed); one
that contradicts a ruling is a Conflict.
- create_system / update_system docstrings: the Rulings section, and that
description replaces the whole text.
- test_guidance_ownership: three topics pinned to their owners.
- Plugin version minted.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
cf4206469b |
feat(lessons): the web editor records which rule a lesson is an instance of (milestone 440, #4658)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 37s
The lesson editor now asks the question the create tool asks, while the writer still has the situation in mind. It offers three answers: - an instance of a rule, ticked from the rules the lesson resembles; - "No rule fits", with a required reason; - leave it open. The rules are fetched before the save from a new GET /api/lessons/rule-candidates. It runs the same rule_candidates search the create door replies with, and passes "search unavailable" through as null, apart from "nothing resembles it". An edit sends an answer only when it changed. Re-sending the same rules would re-stamp their judgments. Worse, a shared editor who cannot see the owner's rule would reject it by sending a list without it. Leaving a linked lesson open sends rule_ids=[], which is what clears the links. "Leave it open" is not offered once "no rule fits" is on file: the server has no way to take that answer back except by naming a rule. When "no rule fits" completes a convergence group, the editor says so in a toast, since the page it goes to does not recompute it. Guards check that the payload uses the names both routes read, that the three answers are offered, that the candidates route sits above the id route and keeps null apart from [], and that the editor reuses .rule-chip. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
75cefe60e4 |
feat(lessons): both records show the link in the web UI (milestone 440 step 7, #4635)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / integration (push) Successful in 1m9s
CI & Build / Python tests (push) Successful in 1m53s
CI & Build / Build & push image (push) Successful in 41s
The lesson page gains an "Instance of" panel that holds one of three answers. Unjudged is the fall-through, so it is stated rather than left blank: - the rule(s) it was judged an instance of; - "No rule — <why>"; - "Not yet judged". Suggested links show what they rest on (distinct situations, and projects when more than one), with Confirm / Not an instance for a reader who can write. Rejected links stay listed with their reason. The rule slide-over lists the lessons that are instances of it, plus the suggestions waiting on a judgment. Each entry links through to the other record, and kind and state wear the existing .rule-chip. The write check moves to utils/permission.ts. The copy on the snippet page looked for "edit", which the server never sends, so shared editors saw a read-only page (#4640 "The snippet page hid its edit controls from shared editors"). Guards pin the client's unions and write levels to the service's, the model's and access.py's own values. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
6d3dca0af5 |
feat(lessons): convergence is named at the write — no-rule lessons that keep landing in one situation suggest a rule (milestone 440 step 5, #4634)
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / Python lint (push) Successful in 2s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Build & push image (push) Successful in 27s
When a lesson is answered "no rule fits" (create_lesson / update_lesson on both doors), the response looks for other no-rule lessons it resembles and, once there are CONVERGENCE_LESSONS (3) of them, carries `convergence`: the members, their incidents and projects, and a hint to draft the missing rule with create_rule (operator approval as always) and point each lesson at it — or to leave them as lessons when no single choice is right every time. - convergence_group is the pure bar: distinct LESSONS count, incidents never stand in for them (one broad lesson cannot trigger it), and a group whose sources all point at one incident is one event written up several times. - convergence_for searches lessons by the new one's claim + trigger (trigger_title) at CONVERGENCE_THRESHOLD 0.65 — above the menu's "worth showing", below the duplicate gate's "same record" — then keeps the ones with a lesson_no_rule answer. Fail-open. No sweep, no timer (#4183). - Defaults stated as defaults (rules 32, 115). - Tests: the bar (pure), the search with stubs, the door, and the no-rule filter against Postgres; conftest stubs convergence_for for unit tests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f34249a2d8 |
test(rules): locate the pre-tool arm by what it does, not by its name (#4633)
CI & Build / integration (push) Successful in 52s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / Python tests (push) Successful in 1m48s
CI & Build / Build & push image (push) Successful in 26s
CI run 712 failed test_neither_rule_arm_logs_its_call_behind_a_results_guard
with "min() iterable argument is empty": the guard walked the function named
build_tool_rule_hint, which since
|
||
|
|
fd2ebf4c13 |
feat(rules): a rule surfaces through the lessons confirmed as its instances (milestone 440 step 4, #4633)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 18s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 55s
CI & Build / Python tests (push) Failing after 1m16s
CI & Build / Build & push image (push) Skipped
On all three rule arms (prompt_rule, pre_tool_rule, write_path_rule), after the direct match, a lesson matching the moment at the notes menu's own bar brings the rule(s) it is CONFIRMED to be an instance of, rendered in rule voice with "Reached through lesson #N “…”, a recorded instance of it." - Confirmed links only (lesson_rules.confirmed_lessons / confirmed_rules_in_scope). A suggested link carrying its rule would manufacture the co-arrival #4637 counts and prove itself. - Scope kept: a lesson surfaces everywhere, but the rule it brings must be global or this project's own — milestone 414's boundary, not reopened through a side door. - Suppression is by RULE: anything the direct band named or the session ledger holds is skipped, whichever lesson reached it. - Its own slot (VIA_LESSON_LIMIT = 1), not a rule slot — a stated default, since the plan's "decide by measurement" has nothing to measure until links are confirmed (#4632). Its own source, rule_via_lesson: registered, RANKED for pull-through, logged whenever it searches. - Nothing searched at all while no lesson carries a confirmed link, which is every install until one is judged. - build_prompt_rule_hint / build_tool_rule_hint are now thin wrappers over the direct arms (_prompt_rule_hint / _tool_rule_hint); the arm leaves its query in `_via_query` only when it ran, so a disabled arm or a blank prompt brings no rule in either, and the key never leaves the server. - Via-lesson rules are not fed to the #4637 co-surfacing recorder: only direct matches are evidence. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
dbbab859ce |
feat(lessons): soft links — a lesson and a rule arriving together in distinct situations are proposed as a link (milestone 440 step 3, #4637)
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Build & push image (push) Successful in 30s
When one hook response puts a lesson and a rule in front of the reader, the pair is recorded as evidence on a SUGGESTED lesson_rule_links row; once the pair has arrived together in PROPOSE_SITUATIONS (3) distinct situations, the next co-arrival carries one line asking the reader to judge it with judge_lesson_link. Nothing about surfacing changes: a suggested link carries no rule anywhere (that is #4633, confirmed links only). - lesson_rules: co_surfaced (fail-open; only pairs the reader could confirm — a lesson they may write, a rule they own; judged pairs gather nothing; one proposal per response; PROPOSE_COOLDOWN 6h between asks), plus the pure counting rules: situation_key, add_evidence, proposal_due, evidence_summary. A situation is the prompt on /retrieve (word tokens, sorted and de-duplicated, so trivial rewordings count once) and the FILE on /prior-art (every edit to one file is one situation). - rules_for_lessons shows a suggested link's evidence counts. - plugin_context: build_autoinject_hint returns lesson_ids, build_prompt_rule_hint returns shown_rule_ids, build_write_path_hint records its own pair; routes/plugin /retrieve records the prompt pair. Shown lines, repeats included: relevance makes a co-arrival, not the session ledger. - Evidence lives in the existing evidence column — no migration, and backup already carries it. - Tests: counting rules (unit), the recorder against Postgres (bar, repeat, judged pairs, ownership, cooldown, evidence kept on confirm); conftest stubs co_surfaced for unit tests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f7d8dc2e55 |
feat(lessons): judged when written — a new lesson is offered its rules, and "no rule fits" is an answer (milestone 440 step 2, #4631)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 53s
CI & Build / Python tests (push) Failing after 1m15s
CI & Build / Build & push image (push) Skipped
"Which rule is this lesson an instance of?" now has three recorded answers: a rule named (a confirmed link, #4630), no rule fits (new), or unjudged. - Model + migration 0112: lesson_no_rule (lesson_id PK, CASCADE from the note; why; judged_at). A table rather than a key in notes.data, because that mirror is re-composed from the body on every edit and would erase it. - Service (lesson_rules): set_no_rule rejects any confirmed link with the reason; a confirmation (set_lesson_rules or judge_link) deletes the answer; require_one_answer refuses both answers in one call before any write; judgments_for_lessons + attach_lesson_rules add rule_judgment (and no_rule) to every lesson payload; list_unjudged lists the open ones; rule_candidates searches rules with the lesson's claim + trigger at the explicit-search bar, None when the search could not run. - MCP: create_lesson/update_lesson take no_rule; an unanswered create returns rule_candidates, rule_judgment and a rule_hint; list_lessons(unjudged=true). - REST: the same on POST/PATCH /api/lessons and GET ?unjudged=1; create returns rule_candidates. - Backup v19: a lesson_no_rule section, export (full and per-user) and import. - Guidance: create_lesson docstring, writing-records.md in using-scribe (owner, pinned in test_guidance_ownership), create_rule docstring on linking the lessons a new rule governs. Plugin version minted. - Tests: door units, integration for the three states, the rejection reason, scoping, cascade; backup registries. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
41e4fbaba1 |
feat(lessons): a lesson names the rule it is an instance of — lesson_rule_links (milestone 440 step 1, #4630)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Failing after 1m13s
CI & Build / Build & push image (push) Skipped
The link between a lesson (one concrete situation) and the rule that governs it, with the operator's soft-then-hard design built into its state: suggested while evidence accumulates, confirmed or rejected once judged. Only confirmed will carry a rule in retrieval (#4633); rejected is kept so the pair is never proposed again. - models/lesson_rule_link.py + migration 0111: one row per (lesson, rule), CASCADE on both ends, indexed both ways, CHECK on state (rule 36), evidence JSONB and judged_at. - services/lesson_rules.py: require_rules (validated before any write, so a bad id leaves nothing half-linked), set_lesson_rules (set-semantics; a dropped rule becomes rejected, not forgotten), judge_link, and the two reads. ACL: write on the lesson (share-aware), ownership of the rule; a reader sees only rules they own. Decorations are fail-open (#4286). - MCP: create_lesson / update_lesson take rule_ids; get/create/update return `rules`; new judge_lesson_link tool. REST: the same on /api/lessons plus PUT /api/lessons/<id>/rules/<rule_id>. Rules: rule_detail carries `lessons`. - Backup v18: export (full and user-scoped, both ends in scope), builder, importer; both column guards register the table. - Tests: integration (states, set-semantics, judge, ACL all-or-nothing, cascade both ways, CHECK, one row per pair); unit (door wiring, judge registered, migration/model state agreement, backup skip and unjudged stays unjudged). conftest stubs the decorations for unit tests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
2e4c2d9493 |
feat(usage): count what happened away from a record's own project — the readout #3735 needs
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m6s
CI & Build / Python tests (push) Successful in 1m50s
CI & Build / Build & push image (push) Successful in 38s
Milestone 385 step 8 (#3735 "a lesson is recalled on a project it was not
written on") is defined by opened-on-another-project. The reader's project has
been recorded on every usage event since
|
||
|
|
582a5a4f48 |
feat(shapes): the practice is written where it is read, and the coverage line measures the slip (milestone 439 step 6)
CI & Build / integration (push) Successful in 47s
CI & Build / Python tests (push) Successful in 1m39s
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Build & push image (push) Successful in 29s
- reusing-code: "Before the turn ends — say what you built" — the four verdicts, the one classify_shapes(repo=…) call, and the component file as a shape. Description names the end-of-turn moment. - shape-accounting: the writer judges; audits are the check that it held. The write path SUGGESTS (no more hook instances); component file rows and whole-file canon described; scoped covers Svelte too. - _INSTRUCTIONS reuse line: "before the turn ends, say what you built (create_snippet the reusable, classify_shapes the rest)" — 1570/1600. - Coverage line: "written-shape check (7d): N turns checked, M asked, K left unjudged", from the Stop hook's recorded outcomes; silent until the question has been put. - test_guidance_ownership pins the new topic on reusing-code. - prior-art hook header no longer says it stamps instance rows. Plugin version minted. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
8b1567cf2c |
feat(shapes): a component file is a candidate shape in its own right (milestone 439 step 5)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 54s
CI & Build / Build & push image (push) Canceled after 0s
CI & Build / Python tests (push) Canceled after 1m42s
A single-file component defines `.card`, `.title` and a `Props`, never anything named after itself — so the ledger could not hold "StatusChip is canon" and could not say a new card was built where one already existed. - coverage.is_file_unit: a file that RENDERS (its markup names classes, or it opens a <script>/<template>/<style> block) and defines nothing named after its stem gets one `file` row. Structural, not a framework list: Svelte and Vue components qualify, a TSX component already has its function's row, a module is accounted for by its definitions. Emitting every file would have put every module of every project in the todo at once. - No body fingerprint on file rows, so an edit never re-asks a judgment. - `file` is its own form and family; derive grouping skips it (`index`, `+page` repeat by convention). Divergence buckets it on its own, so a new component where a component canon dominates is a fair question. - mark_canonicals: a snippet recorded at a path with no symbol makes that file row canonical; it still covers no definition inside. - The write hooks apply the same test before noting a new file for the end-of-turn question. Plugin version minted. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
4cf1c6f042 |
feat(shapes): the write-path hook suggests and no longer stamps (milestone 439 step 4)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m49s
CI & Build / Build & push image (push) Successful in 27s
The hook used to land its evidence as `instance` rows, classified_by=hook — a permanent verdict nobody read, and the source of the weak stamps that poisoned two ledgers (#4608). Now the agent that wrote the code judges it at the end of the turn, and the hook's evidence is what it is shown. - shape_ledger.suggest_write_path_instances (was stamp_write_path_instances): same evidence and form gate, but it writes a PROPOSAL — proposed_snippet_id, basis "reference" (named) or "semantic" (resembles), score — and never a status. A judged row is left alone; a new shape gets an unclassified provisional row so the suggestion reaches the end-of-turn question as "looks like #N". By-name uses edges stay: a fact, not a verdict. - The prior-art line says "looks like #N … when you judge this turn's shapes, say whether it is", and the result key is `suggested`. - write_time_divergence takes `suggested`; _RESEMBLE_MIN re-documented as the suggestion bar and the weak-stamp line. - reusing-code skill no longer says the pull stamps an instance. - Tests moved to the new contract, plus a guard that nothing the hook does writes classified_by="hook". Plugin version minted. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f1fbdf746a |
feat(shapes): the agent judges what it wrote, at the end of the turn (milestone 439 steps 1-3)
CI & Build / Python lint (push) Successful in 5s
CI & Build / Plugin hooks (push) Successful in 18s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m38s
CI & Build / Build & push image (push) Successful in 33s
Recording used to be decided by machinery — the only "record it" prompt fired when a same-named copy already existed (#2664), so a first instance of a reusable piece was never asked about, and judgment arrived only through audits. Now the question is asked where the knowledge is: the end of the turn that wrote the code, of the agent that wrote it. - Write hooks keep `<sid>.written.ids` (path, kind, name) for every definition a write names; a new file adds a `file` line for its stem — a candidate in any language without a framework rule (scribe_written_append). - Stop hook scribe_shape_check.sh sends the ledger to GET /api/plugin/shape-check and blocks once, in the server's words, when anything is unjudged. Same discipline as the report check: never twice, never without a recorded check, another hook's loop left alone; the ledger is kept when the instance cannot be reached. - shape_ledger.unjudged_shapes: no row, unclassified, scoped and hook stamps are unjudged; an agent/audit/import verdict is not. A snippet recorded at the shape answers for it until the refresh stamps it canonical. - services/shape_check owns the reason text and records every outcome in app_logs (passed / blocked / judged_after_block / left_after_block). - classify_shapes(repo=…) judges a shape the ledger has not synced yet via a provisional row under a bound repo; the sync confirms it, or vanishes and revives it with the verdict intact. An unbound repo is refused. Plugin version minted. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
eb5cc6d3a7 |
fix(shapes): weak stamps elect no canon; the arrival line names the review; Svelte scopes by default (#4608)
CI & Build / Python lint (push) Successful in 8s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m3s
CI & Build / Python tests (push) Successful in 1m44s
CI & Build / Build & push image (push) Successful in 1m5s
Librarian and Stash had 132 and 488 write-path stamps written at 0.68-0.72, before the 0.80 floor (#4204). Nothing re-judged them, and dominant_canon counted them, so a YAML CI snippet (#3410) "dominated" Librarian's web directory and every write there was told it diverged from it — 1408 flags on one project, 902 on the other, all skimmed past. - is_weak_stamp: one predicate for "hook-stamped below today's floor". dominant_canon and canon_form skip such rows; stamps_to_review lists them by the same predicate. No stored row changes — they stay for judgment. - flag_divergence withdraws a standing flag whose canon no longer dominates its directory. A flag is a mechanical prompt, not a judgment. - The coverage line ends "N weak stamps · M incoherent canons to judge — stamps_to_review" when either is nonzero, so a project's own session sees the queue on arrival. The shape-accounting skill says what to do with it. - scoped_definitions handles .svelte: every <style> is component-scoped unless <style global> or :global(...); instance-script syms are scoped, module-script ones stay ordinary. Plugin version minted for the skill change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
d5dad587f1 |
refactor(skills): using-scribe keeps the every-turn practices; moment-specific depth moves to reference files (#4398)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m8s
CI & Build / Python tests (push) Successful in 1m44s
CI & Build / Build & push image (push) Successful in 21s
Anthropic's skill guidance: keep SKILL.md under 500 lines, split into reference files linked one level deep as it nears that. using-scribe was 478 and every new practice lands there. - SKILL.md 478 -> 317 lines. It keeps orientation, one copy, the reflexes, scope, the judge section, UI and the process-skill index, plus a "Read these when the moment comes" list naming each file with its moment. - projects.md: binding a non-git directory (.scribe) and project inception. - writing-records.md: where a new rule goes, lesson growth, and notes that carry their own check (reflex 10 keeps a pointer). - missed-retrieval.md: the record-before-dial route, verbatim. - Text moved, not rewritten, except for the seams and one cross-reference. Tests: - tests.helpers.skill_text reads SKILL.md plus its reference files. The ownership registry, the miss-route and the verification tests use it, so a topic stays owned by its skill whichever file holds it. - The force test scans every skill .md on its own, since each file is read on its own. - New test_skill_structure: SKILL.md <= 350 lines, every reference file is linked from SKILL.md, none links another, and one over 100 lines opens with Contents. Each guard is shown to fail. The plugin version is minted. That also clears 4fb53b8's red Plugin hooks lane, which failed only because PACKAGING.md changed without a mint. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
4fb53b844d |
refactor(mcp): _INSTRUCTIONS orients the workflow, not a rulebook (#4389)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Failing after 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m35s
CI & Build / Build & push image (push) Successful in 23s
Spike #4389 read the spec, Claude's docs and a dozen servers: the field is for how the tools fit together, and the field runs ~600-1,600 characters. Ours sat at the 2,048 cap as a keyword index that also carried stance. - JUDGE, REPORT and MISSED leave the index. They fire mid-work, not at session start; using-scribe and reporting-back state them in full, and `placement`/`report_back` cue reporting in-band. No skill text changes. - The rest is rewritten as plain practices (1,503 chars) and keeps every session-start marker the ownership registry pins. - INSTRUCTIONS_BUDGET 2000 -> 1600; the three index markers are dropped from the registry; the miss-route index test now checks that the index keeps what_might_apply and stays off the route. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
c26b7f248e |
feat(mcp): port the server to MCP Python SDK v2 (FastMCP → MCPServer) (#2196)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m37s
CI & Build / Build & push image (push) Successful in 31s
- server.py: `mcp.server.mcpserver.MCPServer`; StrictArgsFastMCP becomes StrictArgsMCPServer, whose call_tool takes and forwards v2's `context` and raises ToolError (the SDK logs anything else as an unexpected crash; the message reaches the caller either way). - stateless_http and transport_security moved from the constructor to `streamable_http_app(...)` in mount_mcp, with their reasons. - Per-request user identity is unchanged: the contextvar set around the ASGI call reaches the handler on both v2 paths (legacy stateless spawns from the request task; the 2026-07-28 modern path opens its task group inside the request). - pyproject: mcp[cli]>=2.2, no ceiling (installs are --locked). uv.lock regenerated with --upgrade-package mcp in the ci-python image: only mcp and its own dependencies moved (106 → 110 packages). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
4502f0a1ae |
feat(dedup): the create gate's similarity bars are settings (#4385, rule 25)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / integration (push) Successful in 56s
CI & Build / Python tests (push) Successful in 1m38s
CI & Build / Build & push image (push) Successful in 38s
gate_bars(user_id, note_type) resolves the block bar and, for notes and tasks, the overlap floor from kb_gate_* settings, with the old constants as defaults. Fail-open on an unreadable value; a block bar clamps at 0.80 and the overlap floor at 0.70 and never above the bar. Five fields in Settings beside the duplicate-report floors. CLAIM_LEASE stays a constant, with the reason written at it: a per-user lease would make one shared task live to one reader and dead to another. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
4a93c8b262 |
fix(task-logs): a collaborator who may write a task may log on it (rule 78)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / integration (push) Successful in 52s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m37s
CI & Build / Build & push image (push) Failing after 24s
create_log filtered on Note.user_id == user_id, a bare owner check, so a collaborator with write access to a shared task was told it did not exist — and, since a log now stamps the claim, could never be seen working it. It now asks can_write_note. Editing and deleting a log still require its author, which is authorship rather than access. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
952e56ee75 |
feat(tasks): the hand-off — SessionEnd releases a session's claims; the practice is written down (milestone 381 step 4)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 1m40s
CI & Build / Build & push image (push) Successful in 24s
Release is the mechanical half: scribe_session_end.sh sends the ending
session's id to /api/plugin/release-session, which releases the claims it
held. A tidy-up, not the guarantee (no SessionEnd on a crash; the lease
covers that), and skipped on /clear so SessionStart(clear) can still
hand the claimed work back.
Saying what happened is the half only the model can do. It is stated as a
practice where it is read: the using-scribe skill owns it ("Hand off before
this session's context stops existing", pinned in test_guidance_ownership),
the static context points at it for the wrap-up moment, and add_task_log's
docstring says a log claims the task. _INSTRUCTIONS is untouched.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|