2b52afcd7290fb8da59ee4534c372df31d149c6c
211
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
dfcf4df2e9 |
feat(moments): mount the corpus by proposal - a pass and an open-after-moment signal, both stopping at the operator (milestone 458 step 7, #4925)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 19s
CI & Build / integration (push) Failing after 43s
CI & Build / Python tests (push) Failing after 46s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Build & push image (push) Skipped
A rule written before moments existed is mounted on nothing. Step 7 records, per (rule, moment), whether it belongs there and who said so: - rule_moment_judgments (migration 0118, backup v23): suggested / confirmed / rejected, from a pass, the signal, or an edit. Moment "" is "no moment fits". - The pass: rules_to_mount lists unjudged rules; propose_rule_moments records suggestions that mount nothing; rule_moment_proposals and judge_rule_moments put them to the operator. A confirm mounts, a reject is kept so the pair is never proposed again. Same service behind REST and a "Waiting on you" panel in Settings > Moments. - Edits are judgments: set_rule_moments, the one mount write path, confirms what was added and rejects what was removed in the same transaction. - The signal: scribe_moment.sh keeps a per-session acts ledger; when a rule is opened, scribe_record_opened.sh sends the last three minutes of it to /api/plugin/rule-opened. The acts resolve through the install's mappings; work.run and work.change are not evidence. Counted per distinct session with lesson_rules' evidence model, and once due the open returns one line asking the reader to offer the mount. - scribe_session_end.sh removes the session's scribe-moment files. Plugin 2026.10.05.2003. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f1fdc4a951 |
feat(moments): skills and stored processes declare the moments they are for (milestone 458 step 5, #4923)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m3s
CI & Build / Python tests (push) Failing after 1m25s
CI & Build / Build & push image (push) Skipped
Loading a procedure now also reaches the moment it is for. Loading the reporting procedure is a report; loading the release procedure is a delivery. - Bundled skills: each SKILL.md declares `metadata: moments:`. The same declaration ships as Skill defaults (BUNDLED_SKILL_MOMENTS), because the server never sees the plugin's files. test_skill_moments holds the two together and pins the plugin name that qualifies the skill. - Stored processes: `moments` on create_process and update_process, stored in the note's data and returned by get_process. A `scribe-proc-<slug>` load resolves its process through the sync manifest at load time. The moments are not copied into the stub, which would go stale mid-session. - reachable_tools lists the skill loader whenever anything is mounted, since a process's moments are known only when it loads. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
b3b616b20a |
fix(moments): the tools import the delivery module, and the tool-list cache leaves the swept directory (milestone 458 step 4a, #4922)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / integration (push) Successful in 49s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m48s
CI & Build / Build & push image (push) Successful in 32s
Two guards caught
|
||
|
|
78653130d6 |
feat(moments): mounted rules arrive when their moment happens, through every door (milestone 458 step 4a, #4922)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 57s
CI & Build / Python tests (push) Failing after 1m19s
CI & Build / Build & push image (push) Skipped
A rule mounted on a moment now reaches the session when an act reaches
that moment, with no semantic match involved:
- run_moment_arm on the pipeline: a lookup, not a ranked search. Each
line names the moment and the act that reached it ("at work.deliver,
reached by `git push`"), so a misfire is visible where it lands and
can be unmapped in-session. A repeat is cited, not quoted; fresh
rules are recorded surfaced under source moment_rule with the moment
in detail. No retrieval_logs row, as for the other lookups, so no
latency is persisted for this arm.
- rule_scope: a rule's home clause, moved out of semantic_search_rules
so the moment lookup scopes by the same one.
- rulebooks.rules_on_moments / mounted_moments.
- The plugin door: a catch-all PreToolUse hook (scribe_moment.sh). It
keeps /moment-tools' answer on disk for five minutes, so a call to a
tool that cannot reach a mounted rule sends nothing, and an install
that has mounted nothing sends one request per window. It shares the
rules ledger with the other arms and fails open silently.
- The MCP door: Scribe's own tools named by the shipped mappings carry
moment_rules in their response, so a client without the plugin gets
them too. The hook skips those tools. A guard pins the attach on
every one.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
cc26054437 |
feat(moments): rules mount on moments, through every rule door (milestone 458 step 3, #4921)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m0s
CI & Build / Python tests (push) Successful in 1m51s
CI & Build / Build & push image (push) Successful in 29s
rule_moments (migration 0117) records which moments a rule arrives at, by catalog name, cascading with the rule. rule_detail, the one seam every rule door already returns through, gains moments beside system_ids: None leaves the mounts alone, a list replaces them. get_rule and both list_rules doors read them back, batched per page. All five MCP rule/preference writes and the three REST ones take moments and validate them before their create or update. An unknown name is refused with the catalog listed and leaves no half-made rule behind; a parity test pins that ordering on every door. Backup v22 carries the mounts as a join table remapped through the rule map; a real-Postgres round trip checks they land on the restored rule. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
cf3de5bae1 |
feat(moments): actions map onto moments, with in-session corrections (milestone 458 step 2, #4920)
CI & Build / Plugin hooks (push) Successful in 18s
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m23s
CI & Build / Python tests (push) Successful in 2m0s
CI & Build / Build & push image (push) Successful in 36s
moment_actions.resolve(tool, input) names every moment a call reaches and the action that reached it. One call can reach several: kubectl apply is a run, a deliver and a reach outside the workspace. Command tools match by how each segment of the line starts, with a word boundary; other tools by field=value arguments. The MCP server prefix and case are ignored. 56 shipped defaults cover the harness tools, Scribe tools and common command shapes. moment_mappings (migration 0116) holds what an install adds and the defaults it switches off. A removal is a stored row, so an upgrade does not switch the default back on. Per the operator ruling, corrections happen in the session: map_action and unmap_action (write tools) return now_reaches so the fix can be confirmed in the same reply. list_moments now shows each moment's actions on this install. REST mirrors both doors, recorded as human. Backup v21 carries the mappings. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
2ff7f2f34f |
feat(moments): the moment catalog rules will mount on, readable in-session (milestone 458 step 1, #4919)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 58s
CI & Build / Python tests (push) Successful in 1m53s
CI & Build / Build & push image (push) Successful in 32s
Fourteen generic moments of work (session.start, work.start … reply.ask) plus the skill.<name> family, each with what it means and the kinds of action that reach it, written for any kind of work rather than software alone. The catalog is code because every install needs the same mount points; which actions reach a moment is per-install data (step 2). require_moment refuses an unknown name with the catalog listed, so a typo cannot become a mount that never fires. list_moments (read-only) and GET /api/retrieval/moments hand out the same catalog. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
d2ac7bf220 |
fix(lessons): a lesson's name is one claim on one line (#4797)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m48s
CI & Build / Build & push image (push) Successful in 28s
#4797 "A lesson's what takes a whole narrative, and the narrative becomes its name". `what` is the title every listing and menu prints. Nothing enforced its documented "one line", so lessons written with the incident in `what` printed up to ~1,500 characters as a menu line, buried the claim, and diluted the trigger in the embedded title. That was 14 of the 39 lessons on this install. - lessons.require_claim refuses a `what` over WHAT_MAX_CHARS (240) or running over several lines. The refusal says the story goes in `insight`. - Both doors call it before writing: - MCP create_lesson / update_lesson; - REST create / update. An update checks only a NEW name, so a lesson stored with a long one can still take the edit that repairs it. - lessons.claim_line is the display half. _menu_name shows an over-long stored name as its first sentence marked " …", because a door guard does not undo rows already stored and an unrepaired install would keep printing them. - The tool docstrings state the bound. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
9909cd2450 |
fix(retrieval): the rules never-opened warning stops sending the reader to a review tool that refuses rules (#4798)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m7s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 26s
#4798 "The rules corpus's surfaced_never_pulled warning sends the reader to menus_to_review, which can only review auto_inject". #4772 gave the warning one remedy for both corpora: "judge a sample with menus_to_review". That is right for notes, where a line carries its passage. For rules it is a dead end: retrieval_review.REVIEWABLE holds only auto_inject, so the tool refuses every rule arm. - The reading is now per corpus (_NEVER_PULLED_READING). For rules, the text says: - the count includes rules that only arrived in a listing; - a rule can rightly be set aside on its trigger alone; - no judged sample exists for the rule arms; - a rule set aside again and again is a trigger to fix (update_rule when_to_apply, then what_might_apply), not a floor. - The tool docstring says the same. - The guard is tied to REVIEWABLE, so it can fail in both directions: if the rules text names menus_to_review while no rule arm is reviewable, or if a rule arm becomes reviewable and the text still says there is no sample. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
5c6d9a7b55 |
fix(mcp): a tool's refusal reaches the agent with its reason (#4794)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 1m52s
CI & Build / Build & push image (push) Successful in 27s
SDK 2.x passes on only a ToolError's text. Any other exception becomes
UnexpectedToolError("Error executing tool X") and its message stays on the
server. ValueError is how every Scribe tool refuses — what was refused, why,
what to do instead — so every refusal reached the agent bare, and it could
only retry blind. Seen live: tune_retrieval(actor="operator") and
judge_menu(verdicts=[]) both answered "Error executing tool …" and nothing
else.
StrictArgsMCPServer.call_tool re-raises an UnexpectedToolError caused by a
ValueError as a ToolError carrying the message, in the SDK's own
"Error executing tool X: <reason>" shape. Other exceptions are crashes and
stay masked. The stale comment claiming the SDK returns ValueError text is
corrected.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
5d0e97576b |
fix(retrieval): the review drops records written after the call it re-runs (#4773)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m8s
CI & Build / Python tests (push) Successful in 1m53s
CI & Build / Build & push image (push) Successful in 36s
The first live sample ranked records the call could never have been offered: the session that made a call writes the decision, often quoting the message, and that record tops the re-run. Judged, it inflates on_point exactly where the budget is decided. - _rerun over-fetches by POSTDATED_SLACK, drops records created after the call before ranking, and names them in `postdated`. - each line carries `changed_since_call` (updated_at or a work log after the call) and `logged` (shown fresh then). - judge_menu shares the re-run, so a post-dated record cannot be judged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
6598c7fa85 |
feat(retrieval): a review pass judges whether injected lines related — menus_to_review, judge_menu and a judged readout (#4772)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 44s
An open rate cannot say whether a menu line related: every line carries its matched passage (#4364), so "not opened" covers unrelated, enough as shown, and already in context. #4772 "Injected notes are never judged". - retrieval_judgments (0115): a reviewer verdict per line of a logged call, on_point / adjacent / unrelated, with its reason, rank, budget side and whether the agent opened it within the hour. - menus_to_review re-runs a random sample of unjudged auto_inject calls with the arm's own parameters, past its budget, passage on every line. judge_menu records verdicts, re-deriving rank from a fresh re-run. - retrieval_telemetry gains a judged block (by rank, within/beyond budget, on_point_unopened). surfaced_never_pulled stops blaming titles. - missed-retrieval guidance names the review before a budget move. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
0a1bb68808 |
feat(rulings): system_usage_events is read back — per-System counts and a telemetry block (#4769)
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Python tests (push) Successful in 1m44s
CI & Build / Build & push image (push) Successful in 36s
#4769 "Rulings are counted where someone will read them": milestone 444 step 4 wrote system_usage_events and nothing read it. - retrieval_telemetry gains a `system_usage` block: surfacings and opens by source, distinct counts, and `by_system` naming the areas most shown. There is deliberately no pull-through ratio, because rulings travel in full in the line and opens are the exception. - usage_for_systems (one GROUP BY) adds `usage` to the REST Systems list and detail, and to MCP get_system. MCP list_systems is unchanged. - The Systems UI shows a "rulings shown N×" chip. - rulings_pre_tool, rulings_write_path and mcp_get_system are now declared registry points; the registry guard covers their recorders. - The Systems store merges a PATCH reply instead of replacing the row. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
556872c039 |
feat(rulings): a command or edit touching an area's files shows its rulings, once per session (milestone 444 step 4, #4757)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m6s
CI & Build / Python tests (push) Failing after 1m22s
CI & Build / Build & push image (push) Skipped
A System's rulings (the Rulings section of its description) now reach the work by path, not by similarity. Both PreToolUse arms resolve the files a command or edit names to the Systems whose path_patterns cover them, and the first touch in a session shows each area's rulings in one line; a repeat is a one-line reference. A lookup, so no floor, no budget, no retrieval_logs row. - services/system_rulings: parse_rulings, command_paths (reads and writes, relative to the repo root from any cwd; flags, URLs, globs skipped), rulings_for_paths - /tool-rules takes root, cwd and seen_ruling_systems; /prior-art takes seen_ruling_systems; both return ruling_system_ids - hooks share <sid>.rulings.ids (cleared on compaction by the ledger naming convention); the Bash hook sends the repo root and cwd - system_usage_events (migration 0114): surfacings by source, pulls from get_system; carried by backup (v20) through the system map - writing-records: rulings also arrive when the area's files are touched Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
a113c72b4f |
feat(systems): a System names its files — path patterns stored, validated and matched (milestone 444 step 3, #4756)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m6s
CI & Build / Python tests (push) Successful in 1m50s
CI & Build / Build & push image (push) Successful in 43s
A System gains path_patterns: globs relative to the repo root (* within one directory, ** across any depth, a plain directory covering everything under it). One service validates them for every door, so the web UI and MCP refuse the same bad pattern with the same message. systems_for_paths resolves paths to every active System that covers them, which step 4 (#4757) uses to deliver an area's rulings when its files are touched. - schema: systems.path_patterns JSONB NOT NULL default [] (migration 0113) - service: normalize_path_patterns, path_matches, systems_for_paths - routes + MCP create_system/update_system accept it; [] clears - web UI: a Files field in the create and edit forms, patterns on the card - backup carries it through export and restore - using-scribe reflex 7: tagging work keeps a System's files current Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
01d8a0b9f1 |
feat(guidance): an operator's ruling lives on the System it governs, and code is read as behaviour, not intent (milestone 444 steps 1-2, #4754 #4755)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m46s
CI & Build / Build & push image (push) Successful in 56s
A Librarian session contradicted a decision the operator had made 13 days
earlier. The ruling ("retry, then replace, never give up on a book") was
kept only as a quote in a work log, beside a session's reading of it that
capped replacements at 3. Three later sessions built on the reading, and one
carried the cap into an option as a "known cost", which the operator then
approved without being asked about it.
- writing-records.md: "A ruling goes on the System it governs". What a
ruling is (the operator decided it; a later change could undo it), how it
differs from a rule, and where it goes: a Rulings section at the end of
the System description, one line each with who, when and the source
record. Written the turn the operator decides; holds what is in force,
not history; a charter line that contradicts a ruling is fixed in the
same edit.
- using-scribe SKILL.md: reflex 1 says code tells you what a thing does,
not what was wanted, and a limit read from code is unconfirmed until a
System's Rulings says otherwise. Reflex 9 points to the ruling section.
- reporting-back: an option that carries existing behaviour says whose call
it was (the operator's ruling, or a past session's never confirmed); one
that contradicts a ruling is a Conflict.
- create_system / update_system docstrings: the Rulings section, and that
description replaces the whole text.
- test_guidance_ownership: three topics pinned to their owners.
- Plugin version minted.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
6d3dca0af5 |
feat(lessons): convergence is named at the write — no-rule lessons that keep landing in one situation suggest a rule (milestone 440 step 5, #4634)
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / Python lint (push) Successful in 2s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Build & push image (push) Successful in 27s
When a lesson is answered "no rule fits" (create_lesson / update_lesson on both doors), the response looks for other no-rule lessons it resembles and, once there are CONVERGENCE_LESSONS (3) of them, carries `convergence`: the members, their incidents and projects, and a hint to draft the missing rule with create_rule (operator approval as always) and point each lesson at it — or to leave them as lessons when no single choice is right every time. - convergence_group is the pure bar: distinct LESSONS count, incidents never stand in for them (one broad lesson cannot trigger it), and a group whose sources all point at one incident is one event written up several times. - convergence_for searches lessons by the new one's claim + trigger (trigger_title) at CONVERGENCE_THRESHOLD 0.65 — above the menu's "worth showing", below the duplicate gate's "same record" — then keeps the ones with a lesson_no_rule answer. Fail-open. No sweep, no timer (#4183). - Defaults stated as defaults (rules 32, 115). - Tests: the bar (pure), the search with stubs, the door, and the no-rule filter against Postgres; conftest stubs convergence_for for unit tests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f7d8dc2e55 |
feat(lessons): judged when written — a new lesson is offered its rules, and "no rule fits" is an answer (milestone 440 step 2, #4631)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 53s
CI & Build / Python tests (push) Failing after 1m15s
CI & Build / Build & push image (push) Skipped
"Which rule is this lesson an instance of?" now has three recorded answers: a rule named (a confirmed link, #4630), no rule fits (new), or unjudged. - Model + migration 0112: lesson_no_rule (lesson_id PK, CASCADE from the note; why; judged_at). A table rather than a key in notes.data, because that mirror is re-composed from the body on every edit and would erase it. - Service (lesson_rules): set_no_rule rejects any confirmed link with the reason; a confirmation (set_lesson_rules or judge_link) deletes the answer; require_one_answer refuses both answers in one call before any write; judgments_for_lessons + attach_lesson_rules add rule_judgment (and no_rule) to every lesson payload; list_unjudged lists the open ones; rule_candidates searches rules with the lesson's claim + trigger at the explicit-search bar, None when the search could not run. - MCP: create_lesson/update_lesson take no_rule; an unanswered create returns rule_candidates, rule_judgment and a rule_hint; list_lessons(unjudged=true). - REST: the same on POST/PATCH /api/lessons and GET ?unjudged=1; create returns rule_candidates. - Backup v19: a lesson_no_rule section, export (full and per-user) and import. - Guidance: create_lesson docstring, writing-records.md in using-scribe (owner, pinned in test_guidance_ownership), create_rule docstring on linking the lessons a new rule governs. Plugin version minted. - Tests: door units, integration for the three states, the rejection reason, scoping, cascade; backup registries. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
c8393975c3 |
fix(mcp): classify judge_lesson_link as a write tool (#4630)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m42s
CI & Build / Build & push image (push) Successful in 31s
The tool-classification guard (test_mcp_auth) failed CI run 705: the new judge_lesson_link tool sat in no set, so a read key would have been silently denied it. It changes a link's state, so it belongs in _WRITE_TOOLS. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
41e4fbaba1 |
feat(lessons): a lesson names the rule it is an instance of — lesson_rule_links (milestone 440 step 1, #4630)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Failing after 1m13s
CI & Build / Build & push image (push) Skipped
The link between a lesson (one concrete situation) and the rule that governs it, with the operator's soft-then-hard design built into its state: suggested while evidence accumulates, confirmed or rejected once judged. Only confirmed will carry a rule in retrieval (#4633); rejected is kept so the pair is never proposed again. - models/lesson_rule_link.py + migration 0111: one row per (lesson, rule), CASCADE on both ends, indexed both ways, CHECK on state (rule 36), evidence JSONB and judged_at. - services/lesson_rules.py: require_rules (validated before any write, so a bad id leaves nothing half-linked), set_lesson_rules (set-semantics; a dropped rule becomes rejected, not forgotten), judge_link, and the two reads. ACL: write on the lesson (share-aware), ownership of the rule; a reader sees only rules they own. Decorations are fail-open (#4286). - MCP: create_lesson / update_lesson take rule_ids; get/create/update return `rules`; new judge_lesson_link tool. REST: the same on /api/lessons plus PUT /api/lessons/<id>/rules/<rule_id>. Rules: rule_detail carries `lessons`. - Backup v18: export (full and user-scoped, both ends in scope), builder, importer; both column guards register the table. - Tests: integration (states, set-semantics, judge, ACL all-or-nothing, cascade both ways, CHECK, one row per pair); unit (door wiring, judge registered, migration/model state agreement, backup skip and unjudged stays unjudged). conftest stubs the decorations for unit tests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
582a5a4f48 |
feat(shapes): the practice is written where it is read, and the coverage line measures the slip (milestone 439 step 6)
CI & Build / integration (push) Successful in 47s
CI & Build / Python tests (push) Successful in 1m39s
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Build & push image (push) Successful in 29s
- reusing-code: "Before the turn ends — say what you built" — the four verdicts, the one classify_shapes(repo=…) call, and the component file as a shape. Description names the end-of-turn moment. - shape-accounting: the writer judges; audits are the check that it held. The write path SUGGESTS (no more hook instances); component file rows and whole-file canon described; scoped covers Svelte too. - _INSTRUCTIONS reuse line: "before the turn ends, say what you built (create_snippet the reusable, classify_shapes the rest)" — 1570/1600. - Coverage line: "written-shape check (7d): N turns checked, M asked, K left unjudged", from the Stop hook's recorded outcomes; silent until the question has been put. - test_guidance_ownership pins the new topic on reusing-code. - prior-art hook header no longer says it stamps instance rows. Plugin version minted. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f1fbdf746a |
feat(shapes): the agent judges what it wrote, at the end of the turn (milestone 439 steps 1-3)
CI & Build / Python lint (push) Successful in 5s
CI & Build / Plugin hooks (push) Successful in 18s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m38s
CI & Build / Build & push image (push) Successful in 33s
Recording used to be decided by machinery — the only "record it" prompt fired when a same-named copy already existed (#2664), so a first instance of a reusable piece was never asked about, and judgment arrived only through audits. Now the question is asked where the knowledge is: the end of the turn that wrote the code, of the agent that wrote it. - Write hooks keep `<sid>.written.ids` (path, kind, name) for every definition a write names; a new file adds a `file` line for its stem — a candidate in any language without a framework rule (scribe_written_append). - Stop hook scribe_shape_check.sh sends the ledger to GET /api/plugin/shape-check and blocks once, in the server's words, when anything is unjudged. Same discipline as the report check: never twice, never without a recorded check, another hook's loop left alone; the ledger is kept when the instance cannot be reached. - shape_ledger.unjudged_shapes: no row, unclassified, scoped and hook stamps are unjudged; an agent/audit/import verdict is not. A snippet recorded at the shape answers for it until the refresh stamps it canonical. - services/shape_check owns the reason text and records every outcome in app_logs (passed / blocked / judged_after_block / left_after_block). - classify_shapes(repo=…) judges a shape the ledger has not synced yet via a provisional row under a bound repo; the sync confirms it, or vanishes and revives it with the verdict intact. An unbound repo is refused. Plugin version minted. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
4fb53b844d |
refactor(mcp): _INSTRUCTIONS orients the workflow, not a rulebook (#4389)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Failing after 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m35s
CI & Build / Build & push image (push) Successful in 23s
Spike #4389 read the spec, Claude's docs and a dozen servers: the field is for how the tools fit together, and the field runs ~600-1,600 characters. Ours sat at the 2,048 cap as a keyword index that also carried stance. - JUDGE, REPORT and MISSED leave the index. They fire mid-work, not at session start; using-scribe and reporting-back state them in full, and `placement`/`report_back` cue reporting in-band. No skill text changes. - The rest is rewritten as plain practices (1,503 chars) and keeps every session-start marker the ownership registry pins. - INSTRUCTIONS_BUDGET 2000 -> 1600; the three index markers are dropped from the registry; the miss-route index test now checks that the index keeps what_might_apply and stays off the route. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
c26b7f248e |
feat(mcp): port the server to MCP Python SDK v2 (FastMCP → MCPServer) (#2196)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m37s
CI & Build / Build & push image (push) Successful in 31s
- server.py: `mcp.server.mcpserver.MCPServer`; StrictArgsFastMCP becomes StrictArgsMCPServer, whose call_tool takes and forwards v2's `context` and raises ToolError (the SDK logs anything else as an unexpected crash; the message reaches the caller either way). - stateless_http and transport_security moved from the constructor to `streamable_http_app(...)` in mount_mcp, with their reasons. - Per-request user identity is unchanged: the contextvar set around the ASGI call reaches the handler on both v2 paths (legacy stateless spawns from the request task; the 2026-07-28 modern path opens its task group inside the request). - pyproject: mcp[cli]>=2.2, no ceiling (installs are --locked). uv.lock regenerated with --upgrade-package mcp in the ci-python image: only mcp and its own dependencies moved (106 → 110 packages). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
952e56ee75 |
feat(tasks): the hand-off — SessionEnd releases a session's claims; the practice is written down (milestone 381 step 4)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 1m40s
CI & Build / Build & push image (push) Successful in 24s
Release is the mechanical half: scribe_session_end.sh sends the ending
session's id to /api/plugin/release-session, which releases the claims it
held. A tidy-up, not the guarantee (no SessionEnd on a crash; the lease
covers that), and skipped on /clear so SessionStart(clear) can still
hand the claimed work back.
Saying what happened is the half only the model can do. It is stated as a
practice where it is read: the using-scribe skill owns it ("Hand off before
this session's context stops existing", pinned in test_guidance_ownership),
the static context points at it for the wrap-up moment, and add_task_log's
docstring says a log claims the task. _INSTRUCTIONS is untouched.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
fc1c463641 |
feat(dedup): a note or task blocks only as a copy; a close match is surfaced for judgement (#4306)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / TypeScript typecheck (push) Successful in 58s
CI & Build / integration (push) Successful in 1m12s
CI & Build / Python tests (push) Failing after 1m29s
CI & Build / Build & push image (push) Skipped
Measured on the live corpus, the 74 note pairs at or above the old 0.90 bar were almost all distinct siblings — consecutive dev-logs, sub-notes of one design, research parts — and the one clear copy sat at 0.997. The block refused the next dev-log and taught force=true, as #4134 found for rules. - The semantic arm blocks notes and tasks only at >= 0.98. The title block stays; processes keep 0.90 (not measured). - 0.87 to 0.98 comes back as `overlaps` on the create reply, from the same per-chunk searches, with a note that leaves the call to the session: fold in and delete if it is the same record, keep both if a sibling. - create_note, create_task, create_records and start_planning's steps all carry it; a batch names the record each overlap belongs to. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
4ae18a9dd9 |
feat(snippets): a snippet has notes; when_to_use is the situation it is ranked on (#4378)
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / Python tests (push) Successful in 1m37s
CI & Build / Python lint (push) Successful in 3s
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Build & push image (push) Successful in 34s
A snippet had no field for prose, so what a session learned about one went into when_to_use — the trigger joined onto every chunk it is embedded as. A sweep found write-ups of up to 3 KB there, headings and all. - notes: stored after the code under `## Notes`, parsed back from the body, carried by every path that rebuilds it (update, merge, un-merge). A snippet with no notes composes the body it always did. - create/update_snippet (MCP) take notes and return trigger_advice when when_to_use is long, headed or multi-paragraph. Advice, not a refusal. - Tool docs, the reusing-code skill and the editor hint describe the trigger as the situation and point the explanation at notes. - Editor gains a Notes field; the detail view renders it as markdown. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
66e21a6c60 |
refactor(notes): a snippet's and lesson's stored title is its name; the trigger joins it only in the embedded document (milestone 427)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / integration (push) Successful in 52s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m35s
CI & Build / Build & push image (push) Successful in 32s
The title was `subject — trigger` because the stored title WAS the embedded one, and the join is what makes these kinds rank on the situation they apply to (#2485). Every surface that shows a title then showed the trigger too -- menus, lists and search rows ran to kilobytes. - embeddings.document_title(title, note_type, data, body) joins the trigger from `data` (body fallback) at embed time. Idempotent: an un-migrated composed title comes out the same, never doubled. The embed path, the startup backfill and the dedup gate's semantic signal all use it, so the embedded text -- and every vector -- is unchanged. - Writers store the subject: snippet create/update (service, REST, MCP) and lesson_document. Both compose_title helpers are removed. - Readers: dedup takes `data`; the menus strip the embedded title from a passage; list rows project `when_to_use`, which SnippetListView reads. - 0108 rewrites existing rows on an exact `' — ' || <own trigger>` suffix with raw SQL, leaving updated_at alone so the backfill does not re-embed the corpus for identical vectors. Downgrade recomposes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
108b12eeb0 |
fix(dedup): a rule or preference create surfaces what it overlaps by meaning (#4134)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / Python tests (push) Failing after 1m16s
CI & Build / Build & push image (push) Skipped
CI & Build / TypeScript typecheck (push) Successful in 56s
CI & Build / integration (push) Successful in 1m4s
find_duplicate_rule was title-only, on the stated premise that rules are not a semantic-retrieval surface - false since rules were embedded. A preference restating a rule under another title passed untouched, and since both kinds share one ranking, the weaker label could arrive alone. find_overlapping_rules queries semantic_search_rules with the rule_document shape, both kinds, in the scope the new record ranks in (global: every rule the caller owns; project: global + that project). All three MCP create doors call it before creating and return overlaps + overlap_note on the reply. It advises rather than blocks, on measurement: across 16 sampled records the nearest DISTINCT neighbour reached 0.853, while a true rewording scored 0.850. No threshold separates the bands, so the floor (0.80) sits below the restatement and the author judges. The stale docstring is corrected. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy |
||
|
|
62f3a485ad |
fix(usage): one seam attaches the surfaced-vs-opened chip, and the Knowledge browse uses it (#4230)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 52s
CI & Build / Python tests (push) Failing after 1m2s
CI & Build / Build & push image (push) Skipped
`usage_for_notes` is named for notes and works on every note row, yet the chip reached snippets and rules only. Notes had it nowhere. Lessons had it collected and shown nowhere a person could reach, because #4196 taught `/api/lessons` to attach it and `KnowledgeView` — the only lesson list in the UI — browses through `/api/knowledge`, so `listLessons` still has no consumer. The cause was not a missing line. SEVEN call sites carried their own copy of the same few lines: two REST lists, two REST details, two MCP lists, one MCP detail. Each read perfectly well alone, so "which doors attach usage?" had no answer anywhere in the code — the same asymmetry test_system_tagging_door_parity.py records for System tagging (#4249), where whichever door nobody exercised for a kind is the one that never grew the feature. `attach_usage(rows, key="id")` is now that answer, and all seven go through it. A detail payload is a one-row list, so the single-record doors share the seam rather than keeping a second shape beside it. Deliberately NO try/except: the fail-open already lives in `usage_for_notes`, which reports through `_report_failure("readout")` and returns the zero-filled map. Wrapping it again would swallow the REPORT as well as the error, and a silently-swallowed readout failure is exactly #2663 — every counter reading zero in production for weeks while the writes landed fine. `/api/knowledge` now attaches usage, which closes both holes at once: it is how notes, lessons and processes are all browsed. `KnowledgeView` renders the badge on the card footer, looking the advice up per row because the feed is mixed. The advice moves to utils/deadWeight.ts. Canon #3460 says each caller owns its own const, and that held while each caller showed ONE kind; a mixed feed would need five of its own and the next surface another five. The canon's actual invariant — advice is kind-specific and never baked into the badge — is kept: it is still a prop. The three existing callers now read the same table, so the sentence has one home rather than four. Recorded against #3460 so the next reader is not left re-litigating it. `_row_id` rejects bools explicitly: `int(True)` is 1, so a row carrying a flag under the key would be credited with note #1's counts, and a wrong chip is worse than no chip because it reads as a measurement. A row with no usable id is skipped rather than failing the page. Tests pin the PROPERTY, not one route: no door calls the aggregate directly (AST, so a comment naming it is not a false positive), and every door that shows usage reaches the seam. Plus the N+1 guard — one aggregate per page, asserted on await_count, because the per-row version reads more naturally and is invisible in review. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy |
||
|
|
43ed7f4db7 |
fix(access): tag a record as the user doing it, not as its owner (#4249)
CI & Build / Python lint (push) Successful in 3s
CI & Build / integration (push) Successful in 1m4s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Build & push image (push) Successful in 26s
`set_record_systems` is not a dumb setter. It runs its own `can_write_note` and then links only the Systems the given user can READ. Four of twenty call sites handed it the record's owner instead of the acting user, which did two quiet things at once: the access check became trivially true, since an owner can always write their own record, and the System filter used the owner's visibility rather than the actor's. On a single-user install neither is observable. With a share it is an editor acting with the owner's reach — the shape rule 47 exists to prevent, and the same reasoning routes/notes.py already spells out for `set_supersedes` two lines away. This is NOT a permission change. Every one of the four sites establishes the caller's write access first: routes/lessons.py and routes/snippets.py call `can_write_note(uid, …)`, mcp/tools/processes.py does the same, and mcp/tools/snippets.py reaches `set_record_systems` only after `update_snippet` has raised PermissionError if the caller may not write. So nobody gains or loses the ability to edit anything. What changes is whose reach the tagging runs with, which is exactly the kind of difference that survives review because every call site reads fine on its own. The four: routes/lessons.py:243 owner_uid -> uid routes/snippets.py:211 owner_uid -> uid mcp/tools/snippets.py:486 note.user_id -> uid mcp/tools/processes.py:196 note.user_id -> uid The last one was written earlier in this same session, an hour before the sweep that found it, with a comment confidently explaining why the owner was correct. That is the argument for the guard rather than for care: the unified stance was known and still got it wrong at the next opportunity. So the guard is the point again. `test_every_tagging_write_acts_as_the_caller` walks every `set_record_systems` call in src/ and asserts the first argument is a bare local named `uid` or `user_id` — an attribute access is a record's owner by construction. Verified against `git show HEAD:` as well as the working tree: clean now, four offenders on the code it replaces. Reads are deliberately untouched. `list_record_systems(owner_uid, …)` gates on reading the NOTE, which the caller can do anyway, so it returns the same list either way; it is a different operation and churning it would add noise without changing behaviour. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy |
||
|
|
e2e1ea3667 |
fix(systems): a record can be filed under a System from whichever door wrote it (#4249)
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Python tests (push) Successful in 1m44s
CI & Build / Build & push image (push) Successful in 29s
A System tag is how `list_system_records` gathers an area's pile, so a
record that cannot be tagged is reachable by search and by nothing else.
`update_lesson` did not take `system_ids` while `create_lesson` did, which
made tagging available exactly once — at the moment of least information.
A lesson is usually written at the end of a piece of work, which is
precisely when that argument gets dropped, and after that the record could
never be filed at all.
Filling in the rest of the table found three more gaps, and the issue's own
generalisation was wrong. It read as "the REST door can do something the
MCP door cannot", on three instances in a row. In fact:
* `create_preference` is the exact MIRROR of the lesson bug — update
takes `system_ids`, create does not. The same capability missing from
the opposite end of the same lifecycle.
* `routes/notes.py` handled `system_ids` NOWHERE, while the MCP door
handled both ends. That runs the opposite way round from the premise.
* `create_process` / `update_process` took it at neither door, though a
process is a note and has always been taggable in the data model.
The real pattern is that whichever door nobody exercised for a kind is the
one that never grew the parameter — which is a better statement of #4248
than the one recorded there, and is not something a reviewer reliably
notices, because each door is only ever read on its own.
Milestones are NOT a fifth gap. `RecordSystem.note_id` is a ForeignKey to
`notes.id` and milestones are their own table, so they cannot be tagged at
any door by construction. Pinned in the test so the next pass does not
re-open it.
So the guard is the point, not the four parameters. `update_lesson` alone
would have left the shape that produced it intact. The new test asserts the
TABLE — every kind taggable anywhere is taggable everywhere it is written —
and keys the registry-coverage check on the SIGNATURE rather than on a
grep, so a module that only names the argument in prose is not swept in and
no hand-kept skip list can go stale. A fifth kind fails there rather than
shipping half-wired, the same reasoning test_derived_mirror_generic_door.py
records for derived mirrors (#3734).
One inconsistency found and deliberately not changed here: for the same
operation `routes/tasks.py` scopes `set_record_systems` by the caller while
`routes/lessons.py` and `routes/snippets.py` scope it by the owner. The new
notes code follows tasks.py and says why in a comment (#47 — an editor-share
holder should tag from what they can see rather than inherit the owner's
reach). Recorded in #4249 rather than fixed as a drive-by.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
|
||
|
|
1fca8c2808 |
feat(retrieval): a System's charter becomes an answer, not just a filter (#4251)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 58s
CI & Build / Python tests (push) Successful in 1m38s
CI & Build / Build & push image (push) Successful in 27s
Step 2 of #4251. A System's `description` is a charter — several hundred words saying what belongs in that area and what does not — and it is the answer to "which part of this codebase does X live in". There was no semantic path to one: `list_systems` enumerates, and `search(system_id=…)` uses a System as a FILTER over notes. So a System could narrow a search and could never be the answer to one, and an agent asking where a record belonged had to read every charter or guess. ITS OWN SEARCH, not a `content_type` over notes, for the reason note 3163 gives about milestones: the row could be shared, the search cannot. A charter competing with the whole note corpus for one top-k is outranked by the records filed under it — the right answer crowded out by its own contents — and "where does this belong?" is a different question from "what prior art is there?", which a caller asking one should not have to read past answers to. So `system_embeddings` (0107) joins note_, rule_ and milestone_embeddings as the fourth sibling, with `system_document`, `upsert_system_embedding`, `semantic_search_systems`, a startup backfill and `search(content_type= "system")`. Scoped like milestones: with a project_id, that project's Systems if the caller can read the project (rule 78); without one, the caller's own. Archived Systems are excluded — an archived area is one the operator has said is no longer where things go, which is exactly the question being asked. `system_document` is the plainest of the four shapes on purpose. A charter is already written as the thing this search has to match, in the words someone asking would use — so there is no trigger to synthesise as `rule_document` must, and no second record to gather as `task_document` must. The stored charter IS the sharp document, the way a snippet's is. `color`, `status` and `order_index` stay out: presentation and bookkeeping, and a vector carrying them would be answering a question nobody asks of a charter. The search publishes `report["best_chunk"]` from the start rather than being retrofitted, which is what #4251 asked of any fourth search. It matters more here than anywhere: a charter runs long and a result shows its NAME, so a match on the paragraph that actually decides where a record belongs would otherwise be previewed by two words that cannot say. The id that comes back is the one `system_id`, `system_ids` and `list_system_records` already take, so the answer to "where does this belong?" is directly usable as "show me what is there" and as "file it here". `embed_system` sits beside `notes.embed_note` at the service for #2056's reason — every door gets it by construction. Not called on delete: that is a soft delete and the search joins through `System`, so the vectors are already unreachable, and leaving them means a restore is findable again immediately. `system_embeddings` is declared in backup's `_NOT_INCLUDED` as derived, beside its three siblings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy |
||
|
|
5fb41af9b0 |
feat(search): the agent's search can ask for every kind the corpus has (#4250)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 58s
CI & Build / Python tests (push) Successful in 1m42s
CI & Build / Build & push image (push) Successful in 28s
The engine took `note_type` and `task_kind` all along. What was missing was a way to say them: the MCP tool's `content_type` knew `note`, `task` and `all`, and `/api/search` knew the same two — so an agent could not ask "has this snippet already been recorded" or "what lessons apply here" without searching everything and reading past the rest. Browse offered nine kinds from the same data. The cause is that each door kept its own map. `_FACETS` in services/knowledge is where a kind is declared, and #3161 made adding one a single edit by generating the SQL filter, the Python predicate and the door's validation from it — but the two search doors were written before that and never joined. So this adds the third dialect, `search_filters_for`, and one composition over it, `content_type_filters`, and both doors now derive instead of listing. Two names keep a meaning of their own, and the docstrings say so: `all` is no filter, and `note` is BROAD — any non-task, snippets and lessons included — where the browse facet of the same name is narrow (`note_type == 'note'`). They are left different deliberately; narrowing this one would stop returning snippets to every caller that already asks this way. An unrecognised kind is now refused rather than answered. Both doors used to fall through: the MCP tool into a filter matching no row, the route into no filter at all, so `?content_type=snippets` returned the whole corpus while looking like a narrowed search. An empty result set is a claim — "the corpus holds nothing like this" — and an agent acts on that claim by building the thing it could not find, so a typo must not be able to make it. The docstring is the agent-facing contract (#2846), and a test now holds it to the table: every kind `_FACETS` declares has to appear in it, because a filter an agent has not been told about is unreachable however well it is wired. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy |
||
|
|
253fb974f3 |
feat(retrieval): every semantic search hands on the passage that matched
CI & Build / Python lint (push) Successful in 8s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Failing after 1m9s
CI & Build / Build & push image (push) Skipped
#4243 fixed one door. Scribe has three semantic searches over three chunk tables, and all three collapsed chunk rows to the best one per record — each of them KNEW which passage earned the hit, and each dropped it. Every surface downstream then previewed the head of the document instead: a span the search had already scored lower, with nothing saying so. Mechanism, one place: - embeddings.record_best_chunk publishes {id: {index, text}} into `report`. Carried in `report`, NOT the return value: all three return list[tuple[float, Record]] and ~30 sites unpack that pair (lesson #4207). - semantic_search_rules and semantic_search_milestones now select chunk_index/chunk_text and publish the winner, as notes already did. semantic_search_milestones gains `report`, which it had no way to take. - services/text.matched_excerpt is the one choice of span, and excerpt_fields the one result block. Doors keep their own field names — the web renders `snippet`, MCP returns `excerpt` — because renaming a field a frontend reads is a different change from fixing what goes in it. Surfaces: - knowledge.query_knowledge, whose own comment calls it "the human's MAIN search surface", was `(note.body or "")[:200]` on every row alike. Now the matched passage on a search, the opening on a browse, and `snippet_is` saying which. KnowledgeView renders that snippet, so this was live. - search(content_type='milestone') gains `matched` — the plan body stays out, but the passage that matched comes along, because recognising a plan means recognising the part you asked about and a description written at the start need not mention it. - The auto-inject menu and the write-path prior-art menu put the passage under their line. Both were title-only, which answers "does this apply?" for a lesson or snippet (the trigger is IN the title) and not at all for an issue or dev-log. No fallback to the body's opening: on a menu that is preamble dressed as a reason, and once indented it cannot be told apart. Left alone deliberately: the rule arms. A rule hint already renders the rule's TRIGGER, which is written to answer exactly "does this apply to me" and beats a matched chunk at it; and that line's budget was measured at #3851. Adding a passage there would duplicate the trigger and spend the budget twice. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy |
||
|
|
fdc07f2a2b |
fix(search): show the passage that matched, not the opening of the body (#4243)
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Failing after 1m5s
CI & Build / Build & push image (push) Skipped
CI & Build / integration (push) Successful in 46s
Raised by the operator: are we limiting what comes back by character count,
and how do we verify the pertinent part is the part displayed?
We were not. mcp/tools/search.py sent (note.body or "")[:240] — a head cut,
with no marker that anything had been removed, so a 240-character preview of
a 4000-character record was indistinguishable from a complete short one.
The opening is the wrong span. The match is semantic and per chunk, and
semantic_search_notes collapses to best-chunk-per-note — its own comment at
the collapse says "the first appearance of a note is its best chunk". So the
system identified the passage that earned the hit and then discarded it:
select(Note, distance) kept no chunk column. A record could rank first on its
sixth paragraph, be previewed by its first, and be judged irrelevant on a
span the search had already scored lower. That biases against long records,
and it is self-concealing — the caller who does not open it never learns the
preview was misleading.
- embeddings: chunk_index/chunk_text ride along in the select, and the
collapse records the winner in report["best_chunk"]. Carried in `report`,
NOT by widening the return tuple: ten callers unpack (score, note) at
~18 sites and nothing would catch the misses (lesson #4207). `report` is
the side-channel this function already uses for best_available_score.
- search(): excerpt / excerpt_is / body_length, and read_full when there is
more. A caller that cannot tell a matched passage from a document opening
cannot judge whether to look deeper, which is the only decision the field
supports.
elide() moves to services/text.py so both callers share one copy, and it
keeps BOTH ends with a stated gap — it is the fallback for when nothing
identifies a better span than "all of it", not the goal.
Also fixes a guard that produced a false failure on the previous commit:
test_pull_telemetry checked `"project_id: int = 0" in body.split("\n")[0]`,
which sees only the first line, so wrapping get_task's signature over four
lines made it report a function that does take the project as one that does
not. Parsed with ast now, and proven to still reject an absent or
wrongly-typed parameter rather than being appeased by reflowing the code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
|
||
|
|
4f2977b848 |
fix(tasks): add_task_log wrote to a surface no agent could read back (#4241)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m4s
CI & Build / Python tests (push) Failing after 1m14s
CI & Build / Build & push image (push) Skipped
The work log reached the web UI through routes/task_logs.py and nothing
else. get_task returned only the body — a claim written once, before the
work — with the record written during it invisible beside it. So a stale
body arrived with nothing to contradict it, and this session rebuilt work
that had already shipped, with the evidence sitting in the task's own logs.
Read side, scoped through the access layer (rule 78):
- logs_for_task / count_logs_for_task / log_counts_for_tasks in
services/task_logs.py. Scoped by who may read the TASK rather than by
who wrote the entry: list_logs filters TaskLog.user_id == user_id,
which hands a shared collaborator an empty list reading as "no work
has been done". The page query folds readable_notes_clause into the
same statement so the permission does not become an N+1.
- get_task returns work_log; list_tasks and get_milestone steps carry
log_count, zero-filled so "none" is a count and not a missing key.
Elision keeps both ends. The newest entry arrives whole to 4000 chars
because it answers "where does this stand"; older ones are shortened from
the MIDDLE, never the head. A head cut selects what a reader sees by
character position, which is uncorrelated with what matters — an entry
closing with "so this shipped in 04775c3" loses the one sentence that
answers the question, and a truncated flag says something went, never
whether it mattered. The gap states how many characters it covers.
conftest gains an autouse stub for the new read arm, same reasoning as
_no_rule_arm: three widely-called tools grew a database read, and the
existing call sites should not each have to learn about it.
Raised while reviewing this: search() has the same shape and worse —
body[:240] with no marker at all, while the chunk that actually matched
sits unused in the row that won. Filed as #4243, not fixed here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
|
||
|
|
f8e53c1c35 |
fix(telemetry): a warning fired on an arm whose decline rate is arithmetic, not evidence (#4232)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 49s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m33s
CI & Build / Build & push image (push) Successful in 26s
Found by reading a live `retrieval_telemetry` readout after milestone 419
deployed, not by inspection. The readout said:
cannot_decline / report_preference — "45 calls, 0 of them returned
nothing. An arm that fires unasked has to be able to say nothing; this one
never has. Check that it applies its floor at all."
And printed, beside it, that arm's band: p10 = p50 = p90 = min = max = 0.791.
FIVE IDENTICAL PERCENTILES IS THE TELL. That is not a ranking, it is one
record at one score on every call — because `report_preference` searches a
fixed string (`reply_preferences.COMPLETION_QUERY`, a module constant, and
deliberately so).
For a fixed query against a stable corpus the top score is a CONSTANT, so the
arm's decline rate is 0% or 100% and never in between; which of the two it is
depends only on where the bar sits relative to that one number. "Never
returned nothing" is therefore arithmetic, not evidence, and the warning's own
remedy — check whether it applies a floor — cannot be answered from it.
The arm already knew this about itself; the warning did not:
"a fixed query makes this arm's score a constant and a floor a hair above
it produces a dead arm no amount of traffic will ever reveal"
— services/reply_preferences.py
THIS CLASS OF BUG ALREADY HAS A GUARD, which is the argument for the shape of
the fix. `Point.logs_unconditionally` exists because of #3497: both rule arms
once logged only their hits, so their zero count was structurally 0 and this
same warning would have fired on a LOGGING property while sending the reader
to move a threshold that was never involved. This is that one step over — a
QUERY-SHAPE property — and gets the same treatment: a declared field on
`Point`, and exclusion rather than trust.
AND THE WARNING THAT WOULD BE INFORMATIVE HERE DID NOT EXIST. For a fixed-query
arm the dangerous state is the mirror image: every call empty, meaning the bar
is above the constant and no further traffic will ever move it. The arm is off
rather than quiet, and nothing in the readout said so — `expects_traffic`
covers an arm with NO calls, not one with calls and a 100% decline rate. That
state is real and reached: `report_preference` once logged 69 consecutive
declines at 0.0006 under its bar.
So `fixed_query_never_clears` sends the reader to `near_miss_samples` and not
to the dial — because that incident is also the one where the statistic and
the correct action pointed opposite ways. Every percentile said lower the
floor; opening the refused record showed it was rule 77 arriving as a false
positive, and lowering it would have delivered that rule on every completion
report ever written.
Guards in tests/test_retrieval_warnings.py, including the falsifier that
matters most here: `cannot_decline` must still fire on an arm whose query
varies, or this change is a disabled check wearing a narrowed one's clothes.
Both boundaries tested from both sides, per that module's own standard.
The new code is documented in the `retrieval_telemetry` tool docstring beside
the others (rule 33) — an undocumented code in a readout is a reader meeting a
verdict with no way to disagree with it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
|
||
|
|
edbc31f8ca |
feat(lessons): the kind whose whole question is "is this trigger right" was the one kind that could not see its own counts (#4196)
CI & Build / Python lint (push) Successful in 8s
CI & Build / Plugin hooks (push) Successful in 17s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 57s
CI & Build / Python tests (push) Successful in 1m41s
CI & Build / Build & push image (push) Successful in 36s
The lesson slot has recorded surfaced-vs-opened since it shipped. Nothing showed it. `get_lesson`'s REST door attached `usage` to the payload and no view rendered it; the listing did not attach it at all, and neither MCP door did. #4196 asks when a lesson that keeps getting followed should become a rule, and names the trap in the same breath: raw frequency cannot separate "this should bind" from "this trigger is too broad", and the second is the commoner reading by a wide margin. Surfaced-AND-opened can separate them. Neither question is answerable by a reader who cannot see the numbers, which is why this is the first step and not the threshold. NO THRESHOLD IS PROPOSED HERE, deliberately. The corpus today is 10 lessons with 7 recorded surfacings and 3 opens, over about fifteen hours of usage data. A promotion rule fitted to that would be fitting noise — #3311's failure, and the warning lesson #4228 was written to carry. `UsageBadge` already declines to render a verdict under three surfacings for the same reason. So #4196 stays open: its subject, the promotion path, is still unbuilt. What lands is the evidence it needs. - REST `GET /api/lessons` and MCP `list_lessons` attach `usage` to every row, from one aggregate per page rather than a per-row read, which would be N+1 by construction. Every row carries the key zero-filled, so "never surfaced" is a state a reader can see rather than a missing field they have to interpret. - MCP `get_lesson` attaches it too, and reads it BEFORE recording its own pull. That door records a pull on every open — it has to, or the kind sits permanently at zero — which makes the order load-bearing in a way it is not for a kind that only counts. The REST detail door already ordered it this way; the two now agree about what the number means. - `LessonDetailView` renders `UsageBadge` (snippet #3460) rather than re-spelling the chip, with the advice keyed to this kind: a lesson that is repeatedly offered and never opened is usually keyed to a situation nobody is in, so it points at re-keying `when_to_apply`, not at deleting the claim. Guards, in the two styles this pair of doors already uses: the MCP side driven behaviourally through mocks, including the call ORDER for `get_lesson`; the REST side on structure like its siblings in test_lesson_rest_door.py, because the route is decorated and returns a Quart response. Rule 167's falsifier is included. KNOWN GAP, not fixed here: `KnowledgeView` is the only lesson LIST in the UI and it reads `/knowledge`, not `/lessons` — so the REST listing change reaches `frontend/src/api/lessons.ts::listLessons`, which currently has no consumer. The agent-facing listing does reach a reader today. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy |
||
|
|
512d0326a0 |
fix(telemetry): a floor that moved inside the window makes the band check a comparison of two populations (#4225)
CI & Build / Python lint (push) Successful in 3s
CI & Build / integration (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / TypeScript typecheck (push) Successful in 1m0s
CI & Build / Build & push image (push) Successful in 33s
`retrieval_telemetry(days=30)` reported, for write_path_rule:
"the weakest tenth of what this arm returns scores 0.6984, only -0.0216
above its floor of 0.72"
A negative distance above something. The tenth percentile of what an arm
RETURNED cannot sit below the floor that gates what it may return — not
inside one population.
MEASURED CAUSE. write_path_rule's floor was 0.68 until 2026-09-02, when
|
||
|
|
bb8013928f |
feat(telemetry): the readout names rules that were opened and changed nothing (#4213)
CI & Build / Python lint (push) Successful in 5s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 56s
CI & Build / integration (push) Successful in 1m6s
CI & Build / Python tests (push) Successful in 1m40s
CI & Build / Build & push image (push) Successful in 35s
Milestone 419 step 2. Step 1 made an outcome recordable; this makes it readable. `retrieval_summary`'s rule block gains `applied`, `departed` and `distinct_rules_acted`, and `_compute_warnings` gains two codes. TWO CODES, NOT ONE WITH A ZERO IN IT. `read_and_unacted` reports rules that were opened and left no outcome, against the ones that did. It only fires once outcomes exist anywhere in the window, because a window with none cannot tell "every rule was ignored" from "nothing calls `rule_outcome` yet" — and on every install the day this ships, the truth is the second. Claiming the first there would be #3311's failure exactly: a statistic that could not vary being read as a fact about the corpus. The cold case gets its own code, `outcomes_never_recorded`, whose prose says in as many words that it does NOT mean the rules were ignored. `applied` AND `departed` ARE NOT SUMMED. A departure carries the reason the agent gave and is evidence about the RULE; an application is evidence about the agent. Folded together they would say only "an outcome exists", which is true of both and useful about neither. `distinct_rules_acted` counts either, because for the unacted arithmetic the distinction does not matter. An outcome is not a pull. The fold branches on OUTCOMES first and never routes an outcome through the surfaced/ambient split: `source` on an outcome row names the door the outcome came through, not a ranker, so the ambient distinction has nothing to say about it. An integration test holds that line — if an outcome leaked into the pull counters the silently-unchanged rule would vanish into a compliant-looking total, which is the confusion #4212 was opened to end. Verified by lifting the shipped `_compute_warnings` out of source with `ast` and exercising it against the six populations the new tests assert: cold instrument, warm instrument, departures-only, full compliance, nothing opened, and a failed read. The integration tests for the new counts run against real Postgres in CI — count(distinct) with an IN over an unconstrained column is a SQL shape a mock would agree with whatever it did, which is what #2663 was. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy |
||
|
|
dfcb000719 |
feat(rules): a surfaced rule gets an outcome, not just a read (#4212)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m3s
CI & Build / Python tests (push) Failing after 1m6s
CI & Build / Build & push image (push) Skipped
Milestone 419 step 1. `rule_usage_events` could say a rule was SURFACED and that it was PULLED. It could not say what happened next, so a rule that fires constantly and is always obeyed and a rule that fires constantly and is never obeyed left byte-identical telemetry. The second is far the more urgent and was the one the readout could not name — measured on a session where three of seven misses were caught by the operator and none by the system. Two new events, `applied` and `departed`, and a `detail` column carrying the why of a departure. No CHECK migration: `event` was created in 0094 as plain Text with no constraint, verified in the migration rather than assumed from the model, so rule 36 does not bite here — said in both places because the next person adding a value will reach for it. THE THIRD STATE IS DERIVED, AND THAT IS THE DESIGN. Read-and-silently- unchanged is the failure this milestone was opened on, and it cannot be reported: an agent that knew it was ignoring a rule would not be ignoring it. So nothing here asks. `applied` and `departed` are reported; the third state is a rule that was opened and left no trace. An `ignored` enum member would collect nothing while reading as though it had measured something, which is #3311's failure — a statistic that could not vary being taken for a finding. `detail` is a column rather than two more bare event strings because a departure stripped of its reason reads back as a miss, so the two states this exists to separate would collapse again one layer down, in the readout, where nobody would see it happen. Nullable: following a rule needs no argument, and an expensive event is one that stops being recorded. `outcome_state` is the single reading of the four states, taking the aggregate `usage_for_rules` already returns, so the badge, the readout and any later session summary cannot disagree about what "followed" means — the drift #3246 found across the rules system. A departure outranks an application: a rule both applied and argued with is a rule someone argued with, and the argument is the half worth surfacing. `rule_outcome` is the MCP door, classed as a WRITE. The read-only set tolerates getters that call record_pulled, but those are reads that leave a trace; this tool's entire effect is the row, and the row carries prose the agent authored. A read-scoped key that can put text in the operator's database is not read-scoped, whatever table it lands in. Backup carries `detail` on both sides. It is the one field here a fresh install cannot re-earn — counts come back by being used again, a stated reason exists once — and #4197 records that the column guard watches the export side only, so the round-trip test is the thing that would catch a one-sided add. Delivery is deliberately not settled here: how an agent gets prompted to record an outcome is step 3's subject, and the same record serves whichever answer that step reaches. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy |
||
|
|
0fe19a8440 |
fix(ledger): a canon may hold a class and the to_dict beside it (#4220)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m41s
CI & Build / Build & push image (push) Successful in 28s
The review surface shipped yesterday reported two canons on its first live day and both were sound. Coherence was "do all judged rows share a form", which #2844 failed at 37/62 = 0.597 for containing a model class and the to_dict the canon's own text says the class must carry, and #2849 failed at 4/7 for pairing sync loop-starters with the async ticks they schedule. A review surface whose whole output is noise is one that stops being read. The obvious repair is a trap, and there is now a test standing in front of it. Grouping by family and keeping a majority test makes the check BLIND: before #2844 was cleaned by hand it held 37 classes and 56 callables, which as families is 56/93 = 0.602 — a clean pass, and the 31 rows that had no business being there (Vue functions, route handlers, a dozen tests) would never have been reported at all. A looser bar in the same shape is worse than the bug. So the verdict is inverted. Instead of asking whether most rows agree, it asks how many rows the canon CANNOT ACCOUNT FOR: a row in the majority family is accounted for; a callable defined in a file that also holds a majority-family `type` row is a method of a member, not a foreign body; and strangers above a fifth of the readable rows make the canon incoherent. The majority vote abstains those methods, so a class's own serialisers cannot outvote the classes and turn the members into the strangers. Measured on the real ledger before it was written, which is why it is this rule and not a nudge to the share: clean #2844 has 0 strangers in 62, #2849 has 0 in 7, and polluted #2844 had 31 in 93 — the same 31 withdrawn by hand this morning, named exactly. The entry now carries `families`, the majority `family`, `attached`, `stranger_count`, `unattended`, and `strangers` — THE ROWS THAT DO NOT FIT, replacing a sample of the first twelve members. The reader's question is which rows are wrong, and a sample of the agreeing majority cannot answer it. `unattended` is the discriminator between a check that is too strict and a ledger full of junk: both canons flagged on day one were entirely audit-judged, and nothing showed that without opening each one. Scoped to the review surface. `canon_form` still answers at the precise form level for stamping and divergence, where a sync helper beside an async canon is a fair question; nothing here changes what the ledger writes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy |
||
|
|
e87bcfa48c |
fix(guidance): the index had two characters of headroom, and I spent 391
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 8s
CI & Build / integration (push) Successful in 44s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m37s
CI & Build / Build & push image (push) Successful in 25s
CI 7098: unit tests red, everything else green. `_INSTRUCTIONS` was 2439 against a 2000 budget. WHAT I DID NOT CHECK. That block is capped because Claude Code injects only the first ~2,048 characters of a server's instructions and cuts the rest mid-word (#2562, observed live — a 20k version delivered ~10% of itself and the Systems guidance never reached a session). The cap is stated in a comment directly above the literal I edited. It was at 1998/2000 before this batch: a shared, nearly-exhausted resource, and I added a six-line entry to it. THE JUDGE LINE STAYS, and paying for it is the decision rather than dropping it. A client with no Agent Skills support receives this index and nothing else, so of everything here, "you are the judge of record" is among the least safe to leave past the fold — an agent that never learns it defers every call to an operator who was never going to make them. So the line is earned by compressing prose AROUND the existing markers, not by removing anyone's entry: RULES loses a clause, RECORD and REPORT lose trailing restatement, PLAN drops a sentence the two markers already imply, and the opening paragraph tightens. Every index marker the ownership registry requires survives verbatim — that is what test_the_index_names_each_reflex_it_points_at checks, and it passes. Back to 1998/2000: the same headroom as before, with one more reflex indexed. The next addition pays the same way. Three guidance modules run green locally (21 tests) — they read files and need no database, so this one did not have to go to CI to be known. Plugin version re-minted; the previous mint is on a commit that never went green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy |
||
|
|
76bf21633e |
feat(guidance): the agent is the judge — stated in the product, not in a rule
CI & Build / Python lint (push) Successful in 8s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 56s
CI & Build / Python tests (push) Failing after 1m10s
CI & Build / Build & push image (push) Skipped
I recorded this as project rule 174 first. That was wrong twice over, and the second reason is the one that matters. RULE 119 SAYS THIS EXACTLY: guidance about how an agent should behave with Scribe belongs in `_INSTRUCTIONS`, `plugin/skills/*` or the adapter's static context, never in the corpus. I read 119 while writing the rule, decided it was "about authority rather than about using Scribe", and wrote it anyway — which is the reasoning preference 29 exists to catch, performed in full. THE REASON THAT MATTERS: a rule in the corpus is true on ONE install. If the agent being the judge is how Scribe works, every install gets it or none does. Baked in, it ships. As a rule it was one operator's private note about a product stance. WHAT IT SAYS. The agent is the judge of record for the work — what a shape is, whether a finding holds, whether something is done. Surfacing a finding for the operator to rule on is the judgment NOT made, however well written up: it reads as diligence and functions as a backlog. Escalate the acts that are genuinely theirs — their money, their infrastructure, anything hard to reverse or facing outward — and keep the decisions. A hard call is still yours; an irreversible act is still theirs. And the half that keeps this from becoming the previous defect: JUDGING IS ATTENDED. An agent reading evidence and recording why is judgment; a threshold or a sweep reclassifying in bulk with nobody reading is the thing that fills a ledger with confident nonsense (#4208, and Portal's 35 rows). When the fix for bad unattended writes is another unattended write, stop. THE PRODUCT WAS TEACHING THE OPPOSITE. reporting-back's Finding row read "Symptom · Cause · Size of the fix · **Offer to fix it**". So the behaviour I was corrected for is the behaviour the skill prescribed — which is the better argument for fixing it here than any rule could be. Three surfaces, per 119 and the ownership registry (#4027): `_INSTRUCTIONS` gets a one-line JUDGE index entry; using-scribe owns the authority and the attended/unattended distinction; reporting-back owns the report shape. Two topics rather than one, registered separately in test_guidance_ownership so trimming one cannot quietly take the other. Plugin version minted — skills only reach a session when the manifest moves (#2209). Rule 174 deleted (trash 074434a2, recoverable). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy |
||
|
|
400253d039 |
feat(ledger): the ledger can say "these look wrong" without acting on it (#4208)
CI & Build / Python lint (push) Failing after 3s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / integration (push) Successful in 1m6s
CI & Build / Python tests (push) Successful in 1m49s
CI & Build / Build & push image (push) Skipped
THE HALF THAT WAS MISSING. #4204 put a floor under what the write-path hook may assert. A floor only guards new writes; every row already stored stands (lesson #4202). Measured after that fix shipped: Portal carried 32 rows under one canon and 3 under another, all stamped on scores of 0.69-0.77 — below the 0.80 floor, so none of them could be written today, and all of them were still there. Scribe's own ledger carries 334 under #2860. `stamps_to_review` reports two things and changes nothing: weak — rows the hook stamped on a resemblance below the current floor, each with its score, signature and derived form. incoherent — canons whose own judged rows do not agree on a form. A canon claims some shapes are the same sort of thing; when its members are a class, three getters and a dozen tests, that claim has stopped being true and every base-rate reading built on it is reading noise. `canon_form` already made such a canon fall silent — nothing made it VISIBLE. IT DELIBERATELY CANNOT FIX ANYTHING, and that is the design, not an omission. The first version of this commit was an automatic sweep that reset rows by score. That is the original defect pointed the other way: what harmed the ledger was not one wrong score, it was a machine recording permanent classifications unattended. Un-recording them unattended is the same act with a wider blast radius. An agent reads the evidence, judges, and records the judgment under its own name through `classify_shapes`. `test_the_service_carries_no_machinery_for_bulk_withdrawal` asserts that structurally, so the next person to reach for an auto-retire has the argument again on purpose rather than in a diff nobody reads. A JUDGMENT IS NEVER LISTED AS WEAK, whatever its age. This is the measured correction to an assumption I nearly shipped: of Scribe's 334 rows under #2860, 302 are in `services/` — the canon's own home — and the ones sampled there are `classified_by="audit"` with no score at all. The legitimate bulk of that canon was never scored; it was judged by an agent in batch. Listing those as weak would invite an agent to withdraw the only real judgments in the ledger. An agent's decision is a different KIND of evidence, not a worse one. THE SCORE NOW HAS A PARSER. It lived only inside a prose sentence, so nothing could ask how strong the evidence for a row was without re-deriving it — which is how 32 rows sat unexamined for nineteen days. Format and reader are one constant apart (`_RESEMBLE_REASON` / `stamp_score`), with a round-trip test and a test pinned to reason strings taken verbatim from the two poisoned ledgers. `live_rows_for` is `live_rows` behind the project read gate, for callers that arrive from outside rather than from a job that already knows who is asking. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy |
||
|
|
e2c3a5c2b5 |
feat(telemetry): retrieval_telemetry says what is wrong (#3431)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 8s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m30s
CI & Build / Build & push image (push) Successful in 23s
The tool returned distributions and left the reading to the caller, so every readout was the same four checks done by hand — #3430's baseline, #3835's rule near-misses, the #1038 rerank gate. Mechanical, and therefore forgettable. Tonight's acceptance pass on #3898 was the case for doing this. Reading it by hand meant catching that two surfaces had `covers_window: false`, that `prompt_rule`'s floor had moved three times inside the window (which made the readout self-contradictory: deliveries at 0.622 beside refusals at 0.7199), and that 15 of 20 near-misses were one record against text no operator wrote. Miss any of those and the obvious conclusion was "the bar is too tight" — a floor change that would have injected one preference into every notification. `warnings` is always present and empty when clean, so its emptiness is an answer rather than a gap. Each entry carries the numbers that produced it: "345 calls, 0 declined" is the analysis, "check write_path_rule" is an instruction to redo it. Five codes — cannot_decline, band_hugs_floor, no_duration, surfaced_never_pulled, unregistered_source. cannot_decline has three guards, each a bug it would otherwise cause. Asked surfaces are exempt (a search returning a list every time is working). An arm not known to log unconditionally is exempt — that is #3497 exactly, where both rule arms recorded only their hits, so a decline count of zero was a LOGGING defect and this warning would have sent the reader to a threshold that was never involved. Unregistered sources get numbers but no verdict. `silent_surfaces` is the half the rows cannot show: an arm that emitted nothing is invisible to every row-based check and looks exactly like an arm that does not exist. It is driven by a new declared registry, `retrieval_registry.POINTS` — deliberately NOT `retrieval_surfaces.SURFACES`, which answers "what can be tuned" and excludes the reserved slots because a budget of 1 is their feature. This answers "what can be measured", and the reserved slots belong in it precisely because they are judgeable without being tunable. A test asserts the two cannot drift apart. The registry test derives sources from the call sites with `ast`, not grep, and the difference is not theoretical: `wide_net` and `report_preference` reach their recorder as `source=SOURCE` through a module constant, so a grep for `source="` is blind to both — the narrowing #3191 warns about. Three sites pass `source` as a variable and are declared in FAN_OUT_SITES; the test pins those sites but not the values they can pass, which is why the `unregistered_source` warning exists to catch the rest at first fire. Thresholds are settings (rule 25) defaulted so a fresh install with almost no data produces no warnings at all (rule 115) — a new user's first readout naming five broken things would be describing the emptiness. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy |
||
|
|
a4aae974a2 |
feat(telemetry): a usage event records which project the reader was in (#4196, #3735)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m5s
CI & Build / Python tests (push) Failing after 1m15s
CI & Build / Build & push image (push) Skipped
`RetrievalLog` has carried `project_id` since it existed, so "this record was SURFACED on project B" was always answerable. `note_usage_events` had none, so "this record was OPENED on project B" was not — and the two cannot be joined to recover it, because there is deliberately no session identity server-side. NoteUsageEvent's own docstring rules that out. That gap sat exactly on the question milestone 385 exists to answer. A lesson's whole claim is that it reaches a session on a project it was not written on, and step 8's acceptance is "retrieved on a different project AND opened". Each half was answerable; the conjunction was not. WHICH project, because the name is ambiguous and the wrong reading makes the column useless: it is the project the READER was in, never the one the record belongs to. The record's own project is already on the note; copying it here would answer a question nobody asked while looking like it answered this one. The surfacing half is free — every arm already holds the scope it just searched, so auto_inject, lesson_slot, the write-path arms and enter_project now record it. process_skill_sync does not and should not: it installs every Process the operator can reach, which is not a project-scoped question, so a project there would be a fiction. The pull half needs the caller, since a getter knows only what it was handed. The five single-record getters take `project_id: int = 0` and pass it through, following the convention `search` and `create_*` already set. Null stays an ordinary answer meaning "not reported" — a pull with no project is still a pull and still counts toward dead weight; it simply cannot speak to transfer. The four REST detail views report none for now: a human opening a record in a browser is a different event from an agent recalling one, and #2245 left that asymmetry deliberately undecided. Guarded the way #2245 and #2476 taught: by source inspection, because a parameter that was never threaded through changes no return value and shows up only as a column that is mysteriously always null. Three guards — the signature, the pass-through, and the arms — plus the can-fail test rule 167 asks for. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy |
||
|
|
26a757ecfe |
feat(lessons): a lesson is yours to keep current too (#4195)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 52s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m49s
CI & Build / Build & push image (push) Successful in 29s
The lesson kind shipped with every mechanism for growing and nothing telling a session to use them. `learned_from` is a list on purpose, the dedup gate hands back an existing id rather than minting a twin, and `update_lesson` already names re-keying a bad trigger as the edit that pays most. None of that was reachable as a habit. The exclusivity claim was the bug. The skill said "a preference is the one record you keep current yourself", and by naming only preferences it put lessons outside the habit. That sentence is now "a preference is yours to keep current", which says the same thing about preferences without saying anything false about lessons. Beside it, a paragraph on what growing a lesson means: another incident added to what taught it, a claim stated more exactly, or a trigger re-keyed to the situation that really fired. Written as a practice rather than a prohibition (rule 165) — the reader is named as the one person placed to judge the trigger, because they are standing in the situation it claims to name. `get_lesson` carries the same prompt at the moment it bites: a session reading a lesson inside the situation it names is the only reader who can tell whether the trigger is keyed to what actually fired. The guidance-ownership registry gains the topic and re-points the preference topic's statement, since the phrase it pinned is the sentence this change rewrites — the module asks for exactly that, in the same commit. No index marker: the index names session-start reflexes and this one fires mid-work, so `_INSTRUCTIONS` stays at 1998/2000. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy |
||
|
|
d36d68a20f |
feat(lessons): the REST door a human can actually reach (#3734)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 8s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m33s
CI & Build / Build & push image (push) Successful in 23s
Step 7, part two. Milestone 385 built the lesson kind through the MCP tools, which is the agent's surface. The Vue app speaks REST, so a lesson was a record a person could not create, read, edit or retire — rule 27 failing at the door rather than in the view. `/api/lessons` now offers list, create, read, update and trash, plus `/api/lessons/taught-by/<id>` — the reverse of `learned_from`, which the task body calls the direction that gets forgotten and arguably the more useful one: a reader opening an old issue wants to know what was learned from it, and until now the relation was only navigable from the lesson's side. `lessons_taught_by` reads `data[taught_by]` through `path_exists`, the same jsonpath dialect the snippet location lookup uses, so both reverse lookups hit the GIN index (0070) the same way rather than scanning bodies. Share-aware via `readable_notes_clause`: it renders beside a record the caller can already see, so a lesson shared with them belongs there exactly as their own does. THE TRIGGER IS REFUSED WHEN EMPTY, at create and at update. This is the one place the door is not a thin wrapper, and it is deliberate: the service will store a triggerless lesson quite happily — it saves, reads correctly in every listing, and never surfaces. There is nothing to notice afterwards, because it looks exactly like a lesson that works. Better to refuse it than to hand back a record that looks finished. The refusal says why, so the next reader does not take it for a nag and delete it. `lesson_to_dict` moves into the service and the MCP tool's `_to_dict` becomes an alias for it. Both doors now return one shape — a payload spelled once per door answers the two of them differently the first time a field is added — and both compose through `services/lessons.py`, so a lesson written from the web ranks identically to one written by an agent. The document IS what ranks, so that parity is the whole reason the door is thin. The dedup gate matches the MCP path: two lessons under one trigger compete in a single ranked list for one reserved slot, so a duplicate here displaces rather than merely clutters. NOT DONE YET: this is the door, not the UI. #3734 stays in_progress until the Vue views, the router entries, the Knowledge browse badge and the both-ways sources panel exist — rule 27 is about the operator being able to touch it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy |