d4de0b9b54ebf3a0a67a53495f784e2d0bb16dc5
1724
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d4de0b9b54 |
Merge pull request 'Merge dev: scope-then-rank searches (#4958, #4961), CI gate on integration, milestone 456 steps 4-6' (#205) from dev into main
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 1m52s
CI & Build / Build & push image (push) Successful in 18s
|
||
|
|
ccbccb025c |
fix(retrieval): every semantic search scopes first, then ranks - the note, milestone and system searches join the rule search on one shared shape, _rank_scoped (#4961)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 17s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m0s
CI & Build / Python tests (push) Successful in 1m54s
CI & Build / Build & push image (push) Successful in 27s
Ordered straight off a *_embeddings table, the planner walks the HNSW index, takes ~ef_search (40) nearest chunks across every owner and project, and only then applies the scope: an in-scope record behind 40 nearer ones the caller cannot see was silently dropped. #4958 fixed the rule search alone; the note, milestone and system searches kept the fault. _scoped_chunks builds the in-scope chunks with their distance; _rank_scoped materializes them as a CTE and ranks exactly. All four searches go through it, and the row shape is unchanged, so callers and mocks are untouched. Tests: a structural guard that every semantic_search_* ranks through _rank_scoped and orders nothing itself (with a replay of the old shape), a compiled-SQL check of the MATERIALIZED CTE, and integration crowd tests - 80 nearer out-of-scope records - for notes (another user; the reader own other project), milestones and systems. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
d2fa723373 |
refactor(retrieval): the rule builders compose one moment - _rule_moment runs the arm and its via-lesson step for all three (milestone 456 step 6, #4908)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m54s
CI & Build / Build & push image (push) Successful in 28s
build_prompt_rule_hint, build_tool_rule_hint and build_write_path_hint each ran the rule arm and then the via-lesson step by hand, the first two passing the query between them through a _via_query key on the payload. They now call one composer, _rule_moment, and the split helpers (_prompt_rule_hint, _tool_rule_hint, _add_rules_via_lessons) and the side channel are deleted. Output shapes are unchanged: the prompt builder returns no checkpoint key, the tool builder always does, and shown_rule_ids stays the direct band. A structural test pins that only _rule_moment runs a rule arm or the via-lesson step, and that the three builders are its only callers. plugin_context.py: 2,418 -> 2,361 lines this step; 3,558 when the milestone began. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
869046dda2 |
refactor(retrieval): the specs are the registry - SURFACES, the ranked POINTS rows and RANKED_SOURCES are read off the pipeline specs (milestone 456 step 5, #4907)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 2m1s
CI & Build / Build & push image (push) Successful in 27s
Each arm spec now carries its tuning pair (Surface, moved verbatim into retrieval_pipeline) and a Declared block - what the telemetry readout must know and cannot read off its rows. The two ranked stages that are not arms (preference_slot, rule_via_lesson) are RankedSource specs, and the note slots carry their own declaration. - retrieval_surfaces.SURFACES = the TUNED_ARMS tuning, same order - retrieval_registry.POINTS ranked rows = one per spec in RANKED; the lookups, asked, ambient and pull rows stay declared there - rule_usage.RANKED_SOURCES = RULE_RANKED_SOURCES + moment_rule Settings keys, defaults, prose and order are unchanged (checked field by field against HEAD). tests/test_retrieval_specs.py pins the derivation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
2b8f41229d |
refactor(retrieval): the notes arms run on the one pipeline - auto_inject, its reuse and lesson slots, the write path by meaning, and rule_via_lesson are specs (milestone 456 step 4, #4906)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m2s
CI & Build / Python tests (push) Successful in 1m52s
CI & Build / Build & push image (push) Successful in 35s
retrieval_pipeline gains the notes half: NoteArm / NoteSlot / NoteMoment / NoteIO / NoteResult and run_note_arm, which writes once the stages both notes arms copied: search, withhold this response's own menu (#3739), fresh/repeat split, the call row before any return (#3497, #3752), the band, the reserved slots in their order (reuse evicts, lesson extends), and the surfacing rows. The note renderer (_record_kind, _menu_name, _menu_passage, the seen pointer, menu_entry) moves with it, and run_via_lesson_arm takes rule_via_lesson. Behaviour-preserving, with flags for today's differences: notes still log BEFORE the band and rules after it (step 7's question). One deliberate change: a notes arm now fails open like the rule arms, so a failing search costs its lines and no longer the whole hook response. The I/O is resolved from plugin_context at call time (_note_io), so the existing patches keep working. The review re-run reads the AUTO_INJECT spec instead of restating it, and its guard now compares the two live searches. The registry declares the pipeline's notes fan-out sites. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
4b1060ae8b |
Revert the deliberate integration failure - the gate was watched rejecting run 8218 (rule 177, #4958)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / integration (push) Successful in 53s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m53s
CI & Build / Build & push image (push) Successful in 19s
Run 8218: integration failure, Build & push skipped, :dev still package
version 11098 (23:36:29Z, run 8217), and no image tagged
|
||
|
|
9a76eb0a34 |
test(ci): TEMPORARY deliberate integration failure, to watch the publish gate reject a red run (rule 177, #4958)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Failing after 1m1s
CI & Build / Python tests (push) Successful in 1m51s
CI & Build / Build & push image (push) Skipped
Reverted next. Expect: integration red, Build & push skipped, :dev unmoved, no image tagged with this commit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
17f4348711 |
ci: the image build waits on the integration lane (rule 177, #4958)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 56s
CI & Build / Python tests (push) Successful in 1m53s
CI & Build / Build & push image (push) Successful in 18s
The integration job was added after the build gate (
|
||
|
|
983fd2c4d1 |
fix(retrieval): scope the rule search before ranking it - an in-scope rule is no longer lost behind other owners' nearer rules (#4958)
CI & Build / Python lint (push) Successful in 5s
CI & Build / Plugin hooks (push) Successful in 20s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m11s
CI & Build / Python tests (push) Successful in 2m0s
CI & Build / Build & push image (push) Successful in 29s
Ordered straight off rule_embeddings, the planner walked the HNSW index, which returns about hnsw.ef_search (40) nearest chunks across every owner and project and filters by home only afterwards. A reader's own rule ranked past the 40th chunk overall was silently dropped - on a shared install, other users' rules fill those 40. This is the likely cause of the intermittent test_integration_rule_scope failures, whose axis-vector fixtures sit far from every real embedding in the graph. The in-scope chunks now go through a MATERIALIZED CTE, which the index cannot order, so the ranking over them is exact. A rulebook is hundreds of chunks; exact is cheap. New test: 80 nearer rules belonging to someone else no longer hide the reader's one rule. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
79cd2341a6 |
test(rules): the scope test says why a rule search returned nothing - an error or an empty scan (#4958)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / integration (push) Successful in 1m6s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 18s
semantic_search_rules fails open, so both failures on run 8209 and main run 8212 read as set() with no trace. Assert report["searched"] and carry the swallowed traceback into the failure. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
25c1c778f1 |
Merge pull request 'Moments: skills teach moments, and a mount that misfires proposes its own removal (milestone 458 steps 8, 7b)' (#204) from dev into main
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 1m53s
CI & Build / Build & push image (push) Successful in 26s
|
||
|
|
4e1320120d |
feat(moments): a mount that keeps arriving where it does not apply proposes its own removal (milestone 458 step 7b, #4955)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m19s
CI & Build / Python tests (push) Successful in 2m3s
CI & Build / Build & push image (push) Successful in 18s
The open-after-moment signal proposes a mount; nothing proposed taking one off, so a wrong mount was noise at every occurrence until someone happened to notice. rule_misfired(rule_id, moment, why, reached_by) records a report against a MOUNTED pair, counted per distinct day (the MCP door carries no session id) on a new rule_moment_judgments.misfire column (migration 0119, backup v24). At three days the response carries a line asking the agent to offer the operator the fix - reject takes the rule off, unmap_action stops the action reaching the moment, confirm keeps the mount and stops the asking - and Settings > Moments lists it as an unmount proposal with the reasons and the actions that reached it. A re-mount clears the count. Taught in moments.md, missed-retrieval.md and the reply hold's wording. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
8afd6da8af |
docs(plugin): reference files point at each other in words, not links - one level deep (milestone 458 step 8, #4926)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 56s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 2m11s
CI & Build / Build & push image (push) Successful in 41s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
a72a422534 |
docs(plugin): the instruction surfaces teach moments - reading a line that arrived at one, correcting a misfire, and giving a new rule its moments (milestone 458 step 8, #4926)
CI & Build / Python lint (push) Successful in 13s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / TypeScript typecheck (push) Successful in 56s
CI & Build / integration (push) Successful in 1m38s
CI & Build / Python tests (push) Failing after 1m59s
CI & Build / Build & push image (push) Skipped
Until now only the tool arguments knew moments existed. The guidance surfaces described rules as reached by resemblance alone: - using-scribe: a short reflex paragraph and a new reference file, moments.md. It covers reading "at <moment>, reached by <action>", the reply held once at reply.report, map_action / unmap_action offered in one line, and the step 7 proposal line answered with judge_rule_moments. - writing-records: asks WHEN a rule applies as well as what it is about. A rule, preference or process about a point in the work gets moments=[...] as it is written, and the trigger stays as the net. - missed-retrieval: a missed WHEN is mounted or mapped, not reworded. A misfire is unmounted or unmapped. - _INSTRUCTIONS: one clause (list_moments; mount rules about WHEN), 1594 of 1600 chars. - static context: injected lines include the rules mounted on a moment that was reached. - test_guidance_ownership: three owned topics, so the text cannot quietly drop out. Plugin minted. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
2b52afcd72 |
Merge pull request 'Mount the corpus by proposal: a pass and an open-after-moment signal (milestone 458 step 7)' (#203) from dev into main
CI & Build / Python lint (push) Successful in 5s
CI & Build / Plugin hooks (push) Successful in 17s
CI & Build / Python tests (push) Successful in 1m59s
CI & Build / TypeScript typecheck (push) Successful in 56s
CI & Build / integration (push) Successful in 1m12s
CI & Build / Build & push image (push) Successful in 22s
|
||
|
|
c404a3127a |
test(retrieval): the proposal endpoints are routed on the app (milestone 458 step 7, #4925)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 19s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m2s
CI & Build / Python tests (push) Successful in 2m17s
CI & Build / Build & push image (push) Successful in 56s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
570d1b6d5a |
test(backup): register the rule_moment_judgments builder in the import column guard (milestone 458 step 7, #4925)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 56s
CI & Build / Python tests (push) Failing after 1m44s
CI & Build / Build & push image (push) Skipped
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
dfcf4df2e9 |
feat(moments): mount the corpus by proposal - a pass and an open-after-moment signal, both stopping at the operator (milestone 458 step 7, #4925)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 19s
CI & Build / integration (push) Failing after 43s
CI & Build / Python tests (push) Failing after 46s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Build & push image (push) Skipped
A rule written before moments existed is mounted on nothing. Step 7 records, per (rule, moment), whether it belongs there and who said so: - rule_moment_judgments (migration 0118, backup v23): suggested / confirmed / rejected, from a pass, the signal, or an edit. Moment "" is "no moment fits". - The pass: rules_to_mount lists unjudged rules; propose_rule_moments records suggestions that mount nothing; rule_moment_proposals and judge_rule_moments put them to the operator. A confirm mounts, a reject is kept so the pair is never proposed again. Same service behind REST and a "Waiting on you" panel in Settings > Moments. - Edits are judgments: set_rule_moments, the one mount write path, confirms what was added and rejects what was removed in the same transaction. - The signal: scribe_moment.sh keeps a per-session acts ledger; when a rule is opened, scribe_record_opened.sh sends the last three minutes of it to /api/plugin/rule-opened. The acts resolve through the install's mappings; work.run and work.change are not evidence. Counted per distinct session with lesson_rules' evidence model, and once due the open returns one line asking the reader to offer the mount. - scribe_session_end.sh removes the session's scribe-moment files. Plugin 2026.10.05.2003. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
2db9e0c0c1 |
Merge pull request 'Skills and processes declare their moments, and the UI for moments (milestone 458 steps 5–6)' (#202) from dev into main
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m9s
CI & Build / Python tests (push) Successful in 1m59s
CI & Build / Build & push image (push) Successful in 19s
|
||
|
|
5bcc603310 |
fix(moments): the store loads rather than fetches, so the deadline guard does not read it as a bare fetch (milestone 458 step 6, #4924)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 52s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 2m2s
CI & Build / Build & push image (push) Successful in 45s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
b73a689849 |
feat(moments): the human door onto mounts, mappings and per-moment telemetry (milestone 458 step 6, #4924)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 18s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Python tests (push) Failing after 1m34s
CI & Build / Build & push image (push) Skipped
CI & Build / integration (push) Successful in 2m7s
Everything an agent can do with moments, a person can now see and change in the app. - Rule editor: a moment picker beside the trigger. Catalog moments are ticked; a named procedure's `skill.<name>` is typed and checked as the server checks it. `moments` is always sent, so unticking the last moment unmounts the rule. - Settings, Moments section (General tab): for each moment, what it means, the actions that reach it on this install (shipped ones can be switched off, the install's own removed), how many rules are mounted on it, and deliveries and agent opens over the window. Below that: named procedures with mounts, switched-off defaults with Restore, and a form to add an action. - retrieval_telemetry.moment_usage: per moment, `delivered`, `rules`, `opened` (agent pulls after the first delivery there; an upper bound, as by_source is) and `last_delivered_at`. No ratio, because a mount is a person's statement, not a ranker's guess. Guarded on its own, and also reported in retrieval_summary as `moment_usage`. - rulebooks.mount_counts; mounted_moments now derives from it. - GET /api/retrieval/moments carries `mounted` and `usage` (?days=). DELETE /moments/mappings also reads the mapping from query parameters, since the browser's DELETE sends no body. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
c489452a2d |
test(moments): an install removal still drops its tool while the skill loader stays listed (milestone 458 step 5, #4923)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 19s
CI & Build / TypeScript typecheck (push) Successful in 1m2s
CI & Build / integration (push) Successful in 1m20s
CI & Build / Python tests (push) Successful in 2m12s
CI & Build / Build & push image (push) Successful in 33s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f1fdc4a951 |
feat(moments): skills and stored processes declare the moments they are for (milestone 458 step 5, #4923)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m3s
CI & Build / Python tests (push) Failing after 1m25s
CI & Build / Build & push image (push) Skipped
Loading a procedure now also reaches the moment it is for. Loading the reporting procedure is a report; loading the release procedure is a delivery. - Bundled skills: each SKILL.md declares `metadata: moments:`. The same declaration ships as Skill defaults (BUNDLED_SKILL_MOMENTS), because the server never sees the plugin's files. test_skill_moments holds the two together and pins the plugin name that qualifies the skill. - Stored processes: `moments` on create_process and update_process, stored in the note's data and returned by get_process. A `scribe-proc-<slug>` load resolves its process through the sync manifest at load time. The moments are not copied into the stub, which would go stale mid-session. - reachable_tools lists the skill loader whenever anything is mounted, since a process's moments are known only when it loads. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
9fc2df080d |
Merge pull request 'Rules mount on moments, and one retrieval pipeline (milestones 458 steps 1–4, 456 steps 1–3)' (#201) from dev into main
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 1m0s
CI & Build / Python tests (push) Successful in 1m54s
CI & Build / Build & push image (push) Successful in 19s
|
||
|
|
b29689d4de |
fix(retrieval): declare the reply moment's surfacing site in FAN_OUT_SITES (milestone 458 step 4b, #4922)
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / Python tests (push) Successful in 1m50s
CI & Build / Build & push image (push) Successful in 40s
moment_delivery records the mounted half's surfacing under source=rp.MOMENT_RULE_SOURCE. That is an attribute the registry's extractor cannot resolve, so the site is now declared with the value it emits, read from the pipeline's own constant. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
c6cdfc2172 |
feat(moments): the reply moment holds a finished reply for one read (milestone 458 step 4b, #4922)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / integration (push) Successful in 1m1s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Python tests (push) Failing after 1m25s
CI & Build / Build & push image (push) Skipped
The reply is the one act no tool call marks, and it is where "let me know if it works" gets said. A new Stop hook (scribe_reply_check.sh) sends the finished reply to POST /api/plugin/reply-rules, which checks it twice: - mounted: every unopened RULE on reply.report, plus reply.ask when the reply asks a question. Deterministic. - semantic: the reply's head and tail against every rule's trigger, on a new ranked surface, reply_rule. It is the backstop for whatever the earlier arms missed. Its floor is its stop bar (default 0.80, budget 1), with its own Settings dials. The new stop_only stage records surfacing for the rule that holds and nothing else, because nothing else reached anyone. Following the operator's ruling from 456 step 8, a rule that holds blocks once, in the server's words. The hook blocks only on a reason it was given, so an unreachable instance never stops a session, and it never holds the rewrite. The ledger is the act checkpoint's own, so a rule holds a session once across both doors and the per-session cap counts both. The turn reader moved from the report check into scribe_defs.sh (scribe_turn_facts / scribe_turn_fact), so the two Stop hooks read a turn the same way. The output was checked identical on a real transcript. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
b3b616b20a |
fix(moments): the tools import the delivery module, and the tool-list cache leaves the swept directory (milestone 458 step 4a, #4922)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / integration (push) Successful in 49s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m48s
CI & Build / Build & push image (push) Successful in 32s
Two guards caught
|
||
|
|
78653130d6 |
feat(moments): mounted rules arrive when their moment happens, through every door (milestone 458 step 4a, #4922)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 57s
CI & Build / Python tests (push) Failing after 1m19s
CI & Build / Build & push image (push) Skipped
A rule mounted on a moment now reaches the session when an act reaches
that moment, with no semantic match involved:
- run_moment_arm on the pipeline: a lookup, not a ranked search. Each
line names the moment and the act that reached it ("at work.deliver,
reached by `git push`"), so a misfire is visible where it lands and
can be unmapped in-session. A repeat is cited, not quoted; fresh
rules are recorded surfaced under source moment_rule with the moment
in detail. No retrieval_logs row, as for the other lookups, so no
latency is persisted for this arm.
- rule_scope: a rule's home clause, moved out of semantic_search_rules
so the moment lookup scopes by the same one.
- rulebooks.rules_on_moments / mounted_moments.
- The plugin door: a catch-all PreToolUse hook (scribe_moment.sh). It
keeps /moment-tools' answer on disk for five minutes, so a call to a
tool that cannot reach a mounted rule sends nothing, and an install
that has mounted nothing sends one request per window. It shares the
rules ledger with the other arms and fails open silently.
- The MCP door: Scribe's own tools named by the shipped mappings carry
moment_rules in their response, so a client without the plugin gets
them too. The hook skips those tools. A guard pins the attach on
every one.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
cc26054437 |
feat(moments): rules mount on moments, through every rule door (milestone 458 step 3, #4921)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m0s
CI & Build / Python tests (push) Successful in 1m51s
CI & Build / Build & push image (push) Successful in 29s
rule_moments (migration 0117) records which moments a rule arrives at, by catalog name, cascading with the rule. rule_detail, the one seam every rule door already returns through, gains moments beside system_ids: None leaves the mounts alone, a list replaces them. get_rule and both list_rules doors read them back, batched per page. All five MCP rule/preference writes and the three REST ones take moments and validate them before their create or update. An unknown name is refused with the catalog listed and leaves no half-made rule behind; a parity test pins that ordering on every door. Backup v22 carries the mounts as a join table remapped through the rule map; a real-Postgres round trip checks they land on the restored rule. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
cf3de5bae1 |
feat(moments): actions map onto moments, with in-session corrections (milestone 458 step 2, #4920)
CI & Build / Plugin hooks (push) Successful in 18s
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m23s
CI & Build / Python tests (push) Successful in 2m0s
CI & Build / Build & push image (push) Successful in 36s
moment_actions.resolve(tool, input) names every moment a call reaches and the action that reached it. One call can reach several: kubectl apply is a run, a deliver and a reach outside the workspace. Command tools match by how each segment of the line starts, with a word boundary; other tools by field=value arguments. The MCP server prefix and case are ignored. 56 shipped defaults cover the harness tools, Scribe tools and common command shapes. moment_mappings (migration 0116) holds what an install adds and the defaults it switches off. A removal is a stored row, so an upgrade does not switch the default back on. Per the operator ruling, corrections happen in the session: map_action and unmap_action (write tools) return now_reaches so the fix can be confirmed in the same reply. list_moments now shows each moment's actions on this install. REST mirrors both doors, recorded as human. Backup v21 carries the mappings. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
2ff7f2f34f |
feat(moments): the moment catalog rules will mount on, readable in-session (milestone 458 step 1, #4919)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 58s
CI & Build / Python tests (push) Successful in 1m53s
CI & Build / Build & push image (push) Successful in 32s
Fourteen generic moments of work (session.start, work.start … reply.ask) plus the skill.<name> family, each with what it means and the kinds of action that reach it, written for any kind of work rather than software alone. The catalog is code because every install needs the same mount points; which actions reach a moment is per-install data (step 2). require_moment refuses an unknown name with the catalog listed, so a typo cannot become a mount that never fires. list_moments (read-only) and GET /api/retrieval/moments hand out the same catalog. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
0720ab6dcf |
refactor(retrieval): the completion-report preferences run through the pipeline (milestone 456 step 3, #4905)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Build & push image (push) Successful in 28s
reply_preferences.completion_preferences was the last hand-written rule search + record_retrieval + record_rule_surfaced triple outside the pipeline. It is now REPORT_PREFERENCE, a RuleArm with kind="preference" over COMPLETION_QUERY, read back as records rather than as lines. - RuleArm gains `kind`. _ranked asks the ranker for that kind and drops anything else on the way out, before the row is logged. - RuleResult gains `shown`, the (score, rule) pairs behind the lines, for callers that return records. - RuleMoment.project_id may be None: logged as given, searched as `project_id or None`. report_preference rows keep their NULL project. - Guards: reply_preferences now has 0 direct rule searches. The registry constant-source example moves from report_preference to preference_slot, the pipeline slot that records through a module constant. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
b72c9a92e2 |
test(retrieval): two source guards read the pipeline where the rule arms now live (milestone 456 step 2, #4904)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 49s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Build & push image (push) Successful in 30s
CI run 8137 had two failures, both source-scanning guards whose property moved to the pipeline. The property itself still held: - test_every_surface_name_is_a_real_telemetry_source looked for source="<name>" literals. The rule arms now record under their spec field, so the spec is the join key, read from RULE_ARMS. - test_every_hook_rule_search_says_which_project_it_is_for counted 4 direct searches in plugin_context. It now expects 0 there, so a copied arm fails the count. It also asserts that the pipeline has exactly one search and that its keyword literal carries project_id and never everywhere. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f6b824b214 |
refactor(retrieval): the three rule arms run through one pipeline (milestone 456 step 2, #4904)
CI & Build / Plugin hooks (push) Successful in 19s
CI & Build / Python lint (push) Successful in 2s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / Python tests (push) Failing after 1m28s
CI & Build / Build & push image (push) Skipped
CI & Build / integration (push) Successful in 1m15s
prompt_rule, pre_tool_rule and the rule half of the write path each wrote the same steps out by hand: search, band, split fresh from repeats, log the call before any early return, reserve a slot, render, record surfacings. The copies drifted, and #3497, #3750 and #3752 were fixed one copy at a time. - New services/retrieval_pipeline.py. run_rule_arm runs the stages in one order. RuleArm is the spec (source, band, compact_tail, checkpoint, preference_slot). RuleMoment is the query and the session ledger. - The I/O (ranker and recorders) is passed in as RuleIO. plugin_context resolves it from its own names at call time, so existing patch points still apply. - _rule_band, _rule_hint_line, checkpoint_for, checkpoint_reason, the band constant and the preference slot moved into the pipeline unchanged. plugin_context re-exports them. - A recorder that raises now costs only its row, never a rendered line. - The flags reproduce today exactly. Whether prompt_rule should band, and whether the act arms should reserve a preference, are step 7. - Registry: the pipeline call sites are declared in FAN_OUT_SITES, with values read from the specs. - The #3497 structural guard now checks the one implementation, and that plugin_context writes no rule-source row of its own. plugin_context.py: 3,558 -> 2,980 lines. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
be62abf142 |
perf(retrieval): one embedding per query per plugin request (milestone 456 step 1, #4903)
CI & Build / Plugin hooks (push) Successful in 17s
CI & Build / Python lint (push) Successful in 2s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m25s
CI & Build / Python tests (push) Successful in 2m6s
CI & Build / Build & push image (push) Successful in 1m25s
A single plugin request fans one query out to several ranked arms. The operator message is searched by auto_inject, its reuse and lesson slots, prompt_rule, the preference slot and rule_via_lesson, and each arm called get_embedding on its own: up to six model calls for one vector. - embeddings.query_embedding_memo(): a request-scoped ContextVar memo that get_embedding consults. Outside a scope the default is None, so every other caller is unchanged. A failed embedding is not remembered. - memoized_query_embeddings decorates /retrieve, /tool-rules and /prior-art. - get_embeddings (document chunks) is untouched. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
802ead748b |
Merge pull request 'Rules never-opened advice per corpus (#4798); a lesson's name is one line (#4797)' (#200) from dev into main
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 53s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Build & push image (push) Successful in 22s
|
||
|
|
d2ac7bf220 |
fix(lessons): a lesson's name is one claim on one line (#4797)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m48s
CI & Build / Build & push image (push) Successful in 28s
#4797 "A lesson's what takes a whole narrative, and the narrative becomes its name". `what` is the title every listing and menu prints. Nothing enforced its documented "one line", so lessons written with the incident in `what` printed up to ~1,500 characters as a menu line, buried the claim, and diluted the trigger in the embedded title. That was 14 of the 39 lessons on this install. - lessons.require_claim refuses a `what` over WHAT_MAX_CHARS (240) or running over several lines. The refusal says the story goes in `insight`. - Both doors call it before writing: - MCP create_lesson / update_lesson; - REST create / update. An update checks only a NEW name, so a lesson stored with a long one can still take the edit that repairs it. - lessons.claim_line is the display half. _menu_name shows an over-long stored name as its first sentence marked " …", because a door guard does not undo rows already stored and an unrepaired install would keep printing them. - The tool docstrings state the bound. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
9909cd2450 |
fix(retrieval): the rules never-opened warning stops sending the reader to a review tool that refuses rules (#4798)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 15s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m7s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 26s
#4798 "The rules corpus's surfaced_never_pulled warning sends the reader to menus_to_review, which can only review auto_inject". #4772 gave the warning one remedy for both corpora: "judge a sample with menus_to_review". That is right for notes, where a line carries its passage. For rules it is a dead end: retrieval_review.REVIEWABLE holds only auto_inject, so the tool refuses every rule arm. - The reading is now per corpus (_NEVER_PULLED_READING). For rules, the text says: - the count includes rules that only arrived in a listing; - a rule can rightly be set aside on its trigger alone; - no judged sample exists for the rule arms; - a rule set aside again and again is a trigger to fix (update_rule when_to_apply, then what_might_apply), not a floor. - The tool docstring says the same. - The guard is tied to REVIEWABLE, so it can fail in both directions: if the rules text names menus_to_review while no rule arm is reviewable, or if a rule arm becomes reviewable and the text still says there is no sample. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
8b9ff0c5c1 |
Merge pull request 'Auto-inject looks up records named by number (#4796)' (#199) from dev into main
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 1m47s
CI & Build / Build & push image (push) Successful in 18s
|
||
|
|
f95ffae972 |
feat(retrieval): a record the operator names by number reaches the menu (#4796)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m47s
CI & Build / Build & push image (push) Successful in 37s
#4796 "A record the operator names by number reaches auto-inject only if its wording happens to match". "yes go ahead with 4448" says which record is meant, but a number means nothing to an embedding, so the prompt menu filled with records resembling the words around it. - record_refs.named_record_ids reads the operator's raw prompt, never the reply-enriched query: - `#N`, unless the word before it marks another numbering (PR, CI, rule, milestone, system, log...); - a bare number of 3 or more digits that opens the message, follows a reference word or continues a list one started; - never a quantity ("300 seconds"), a date, version, path or fenced code. - Each id is resolved through the ACL check; trashed or inaccessible ids are dropped. - A "Named in your message" block leads the menu: the kind, the System, the name and the opening of the body, or the seen pointer if the record is already on the ledger. It takes no share of top_k, and the semantic lines leave those ids out. - The block is booked under the new `named_ref` source, registered as an unbidden lookup that is allowed to be quiet. A named id that also ranked counts as suppressed in the auto_inject row, so #3668's identity holds. - Named records now arrive even when the search finds nothing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
a19c344770 |
Merge pull request 'No heading-only chunks, agent-message skip, refusals reach the agent (#4784, #4785, #4794)' (#198) from dev into main
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m14s
CI & Build / Python tests (push) Successful in 2m2s
CI & Build / Build & push image (push) Successful in 18s
|
||
|
|
5c6d9a7b55 |
fix(mcp): a tool's refusal reaches the agent with its reason (#4794)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 1m52s
CI & Build / Build & push image (push) Successful in 27s
SDK 2.x passes on only a ToolError's text. Any other exception becomes
UnexpectedToolError("Error executing tool X") and its message stays on the
server. ValueError is how every Scribe tool refuses — what was refused, why,
what to do instead — so every refusal reached the agent bare, and it could
only retry blind. Seen live: tune_retrieval(actor="operator") and
judge_menu(verdicts=[]) both answered "Error executing tool …" and nothing
else.
StrictArgsMCPServer.call_tool re-raises an UnexpectedToolError caused by a
ValueError as a ToolError carrying the message, in the SDK's own
"Error executing tool X: <reason>" shape. Other exceptions are crashes and
stay masked. The stale comment claiming the SDK returns ValueError text is
corrected.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
df4b37673f |
fix(plugin): a subagent hand-back is not an operator prompt (#4785)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 54s
CI & Build / TypeScript typecheck (push) Successful in 56s
CI & Build / Python tests (push) Successful in 1m49s
CI & Build / Build & push image (push) Successful in 29s
`<agent-message …>` turns were scored by auto-inject and the prompt rule arm as if the operator had typed them, spending a retrieval and writing a retrieval_logs row that reads as a real message (log 54355, found by the #4772 review). Same class as #4142. Matched up to the tag name, since the tag carries attributes; `<agent-messages>` and the tag named mid-sentence are still retrieved against. Plugin 2026.10.04.0125. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
78fa01b025 |
fix(embeddings): no chunk is a heading alone (#4784)
Judging live auto-inject menus (#4772) found "## Work log — 2026-08-21", "## Findings worth carrying forward" and "Two dev→main PRs this session." each embedded as a whole chunk. Text about nothing embeds close to everything, so they took ranks 1–3 on vague queries and pushed real records down. Three ways the chunker made them, each closed: - _split_paragraphs flushed a heading on its own when the paragraph under it was long. A thin lead is now carried into the paragraph that follows, and a hard cut never lands in the first quarter of the budget, where the newline it finds closes a heading. - A heading with no body (above its ### parts) became a section of its own when the section before it was full. A thin section now takes the next. - A short lead-in or a short last section stood alone. The lead-in joins what follows; a thin tail joins what precedes it. A continuation piece now repeats the LAST heading before it, since one section can now hold several. CHUNKER_VERSION 2 → 3, so the startup backfill re-embeds; tuned dials will report their shape as changed, which is true. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
c389ed61d2 |
Merge pull request 'The review drops records written after the call it re-runs (#4773)' (#197) from dev into main
CI & Build / Python lint (push) Successful in 5s
CI & Build / Plugin hooks (push) Successful in 18s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m59s
CI & Build / integration (push) Successful in 1m5s
CI & Build / Build & push image (push) Successful in 23s
|
||
|
|
5d0e97576b |
fix(retrieval): the review drops records written after the call it re-runs (#4773)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m8s
CI & Build / Python tests (push) Successful in 1m53s
CI & Build / Build & push image (push) Successful in 36s
The first live sample ranked records the call could never have been offered: the session that made a call writes the decision, often quoting the message, and that record tops the re-run. Judged, it inflates on_point exactly where the budget is decided. - _rerun over-fetches by POSTDATED_SLACK, drops records created after the call before ranking, and names them in `postdated`. - each line carries `changed_since_call` (updated_at or a work log after the call) and `logged` (shown fresh then). - judge_menu shares the re-run, so a post-dated record cannot be judged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
c50c1beb83 |
Merge pull request 'A review pass judges whether injected lines related (#4772)' (#196) from dev into main
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / integration (push) Successful in 1m19s
CI & Build / Python tests (push) Successful in 2m6s
CI & Build / Build & push image (push) Successful in 19s
|
||
|
|
6598c7fa85 |
feat(retrieval): a review pass judges whether injected lines related — menus_to_review, judge_menu and a judged readout (#4772)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 44s
An open rate cannot say whether a menu line related: every line carries its matched passage (#4364), so "not opened" covers unrelated, enough as shown, and already in context. #4772 "Injected notes are never judged". - retrieval_judgments (0115): a reviewer verdict per line of a logged call, on_point / adjacent / unrelated, with its reason, rank, budget side and whether the agent opened it within the hour. - menus_to_review re-runs a random sample of unjudged auto_inject calls with the arm's own parameters, past its budget, passage on every line. judge_menu records verdicts, re-deriving rank from a fresh re-run. - retrieval_telemetry gains a judged block (by rank, within/beyond budget, on_point_unopened). surfaced_never_pulled stops blaming titles. - missed-retrieval guidance names the review before a budget move. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
b4c637a7df |
Merge pull request 'Rulings readout and the prior-art 500 fix' (#195) from dev into main
CI & Build / Plugin hooks (push) Successful in 16s
CI & Build / Python lint (push) Successful in 2s
CI & Build / TypeScript typecheck (push) Successful in 59s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m43s
CI & Build / Build & push image (push) Successful in 20s
|
||
|
|
0a1bb68808 |
feat(rulings): system_usage_events is read back — per-System counts and a telemetry block (#4769)
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Python tests (push) Successful in 1m44s
CI & Build / Build & push image (push) Successful in 36s
#4769 "Rulings are counted where someone will read them": milestone 444 step 4 wrote system_usage_events and nothing read it. - retrieval_telemetry gains a `system_usage` block: surfacings and opens by source, distinct counts, and `by_system` naming the areas most shown. There is deliberately no pull-through ratio, because rulings travel in full in the line and opens are the exception. - usage_for_systems (one GROUP BY) adds `usage` to the REST Systems list and detail, and to MCP get_system. MCP list_systems is unchanged. - The Systems UI shows a "rulings shown N×" chip. - rulings_pre_tool, rulings_write_path and mcp_get_system are now declared registry points; the registry guard covers their recorders. - The Systems store merges a PATCH reply instead of replacing the row. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |