c6cdfc2172203f698fe469a93f1d19c31e1fb00e
9
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c6cdfc2172 |
feat(moments): the reply moment holds a finished reply for one read (milestone 458 step 4b, #4922)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / integration (push) Successful in 1m1s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Python tests (push) Failing after 1m25s
CI & Build / Build & push image (push) Skipped
The reply is the one act no tool call marks, and it is where "let me know if it works" gets said. A new Stop hook (scribe_reply_check.sh) sends the finished reply to POST /api/plugin/reply-rules, which checks it twice: - mounted: every unopened RULE on reply.report, plus reply.ask when the reply asks a question. Deterministic. - semantic: the reply's head and tail against every rule's trigger, on a new ranked surface, reply_rule. It is the backstop for whatever the earlier arms missed. Its floor is its stop bar (default 0.80, budget 1), with its own Settings dials. The new stop_only stage records surfacing for the rule that holds and nothing else, because nothing else reached anyone. Following the operator's ruling from 456 step 8, a rule that holds blocks once, in the server's words. The hook blocks only on a reason it was given, so an unreachable instance never stops a session, and it never holds the rewrite. The ledger is the act checkpoint's own, so a rule holds a session once across both doors and the per-session cap counts both. The turn reader moved from the report check into scribe_defs.sh (scribe_turn_facts / scribe_turn_fact), so the two Stop hooks read a turn the same way. The output was checked identical on a real transcript. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
78653130d6 |
feat(moments): mounted rules arrive when their moment happens, through every door (milestone 458 step 4a, #4922)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 57s
CI & Build / Python tests (push) Failing after 1m19s
CI & Build / Build & push image (push) Skipped
A rule mounted on a moment now reaches the session when an act reaches
that moment, with no semantic match involved:
- run_moment_arm on the pipeline: a lookup, not a ranked search. Each
line names the moment and the act that reached it ("at work.deliver,
reached by `git push`"), so a misfire is visible where it lands and
can be unmapped in-session. A repeat is cited, not quoted; fresh
rules are recorded surfaced under source moment_rule with the moment
in detail. No retrieval_logs row, as for the other lookups, so no
latency is persisted for this arm.
- rule_scope: a rule's home clause, moved out of semantic_search_rules
so the moment lookup scopes by the same one.
- rulebooks.rules_on_moments / mounted_moments.
- The plugin door: a catch-all PreToolUse hook (scribe_moment.sh). It
keeps /moment-tools' answer on disk for five minutes, so a call to a
tool that cannot reach a mounted rule sends nothing, and an install
that has mounted nothing sends one request per window. It shares the
rules ledger with the other arms and fails open silently.
- The MCP door: Scribe's own tools named by the shipped mappings carry
moment_rules in their response, so a client without the plugin gets
them too. The hook skips those tools. A guard pins the attach on
every one.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
0720ab6dcf |
refactor(retrieval): the completion-report preferences run through the pipeline (milestone 456 step 3, #4905)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Build & push image (push) Successful in 28s
reply_preferences.completion_preferences was the last hand-written rule search + record_retrieval + record_rule_surfaced triple outside the pipeline. It is now REPORT_PREFERENCE, a RuleArm with kind="preference" over COMPLETION_QUERY, read back as records rather than as lines. - RuleArm gains `kind`. _ranked asks the ranker for that kind and drops anything else on the way out, before the row is logged. - RuleResult gains `shown`, the (score, rule) pairs behind the lines, for callers that return records. - RuleMoment.project_id may be None: logged as given, searched as `project_id or None`. report_preference rows keep their NULL project. - Guards: reply_preferences now has 0 direct rule searches. The registry constant-source example moves from report_preference to preference_slot, the pipeline slot that records through a module constant. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f6b824b214 |
refactor(retrieval): the three rule arms run through one pipeline (milestone 456 step 2, #4904)
CI & Build / Plugin hooks (push) Successful in 19s
CI & Build / Python lint (push) Successful in 2s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / Python tests (push) Failing after 1m28s
CI & Build / Build & push image (push) Skipped
CI & Build / integration (push) Successful in 1m15s
prompt_rule, pre_tool_rule and the rule half of the write path each wrote the same steps out by hand: search, band, split fresh from repeats, log the call before any early return, reserve a slot, render, record surfacings. The copies drifted, and #3497, #3750 and #3752 were fixed one copy at a time. - New services/retrieval_pipeline.py. run_rule_arm runs the stages in one order. RuleArm is the spec (source, band, compact_tail, checkpoint, preference_slot). RuleMoment is the query and the session ledger. - The I/O (ranker and recorders) is passed in as RuleIO. plugin_context resolves it from its own names at call time, so existing patch points still apply. - _rule_band, _rule_hint_line, checkpoint_for, checkpoint_reason, the band constant and the preference slot moved into the pipeline unchanged. plugin_context re-exports them. - A recorder that raises now costs only its row, never a rendered line. - The flags reproduce today exactly. Whether prompt_rule should band, and whether the act arms should reserve a preference, are step 7. - Registry: the pipeline call sites are declared in FAN_OUT_SITES, with values read from the specs. - The #3497 structural guard now checks the one implementation, and that plugin_context writes no rule-source row of its own. plugin_context.py: 3,558 -> 2,980 lines. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f95ffae972 |
feat(retrieval): a record the operator names by number reaches the menu (#4796)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m47s
CI & Build / Build & push image (push) Successful in 37s
#4796 "A record the operator names by number reaches auto-inject only if its wording happens to match". "yes go ahead with 4448" says which record is meant, but a number means nothing to an embedding, so the prompt menu filled with records resembling the words around it. - record_refs.named_record_ids reads the operator's raw prompt, never the reply-enriched query: - `#N`, unless the word before it marks another numbering (PR, CI, rule, milestone, system, log...); - a bare number of 3 or more digits that opens the message, follows a reference word or continues a list one started; - never a quantity ("300 seconds"), a date, version, path or fenced code. - Each id is resolved through the ACL check; trashed or inaccessible ids are dropped. - A "Named in your message" block leads the menu: the kind, the System, the name and the opening of the body, or the seen pointer if the record is already on the ledger. It takes no share of top_k, and the semantic lines leave those ids out. - The block is booked under the new `named_ref` source, registered as an unbidden lookup that is allowed to be quiet. A named id that also ranked counts as suppressed in the auto_inject row, so #3668's identity holds. - Named records now arrive even when the search finds nothing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
0a1bb68808 |
feat(rulings): system_usage_events is read back — per-System counts and a telemetry block (#4769)
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Python tests (push) Successful in 1m44s
CI & Build / Build & push image (push) Successful in 36s
#4769 "Rulings are counted where someone will read them": milestone 444 step 4 wrote system_usage_events and nothing read it. - retrieval_telemetry gains a `system_usage` block: surfacings and opens by source, distinct counts, and `by_system` naming the areas most shown. There is deliberately no pull-through ratio, because rulings travel in full in the line and opens are the exception. - usage_for_systems (one GROUP BY) adds `usage` to the REST Systems list and detail, and to MCP get_system. MCP list_systems is unchanged. - The Systems UI shows a "rulings shown N×" chip. - rulings_pre_tool, rulings_write_path and mcp_get_system are now declared registry points; the registry guard covers their recorders. - The Systems store merges a PATCH reply instead of replacing the row. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
fd2ebf4c13 |
feat(rules): a rule surfaces through the lessons confirmed as its instances (milestone 440 step 4, #4633)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 18s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 55s
CI & Build / Python tests (push) Failing after 1m16s
CI & Build / Build & push image (push) Skipped
On all three rule arms (prompt_rule, pre_tool_rule, write_path_rule), after the direct match, a lesson matching the moment at the notes menu's own bar brings the rule(s) it is CONFIRMED to be an instance of, rendered in rule voice with "Reached through lesson #N “…”, a recorded instance of it." - Confirmed links only (lesson_rules.confirmed_lessons / confirmed_rules_in_scope). A suggested link carrying its rule would manufacture the co-arrival #4637 counts and prove itself. - Scope kept: a lesson surfaces everywhere, but the rule it brings must be global or this project's own — milestone 414's boundary, not reopened through a side door. - Suppression is by RULE: anything the direct band named or the session ledger holds is skipped, whichever lesson reached it. - Its own slot (VIA_LESSON_LIMIT = 1), not a rule slot — a stated default, since the plan's "decide by measurement" has nothing to measure until links are confirmed (#4632). Its own source, rule_via_lesson: registered, RANKED for pull-through, logged whenever it searches. - Nothing searched at all while no lesson carries a confirmed link, which is every install until one is judged. - build_prompt_rule_hint / build_tool_rule_hint are now thin wrappers over the direct arms (_prompt_rule_hint / _tool_rule_hint); the arm leaves its query in `_via_query` only when it ran, so a disabled arm or a blank prompt brings no rule in either, and the key never leaves the server. - Via-lesson rules are not fed to the #4637 co-surfacing recorder: only direct matches are evidence. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f8e53c1c35 |
fix(telemetry): a warning fired on an arm whose decline rate is arithmetic, not evidence (#4232)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / integration (push) Successful in 49s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m33s
CI & Build / Build & push image (push) Successful in 26s
Found by reading a live `retrieval_telemetry` readout after milestone 419
deployed, not by inspection. The readout said:
cannot_decline / report_preference — "45 calls, 0 of them returned
nothing. An arm that fires unasked has to be able to say nothing; this one
never has. Check that it applies its floor at all."
And printed, beside it, that arm's band: p10 = p50 = p90 = min = max = 0.791.
FIVE IDENTICAL PERCENTILES IS THE TELL. That is not a ranking, it is one
record at one score on every call — because `report_preference` searches a
fixed string (`reply_preferences.COMPLETION_QUERY`, a module constant, and
deliberately so).
For a fixed query against a stable corpus the top score is a CONSTANT, so the
arm's decline rate is 0% or 100% and never in between; which of the two it is
depends only on where the bar sits relative to that one number. "Never
returned nothing" is therefore arithmetic, not evidence, and the warning's own
remedy — check whether it applies a floor — cannot be answered from it.
The arm already knew this about itself; the warning did not:
"a fixed query makes this arm's score a constant and a floor a hair above
it produces a dead arm no amount of traffic will ever reveal"
— services/reply_preferences.py
THIS CLASS OF BUG ALREADY HAS A GUARD, which is the argument for the shape of
the fix. `Point.logs_unconditionally` exists because of #3497: both rule arms
once logged only their hits, so their zero count was structurally 0 and this
same warning would have fired on a LOGGING property while sending the reader
to move a threshold that was never involved. This is that one step over — a
QUERY-SHAPE property — and gets the same treatment: a declared field on
`Point`, and exclusion rather than trust.
AND THE WARNING THAT WOULD BE INFORMATIVE HERE DID NOT EXIST. For a fixed-query
arm the dangerous state is the mirror image: every call empty, meaning the bar
is above the constant and no further traffic will ever move it. The arm is off
rather than quiet, and nothing in the readout said so — `expects_traffic`
covers an arm with NO calls, not one with calls and a 100% decline rate. That
state is real and reached: `report_preference` once logged 69 consecutive
declines at 0.0006 under its bar.
So `fixed_query_never_clears` sends the reader to `near_miss_samples` and not
to the dial — because that incident is also the one where the statistic and
the correct action pointed opposite ways. Every percentile said lower the
floor; opening the refused record showed it was rule 77 arriving as a false
positive, and lowering it would have delivered that rule on every completion
report ever written.
Guards in tests/test_retrieval_warnings.py, including the falsifier that
matters most here: `cannot_decline` must still fire on an arm whose query
varies, or this change is a disabled check wearing a narrowed one's clothes.
Both boundaries tested from both sides, per that module's own standard.
The new code is documented in the `retrieval_telemetry` tool docstring beside
the others (rule 33) — an undocumented code in a readout is a reader meeting a
verdict with no way to disagree with it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
|
||
|
|
e2c3a5c2b5 |
feat(telemetry): retrieval_telemetry says what is wrong (#3431)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 8s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m30s
CI & Build / Build & push image (push) Successful in 23s
The tool returned distributions and left the reading to the caller, so every readout was the same four checks done by hand — #3430's baseline, #3835's rule near-misses, the #1038 rerank gate. Mechanical, and therefore forgettable. Tonight's acceptance pass on #3898 was the case for doing this. Reading it by hand meant catching that two surfaces had `covers_window: false`, that `prompt_rule`'s floor had moved three times inside the window (which made the readout self-contradictory: deliveries at 0.622 beside refusals at 0.7199), and that 15 of 20 near-misses were one record against text no operator wrote. Miss any of those and the obvious conclusion was "the bar is too tight" — a floor change that would have injected one preference into every notification. `warnings` is always present and empty when clean, so its emptiness is an answer rather than a gap. Each entry carries the numbers that produced it: "345 calls, 0 declined" is the analysis, "check write_path_rule" is an instruction to redo it. Five codes — cannot_decline, band_hugs_floor, no_duration, surfaced_never_pulled, unregistered_source. cannot_decline has three guards, each a bug it would otherwise cause. Asked surfaces are exempt (a search returning a list every time is working). An arm not known to log unconditionally is exempt — that is #3497 exactly, where both rule arms recorded only their hits, so a decline count of zero was a LOGGING defect and this warning would have sent the reader to a threshold that was never involved. Unregistered sources get numbers but no verdict. `silent_surfaces` is the half the rows cannot show: an arm that emitted nothing is invisible to every row-based check and looks exactly like an arm that does not exist. It is driven by a new declared registry, `retrieval_registry.POINTS` — deliberately NOT `retrieval_surfaces.SURFACES`, which answers "what can be tuned" and excludes the reserved slots because a budget of 1 is their feature. This answers "what can be measured", and the reserved slots belong in it precisely because they are judgeable without being tunable. A test asserts the two cannot drift apart. The registry test derives sources from the call sites with `ast`, not grep, and the difference is not theoretical: `wide_net` and `report_preference` reach their recorder as `source=SOURCE` through a module constant, so a grep for `source="` is blind to both — the narrowing #3191 warns about. Three sites pass `source` as a variable and are declared in FAN_OUT_SITES; the test pins those sites but not the values they can pass, which is why the `unregistered_source` warning exists to catch the rest at first fire. Thresholds are settings (rule 25) defaulted so a fresh install with almost no data produces no warnings at all (rule 115) — a new user's first readout naming five broken things would be describing the emptiness. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy |