feat(plugin): a recognized retrieval miss has a route, and the record comes first (#4133)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Build & push image (push) Successful in 43s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Build & push image (push) Successful in 43s
The tuning loop shipped in steps 4 and 6 and logged zero events in its lifetime. `tune_retrieval`, `retrieval_telemetry` and `retrieval_surfaces` appeared on no instruction surface at all — not the skills, not the hooks, not the MCP instructions — so the decision that "the model should be the thing handling it 9 times out of 10" could not begin to happen. What was missing was not an auditor but a route. `using-scribe` now carries it, ordered: read the refused records, fix the trigger, and only then consider the dial. The order is the content. A rule's `when_to_apply` IS the text its score is computed against, so a miss is evidence about that text first; rewording one trigger changes one rule's reach, while moving a floor changes what every record on the surface does and cannot tell a badly-worded trigger from a genuinely distant one. Measured, and the reason the order is asserted rather than suggested: rule 1 scored 0.6515 and ranked 5th for the moment it governed, behind three rules that restrained the same act. Every percentile said "lower the floor"; at 0.60 the arm delivered those three restraints and still not rule 1. Rewriting the trigger to lead with the symptom put it 1st at 0.7130. Also: the create path gets a precondition. A new record is itself a retrieval-affecting act, so before writing one, what_might_apply asks what already covers that moment — fifty candidates and no bar, because a bar is what lets the existing record hide. `_INSTRUCTIONS` gets one index line, not the route: 1,986 of 2,000 characters, since Claude Code cuts the rest mid-word (#2562). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
@@ -142,6 +142,15 @@ TOPICS: tuple[Topic, ...] = (
|
||||
Topic("where a new rule goes, and its trigger", U, ("create_project_rule", "when_to_apply"),
|
||||
"whichever home it gets"),
|
||||
Topic("a rule vs the other entities", U, ("standing instruction",), "first ask whether it's a rule at all"),
|
||||
# Milestone 416 step 9: the tuning tools shipped in #4102/#4104 and were
|
||||
# named on NO instruction surface — measured, `retrieval_tuning_history`
|
||||
# returned zero events. Machinery with no route to it.
|
||||
Topic("a missed rule is a trigger to fix before a floor to move", U,
|
||||
("retrieval_telemetry", "tune_retrieval", "retrieval_surfaces"),
|
||||
"take it to the record first and the dial second",
|
||||
index=("retrieval_telemetry",)),
|
||||
Topic("ask what already covers a moment before writing a record", U,
|
||||
("what_might_apply",), "ask what already covers that moment"),
|
||||
Topic("reference notes update in place; dev-logs don't", U, ("reference note",),
|
||||
"state updates in place; chronicles don't"),
|
||||
# ── process arcs — owned by their skills ──
|
||||
|
||||
Reference in New Issue
Block a user