feat(plugin): a recognized retrieval miss has a route, and the record comes first (#4133)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Build & push image (push) Successful in 43s

The tuning loop shipped in steps 4 and 6 and logged zero events in its
lifetime. `tune_retrieval`, `retrieval_telemetry` and `retrieval_surfaces`
appeared on no instruction surface at all — not the skills, not the hooks,
not the MCP instructions — so the decision that "the model should be the
thing handling it 9 times out of 10" could not begin to happen.

What was missing was not an auditor but a route. `using-scribe` now carries
it, ordered: read the refused records, fix the trigger, and only then
consider the dial. The order is the content. A rule's `when_to_apply` IS
the text its score is computed against, so a miss is evidence about that
text first; rewording one trigger changes one rule's reach, while moving a
floor changes what every record on the surface does and cannot tell a
badly-worded trigger from a genuinely distant one.

Measured, and the reason the order is asserted rather than suggested: rule
1 scored 0.6515 and ranked 5th for the moment it governed, behind three
rules that restrained the same act. Every percentile said "lower the
floor"; at 0.60 the arm delivered those three restraints and still not rule
1. Rewriting the trigger to lead with the symptom put it 1st at 0.7130.

Also: the create path gets a precondition. A new record is itself a
retrieval-affecting act, so before writing one, what_might_apply asks what
already covers that moment — fifty candidates and no bar, because a bar is
what lets the existing record hide.

`_INSTRUCTIONS` gets one index line, not the route: 1,986 of 2,000
characters, since Claude Code cuts the rest mid-word (#2562).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
2026-09-18 00:37:04 -04:00
co-authored by Claude Opus 5
parent a7d736860f
commit 104c1d6f37
5 changed files with 226 additions and 1 deletions
+9
View File
@@ -142,6 +142,15 @@ TOPICS: tuple[Topic, ...] = (
Topic("where a new rule goes, and its trigger", U, ("create_project_rule", "when_to_apply"),
"whichever home it gets"),
Topic("a rule vs the other entities", U, ("standing instruction",), "first ask whether it's a rule at all"),
# Milestone 416 step 9: the tuning tools shipped in #4102/#4104 and were
# named on NO instruction surface — measured, `retrieval_tuning_history`
# returned zero events. Machinery with no route to it.
Topic("a missed rule is a trigger to fix before a floor to move", U,
("retrieval_telemetry", "tune_retrieval", "retrieval_surfaces"),
"take it to the record first and the dial second",
index=("retrieval_telemetry",)),
Topic("ask what already covers a moment before writing a record", U,
("what_might_apply",), "ask what already covers that moment"),
Topic("reference notes update in place; dev-logs don't", U, ("reference note",),
"state updates in place; chronicles don't"),
# ── process arcs — owned by their skills ──
+158
View File
@@ -0,0 +1,158 @@
"""A recognized retrieval miss has a route, and the route is ordered (#4133).
WHY THIS EXISTS
Milestone 416 built the tuning loop across steps 4 and 6 — `retrieval_surfaces`
to read a dial, `tune_retrieval` to move one with its argument attached,
`migrate_retrieval_floor` to carry it across a model change — and then measured,
on 2026-09-17, that `retrieval_tuning_history` returned **zero events**. Not one
call, ever.
The cause was not a tool being wrong. The three tools appeared on NO instruction
surface: not the skills, not the hooks, not the MCP `_INSTRUCTIONS`. They were
classified in `server.py`'s read/write lists and named nowhere a session would
read them. The operator's decision for #4102 — *"the model should be the thing
handling it 9 times out of 10"* — could not begin to happen.
So this file guards a ROUTE, not a behaviour: that a session which notices a
rule missed the moment it governed has somewhere written down to take it.
WHY THE ORDER IS THE LOAD-BEARING PART
`test_the_record_comes_before_the_dial` is the test that matters here, and it is
positional on purpose. Both halves of the route are individually reasonable —
read the telemetry, move the floor — and stated in either order they read as a
menu. They are not a menu. A rule's `when_to_apply` IS the text its score is
computed against, so a miss is evidence about that text first; a floor is one
number governing every record on the surface and cannot tell a badly-worded
trigger from a genuinely distant one.
Measured on this install, which is why the order is asserted rather than
suggested: rule 1 "`dev` is home" scored 0.6515 and ranked 5th for the moment it
governed, behind three rules that RESTRAINED the same act. Every percentile said
"lower the floor" — and at 0.60 the arm delivered those three restraints and
still not rule 1. Rewriting the trigger to lead with the symptom put the same
rule 1st at 0.7130. The dial-first path was available, statistically supported,
and wrong.
"""
from __future__ import annotations
import pathlib
import re
ROOT = pathlib.Path(__file__).resolve().parents[1]
SKILL = ROOT / "plugin/skills/using-scribe/SKILL.md"
# The three tools that shipped with no route to them. Named together because
# the gap was all three at once, and a partial fix would leave the loop broken
# at whichever step lost its pointer.
TUNING_TOOLS = ("retrieval_telemetry", "update_rule", "retrieval_surfaces",
"tune_retrieval")
def _skill() -> str:
return SKILL.read_text()
def _instructions() -> str:
src = (ROOT / "src/scribe/mcp/server.py").read_text()
match = re.search(r'_INSTRUCTIONS = """(.*?)"""', src, re.S)
assert match, "server.py no longer defines _INSTRUCTIONS as a literal"
return match.group(1)
def test_the_route_names_every_tool_it_depends_on():
"""The measured gap: machinery with no instruction surface pointing at it."""
text = _skill()
missing = [t for t in TUNING_TOOLS if t not in text]
assert not missing, (
f"using-scribe no longer names {missing}. These tools reach a session "
f"through this skill and nowhere else — when they were named on no "
f"surface at all, the tuning loop logged zero events in its lifetime."
)
def test_the_record_comes_before_the_dial():
"""THE GUARD. Stated in the wrong order, the route becomes a menu.
Positional rather than a presence check, because presence is exactly what
a reordering preserves. Asserted on first occurrence: what a reader meets
first is what they act on.
"""
text = _skill()
read_records = text.index("retrieval_telemetry")
fix_trigger = text.index("update_rule(when_to_apply")
move_dial = text.index("tune_retrieval")
assert read_records < move_dial, (
"the route reaches tune_retrieval before retrieval_telemetry. Reading "
"the refused records is the step that separates a real miss from a bar "
"doing its job — a percentile cannot, and has been measured pointing "
"the wrong way."
)
assert fix_trigger < move_dial, (
"the route reaches tune_retrieval before update_rule. A miss is "
"evidence about the trigger first: rewording one trigger changes one "
"rule's reach, while moving a floor changes every record on the surface."
)
def test_the_route_says_what_a_floor_costs_that_a_trigger_does_not():
"""The REASON for the order, not just the order.
Without it the ordering is arbitrary and the first session under deadline
will invert it. The asymmetry is the whole argument: one trigger against
every record on the surface.
"""
text = " ".join(_skill().split()).lower()
assert "take it to the record first and the dial second" in text
assert "changes one rule's reach" in text
assert "changes what every record on" in text
def test_a_session_can_notice_its_own_miss():
"""Both noticers, or the route only ever fires when the operator complains.
The self-noticed case is the one that scales — it does not need somebody
watching — and it is the one a surface written only around operator
feedback silently omits.
"""
text = " ".join(_skill().split()).lower()
assert "either noticer" in text
assert "reached for a rule nobody offered you" in text
# Both directions too: a rule that never arrives and one that always does
# are the same misjudgment, and they point at different dials.
assert "either direction counts" in text
def test_writing_a_new_record_asks_what_already_covers_the_moment():
"""The create path is retrieval-affecting, and nothing said so.
A second record on one moment splits a single budget between two claims,
and when the copies differ in force the softer one teaches a session that
binding guidance is optional. `what_might_apply` rather than `search`
because a bar is exactly what lets the existing record hide.
"""
text = " ".join(_skill().split()).lower()
assert "ask what already covers that moment" in text
assert "what_might_apply" in text
assert "no bar" in text
def test_the_index_points_at_the_route_without_restating_it():
"""`_INSTRUCTIONS` is the only surface every MCP client gets (decision #4027).
It indexes; using-scribe states. A client with no Agent Skills support still
learns the ordering exists and which tool opens it.
Whitespace-flattened before matching. The block is hard-wrapped to fit a
2,000-character budget, so a phrase straddles a line break the moment
anything before it changes length — #4103 shipped a guard that broke
exactly that way, on text nobody had touched.
"""
text = " ".join(_instructions().split())
assert "retrieval_telemetry" in text
assert "not a floor to move" in text
# The route itself must NOT be here — 2,000 characters is the whole budget
# and Claude Code cuts the rest mid-word (#2562).
assert "take it to the record first" not in text.lower()