CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m8s
CI & Build / Python tests (push) Successful in 1m44s
CI & Build / Build & push image (push) Successful in 21s
Anthropic's skill guidance: keep SKILL.md under 500 lines, split into reference files linked one level deep as it nears that. using-scribe was 478 and every new practice lands there. - SKILL.md 478 -> 317 lines. It keeps orientation, one copy, the reflexes, scope, the judge section, UI and the process-skill index, plus a "Read these when the moment comes" list naming each file with its moment. - projects.md: binding a non-git directory (.scribe) and project inception. - writing-records.md: where a new rule goes, lesson growth, and notes that carry their own check (reflex 10 keeps a pointer). - missed-retrieval.md: the record-before-dial route, verbatim. - Text moved, not rewritten, except for the seams and one cross-reference. Tests: - tests.helpers.skill_text reads SKILL.md plus its reference files. The ownership registry, the miss-route and the verification tests use it, so a topic stays owned by its skill whichever file holds it. - The force test scans every skill .md on its own, since each file is read on its own. - New test_skill_structure: SKILL.md <= 350 lines, every reference file is linked from SKILL.md, none links another, and one over 100 lines opens with Contents. Each guard is shown to fail. The plugin version is minted. That also clears 4fb53b8's red Plugin hooks lane, which failed only because PACKAGING.md changed without a mint. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
160 lines
7.0 KiB
Python
160 lines
7.0 KiB
Python
"""A recognized retrieval miss has a route, and the route is ordered (#4133).
|
|
|
|
WHY THIS EXISTS
|
|
|
|
Milestone 416 built the tuning loop across steps 4 and 6 — `retrieval_surfaces`
|
|
to read a dial, `tune_retrieval` to move one with its argument attached,
|
|
`migrate_retrieval_floor` to carry it across a model change — and then measured,
|
|
on 2026-09-17, that `retrieval_tuning_history` returned **zero events**. Not one
|
|
call, ever.
|
|
|
|
The cause was not a tool being wrong. The three tools appeared on NO instruction
|
|
surface: not the skills, not the hooks, not the MCP `_INSTRUCTIONS`. They were
|
|
classified in `server.py`'s read/write lists and named nowhere a session would
|
|
read them. The operator's decision for #4102 — *"the model should be the thing
|
|
handling it 9 times out of 10"* — could not begin to happen.
|
|
|
|
So this file guards a ROUTE, not a behaviour: that a session which notices a
|
|
rule missed the moment it governed has somewhere written down to take it.
|
|
|
|
WHY THE ORDER IS THE LOAD-BEARING PART
|
|
|
|
`test_the_record_comes_before_the_dial` is the test that matters here, and it is
|
|
positional on purpose. Both halves of the route are individually reasonable —
|
|
read the telemetry, move the floor — and stated in either order they read as a
|
|
menu. They are not a menu. A rule's `when_to_apply` IS the text its score is
|
|
computed against, so a miss is evidence about that text first; a floor is one
|
|
number governing every record on the surface and cannot tell a badly-worded
|
|
trigger from a genuinely distant one.
|
|
|
|
Measured on this install, which is why the order is asserted rather than
|
|
suggested: rule 1 "`dev` is home" scored 0.6515 and ranked 5th for the moment it
|
|
governed, behind three rules that RESTRAINED the same act. Every percentile said
|
|
"lower the floor" — and at 0.60 the arm delivered those three restraints and
|
|
still not rule 1. Rewriting the trigger to lead with the symptom put the same
|
|
rule 1st at 0.7130. The dial-first path was available, statistically supported,
|
|
and wrong.
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
import pathlib
|
|
import re
|
|
|
|
ROOT = pathlib.Path(__file__).resolve().parents[1]
|
|
|
|
# The three tools that shipped with no route to them. Named together because
|
|
# the gap was all three at once, and a partial fix would leave the loop broken
|
|
# at whichever step lost its pointer.
|
|
TUNING_TOOLS = ("retrieval_telemetry", "update_rule", "retrieval_surfaces",
|
|
"tune_retrieval")
|
|
|
|
|
|
def _skill() -> str:
|
|
# The route lives in missed-retrieval.md, a reference file of using-scribe
|
|
# (#4398); the skill is SKILL.md and its references read together.
|
|
from tests.helpers import skill_text
|
|
return skill_text("using-scribe")
|
|
|
|
|
|
def _instructions() -> str:
|
|
src = (ROOT / "src/scribe/mcp/server.py").read_text()
|
|
match = re.search(r'_INSTRUCTIONS = """(.*?)"""', src, re.S)
|
|
assert match, "server.py no longer defines _INSTRUCTIONS as a literal"
|
|
return match.group(1)
|
|
|
|
|
|
def test_the_route_names_every_tool_it_depends_on():
|
|
"""The measured gap: machinery with no instruction surface pointing at it."""
|
|
text = _skill()
|
|
missing = [t for t in TUNING_TOOLS if t not in text]
|
|
assert not missing, (
|
|
f"using-scribe no longer names {missing}. These tools reach a session "
|
|
f"through this skill and nowhere else — when they were named on no "
|
|
f"surface at all, the tuning loop logged zero events in its lifetime."
|
|
)
|
|
|
|
|
|
def test_the_record_comes_before_the_dial():
|
|
"""THE GUARD. Stated in the wrong order, the route becomes a menu.
|
|
|
|
Positional rather than a presence check, because presence is exactly what
|
|
a reordering preserves. Asserted on first occurrence: what a reader meets
|
|
first is what they act on.
|
|
"""
|
|
text = _skill()
|
|
read_records = text.index("retrieval_telemetry")
|
|
fix_trigger = text.index("update_rule(when_to_apply")
|
|
move_dial = text.index("tune_retrieval")
|
|
|
|
assert read_records < move_dial, (
|
|
"the route reaches tune_retrieval before retrieval_telemetry. Reading "
|
|
"the refused records is the step that separates a real miss from a bar "
|
|
"doing its job — a percentile cannot, and has been measured pointing "
|
|
"the wrong way."
|
|
)
|
|
assert fix_trigger < move_dial, (
|
|
"the route reaches tune_retrieval before update_rule. A miss is "
|
|
"evidence about the trigger first: rewording one trigger changes one "
|
|
"rule's reach, while moving a floor changes every record on the surface."
|
|
)
|
|
|
|
|
|
def test_the_route_says_what_a_floor_costs_that_a_trigger_does_not():
|
|
"""The REASON for the order, not just the order.
|
|
|
|
Without it the ordering is arbitrary and the first session under deadline
|
|
will invert it. The asymmetry is the whole argument: one trigger against
|
|
every record on the surface.
|
|
"""
|
|
text = " ".join(_skill().split()).lower()
|
|
assert "take it to the record first and the dial second" in text
|
|
assert "changes one rule's reach" in text
|
|
assert "changes what every record on" in text
|
|
|
|
|
|
def test_a_session_can_notice_its_own_miss():
|
|
"""Both noticers, or the route only ever fires when the operator complains.
|
|
|
|
The self-noticed case is the one that scales — it does not need somebody
|
|
watching — and it is the one a surface written only around operator
|
|
feedback silently omits.
|
|
"""
|
|
text = " ".join(_skill().split()).lower()
|
|
assert "either noticer" in text
|
|
assert "reached for a rule nobody offered you" in text
|
|
# Both directions too: a rule that never arrives and one that always does
|
|
# are the same misjudgment, and they point at different dials.
|
|
assert "either direction counts" in text
|
|
|
|
|
|
def test_writing_a_new_record_asks_what_already_covers_the_moment():
|
|
"""The create path is retrieval-affecting, and nothing said so.
|
|
|
|
A second record on one moment splits a single budget between two claims,
|
|
and when the copies differ in force the softer one teaches a session that
|
|
binding guidance is optional. `what_might_apply` rather than `search`
|
|
because a bar is exactly what lets the existing record hide.
|
|
"""
|
|
text = " ".join(_skill().split()).lower()
|
|
assert "ask what already covers that moment" in text
|
|
assert "what_might_apply" in text
|
|
assert "no bar" in text
|
|
|
|
|
|
def test_the_route_stays_off_the_index():
|
|
"""`_INSTRUCTIONS` is the only surface every MCP client gets (decision #4027).
|
|
|
|
It carried a one-line MISSED pointer to this route until #4389, which found
|
|
the index was being used as a rulebook: a missed rule is noticed mid-work,
|
|
not at session start, and the spec asks server instructions for the
|
|
workflow across tools, not stance. using-scribe states the route (the tests
|
|
above pin it); the index keeps what_might_apply, the tool that opens it.
|
|
|
|
Whitespace-flattened before matching: the block is hard-wrapped, so a
|
|
phrase straddles a line break the moment anything before it changes length
|
|
(#4103).
|
|
"""
|
|
text = " ".join(_instructions().split()).lower()
|
|
assert "what_might_apply" in text
|
|
assert "take it to the record first" not in text
|