feat(retrieval): the wide net becomes a pull — fifty candidates, no bar (#4103)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Failing after 1m8s
CI & Build / Build & push image (push) Skipped
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Failing after 1m8s
CI & Build / Build & push image (push) Skipped
Milestone 416 step 5. The operator's compromise — "if you're worried about excluding potentially important data let limit it to 50 entries" — moved to the surface where it is safe. Fifty in the push would be milestone 394 with extra steps; fifty in a pull crowds nothing out. `what_might_apply(query)` returns ranked rule candidates with NO threshold. The moment it serves is the one where the caller does not trust a bar to decide for them, so it does not have one — every row carries its score and the reader judges. WHY IT IS NOT A BIGGER `limit` ON `search` `_search_rules` returns statement, why and how_to_apply in full, on the stated reasoning that a caller who went looking deserves the whole record. That is the DEEP pull and should stay that way. This is the SHALLOW one — many candidates, each just enough to decide whether to open it. Opposite trade-offs, so it is a second tool. THE PREMISE THE STEP GOT WRONG The task said fifty "costs nothing". `_rule_hint_line` had already measured otherwise: ~143 tokens per line once the trigger is rendered, and #3855 tripled trigger lengths across the corpus. Fifty is ~7,000 tokens — cheap next to an arm firing before every Bash call, but not free, and a tool promising a free wide net gets reached for casually and then regretted. So it reuses the graduated shape #3851 measured for the push: the top few carry their trigger whole, the rest carry a cut of it. TRUNCATED, never dropped — the trigger is what lets a reader judge without opening, and a teaser without one is just an id. The cut borrows `_goal_line`'s technique including the fallback that matters (#4036): `textwrap.shorten` returns a bare "…" for one unbroken word. TELEMETRY Logged under its own pull source, asserted absent from both the tunable push registry and AMBIENT_SOURCES. That guard is load-bearing right now: the push arms' near-miss distributions are the evidence #4121 argues from, and a pull folded into them would move those numbers. INSTRUCTION SURFACES The using-scribe reflex and the MCP instructions both pointed at `search(content_type="rule")` for the consequential moment — the deep tool, at the moment you want breadth. They now point here, and keep `search` for reading a rule you already suspect. Written as a practice rather than a prohibition (rule 165). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
@@ -1,7 +1,7 @@
|
||||
{
|
||||
"name": "scribe",
|
||||
"description": "Scribe for Claude Code: connects the scribe MCP server, adds the hooks that deliver live project state and relevant records at the right moment, ships the shared client-neutral Scribe skills (using-scribe, writing-plans, reporting-back, systematic-debugging, verification, brainstorming, reusing-code, shape-accounting), and syncs your saved Scribe Processes as skills (/scribe:sync).",
|
||||
"version": "2026.09.17.0111",
|
||||
"version": "2026.09.17.1618",
|
||||
"author": {
|
||||
"name": "Bryan Van Deusen"
|
||||
},
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
name: using-scribe
|
||||
description: Use at the START of every session, and before answering anything about the operator's work or starting any task — establishes the Scribe-first reflex. You hold none of the operator's rules: they arrive by retrieval when your work matches one, and search(content_type="rule") is how you ask before a consequential act. Call enter_project when a repo/project is in scope. Then recall before acting, update over duplicate, plan in Scribe not in files.
|
||||
description: Use at the START of every session, and before answering anything about the operator's work or starting any task — establishes the Scribe-first reflex. You hold none of the operator's rules: they arrive by retrieval when your work matches one, and what_might_apply is how you ask before a consequential act — it returns the wide net of candidates with no bar. Call enter_project when a repo/project is in scope. Then recall before acting, update over duplicate, plan in Scribe not in files.
|
||||
---
|
||||
|
||||
# Using Scribe
|
||||
@@ -73,9 +73,19 @@ Two constraints on *how* that's achieved:
|
||||
|
||||
So "no rule arrived" means "nothing matched", never "no rule exists" — an
|
||||
empty session is not evidence of an empty rulebook. Retrieval fires when
|
||||
something asks: before a consequential act, `search(content_type="rule")` on
|
||||
what you are about to do, and pull a record's full statement with
|
||||
`get_rule(id)` when it is about to bite. When a project is in scope, pass
|
||||
something asks, so ask — and reach for the tool that fits the moment:
|
||||
|
||||
- **Before a consequential act**, and before handing work back because you
|
||||
are unsure you may finish it: `what_might_apply("what you are about to
|
||||
do")`. It returns up to fifty ranked candidates with NO bar. The arms that
|
||||
push rules at you spend a budget of three and say nothing about what sat
|
||||
just underneath — right for something firing before every command, wrong
|
||||
for the one moment you actually want to be sure. Expect the tail to be
|
||||
noise; you are reading for the one record you would have missed.
|
||||
- **When you already suspect a particular rule**: `search(content_type=
|
||||
"rule")` for full statements, or `get_rule(id)` when it is about to bite.
|
||||
|
||||
When a project is in scope, pass
|
||||
its `project_id`: the answer is then the global rules plus that project's
|
||||
own, never another project's. `enter_project(id)` lists the project's own
|
||||
rules by title.
|
||||
|
||||
@@ -44,8 +44,11 @@ client reads Agent Skills) and in each tool's description. The index:
|
||||
system. An `inception` key: ask what it inherits, then
|
||||
decide_project_inception.
|
||||
- RULES: nothing preloads; a rule arrives when your work matches it. Before a
|
||||
consequential act, search(content_type="rule"). Silence means
|
||||
nothing matched, not none. Rules bind; preferences guide.
|
||||
consequential act — or before handing work back unsure you may finish it —
|
||||
what_might_apply("what you are about to do"): the wide net, fifty ranked
|
||||
candidates, no bar. search(content_type="rule") reads one you already
|
||||
suspect. Silence means nothing matched, not none. Rules bind; preferences
|
||||
guide.
|
||||
- RECALL: search before acting, scoped with the active project_id.
|
||||
- RECORD: create_task; a fix is kind="issue". add_task_log as you go; status
|
||||
in_progress on start, done on finish. Tag system_ids as you write.
|
||||
@@ -137,6 +140,11 @@ _READ_ONLY_TOOLS = frozenset({
|
||||
# usual for these two: a session that cannot see the bar in force, or the
|
||||
# reason it was last moved, is a session that will move it again blind.
|
||||
"retrieval_surfaces", "retrieval_tuning_history",
|
||||
# The wide net (#4103) — ranked rule candidates with no bar, for the
|
||||
# moment before a consequential act. A pure read, and one a read key needs
|
||||
# most: it is the surface the "ask before acting" reflex calls, and a key
|
||||
# that could not reach it would be denied exactly the check it should run.
|
||||
"what_might_apply",
|
||||
})
|
||||
|
||||
# Every tool that WRITES, by name. Nothing reads this set at runtime — a tool
|
||||
|
||||
@@ -6,6 +6,7 @@ from `mcp.server.build_mcp_server`.
|
||||
"""
|
||||
from scribe.mcp.tools import (
|
||||
design_systems, milestones, notes, processes, projects, recent, repos, retrieval_tuning,
|
||||
wide_net,
|
||||
rulebooks, search, shapes, snippets, systems, tags, tasks, trash,
|
||||
)
|
||||
|
||||
@@ -14,6 +15,7 @@ def register_all(mcp) -> None:
|
||||
"""Register every tool module's tools on the given FastMCP instance."""
|
||||
search.register(mcp)
|
||||
retrieval_tuning.register(mcp)
|
||||
wide_net.register(mcp)
|
||||
notes.register(mcp)
|
||||
tasks.register(mcp)
|
||||
projects.register(mcp)
|
||||
|
||||
@@ -0,0 +1,218 @@
|
||||
"""The wide net — every candidate that might govern what you are about to do (#4103).
|
||||
|
||||
WHY THIS EXISTS
|
||||
|
||||
Milestone 416 step 5, from the operator's compromise:
|
||||
|
||||
"if you're worried about excluding potentially important data let limit it
|
||||
to 50 entries or something like that"
|
||||
|
||||
Fifty in the PUSH would be milestone 394 with extra steps — a wall the reader
|
||||
skims, with the governing record indistinguishable from forty-nine others.
|
||||
Fifty in a PULL is a different object: it arrives only when a session asks, it
|
||||
crowds nothing out, and it is the honest home for "do not exclude potentially
|
||||
important data".
|
||||
|
||||
WHY IT IS NOT `search(content_type="rule")`
|
||||
|
||||
That tool is the DEEP pull and should stay that way — `_search_rules` returns
|
||||
`statement`, `why` and `how_to_apply` in full, on the reasoning that a caller
|
||||
who went looking deserves the whole record rather than a summary to re-fetch.
|
||||
Six results already run to thousands of tokens.
|
||||
|
||||
This is the SHALLOW pull: many candidates, each just enough to decide whether
|
||||
to open it. Opposite trade-off, so it is a second tool rather than a bigger
|
||||
`limit` on the first.
|
||||
|
||||
WHY THE TAIL IS TRUNCATED RATHER THAN FULL
|
||||
|
||||
Because the step's premise needed correcting. The task said fifty "costs
|
||||
nothing", and `_rule_hint_line` had already measured otherwise: a line runs
|
||||
~143 tokens once its trigger is rendered, and #3855 tripled trigger lengths
|
||||
across the corpus. Fifty of those is ~7,000 tokens — cheap next to an arm that
|
||||
fires before every Bash call, but not free, and a tool that promises a free
|
||||
wide net gets reached for casually and then regretted.
|
||||
|
||||
So this reuses the graduated shape #3851 measured for the push: the top few
|
||||
carry their trigger in full, the rest carry a cut of it. TRUNCATED, never
|
||||
dropped — the trigger is what lets a reader judge relevance without opening
|
||||
the record, and a teaser without one is just an id.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import textwrap
|
||||
import time
|
||||
|
||||
from scribe.mcp._context import current_user_id
|
||||
from scribe.services.embeddings import semantic_search_rules
|
||||
from scribe.services.retrieval_telemetry import record_retrieval
|
||||
|
||||
# The telemetry `source`, and a PULL — never added to AMBIENT_SOURCES (those
|
||||
# are deliveries nobody chose) and never confused with a push arm. Keeping it
|
||||
# separate is load-bearing right now: the push arms' near-miss distributions
|
||||
# are the evidence #4121 rests on, and a pull mixed into them would move the
|
||||
# very numbers that step is arguing from.
|
||||
SOURCE = "wide_net"
|
||||
|
||||
# No bar. That is the point of the tool rather than an oversight — the caller
|
||||
# asked for the wide net precisely because they do not trust a bar to decide
|
||||
# for them here. Every row carries its score, so the reader sees where the
|
||||
# ranking falls off and judges it themselves.
|
||||
THRESHOLD = 0.0
|
||||
|
||||
MAX_LIMIT = 50
|
||||
DEFAULT_LIMIT = 25
|
||||
|
||||
# How many candidates carry their trigger in full before the rest are cut.
|
||||
DEFAULT_DETAIL = 5
|
||||
|
||||
# The cut length for the tail. Long enough to carry the first clause of a
|
||||
# trigger — which is where these state the act they are about — short enough
|
||||
# that forty-five of them stay affordable.
|
||||
_TEASER_CHARS = 140
|
||||
|
||||
|
||||
def _teaser(text: str) -> tuple[str, bool]:
|
||||
"""Shorten at a word break with a visible cut. Returns (text, was_cut).
|
||||
|
||||
The technique is `plugin_context._goal_line`'s, and the fallback is the
|
||||
part worth copying: `textwrap.shorten` returns a bare "…" when the string
|
||||
is one unbroken word longer than the cap, and a raw slice ends mid-word
|
||||
claiming to be the whole thing (#4036). Kept local rather than shared
|
||||
because the two callers wrap it in different sentences; if a third appears,
|
||||
that is the moment to extract it rather than now.
|
||||
"""
|
||||
flat = " ".join((text or "").split())
|
||||
if len(flat) <= _TEASER_CHARS:
|
||||
return flat, False
|
||||
short = textwrap.shorten(flat, width=_TEASER_CHARS, placeholder="…")
|
||||
if short == "…":
|
||||
short = flat[: _TEASER_CHARS - 1] + "…"
|
||||
return short, True
|
||||
|
||||
|
||||
async def what_might_apply(
|
||||
query: str,
|
||||
limit: int = DEFAULT_LIMIT,
|
||||
detail: int = DEFAULT_DETAIL,
|
||||
kind: str = "",
|
||||
project_id: int = 0,
|
||||
) -> dict:
|
||||
"""Every rule that might bear on what you are about to do, ranked, with no bar.
|
||||
|
||||
REACH FOR THIS BEFORE A CONSEQUENTIAL OR IRREVERSIBLE ACT — a push, a
|
||||
merge, a delete, a deploy, anything outward-facing — and before handing
|
||||
work back because you are unsure whether you are allowed to finish it.
|
||||
|
||||
It is the tool the "ask before a consequential act" reflex should call.
|
||||
A rule reaches a session by retrieval, and the arms that push rules at you
|
||||
have a small budget: they deliver the few highest-scoring candidates and
|
||||
say nothing about what sat just underneath. That is right for an arm that
|
||||
fires before every command and wrong for the one moment you actually want
|
||||
to be sure. This is that moment's tool.
|
||||
|
||||
WHAT IT RETURNS, AND WHY THE TAIL LOOKS LIKE NOISE
|
||||
|
||||
There is no threshold. You get the `limit` nearest candidates whatever they
|
||||
score, ordered, each with its score — so the tail IS expected to be
|
||||
irrelevant, and that is the design. You are not reading the list for its
|
||||
average quality; you are reading it for the one record you would otherwise
|
||||
have missed. Scan the triggers, open what looks live with `get_rule(id)`,
|
||||
and ignore the rest.
|
||||
|
||||
The first few carry their trigger in full; the rest carry a cut of it,
|
||||
marked `truncated`. A cut trigger is still enough to decide whether to
|
||||
open the record, which is the whole job of a teaser.
|
||||
|
||||
`kind` is on every row and is never something to infer: a **rule** must be
|
||||
followed, a **preference** records how the operator wants work done.
|
||||
Missing a rule is a mistake; missing a preference costs consistency.
|
||||
|
||||
A NOTE ON WHAT THIS CANNOT DO. It is a pull, so it only helps if you ask.
|
||||
The failure it was built for — a session withholding a routine action
|
||||
because the rule permitting it never arrived — produces no tool call of its
|
||||
own, so nothing will prompt you. Asking is the habit; this is where to put
|
||||
it.
|
||||
|
||||
Args:
|
||||
query: what you are about to do, in the words you would use to
|
||||
describe it — "push to dev after committing", "delete the staging
|
||||
database", "merge dev to main". A command string works; a sentence
|
||||
usually works better, because triggers are written as prose.
|
||||
limit: how many candidates, default 25, capped at 50.
|
||||
detail: how many carry their trigger in FULL before the rest are cut,
|
||||
default 5.
|
||||
kind: "rule" or "preference" to restrict; omit for both.
|
||||
project_id: scope to one project — its own rules plus every global
|
||||
one. Omit to ask the whole rulebook.
|
||||
"""
|
||||
uid = current_user_id()
|
||||
limit = max(1, min(int(limit), MAX_LIMIT))
|
||||
detail = max(0, min(int(detail), limit))
|
||||
|
||||
report: dict = {}
|
||||
started = time.perf_counter()
|
||||
if project_id:
|
||||
raw = await semantic_search_rules(
|
||||
uid, query, limit=limit, threshold=THRESHOLD, kind=kind or None,
|
||||
report=report, project_id=project_id,
|
||||
)
|
||||
else:
|
||||
raw = await semantic_search_rules(
|
||||
uid, query, limit=limit, threshold=THRESHOLD, kind=kind or None,
|
||||
report=report, everywhere=True,
|
||||
)
|
||||
duration_ms = (time.perf_counter() - started) * 1000
|
||||
|
||||
candidates = []
|
||||
for rank, (score, rule) in enumerate(raw):
|
||||
full = rank < detail
|
||||
trigger, was_cut = (
|
||||
(" ".join((rule.when_to_apply or "").split()), False)
|
||||
if full else _teaser(rule.when_to_apply or "")
|
||||
)
|
||||
candidates.append({
|
||||
"id": rule.id,
|
||||
"title": rule.title,
|
||||
# Force, not topic. See the docstring — this is never inferred.
|
||||
"kind": rule.kind,
|
||||
# A rule in a rulebook topic is global; one on a project binds
|
||||
# there alone (milestone 414). Which it is changes how far a
|
||||
# reader should generalise from it.
|
||||
"scope": "project" if rule.project_id else "global",
|
||||
"when_to_apply": trigger,
|
||||
"truncated": was_cut,
|
||||
"score": round(float(score), 4),
|
||||
})
|
||||
|
||||
# Logged as a pull, with `searched` honoured: a search that never ran must
|
||||
# not be recorded as a ranker declining (#3765).
|
||||
record_retrieval(
|
||||
user_id=uid,
|
||||
source=SOURCE,
|
||||
query=query,
|
||||
threshold=THRESHOLD,
|
||||
limit=limit,
|
||||
project_id=project_id or None,
|
||||
is_task=None,
|
||||
results=raw,
|
||||
duration_ms=duration_ms,
|
||||
best_available=report.get("best_available_score"),
|
||||
best_available_id=report.get("best_available_id"),
|
||||
searched=report.get("searched", True),
|
||||
)
|
||||
|
||||
return {
|
||||
"candidates": candidates,
|
||||
"returned": len(candidates),
|
||||
"detailed": min(detail, len(candidates)),
|
||||
# Present so a caller can tell "the corpus offered nothing" from "the
|
||||
# search never ran" — an empty query, an unavailable embedder and a
|
||||
# failed query all return zero rows and mean different things (#3670).
|
||||
"searched": bool(report.get("searched", True)),
|
||||
"open_with": "get_rule(id)",
|
||||
}
|
||||
|
||||
|
||||
def register(mcp) -> None:
|
||||
mcp.tool(name="what_might_apply")(what_might_apply)
|
||||
@@ -186,14 +186,10 @@ def test_the_tool_teaches_reading_the_records_not_the_percentile():
|
||||
def test_all_three_tools_are_registered():
|
||||
from scribe.mcp.tools import retrieval_tuning as tool
|
||||
|
||||
names = []
|
||||
from tests.helpers import FakeMCP
|
||||
|
||||
class _MCP:
|
||||
def tool(self, name):
|
||||
names.append(name)
|
||||
return lambda fn: fn
|
||||
|
||||
tool.register(_MCP())
|
||||
assert names == [
|
||||
mcp = FakeMCP()
|
||||
tool.register(mcp)
|
||||
assert mcp.names == [
|
||||
"retrieval_surfaces", "tune_retrieval", "retrieval_tuning_history",
|
||||
]
|
||||
|
||||
@@ -0,0 +1,243 @@
|
||||
"""The wide net: many candidates, no bar, and a cost that stays bounded (#4103).
|
||||
|
||||
WHY THIS EXISTS
|
||||
|
||||
Milestone 416 step 5 moves "do not exclude potentially important data" off the
|
||||
push, where fifty candidates would be a wall nobody reads, and onto a pull,
|
||||
where they arrive only when a session asks for them.
|
||||
|
||||
Three things have to hold, and each has a way of quietly failing:
|
||||
|
||||
1. **No bar.** The moment this tool serves is the one where the caller does
|
||||
not trust a threshold to decide for them. A default floor creeping in
|
||||
here would turn the wide net into the narrow one and nothing would look
|
||||
wrong — the results would simply be fewer.
|
||||
2. **A bounded cost.** The step was filed claiming fifty "costs nothing".
|
||||
`_rule_hint_line` had already measured ~143 tokens for a line with its
|
||||
trigger rendered, so fifty is ~7,000. The graduated shape (#3851) is what
|
||||
keeps this affordable, and a change that renders every trigger in full
|
||||
would pass every other test here.
|
||||
3. **A pull, logged as one.** The push arms' near-miss distributions are the
|
||||
evidence #4121 argues from. A pull mixed into them moves those numbers.
|
||||
"""
|
||||
from unittest.mock import AsyncMock, patch
|
||||
|
||||
import pytest
|
||||
|
||||
from scribe.mcp.tools import wide_net
|
||||
from tests.helpers import FakeMCP, fake_rule, tool_doc
|
||||
|
||||
|
||||
def _hits(n=3, trigger="Running git push", **over):
|
||||
"""n (score, rule) pairs, descending, the shape the service returns."""
|
||||
return [
|
||||
(0.9 - i * 0.01,
|
||||
fake_rule(id=i + 1, title=f"rule {i + 1}", when_to_apply=trigger, **over))
|
||||
for i in range(n)
|
||||
]
|
||||
|
||||
|
||||
def _patched(hits, report=None):
|
||||
"""Patch the search and the telemetry sink; hand back the search mock."""
|
||||
async def _search(*a, **kw):
|
||||
if report is not None and "report" in kw and kw["report"] is not None:
|
||||
kw["report"].update(report)
|
||||
return hits
|
||||
return patch.object(wide_net, "semantic_search_rules", AsyncMock(side_effect=_search))
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_no_bar_reaches_the_search():
|
||||
"""THE POINT OF THE TOOL. A floor here would narrow the net silently."""
|
||||
with _patched(_hits()) as search, \
|
||||
patch.object(wide_net, "record_retrieval"), \
|
||||
patch.object(wide_net, "current_user_id", lambda: 1):
|
||||
await wide_net.what_might_apply("push to dev")
|
||||
|
||||
assert search.await_args.kwargs["threshold"] == 0.0
|
||||
assert wide_net.THRESHOLD == 0.0
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
@pytest.mark.parametrize("asked, expected", [(999, 50), (0, 1), (-5, 1), (25, 25)])
|
||||
async def test_the_limit_is_clamped_to_the_cap(asked, expected):
|
||||
with _patched(_hits()) as search, \
|
||||
patch.object(wide_net, "record_retrieval"), \
|
||||
patch.object(wide_net, "current_user_id", lambda: 1):
|
||||
await wide_net.what_might_apply("q", limit=asked)
|
||||
|
||||
assert search.await_args.kwargs["limit"] == expected
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_head_carries_its_trigger_whole_and_the_tail_is_cut():
|
||||
"""The graduated shape (#3851), which is what makes fifty affordable.
|
||||
|
||||
A regression that rendered every trigger in full would satisfy every other
|
||||
assertion in this file, so the cut is pinned on both sides: the head is
|
||||
NOT marked truncated and the tail IS.
|
||||
"""
|
||||
long_trigger = " ".join(["running a git command before pushing anything"] * 12)
|
||||
with _patched(_hits(6, trigger=long_trigger)), \
|
||||
patch.object(wide_net, "record_retrieval"), \
|
||||
patch.object(wide_net, "current_user_id", lambda: 1):
|
||||
out = await wide_net.what_might_apply("q", detail=2)
|
||||
|
||||
head, tail = out["candidates"][:2], out["candidates"][2:]
|
||||
assert all(c["truncated"] is False for c in head)
|
||||
assert all(c["when_to_apply"] == " ".join(long_trigger.split()) for c in head)
|
||||
assert all(c["truncated"] is True for c in tail)
|
||||
assert all(len(c["when_to_apply"]) <= wide_net._TEASER_CHARS + 1 for c in tail)
|
||||
assert out["detailed"] == 2
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_short_trigger_is_never_marked_truncated():
|
||||
"""`truncated` is a claim about this row, not about its rank. A tail row
|
||||
whose trigger already fits must not claim a cut that did not happen."""
|
||||
with _patched(_hits(4, trigger="Running git push")), \
|
||||
patch.object(wide_net, "record_retrieval"), \
|
||||
patch.object(wide_net, "current_user_id", lambda: 1):
|
||||
out = await wide_net.what_might_apply("q", detail=1)
|
||||
|
||||
assert all(c["truncated"] is False for c in out["candidates"])
|
||||
|
||||
|
||||
def test_the_cut_breaks_on_a_word_and_says_it_was_cut():
|
||||
"""#4036's lesson, borrowed: a raw slice ends mid-word and reads as the
|
||||
whole thing."""
|
||||
text, cut = wide_net._teaser("alpha beta gamma delta " * 40)
|
||||
assert cut is True
|
||||
assert text.endswith("…")
|
||||
# Broke on a word, so no partial token sits before the marker.
|
||||
assert not text.removesuffix("…").rstrip().endswith(("alph", "bet", "gam"))
|
||||
|
||||
|
||||
def test_one_unbroken_word_still_yields_text_rather_than_a_bare_marker():
|
||||
"""`textwrap.shorten` returns just "…" here, which would render a teaser
|
||||
carrying no information at all."""
|
||||
text, cut = wide_net._teaser("x" * 500)
|
||||
assert cut is True and text != "…"
|
||||
assert len(text) == wide_net._TEASER_CHARS
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_every_row_says_its_force_and_its_scope():
|
||||
"""`kind` is never inferred — a rule must be followed, a preference guides
|
||||
— and `scope` says how far a reader should generalise from it."""
|
||||
hits = [
|
||||
(0.8, fake_rule(id=1, kind="rule", project_id=None)),
|
||||
(0.7, fake_rule(id=2, kind="preference", project_id=44)),
|
||||
]
|
||||
with _patched(hits), patch.object(wide_net, "record_retrieval"), \
|
||||
patch.object(wide_net, "current_user_id", lambda: 1):
|
||||
out = await wide_net.what_might_apply("q")
|
||||
|
||||
assert [c["kind"] for c in out["candidates"]] == ["rule", "preference"]
|
||||
assert [c["scope"] for c in out["candidates"]] == ["global", "project"]
|
||||
|
||||
|
||||
# ── telemetry: a pull, and never mistaken for a push ────────────────────────
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_the_call_is_logged_under_its_own_pull_source():
|
||||
with _patched(_hits()), \
|
||||
patch.object(wide_net, "record_retrieval") as rec, \
|
||||
patch.object(wide_net, "current_user_id", lambda: 1):
|
||||
await wide_net.what_might_apply("push to dev")
|
||||
|
||||
kw = rec.call_args.kwargs
|
||||
assert kw["source"] == wide_net.SOURCE == "wide_net"
|
||||
assert kw["threshold"] == 0.0
|
||||
|
||||
|
||||
def test_the_wide_net_is_not_one_of_the_tunable_push_surfaces():
|
||||
"""THE GUARD that keeps #4121's evidence clean (rule 167).
|
||||
|
||||
The registry holds the PUSH arms — the ones with a floor and a budget the
|
||||
model tunes. This source must not appear there: a pull folded into those
|
||||
rows would move the near-miss distributions that step is arguing from, and
|
||||
would offer a floor to tune on a tool whose whole point is not having one.
|
||||
"""
|
||||
from scribe.services.retrieval_surfaces import SURFACES
|
||||
|
||||
assert wide_net.SOURCE not in SURFACES
|
||||
|
||||
|
||||
def test_the_wide_net_is_not_ambient_either():
|
||||
"""Ambient means a delivery nobody chose. This one is chosen by definition
|
||||
— somebody called the tool — so counting it as ambient would make a
|
||||
deliberate ask read as a bulk hand-over."""
|
||||
from scribe.services.note_usage import AMBIENT_SOURCES
|
||||
|
||||
assert wide_net.SOURCE not in AMBIENT_SOURCES
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_search_that_never_ran_is_not_reported_as_a_decline():
|
||||
"""#3765: an empty query, a dead embedder and a failed query all return
|
||||
nothing, and none of them is a ranker declining."""
|
||||
with _patched([], report={"searched": False, "best_available_score": None}), \
|
||||
patch.object(wide_net, "record_retrieval") as rec, \
|
||||
patch.object(wide_net, "current_user_id", lambda: 1):
|
||||
out = await wide_net.what_might_apply("")
|
||||
|
||||
assert rec.call_args.kwargs["searched"] is False
|
||||
assert out["searched"] is False
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_what_the_bar_turned_away_is_carried_through():
|
||||
"""There is no bar here, but `best_available` still answers "did the corpus
|
||||
have anything at all" for a call that came back empty (#3670)."""
|
||||
with _patched([], report={"searched": True, "best_available_score": 0.31,
|
||||
"best_available_id": 7}), \
|
||||
patch.object(wide_net, "record_retrieval") as rec, \
|
||||
patch.object(wide_net, "current_user_id", lambda: 1):
|
||||
await wide_net.what_might_apply("q")
|
||||
|
||||
assert rec.call_args.kwargs["best_available"] == 0.31
|
||||
assert rec.call_args.kwargs["best_available_id"] == 7
|
||||
|
||||
|
||||
# ── the contract a session actually reads ───────────────────────────────────
|
||||
|
||||
def test_the_docstring_says_when_to_reach_for_it():
|
||||
"""The step's done-when, and the load-bearing half of this tool.
|
||||
|
||||
A wide net nobody knows to call is worth nothing, so the docstring has to
|
||||
name the MOMENT, not just the parameters — including the one that prompted
|
||||
it, where a session hands work back rather than finishing it.
|
||||
"""
|
||||
doc = tool_doc("scribe.mcp.tools.wide_net", "what_might_apply").lower()
|
||||
|
||||
assert "consequential" in doc
|
||||
assert "handing work back" in doc
|
||||
# Says the tail is expected to be noise — otherwise the first caller reads
|
||||
# a low-scoring list as the tool being broken.
|
||||
assert "noise" in doc
|
||||
# And names its own limit: a pull only helps if somebody asks.
|
||||
assert "only helps if you ask" in doc
|
||||
|
||||
|
||||
def test_it_is_registered_and_readable_with_a_read_key():
|
||||
from scribe.mcp.server import _READ_ONLY_TOOLS
|
||||
|
||||
mcp = FakeMCP()
|
||||
wide_net.register(mcp)
|
||||
assert mcp.names == ["what_might_apply"]
|
||||
assert "what_might_apply" in _READ_ONLY_TOOLS
|
||||
|
||||
|
||||
def test_the_instruction_surfaces_point_at_it():
|
||||
"""Rule 119: the instruction surfaces ARE the specification for product
|
||||
behaviour, so a tool the reflex never learns about is not shipped."""
|
||||
import pathlib
|
||||
|
||||
from scribe.mcp import server
|
||||
|
||||
assert "what_might_apply" in server._INSTRUCTIONS
|
||||
|
||||
skill = (pathlib.Path(__file__).resolve().parents[1]
|
||||
/ "plugin" / "skills" / "using-scribe" / "SKILL.md").read_text()
|
||||
assert "what_might_apply" in skill
|
||||
Reference in New Issue
Block a user