Files
FabledScribe/tests/test_services_plugin_context.py
T
bvandeusenandClaude Opus 5 238510080e
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 7s
CI & Build / TypeScript typecheck (push) Successful in 11s
CI & Build / integration (push) Successful in 31s
CI & Build / Python tests (push) Successful in 1m4s
CI & Build / Build & push image (push) Successful in 35s
feat(retrieval): the standing-rule arm gets its own bar, and asks for one rule not two (#3318)
Milestone 333 step 4 — the split #2223 made one surface down, now made for the
third corpus. The arm inherited WRITEPATH_DEFAULT_THRESHOLD = 0.68, a number
measured against code-vs-note-PROSE and never re-derived for code-vs-RULE-TEXT.

THE DEFAULT IS ARGUED STRUCTURALLY, NOT READ OFF A HISTOGRAM (rule 115). Two
facts hold on any install, including one with six rules and no telemetry:

- The eligible corpus is tiny — conditional rules only, a handful to a few
  dozen against thousands of notes. A top-k over forty candidates always
  returns something, so "the best match cleared the bar" stops meaning "a good
  match exists". A bar calibrated for best-of-thousands is cleared by
  best-of-forty as arithmetic, not relevance.
- Rules are short imperative technical English, far more homogeneous than note
  prose. #2223 put the code-vs-prose floor at 0.55-0.63 and set 0.68 above it;
  a more homogeneous corpus has a HIGHER floor, so 0.68 is not merely
  inherited, it sits below where this corpus's noise lives.

0.72 errs deliberately toward silence on an asymmetry that is also structural:
this hint fires on EVERY write. A missed rule is recoverable — it is still in
Scribe and the agent can search it. A hint that cries wolf is not: it teaches
the reader to skip the whole block, and the true positives go with it. The
arm's own comment already said "noise on a hint that fires on every write is
how a hint gets ignored".

Pinned as an INEQUALITY, not a value: test_the_rule_bar_defaults_above_the_code_bar
asserts RULEHINT > WRITEPATH, so tuning the number stays free while inverting
the relationship — which would silently reinstate #3311 — does not.

RULEHINT_LIMIT = 1, and deliberately not a knob. With a corpus this small, k=2
means the second line is almost always the second-best noise wearing the same
confident framing as the first; halving k halves that regardless of the bar.
It stays a constant because it is a decision about how loud one hint may be,
not a per-install tuning question — and a knob nobody turns only adds a way to
misconfigure the surface.

Reachable from Settings, no restart (rule 25), with copy that says which way to
move it and points at retrieval_telemetry's rule pull-through — which step 3
made readable — to tell "arriving unread" from "never arrived".

Every config stand-in in the suite gained the key, not just the one that
noticed. The arm reads `rule_threshold` while BUILDING its search arguments, so
a missing key raises inside its fail-open except and turns the arm into a
silent no-op — indistinguishable from it running and finding nothing. That is
the same vacuous-pass shape that bit step 2, one layer down (rule 33).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TcCs1CcQ1ormdnzSshKqvN
2026-09-02 18:05:00 -04:00

473 lines
22 KiB
Python

from unittest.mock import AsyncMock, MagicMock, patch
import pytest
from tests.helpers import fake_note
@pytest.fixture(autouse=True)
def _no_exclusions():
"""build_session_context asks for the bound project's always-on
exclusions (milestone 297); these tests script the rules only."""
with patch("scribe.services.plugin_context.rulebooks_svc.excluded_always_on_rulebooks",
AsyncMock(return_value=[])):
yield
pytestmark = pytest.mark.usefixtures("_no_supersession")
def _rule(rid, title, topic_id):
r = MagicMock()
r.id, r.title, r.topic_id = rid, title, topic_id
r.statement = "FULL STATEMENT SHOULD NOT BE INJECTED"
return r
# ─── knowledge auto-inject (Path A) ──────────────────────────────────────────
@pytest.mark.asyncio
async def test_get_autoinject_config_defaults_and_clamps():
from scribe.services import plugin_context as pc
# No settings stored → defaults.
with patch.object(pc, "get_setting", AsyncMock(side_effect=lambda uid, k, d: d)):
cfg = await pc.get_autoinject_config(1)
assert cfg == {
"enabled": pc.AUTOINJECT_DEFAULT_ENABLED,
"threshold": pc.AUTOINJECT_DEFAULT_THRESHOLD,
"top_k": pc.AUTOINJECT_DEFAULT_TOP_K,
}
# Out-of-range values are clamped; top_k capped at the hard ceiling.
stored = {
pc.AUTOINJECT_ENABLED_KEY: "false",
pc.AUTOINJECT_THRESHOLD_KEY: "5",
pc.AUTOINJECT_TOP_K_KEY: "999",
}
with patch.object(pc, "get_setting",
AsyncMock(side_effect=lambda uid, k, d: stored.get(k, d))):
cfg = await pc.get_autoinject_config(1)
assert cfg["enabled"] is False
assert cfg["threshold"] == 1.0
assert cfg["top_k"] == pc._AUTOINJECT_MAX_TOP_K
@pytest.mark.asyncio
async def test_build_autoinject_hint_disabled_returns_empty_and_skips_search():
from scribe.services import plugin_context as pc
search = AsyncMock()
with patch.object(pc, "get_autoinject_config",
AsyncMock(return_value={"enabled": False, "threshold": 0.55, "top_k": 3})), \
patch.object(pc, "semantic_search_notes", search), \
patch.object(pc, "record_retrieval", MagicMock()):
out = await pc.build_autoinject_hint(1, "anything")
assert out["context"] == "" and out["note_ids"] == []
search.assert_not_called() # disabled → no retrieval at all
@pytest.mark.asyncio
async def test_build_autoinject_hint_titles_only_with_margin_gate():
from scribe.services import plugin_context as pc
# top=0.80; 0.74 within band (0.10), 0.61 outside → dropped.
hits = [(0.80, fake_note(id=11, title="Pool sizing decision", user_id=1)),
(0.74, fake_note(id=22, title="run_maintenance thresholds", user_id=1)),
(0.61, fake_note(id=33, title="unrelated-ish", user_id=1))]
rec = MagicMock()
with patch.object(pc, "get_autoinject_config",
AsyncMock(return_value={"enabled": True, "threshold": 0.55, "top_k": 3})), \
patch.object(pc, "semantic_search_notes", AsyncMock(return_value=hits)), \
patch.object(pc, "record_retrieval", rec):
out = await pc.build_autoinject_hint(1, "postgres pool", project_id=2,
exclude_ids=[99])
# Margin gate kept the top two, dropped the straggler.
assert out["note_ids"] == [11, 22]
assert '#11 [note] "Pool sizing decision" (0.80)' in out["context"]
assert "#33" not in out["context"]
# Title-first: no body text, ever.
assert "get_note(id)" in out["context"]
# Telemetry fired for BOTH retrievals this path runs: the scored menu and
# the reuse-slot query competing against it. The slot's query used to be
# the one unlogged retrieval on this path — the hit it displaced was in
# retrieval_logs, the query that displaced it was not (#2463).
sources = [c.kwargs["source"] for c in rec.call_args_list]
assert sources == ["auto_inject", "reuse_slot"]
@pytest.mark.asyncio
async def test_build_autoinject_hint_blank_query_returns_empty():
from scribe.services import plugin_context as pc
search = AsyncMock()
with patch.object(pc, "get_autoinject_config",
AsyncMock(return_value={"enabled": True, "threshold": 0.55, "top_k": 3})), \
patch.object(pc, "semantic_search_notes", search):
out = await pc.build_autoinject_hint(1, " ")
assert out["context"] == ""
search.assert_not_called()
@pytest.mark.asyncio
async def test_build_session_context_renders_titles_grouped_by_topic():
rules = [
_rule(1, "`dev` is home", 1),
_rule(2, "Release — never without explicit request", 1),
_rule(3, "No GitHub — Fabled-Git only", 2),
]
with patch("scribe.services.plugin_context.rulebooks_svc.list_always_on_rules",
AsyncMock(return_value=rules)), \
patch("scribe.services.plugin_context._topic_titles",
AsyncMock(return_value={1: "git-workflow", 2: "fabled-git"})):
from scribe.services.plugin_context import build_session_context
out = await build_session_context(user_id=7, project_id=0)
ctx = out["context"]
assert out["rule_count"] == 3
assert out["project"] is None
# Titles present, grouped under topic headings
assert "### git-workflow" in ctx
assert "### fabled-git" in ctx
assert "- [1] `dev` is home" in ctx
assert "- [3] No GitHub — Fabled-Git only" in ctx
# Full statements must NOT be dumped (push channel injects titles only)
assert "FULL STATEMENT SHOULD NOT BE INJECTED" not in ctx
@pytest.mark.asyncio
async def test_build_session_context_includes_project_when_scoped():
# design_system_id explicitly None: a bare MagicMock would hand back a truthy
# auto-attribute and send this through the design branch, which is the
# opposite of what this test is about.
project = MagicMock(id=2, title="FabledScribe", goal="ship it",
design_system_id=None)
with patch("scribe.services.plugin_context.rulebooks_svc.list_always_on_rules",
AsyncMock(return_value=[_rule(1, "rule", 1)])), \
patch("scribe.services.plugin_context._topic_titles",
AsyncMock(return_value={1: "git-workflow"})), \
patch("scribe.services.plugin_context.projects_svc.get_project",
AsyncMock(return_value=project)), \
patch("scribe.services.plugin_context.notes_svc.list_notes",
AsyncMock(return_value=([], 4))):
from scribe.services.plugin_context import build_session_context
out = await build_session_context(user_id=7, project_id=2)
assert out["project"] == {"id": 2, "title": "FabledScribe"}
assert "## Active project: FabledScribe (id 2)" in out["context"]
assert "Open todo tasks: 4" in out["context"]
# No design system on the project -> no design block at all. An install with
# none is the ordinary case, not a degraded one.
assert "## Design system" not in out["context"]
@pytest.mark.asyncio
async def test_build_session_context_pushes_the_projects_design_system():
"""The gap this closes: a design system had no push channel, so its
standards reached a session only if the agent already knew to go looking —
the same silent failure as a token nobody declares."""
project = MagicMock(id=2, title="App", goal="", design_system_id=9)
design = {
"id": 9, "title": "App kit", "description": "",
"inherits_from": ["House"],
"guidance": [], "token_count": 95,
"token_groups": ["accent", "surface", "type"],
}
with patch("scribe.services.plugin_context.rulebooks_svc.list_always_on_rules",
AsyncMock(return_value=[_rule(1, "rule", 1)])), \
patch("scribe.services.plugin_context._topic_titles",
AsyncMock(return_value={1: "git-workflow"})), \
patch("scribe.services.plugin_context.projects_svc.get_project",
AsyncMock(return_value=project)), \
patch("scribe.services.plugin_context.notes_svc.list_notes",
AsyncMock(return_value=([], 0))), \
patch("scribe.services.plugin_context.design_systems_svc.design_context",
AsyncMock(return_value=design)):
from scribe.services.plugin_context import build_session_context
out = await build_session_context(user_id=7, project_id=2)
ctx = out["context"]
assert "## Design system: App kit (id 9) (inherits House)" in ctx
assert "95 tokens across accent, surface, type" in ctx
# A pointer to the values, never the values themselves — a hundred token
# declarations would crowd out the context they are meant to inform. The
# summary carries counts and group NAMES only, which is why design_context
# returns those rather than the resolved tokens.
assert "resolve_design_system(9)" in ctx
assert "get_design_system_stylesheet(9)" in ctx
@pytest.mark.asyncio
async def test_build_session_context_survives_an_unreadable_design_system():
"""design_context returns None when the caller may not read the system.
That must degrade to "no design block", not to a crash that costs the
session its rules too."""
project = MagicMock(id=2, title="App", goal="", design_system_id=9)
with patch("scribe.services.plugin_context.rulebooks_svc.list_always_on_rules",
AsyncMock(return_value=[_rule(1, "rule", 1)])), \
patch("scribe.services.plugin_context._topic_titles",
AsyncMock(return_value={1: "git-workflow"})), \
patch("scribe.services.plugin_context.projects_svc.get_project",
AsyncMock(return_value=project)), \
patch("scribe.services.plugin_context.notes_svc.list_notes",
AsyncMock(return_value=([], 0))), \
patch("scribe.services.plugin_context.design_systems_svc.design_context",
AsyncMock(return_value=None)):
from scribe.services.plugin_context import build_session_context
out = await build_session_context(user_id=7, project_id=2)
assert "## Design system" not in out["context"]
assert "## Active project: App (id 2)" in out["context"]
@pytest.mark.asyncio
async def test_build_session_context_unbound_repo_emits_bind_hint():
with patch("scribe.services.plugin_context.rulebooks_svc.list_always_on_rules",
AsyncMock(return_value=[_rule(1, "rule", 1)])), \
patch("scribe.services.plugin_context._topic_titles",
AsyncMock(return_value={1: "git-workflow"})):
from scribe.services.plugin_context import build_session_context
out = await build_session_context(
user_id=7, project_id=0, unbound_repo="host/owner/repo",
)
ctx = out["context"]
assert out["project"] is None
assert "## Repository not yet bound" in ctx
assert 'bind_repo(repo_url="host/owner/repo"' in ctx
# No project block when unbound.
assert "## Active project" not in ctx
@pytest.mark.asyncio
async def test_build_process_manifest_renders_stub_specs():
items = [
{"id": 5, "title": "Drift Audit", "tags": [], "snippet": "Find drifted docs."},
{"id": 9, "title": "DRY Pass", "tags": [], "snippet": ""},
]
with patch("scribe.services.plugin_context.knowledge_svc.query_knowledge",
AsyncMock(return_value=(items, 2))):
from scribe.services.plugin_context import build_process_manifest
out = await build_process_manifest(user_id=7)
assert out["total"] == 2
drift = out["processes"][0]
assert drift["id"] == 5
assert drift["name"] == "Drift Audit"
assert drift["slug"] == "drift-audit" # kebab-cased
assert "Drift Audit" in drift["description"] # auto-surface trigger
assert "Find drifted docs." in drift["description"] # preview folded in
# No-snippet process still gets a usable description.
assert "DRY Pass" in out["processes"][1]["description"]
@pytest.mark.asyncio
async def test_build_process_manifest_dedupes_slugs_and_skips_blank_titles():
items = [
{"id": 1, "title": "My Process", "tags": [], "snippet": "a"},
{"id": 2, "title": "my process", "tags": [], "snippet": "b"}, # same slug
{"id": 3, "title": " ", "tags": [], "snippet": "skip me"}, # blank title
]
with patch("scribe.services.plugin_context.knowledge_svc.query_knowledge",
AsyncMock(return_value=(items, 3))):
from scribe.services.plugin_context import build_process_manifest
out = await build_process_manifest(user_id=7)
slugs = [p["slug"] for p in out["processes"]]
assert slugs == ["my-process", "my-process-2"] # collision suffixed with id
assert out["total"] == 2 # blank-title entry dropped
@pytest.mark.asyncio
async def test_build_process_manifest_truncates_long_preview():
items = [{"id": 1, "title": "Big", "tags": [], "snippet": "x" * 500}]
with patch("scribe.services.plugin_context.knowledge_svc.query_knowledge",
AsyncMock(return_value=(items, 1))):
from scribe.services.plugin_context import build_process_manifest
out = await build_process_manifest(user_id=7)
assert "…" in out["processes"][0]["description"]
assert "x" * 500 not in out["processes"][0]["description"]
@pytest.mark.asyncio
async def test_build_session_context_caps_length():
many = [_rule(i, "x" * 200, 1) for i in range(200)]
with patch("scribe.services.plugin_context.rulebooks_svc.list_always_on_rules",
AsyncMock(return_value=many)), \
patch("scribe.services.plugin_context._topic_titles",
AsyncMock(return_value={1: "git-workflow"})):
from scribe.services.plugin_context import build_session_context
out = await build_session_context(user_id=7)
assert len(out["context"]) <= 9000 + 60 # cap + truncation note
assert "truncated" in out["context"]
# --- the reuse slot (#2246) --------------------------------------------------
_CFG = {"enabled": True, "threshold": 0.55, "top_k": 3}
async def _autoinject(main_hits, reuse_hits, cfg=None):
"""Run build_autoinject_hint with the two semantic queries stubbed in order:
the unscoped pool first, then the reserved reuse query."""
calls: list[dict] = []
async def fake_search(*_a, **kw):
calls.append(kw)
return reuse_hits if kw.get("note_type") else main_hits
with patch("scribe.services.plugin_context.get_autoinject_config",
AsyncMock(return_value=dict(cfg or _CFG))), \
patch("scribe.services.plugin_context.semantic_search_notes",
AsyncMock(side_effect=fake_search)), \
patch("scribe.services.plugin_context.record_retrieval", MagicMock()), \
patch("scribe.services.plugin_context.record_surfaced", MagicMock()), \
patch("scribe.services.plugin_context.owner_names_for",
AsyncMock(return_value={})):
from scribe.services.plugin_context import build_autoinject_hint
out = await build_autoinject_hint(1, "write a debounce helper")
return out, calls
@pytest.mark.asyncio
async def test_a_snippet_takes_the_last_slot_when_none_won_on_score():
"""The measured failure: a prompt asking for a helper returned three project
records ABOUT building the retrieval system, and zero snippets. Scribe's own
records are about software work, so they share vocabulary with any coding
prompt while answering none of them."""
main = [(0.66, fake_note(id=1, title="Step 3: title-first auto-inject", user_id=1)),
(0.65, fake_note(id=2, title="Task-reminder dedup query crashes", user_id=1, is_task=True)),
(0.64, fake_note(id=3, title="Drafter hardening · write-path trigger", user_id=1, is_task=True))]
reuse = [(0.58, fake_note(id=9, title="debounce — collapse rapid calls", user_id=1, note_type="snippet"))]
out, calls = await _autoinject(main, reuse)
assert out["note_ids"] == [1, 2, 9] # last slot displaced, not the first
assert "#9" in out["context"] and "[snippet]" in out["context"]
# The reserved query is scoped to the reuse kinds and asks for exactly one.
assert calls[1]["note_type"] == ("snippet", "process")
assert calls[1]["limit"] == 1
@pytest.mark.asyncio
async def test_the_reserved_query_is_skipped_when_a_snippet_already_won():
"""No second query, and no slot spent twice, when ranking already did the
right thing — the fix must be invisible in the case it isn't needed."""
main = [(0.81, fake_note(id=9, title="debounce helper", user_id=1, note_type="snippet")),
(0.80, fake_note(id=1, title="some task", user_id=1, is_task=True))]
out, calls = await _autoinject(main, [])
assert len(calls) == 1 # reserved query never ran
assert out["note_ids"] == [9, 1]
@pytest.mark.asyncio
async def test_a_weak_snippet_does_not_buy_the_slot():
"""The reserved hit skips the MARGIN band — that band is what snippets lose
to — but never the threshold. Silence stays the default; a slot spent on an
irrelevant snippet is how a menu teaches people to ignore it."""
main = [(0.66, fake_note(id=1, title="a task", user_id=1, is_task=True))]
out, calls = await _autoinject(main, []) # threshold returned nothing
assert out["note_ids"] == [1]
assert len(calls) == 2 # it asked, and got nothing
@pytest.mark.asyncio
async def test_the_reserved_hit_is_not_held_to_the_margin_band():
"""0.58 is 0.08 below the top hit. Under the band it would survive; the point
is that it must survive even when it wouldn't — a snippet losing to a
same-vocabulary project record by a wide margin is the whole bug."""
main = [(0.90, fake_note(id=1, title="a task", user_id=1, is_task=True))]
reuse = [(0.58, fake_note(id=9, title="debounce", user_id=1, note_type="snippet"))]
out, _calls = await _autoinject(main, reuse)
assert 9 in out["note_ids"]
@pytest.mark.asyncio
async def test_a_process_counts_as_reuse_too():
"""A stored process answers 'how do we do X here' the same way a snippet
answers 'what do we already have' — both lose to the same project records."""
main = [(0.70, fake_note(id=1, title="a task", user_id=1, is_task=True))]
reuse = [(0.60, fake_note(id=8, title="DRY pass process", user_id=1, note_type="process"))]
out, _ = await _autoinject(main, reuse)
assert 8 in out["note_ids"]
# …and one already on the menu suppresses the reserved query.
out2, calls2 = await _autoinject(
[(0.70, fake_note(id=8, title="DRY pass process", user_id=1, note_type="process"))], [])
assert len(calls2) == 1
# --- write-path widened beyond snippets (#2246, the mirror half) -------------
@pytest.mark.asyncio
async def test_write_path_semantic_arm_asks_for_experience_not_just_snippets():
"""The inverse of auto-inject's mistake. This arm was snippets-only, so an
issue saying "we tried this and it deadlocked" could never reach the moment
that code was about to be written — arguably the better prior art, because
it says what NOT to do."""
from scribe.services import plugin_context as pc
hits = [(0.72, fake_note(id=9, title="debounce helper", user_id=1, note_type="snippet")),
(0.70, fake_note(id=7, title="Debounce dropped the trailing call", user_id=1, is_task=True, task_kind="issue"))]
search = AsyncMock(return_value=hits)
rec = MagicMock()
with patch.object(pc, "get_writepath_config",
AsyncMock(return_value={"enabled": True, "threshold": 0.6,
"top_k": 3,
"rule_threshold": 0.72})), \
patch.object(pc.snippets_svc, "list_snippets",
AsyncMock(return_value=([], 0))), \
patch.object(pc, "semantic_search_notes", search), \
patch.object(pc, "record_retrieval", rec), \
patch.object(pc, "record_surfaced", MagicMock()), \
patch.object(pc, "owner_names_for", AsyncMock(return_value={})), \
patch.object(pc, "concept_query", MagicMock(return_value="debounce a callback")):
out = await pc.build_write_path_hint(
1, "src/utils/debounce.ts", code="x" * 400,
)
kw = search.await_args.kwargs
assert kw["note_type"] == ("snippet", "note")
# An open to-do resembling the code answers nothing; an ISSUE carries a root
# cause and a NOTE carries durable knowledge. Only the todo is excluded.
assert kw["task_kind"] == "issue"
# Telemetry must not claim this was a notes-only retrieval any more.
assert rec.call_args.kwargs["is_task"] is None
@pytest.mark.asyncio
async def test_write_path_labels_a_non_snippet_hit_with_its_kind():
"""An unlabelled issue on this menu reads as "here is code to reuse", which
is the opposite of what it says. A snippet stays unlabelled — it is the
menu's default and the header's default reading."""
from scribe.services import plugin_context as pc
hits = [(0.72, fake_note(id=9, title="debounce helper", user_id=1, note_type="snippet")),
(0.71, fake_note(id=7, title="Debounce dropped the trailing call", user_id=1, is_task=True, task_kind="issue"))]
with patch.object(pc, "get_writepath_config",
AsyncMock(return_value={"enabled": True, "threshold": 0.6,
"top_k": 3,
"rule_threshold": 0.72})), \
patch.object(pc.snippets_svc, "list_snippets",
AsyncMock(return_value=([], 0))), \
patch.object(pc, "semantic_search_notes", AsyncMock(return_value=hits)), \
patch.object(pc, "record_retrieval", MagicMock()), \
patch.object(pc, "record_surfaced", MagicMock()), \
patch.object(pc, "owner_names_for", AsyncMock(return_value={})), \
patch.object(pc, "concept_query", MagicMock(return_value="debounce a callback")):
out = await pc.build_write_path_hint(
1, "src/utils/debounce.ts", code="x" * 400,
)
ctx = out["context"]
assert "· issue]" in ctx # the issue says what it is
assert '[similar 0.72] "debounce helper"' in ctx # the snippet does not
# The header now names the right opener for each kind.
assert "get_task(id)" in ctx and "get_snippet(id)" in ctx