feat(systems): read-side teeth — the vocabulary at session start, a search filter, and the state/chronicle instructions
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / TypeScript typecheck (push) Successful in 10s
CI & Build / integration (push) Successful in 18s
CI & Build / Python tests (push) Successful in 46s
CI & Build / Build & push image (push) Successful in 25s

Step 4 of #278, product half. The audit that motivated it: one System in
project 2, thirty records tagged, nothing since July 28 — three days after the
feature landed. Not a discipline failure; retrieval was completely blind to
the association (zero references in embeddings, knowledge, search, auto-inject,
or enter_project), so tagging was a write-side label with no read-side payoff,
and labels nobody reads don't get maintained.

Three changes, ordered by what makes the others workable:

1. enter_project returns the project's Systems (id, name, first line of the
   charter). Load-bearing for the tagging instruction: you cannot ask an agent
   to check a record against a vocabulary it never sees. Trimmed because it
   rides on every session start; the full charter stays get_system's job.
   Present-and-empty rather than absent when a project has none — "no named
   areas yet" is information the create-the-System instruction acts on.

2. search accepts system_id, MCP and REST (#33). Implemented once in
   semantic_search_notes as an EXISTS against record_systems — an association
   filter deciding candidate-set membership before scoring, like project_id,
   not a ranking signal. The REST route's missing project filter stays #2463's:
   it carries a default-scope UI decision this change must not preempt.

3. The instructions (#119, _INSTRUCTIONS + using-scribe skill; plugin 0.1.25
   for the cache):
   - Tag as you write, with an executable test — "would someone investigating
     that subsystem want this in the pile list_system_records returns?" —
     rather than "tag appropriately", which is what died.
   - Create the System when the area has no record: the two-or-more test
     snippets use, plus "don't wait to be asked to name an area that plainly
     exists", because the agent's default was leaving un-modelled areas
     un-modelled forever.
   - State vs chronicle: dev-logs are written once and never rewritten; durable
     findings live in the System's reference note, updated in place — safe
     because note versions are the changelog, which has existed since the
     feature shipped and was never named as one.

list_system_records' docstring now sells it as the way to READ a subsystem,
reference note first. No auto-inject boost by System — vocabulary and filter
first, measure before adding ranking behaviour (the #2486 lesson).

Refs #278, #2546
This commit is contained in:
2026-08-08 18:19:58 -04:00
parent 6c4c1bccfc
commit 3f1523b19f
9 changed files with 154 additions and 7 deletions
+54
View File
@@ -17,6 +17,18 @@ def _bind_user():
_user_id_ctx.reset(token)
@pytest.fixture(autouse=True)
def _no_systems():
"""enter_project now surfaces the project's Systems as the tagging
vocabulary (#2546). These are tool-layer unit tests with no database, so
the lookup is stubbed to the common case — a project with none. The
populated shape is asserted in its own test below.
"""
with patch("scribe.mcp.tools.projects.systems_svc.list_systems",
AsyncMock(return_value=[])):
yield
def _fake_project(design_system_id=None, **overrides) -> MagicMock:
p = MagicMock()
base = {"id": 1, "title": "P", "description": "", "goal": "",
@@ -184,6 +196,48 @@ async def test_enter_project_composes_full_context():
# absent. A caller that has to distinguish "no key" from "no system" will
# eventually get it wrong.
assert out["design_system"] is None
# No Systems -> present-and-empty, NOT absent: this key is the tagging
# vocabulary, and "this project has no named areas yet" is information the
# create-the-System instruction acts on.
assert out["systems"] == []
@pytest.mark.asyncio
async def test_enter_project_surfaces_the_systems_vocabulary():
"""The tagging instruction is only executable if the vocabulary is in
front of the agent when it writes. It never was, and tagging stopped three
days after the feature landed — one System, nothing tagged since July 28
(#2546's audit). Trimmed to id/name/first-line: it rides on every session
start, and the full charter is get_system's job."""
p = _fake_project(id=5)
sys1 = MagicMock()
sys1.id = 3
sys1.name = "retrieval"
sys1.description = "Embeddings, ranking, auto-inject.\nLong detail below."
with patch(
"scribe.mcp.tools.projects.projects_svc.get_project",
AsyncMock(return_value=p),
), patch(
"scribe.mcp.tools.projects.rulebooks_svc.get_applicable_rules",
AsyncMock(return_value={"rules": [], "truncated": False,
"subscribed_rulebooks": []}),
), patch(
"scribe.mcp.tools.projects.milestones_svc.get_project_milestone_summary",
AsyncMock(return_value=[]),
), patch(
"scribe.mcp.tools.projects.notes_svc.list_notes",
AsyncMock(side_effect=[([], 0), ([], 0)]),
), patch(
"scribe.mcp.tools.projects.systems_svc.list_systems",
AsyncMock(return_value=[sys1]),
):
out = await enter_project(project_id=5)
assert out["systems"] == [
{"id": 3, "name": "retrieval",
"description": "Embeddings, ranking, auto-inject."}
]
@pytest.mark.asyncio