CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 58s
CI & Build / Python tests (push) Successful in 1m38s
CI & Build / Build & push image (push) Successful in 27s
Step 2 of #4251. A System's `description` is a charter — several hundred words saying what belongs in that area and what does not — and it is the answer to "which part of this codebase does X live in". There was no semantic path to one: `list_systems` enumerates, and `search(system_id=…)` uses a System as a FILTER over notes. So a System could narrow a search and could never be the answer to one, and an agent asking where a record belonged had to read every charter or guess. ITS OWN SEARCH, not a `content_type` over notes, for the reason note 3163 gives about milestones: the row could be shared, the search cannot. A charter competing with the whole note corpus for one top-k is outranked by the records filed under it — the right answer crowded out by its own contents — and "where does this belong?" is a different question from "what prior art is there?", which a caller asking one should not have to read past answers to. So `system_embeddings` (0107) joins note_, rule_ and milestone_embeddings as the fourth sibling, with `system_document`, `upsert_system_embedding`, `semantic_search_systems`, a startup backfill and `search(content_type= "system")`. Scoped like milestones: with a project_id, that project's Systems if the caller can read the project (rule 78); without one, the caller's own. Archived Systems are excluded — an archived area is one the operator has said is no longer where things go, which is exactly the question being asked. `system_document` is the plainest of the four shapes on purpose. A charter is already written as the thing this search has to match, in the words someone asking would use — so there is no trigger to synthesise as `rule_document` must, and no second record to gather as `task_document` must. The stored charter IS the sharp document, the way a snippet's is. `color`, `status` and `order_index` stay out: presentation and bookkeeping, and a vector carrying them would be answering a question nobody asks of a charter. The search publishes `report["best_chunk"]` from the start rather than being retrofitted, which is what #4251 asked of any fourth search. It matters more here than anywhere: a charter runs long and a result shows its NAME, so a match on the paragraph that actually decides where a record belongs would otherwise be previewed by two words that cannot say. The id that comes back is the one `system_id`, `system_ids` and `list_system_records` already take, so the answer to "where does this belong?" is directly usable as "show me what is there" and as "file it here". `embed_system` sits beside `notes.embed_note` at the service for #2056's reason — every door gets it by construction. Not called on delete: that is a soft delete and the search joins through `System`, so the vectors are already unreachable, and leaving them means a restore is findable again immediately. `system_embeddings` is declared in backup's `_NOT_INCLUDED` as derived, beside its three siblings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
410 lines
17 KiB
Python
410 lines
17 KiB
Python
"""search tool — proves the tool pattern (context + service call + dict shape).
|
|
|
|
Service call is mocked; no DB needed."""
|
|
from unittest.mock import AsyncMock, MagicMock, patch
|
|
|
|
import pytest
|
|
|
|
from scribe.mcp._context import _user_id_ctx
|
|
from scribe.mcp.tools.search import search
|
|
from tests.helpers import fake_note
|
|
|
|
|
|
@pytest.fixture(autouse=True)
|
|
def _reset_user_ctx():
|
|
"""Each test starts with no MCP context. Tests set it explicitly."""
|
|
token = _user_id_ctx.set(None)
|
|
yield
|
|
_user_id_ctx.reset(token)
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_fable_search_raises_without_context():
|
|
with pytest.raises(RuntimeError, match="no MCP user context"):
|
|
await search(q="anything")
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_fable_search_returns_repackaged_results():
|
|
_user_id_ctx.set(7)
|
|
fake = fake_note(id=1, title="kafka rebalance", is_task=False, body="HPA details", tags=["ops"])
|
|
with patch(
|
|
"scribe.mcp.tools.search.semantic_search_notes",
|
|
AsyncMock(return_value=[(0.93, fake)]),
|
|
):
|
|
out = await search(q="kafka")
|
|
|
|
assert out["total"] == 1
|
|
assert len(out["results"]) == 1
|
|
r = out["results"][0]
|
|
assert r["id"] == 1
|
|
assert r["title"] == "kafka rebalance"
|
|
# The whole body, because it fits — and named as the opening, not as the
|
|
# passage that matched, since this call patched the search and so carries
|
|
# no chunk.
|
|
assert r["excerpt"] == "HPA details"
|
|
assert r["excerpt_is"] == "body_opening"
|
|
assert r["body_length"] == len("HPA details")
|
|
assert r["is_task"] is False
|
|
assert r["tags"] == ["ops"]
|
|
assert r["similarity"] == pytest.approx(0.93)
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_search_shows_the_passage_that_matched_not_the_opening():
|
|
"""The point of #4243. A record can rank on its sixth paragraph; showing
|
|
its first is showing the caller a span the search already judged less
|
|
relevant, and letting them decide from it."""
|
|
_user_id_ctx.set(7)
|
|
body = "Chapter one, about nothing. " * 40 + " THE ANSWER IS 04775c3."
|
|
fake = fake_note(id=1, title="t", body=body)
|
|
|
|
async def _search(*a, **kw):
|
|
kw["report"]["best_chunk"] = {
|
|
1: {"index": 3, "text": "THE ANSWER IS 04775c3."}
|
|
}
|
|
return [(0.5, fake)]
|
|
|
|
with patch("scribe.mcp.tools.search.semantic_search_notes", _search):
|
|
out = await search(q="answer")
|
|
r = out["results"][0]
|
|
assert r["excerpt"] == "THE ANSWER IS 04775c3."
|
|
assert r["excerpt_is"] == "matched_passage"
|
|
assert r["chunk_index"] == 3
|
|
# And the caller is told there is more record behind the passage.
|
|
assert r["body_length"] == len(body)
|
|
assert "read_full" in r
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_a_result_says_whether_its_excerpt_is_the_match_or_the_opening():
|
|
"""A caller that cannot tell the two apart cannot judge whether looking
|
|
deeper is worth it, which is the only decision this field supports."""
|
|
_user_id_ctx.set(7)
|
|
fake = fake_note(id=1, title="t", body="short body")
|
|
with patch(
|
|
"scribe.mcp.tools.search.semantic_search_notes",
|
|
AsyncMock(return_value=[(0.5, fake)]),
|
|
):
|
|
out = await search(q="x")
|
|
assert out["results"][0]["excerpt_is"] == "body_opening"
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_a_long_excerpt_keeps_both_ends_and_says_how_much_went():
|
|
"""The fallback is still an elision, and an elision that drops the tail
|
|
drops wherever the conclusion was."""
|
|
_user_id_ctx.set(7)
|
|
body = "OPENING. " + ("m" * 3000) + " CLOSING."
|
|
fake = fake_note(id=1, title="t", body=body)
|
|
with patch(
|
|
"scribe.mcp.tools.search.semantic_search_notes",
|
|
AsyncMock(return_value=[(0.5, fake)]),
|
|
):
|
|
out = await search(q="x")
|
|
excerpt = out["results"][0]["excerpt"]
|
|
assert excerpt.startswith("OPENING.")
|
|
assert excerpt.rstrip().endswith("CLOSING.")
|
|
assert "characters omitted" in excerpt
|
|
assert out["results"][0]["body_length"] == len(body)
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_a_short_record_arrives_whole_and_unmarked():
|
|
"""Fragmenting a 200-character note serves nobody."""
|
|
_user_id_ctx.set(7)
|
|
fake = fake_note(id=1, title="t", body="all of it")
|
|
with patch(
|
|
"scribe.mcp.tools.search.semantic_search_notes",
|
|
AsyncMock(return_value=[(0.5, fake)]),
|
|
):
|
|
out = await search(q="x")
|
|
assert out["results"][0]["excerpt"] == "all of it"
|
|
assert "read_full" not in out["results"][0]
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_fable_search_content_type_filters_at_service_layer():
|
|
"""The three historical values keep meaning exactly what they meant.
|
|
|
|
`note` is the one that could plausibly have been tightened when the
|
|
specific kinds arrived — it is BROAD here (any non-task, snippets and
|
|
lessons included) while the web facet of the same name is narrow. Pinning
|
|
it stops a later tidy-up from silently removing snippets from every caller
|
|
that already asks this way (#4250)."""
|
|
_user_id_ctx.set(7)
|
|
mock_search = AsyncMock(return_value=[])
|
|
with patch("scribe.mcp.tools.search.semantic_search_notes", mock_search):
|
|
await search(q="x", content_type="task")
|
|
assert mock_search.call_args.kwargs["is_task"] is True
|
|
assert mock_search.call_args.kwargs.get("task_kind") is None
|
|
|
|
mock_search.reset_mock()
|
|
await search(q="x", content_type="note")
|
|
assert mock_search.call_args.kwargs["is_task"] is False
|
|
assert mock_search.call_args.kwargs.get("note_type") is None
|
|
|
|
mock_search.reset_mock()
|
|
await search(q="x", content_type="all")
|
|
assert mock_search.call_args.kwargs["is_task"] is None
|
|
|
|
|
|
# --- the specific kinds: the engine already supported them (#4250) -----------
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
@pytest.mark.parametrize(
|
|
"content_type,expected",
|
|
[
|
|
("snippet", {"is_task": False, "note_type": "snippet"}),
|
|
("lesson", {"is_task": False, "note_type": "lesson"}),
|
|
("process", {"is_task": False, "note_type": "process"}),
|
|
("issue", {"is_task": True, "task_kind": "issue"}),
|
|
("spike", {"is_task": True, "task_kind": "spike"}),
|
|
("work", {"is_task": True, "task_kind": "work"}),
|
|
("plan", {"is_task": True, "task_kind": "plan"}),
|
|
],
|
|
)
|
|
async def test_each_specific_kind_reaches_the_engine_filter_it_names(
|
|
content_type, expected
|
|
):
|
|
"""`note_type` and `task_kind` were parameters of the search all along —
|
|
what was missing was a way to ask for them from the agent's door. Asking
|
|
for a snippet must narrow to snippets, not merely to non-tasks."""
|
|
_user_id_ctx.set(7)
|
|
mock_search = AsyncMock(return_value=[])
|
|
with patch("scribe.mcp.tools.search.semantic_search_notes", mock_search):
|
|
await search(q="x", content_type=content_type)
|
|
kwargs = mock_search.call_args.kwargs
|
|
for key, value in expected.items():
|
|
assert kwargs[key] == value, f"{content_type}: {key}"
|
|
|
|
|
|
def test_the_vocabulary_is_derived_from_the_facet_table_not_recopied():
|
|
"""#3161's property, held at this door too: adding a kind to `_FACETS` is
|
|
one edit. A hand-kept list here is exactly how the agent's search came to
|
|
offer two kinds while the web's offered nine."""
|
|
from scribe.services.knowledge import FACET_TYPES, content_type_filters
|
|
|
|
for facet in FACET_TYPES:
|
|
filters = content_type_filters(facet) # raises if a kind is unreachable
|
|
assert "is_task" in filters, facet
|
|
|
|
# The guard can fail: a name absent from the table is refused, so this is
|
|
# membership in the table and not "every string works" (#167).
|
|
with pytest.raises(ValueError):
|
|
content_type_filters("a-kind-that-is-not-in-the-facet-table")
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_an_unknown_content_type_is_refused_rather_than_returning_nothing():
|
|
"""An empty result set is a CLAIM — "the corpus holds nothing like this" —
|
|
and an agent acts on it by writing the thing it could not find. A typo must
|
|
not be able to make that claim, so the door raises with the vocabulary
|
|
instead of falling through to a filter that matches no row."""
|
|
_user_id_ctx.set(7)
|
|
mock_search = AsyncMock(return_value=[])
|
|
with patch("scribe.mcp.tools.search.semantic_search_notes", mock_search):
|
|
with pytest.raises(ValueError) as err:
|
|
await search(q="x", content_type="snippets") # plural typo
|
|
mock_search.assert_not_awaited()
|
|
message = str(err.value)
|
|
assert "snippets" in message
|
|
assert "snippet" in message and "lesson" in message # names the valid ones
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_the_docstring_names_every_kind_the_tool_accepts():
|
|
"""The docstring IS the agent-facing contract (#2846) — a filter an agent
|
|
has not been told about is unreachable however well it is wired."""
|
|
from scribe.mcp.tools.search import search as search_tool
|
|
from scribe.services.knowledge import FACET_TYPES
|
|
|
|
doc = search_tool.__doc__ or ""
|
|
for facet in FACET_TYPES:
|
|
assert f"'{facet}'" in doc, f"{facet} is accepted but never documented"
|
|
for own in ("'rule'", "'milestone'", "'system'", "'all'"):
|
|
assert own in doc, f"{own} dispatches somewhere and is never documented"
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_an_explicit_search_reaches_a_lesson_from_any_project():
|
|
"""The wiring half of milestone 385 step 3.
|
|
|
|
A lesson records an insight that transfers, so a project filter that hid it
|
|
would hide it precisely on the project that has not learned it yet. This is
|
|
the EXPLICIT search — the operator asked — so it opts in; the unasked-for
|
|
injection arms decide their own budget separately (step 5).
|
|
|
|
Asserted on the kwarg rather than on results, because what can regress here
|
|
is the wiring: the service grew the capability and a call site that never
|
|
passes it leaves the whole kind unreachable, with every unit test still
|
|
green.
|
|
"""
|
|
_user_id_ctx.set(7)
|
|
mock_search = AsyncMock(return_value=[])
|
|
with patch("scribe.mcp.tools.search.semantic_search_notes", mock_search):
|
|
await search(q="x", project_id=3)
|
|
|
|
assert mock_search.call_args.kwargs["include_global_kinds"] is True
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_fable_search_limit_is_clamped():
|
|
_user_id_ctx.set(7)
|
|
mock_search = AsyncMock(return_value=[])
|
|
with patch("scribe.mcp.tools.search.semantic_search_notes", mock_search):
|
|
await search(q="x", limit=999)
|
|
assert mock_search.call_args.kwargs["limit"] == 50
|
|
|
|
mock_search.reset_mock()
|
|
await search(q="x", limit=0)
|
|
assert mock_search.call_args.kwargs["limit"] == 1
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
@pytest.mark.parametrize("project_id, scope", [
|
|
(5, {"project_id": 5}),
|
|
# No project: the explicit question is asked of the whole rulebook.
|
|
(0, {"everywhere": True}),
|
|
])
|
|
async def test_rule_search_scopes_to_the_project_it_is_given(project_id, scope):
|
|
"""With a project: global rules plus that project's (milestone 414).
|
|
Without one: every rule, because an unscoped "is there a rule about this"
|
|
is asking the whole rulebook — unlike a hook, which speaks unasked."""
|
|
_user_id_ctx.set(7)
|
|
found = AsyncMock(return_value=[])
|
|
with patch("scribe.mcp.tools.search.semantic_search_rules", found):
|
|
await search(q="release tagging", content_type="rule", project_id=project_id)
|
|
kwargs = found.await_args.kwargs
|
|
assert {k: kwargs[k] for k in scope} == scope
|
|
assert set(kwargs) & {"project_id", "everywhere"} == set(scope)
|
|
|
|
|
|
def test_a_milestone_is_embedded_by_what_it_is_for_then_its_plan():
|
|
from scribe.services.embeddings import milestone_document
|
|
|
|
assert milestone_document("M3", "metadata providers", "## Goal\nx") == (
|
|
"M3 — metadata providers", "metadata providers\n\n## Goal\nx")
|
|
# A roadmap milestone written with no description is still findable by its plan.
|
|
assert milestone_document("M3", None, "the plan") == ("M3", "the plan")
|
|
assert milestone_document(None, None, None) == (None, None)
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_milestone_search_is_its_own_shape_and_scopes_to_the_project():
|
|
"""milestone 415: 'is there already a plan for this?' has a tool."""
|
|
from unittest.mock import MagicMock
|
|
|
|
_user_id_ctx.set(7)
|
|
# `body` is a REAL string, not left to MagicMock's attribute autovivication:
|
|
# the result now reads it, and a mock body would make `matched` a mock
|
|
# object and `body_length` zero while the assertion still looked green
|
|
# (lesson #2833).
|
|
ms = MagicMock(id=339, title="M3 — Metadata", description="works, editions",
|
|
status="active", project_id=30,
|
|
body="Step 4 — the metadata editions table.")
|
|
|
|
summary = AsyncMock(return_value=[{"id": 339, "total": 0, "completed": 0}])
|
|
seen: dict = {}
|
|
|
|
async def _milestone_search(*_a, **kw):
|
|
seen.update(kw)
|
|
kw["report"]["best_chunk"] = {
|
|
339: {"index": 1, "text": "Step 4 — the metadata editions table."}
|
|
}
|
|
return [(0.81, ms)]
|
|
|
|
with patch("scribe.mcp.tools.search.semantic_search_milestones",
|
|
_milestone_search), \
|
|
patch("scribe.services.milestones.get_project_milestone_summary", summary):
|
|
out = await search(q="book metadata", content_type="milestone", project_id=30)
|
|
assert seen["project_id"] == 30
|
|
assert out["results"] == [{
|
|
"id": 339, "title": "M3 — Metadata", "description": "works, editions",
|
|
# The plan body stays out; the passage that matched comes along, and
|
|
# says which of the two it is (#4243).
|
|
"matched": "Step 4 — the metadata editions table.",
|
|
"matched_is": "matched_passage",
|
|
"body_length": len("Step 4 — the metadata editions table."),
|
|
"status": "active", "project_id": 30, "total": 0, "completed": 0,
|
|
"similarity": 0.81,
|
|
}]
|
|
|
|
|
|
# --- systems: the charter as an answer, not a filter (#4251) -----------------
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_system_search_is_its_own_shape_and_carries_the_matched_passage():
|
|
"""A System's charter runs to several hundred words and a result shows its
|
|
NAME — so a match on the paragraph that actually decides where a record
|
|
belongs would be previewed by two words that cannot. The passage comes
|
|
along, marked as the passage (#4243)."""
|
|
_user_id_ctx.set(7)
|
|
charter = (
|
|
"Opening sentence about the area in general. "
|
|
+ "x" * 400
|
|
+ " Retrieval telemetry and the floors it judges belong here."
|
|
)
|
|
fake = MagicMock(id=3, project_id=2, description=charter)
|
|
# `name` is MagicMock's own constructor kwarg — passed in, it names the
|
|
# mock and leaves `.name` a mock object, which is note #2833's whole point.
|
|
fake.name = "Retrieval & recall"
|
|
|
|
async def _system_search(uid, q, **kwargs):
|
|
kwargs["report"]["best_chunk"] = {
|
|
3: {"index": 2, "text": "Retrieval telemetry and the floors it judges belong here."}
|
|
}
|
|
return [(0.81, fake)]
|
|
|
|
with patch("scribe.mcp.tools.search.semantic_search_systems", _system_search):
|
|
out = await search(q="where does retrieval telemetry go?",
|
|
content_type="system", project_id=2)
|
|
|
|
assert out["total"] == 1
|
|
row = out["results"][0]
|
|
assert row["id"] == 3
|
|
assert row["name"] == "Retrieval & recall"
|
|
assert row["project_id"] == 2
|
|
assert row["similarity"] == 0.81
|
|
# The passage that won, not the charter's opening — and SAID to be that.
|
|
assert row["matched"] == (
|
|
"Retrieval telemetry and the floors it judges belong here."
|
|
)
|
|
assert row["matched_is"] == "matched_passage"
|
|
assert row["body_length"] == len(charter)
|
|
assert "read_full" in row, "a 56-character span of a 500-character charter"
|
|
# The charter itself does not ride along: get_system reads it.
|
|
assert "description" not in row
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_system_search_scopes_to_the_project_it_was_given():
|
|
"""Systems are per-project and a charter from another project is not an
|
|
answer to "where does this belong here?"."""
|
|
_user_id_ctx.set(7)
|
|
mock = AsyncMock(return_value=[])
|
|
with patch("scribe.mcp.tools.search.semantic_search_systems", mock):
|
|
await search(q="x", content_type="system", project_id=30)
|
|
assert mock.call_args.kwargs["project_id"] == 30
|
|
|
|
mock.reset_mock()
|
|
with patch("scribe.mcp.tools.search.semantic_search_systems", mock):
|
|
await search(q="x", content_type="system")
|
|
assert mock.call_args.kwargs["project_id"] is None
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_a_system_search_does_not_go_through_the_note_search():
|
|
"""Its own search, not a note_type. A charter competing with the whole note
|
|
corpus for one top-k is outranked by the records filed under it."""
|
|
_user_id_ctx.set(7)
|
|
notes = AsyncMock(return_value=[])
|
|
with (
|
|
patch("scribe.mcp.tools.search.semantic_search_notes", notes),
|
|
patch("scribe.mcp.tools.search.semantic_search_systems", AsyncMock(return_value=[])),
|
|
):
|
|
await search(q="x", content_type="system")
|
|
notes.assert_not_awaited()
|