feat(search): the agent's search can ask for every kind the corpus has (#4250)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 58s
CI & Build / Python tests (push) Successful in 1m42s
CI & Build / Build & push image (push) Successful in 28s

The engine took `note_type` and `task_kind` all along. What was missing was a
way to say them: the MCP tool's `content_type` knew `note`, `task` and `all`,
and `/api/search` knew the same two — so an agent could not ask "has this
snippet already been recorded" or "what lessons apply here" without searching
everything and reading past the rest. Browse offered nine kinds from the same
data.

The cause is that each door kept its own map. `_FACETS` in services/knowledge
is where a kind is declared, and #3161 made adding one a single edit by
generating the SQL filter, the Python predicate and the door's validation from
it — but the two search doors were written before that and never joined. So
this adds the third dialect, `search_filters_for`, and one composition over it,
`content_type_filters`, and both doors now derive instead of listing.

Two names keep a meaning of their own, and the docstrings say so: `all` is no
filter, and `note` is BROAD — any non-task, snippets and lessons included —
where the browse facet of the same name is narrow (`note_type == 'note'`).
They are left different deliberately; narrowing this one would stop returning
snippets to every caller that already asks this way.

An unrecognised kind is now refused rather than answered. Both doors used to
fall through: the MCP tool into a filter matching no row, the route into no
filter at all, so `?content_type=snippets` returned the whole corpus while
looking like a narrowed search. An empty result set is a claim — "the corpus
holds nothing like this" — and an agent acts on that claim by building the
thing it could not find, so a typo must not be able to make it.

The docstring is the agent-facing contract (#2846), and a test now holds it to
the table: every kind `_FACETS` declares has to appear in it, because a filter
an agent has not been told about is unreachable however well it is wired.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
2026-09-21 11:12:48 -04:00
co-authored by Claude Opus 5
parent 0ec499d9b4
commit 5fb41af9b0
5 changed files with 242 additions and 27 deletions
+63
View File
@@ -376,6 +376,69 @@ def matches_facet(note, note_type: str | None) -> bool:
return not note.is_task and note.note_type == value
def search_filters_for(facet: str) -> dict:
"""The semantic-search kwargs one facet implies: is_task + note_type/task_kind.
A third dialect of the same table, for the arm that has no SQL statement to
narrow and no fetched row to test — it is passing filters INTO
`semantic_search_notes`. `_apply_type_filter` is the SQL dialect and
`matches_facet` the Python one; all three read `_FACETS`, which is what
keeps "adding a kind" a single edit (#3161).
Returns kwargs rather than a tuple so a caller splats it and cannot pair
`task_kind` with `is_task=False` by writing the positions out of order.
"""
is_task, value = _facet(facet)
if is_task:
return {"is_task": True, "task_kind": value}
return {"is_task": False, "note_type": value}
# The two names a `content_type` parameter carries that are not facets. Both
# search doors that take one — the MCP tool and /api/search — have always
# spelled them this way, so they are the contract rather than a convenience:
#
# 'all' (and empty) — no kind filter at all.
# 'note' — ANY non-task record, so snippets, lessons and processes
# are all still in scope. The BROWSE facet of the same
# name is narrower (`note_type == 'note'` exactly). The
# two are deliberately different: narrowing this one
# would stop returning snippets to every caller that
# already asks this way, and the doors document the
# containment instead.
_CONTENT_TYPE_BROAD = {
"": {"is_task": None},
"all": {"is_task": None},
"note": {"is_task": False},
}
def content_type_filters(content_type: str, extra: tuple = ()) -> dict:
"""The search kwargs a door's `content_type` parameter implies.
Lives here, beside `_FACETS`, because a door that keeps its own map is how
this went wrong: the agent's search offered two kinds and browse offered
nine, and a third copy in `/api/search` offered two more quietly still
(#4250). Derived, so adding a kind stays one edit (#3161).
Raises on anything unrecognised. `_facet` alone would fall back to reading
an unknown string as a note_type, which matches no row — so a typo comes
back as a confident empty result, and an empty result is a CLAIM that the
corpus holds nothing of the sort. A door has to be able to say "that is not
a kind" instead. `extra` names kinds the caller dispatches elsewhere (the
MCP tool's 'rule' and 'milestone'), so the refusal lists what that door
really accepts and not a vocabulary from some other door.
"""
if content_type in _CONTENT_TYPE_BROAD:
return dict(_CONTENT_TYPE_BROAD[content_type])
if content_type in FACET_TYPES:
return search_filters_for(content_type)
raise ValueError(
f"unknown content_type {content_type!r}. Valid: "
+ ", ".join(sorted({"all", *FACET_TYPES, *extra}))
)
def _apply_type_filter(stmt, note_type: str | None):
"""Apply the type facet to a Note select. Trashed rows are always excluded."""
stmt = stmt.where(Note.deleted_at.is_(None))