CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 58s
CI & Build / Python tests (push) Successful in 1m42s
CI & Build / Build & push image (push) Successful in 28s
The engine took `note_type` and `task_kind` all along. What was missing was a way to say them: the MCP tool's `content_type` knew `note`, `task` and `all`, and `/api/search` knew the same two — so an agent could not ask "has this snippet already been recorded" or "what lessons apply here" without searching everything and reading past the rest. Browse offered nine kinds from the same data. The cause is that each door kept its own map. `_FACETS` in services/knowledge is where a kind is declared, and #3161 made adding one a single edit by generating the SQL filter, the Python predicate and the door's validation from it — but the two search doors were written before that and never joined. So this adds the third dialect, `search_filters_for`, and one composition over it, `content_type_filters`, and both doors now derive instead of listing. Two names keep a meaning of their own, and the docstrings say so: `all` is no filter, and `note` is BROAD — any non-task, snippets and lessons included — where the browse facet of the same name is narrow (`note_type == 'note'`). They are left different deliberately; narrowing this one would stop returning snippets to every caller that already asks this way. An unrecognised kind is now refused rather than answered. Both doors used to fall through: the MCP tool into a filter matching no row, the route into no filter at all, so `?content_type=snippets` returned the whole corpus while looking like a narrowed search. An empty result set is a claim — "the corpus holds nothing like this" — and an agent acts on that claim by building the thing it could not find, so a typo must not be able to make it. The docstring is the agent-facing contract (#2846), and a test now holds it to the table: every kind `_FACETS` declares has to appear in it, because a filter an agent has not been told about is unreachable however well it is wired. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
86 lines
3.4 KiB
Python
86 lines
3.4 KiB
Python
import time
|
|
|
|
from quart import Blueprint, jsonify, request
|
|
|
|
from scribe.auth import login_required, get_current_user_id
|
|
from scribe.services.access import owner_names_for
|
|
from scribe.services.embeddings import (
|
|
INTERACTIVE_SEARCH_THRESHOLD as _REST_SEARCH_THRESHOLD,
|
|
)
|
|
from scribe.services.embeddings import semantic_search_notes
|
|
from scribe.services.knowledge import content_type_filters
|
|
from scribe.services.retrieval_telemetry import record_retrieval
|
|
|
|
# The interactive floor lives in embeddings.py now, shared with Browse search
|
|
# and the list views' semantic `q` — one number, one rationale (#2463).
|
|
|
|
search_bp = Blueprint("search", __name__, url_prefix="/api/search")
|
|
|
|
|
|
@search_bp.route("", methods=["GET"])
|
|
@login_required
|
|
async def search_route():
|
|
uid = get_current_user_id()
|
|
q = (request.args.get("q") or "").strip()
|
|
if not q:
|
|
return jsonify({"error": "q is required"}), 400
|
|
|
|
limit = min(request.args.get("limit", 10, type=int), 50)
|
|
# Every kind the facet table declares, derived rather than mapped here —
|
|
# this route used to know exactly two and read anything else as "no
|
|
# filter", so `?content_type=snippets` silently returned the whole corpus
|
|
# (#4250). An unknown kind is now a 400 naming the ones that exist: a
|
|
# result set is an answer, and it should not be able to answer a question
|
|
# nobody asked.
|
|
try:
|
|
filters = content_type_filters(request.args.get("content_type", "all"))
|
|
except ValueError as exc:
|
|
return jsonify({"error": str(exc)}), 400
|
|
is_task = filters.get("is_task")
|
|
# Same association filters the MCP tool takes (#33). Optional, default
|
|
# global: this route has NO frontend consumer today (measured 2026-08-08 —
|
|
# the web UI searches through /api/knowledge), so it serves API callers,
|
|
# and an API caller states its scope explicitly.
|
|
system_id = request.args.get("system_id", type=int)
|
|
project_id = request.args.get("project_id", type=int)
|
|
|
|
t0 = time.perf_counter()
|
|
report: dict = {}
|
|
results = await semantic_search_notes(
|
|
uid, q, limit=limit, **filters, threshold=_REST_SEARCH_THRESHOLD,
|
|
project_id=project_id, system_id=system_id,
|
|
# The user typed this, so it reaches everything they may read.
|
|
scope="read",
|
|
report=report,
|
|
)
|
|
record_retrieval(
|
|
user_id=uid, source="rest_search", query=q,
|
|
threshold=_REST_SEARCH_THRESHOLD, limit=limit,
|
|
project_id=project_id, is_task=is_task, results=results,
|
|
duration_ms=(time.perf_counter() - t0) * 1000.0,
|
|
best_available=report.get("best_available_score"),
|
|
best_available_id=report.get("best_available_id"),
|
|
searched=bool(report.get("searched", True)),
|
|
)
|
|
owners = await owner_names_for(
|
|
{int(note.user_id) for _s, note in results if note.user_id != uid}
|
|
)
|
|
return jsonify({
|
|
"results": [
|
|
{
|
|
"id": note.id,
|
|
"title": note.title,
|
|
"body": note.body or "",
|
|
"is_task": note.is_task,
|
|
"tags": note.tags or [],
|
|
"similarity": score,
|
|
**(
|
|
{"shared": True, "owner": owners.get(int(note.user_id))}
|
|
if note.user_id != uid else {}
|
|
),
|
|
}
|
|
for score, note in results # semantic_search_notes returns list[tuple[float, Note]]
|
|
],
|
|
"total": len(results),
|
|
})
|