2b85443dd1
CI & Build / Python lint (push) Successful in 2s
CI & Build / integration (push) Successful in 31s
CI & Build / TypeScript typecheck (push) Successful in 34s
CI & Build / Python tests (push) Successful in 56s
CI & Build / Build & push image (push) Successful in 44s
retrieval_logs answers "what did the ranker return, at what scores" — the right substrate for tuning a threshold. It cannot answer the question the snippet corpus actually needs: did anyone open this? A snippet nobody opens is not neutral. It takes a slot in every future auto-inject menu and crowds out something useful. Adds note_usage_events (migration 0071): one row per note per event, either 'surfaced' (we put its title in front of an agent) or 'pulled' (someone opened it in full), tagged with which surface produced it. Closes the gap #2082 recorded against this work. The write-path PLACE arm carries no score, so it has no home in retrieval_logs — folding it in would corrupt the score distribution that table exists to capture. The result was that the arm firing on the STRONGEST claim ("there is already a canonical helper in this exact file") was the one arm nobody could measure. Both arms now emit usage events under distinct sources, so their pull-through rates are finally comparable. Deliberate departures from the task as written: - Not in-session correlation. The original framing was "correlate result_ids against a later get_note in the same session." There is no session identity server-side — the MCP endpoint is stateless and the hooks send no session id — and adding one would mean threading an opaque client-supplied token through every read path. Two independent counters answer the question without it: surfaced 40×, pulled 0 is dead weight regardless of how those events distribute across sessions. - Pulls record at the ENTRY POINTS (MCP tools, REST detail route), not in snippets_svc.get_snippet, which update and merge also reach. Counting those would inflate precisely the number meant to say "someone chose to look at this." - get_note records for every note kind, not just snippets. The auto-inject menu surfaces tasks and processes too; scoping this to snippets would pin those at zero pulls forever and make them read as dead weight next to snippets that merely had a counter. Surfaced in the Snippets list as an "N/M used" badge, warning-toned once a record has been offered 3+ times and never opened, with the tooltip saying what to do about it (usually: its "when to reach for it" doesn't say when). No badge at all below one surfacing — "0/0" reads as a verdict when it's an absence of evidence. Also returned from MCP list_snippets so the agent can see dead weight without opening the UI. Telemetry keeps the retrieval_telemetry contract throughout: writes are fire-and-forget, reads degrade to zeroes, and no path can raise into the surface it observes. Refs #2085 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UaYUaouG9jjhATyuxCKrQs
557 lines
24 KiB
Python
557 lines
24 KiB
Python
"""Session-context rendering for the Scribe plugin's SessionStart hook.
|
|
|
|
The plugin's hook curls `GET /api/plugin/context` at session start and injects
|
|
the returned text as `additionalContext`, giving Scribe the same push channel
|
|
that superpowers and file-memory have. This module renders that text.
|
|
|
|
Design note — altitude: we inject rule *titles* grouped by topic (a compact
|
|
index), NOT every rule's full statement. The 48 always-on statements run well
|
|
past the 10k-char `additionalContext` cap, and the push channel's job is to make
|
|
Claude *aware* the rules exist and *reach* for them — not to dump them. Full
|
|
text stays one `get_rule(id)` / `list_always_on_rules()` call away. Titles are
|
|
mostly self-describing ("`dev` is home", "No GitHub — Fabled-Git only"), so the
|
|
index alone already steers behavior.
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
import logging
|
|
import re
|
|
import time
|
|
|
|
from sqlalchemy import select
|
|
|
|
from scribe.models import async_session
|
|
from scribe.models.rulebook import RulebookTopic
|
|
from scribe.services import knowledge as knowledge_svc
|
|
from scribe.services import notes as notes_svc
|
|
from scribe.services import projects as projects_svc
|
|
from scribe.services import rulebooks as rulebooks_svc
|
|
from scribe.services import snippets as snippets_svc
|
|
from scribe.services.access import label_shared_items, owner_names_for
|
|
from scribe.services.embeddings import semantic_search_notes
|
|
from scribe.services.note_usage import record_surfaced
|
|
from scribe.services.retrieval_telemetry import record_retrieval
|
|
from scribe.services.settings import get_setting
|
|
|
|
logger = logging.getLogger(__name__)
|
|
|
|
# Defensive cap below Claude Code's 10k additionalContext limit.
|
|
_MAX_CHARS = 9000
|
|
|
|
# Max chars of a Process body to fold into the auto-surface description.
|
|
_PROC_PREVIEW_CHARS = 200
|
|
|
|
# --- Knowledge auto-inject (Path A: per-turn awareness push) -----------------
|
|
# Per-user settings (keys live in the generic settings table). The threshold is
|
|
# deliberately STRICTER than the pull-search default (embeddings
|
|
# DEFAULT_SIMILARITY_THRESHOLD = 0.45): an unsolicited per-turn inject must clear
|
|
# a higher bar than a search the agent chose to run. Defaults start conservative
|
|
# and are meant to be tuned from retrieval_logs (source='auto_inject') once data
|
|
# accrues — they're exposed in the Settings UI, no restart needed.
|
|
AUTOINJECT_ENABLED_KEY = "kb_autoinject_enabled"
|
|
AUTOINJECT_THRESHOLD_KEY = "kb_autoinject_threshold"
|
|
AUTOINJECT_TOP_K_KEY = "kb_autoinject_top_k"
|
|
|
|
AUTOINJECT_DEFAULT_ENABLED = True
|
|
AUTOINJECT_DEFAULT_THRESHOLD = 0.55
|
|
AUTOINJECT_DEFAULT_TOP_K = 3
|
|
|
|
# The write-path trigger (#2082) gets its own on/off switch but SHARES the
|
|
# threshold and top-k above. One knob for "how loud may Scribe be" is easier to
|
|
# reason about than two that drift; the switch is separate because wanting prior
|
|
# art on writes while keeping prompts quiet (or the reverse) is a real
|
|
# preference, and it's the only part of the gate an operator can judge without
|
|
# data. The two surfaces log under different `source` values, so if the
|
|
# telemetry ever shows they want different thresholds, splitting them is a
|
|
# data-backed change rather than a guess made up front.
|
|
WRITEPATH_ENABLED_KEY = "kb_writepath_enabled"
|
|
WRITEPATH_DEFAULT_ENABLED = True
|
|
|
|
# Margin gate: drop any hit more than this far below the top hit's score, so a
|
|
# single strong match doesn't drag in a wall of barely-passing neighbours.
|
|
_AUTOINJECT_BAND = 0.10
|
|
# Hard ceiling on top-k regardless of the user's setting — this is an
|
|
# awareness menu (titles only), never a content dump.
|
|
_AUTOINJECT_MAX_TOP_K = 10
|
|
|
|
|
|
def _slugify(text: str) -> str:
|
|
"""kebab-case slug for a skill directory name (a-z0-9 + single hyphens)."""
|
|
s = re.sub(r"[^a-z0-9]+", "-", (text or "").lower()).strip("-")
|
|
return s or "process"
|
|
|
|
|
|
async def build_process_manifest(user_id: int) -> dict:
|
|
"""List the user's stored Processes as auto-surfacing skill-stub specs.
|
|
|
|
The plugin's sync script (scribe_sync_processes.sh) writes one
|
|
~/.claude/skills/scribe-proc-<slug>/SKILL.md per entry — `description` is the
|
|
auto-surface trigger, and the stub body calls get_process(name) for the live
|
|
procedure (single source of truth in the DB). Reuses the list_processes query
|
|
(note_type='process'). Instance-agnostic: derived from whatever Processes the
|
|
calling install owns, no operator-specific coupling.
|
|
|
|
SCOPE: this is the most consequential passive surface Scribe has — every
|
|
entry becomes a skill file on the operator's machine that auto-surfaces and
|
|
is followed as written. It therefore uses the BROWSE scope (via the
|
|
no-query knowledge list): a Process shared directly with the operator is
|
|
never installed here, only one they own or reach through a shared project
|
|
(decision note 2094). Project-shared entries are labelled with their owner so
|
|
the stub can't pass off someone else's procedure as the operator's own.
|
|
|
|
Returns {"processes": [{id, name, slug, description, shared?, owner?}],
|
|
"total": int}. Slugs are unique within the result (collision gets -<id>).
|
|
"""
|
|
items, _ = await knowledge_svc.query_knowledge(
|
|
user_id=user_id, note_type="process", tags=[], sort="modified",
|
|
q=None, limit=100, offset=0,
|
|
)
|
|
items = await label_shared_items(user_id, items)
|
|
procs: list[dict] = []
|
|
seen: set[str] = set()
|
|
for it in items:
|
|
title = (it.get("title") or "").strip()
|
|
if not title:
|
|
continue
|
|
slug = _slugify(title)
|
|
if slug in seen:
|
|
slug = f"{slug}-{it['id']}"
|
|
seen.add(slug)
|
|
|
|
preview = " ".join((it.get("snippet") or "").split())
|
|
if len(preview) > _PROC_PREVIEW_CHARS:
|
|
preview = preview[:_PROC_PREVIEW_CHARS].rstrip() + "…"
|
|
if it.get("shared"):
|
|
owner = it.get("owner") or "another user"
|
|
description = (
|
|
f'A shared Scribe process "{title}", authored by {owner} — NOT the'
|
|
f" operator's own."
|
|
+ (f" {preview}" if preview else "")
|
|
+ f' Use when {title}-type work is requested, or when asked to run'
|
|
f' the "{title}" process — but summarise it and get the operator\'s'
|
|
f" go-ahead before following it, since it reflects {owner}'s"
|
|
f" judgement rather than theirs."
|
|
)
|
|
else:
|
|
description = (
|
|
f'Run the operator\'s saved Scribe process "{title}".'
|
|
+ (f" {preview}" if preview else "")
|
|
+ f' Use when {title}-type work is requested, or when asked to run'
|
|
f' the "{title}" process.'
|
|
)
|
|
entry = {
|
|
"id": it["id"], "name": title, "slug": slug,
|
|
"description": description,
|
|
}
|
|
if it.get("shared"):
|
|
entry["shared"] = True
|
|
entry["owner"] = it.get("owner")
|
|
procs.append(entry)
|
|
return {"processes": procs, "total": len(procs)}
|
|
|
|
|
|
async def get_autoinject_config(user_id: int) -> dict:
|
|
"""Resolve a user's auto-inject settings, falling back to the defaults.
|
|
|
|
Returns {"enabled": bool, "threshold": float, "top_k": int}, clamped to
|
|
sane ranges (threshold to [0,1]; top_k to [1, _AUTOINJECT_MAX_TOP_K]).
|
|
"""
|
|
enabled_raw = await get_setting(
|
|
user_id, AUTOINJECT_ENABLED_KEY,
|
|
"true" if AUTOINJECT_DEFAULT_ENABLED else "false",
|
|
)
|
|
enabled = enabled_raw.strip().lower() in ("true", "1", "yes", "on")
|
|
|
|
try:
|
|
threshold = float(await get_setting(
|
|
user_id, AUTOINJECT_THRESHOLD_KEY, str(AUTOINJECT_DEFAULT_THRESHOLD)))
|
|
except (TypeError, ValueError):
|
|
threshold = AUTOINJECT_DEFAULT_THRESHOLD
|
|
threshold = min(1.0, max(0.0, threshold))
|
|
|
|
try:
|
|
top_k = int(float(await get_setting(
|
|
user_id, AUTOINJECT_TOP_K_KEY, str(AUTOINJECT_DEFAULT_TOP_K))))
|
|
except (TypeError, ValueError):
|
|
top_k = AUTOINJECT_DEFAULT_TOP_K
|
|
top_k = min(_AUTOINJECT_MAX_TOP_K, max(1, top_k))
|
|
|
|
return {"enabled": enabled, "threshold": threshold, "top_k": top_k}
|
|
|
|
|
|
def _record_kind(note) -> str:
|
|
"""The one-word kind marker for an injected menu line.
|
|
|
|
The menu is drawn from every record that carries an embedding, so a snippet,
|
|
a stored process, an issue and a stray dev-log all arrive looking identical.
|
|
Recorded prior art only stands out if the line says what it is — and the kind
|
|
is also what tells the reader which tool opens it.
|
|
|
|
Task-ness wins over `note_type` because it's the more useful distinction at a
|
|
glance: "there's an open issue about this" beats "there's a note about this".
|
|
"""
|
|
if note.is_task:
|
|
return "issue" if note.task_kind == "issue" else "task"
|
|
return note.note_type or "note"
|
|
|
|
|
|
async def build_autoinject_hint(
|
|
user_id: int,
|
|
query: str,
|
|
project_id: int = 0,
|
|
exclude_ids: list[int] | None = None,
|
|
) -> dict:
|
|
"""Title-first awareness hint for the plugin's UserPromptSubmit hook.
|
|
|
|
The four anti-bloat gates (see the module + milestone-93 design):
|
|
1. high-confidence threshold (stricter than pull) — set per-user;
|
|
2. margin gate — keep only hits within _AUTOINJECT_BAND of the top score;
|
|
3. session dedup — caller passes already-injected ids as `exclude_ids`;
|
|
4. title-first payload — id + kind + title + score only, never bodies.
|
|
Disabled, blank-query, or nothing-clears-the-gates all return empty context,
|
|
so most turns inject nothing.
|
|
|
|
Returns {"context": str, "note_ids": list[int], "config": dict}. Every
|
|
retrieval (even empty) is logged to retrieval_logs as source='auto_inject'
|
|
so the threshold can be tuned from data.
|
|
"""
|
|
cfg = await get_autoinject_config(user_id)
|
|
empty = {"context": "", "note_ids": [], "config": cfg}
|
|
q = (query or "").strip()
|
|
if not cfg["enabled"] or not q:
|
|
return empty
|
|
|
|
t0 = time.perf_counter()
|
|
hits = await semantic_search_notes(
|
|
user_id, q,
|
|
limit=cfg["top_k"],
|
|
threshold=cfg["threshold"],
|
|
project_id=(project_id or None),
|
|
exclude_ids=set(exclude_ids or []),
|
|
# Injection is the one retrieval nobody asked for, so it takes the BROWSE
|
|
# scope: never a record shared one-to-one with the operator. What can
|
|
# still appear is a collaborator's note inside a shared project — legible
|
|
# only because the line below names its owner.
|
|
scope="browse",
|
|
)
|
|
record_retrieval(
|
|
user_id=user_id, source="auto_inject", query=q,
|
|
threshold=cfg["threshold"], limit=cfg["top_k"],
|
|
project_id=(project_id or None), is_task=None, results=hits,
|
|
duration_ms=(time.perf_counter() - t0) * 1000.0,
|
|
)
|
|
if not hits:
|
|
return empty
|
|
|
|
# Margin gate: keep only hits close to the strongest one.
|
|
top_score = hits[0][0]
|
|
kept = [(s, n) for s, n in hits if s >= top_score - _AUTOINJECT_BAND]
|
|
|
|
# A collaborator's note can reach this menu via a shared project, and the
|
|
# operator never asked for it — so say whose it is. Unattributed, it reads as
|
|
# something they wrote and settled.
|
|
owners = await owner_names_for({
|
|
int(n.user_id) for _s, n in kept if n.user_id != user_id
|
|
})
|
|
|
|
# "records", not "notes" — the menu can hold snippets, processes and tasks
|
|
# too, and the kind marker on each line is only legible if the header doesn't
|
|
# already claim they're all one thing.
|
|
lines = [
|
|
"> Possibly relevant from your Scribe records — open any in full with "
|
|
"`get_note(id)`, or `get_snippet` / `get_process` for those kinds "
|
|
"(titles only; injected once per session):",
|
|
]
|
|
note_ids: list[int] = []
|
|
for score, note in kept:
|
|
note_ids.append(int(note.id))
|
|
title = (note.title or "(untitled)").replace("\n", " ").strip()
|
|
line = f"> - #{note.id} [{_record_kind(note)}] \"{title}\" ({score:.2f})"
|
|
if note.user_id != user_id:
|
|
who = owners.get(int(note.user_id)) or "another user"
|
|
line += f" — shared by {who}, treat as a suggestion"
|
|
lines.append(line)
|
|
|
|
# Records what SURVIVED the margin gate, not what the ranker returned — the
|
|
# menu the agent actually saw. retrieval_logs already holds the full
|
|
# candidate set for threshold tuning; conflating the two would make
|
|
# "surfaced" mean two different things depending on the surface (#2085).
|
|
record_surfaced(user_id=user_id, note_ids=note_ids, source="auto_inject")
|
|
|
|
return {"context": "\n".join(lines), "note_ids": note_ids, "config": cfg}
|
|
|
|
|
|
# --- Write-path trigger (#2082): prior art at the moment code is written ------
|
|
# Auto-inject above fires on the operator's prompt. The moment reuse is actually
|
|
# lost is later — when the AGENT decides mid-task to write a helper — and nothing
|
|
# fired there. This is that trigger: the plugin's PreToolUse hook on Write/Edit
|
|
# asks what prior art is already recorded for the file being written.
|
|
#
|
|
# Two arms, deliberately different in kind:
|
|
# - BY PLACE — a snippet recorded at this path (or in its directory) is prior
|
|
# art by definition, not by resemblance, so it isn't scored or thresholded.
|
|
# This is what the reverse lookup (#2083) was built to answer.
|
|
# - BY MEANING — semantic search over snippets only, using the code about to be
|
|
# written, under the same gates as auto-inject.
|
|
# Place beats meaning in the menu because "there is already a canonical helper in
|
|
# this exact file" is a stronger claim than "this resembles something".
|
|
|
|
|
|
def _prior_art_line(item: dict, marker: str, owner: str | None) -> str:
|
|
"""One menu line: `- #12 [here] "title"`, attributed when it isn't yours."""
|
|
title = (item.get("title") or "(untitled)").replace("\n", " ").strip()
|
|
line = f"> - #{item['id']} [{marker}] \"{title}\""
|
|
if owner:
|
|
line += f" — shared by {owner}, treat as a suggestion"
|
|
return line
|
|
|
|
|
|
async def get_writepath_config(user_id: int) -> dict:
|
|
"""Write-path trigger settings: its own `enabled`, auto-inject's gates."""
|
|
cfg = await get_autoinject_config(user_id)
|
|
enabled_raw = await get_setting(
|
|
user_id, WRITEPATH_ENABLED_KEY,
|
|
"true" if WRITEPATH_DEFAULT_ENABLED else "false",
|
|
)
|
|
return {
|
|
**cfg,
|
|
"enabled": enabled_raw.strip().lower() in ("true", "1", "yes", "on"),
|
|
}
|
|
|
|
|
|
async def build_write_path_hint(
|
|
user_id: int,
|
|
path: str,
|
|
code: str = "",
|
|
project_id: int = 0,
|
|
exclude_ids: list[int] | None = None,
|
|
) -> dict:
|
|
"""Prior-art hint for the plugin's PreToolUse hook on Write/Edit.
|
|
|
|
`path` is the file about to be written, REPO-RELATIVE — matching the
|
|
convention snippet locations are recorded in. `code` is what's about to be
|
|
written, used only as the semantic query.
|
|
|
|
Same four anti-bloat gates as auto-inject (threshold, margin, session dedup
|
|
via `exclude_ids`, titles-never-bodies) plus the shared top-k cap across BOTH
|
|
arms — so a file with a lot of recorded history can't turn one edit into a
|
|
wall of text. Returns empty context when disabled, when there's no path, or
|
|
when nothing is recorded — which is the common case, and the point.
|
|
|
|
Note the repo↔project mapping is deliberately one-way: the hook sends a git
|
|
remote, which the ROUTE resolves to `project_id` through the repo bindings.
|
|
It is never used as the location `repo` filter — a snippet's `repo` is a
|
|
free-text label the operator typed ("Scribe"), not a remote URL, and matching
|
|
one against the other would silently return nothing.
|
|
|
|
Returns {"context": str, "note_ids": list[int], "config": dict}. The semantic
|
|
arm is logged to retrieval_logs as source='write_path' — its own source, so
|
|
its precision is tunable separately from auto-inject's.
|
|
|
|
Location hits still carry no score and so stay out of retrieval_logs, whose
|
|
score distribution they would corrupt. What closed the gap (#2085) is that
|
|
un-scored surfacing now has its own home: BOTH arms emit note_usage_events,
|
|
tagged 'write_path_place' vs 'write_path_semantic', so the place arm is
|
|
finally measurable — and the two arms' pull-through rates are comparable,
|
|
which is the number that says whether place really does beat meaning here.
|
|
"""
|
|
cfg = await get_writepath_config(user_id)
|
|
empty = {"context": "", "note_ids": [], "config": cfg}
|
|
path = (path or "").strip()
|
|
if not cfg["enabled"] or not path:
|
|
return empty
|
|
|
|
top_k = cfg["top_k"]
|
|
excluded = set(exclude_ids or [])
|
|
scope_project = project_id or None
|
|
|
|
# --- arm 1: by place ---
|
|
here: list[dict] = []
|
|
nearby: list[dict] = []
|
|
try:
|
|
here, _ = await snippets_svc.list_snippets(
|
|
user_id, path=path, limit=top_k, project_id=scope_project,
|
|
)
|
|
directory = path.rsplit("/", 1)[0] if "/" in path else ""
|
|
if directory and len(here) < top_k:
|
|
nearby, _ = await snippets_svc.list_snippets(
|
|
user_id, path=directory, limit=top_k, project_id=scope_project,
|
|
)
|
|
except Exception:
|
|
logger.warning("Write-path location lookup failed", exc_info=True)
|
|
|
|
seen: set[int] = set(excluded)
|
|
placed: list[tuple[str, dict]] = []
|
|
for marker, items in (("here", here), ("nearby", nearby)):
|
|
for item in items:
|
|
nid = int(item["id"])
|
|
if nid in seen:
|
|
continue
|
|
seen.add(nid)
|
|
placed.append((marker, item))
|
|
|
|
# --- arm 2: by meaning ---
|
|
scored: list[tuple[str, dict]] = []
|
|
remaining = top_k - len(placed)
|
|
query = (code or "").strip()
|
|
if remaining > 0 and query:
|
|
t0 = time.perf_counter()
|
|
hits = await semantic_search_notes(
|
|
user_id, query,
|
|
limit=remaining,
|
|
threshold=cfg["threshold"],
|
|
project_id=scope_project,
|
|
exclude_ids=seen,
|
|
note_type="snippet",
|
|
# Same reasoning as auto-inject: nobody asked for this, so it takes
|
|
# the browse scope and never surfaces a one-to-one direct share.
|
|
scope="browse",
|
|
)
|
|
record_retrieval(
|
|
user_id=user_id, source="write_path", query=query,
|
|
threshold=cfg["threshold"], limit=remaining,
|
|
project_id=scope_project, is_task=False, results=hits,
|
|
duration_ms=(time.perf_counter() - t0) * 1000.0,
|
|
)
|
|
if hits:
|
|
top_score = hits[0][0]
|
|
for score, note in hits:
|
|
if score < top_score - _AUTOINJECT_BAND:
|
|
continue
|
|
scored.append((
|
|
f"similar {score:.2f}",
|
|
{"id": int(note.id), "title": note.title, "user_id": note.user_id},
|
|
))
|
|
|
|
menu = (placed + scored)[:top_k]
|
|
if not menu:
|
|
return empty
|
|
|
|
owners = await owner_names_for({
|
|
int(it["user_id"]) for _m, it in menu
|
|
if it.get("user_id") is not None and int(it["user_id"]) != user_id
|
|
})
|
|
|
|
lines = [
|
|
f"> Prior art already recorded in Scribe for `{path}` — open one with "
|
|
"`get_snippet(id)` and reuse it rather than writing a fresh one-off "
|
|
"(titles only; shown once per session):",
|
|
]
|
|
note_ids: list[int] = []
|
|
for marker, item in menu:
|
|
note_ids.append(int(item["id"]))
|
|
owner_id = item.get("user_id")
|
|
owner = None
|
|
if owner_id is not None and int(owner_id) != user_id:
|
|
owner = owners.get(int(owner_id)) or "another user"
|
|
lines.append(_prior_art_line(item, marker, owner))
|
|
|
|
# Split by arm, which is the whole reason this table exists. The place arm
|
|
# carries no score and so has no home in retrieval_logs; before #2085 a
|
|
# snippet surfaced BY PLACE left no trace anywhere, making the arm that
|
|
# fires on the strongest possible claim ("there is already a canonical
|
|
# helper in this exact file") the one arm nobody could measure.
|
|
by_arm: dict[str, list[int]] = {}
|
|
for marker, item in menu:
|
|
arm = "write_path_place" if marker in ("here", "nearby") else "write_path_semantic"
|
|
by_arm.setdefault(arm, []).append(int(item["id"]))
|
|
for arm, ids in by_arm.items():
|
|
record_surfaced(user_id=user_id, note_ids=ids, source=arm)
|
|
|
|
return {"context": "\n".join(lines), "note_ids": note_ids, "config": cfg}
|
|
|
|
|
|
async def _topic_titles(topic_ids: set[int]) -> dict[int, str]:
|
|
"""Map topic_id -> title for the given ids (live topics only)."""
|
|
if not topic_ids:
|
|
return {}
|
|
async with async_session() as session:
|
|
rows = await session.execute(
|
|
select(RulebookTopic.id, RulebookTopic.title).where(
|
|
RulebookTopic.id.in_(topic_ids),
|
|
RulebookTopic.deleted_at.is_(None),
|
|
)
|
|
)
|
|
return {tid: title for tid, title in rows.all()}
|
|
|
|
|
|
async def build_session_context(
|
|
user_id: int, project_id: int = 0, unbound_repo: str = ""
|
|
) -> dict:
|
|
"""Render the SessionStart context for a user, optionally project-scoped.
|
|
|
|
Args:
|
|
user_id: the operator.
|
|
project_id: the resolved active project (0 = none). The endpoint
|
|
resolves this from the working repo's remote, not from config.
|
|
unbound_repo: when the hook sent a repo remote that maps to no project,
|
|
its normalized key — triggers a one-line "bind this repo" hint so
|
|
the binding is self-healing.
|
|
|
|
Returns {"context": str, "rule_count": int, "project": dict | None}.
|
|
`context` is markdown ready to drop into `additionalContext`; it is capped
|
|
at _MAX_CHARS with an explicit truncation note so the hook can pass it
|
|
through verbatim.
|
|
"""
|
|
rules = await rulebooks_svc.list_always_on_rules(user_id)
|
|
topic_map = await _topic_titles({r.topic_id for r in rules if r.topic_id})
|
|
|
|
lines: list[str] = [
|
|
"# Scribe — standing session context (auto-injected by the Scribe plugin)",
|
|
"",
|
|
"You are working with Scribe, the operator's self-hosted second brain. "
|
|
"The always-on rules below are BINDING this session. Titles only — full "
|
|
"text via `list_always_on_rules()` or `get_rule(id)`.",
|
|
"",
|
|
"## Always-on rules (by topic)",
|
|
]
|
|
|
|
# rules already arrive ordered by rulebook/topic/order, so grouping by
|
|
# consecutive topic_id preserves the intended sequence.
|
|
current_topic: int | None = object() # sentinel distinct from any id/None
|
|
for r in rules:
|
|
if r.topic_id != current_topic:
|
|
current_topic = r.topic_id
|
|
heading = topic_map.get(r.topic_id, "ungrouped") if r.topic_id else "ungrouped"
|
|
lines.append(f"### {heading}")
|
|
lines.append(f"- [{r.id}] {r.title}")
|
|
|
|
project_dict: dict | None = None
|
|
if project_id:
|
|
project = await projects_svc.get_project(user_id, project_id)
|
|
if project is not None:
|
|
_, open_count = await notes_svc.list_notes(
|
|
user_id, is_task=True, status="todo", project_id=project_id, limit=1,
|
|
)
|
|
goal = (getattr(project, "goal", "") or "").strip()
|
|
project_dict = {"id": project.id, "title": project.title}
|
|
lines += [
|
|
"",
|
|
f"## Active project: {project.title} (id {project.id})",
|
|
f"Goal: {goal[:200]}" if goal else "",
|
|
f"Open todo tasks: {open_count}",
|
|
]
|
|
elif unbound_repo:
|
|
lines += [
|
|
"",
|
|
"## Repository not yet bound",
|
|
f"This repo (`{unbound_repo}`) isn't mapped to a Scribe project, so "
|
|
"no project context was loaded. Bind it once with "
|
|
f'`bind_repo(repo_url="{unbound_repo}", project_id=<id>)` '
|
|
"(call `list_projects` to find the id) and future sessions here will "
|
|
"auto-load that project's context.",
|
|
]
|
|
|
|
lines += [
|
|
"",
|
|
"Reflex: search Scribe (search / list_tasks / list_notes, scoped to the "
|
|
"active project) before answering or starting work; prefer UPDATING an "
|
|
"existing note/rule over creating a new one.",
|
|
]
|
|
|
|
context = "\n".join(line for line in lines if line is not None)
|
|
if len(context) > _MAX_CHARS:
|
|
context = context[:_MAX_CHARS].rstrip() + "\n\n…(truncated — call list_always_on_rules())"
|
|
|
|
return {"context": context, "rule_count": len(rules), "project": project_dict}
|