Merge pull request 'Guidance by ownership, placement and report-back, and no more guessed record ids' (#155) from dev into main
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Build & push image (push) Successful in 14s
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Build & push image (push) Successful in 14s
This commit was merged in pull request #155.
This commit is contained in:
@@ -1,7 +1,7 @@
|
||||
{
|
||||
"name": "scribe",
|
||||
"description": "Scribe system-of-record for Claude Code: MCP tools over your notes/tasks/projects/rules, a session-start push channel that surfaces your active-project context, process-skills (writing-plans, systematic-debugging, verification, brainstorming, reusing-code), and your saved Scribe Processes auto-surfaced as skills (/scribe:sync). Replaces superpowers + file-memory with one app-backed plugin.",
|
||||
"version": "2026.09.11.2026",
|
||||
"description": "Scribe for Claude Code: connects the scribe MCP server, adds the hooks that deliver live project state and relevant records at the right moment, ships the shared client-neutral Scribe skills (using-scribe, writing-plans, reporting-back, systematic-debugging, verification, brainstorming, reusing-code, shape-accounting), and syncs your saved Scribe Processes as skills (/scribe:sync).",
|
||||
"version": "2026.09.14.1723",
|
||||
"author": {
|
||||
"name": "Bryan Van Deusen"
|
||||
},
|
||||
|
||||
@@ -0,0 +1,62 @@
|
||||
# Packaging Scribe for an agent client
|
||||
|
||||
This is the maintainer contract for shipping Scribe into an agent client:
|
||||
what every client package shares, what each one adds, and where a new client's
|
||||
files go. It follows from decision #4027 (milestone 410): **the server orients,
|
||||
the skills hold the depth, and a client adapter only times delivery and says
|
||||
what that one client needs said.**
|
||||
|
||||
Claude Code is the only client built today. The layout below is kept so adding
|
||||
another one means adding files, not moving or rewriting any.
|
||||
|
||||
## Shared by every client package
|
||||
|
||||
| Piece | Where | What it is |
|
||||
|---|---|---|
|
||||
| **The skills** | `plugin/skills/*/SKILL.md` | Agent Skills (the open SKILL.md format). They state every Scribe reflex in full and name no client. `tests/test_guidance_ownership.py` fails if a skill names a particular client, or references anything outside its own folder. Every client package ships this folder verbatim. |
|
||||
| **The MCP server** | `<base URL>/mcp` | HTTP, `Authorization: Bearer <fmcp_ key>`. Its `_INSTRUCTIONS` is a client-neutral index (≤2,000 chars); each tool's description carries its contract; in-band responses (`placement`, `report_back`, `systems_hint`, the duplicate gate, the guessed-id refusal) fire in every client. |
|
||||
| **The adapter API** | `<base URL>/api/plugin/*` | Plain `GET` endpoints any client's hooks can call with the same key (read scope is enough): `context` (live session state), `retrieve` (rules, preferences and notes for a message), `prior-art` (records and shape-ledger hints for code being written), `tool-rules` (rules for a command about to run), `processes` (stored Processes to expose as skills). |
|
||||
| **The API key** | Scribe → Settings → API Keys | One `fmcp_` key per install. Read scope for hooks; write scope for the MCP tools. |
|
||||
|
||||
## Added by each client
|
||||
|
||||
| Piece | What it holds | Rule of thumb |
|
||||
|---|---|---|
|
||||
| **A manifest** | Name, version, how to reach the MCP server with the key, where the skills and hooks are | Only that client's format. |
|
||||
| **Hooks** | Scripts in that client's hook format that call the adapter API at the right moment and print what comes back | Timing and transport only. A hook may print status lines (unconfigured, unreachable, running version) and pointers; guidance text comes from the server or the adapter's own static text. |
|
||||
| **Commands** | That client's command wrappers, e.g. a process sync | Optional. |
|
||||
| **Adapter static text** | What only that client needs said: its local memory files, its compaction command, its commands | Never a copy of a skill or of `_INSTRUCTIONS`. The loss guard and ownership registry in `tests/test_guidance_ownership.py` cover it. |
|
||||
|
||||
## The worked example: Claude Code
|
||||
|
||||
| File | Role |
|
||||
|---|---|
|
||||
| `.claude-plugin/plugin.json` | Manifest: `mcpServers.scribe` (HTTP, `Authorization: Bearer ${user_config.api_token}`), `userConfig` for the base URL and key, version (minted, never hand-edited; see README). |
|
||||
| `../.claude-plugin/marketplace.json` (repo root) | The marketplace entry pointing at `./plugin`. |
|
||||
| `hooks/hooks.json` | Wires the scripts to Claude Code events. |
|
||||
| `hooks/scribe_session_context.sh` | SessionStart: adapter static text + running version + `GET /api/plugin/context`; a reload banner after compaction. |
|
||||
| `hooks/scribe_autoinject.sh` | UserPromptSubmit: `GET /api/plugin/retrieve` for the message. |
|
||||
| `hooks/scribe_prior_art.sh` | PreToolUse on editor writes: `GET /api/plugin/prior-art`. |
|
||||
| `hooks/scribe_after_write.sh` | PostToolUse on shell commands: the same check for code written through the shell. |
|
||||
| `hooks/scribe_tool_rules.sh` | PreToolUse on shell commands: `GET /api/plugin/tool-rules`. |
|
||||
| `hooks/scribe_sync_processes.sh` + `commands/sync.md` | `GET /api/plugin/processes` → `~/.claude/skills/scribe-proc-*` stubs; `/scribe:sync` on demand. |
|
||||
| `hooks/scribe_defs.sh` | Shared shell helpers: config, dedup ledgers, outage line. |
|
||||
| `hooks/scribe_static_context.md` | The adapter static text. |
|
||||
|
||||
Hook config arrives as `CLAUDE_PLUGIN_OPTION_API_ENDPOINT` / `CLAUDE_PLUGIN_OPTION_API_TOKEN` (uppercased by Claude Code, #2198), with `SCRIBE_URL` / `SCRIBE_TOKEN` as an override.
|
||||
|
||||
## How other clients would map
|
||||
|
||||
**Researched, not tested** (spike #4023, September 2026). Re-check each client's current docs before building.
|
||||
|
||||
- **Agent Plugins 1.0** (Codex, Cursor, GitHub Copilot / VS Code, Kiro, ChatGPT): a root `plugin.json`, a `skills/` folder, and an `mcp.json`. That is the same `plugin/skills/` plus two small files. Hooks, commands and rules are not part of v1; each client adds its own under a reverse-domain directory (e.g. `com.<client>/`) that other clients ignore.
|
||||
- **Gemini CLI**: `gemini-extension.json` configuring the MCP server, a context file, and bundled skills.
|
||||
- **Client-native formats also exist** (`.codex-plugin/`, `.cursor-plugin/`) where a client wants more than the shared standard carries.
|
||||
- **Claude Code** does not read Agent Plugins: only `.claude-plugin/plugin.json`.
|
||||
|
||||
## Open questions for whoever builds the second package
|
||||
|
||||
1. **Passing the key to the MCP server.** Claude Code substitutes `${user_config.api_token}` into the header. Agent Plugins' `mcp.json` and Gemini's extension config have their own variable and secret handling. Decide per client, and keep the key out of files that get committed.
|
||||
2. **Coexistence in one folder.** Whether a root `plugin.json` (Agent Plugins) beside `.claude-plugin/plugin.json` changes how Claude Code loads the plugin is untested. Test it before shipping both from `plugin/`, or give the second client its own package directory that reuses `plugin/skills/`.
|
||||
3. **Hook parity.** Other clients' hook events differ and are partly implemented. Map each Scribe hook to the nearest event, and where none exists, leave that timing to the in-band server responses rather than writing a copy of the guidance.
|
||||
4. **Process skills.** The Claude Code sync writes `~/.claude/skills`. Another client needs its own skills location, or the MCP skills extension (SEP-2640) once a client supports it.
|
||||
+17
-5
@@ -5,13 +5,17 @@ instance into a first-class Claude Code extension:
|
||||
|
||||
- **MCP tools** over your notes, tasks, projects, milestones, systems, and
|
||||
rulebook (the `scribe` server).
|
||||
- **Session-start push channel** — a `SessionStart` hook injects your
|
||||
rules + active-project context so Scribe surfaces *without being asked*.
|
||||
- **Session-start push channel** — a `SessionStart` hook injects the
|
||||
active project's live state (from the server) and this adapter's short
|
||||
Claude Code guidance, so Scribe surfaces *without being asked*. Rules are
|
||||
never preloaded; they arrive by retrieval when your work matches one.
|
||||
- **Prior-art recall on writes** — a `PreToolUse` hook on Write/Edit checks the
|
||||
file about to be written against your recorded snippets (what's kept at that
|
||||
path, and what resembles the code) and offers them before the helper is
|
||||
rewritten. Titles only, never blocks the edit.
|
||||
- **Universal process-skills** — using-scribe, writing-plans,
|
||||
- **The shared Scribe skills** — client-neutral Agent Skills, the same files
|
||||
any client's package would ship: using-scribe, writing-plans, reporting-back
|
||||
(reply to the operator in a shape that says where the work stands),
|
||||
systematic-debugging, verification, brainstorming, reusing-code (record and
|
||||
recall reusable code as snippets). Replaces superpowers.
|
||||
- **Your Scribe Processes as skills** — saved Processes are synced into local
|
||||
@@ -19,8 +23,16 @@ instance into a first-class Claude Code extension:
|
||||
stub fetches the live procedure via `get_process`. Refreshed each session and
|
||||
on demand with `/scribe:sync`.
|
||||
|
||||
It is designed so you can uninstall `superpowers` and disable auto-memory and
|
||||
depend on neither.
|
||||
It is designed so you can uninstall `superpowers` and depend on Scribe instead
|
||||
of auto-memory — leave auto-memory at its default; Scribe replaces its job by
|
||||
holding the one copy, not by switching it off.
|
||||
|
||||
**How the pieces divide the work** (decision #4027): the Scribe server orients
|
||||
every MCP client and serves live state; the skills in `skills/` state every
|
||||
reflex in full and name no client; this plugin is the Claude Code adapter —
|
||||
hooks that deliver at the right moment, `/scribe:sync`, and the few things only
|
||||
Claude Code needs said (`hooks/scribe_static_context.md`). Packaging Scribe for
|
||||
another client: see [PACKAGING.md](PACKAGING.md).
|
||||
|
||||
## Install
|
||||
|
||||
|
||||
@@ -1,11 +1,12 @@
|
||||
#!/usr/bin/env bash
|
||||
# Scribe plugin — SessionStart push channel (two tiers + compaction re-grounding).
|
||||
#
|
||||
# Tier 1 (STATIC, always fires, no auth, no network): injects a bundled
|
||||
# behavioral mandate (scribe_static_context.md) so a fresh session knows to
|
||||
# reach for Scribe — record work, recall before acting — even when the instance
|
||||
# is unreachable or unconfigured. The static tier is the load-bearing floor that
|
||||
# does not depend on the key or the network.
|
||||
# Tier 1 (STATIC, always fires, no auth, no network): injects the Claude Code
|
||||
# adapter's own guidance (scribe_static_context.md) — where Scribe's reflexes
|
||||
# are stated (the using-scribe skill), Claude Code's memory files, /compact,
|
||||
# /scribe:sync, and what to do when Scribe is unavailable. Since milestone 410
|
||||
# (decision #4027) it carries only what this client needs said; the reflexes
|
||||
# themselves are owned by the shared skills and the server's index.
|
||||
#
|
||||
# Tier 2 (DYNAMIC, best-effort enrichment): curls the operator's Scribe instance
|
||||
# for active-project context and appends it. Config comes from
|
||||
@@ -160,7 +161,7 @@ if [ -n "$url" ] && [ -n "$token" ] && command -v curl >/dev/null 2>&1; then
|
||||
# (milestone 394). Nothing is preloaded, so there is no set whose
|
||||
# drift a later write could be told about — a rule is retrieved at
|
||||
# the moment it applies, which cannot be stale.
|
||||
[ -z "$dyn" ] && status="> ⚠️ Scribe: live project context could not be loaded this session (instance unreachable or request failed). The standing guidance above still applies — ask for rules with \`search(content_type=\"rule\")\` and project context with \`enter_project()\` as needed."
|
||||
[ -z "$dyn" ] && status="> ⚠️ Scribe: live project context could not be loaded this session (instance unreachable or request failed). The using-scribe skill still applies — ask for rules with \`search(content_type=\"rule\")\` and project context with \`enter_project()\` as needed."
|
||||
elif [ -n "$url" ] && [ -z "$token" ]; then
|
||||
status="> ⚠️ Scribe: live context disabled this session — the API key is not configured (Scribe base URL is). Set it with \`/plugin\` → Scribe → configure, or export SCRIBE_TOKEN. Tools still work; ask for rules with \`search(content_type=\"rule\")\` and project context with \`enter_project()\`."
|
||||
elif [ -z "$url" ] && [ -z "$token" ]; then
|
||||
|
||||
@@ -1,121 +1,31 @@
|
||||
# Scribe — your system of record
|
||||
# Scribe — Claude Code adapter
|
||||
|
||||
This environment has the **Scribe** plugin: the operator's self-hosted system
|
||||
of record (notes, tasks, projects, milestones, rules) reachable through the
|
||||
`scribe` MCP tools. Treat Scribe — **not local files** — as the source of truth
|
||||
for the operator's work, and as your own working memory across sessions.
|
||||
This session is connected to **Scribe**, the operator's system of record for
|
||||
their work, through the `scribe` MCP tools. How to work with Scribe is stated
|
||||
once, in the **`using-scribe`** skill: reach for it at the start of the session
|
||||
and whenever you are unsure what Scribe expects. Each tool's contract is in its
|
||||
description, and the process skills (writing-plans, reporting-back,
|
||||
reusing-code, systematic-debugging, verification, brainstorming,
|
||||
shape-accounting) carry their arcs.
|
||||
|
||||
**At the start of this session:**
|
||||
- You hold none of the operator's rules, and there is no call that loads them
|
||||
all. Rules arrive when something you are about to do matches one.
|
||||
- If the working repo maps to a Scribe project (check `list_repo_bindings`),
|
||||
call `enter_project(<id>)` to load that project's rules, open tasks, and
|
||||
recent notes in one shot.
|
||||
What only Claude Code needs said:
|
||||
|
||||
**While you work:**
|
||||
- **Operator rules govern consequential actions** — before any git branch /
|
||||
commit / push, or any other hard-to-reverse or outward-facing action, the
|
||||
operator's Scribe rules decide what to do — NOT generic conventions baked
|
||||
into the harness or your defaults (e.g. "branch before committing," "open a
|
||||
feature branch per task," "push to a fork"). If you have not loaded the
|
||||
no rule has arrived for the act in front of you, `search(content_type=
|
||||
"rule")` BEFORE acting rather than falling back on a default habit. When a
|
||||
retrieved rule and a default habit disagree, the rule wins; if no rule
|
||||
speaks to it, ask rather than assume.
|
||||
- **Rules bind; preferences do not.** A record's `kind` says which. A **rule**
|
||||
must be followed — ignoring it breaks something or crosses a boundary. A
|
||||
**preference** is how the operator wants work done: worth following for
|
||||
consistency, not a defect to miss. Injected lines name the kind in their
|
||||
opening words. A preference is also yours to keep current when they correct
|
||||
you (`update_preference`); a rule waits for them.
|
||||
- **Silence is not absence.** Nothing is preloaded: every rule is RETRIEVED,
|
||||
when what you are doing resembles what the rule is about. Most turns
|
||||
retrieve none, and a rule you were never handed binds exactly as hard as one
|
||||
you were. So before a consequential act, `search` for a rule about it
|
||||
(`content_type="rule"`) rather than concluding from an empty session that
|
||||
nothing applies. "I was not told" is not the same as "there is no rule," and
|
||||
only one of those is checkable.
|
||||
This bites hardest on which TOOL to reach for — curling an API that has an
|
||||
MCP client, standing up a local stack, running a suite CI owns. Those feel
|
||||
like mechanics rather than decisions, so they raise no doubt and generate no
|
||||
query; the moment you are most confident is the moment to look.
|
||||
- **Recall before acting** — before you answer anything about the operator's
|
||||
work or start a task, `search` Scribe first; assume a related note, task, or
|
||||
decision already exists. Concretely, reach for recall whenever a request
|
||||
touches the operator's projects, people, places, prior decisions, or existing
|
||||
work: check for an existing task before opening a new one, and for a prior
|
||||
note/decision before re-deriving one. When a project is in scope (you entered
|
||||
one), pass its id to `search` so results stay scoped to it. Treating Scribe as
|
||||
the first place you look — not just somewhere you write — is what makes it a
|
||||
trustworthy record.
|
||||
- **Record as you go** — track work as Scribe tasks and log progress with
|
||||
`add_task_log`. Always log when you **complete a task** and when you **hit or
|
||||
discover a problem** — so changes of direction are captured, not just
|
||||
successes. Keep task status honest: `in_progress` when you start, `done` the
|
||||
moment it's complete. When you **fix** something — even in passing — record it
|
||||
as its own issue (`create_task(kind="issue")`), not as a work-log line on an
|
||||
unrelated open task.
|
||||
- **Tag to Systems as you write** — `enter_project` lists the project's
|
||||
Systems (its named subsystems/areas). When you create or meaningfully update
|
||||
a record, ask which areas it is about and pass `system_ids`; if an area has
|
||||
no System yet, create it with `create_system` (name + a one-paragraph
|
||||
charter) rather than leaving it unmodelled. Cross-cutting records — audits,
|
||||
sweeps, reviews — take SEVERAL tags, and are the best moment to DISCOVER
|
||||
missing Systems: a pass that walks the subsystems has just enumerated the
|
||||
vocabulary, so mint what it names. Create liberally; the duplicate gate on
|
||||
`create_system` (and reviewing the existing list) is the guardrail against
|
||||
sprawl, not restraint. Every read and write of a project record shows its
|
||||
`systems` — that is the "am I in a System's territory?" signal, and
|
||||
`list_system_records` reads that territory's whole pile before you work in
|
||||
it. An untagged project record carries the `systems_hint` question instead,
|
||||
on creates, updates, and work-logs alike — treat it as the tagging question
|
||||
asked at the moment of work, not as noise to skip past.
|
||||
- **The pattern library: start from recorded shapes, and record every shape
|
||||
at first build** — recorded **snippets** are the project's pattern library,
|
||||
not a dedup net. Before building ANY shape — a button, an input field, a
|
||||
modal, a route handler, a service class, a test scaffold, up through complex
|
||||
subsystem patterns — search snippets and START from the recorded shape; a
|
||||
deliberate departure is recorded as its own named variant, never left as
|
||||
silent drift. And the FIRST time a shape is built, record it with
|
||||
`create_snippet` (name, when-to-reach-for-it, location, code) in the same
|
||||
breath — do not judge whether it "might recur": the builder of the first
|
||||
instance can never know, and a missed record is invisible until it
|
||||
resurfaces as an uninformed duplicate. A mature project's snippet corpus
|
||||
should read as a map of every shape in it. The backstop still holds:
|
||||
noticing the second copy of anything, or consolidating copies into a shared
|
||||
X, means X gets recorded before that work is finished — which is how a
|
||||
codebase is kept from growing four `.btn-primary` definitions. The write-path
|
||||
hooks (before a Write/Edit, and after any Bash call that changed the tree)
|
||||
name a known duplicate family or a canon elsewhere for what was just
|
||||
written — act on that line at the write, not at the next audit.
|
||||
- Do **not** keep the operator's rules, plans, or project notes in local
|
||||
memory / CLAUDE.md in parallel with Scribe — Scribe holds the single copy.
|
||||
- **Compact at clean seams** — because you record as you go, a context
|
||||
compaction is safe: the durable record lives in Scribe, not the transcript.
|
||||
After finishing a block of work in a long session, make sure in-flight state
|
||||
is logged to Scribe, then tell the operator it's a good, safe moment to
|
||||
`/compact` (name what you logged). You can't run it yourself — surface the
|
||||
recommendation and let them decide. Suggest it at seams, not every turn.
|
||||
|
||||
**How the instruction surfaces divide the work:** this file carries the
|
||||
session-level reflexes (WHEN to reach for Scribe); each tool's own description
|
||||
carries its full contract (HOW to call it — read it when you load the tool);
|
||||
the bundled skills carry process arcs (planning, debugging, verification). The
|
||||
MCP server's instruction block is deliberately only a map — the client injects
|
||||
roughly its first 2,000 characters and silently cuts the rest, so nothing
|
||||
load-bearing lives below that fold.
|
||||
|
||||
**If two Scribe instruction surfaces disagree** — this file, the MCP server's
|
||||
tool instructions, the `using-scribe` skill — **follow the one that assumes
|
||||
least about its own delivery.** This file is the floor: it ships with the
|
||||
plugin and needs no API key and no network, so it still applies in exactly the
|
||||
session where the others never arrived. The others may elaborate on what is
|
||||
written here; they must not contradict it. Weigh a disagreement by which way it
|
||||
fails, not by which surface said more: doing something a push would also have
|
||||
covered costs one redundant call, while skipping it because you expected a push
|
||||
that never came means working without the operator's rules and not knowing.
|
||||
A contradiction between surfaces is a defect in the product — say so, so it
|
||||
gets recorded and fixed rather than silently arbitrated again next session.
|
||||
|
||||
If the Scribe tools are unavailable, say so rather than silently falling back
|
||||
to local notes.
|
||||
- **Keep one copy — in Scribe, not Claude Code's local memory.** The
|
||||
operator's rules, plans and project notes go to Scribe, not also to
|
||||
`CLAUDE.md` or auto-memory. Leave auto-memory at its default setting: you
|
||||
replace its job by doing the work in Scribe, not by switching it off.
|
||||
- **Lines injected beside your work are retrieval.** When the operator sends a
|
||||
message, and before a write or a command, Scribe may add rules, preferences,
|
||||
notes and prior art that resemble what you are doing. Open the ones that
|
||||
apply. They are what matched, never the whole set — using-scribe says how to
|
||||
ask for the rest.
|
||||
- **Compact at clean seams.** Because work is recorded as you go, a compaction
|
||||
is safe once in-flight state is logged. After finishing a block of work in a
|
||||
long session, log it to Scribe, then tell the operator it's a good moment to
|
||||
`/compact` and name what you logged. You can't run it yourself; suggest it at
|
||||
seams, not every turn.
|
||||
- **Stored Processes arrive as skills** (`scribe-proc-*`), refreshed at session
|
||||
start. After a Process is added or edited, `/scribe:sync` makes it available
|
||||
straight away.
|
||||
- **If the Scribe tools are unavailable, say so** rather than silently falling
|
||||
back to local notes.
|
||||
|
||||
@@ -0,0 +1,122 @@
|
||||
---
|
||||
name: reporting-back
|
||||
description: Use when you are about to write the reply the operator will read — work finished, a task marked done, stopping on a blocker, asking them to decide or to do something, answering "where are we" / "what's next", or proposing an approach. Shapes the reply around where the work stands (which task, what changed, what needs them, what comes next) instead of the order you did things in. Triggers on reporting completion, handing off, asking a question, or summarising progress.
|
||||
---
|
||||
|
||||
# Reporting back
|
||||
|
||||
Your reply is where the operator finds out what happened. They were not there
|
||||
while you worked: they don't hold the files you read, the names you used or
|
||||
the order you did things in. A reply that follows *your* path is accurate and
|
||||
still unreadable to them. Shape it around **where the work stands**.
|
||||
|
||||
Pick the kind of reply first (the tables below), then fill its sections. The
|
||||
sections are what lets the operator find things at a glance, so keep them even
|
||||
when one is short — "**Needs you:** nothing" is an answer they were looking for.
|
||||
|
||||
## Every reply
|
||||
|
||||
- **Conclusion first.** The verdict, the result, or the question — before the
|
||||
reasoning that supports it.
|
||||
- **One topic per section.** Two things the operator raised get two sections.
|
||||
- **Make priority visible.** Bold the few things that matter; let the rest be
|
||||
plain.
|
||||
- **End with the ask, in bold** — the one thing they need to decide or do. If
|
||||
there is nothing, say so.
|
||||
- **Plain words.** Use the operator's vocabulary, not the names you coined while
|
||||
working. If a term has to appear, explain it once.
|
||||
- **Place the work in Scribe.** Name the task, issue or milestone it belongs to,
|
||||
by id *and* title (using-scribe: "Name the record, never just its number").
|
||||
|
||||
## Take the placement from the record
|
||||
|
||||
A remembered milestone title or "next step" reads exactly like a real one when
|
||||
it is wrong. So take placement from Scribe:
|
||||
|
||||
- `update_task` and `create_task` return a **`placement`** block — the project,
|
||||
the milestone, `position` (step N of M), `progress`, and `next` (the next open
|
||||
step). Use those values as they came back.
|
||||
- For a wider view, `get_milestone` (a plan and its steps) or `enter_project`
|
||||
(the whole project).
|
||||
- Work with no task behind it: say so plainly — "this wasn't tracked as a
|
||||
task" — and offer to record it. An honest "untracked" is a placement too.
|
||||
|
||||
## Reports — work happened
|
||||
|
||||
| Kind | Sections |
|
||||
|---|---|
|
||||
| **Completion** | Where this sits · What now works · How / why · Needs you · Next |
|
||||
| **Finding** (a problem you found and did not fix) | Symptom · Cause · Size of the fix · **Offer to fix it** |
|
||||
| **Blocked / failed** | What stopped · What you tried · What you need from them |
|
||||
| **Progress** (mid-work) | One or two lines: where things are, what's next, any blocker |
|
||||
| **Where are we** | The milestone and its progress · Done · Open · Needs you · Next |
|
||||
|
||||
## Asks — the operator needs to act or decide
|
||||
|
||||
| Kind | Sections |
|
||||
|---|---|
|
||||
| **Decision** | The question first · 2–4 options, each with what it changes · recommendation first |
|
||||
| **Clarification** | "My reading is X · the gap is Y · unless you say otherwise I'll do Z" |
|
||||
| **Handoff** (only they can do it) | The action · why it needs them · what it unblocks · what you'll do after · any way to skip it |
|
||||
| **Conflict** (what you're about to do clashes with a rule, a plan or an earlier decision) | What it says · what you were about to do · where they clash · A or B? |
|
||||
|
||||
Before asking, check whether you can find the answer yourself — something that
|
||||
can be read or looked up is a fact to check, not a question to send.
|
||||
|
||||
## Answers — the operator asked something
|
||||
|
||||
| Kind | Sections |
|
||||
|---|---|
|
||||
| **Explanation** | The answer first · then the evidence, pointing at what they could open to check it |
|
||||
| **Evaluation** ("can we / should we") | Verdict · What exists · The gaps · Recommendation |
|
||||
|
||||
## Proposals — shaping future work
|
||||
|
||||
| Kind | Sections |
|
||||
|---|---|
|
||||
| **Options** | 2–3 approaches · the trade-off of each · one recommendation |
|
||||
| **Plan** | Goal · Steps · Open questions — for review before starting (writing-plans) |
|
||||
| **Review** | Findings ranked by how much they matter, one per item |
|
||||
|
||||
## The completion report, in full
|
||||
|
||||
The most common reply, and the one most often written in the order the work
|
||||
happened. The shape:
|
||||
|
||||
> **Where this sits:** milestone 12 "Move the backups offsite", step 3 of 5.
|
||||
> Task #340 "Schedule the nightly sync" is done.
|
||||
>
|
||||
> **What now works**
|
||||
> - The nightly sync runs at 02:00 and copies the photo library to the remote
|
||||
> store.
|
||||
>
|
||||
> **How / why**
|
||||
> - Used the scheduler the other jobs already use, so there is one place to
|
||||
> look when a job doesn't run.
|
||||
> - Verified by running it once by hand and checking the remote copy's size
|
||||
> matches the source.
|
||||
>
|
||||
> **Needs you:** nothing.
|
||||
>
|
||||
> **Next:** #341 "Alert when a sync fails". Starting it unless you redirect.
|
||||
|
||||
Notes on each section:
|
||||
|
||||
- **Where this sits** — from `placement`. If the work isn't under a milestone,
|
||||
the task alone is enough.
|
||||
- **What now works** — outcomes the operator would notice: "You can now…",
|
||||
"X no longer…". The files and steps behind them belong in the task's log.
|
||||
- **How / why** — only the decisions worth knowing, plus **how it was
|
||||
verified**. If something could not be verified, say what and why here rather
|
||||
than letting it read as passed.
|
||||
- **Needs you** — an action, an approval, a decision, or "nothing". If it's an
|
||||
action, give the reason with it.
|
||||
- **Next** — from `placement.next`, or say the milestone is finished. If you
|
||||
found something you didn't fix, the offer to fix it goes here.
|
||||
|
||||
## Before sending
|
||||
|
||||
Read the reply as the operator will: someone who wasn't there, reading
|
||||
quickly. Can they tell **what was done, whether anything needs them, and what
|
||||
happens next** without asking a follow-up? If not, the sections are what's
|
||||
missing — not more detail.
|
||||
@@ -47,8 +47,8 @@ through recall/auto-inject; this skill is the active reflex around that.
|
||||
- **A `Shape ledger at …` line is the ledger speaking, not the record.** It
|
||||
names a duplicate family ("identical body in N other files, no canon"), a
|
||||
repeated name ("defined in N other files") or a canon elsewhere for a name
|
||||
you just wrote — for edits made through Bash
|
||||
(sed, heredocs, scripts) as much as through Write/Edit. Derive the family or
|
||||
you just wrote — for edits made through a shell
|
||||
(sed, heredocs, scripts) as much as through your editor tools. Derive the family or
|
||||
reuse the canon *now*; a family that is convention rather than copies is
|
||||
dismissed with `classify_shapes(..., status="exempt",
|
||||
reason_code="convention-plumbing")`, never ignored.
|
||||
|
||||
@@ -47,12 +47,12 @@ need a hand judgment:
|
||||
- **The sync** stamps a snippet's own reference location `canonical`
|
||||
(`classified_by: mechanical`).
|
||||
- **The write path** stamps instances as you work: when you `get_snippet` a
|
||||
canon and then Write/Edit code that references or resembles it, the
|
||||
canon and then write code that references or resembles it, the
|
||||
definitions being written land as `instance` rows (`classified_by: hook`,
|
||||
the evidence in `reason`), and the prior-art hook tells you what landed
|
||||
the evidence in `reason`), and the prior-art hint tells you what landed
|
||||
("Shape accounting: recorded at … → instance of #N"). Offered-but-unopened
|
||||
snippets stamp nothing — so *pull the canon you are instantiating*; that
|
||||
pull is what turns your reuse into accounting. A hook row is evidence, not
|
||||
pull is what turns your reuse into accounting. A `hook` row is evidence, not
|
||||
judgment: it never overrides a classification you made, and a
|
||||
`classify_shapes` call overrides it.
|
||||
|
||||
@@ -91,8 +91,8 @@ The catalogue exists so the codebase is DRY **from inception**, not as DRY as
|
||||
the last sweep left it. Three surfaces say so without anyone running an audit
|
||||
(milestone 299):
|
||||
|
||||
- **At the write** — the prior-art hint (the Write/Edit hook, and since
|
||||
0.1.39 the after-write hook on Bash, so sed/heredoc/script edits count too)
|
||||
- **At the write** — the prior-art hint (delivered beside a write where your
|
||||
client supports it; shell edits count as much as editor writes)
|
||||
carries a `Shape ledger at <path>` line when a name just written is a known
|
||||
**duplicate family** ("identical body in N other files, no canon"), a
|
||||
**repeated name** ("defined in N other files, no canon") or a
|
||||
@@ -144,7 +144,7 @@ Three questions the ledger answers mechanically (#2793):
|
||||
proposer did not match to that canon. `diverges_from` names the canon.
|
||||
Judge it: `instance` if it should be built from the canon (and rebuild
|
||||
it), `variant` with the why if the departure is deliberate. The write-path
|
||||
hook asks the same question in-band the moment such a shape is written.
|
||||
hint asks the same question in-band the moment such a shape is written.
|
||||
- **History** — `shape_history(project_id, path, symbol?)`: the current rows
|
||||
plus every `classified` / `vanished` / `reappeared` / `drifted` event with
|
||||
its commit — "instance of #N from <date>, re-judged variant of #M because
|
||||
|
||||
@@ -27,33 +27,29 @@ If the working repo maps to a Scribe project (you're in a known repo, or
|
||||
project plus the rules bound to the areas it works in, open tasks, and recent
|
||||
notes in one shot.
|
||||
|
||||
Do this actively. A SessionStart hook *may* also inject a rule index, but treat
|
||||
that as a bonus, not a precondition: it can be absent (e.g. when the instance is
|
||||
unreachable, or the token didn't reach the hook), so the reliable path is this
|
||||
explicit pull. Rules loaded this way are **binding** for the session.
|
||||
Do this actively. Nothing is handed to a session up front to stand in for it —
|
||||
rules arrive by retrieval, when your work or the operator's message matches
|
||||
one — so asking and entering the project are the reliable path.
|
||||
|
||||
## Scribe holds these functions — don't keep a second copy
|
||||
## Scribe holds these functions — keep one copy
|
||||
|
||||
This plugin makes Scribe the home for the operator's **rules, recall, and
|
||||
planning** — the jobs Claude's native auto-memory would otherwise do. When the
|
||||
plugin is present, route those jobs to Scribe and **do not also write them to
|
||||
native memory**: codify rules with `create_rule` / `create_project_rule`,
|
||||
capture durable knowledge as Scribe notes, and keep plans in Scribe milestones
|
||||
(via `start_planning`) — not in `MEMORY.md` or `CLAUDE.md`. One copy, in Scribe; let any existing local
|
||||
memory shrink as Scribe takes over. Don't maintain both stores in parallel.
|
||||
Scribe is the home for the operator's **rules, recall, and planning** — the
|
||||
jobs your client's own local memory files would otherwise do. Route those jobs
|
||||
to Scribe instead of also writing them locally: codify rules with
|
||||
`create_rule` / `create_project_rule`, capture durable knowledge as Scribe
|
||||
notes, and keep plans in Scribe milestones (via `start_planning`). One copy, in
|
||||
Scribe; let any existing local memory shrink as Scribe takes over.
|
||||
|
||||
Two constraints on *how* that's achieved:
|
||||
|
||||
- **Steer behavior; never flip a native switch.** The plugin must work with
|
||||
native auto-memory at its default (ON). Never tell the operator to set
|
||||
`autoMemoryEnabled:false` or otherwise disable a built-in function to make
|
||||
Scribe "win" — a setting the operator may not know was changed (and wouldn't
|
||||
know to restore) is exactly the hidden breakage to avoid. You replace memory's
|
||||
functions by *doing the work in Scribe*, not by turning memory off.
|
||||
- **A Scribe-shaped hole is acceptable.** If the plugin is later removed, the
|
||||
operator recovers context over time — that's fine. You do **not** need to keep
|
||||
native memory as a self-sufficient fallback. The only thing to avoid is
|
||||
breakage caused by a settings change the operator didn't make knowingly.
|
||||
- **Steer behaviour; leave the client's own settings as they are.** Scribe works
|
||||
alongside a client's built-in memory at its defaults. You replace memory's
|
||||
functions by *doing the work in Scribe*, so there is no reason to ask the
|
||||
operator to switch a built-in feature off — a setting they didn't knowingly
|
||||
change is breakage they won't know to restore.
|
||||
- **A Scribe-shaped hole is acceptable.** If Scribe is later removed, the
|
||||
operator recovers context over time — that's fine. Local memory doesn't need
|
||||
to be kept as a self-sufficient fallback.
|
||||
|
||||
## The reflex
|
||||
|
||||
@@ -88,17 +84,23 @@ Two constraints on *how* that's achieved:
|
||||
asks. If what you learned is that something MUST be done a certain way,
|
||||
that is a rule to propose, not a preference to harden in place.
|
||||
|
||||
Rules come in two tiers. **Always-on** rules are delivered — they arrive
|
||||
whether or not you ask. **Conditional** rules are RETRIEVED, and one binds
|
||||
just as hard for never having been handed to you. So before a consequential
|
||||
act, `search(content_type="rule")` on what you are about to do. An empty
|
||||
loaded set is not evidence that no rule applies; it is only evidence that
|
||||
none was pushed, and those are different claims.
|
||||
Every rule is RETRIEVED: one reaches you when a command, the code you are
|
||||
writing or the operator's message resembles what it is about, and a rule
|
||||
binds just as hard for never having been handed to you. So before a
|
||||
consequential act, `search(content_type="rule")` on what you are about to
|
||||
do. An empty session is not evidence that no rule applies; it is only
|
||||
evidence that nothing has matched yet, and those are different claims.
|
||||
|
||||
The tier split exists because delivery does not scale: every resident rule
|
||||
costs tokens in every session forever, so a rulebook that grows past a few
|
||||
dozen either stops growing or stops fitting. Retrieval is what lets the
|
||||
rulebook keep growing — but retrieval only fires if something asks.
|
||||
Retrieval is what lets a rulebook keep growing — a rule costs nothing in a
|
||||
session it has nothing to do with — but retrieval only fires if something
|
||||
asks.
|
||||
|
||||
**A retrieved rule outranks a default habit.** Before a hard-to-reverse or
|
||||
outward-facing act — changing shared state, publishing, deleting, sending
|
||||
something outside the session — the operator's rules decide what to do, not
|
||||
the generic conventions your client or your training suggest. When a rule
|
||||
and a habit disagree, the rule wins; when no rule speaks to it, ask rather
|
||||
than assume.
|
||||
|
||||
**Ask hardest where you feel most certain.** Rules about which TOOL to reach
|
||||
for — use the forge's MCP client rather than curling its API, don't stand up
|
||||
@@ -120,7 +122,9 @@ Two constraints on *how* that's achieved:
|
||||
plans/specs to local `.md` files. See the **writing-plans** skill.
|
||||
|
||||
5. **Keep state honest.** Set a task `in_progress` when you start it, `done` the
|
||||
moment it's complete; log progress as you go.
|
||||
moment it's complete; log progress as you go. Always log when you
|
||||
**complete** a task and when you **hit or discover a problem**, so a change
|
||||
of direction is on the record and not only the successes.
|
||||
|
||||
6. **Fixes are issues, not work-logs.** When you fix a problem — even one solved
|
||||
in passing — record it as its own issue (`create_task(kind="issue")`) with
|
||||
@@ -143,6 +147,11 @@ Two constraints on *how* that's achieved:
|
||||
not restraint. Only a record genuinely about no particular area goes
|
||||
untagged.
|
||||
|
||||
Every read and write of a project record shows its `systems`. An untagged
|
||||
one carries a `systems_hint` question instead — on creates, updates and
|
||||
work-logs alike. Treat it as the tagging question asked at the moment of
|
||||
work: tag the record, create the missing System, or deliberately leave it.
|
||||
|
||||
8. **Name the record, never just its number.** Whenever you refer to a Scribe
|
||||
record — in a message to the operator, a commit message, a task body, a
|
||||
work-log — write the id *and* its title: `#3244 "the staleness signal"`,
|
||||
@@ -161,6 +170,13 @@ Two constraints on *how* that's achieved:
|
||||
context than the operator has now: commit messages, task bodies, and any
|
||||
record that cites another.
|
||||
|
||||
**An id exists only once a create call returns it.** Never write the id you
|
||||
expect a record to get — every session and user draws from one sequence,
|
||||
so the number goes to whoever creates next, and Scribe refuses a body that
|
||||
cites an id not yet assigned. Records that must cite each other are created
|
||||
together — `create_records` or `start_planning(steps=...)` — with
|
||||
`{{ref:N}}` where the Nth record's id belongs.
|
||||
|
||||
9. **State updates in place; chronicles don't.** A dev-log records what
|
||||
*happened* — write it once, never rewrite it. A durable finding (how a
|
||||
subsystem works, a measured number) lives in that System's **reference
|
||||
@@ -213,6 +229,15 @@ Two constraints on *how* that's achieved:
|
||||
either: `verify_snippet` compares the recorded location and code against the
|
||||
repo, which is richer and already wired to drift detection.
|
||||
|
||||
11. **Report back in a shape the operator can read.** They were not there while
|
||||
you worked, so organise the reply around **where the work stands**, not the
|
||||
order you did things in: which task or milestone it belongs to, what now
|
||||
works, what needs them, and what comes next. Take the placement from the
|
||||
`placement` block that `create_task` / `update_task` return — the milestone,
|
||||
step N of M, the next open step — rather than from memory. The
|
||||
`reporting-back` skill holds the shape for each kind of reply: completions,
|
||||
findings, decisions, handoffs, "where are we".
|
||||
|
||||
## Stay inside the active project's scope
|
||||
|
||||
Once a project is in scope — you called `enter_project`, or the working repo is
|
||||
@@ -293,6 +318,6 @@ nothing will tell you it drifted.
|
||||
|
||||
## Other Scribe process-skills
|
||||
|
||||
This plugin also ships focused process-skills — writing-plans, systematic
|
||||
debugging, verification, and brainstorming. Reach for the matching one when its
|
||||
Scribe also ships focused process-skills — writing-plans, reporting-back,
|
||||
systematic debugging, verification, and brainstorming. Reach for the matching one when its
|
||||
situation arises, the same way you reach for this skill.
|
||||
|
||||
@@ -34,13 +34,26 @@ already done. It creates a **milestone** (the plan container) seeded with a
|
||||
design template and returns the milestone id plus the project's applicable
|
||||
rules. The plan lives in that milestone:
|
||||
|
||||
- The **design/intent** goes in the milestone `body` — edit it with
|
||||
- The **design/intent** goes in the milestone `body` — pass it as
|
||||
`start_planning(..., body=...)`, or edit it later with
|
||||
`update_milestone(milestone_id, body=...)`.
|
||||
- Each **step** is its own task under the milestone — create it with
|
||||
`create_task(milestone_id=<that milestone>)` and track it with status +
|
||||
`add_task_log`. Steps are first-class tasks, **not** checkboxes in the body.
|
||||
- Each **step** is its own task under the milestone. When you know the steps,
|
||||
pass them in the same call — `start_planning(..., steps=[{title, body}, …])`
|
||||
— and the milestone and every step are created together. Add steps later
|
||||
with `create_records(milestone_id=<that milestone>, records=[…])`, or
|
||||
`create_task` for one. Track each with status + `add_task_log`. Steps are
|
||||
first-class tasks, **not** checkboxes in the body.
|
||||
- Read the whole plan back with `get_milestone` (body + its step-tasks).
|
||||
|
||||
**Never write an id you have not been given.** A plan body that says "see
|
||||
#4012" before #4012 exists is a guess, and other sessions and users are
|
||||
creating records from the same sequence at the same time — the number goes to
|
||||
whoever creates next, and the plan then points at their record. Scribe refuses
|
||||
a body citing an id that has not been assigned. Where the plan and its steps
|
||||
need to cite each other, write a placeholder instead: `{{ref:N}}` is the Nth
|
||||
step in the list, `{{ref:milestone}}` is the milestone. They are replaced with
|
||||
the real id and title as the records are created.
|
||||
|
||||
**Do not** write plans or specs to local `.md` files — the milestone is the
|
||||
record, not a file on disk. (The old `kind=plan` task is retired; `start_planning`
|
||||
no longer creates one.)
|
||||
|
||||
+48
-92
@@ -7,104 +7,60 @@ from mcp.server.fastmcp import FastMCP
|
||||
from mcp.server.transport_security import TransportSecuritySettings
|
||||
from quart import Quart
|
||||
|
||||
## The delivery budget — read before editing this block
|
||||
## What this block is — read before editing it
|
||||
#
|
||||
# Claude Code injects only the FIRST ~2,048 CHARACTERS of an MCP server's
|
||||
# instructions into the system prompt; the rest is silently cut mid-word
|
||||
# (#2562 — the cut was observed live at exactly offset 2,048, and ~90% of the
|
||||
# previous 20k-char version of this block never reached any session). So this
|
||||
# block is deliberately a MAP, not a manual, and a test pins it under the
|
||||
# fold (test_instruction_surfaces_agree.py::test_instructions_fit_the_fold).
|
||||
# ONE OWNER PER PIECE OF GUIDANCE (decision #4027, milestone 410). Scribe's
|
||||
# guidance to agents lives where it can be delivered, and each topic is stated
|
||||
# in full exactly once:
|
||||
# - Tool docstrings: each tool's contract, delivered with its schema.
|
||||
# - In-band tool responses: behaviour prose cannot be trusted to trigger —
|
||||
# the duplicate gate, the guessed-id refusal, `systems_hint`,
|
||||
# `placement` and `report_back` — at the moment it applies.
|
||||
# - The bundled skills (Agent Skills, client-neutral): every reflex in full.
|
||||
# `using-scribe` owns the working reflexes; the process skills own arcs.
|
||||
# - Client adapters (the Claude Code plugin today): timing and that
|
||||
# client's own conventions, never a copy of the above.
|
||||
# tests/test_guidance_ownership.py holds the topic registry that enforces it.
|
||||
#
|
||||
# Where the detail lives instead — each surface has one job:
|
||||
# - Tool docstrings: the per-tool HOW. Delivered with the tool schema, at
|
||||
# reach-for time when the client defers tools. Guidance about one tool
|
||||
# belongs there, not here.
|
||||
# - Plugin static context (plugin/hooks/scribe_static_context.md): the
|
||||
# session-level reflexes (recall-first, record-as-you-go, tag-to-Systems,
|
||||
# compaction). Always delivered in full; needs no key and no network.
|
||||
# - Plugin skills: process arcs (planning, debugging, verification…).
|
||||
# Their listing line is the always-visible trigger; the body loads on
|
||||
# match. Stored Processes become skills via /scribe:sync.
|
||||
# - The server itself: behaviors prose can't be trusted to fire (the
|
||||
# duplicate gate, the untagged-record systems_hint) act in-band in tool
|
||||
# responses, at the moment they apply.
|
||||
# Grow one of those, not this block.
|
||||
# BUDGET: ~1980 of the client's ~2048-char cap (#2562). Everything below is
|
||||
# competing for the last ~68 characters, so an addition here is a trade, never
|
||||
# an append.
|
||||
# THIS BLOCK IS THE SERVER'S ORIENTATION, WRITTEN AS AN INDEX. It reaches
|
||||
# every MCP client, so it names no client, and it points at where each reflex
|
||||
# is stated rather than restating it. A new topic gets a line here only if it
|
||||
# is a session-start reflex; its full statement goes to its owner.
|
||||
#
|
||||
# Milestone 317 (a note's own verify_with / expires_when, and the sweep over
|
||||
# them) was DECLINED a line, deliberately, by the operator — not overlooked.
|
||||
# The reasoning, so it is not re-litigated blind: this is a map, and its own
|
||||
# closing line says each tool's description carries the full contract. The
|
||||
# sweep is a curation act, not a session-start reflex like enter_project.
|
||||
# Spending the last of the budget on it would leave the
|
||||
# map unable to grow for something more central later.
|
||||
#
|
||||
# The accepted cost: an agent that never opens create_note's docstring never
|
||||
# learns the field exists. Guidance lives in the create_note / update_note
|
||||
# docstrings and the using-scribe skill instead.
|
||||
#
|
||||
# Milestone 333 step 3 (2026-09-04) bought the HOW bullet's second clause —
|
||||
# search(content_type="rule") before a consequential act — by TRADING OUT
|
||||
# "Processes are saved procedures (follow verbatim)" and "Deletes are
|
||||
# trash-recoverable". Recorded so the trade is not silently reversed:
|
||||
# - Both were already in test_instruction_surfaces_agree's DISPLACED_TOPICS
|
||||
# and already stated on a delivered surface, so nothing fell off: the
|
||||
# process reflex is in every scribe-proc-* skill listing (each says the
|
||||
# process governs and is followed verbatim), and trash recovery is in the
|
||||
# delete_*/list_trash/restore docstrings, which is where per-tool guidance
|
||||
# belongs by this block's own doctrine.
|
||||
# - What it bought is not per-tool guidance and has nowhere else to live at
|
||||
# session-start altitude. Rules were retrievable only by RESIDENCY: the
|
||||
# always-on preload put them in front of the agent, and nothing told a
|
||||
# session to go looking for one it had not been handed. That preload is
|
||||
# gone (milestone 394), which makes this line LOAD-BEARING rather than
|
||||
# supplementary: retrieval is now the only delivery, and retrieval fires
|
||||
# only if something asks. A session that waits to be handed a rule is
|
||||
# handed nothing. A tool-choice reflex asks least of all (#3476, #161).
|
||||
# - It also has to carry what absence MEANS. "No rule arrived" is now the
|
||||
# ordinary state rather than the exceptional one, and reading it as
|
||||
# "there is no rule" is the #3720 defect at session scale. Rule 119 makes
|
||||
# these surfaces the
|
||||
# specification, so the same sentence lands on all three session-start
|
||||
# surfaces, and test_instruction_surfaces_agree pins it.
|
||||
# BUDGET: at most 2,000 characters (test_instructions_fit_the_fold). Claude
|
||||
# Code injects only the first ~2,048 characters of a server's instructions and
|
||||
# cuts the rest mid-word (#2562, observed live), and other clients differ, so
|
||||
# nothing load-bearing may sit past the fold. The history of what was traded
|
||||
# for space before the ownership split (milestones 317, 333, 409) is in
|
||||
# decision #4027 and the notes it supersedes.
|
||||
_INSTRUCTIONS = """
|
||||
Scribe is the operator's self-hosted second brain and system of record — and
|
||||
yours: recall from it before acting, record as you go. Keep no parallel copy
|
||||
in local files (CLAUDE.md, auto-memory); Scribe holds the single copy.
|
||||
Scribe is the operator's system of record for their work, and yours: recall
|
||||
from it before acting, record in it as you go, and keep one copy here rather
|
||||
than in local memory files.
|
||||
|
||||
Hierarchy: Project -> Milestone -> Task/Note. The map, by purpose:
|
||||
- ORIENT: enter_project(id) at session start — rules, open tasks, recent
|
||||
notes, Systems, design system. `inception`: ask what the project
|
||||
inherits, then decide_project_inception.
|
||||
- DO: create_task. Fixed a problem? kind="issue" (symptom -> root cause ->
|
||||
fix), never a work-log line on an unrelated task. Log with add_task_log;
|
||||
keep status honest — in_progress on start, done on finish.
|
||||
- PLAN work with an arc: start_planning. The plan IS a milestone; each step is
|
||||
a child task, not a checkbox. No local plan .md files.
|
||||
- CAPTURE: create_note. RECALL: search first — prior art exists; pass the
|
||||
active project_id to stay in scope.
|
||||
- WHERE work happens: Systems. Tag records with system_ids as you write;
|
||||
create_system when the area is unmodelled.
|
||||
- HOW: rules bind; preferences guide. Nothing preloads — a rule arrives
|
||||
when your work matches it. Before a consequential act,
|
||||
search(content_type="rule"); silence means nothing matched, not none.
|
||||
- UI: the project's design system is binding — resolve_design_system /
|
||||
get_design_system_stylesheet before hand-writing a value.
|
||||
- REUSE: search snippets before writing a helper; record what you build with
|
||||
create_snippet; classify shapes against canon (classify_shapes) — a
|
||||
consumer map is rows, never prose.
|
||||
Every reflex below is stated in full in the using-scribe skill (if your
|
||||
client reads Agent Skills) and in each tool's description. The index:
|
||||
- ORIENT: enter_project(id) loads the project, open work, Systems and design
|
||||
system. An `inception` key: ask what it inherits, then
|
||||
decide_project_inception.
|
||||
- RULES: nothing preloads; a rule arrives when your work matches it. Before a
|
||||
consequential act, search(content_type="rule"). Silence means
|
||||
nothing matched, not none. Rules bind; preferences guide.
|
||||
- RECALL: search before acting, scoped with the active project_id.
|
||||
- RECORD: create_task; a fix is kind="issue". add_task_log as you go; status
|
||||
in_progress on start, done on finish. Tag system_ids as you write.
|
||||
- PLAN work with an arc: start_planning(steps=[...]). The plan is a milestone
|
||||
and each step a task.
|
||||
- IDS exist only once a create returns them. Records that cite each other go
|
||||
through create_records, writing {{ref:N}} for the Nth record.
|
||||
- REUSE: search snippets before building; create_snippet what you build.
|
||||
- UI: the project's design system binds; resolve_design_system before
|
||||
hand-writing a value.
|
||||
- REPORT back from the `placement` a task write returns: where the work sits,
|
||||
what changed, what needs the operator, what comes next.
|
||||
|
||||
A task is a note with status (*_note vs *_task tools).
|
||||
Creates are duplicate-gated: a near-match BLOCKS and returns the existing
|
||||
id — update it, don't force. shared:true records are another user's — a
|
||||
suggestion, not the operator's settled practice.
|
||||
|
||||
This is only a map — the client injects ~2k chars and cuts the rest. Each
|
||||
tool's description carries its full contract: read it when you load the
|
||||
tool, and trust it over habit.
|
||||
Creates are duplicate-gated: a near-match returns the existing id to update.
|
||||
shared:true records are another user's suggestion, not settled practice.
|
||||
"""
|
||||
|
||||
|
||||
|
||||
@@ -16,6 +16,7 @@ from scribe.services import milestones as milestones_svc
|
||||
from scribe.services import notes as notes_svc
|
||||
from scribe.services import rulebooks as rulebooks_svc
|
||||
from scribe.services import trash as trash_svc
|
||||
from scribe.services.record_refs import refuse_guessed_ids
|
||||
|
||||
|
||||
async def list_milestones(project_id: int) -> dict:
|
||||
@@ -83,6 +84,7 @@ async def create_milestone(
|
||||
status: active (default) or done.
|
||||
"""
|
||||
uid = current_user_id()
|
||||
await refuse_guessed_ids(title, description, body)
|
||||
milestone = await milestones_svc.create_milestone(
|
||||
uid,
|
||||
project_id=project_id,
|
||||
@@ -127,6 +129,7 @@ async def update_milestone(
|
||||
fields["status"] = status
|
||||
if order_index >= 0:
|
||||
fields["order_index"] = order_index
|
||||
await refuse_guessed_ids(title, description, body)
|
||||
milestone = await milestones_svc.update_milestone(uid, milestone_id, **fields)
|
||||
if milestone is None:
|
||||
raise ValueError(f"milestone {milestone_id} not found")
|
||||
|
||||
@@ -22,6 +22,7 @@ from scribe.services import supersession as supersession_svc
|
||||
from scribe.services import systems as systems_svc
|
||||
from scribe.services import trash as trash_svc
|
||||
from scribe.services.note_usage import record_pulled
|
||||
from scribe.services.record_refs import refuse_guessed_ids
|
||||
|
||||
|
||||
async def list_notes(
|
||||
@@ -169,6 +170,11 @@ async def create_note(
|
||||
instead (no duplicate bloat / no stale RAG copies). Set true only
|
||||
when you're sure this is a genuinely distinct note.
|
||||
|
||||
AN ID EXISTS ONLY ONCE A CREATE RETURNS IT. A body citing a `#N` that has
|
||||
not been assigned yet is refused — every session and user draws from one
|
||||
sequence, so an expected id is taken by whoever creates next. Records that
|
||||
must cite each other: create_records, with {{ref:N}} placeholders.
|
||||
|
||||
Returns the created note object including its assigned id, OR — when a
|
||||
near-duplicate is found and force is false — {"duplicate": true,
|
||||
"existing_id": ..., "message": ...} and nothing is created. A tagged
|
||||
@@ -177,6 +183,7 @@ async def create_note(
|
||||
create the missing System, or deliberately leave it untagged.
|
||||
"""
|
||||
uid = current_user_id()
|
||||
await refuse_guessed_ids(title, body)
|
||||
if not force:
|
||||
dup = await dedup_svc.find_duplicate_note(
|
||||
uid, title, body, project_id=project_id or None,
|
||||
@@ -269,6 +276,7 @@ async def update_note(
|
||||
fields["verify_with"] = verify_with
|
||||
if expires_when:
|
||||
fields["expires_when"] = expires_when
|
||||
await refuse_guessed_ids(title, body)
|
||||
note = await notes_svc.update_note(
|
||||
uid, note_id, clear=clear or (), **fields
|
||||
)
|
||||
@@ -376,8 +384,8 @@ async def notes_due_for_verification(
|
||||
thing there is. 0 = no age filter.
|
||||
project_id: narrow to one project. 0 = every project. Unlike the rules
|
||||
sweep, this filter is safe: a note belongs to at most one project
|
||||
outright, with none of the subscription and always-on paths that
|
||||
would make a project filter UNDER-report a rule.
|
||||
outright, with none of the subscription paths that would make a
|
||||
project filter UNDER-report a rule.
|
||||
never_only: only notes nobody has ever verified.
|
||||
"""
|
||||
uid = current_user_id()
|
||||
|
||||
@@ -83,7 +83,7 @@ async def enter_project(project_id: int) -> dict:
|
||||
|
||||
`inception` (milestone 297) appears ONLY when the project is yours and
|
||||
nobody has decided what it inherits: it carries the current defaults
|
||||
(which always-on rulebooks bind, design system, Systems), what to ask the
|
||||
(the rulebooks it could subscribe to, design system, Systems), what to ask the
|
||||
operator — once — and the decide_project_inception call that answers it;
|
||||
it repeats on every enter until a decision is recorded.
|
||||
|
||||
@@ -149,7 +149,7 @@ async def enter_project(project_id: int) -> dict:
|
||||
)
|
||||
|
||||
# The inception ask (milestone 297): a project nobody has decided on
|
||||
# inherits its defaults silently — always-on rulebooks, no design system,
|
||||
# inherits nothing, silently — no rulebook subscriptions, no design system,
|
||||
# no Systems. Owner-only (deciding is the owner's), and only until a
|
||||
# decision is recorded; the key is ABSENT otherwise (#2483).
|
||||
inception_ask = None
|
||||
@@ -288,8 +288,8 @@ async def create_project(
|
||||
before calling, ask the operator the four inception questions and pass
|
||||
the answers; a project created without any of them is UNDECIDED and
|
||||
enter_project will ask until decide_project_inception records it.
|
||||
Defaults if nobody decides: every always-on rulebook binds, nothing is
|
||||
subscribed, no design system, no Systems.
|
||||
Defaults if nobody decides: no rulebook subscriptions, no design system,
|
||||
no Systems.
|
||||
|
||||
Args:
|
||||
title: Project name (required).
|
||||
@@ -297,11 +297,10 @@ async def create_project(
|
||||
goal: The desired outcome or definition of done for the project.
|
||||
status: one of active (default), paused, completed, archived.
|
||||
color: Optional hex colour for the project card (e.g. "#6366f1").
|
||||
subscribe_rulebooks: rulebook ids this project opts into. Since
|
||||
milestone 394 subscription is the only way a rulebook binds a
|
||||
project, so there is no automatic tier left to decline. Was
|
||||
NOT inherit ([] = inherit them all). list_rulebooks shows which are
|
||||
subscribe_rulebooks: rulebook ids to subscribe (the non-always-on ones).
|
||||
subscribe_rulebooks: rulebook ids this project opts into.
|
||||
Subscription is the only way a rulebook binds a project, so a
|
||||
rulebook left out simply does not apply. list_rulebooks shows
|
||||
which exist.
|
||||
design_system_id: the design system this project's UI is built from
|
||||
(list_design_systems); -1 = explicitly none; 0 = not stated.
|
||||
seed_systems: true mints the standard starter Systems (CI & Release,
|
||||
@@ -351,9 +350,9 @@ async def decide_project_inception(
|
||||
unsubscribe_project_from_rulebook to undo one), replaces the design
|
||||
system, and never re-seeds Systems a project already has.
|
||||
|
||||
Args: as create_project's inception args. Passing nothing records an
|
||||
inherit-all decision (every always-on rulebook binds, no subscriptions,
|
||||
no design system, no seed) — a valid answer, stated.
|
||||
Args: as create_project's inception args. Passing nothing records a
|
||||
decision to take nothing (no subscriptions, no design system, no seed) —
|
||||
a valid answer, stated.
|
||||
"""
|
||||
uid = current_user_id()
|
||||
choices = _inception_choices(
|
||||
|
||||
@@ -308,13 +308,12 @@ async def create_rule(
|
||||
and let the answer stand; re-raising a declined proposal argues a rule
|
||||
into existence, which is the thing this whole loop exists to prevent.
|
||||
|
||||
A rulebook rule is shared by every project that gets the rulebook: an
|
||||
A subscribed rulebook binds the
|
||||
projects that opt in. So a rulebook rule must read as a general standard —
|
||||
A rulebook rule is shared by every project subscribed to the rulebook, so
|
||||
it must read as a general standard —
|
||||
never pin it to one project's files, paths, or quirks. For a rule that
|
||||
applies to a single project only, use create_project_rule instead (no
|
||||
rulebook+topic ceremony). If it's a standard a CATEGORY of projects shares,
|
||||
put it in a themed subscribed rulebook, not the always-on one.
|
||||
put it in a rulebook for that category and subscribe those projects to it.
|
||||
|
||||
Write it general WITHOUT hedging for the exceptions. A project that needs
|
||||
to strengthen, narrow or replace this rule writes its own and links it
|
||||
@@ -1048,8 +1047,8 @@ async def rules_due_for_verification(
|
||||
never_only: only rules nobody has ever verified.
|
||||
|
||||
NOT filterable by project, deliberately: a project reaches rules through
|
||||
project scope, subscriptions, always-on rulebooks and exclusions, and a
|
||||
filter that missed one of those paths would UNDER-report — which is the
|
||||
project scope and rulebook subscriptions, and a filter that missed one of
|
||||
those paths would UNDER-report — which is the
|
||||
exact failure this whole surface exists to prevent. Read the whole list.
|
||||
"""
|
||||
uid = current_user_id()
|
||||
|
||||
@@ -280,19 +280,20 @@ It is an UPPER BOUND per surface: a pull records the door it came
|
||||
|
||||
`surfaced` VS `ambient` IS THE READING THAT MATTERS HERE. `surfaced` counts
|
||||
rules a ranker chose — today only the write-path arm — and those are claims
|
||||
a pull can settle. `ambient` counts BULK DELIVERIES: the SessionStart
|
||||
preload and the `rules_payload` surfaces
|
||||
(`enter_project`, `get_project`, `get_milestone`, `start_planning`,
|
||||
`get_task`), which hand over the whole applicable set at once with nobody
|
||||
choosing anything. A large `ambient` says the resident set is big and
|
||||
arrives often — never that it is useful, and never that it is read.
|
||||
a pull can settle. `ambient` counts BULK DELIVERIES: the `rules_payload`
|
||||
surfaces (`enter_project`, `get_project`, `get_milestone`,
|
||||
`start_planning`, `get_task`), which hand over the whole applicable set at
|
||||
once with nobody choosing anything — plus, in rows older than milestone
|
||||
394, the SessionStart preload it removed. A large `ambient` says a bulk set
|
||||
is big and arrives often — never that it is useful, and never that it is
|
||||
read.
|
||||
|
||||
`pull_through` therefore divides by `surfaced` alone. Fold the preload in
|
||||
and growing the always-on set would depress the arm's measured precision
|
||||
while trimming it would flatter it, for reasons having nothing to do with
|
||||
the arm. To judge the PRELOAD instead, compare `ambient` against pulls of
|
||||
those same rules over time: a resident set surfaced thousands of times and
|
||||
opened never is the dead-weight signal, one tier up.
|
||||
and growing a bulk set would depress the arm's measured precision while
|
||||
trimming it would flatter it, for reasons having nothing to do with the
|
||||
arm. To judge a BULK surface instead, compare `ambient` against pulls of
|
||||
those same rules over time: a set surfaced thousands of times and opened
|
||||
never is the dead-weight signal, one level up.
|
||||
|
||||
Read it against `sources["write_path_rule"]`. That arm was once believed
|
||||
never to decline — the reading that scoped #3311 — but it was the arm's
|
||||
|
||||
@@ -28,12 +28,15 @@ from scribe.services import notes as notes_svc
|
||||
# avoid the database would otherwise stub the validation too — turning a
|
||||
# guard into a MagicMock that approves anything.
|
||||
from scribe.services.notes import minted_kind
|
||||
from scribe.services import placement as placement_svc
|
||||
from scribe.services import planning as planning_svc
|
||||
from scribe.services import record_batch as batch_svc
|
||||
from scribe.services import rulebooks as rulebooks_svc
|
||||
from scribe.services import systems as systems_svc
|
||||
from scribe.services import task_logs as task_logs_svc
|
||||
from scribe.services import trash as trash_svc
|
||||
from scribe.services.note_usage import record_pulled
|
||||
from scribe.services.record_refs import refuse_guessed_ids
|
||||
|
||||
|
||||
async def list_tasks(
|
||||
@@ -174,9 +177,16 @@ async def create_task(
|
||||
BLOCKED and the existing task's id is returned so you update it
|
||||
instead. Set true only for a genuinely distinct task.
|
||||
|
||||
AN ID EXISTS ONLY ONCE A CREATE RETURNS IT. Never write the id you expect
|
||||
a record to get: every session and user draws from one sequence, so the
|
||||
number is taken by whoever creates next. A body citing a `#N` that has not
|
||||
been assigned yet is refused. Creating several records that cite each
|
||||
other? Use create_records (or start_planning(steps=...)) and write
|
||||
{{ref:N}} — the real ids are filled in as they are created.
|
||||
|
||||
Returns the created task, OR — when a near-duplicate is found and force is
|
||||
false — {"duplicate": true, "existing_id": ..., "message": ...} (nothing
|
||||
created). A tagged record shows its `systems`; created untagged in a
|
||||
created). A task in a project carries `placement` (see update_task). A tagged record shows its `systems`; created untagged in a
|
||||
project, the response carries the `systems_hint` question instead —
|
||||
answer it: tag the record, create the missing System, or deliberately
|
||||
leave it untagged.
|
||||
@@ -189,6 +199,7 @@ async def create_task(
|
||||
"milestone + seeds the design), then create each step as its own "
|
||||
"task with create_task(milestone_id=<that milestone>)."
|
||||
)
|
||||
await refuse_guessed_ids(title, body)
|
||||
if not force:
|
||||
dup = await dedup_svc.find_duplicate_note(
|
||||
uid, title, body, project_id=project_id or None,
|
||||
@@ -213,7 +224,7 @@ async def create_task(
|
||||
await systems_svc.set_record_systems(uid, note.id, system_ids)
|
||||
data = note.to_dict()
|
||||
await systems_tools.attach_systems(uid, uid, data, note.id, project_id or None)
|
||||
return data
|
||||
return await placement_svc.attach_placement(uid, data, note)
|
||||
|
||||
|
||||
async def update_task(
|
||||
@@ -252,6 +263,17 @@ async def update_task(
|
||||
work that becomes an investigation should say so. 'plan' is
|
||||
refused: plans are milestones (start_planning), and the value
|
||||
survives only so historical plan-tasks stay writable.
|
||||
|
||||
The response carries `placement` for a task in a project: its `project`,
|
||||
and for a task in a milestone its `milestone`, `position` ({step, of}),
|
||||
`progress` ({completed, total, pct}) and `next` (the next open step, or
|
||||
null). These are the facts to use when telling the operator where the
|
||||
work sits and what comes next — read them from here rather than
|
||||
reconstructing them, because a remembered milestone title or "next step"
|
||||
reads exactly like a real one when it is wrong.
|
||||
|
||||
Closing a task (done or cancelled) also returns `report_back`: a one-line
|
||||
reminder of what the reply to the operator should cover.
|
||||
"""
|
||||
uid = current_user_id()
|
||||
fields: dict = {}
|
||||
@@ -279,6 +301,7 @@ async def update_task(
|
||||
fields["arose_from_id"] = arose_from_id
|
||||
if kind:
|
||||
fields["task_kind"] = minted_kind(kind)
|
||||
await refuse_guessed_ids(title, body)
|
||||
note = await notes_svc.update_note(uid, task_id, **fields)
|
||||
if note is None:
|
||||
raise ValueError(f"task {task_id} not found")
|
||||
@@ -288,6 +311,9 @@ async def update_task(
|
||||
await systems_tools.attach_systems(
|
||||
uid, getattr(note, "user_id", uid) or uid, data, task_id, note.project_id
|
||||
)
|
||||
await placement_svc.attach_placement(uid, data, note)
|
||||
if status in _CLOSING_STATUSES:
|
||||
data["report_back"] = REPORT_BACK_CUE
|
||||
return data
|
||||
|
||||
|
||||
@@ -324,7 +350,128 @@ async def add_task_log(task_id: int, content: str) -> dict:
|
||||
return data
|
||||
|
||||
|
||||
async def start_planning(project_id: int, title: str) -> dict:
|
||||
# The in-band half of milestone 409 step 3. The reporting-back skill and the
|
||||
# static context carry the full shapes, but both live only in the Claude Code
|
||||
# plugin; a tool response reaches every MCP client, at the moment a piece of
|
||||
# work closes, which is exactly when the report is about to be written. One
|
||||
# line on purpose: a template here would be read as the reply itself.
|
||||
_CLOSING_STATUSES = ("done", "cancelled")
|
||||
REPORT_BACK_CUE = (
|
||||
"Reporting this to the operator? Say where it sits (from `placement`), "
|
||||
"what now works, what needs them, and what comes next."
|
||||
)
|
||||
|
||||
_ITEM_KEYS = {"title", "body", "type", "status", "priority", "kind", "tags", "system_ids"}
|
||||
|
||||
|
||||
def _batch_items(records: list[dict], *, what: str = "record") -> list[batch_svc.BatchItem]:
|
||||
"""Parse the door's plain dicts into BatchItems, refusing unknown keys.
|
||||
|
||||
Strict for the same reason StrictArgsFastMCP is (#2709): a misspelt key
|
||||
silently dropped — `milestone` for a per-record milestone, `desc` for body —
|
||||
creates a record that looks right and is missing what the caller sent.
|
||||
"""
|
||||
items: list[batch_svc.BatchItem] = []
|
||||
for i, raw in enumerate(records or [], start=1):
|
||||
if not isinstance(raw, dict):
|
||||
raise ValueError(f"{what} {i} must be an object with at least a title")
|
||||
unknown = set(raw) - _ITEM_KEYS
|
||||
if unknown:
|
||||
raise ValueError(
|
||||
f"{what} {i} has unknown field(s) {sorted(unknown)}; "
|
||||
f"allowed: {sorted(_ITEM_KEYS)}"
|
||||
)
|
||||
rtype = raw.get("type") or "task"
|
||||
if rtype not in ("task", "note"):
|
||||
raise ValueError(f"{what} {i}: type must be 'task' or 'note', got {rtype!r}")
|
||||
items.append(batch_svc.BatchItem(
|
||||
title=raw.get("title") or "",
|
||||
body=raw.get("body") or "",
|
||||
is_task=rtype == "task",
|
||||
status=raw.get("status") or "todo",
|
||||
priority=raw.get("priority") or None,
|
||||
task_kind=raw.get("kind") or "work",
|
||||
tags=list(raw.get("tags") or []),
|
||||
system_ids=list(raw.get("system_ids") or []),
|
||||
))
|
||||
return items
|
||||
|
||||
|
||||
async def _first_duplicate(uid: int, items: list, project_id: int | None) -> dict | None:
|
||||
"""The duplicate gate over a whole batch — the first hit blocks all of it."""
|
||||
for i, item in enumerate(items, start=1):
|
||||
dup = await dedup_svc.find_duplicate_note(
|
||||
uid, item.title, item.body, project_id=project_id,
|
||||
is_task=item.is_task, note_type="note",
|
||||
)
|
||||
if dup is not None:
|
||||
payload = dedup_svc.duplicate_response(dup, "task" if item.is_task else "note")
|
||||
payload["record"] = i
|
||||
payload["message"] = f"Record {i} of the batch: {payload['message']} Nothing in the batch was created."
|
||||
return payload
|
||||
return None
|
||||
|
||||
|
||||
async def create_records(
|
||||
records: list[dict],
|
||||
project_id: int = 0,
|
||||
milestone_id: int = 0,
|
||||
force: bool = False,
|
||||
) -> dict:
|
||||
"""Create several tasks and/or notes in ONE call, so they can cite each other.
|
||||
|
||||
Reach for this whenever records you are about to create need to reference
|
||||
one another — a set of tasks that name their siblings, a reference note
|
||||
listing the tasks it indexes. Writing the ids you EXPECT them to get is
|
||||
always wrong: every session and user draws from one sequence, so another
|
||||
create takes those numbers, and a body citing an unassigned `#N` is refused.
|
||||
A plan (a milestone plus its steps) is start_planning(steps=...), which
|
||||
uses the same mechanism. One record with nothing to cite is create_task or
|
||||
create_note.
|
||||
|
||||
In any record's body write {{ref:N}} where the Nth record's id belongs
|
||||
(1-based, in the order listed). It is replaced with `#<id> "<title>"`
|
||||
once the ids exist. Everything is created in one transaction: a bad
|
||||
placeholder, an invalid field or a duplicate creates NOTHING.
|
||||
|
||||
Args:
|
||||
records: The records, in order. Each is an object with `title`
|
||||
(required) and optionally `body`, `type` ('task', the default, or
|
||||
'note'), and for tasks `status`, `priority` and `kind`
|
||||
('work' | 'issue' | 'spike'); plus `tags` and `system_ids`.
|
||||
project_id: The project every record belongs to (0 = none, or taken
|
||||
from milestone_id).
|
||||
milestone_id: File every record under this existing milestone (0 = none).
|
||||
force: Bypass the near-duplicate gate for the whole batch. By default
|
||||
the first record that near-duplicates an existing one BLOCKS the
|
||||
batch, and its existing id comes back so you can update it instead.
|
||||
|
||||
Returns {"ids": [...], "records": [...]} in the order given — the ids
|
||||
cite-able from here on — OR a duplicate payload naming which `record`
|
||||
matched, with nothing created. The ids are the real ones and are not
|
||||
necessarily consecutive: other sessions keep creating meanwhile, and
|
||||
nothing here depends on the numbers being adjacent.
|
||||
"""
|
||||
uid = current_user_id()
|
||||
items = _batch_items(records)
|
||||
await refuse_guessed_ids(*[t for item in items for t in (item.title, item.body)])
|
||||
if not force:
|
||||
dup = await _first_duplicate(uid, items, project_id or None)
|
||||
if dup is not None:
|
||||
return dup
|
||||
_ms, notes = await batch_svc.create_batch(
|
||||
uid, items, project_id=project_id or None, milestone_id=milestone_id or None,
|
||||
)
|
||||
return {"ids": [n.id for n in notes], "records": [n.to_dict() for n in notes]}
|
||||
|
||||
|
||||
async def start_planning(
|
||||
project_id: int,
|
||||
title: str,
|
||||
body: str = "",
|
||||
steps: list[dict] | None = None,
|
||||
force: bool = False,
|
||||
) -> dict:
|
||||
"""Begin a plan in Scribe (the preferred home for plans — not a local .md file).
|
||||
|
||||
Reach for this when the work has an ARC — several steps toward one goal,
|
||||
@@ -341,11 +488,17 @@ async def start_planning(project_id: int, title: str) -> dict:
|
||||
context. The milestone is the plan container — the individual steps live as
|
||||
first-class child tasks under it, not as checkboxes in the body.
|
||||
|
||||
Afterwards:
|
||||
- Edit the plan/design with update_milestone(milestone_id, body=...).
|
||||
- Create each step as its own task with create_task(milestone_id=<this id>);
|
||||
track it with status + add_task_log. Do NOT put steps as checkboxes in the
|
||||
milestone body.
|
||||
PASS THE STEPS HERE when you already know them. The milestone and every
|
||||
step are then created in one transaction, and the plan body and the steps
|
||||
can cite each other with placeholders that become real ids: {{ref:N}} is
|
||||
the Nth step (1-based) and {{ref:milestone}} is this milestone. Never write
|
||||
the ids you expect records to get — other sessions and users draw from the
|
||||
same sequence, and a body citing an unassigned `#N` is refused.
|
||||
|
||||
Without steps: edit the design afterwards with update_milestone(
|
||||
milestone_id, body=...), and add steps with create_records(milestone_id=
|
||||
<this id>, ...) — or create_task for a single one. Track each with status
|
||||
+ add_task_log. Do NOT put steps as checkboxes in the milestone body.
|
||||
|
||||
(kind=plan tasks are retired — use this instead. Existing historical
|
||||
plan-tasks remain readable but new planning goes through milestones.)
|
||||
@@ -353,10 +506,29 @@ async def start_planning(project_id: int, title: str) -> dict:
|
||||
Args:
|
||||
project_id: The project this plan is for.
|
||||
title: A short title for the plan/milestone.
|
||||
body: The plan's design (markdown). Omit to seed the Goal/Approach/
|
||||
Verification template.
|
||||
steps: The plan's step-tasks, in order — each an object with `title`
|
||||
(required) and optionally `body`, `status`, `priority`, `kind`
|
||||
('work' | 'issue' | 'spike'), `tags`, `system_ids`.
|
||||
force: Bypass the near-duplicate gate on the steps. By default a step
|
||||
that near-duplicates an existing task blocks the whole plan, and
|
||||
nothing — milestone included — is created.
|
||||
|
||||
Returns the milestone, the project's applicable rules and brief context,
|
||||
plus `steps` (the created tasks, in order) when steps were given — OR a
|
||||
duplicate payload naming the `record` that matched, with nothing created.
|
||||
"""
|
||||
uid = current_user_id()
|
||||
items = _batch_items(steps or [], what="step")
|
||||
await refuse_guessed_ids(body, *[t for item in items for t in (item.title, item.body)])
|
||||
if items and not force:
|
||||
dup = await _first_duplicate(uid, items, project_id or None)
|
||||
if dup is not None:
|
||||
return dup
|
||||
return await planning_svc.start_planning(
|
||||
user_id=uid, project_id=project_id, title=title,
|
||||
body=body or None, steps=items or None,
|
||||
)
|
||||
|
||||
|
||||
@@ -388,6 +560,7 @@ def register(mcp) -> None:
|
||||
list_tasks,
|
||||
get_task,
|
||||
create_task,
|
||||
create_records,
|
||||
update_task,
|
||||
add_task_log,
|
||||
start_planning,
|
||||
|
||||
@@ -6,6 +6,7 @@ from scribe.auth import login_required, get_current_user_id
|
||||
from scribe.models.note import TaskPriority, TaskStatus
|
||||
from scribe.routes.utils import not_found, parse_iso_date, parse_pagination
|
||||
from scribe.services.access import can_write_note
|
||||
from scribe.services import placement as placement_svc
|
||||
from scribe.services import systems as systems_svc
|
||||
from scribe.services.notes import (
|
||||
create_note,
|
||||
@@ -151,6 +152,7 @@ async def create_task_route():
|
||||
await systems_svc.set_record_systems(uid, task.id, data["system_ids"])
|
||||
out = task.to_dict()
|
||||
out["systems"] = [s.to_dict() for s in await systems_svc.list_record_systems(uid, task.id)]
|
||||
await placement_svc.attach_placement(uid, out, task)
|
||||
return jsonify(out), 201
|
||||
|
||||
|
||||
@@ -267,6 +269,7 @@ async def update_task_route(task_id: int):
|
||||
await systems_svc.set_record_systems(uid, task_id, data["system_ids"])
|
||||
out = task.to_dict()
|
||||
out["systems"] = [s.to_dict() for s in await systems_svc.list_record_systems(uid, task_id)]
|
||||
await placement_svc.attach_placement(uid, out, task)
|
||||
return jsonify(out)
|
||||
|
||||
|
||||
|
||||
@@ -155,6 +155,77 @@ async def create_note(
|
||||
verify_with: str | None = None,
|
||||
expires_when: str | None = None,
|
||||
) -> Note:
|
||||
# Auto-populate project_id from milestone when not explicitly provided
|
||||
if milestone_id is not None and project_id is None:
|
||||
from scribe.models.milestone import Milestone
|
||||
async with async_session() as lookup:
|
||||
result = await lookup.execute(
|
||||
select(Milestone).where(Milestone.id == milestone_id, Milestone.user_id == user_id)
|
||||
)
|
||||
ms = result.scalars().first()
|
||||
if ms is not None:
|
||||
project_id = ms.project_id
|
||||
|
||||
note = build_note(
|
||||
user_id,
|
||||
title=title,
|
||||
body=body,
|
||||
description=description,
|
||||
tags=tags,
|
||||
parent_id=parent_id,
|
||||
project_id=project_id,
|
||||
milestone_id=milestone_id,
|
||||
status=status,
|
||||
priority=priority,
|
||||
due_date=due_date,
|
||||
recurrence_rule=recurrence_rule,
|
||||
note_type=note_type,
|
||||
task_kind=task_kind,
|
||||
arose_from_id=arose_from_id,
|
||||
data=data,
|
||||
verify_with=verify_with,
|
||||
expires_when=expires_when,
|
||||
)
|
||||
async with async_session() as session:
|
||||
session.add(note)
|
||||
await session.commit()
|
||||
await session.refresh(note)
|
||||
|
||||
embed_note(note)
|
||||
|
||||
if project_id is not None:
|
||||
await _maybe_reactivate_project(project_id)
|
||||
|
||||
return note
|
||||
|
||||
|
||||
def build_note(
|
||||
user_id: int,
|
||||
title: str = "",
|
||||
body: str = "",
|
||||
description: str | None = None,
|
||||
tags: list[str] | None = None,
|
||||
parent_id: int | None = None,
|
||||
project_id: int | None = None,
|
||||
milestone_id: int | None = None,
|
||||
status: str | None = None,
|
||||
priority: str | None = None,
|
||||
due_date: date | None = None,
|
||||
recurrence_rule: dict | None = None,
|
||||
note_type: str = "note",
|
||||
task_kind: str = "work",
|
||||
arose_from_id: int | None = None,
|
||||
data: dict | None = None,
|
||||
verify_with: str | None = None,
|
||||
expires_when: str | None = None,
|
||||
) -> Note:
|
||||
"""A validated, UNSAVED Note — every check create_note makes, no write.
|
||||
|
||||
Split out of create_note so a caller that must add several records in one
|
||||
transaction (services/record_batch.py) validates them exactly as a single
|
||||
create would, instead of carrying a second copy of the rules that drifts.
|
||||
Raises before anything is built, so nothing illegal reaches a session.
|
||||
"""
|
||||
# Empty means empty (NULLABLE_NOTE_TEXT), then the invariant. Both run
|
||||
# before anything is written, so an illegal shape never reaches the table.
|
||||
verify_with = verify_with or None
|
||||
@@ -176,48 +247,26 @@ async def create_note(
|
||||
except ValueError:
|
||||
raise ValueError(f"Invalid priority: {priority!r}. Must be one of: {[p.value for p in TaskPriority]}")
|
||||
|
||||
# Auto-populate project_id from milestone when not explicitly provided
|
||||
if milestone_id is not None and project_id is None:
|
||||
from scribe.models.milestone import Milestone
|
||||
async with async_session() as lookup:
|
||||
result = await lookup.execute(
|
||||
select(Milestone).where(Milestone.id == milestone_id, Milestone.user_id == user_id)
|
||||
)
|
||||
ms = result.scalars().first()
|
||||
if ms is not None:
|
||||
project_id = ms.project_id
|
||||
|
||||
async with async_session() as session:
|
||||
note = Note(
|
||||
user_id=user_id,
|
||||
title=title,
|
||||
body=body,
|
||||
description=description,
|
||||
tags=_normalize_tags(tags or []),
|
||||
parent_id=parent_id,
|
||||
project_id=project_id,
|
||||
milestone_id=milestone_id,
|
||||
status=status,
|
||||
priority=priority,
|
||||
due_date=due_date,
|
||||
recurrence_rule=recurrence_rule,
|
||||
note_type=note_type,
|
||||
task_kind=task_kind,
|
||||
arose_from_id=arose_from_id,
|
||||
data=data,
|
||||
verify_with=verify_with,
|
||||
expires_when=expires_when,
|
||||
)
|
||||
session.add(note)
|
||||
await session.commit()
|
||||
await session.refresh(note)
|
||||
|
||||
embed_note(note)
|
||||
|
||||
if project_id is not None:
|
||||
await _maybe_reactivate_project(project_id)
|
||||
|
||||
return note
|
||||
return Note(
|
||||
user_id=user_id,
|
||||
title=title,
|
||||
body=body,
|
||||
description=description,
|
||||
tags=_normalize_tags(tags or []),
|
||||
parent_id=parent_id,
|
||||
project_id=project_id,
|
||||
milestone_id=milestone_id,
|
||||
status=status,
|
||||
priority=priority,
|
||||
due_date=due_date,
|
||||
recurrence_rule=recurrence_rule,
|
||||
note_type=note_type,
|
||||
task_kind=task_kind,
|
||||
arose_from_id=arose_from_id,
|
||||
data=data,
|
||||
verify_with=verify_with,
|
||||
expires_when=expires_when,
|
||||
)
|
||||
|
||||
|
||||
async def get_note(user_id: int, note_id: int) -> Note | None:
|
||||
|
||||
@@ -0,0 +1,145 @@
|
||||
"""Where a task sits — its project, its milestone, its step position, what is next.
|
||||
|
||||
WHY THIS EXISTS (milestone 409 step 1)
|
||||
|
||||
An agent reporting finished work to the operator is asked to place it: which
|
||||
milestone, which step of how many, what comes next. Without those facts in
|
||||
hand it reconstructs them from memory, and a reconstruction reads exactly like
|
||||
the real thing while being wrong — the drafting of the feature note that
|
||||
started this milestone invented a milestone title and named a "next" step that
|
||||
was already done. So the facts come back on the write that changes a task,
|
||||
where the report is about to be written, instead of being left to recall.
|
||||
|
||||
THE SHAPE
|
||||
|
||||
{"project": {"id", "title"},
|
||||
"milestone": {"id", "title", "status"},
|
||||
"position": {"step": 3, "of": 6},
|
||||
"progress": {"completed", "total", "pct"},
|
||||
"next": {"id", "title", "status"} | None}
|
||||
|
||||
`milestone`, `position`, `progress` and `next` appear only for a task in a
|
||||
milestone; a task with a project and no milestone gets `project` alone; a task
|
||||
with neither gets no placement at all (None), which the doors omit rather than
|
||||
send empty.
|
||||
|
||||
WHAT "STEP" AND "NEXT" MEAN
|
||||
|
||||
Notes carry no order column, so a milestone's steps are in CREATION order
|
||||
(created_at, then id) — the order a plan's steps are written in, and the order
|
||||
a batch create inserts them. Deliberately NOT get_milestone's listing order
|
||||
(status, then last update): that is a display choice, and it reshuffles every
|
||||
time anything is touched.
|
||||
|
||||
`next` is the first open step (todo or in_progress) AFTER this one; when every
|
||||
later step is closed it falls back to the earliest open step before it, since
|
||||
that is still the milestone's next piece of work. None when nothing else is open.
|
||||
|
||||
ACCESS
|
||||
|
||||
Siblings are read through access.readable_notes_clause, so a collaborator on a
|
||||
shared task is never shown the title of a step they cannot open. Position and
|
||||
progress are computed over that same readable set — one list, so the counts can
|
||||
never disagree with the titles they sit beside. For an owner the readable set
|
||||
is every step, and the numbers equal get_milestone's.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
|
||||
from sqlalchemy import select
|
||||
|
||||
from scribe.models import async_session
|
||||
from scribe.models.milestone import Milestone
|
||||
from scribe.models.note import Note
|
||||
from scribe.models.project import Project
|
||||
from scribe.services import access as access_svc
|
||||
from scribe.services.milestones import _progress_from_counts
|
||||
|
||||
_OPEN = ("todo", "in_progress")
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
def _next_open(steps: list, current_id: int) -> dict | None:
|
||||
"""The next open step after `current_id`, else the earliest open one before it."""
|
||||
index = next((i for i, s in enumerate(steps) if s.id == current_id), -1)
|
||||
later = [s for s in steps[index + 1:] if s.status in _OPEN]
|
||||
earlier = [s for s in steps[:max(index, 0)] if s.status in _OPEN]
|
||||
pick = (later or earlier or [None])[0]
|
||||
if pick is None:
|
||||
return None
|
||||
return {"id": pick.id, "title": pick.title, "status": pick.status}
|
||||
|
||||
|
||||
async def task_placement(user_id: int, task) -> dict | None:
|
||||
"""Placement for `task` as `user_id` may see it, or None when it has none.
|
||||
|
||||
`task` is the Note the caller already holds — it was just written or read —
|
||||
so its own row is not fetched again.
|
||||
"""
|
||||
project_id = getattr(task, "project_id", None)
|
||||
milestone_id = getattr(task, "milestone_id", None)
|
||||
if not project_id and not milestone_id:
|
||||
return None
|
||||
|
||||
async with async_session() as session:
|
||||
milestone = None
|
||||
if milestone_id:
|
||||
milestone = (await session.execute(
|
||||
select(Milestone).where(
|
||||
Milestone.id == milestone_id, Milestone.deleted_at.is_(None),
|
||||
)
|
||||
)).scalars().first()
|
||||
project_id = project_id or (milestone.project_id if milestone else None)
|
||||
|
||||
project = None
|
||||
if project_id and await access_svc.can_read_project(user_id, project_id):
|
||||
project = await session.get(Project, project_id)
|
||||
if project is not None and project.deleted_at is not None:
|
||||
project = None
|
||||
|
||||
steps: list = []
|
||||
if milestone is not None:
|
||||
steps = list((await session.execute(
|
||||
select(Note).where(
|
||||
Note.milestone_id == milestone.id,
|
||||
Note.status.isnot(None),
|
||||
Note.deleted_at.is_(None),
|
||||
access_svc.readable_notes_clause(user_id),
|
||||
).order_by(Note.created_at.asc(), Note.id.asc())
|
||||
)).scalars().all())
|
||||
|
||||
out: dict = {}
|
||||
if project is not None:
|
||||
out["project"] = {"id": project.id, "title": project.title}
|
||||
# A milestone is shown only to someone who can read its project (or who
|
||||
# owns it): the task being readable does not make its plan readable.
|
||||
if milestone is not None and (project is not None or milestone.user_id == user_id):
|
||||
counts: dict[str, int] = {}
|
||||
for step in steps:
|
||||
counts[step.status] = counts.get(step.status, 0) + 1
|
||||
progress = _progress_from_counts(counts)
|
||||
position = next((i for i, s in enumerate(steps, start=1) if s.id == task.id), None)
|
||||
out["milestone"] = {"id": milestone.id, "title": milestone.title, "status": milestone.status}
|
||||
out["position"] = {"step": position, "of": len(steps)}
|
||||
out["progress"] = {k: progress[k] for k in ("completed", "total", "pct")}
|
||||
out["next"] = _next_open(steps, task.id)
|
||||
return out or None
|
||||
|
||||
|
||||
async def attach_placement(user_id: int, data: dict, task) -> dict:
|
||||
"""Add `placement` to a task payload the door is about to return.
|
||||
|
||||
Fail-open, like every in-band decoration: the write it rides on has
|
||||
already happened, and a placement lookup that errors must not turn a
|
||||
successful update into a reported failure. Omitted, never sent empty.
|
||||
"""
|
||||
try:
|
||||
placement = await task_placement(user_id, task)
|
||||
except Exception: # noqa: BLE001 - a decoration never breaks its payload
|
||||
logger.warning("placement lookup failed for task %s", getattr(task, "id", None), exc_info=True)
|
||||
return data
|
||||
if placement:
|
||||
data["placement"] = placement
|
||||
return data
|
||||
@@ -6,13 +6,19 @@ The milestone IS the plan: its `body` holds the design/intent (Goal/Approach/
|
||||
Verification), and the individual steps live as first-class child tasks
|
||||
(milestone_id) rather than checkboxes crammed into one body. The legacy
|
||||
kind=plan task is retired going forward — start_planning never creates one.
|
||||
|
||||
Steps passed up front are created in the SAME transaction as the milestone
|
||||
(services/record_batch.py), so the plan body and the steps can cite each other
|
||||
through `{{ref:N}}` placeholders instead of predicted ids (#4016).
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from scribe.services import milestones as milestones_svc
|
||||
from scribe.services import notes as notes_svc
|
||||
from scribe.services import projects as projects_svc
|
||||
from scribe.services import record_batch as batch_svc
|
||||
from scribe.services import rulebooks as rulebooks_svc
|
||||
from scribe.services.record_refs import placeholder_keys
|
||||
|
||||
# The plan body template — design only. Steps are NOT checkboxes here; each
|
||||
# step becomes its own child task under this milestone (status, work-logs,
|
||||
@@ -25,10 +31,19 @@ PLAN_TEMPLATE = """## Goal
|
||||
"""
|
||||
|
||||
|
||||
async def start_planning(user_id: int, project_id: int, title: str) -> dict:
|
||||
async def start_planning(
|
||||
user_id: int,
|
||||
project_id: int,
|
||||
title: str,
|
||||
body: str | None = None,
|
||||
steps: list[batch_svc.BatchItem] | None = None,
|
||||
) -> dict:
|
||||
"""Create a milestone seeded as a plan container and return it with the
|
||||
project's applicable rules + brief context.
|
||||
|
||||
`body` replaces the seeded template; `steps` creates the step-tasks with
|
||||
the milestone, atomically.
|
||||
|
||||
Returns:
|
||||
{
|
||||
"milestone": <milestone dict>,
|
||||
@@ -37,19 +52,35 @@ async def start_planning(user_id: int, project_id: int, title: str) -> dict:
|
||||
"applicable_rules_truncated": bool,
|
||||
"project_goal": str,
|
||||
"open_task_count": int,
|
||||
"steps": [<task dict>, ...], # only when steps were given
|
||||
}
|
||||
"""
|
||||
project = await projects_svc.get_project(user_id, project_id)
|
||||
if project is None:
|
||||
raise ValueError(f"project {project_id} not found")
|
||||
|
||||
milestone = await milestones_svc.create_milestone(
|
||||
user_id,
|
||||
project_id=project_id,
|
||||
title=title,
|
||||
body=PLAN_TEMPLATE,
|
||||
status="active",
|
||||
)
|
||||
plan_body = body or PLAN_TEMPLATE
|
||||
step_notes = []
|
||||
if steps:
|
||||
milestone, step_notes = await batch_svc.create_batch(
|
||||
user_id, steps, project_id=project_id,
|
||||
milestone=batch_svc.BatchMilestone(title=title, body=plan_body),
|
||||
)
|
||||
else:
|
||||
# A placeholder with no batch behind it has nothing to resolve to, and
|
||||
# storing `{{ref:1}}` verbatim would be a reference to nothing.
|
||||
if placeholder_keys(plan_body):
|
||||
raise ValueError(
|
||||
"the plan body uses {{ref:...}} placeholders but no steps were "
|
||||
"given — pass steps=[...] so they have records to name"
|
||||
)
|
||||
milestone = await milestones_svc.create_milestone(
|
||||
user_id,
|
||||
project_id=project_id,
|
||||
title=title,
|
||||
body=plan_body,
|
||||
status="active",
|
||||
)
|
||||
|
||||
applicable = await rulebooks_svc.get_applicable_rules(
|
||||
project_id=project_id, user_id=user_id,
|
||||
@@ -58,9 +89,12 @@ async def start_planning(user_id: int, project_id: int, title: str) -> dict:
|
||||
user_id, is_task=True, status="todo", project_id=project_id, limit=1,
|
||||
)
|
||||
|
||||
return {
|
||||
result = {
|
||||
"milestone": milestone.to_dict(),
|
||||
**rulebooks_svc.rules_payload(applicable, user_id=user_id, source="start_planning"),
|
||||
"project_goal": getattr(project, "goal", "") or "",
|
||||
"open_task_count": open_count,
|
||||
}
|
||||
if steps:
|
||||
result["steps"] = [n.to_dict() for n in step_notes]
|
||||
return result
|
||||
|
||||
@@ -2160,30 +2160,18 @@ async def build_session_context(
|
||||
on each write so the server could say whether the resident rules had
|
||||
moved; nothing is resident now, so nothing can have moved, and a rule is
|
||||
re-retrieved at the moment it applies rather than held and aged.
|
||||
`context` is markdown ready to drop into `additionalContext`; it is capped
|
||||
at _MAX_CHARS with an explicit truncation note so the hook can pass it
|
||||
through verbatim.
|
||||
"""
|
||||
lines: list[str] = [
|
||||
"# Scribe — standing session context (auto-injected by the Scribe plugin)",
|
||||
"",
|
||||
"You are working with Scribe, the operator's self-hosted second brain.",
|
||||
"",
|
||||
"## You are not holding the operator's rules",
|
||||
"",
|
||||
"No rule has been loaded into this session, and that is deliberate. "
|
||||
"Rules arrive when something you are about to do makes one relevant — "
|
||||
"a command you are about to run, code you are writing, or what the "
|
||||
"operator just asked for. On most turns none will, and that is the "
|
||||
"surface working rather than failing.",
|
||||
"",
|
||||
"**\"No rule arrived\" means \"nothing matched\" — never \"there is no "
|
||||
"rule.\"** Before a consequential act, one that is hard to reverse or "
|
||||
"outward-facing, `search(content_type=\"rule\")` is how you ask. "
|
||||
"Retrieval runs on its own and is a convenience; asking is what you do "
|
||||
"when it matters and nothing has spoken.",
|
||||
]
|
||||
`context` is markdown an adapter can drop into its session verbatim; it
|
||||
is capped at _MAX_CHARS with an explicit truncation note.
|
||||
|
||||
LIVE STATE ONLY (decision #4027, milestone 410). This used to open with the
|
||||
rules reflex and close with a recall reflex — a fifth copy of guidance the
|
||||
using-scribe skill owns, arriving in every session beside the other four.
|
||||
It now says only what the server alone knows about THIS session: the
|
||||
active project, its open work, its design system, or that the working repo
|
||||
is unbound. How to work with Scribe is the skill's to say, and it says it
|
||||
once.
|
||||
"""
|
||||
lines: list[str] = ["# Scribe — live session state"]
|
||||
|
||||
project_dict: dict | None = None
|
||||
if project_id:
|
||||
@@ -2222,9 +2210,7 @@ async def build_session_context(
|
||||
f"(id {design['id']}){inherits}",
|
||||
f"{design['token_count']} tokens"
|
||||
+ (f" across {groups}" if groups else "")
|
||||
+ ". This project's UI is built from these, not from "
|
||||
"literals — reach for a token before writing a colour, "
|
||||
"size, radius or duration by hand.",
|
||||
+ ".",
|
||||
f"Values: `resolve_design_system({design['id']})` · "
|
||||
f"stylesheet: `get_design_system_stylesheet({design['id']})` "
|
||||
f"· the prose (aesthetic, voice, where the accent may "
|
||||
@@ -2241,18 +2227,12 @@ async def build_session_context(
|
||||
"(call `list_projects` to find the id) and future sessions here will "
|
||||
"auto-load that project's context.",
|
||||
]
|
||||
|
||||
lines += [
|
||||
"",
|
||||
"Reflex: search Scribe (search / list_tasks / list_notes, scoped to the "
|
||||
"active project) before answering or starting work; prefer UPDATING an "
|
||||
"existing note/rule over creating a new one.",
|
||||
]
|
||||
else:
|
||||
lines += ["", "No Scribe project is bound to this working directory."]
|
||||
|
||||
context = "\n".join(line for line in lines if line is not None)
|
||||
if len(context) > _MAX_CHARS:
|
||||
context = context[:_MAX_CHARS].rstrip() + \
|
||||
"\n\n…(truncated — ask with search(content_type=\"rule\"))"
|
||||
context = context[:_MAX_CHARS].rstrip() + "\n\n…(truncated)"
|
||||
|
||||
return {
|
||||
"context": context,
|
||||
|
||||
@@ -0,0 +1,193 @@
|
||||
"""Create several records at once, so they can cite each other without guessing ids.
|
||||
|
||||
Why this exists is in services/record_refs.py (#4016): a session that needs
|
||||
records to reference one another used to create them one by one and predict
|
||||
the ids of the ones not yet made — and any concurrent create, from any session
|
||||
or any user, took those numbers.
|
||||
|
||||
Here the whole batch is ONE transaction: insert every record, flush so the
|
||||
sequence assigns the real ids, rewrite each `{{ref:N}}` with the id and title
|
||||
it names, commit. Other sessions keep creating throughout and cannot interfere
|
||||
— the sequence never hands out the same number twice, so another session's
|
||||
insert just takes a different one. The ids a batch receives are therefore NOT
|
||||
guaranteed consecutive, and nothing here needs them to be: the placeholders
|
||||
are filled with whatever ids came back. Forcing consecutive ids would take a
|
||||
table lock that stalls every user's writes; nobody needs that.
|
||||
|
||||
All or nothing: a bad placeholder, an invalid status, or a failure mid-insert
|
||||
leaves no partial batch behind — which is also why there are no "reserved"
|
||||
stub records to clean up after a session that dies.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass, field
|
||||
|
||||
from sqlalchemy import select
|
||||
|
||||
from scribe.models import async_session
|
||||
from scribe.models.milestone import Milestone
|
||||
from scribe.models.note import Note
|
||||
from scribe.services import access as access_svc
|
||||
from scribe.services import notes as notes_svc
|
||||
from scribe.services import systems as systems_svc
|
||||
from scribe.services.record_refs import placeholder_keys, resolve_placeholders
|
||||
|
||||
# A batch is a plan's steps or a handful of linked records, and every item runs
|
||||
# the near-duplicate gate (an embedding search) before the transaction. 50 is
|
||||
# far past any real plan and keeps one call from becoming a bulk import, which
|
||||
# is a different job with a different door.
|
||||
MAX_BATCH = 50
|
||||
|
||||
|
||||
@dataclass
|
||||
class BatchItem:
|
||||
"""One record in a batch. `is_task` False makes it a plain note."""
|
||||
|
||||
title: str
|
||||
body: str = ""
|
||||
is_task: bool = True
|
||||
status: str = "todo"
|
||||
priority: str | None = None
|
||||
task_kind: str = "work"
|
||||
tags: list[str] = field(default_factory=list)
|
||||
system_ids: list[int] = field(default_factory=list)
|
||||
|
||||
|
||||
@dataclass
|
||||
class BatchMilestone:
|
||||
"""The milestone start_planning creates in the same transaction as its steps."""
|
||||
|
||||
title: str
|
||||
body: str
|
||||
description: str | None = None
|
||||
|
||||
|
||||
def check_batch(items: list[BatchItem], milestone: BatchMilestone | None = None) -> None:
|
||||
"""Refuse a batch that cannot be written in full, before any write.
|
||||
|
||||
Every placeholder must name something the batch will actually create;
|
||||
otherwise the rewrite has nothing to put there, and the alternatives —
|
||||
leaving `{{ref:9}}` in a stored body, or dropping it — both store a
|
||||
reference that points nowhere.
|
||||
"""
|
||||
if not items and milestone is None:
|
||||
raise ValueError("a batch needs at least one record")
|
||||
if len(items) > MAX_BATCH:
|
||||
raise ValueError(f"a batch holds at most {MAX_BATCH} records; got {len(items)}")
|
||||
for i, item in enumerate(items, start=1):
|
||||
if not (item.title or "").strip():
|
||||
raise ValueError(f"record {i} has no title")
|
||||
valid = {str(i) for i in range(1, len(items) + 1)}
|
||||
if milestone is not None:
|
||||
valid.add("milestone")
|
||||
texts = [item.body for item in items] + ([milestone.body] if milestone else [])
|
||||
unknown = placeholder_keys(*texts) - valid
|
||||
if unknown:
|
||||
named = ", ".join("{{ref:%s}}" % k for k in sorted(unknown))
|
||||
raise ValueError(
|
||||
f"{named} names no record in this batch. {{{{ref:N}}}} is the Nth "
|
||||
f"record listed (1 to {len(items)})"
|
||||
+ (", and {{ref:milestone}} is the milestone being created." if milestone else ".")
|
||||
)
|
||||
|
||||
|
||||
async def create_batch(
|
||||
user_id: int,
|
||||
items: list[BatchItem],
|
||||
*,
|
||||
project_id: int | None = None,
|
||||
milestone_id: int | None = None,
|
||||
milestone: BatchMilestone | None = None,
|
||||
) -> tuple[Milestone | None, list[Note]]:
|
||||
"""Create `items` (and optionally a new milestone they belong to) atomically.
|
||||
|
||||
`milestone_id` files the items under an EXISTING milestone the caller owns;
|
||||
`milestone` creates a new one alongside them and files them there. Passing
|
||||
both is refused. Returns (the new milestone or None, the notes in input
|
||||
order). Near-duplicate gating is the door's job, exactly as for a single
|
||||
create.
|
||||
"""
|
||||
if milestone is not None and milestone_id:
|
||||
raise ValueError("pass milestone_id (an existing milestone) or a new milestone, not both")
|
||||
check_batch(items, milestone)
|
||||
|
||||
if milestone_id:
|
||||
async with async_session() as session:
|
||||
existing = (await session.execute(
|
||||
select(Milestone).where(
|
||||
Milestone.id == milestone_id, Milestone.deleted_at.is_(None),
|
||||
)
|
||||
)).scalars().first()
|
||||
if existing is None:
|
||||
raise ValueError(f"milestone {milestone_id} not found")
|
||||
if project_id and project_id != existing.project_id:
|
||||
raise ValueError(
|
||||
f"milestone {milestone_id} belongs to project {existing.project_id}, not {project_id}"
|
||||
)
|
||||
project_id = existing.project_id
|
||||
if milestone is not None and not project_id:
|
||||
raise ValueError("a new milestone needs a project_id")
|
||||
# Share-aware (rule 78): a collaborator with write access to a shared
|
||||
# project may plan in it; a bare owner filter would refuse them.
|
||||
if project_id and not await access_svc.can_write_project(user_id, project_id):
|
||||
raise ValueError(f"project {project_id} not found")
|
||||
|
||||
async with async_session() as session:
|
||||
# Validate every record before adding any: build_note raises on a bad
|
||||
# status/priority, and raising here writes nothing.
|
||||
notes = [
|
||||
notes_svc.build_note(
|
||||
user_id,
|
||||
title=item.title,
|
||||
body=item.body,
|
||||
tags=item.tags,
|
||||
project_id=project_id,
|
||||
milestone_id=milestone_id,
|
||||
status=item.status if item.is_task else None,
|
||||
priority=item.priority if item.is_task else None,
|
||||
task_kind=notes_svc.minted_kind(item.task_kind) if item.is_task else "work",
|
||||
)
|
||||
for item in items
|
||||
]
|
||||
|
||||
new_ms = None
|
||||
if milestone is not None:
|
||||
new_ms = Milestone(
|
||||
user_id=user_id, project_id=project_id, title=milestone.title,
|
||||
description=milestone.description, body=milestone.body, status="active",
|
||||
)
|
||||
session.add(new_ms)
|
||||
await session.flush()
|
||||
for note in notes:
|
||||
note.milestone_id = new_ms.id
|
||||
|
||||
session.add_all(notes)
|
||||
# The flush is where the sequence assigns ids. Nothing is visible to
|
||||
# any other session until the commit below, and if anything between
|
||||
# here and there raises, the context manager rolls it all back.
|
||||
await session.flush()
|
||||
|
||||
refs = {str(i): f'#{n.id} "{n.title}"' for i, n in enumerate(notes, start=1)}
|
||||
if new_ms is not None:
|
||||
refs["milestone"] = f'milestone {new_ms.id} "{new_ms.title}"'
|
||||
new_ms.body = resolve_placeholders(new_ms.body, refs)
|
||||
for note in notes:
|
||||
note.body = resolve_placeholders(note.body, refs)
|
||||
|
||||
await session.commit()
|
||||
for note in notes:
|
||||
await session.refresh(note)
|
||||
if new_ms is not None:
|
||||
await session.refresh(new_ms)
|
||||
|
||||
# After the commit, as a single create does: embedding and System tags are
|
||||
# enrichment on records that now exist, and a failure in either must not
|
||||
# un-create them.
|
||||
for note, item in zip(notes, items):
|
||||
notes_svc.embed_note(note)
|
||||
if item.system_ids:
|
||||
await systems_svc.set_record_systems(user_id, note.id, item.system_ids)
|
||||
if project_id is not None:
|
||||
await notes_svc._maybe_reactivate_project(project_id)
|
||||
|
||||
return new_ms, notes
|
||||
@@ -0,0 +1,131 @@
|
||||
"""Record references written into prose: refusing guessed ids, resolving placeholders.
|
||||
|
||||
THE DEFECT THIS EXISTS FOR (#4016)
|
||||
|
||||
A session about to create several records guesses the ids they will get — the
|
||||
last id it saw, plus one, plus two — and writes those guesses into a plan body
|
||||
or a reference note before the records exist. The database never collides:
|
||||
ids come from one Postgres sequence per table, handed out once each. But that
|
||||
sequence is shared by every session and every user, so any concurrent create
|
||||
takes a guessed number, and the reference silently points at someone else's
|
||||
record. More users make it worse, never better.
|
||||
|
||||
Two halves, and they are not alternatives:
|
||||
|
||||
- `{{ref:N}}` PLACEHOLDERS (services/record_batch.py) remove the reason to
|
||||
guess. A batch create inserts its records, receives their real ids and
|
||||
rewrites the placeholders inside ONE transaction, so there is no moment at
|
||||
which an id has to be predicted.
|
||||
- REFUSING A GUESSED ID is the backstop for a session that guesses anyway. It
|
||||
is a refusal, not a warning, because a warning is an instruction and the
|
||||
operator's point was exactly that instructions are not always followed.
|
||||
|
||||
HOW A GUESS IS RECOGNISED
|
||||
|
||||
A guess is always just ABOVE the highest id that exists — that is the only
|
||||
place "the next id" can be. So a `#N` is refused when it sits in
|
||||
(max_id, max_id + GUESS_WINDOW]. Numbers below the max are real records (or
|
||||
were), and numbers far above it are not Scribe ids at all — a PR, a forge
|
||||
issue — so both pass. What this cannot catch, by construction: a guess that
|
||||
another session has ALREADY used. That reference points at a real record, and
|
||||
nothing in the text distinguishes it from a deliberate one. The placeholders
|
||||
are the fix for that case; this is not.
|
||||
|
||||
WHERE IT IS ENFORCED
|
||||
|
||||
At the MCP door (the create/update tools), not in the services. The defect is
|
||||
an agent predicting ids; a person typing `#123` in the web editor has read that
|
||||
number off a record. Enforcing it in `notes.create_note` would also reach
|
||||
restores and internal writers that legitimately carry numbers from elsewhere.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
|
||||
from sqlalchemy import func, select
|
||||
|
||||
from scribe.models import async_session
|
||||
from scribe.models.note import Note
|
||||
|
||||
# `#123` as a standalone token. The lookbehind keeps out `{` entities,
|
||||
# URL fragments (`page#12`) and doubled `##`; `\b` after the digits keeps out a
|
||||
# hex colour like `#123abc`.
|
||||
REF_RE = re.compile(r"(?<![\w&#/])#(\d+)\b")
|
||||
|
||||
# How far above the highest existing id a number still reads as a guess. It
|
||||
# only has to cover how far ahead a session predicts, which is the size of the
|
||||
# batch it is planning — a plan's steps, a handful of cross-linked notes. 50 is
|
||||
# generous for that and still narrow enough that an unrelated number (a PR, an
|
||||
# issue on a forge) landing inside it is rare. Widening it trades exactly that.
|
||||
GUESS_WINDOW = 50
|
||||
|
||||
# `{{ref:3}}` names the third record of the same batch; `{{ref:milestone}}` names
|
||||
# the milestone a start_planning call creates alongside its steps.
|
||||
PLACEHOLDER_RE = re.compile(r"\{\{\s*ref:\s*(\d+|milestone)\s*\}\}")
|
||||
|
||||
|
||||
def cited_ids(*texts: str | None) -> set[int]:
|
||||
"""Every `#N` written in these texts."""
|
||||
found: set[int] = set()
|
||||
for text in texts:
|
||||
if text:
|
||||
found.update(int(m) for m in REF_RE.findall(text))
|
||||
return found
|
||||
|
||||
|
||||
async def _max_note_id() -> int:
|
||||
# Across ALL users and including trashed rows: the sequence is global, so
|
||||
# "the next id" is global too. Only the number is read, never a row.
|
||||
async with async_session() as session:
|
||||
return (await session.execute(select(func.max(Note.id)))).scalar() or 0
|
||||
|
||||
|
||||
async def guessed_ids(*texts: str | None) -> list[int]:
|
||||
"""The `#N` references in these texts that can only be predictions."""
|
||||
cited = cited_ids(*texts)
|
||||
if not cited:
|
||||
return []
|
||||
ceiling = await _max_note_id()
|
||||
return sorted(n for n in cited if ceiling < n <= ceiling + GUESS_WINDOW)
|
||||
|
||||
|
||||
async def refuse_guessed_ids(*texts: str | None) -> None:
|
||||
"""Raise if any text cites a note/task id that has not been assigned yet.
|
||||
|
||||
Raised before anything is written, so the caller's create or update does
|
||||
not happen. The message names the fix, because the agent reading it is
|
||||
mid-turn and can apply it immediately.
|
||||
"""
|
||||
guessed = await guessed_ids(*texts)
|
||||
if not guessed:
|
||||
return
|
||||
listed = ", ".join(f"#{n}" for n in guessed)
|
||||
raise ValueError(
|
||||
f"{listed} does not exist yet — it reads as a guessed id, and a guess "
|
||||
"is taken by whatever any other session creates next. Nothing was "
|
||||
"written. An id exists only once a create call RETURNS it: create the "
|
||||
"records first and cite the ids they come back with, or create them "
|
||||
"together with create_records / start_planning(steps=...) and write "
|
||||
"{{ref:N}} where the Nth record's id belongs. If this number is not a "
|
||||
"Scribe record (a PR, an issue elsewhere), write it without the '#'."
|
||||
)
|
||||
|
||||
|
||||
def placeholder_keys(*texts: str | None) -> set[str]:
|
||||
"""Every `{{ref:...}}` key in these texts: "1", "2", …, "milestone"."""
|
||||
keys: set[str] = set()
|
||||
for text in texts:
|
||||
if text:
|
||||
keys.update(PLACEHOLDER_RE.findall(text))
|
||||
return keys
|
||||
|
||||
|
||||
def resolve_placeholders(text: str | None, refs: dict[str, str]) -> str | None:
|
||||
"""Rewrite `{{ref:...}}` using `refs` (key -> rendered reference).
|
||||
|
||||
Callers validate the keys first; an unknown key here is a bug in the
|
||||
caller, so it raises rather than leaving a placeholder in a stored record.
|
||||
"""
|
||||
if not text:
|
||||
return text
|
||||
return PLACEHOLDER_RE.sub(lambda m: refs[m.group(1)], text)
|
||||
@@ -69,6 +69,21 @@ async def _dispose_engine():
|
||||
await engine.dispose()
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def _no_embedding():
|
||||
"""Stub the fire-and-forget embedding refresh a note write detaches.
|
||||
|
||||
For integration tests about ids, transactions and access rather than
|
||||
recall: `embed_note` spawns a task that loads the embedding model, which
|
||||
outlives the test's event loop and makes the lane slower for nothing. Opt
|
||||
in alongside `_dispose_engine`.
|
||||
"""
|
||||
from unittest.mock import MagicMock
|
||||
|
||||
with patch("scribe.services.notes.embed_note", MagicMock()):
|
||||
yield
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def _no_supersession():
|
||||
"""Stub the auto-inject menu's "which lines are superseded?" lookup (#278).
|
||||
|
||||
@@ -36,6 +36,7 @@ from tests.helpers import tool_doc as _doc
|
||||
_SURFACES = [
|
||||
("scribe.mcp.tools.notes", "create_note"),
|
||||
("scribe.mcp.tools.tasks", "create_task"),
|
||||
("scribe.mcp.tools.tasks", "create_records"),
|
||||
("scribe.mcp.tools.tasks", "start_planning"),
|
||||
("scribe.mcp.tools.snippets", "create_snippet"),
|
||||
("scribe.mcp.tools.processes", "create_process"),
|
||||
|
||||
@@ -0,0 +1,346 @@
|
||||
"""Every piece of agent guidance has exactly one owner that states it.
|
||||
|
||||
WHY THIS EXISTS (milestone 410, decision #4027)
|
||||
|
||||
Scribe's guidance to agents was written up to five times over: the MCP
|
||||
`_INSTRUCTIONS`, the plugin's static session context, the live session context
|
||||
the server builds, the `using-scribe` skill, and tool docstrings. An earlier
|
||||
design (#2494) made that deliberate — insurance against any one surface failing
|
||||
silently — and the copies drifted apart instead (#2497, #4022).
|
||||
|
||||
Decision #4027 replaced the redundancy with ownership: the server orients with
|
||||
a short index, the skills hold the depth, tool docstrings hold each tool's
|
||||
contract, and each client adapter holds only its own timing and conventions.
|
||||
This module is the registry that decision is enforced from.
|
||||
|
||||
WHAT IT PINS
|
||||
|
||||
1. OWNERS STATE THEIR TOPIC. Every topic's markers and its distinctive
|
||||
statement appear on its owner — so moving or trimming text cannot drop a
|
||||
topic without failing here.
|
||||
2. NOWHERE ELSE STATES IT IN FULL. The statement — a phrase distinctive to
|
||||
the owner's full wording — must not appear on any other session surface
|
||||
(the index, the adapter's static text and commands, the live context, the
|
||||
other skills). A topic legitimately stated in two places at two different
|
||||
moments declares that in `shared_with`, with the reason beside it.
|
||||
3. THE INDEX NAMES THE SESSION-START REFLEXES. `_INSTRUCTIONS` is the one
|
||||
surface every MCP client receives, so each reflex it indexes keeps its
|
||||
`index` markers there — a one-line pointer, not a copy.
|
||||
|
||||
WHAT IT CANNOT SEE
|
||||
|
||||
A copy reworded so it no longer contains the statement phrase passes. The
|
||||
guard catches the ordinary way duplication happens — pasting a paragraph
|
||||
into a second surface — not a determined paraphrase. Tool docstrings are not
|
||||
scanned for copies: a tool's contract may elaborate the reflex that calls it.
|
||||
|
||||
MARKERS AND STATEMENTS ARE PHRASES, NOT WORDS
|
||||
|
||||
Reword a topic deliberately and update its markers/statement in the same
|
||||
commit; each failure message names the topic, the surface and the phrase.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import pathlib
|
||||
import re
|
||||
from typing import NamedTuple
|
||||
|
||||
ROOT = pathlib.Path(__file__).resolve().parents[1]
|
||||
|
||||
|
||||
def _norm(text: str) -> str:
|
||||
# Whitespace-flattened and lowercased: prose is hard-wrapped, so a phrase
|
||||
# can straddle a line break without the guidance having changed.
|
||||
return " ".join(text.split()).lower()
|
||||
|
||||
|
||||
def _live_session_context_source() -> str:
|
||||
"""The source of build_session_context — the prose it emits lives there.
|
||||
|
||||
Read as text rather than called: the function needs a database, and what
|
||||
this module checks is what the product SAYS, which is the literal strings.
|
||||
"""
|
||||
src = (ROOT / "src/scribe/services/plugin_context.py").read_text()
|
||||
start = src.index("async def build_session_context")
|
||||
nxt = re.search(r"\n(?:async def|def) ", src[start + 1:])
|
||||
return src[start:start + 1 + nxt.start()] if nxt else src[start:]
|
||||
|
||||
|
||||
def delivered_surfaces() -> dict[str, str]:
|
||||
"""Every surface a session receives as guidance, by label.
|
||||
|
||||
The one definition of "delivered" for this module:
|
||||
- `instructions` — the MCP server's `_INSTRUCTIONS` (every MCP client)
|
||||
- `docstrings` — the MCP tool modules (tool descriptions, every client)
|
||||
- `skill:<name>` — each bundled Agent Skill
|
||||
- `static` — the Claude Code adapter's static session context
|
||||
- `commands` — the Claude Code adapter's slash commands
|
||||
- `live` — the live session context the server builds
|
||||
"""
|
||||
server = (ROOT / "src/scribe/mcp/server.py").read_text()
|
||||
match = re.search(r'_INSTRUCTIONS = """(.*?)"""', server, re.S)
|
||||
assert match, "server.py no longer defines _INSTRUCTIONS as a triple-quoted literal"
|
||||
surfaces = {
|
||||
"instructions": match.group(1),
|
||||
"docstrings": "".join(p.read_text() for p in sorted((ROOT / "src/scribe/mcp/tools").glob("*.py"))),
|
||||
"static": (ROOT / "plugin/hooks/scribe_static_context.md").read_text(),
|
||||
"commands": "".join(p.read_text() for p in sorted((ROOT / "plugin/commands").glob("*.md"))),
|
||||
"live": _live_session_context_source(),
|
||||
}
|
||||
for skill in sorted((ROOT / "plugin/skills").glob("*/SKILL.md")):
|
||||
surfaces[f"skill:{skill.parent.name}"] = skill.read_text()
|
||||
return {label: _norm(text) for label, text in surfaces.items()}
|
||||
|
||||
|
||||
# Not scanned for copies (see the module docstring).
|
||||
NOT_COPY_SCANNED = frozenset({"docstrings"})
|
||||
|
||||
|
||||
class Topic(NamedTuple):
|
||||
key: str
|
||||
owner: str # a label from delivered_surfaces()
|
||||
markers: tuple[str, ...] # what the topic is about; all on the owner
|
||||
statement: str # distinctive to the owner's full wording
|
||||
index: tuple[str, ...] = () # required in _INSTRUCTIONS, if it indexes this
|
||||
shared_with: tuple[str, ...] = () # other surfaces allowed the statement, with a reason
|
||||
|
||||
|
||||
U = "skill:using-scribe"
|
||||
|
||||
TOPICS: tuple[Topic, ...] = (
|
||||
# ── the working reflexes — owned by the using-scribe skill ──
|
||||
Topic("scribe is the system of record; keep one copy", U, ("one copy",),
|
||||
"let any existing local memory shrink", index=("one copy",)),
|
||||
Topic("orient: enter the project, check repo bindings", U, ("enter_project", "list_repo_bindings"),
|
||||
"returns the project plus the rules bound to the areas it works in",
|
||||
index=("enter_project",)),
|
||||
Topic("rules are retrieved; ask before a consequential act", U, ('content_type="rule"', "nothing matched"),
|
||||
"an empty session is not evidence of an empty rulebook",
|
||||
index=('content_type="rule"', "nothing matched")),
|
||||
Topic("rules bind, preferences guide and are kept current", U, ("preference", "update_preference"),
|
||||
"a preference is the one record you keep current yourself", index=("preferences guide",)),
|
||||
Topic("recall before acting", U, ("recall before acting",), "for related prior work", index=("recall",)),
|
||||
Topic("stay inside the active project's scope", U, ("stay inside the active project", "cross-project"),
|
||||
"stay inside the active project's scope", index=("project_id",)),
|
||||
Topic("record as you go; honest status; fixes are issues", U, ("add_task_log", "in_progress", 'kind="issue"'),
|
||||
"fixes are issues, not work-logs", index=("add_task_log", 'kind="issue"')),
|
||||
Topic("an id exists only once a create returns it", U, ("exists only once a create", "{{ref:"),
|
||||
"exists only once a create call returns it", index=("create_records", "{{ref:n}}")),
|
||||
Topic("tag records to systems as you write", U, ("system_ids", "create_system"),
|
||||
"would someone investigating that subsystem want this record", index=("system_ids",)),
|
||||
Topic("answer the systems_hint at the moment of work", U, ("systems_hint",),
|
||||
"treat it as the tagging question asked at the moment of work"),
|
||||
Topic("a retrieved rule outranks a default habit", U, ("outranks a default habit",),
|
||||
"a retrieved rule outranks a default habit"),
|
||||
Topic("log on completion and on a problem", U, ("hit or discover a problem",), "hit or discover a problem"),
|
||||
Topic("the project's design system binds ui", U, ("resolve_design_system",),
|
||||
"building ui: the project's design system binds", index=("resolve_design_system",)),
|
||||
Topic("name the record, never just its number", U, ("name the record",),
|
||||
"a bare id reads as complete to you and as homework to them"),
|
||||
Topic("project inception is a decision", U, ("decide_project_inception",),
|
||||
"starting a project: decide what it inherits", index=("decide_project_inception",)),
|
||||
Topic("where a new rule goes, and its trigger", U, ("create_project_rule", "when_to_apply"),
|
||||
"whichever home it gets"),
|
||||
Topic("a rule vs the other entities", U, ("standing instruction",), "first ask whether it's a rule at all"),
|
||||
Topic("reference notes update in place; dev-logs don't", U, ("reference note",),
|
||||
"state updates in place; chronicles don't"),
|
||||
# ── process arcs — owned by their skills ──
|
||||
Topic("plan in a milestone, steps created together", "skill:writing-plans", ("start_planning", "{{ref:"),
|
||||
"a milestone earns its place when the work has an arc", index=("start_planning",)),
|
||||
Topic("reuse recorded shapes; record at first build", "skill:reusing-code",
|
||||
("create_snippet", "when_to_use", "first build", "second copy"),
|
||||
"prior art offered beside a write is not noise", index=("create_snippet",)),
|
||||
Topic("report back where the work stands", "skill:reporting-back", ("reporting-back", "placement"),
|
||||
"take the placement from the record", index=("placement",)),
|
||||
# ── per-tool contracts and in-band behaviour — owned by the server ──
|
||||
Topic("closing a task cues the report", "docstrings", ("report_back",), "reporting this to the operator?"),
|
||||
Topic("a note that asserts a fact carries its check", "docstrings", ("verify_with", "expires_when"),
|
||||
"could this note become false without anyone editing it",
|
||||
# Stated at two different moments on purpose: the tool contract is
|
||||
# read when the field is about to be filled, the skill while deciding
|
||||
# what to write at all (test_verification_guidance_survives pins both).
|
||||
shared_with=(U,)),
|
||||
Topic("supersession demotes, never hides", "docstrings", ("supersedes",),
|
||||
"it simply stops competing with this one"),
|
||||
Topic("deletes are recoverable from the trash", "docstrings", ("deleted_batch_id",), "deleted_batch_id"),
|
||||
Topic("creates are duplicate-gated", "docstrings", ("near-duplicate",), "bypass the near-duplicate gate",
|
||||
index=("duplicate-gated",)),
|
||||
Topic("shared records are another user's suggestion", "docstrings", ("shared: true",),
|
||||
"belongs to another user", index=("shared:true",)),
|
||||
Topic("stored processes are followed as written", "docstrings", ("get_process",), "follow the returned body"),
|
||||
Topic("a project is never guessed", "docstrings", ("never guessing a project",), "never guessing a project"),
|
||||
Topic("an unbound repo gets a bind hint", "live", ("bind_repo",), "isn't mapped to a scribe project"),
|
||||
# ── the Claude Code adapter's own conventions ──
|
||||
Topic("compact at clean seams", "static", ("/compact",), "compact at clean seams"),
|
||||
Topic("stored processes sync into local skills", "commands", ("scribe-proc-",), "regenerate the local skill stubs"),
|
||||
Topic("say so when scribe's tools are unavailable", "static", ("tools are unavailable",),
|
||||
"rather than silently falling back to local notes"),
|
||||
)
|
||||
|
||||
|
||||
def owner_gaps(topics, surfaces: dict[str, str]) -> list[str]:
|
||||
"""Topics whose owner no longer carries every marker and the statement."""
|
||||
gaps = []
|
||||
for t in topics:
|
||||
text = surfaces.get(t.owner, "")
|
||||
absent = [p for p in (*t.markers, t.statement) if p.lower() not in text]
|
||||
if absent:
|
||||
elsewhere = [label for label, other in surfaces.items() if t.statement.lower() in other]
|
||||
gaps.append(f"{t.key!r} on {t.owner}: missing {absent}; statement found on {elsewhere or 'nothing'}")
|
||||
return gaps
|
||||
|
||||
|
||||
def copies(topics, surfaces: dict[str, str]) -> list[str]:
|
||||
"""Topics whose full statement also appears on a surface that doesn't own it."""
|
||||
found = []
|
||||
for t in topics:
|
||||
allowed = {t.owner, *t.shared_with}
|
||||
extra = [label for label, text in surfaces.items()
|
||||
if label not in NOT_COPY_SCANNED and label not in allowed and t.statement.lower() in text]
|
||||
if extra:
|
||||
found.append(f"{t.key!r} (owner {t.owner}) is also stated on {extra}")
|
||||
return found
|
||||
|
||||
|
||||
def index_gaps(topics, surfaces: dict[str, str]) -> list[str]:
|
||||
text = surfaces.get("instructions", "")
|
||||
return [f"{t.key!r}: {[m for m in t.index if m.lower() not in text]}"
|
||||
for t in topics if any(m.lower() not in text for m in t.index)]
|
||||
|
||||
|
||||
def test_every_topic_is_stated_by_its_owner():
|
||||
gaps = owner_gaps(TOPICS, delivered_surfaces())
|
||||
assert not gaps, (
|
||||
"these topics are no longer stated in full by their owner:\n "
|
||||
+ "\n ".join(gaps) + "\nEvery guidance topic has one owner (decision "
|
||||
"#4027). If the text moved on purpose, move the topic's owner here in the "
|
||||
"same commit; if it was reworded, update its markers and statement."
|
||||
)
|
||||
|
||||
|
||||
def test_no_topic_is_stated_in_full_off_its_owner():
|
||||
found = copies(TOPICS, delivered_surfaces())
|
||||
assert not found, (
|
||||
"guidance stated in full on a surface that doesn't own it:\n "
|
||||
+ "\n ".join(found) + "\nOne owner per topic (decision #4027): replace "
|
||||
"the copy with a one-line pointer to the owner. If both places genuinely "
|
||||
"need it at different moments, declare it in `shared_with` with the reason."
|
||||
)
|
||||
|
||||
|
||||
def test_the_index_names_each_reflex_it_points_at():
|
||||
gaps = index_gaps(TOPICS, delivered_surfaces())
|
||||
assert not gaps, (
|
||||
f"_INSTRUCTIONS no longer carries the index line for: {gaps}. It is the "
|
||||
f"one guidance surface every MCP client receives; keep a one-line pointer "
|
||||
f"per session-start reflex (see the comment above _INSTRUCTIONS)."
|
||||
)
|
||||
|
||||
|
||||
def test_every_owner_and_sharer_is_a_surface_that_exists():
|
||||
labels = set(delivered_surfaces())
|
||||
unknown = [(t.key, s) for t in TOPICS for s in (t.owner, *t.shared_with) if s not in labels]
|
||||
assert not unknown, f"owners or sharers that name no delivered surface: {unknown}"
|
||||
|
||||
|
||||
def test_topic_keys_are_unique():
|
||||
keys = [t.key for t in TOPICS]
|
||||
assert len(keys) == len(set(keys))
|
||||
|
||||
|
||||
def test_the_ownership_guards_can_fail():
|
||||
"""Rule 167: each guard is shown turning red once."""
|
||||
surfaces = {
|
||||
"instructions": "use the widget",
|
||||
"docstrings": "the widget owner statement lives here",
|
||||
"skill:a": "widget owner statement lives here, and widget tool",
|
||||
"skill:b": "a pasted copy: widget owner statement lives here",
|
||||
"static": "",
|
||||
}
|
||||
topic = Topic("widget", "skill:a", ("widget tool",), "widget owner statement lives here",
|
||||
index=("use the widget", "not in the index"))
|
||||
moved = topic._replace(owner="static")
|
||||
assert owner_gaps((topic,), surfaces) == []
|
||||
assert owner_gaps((moved,), surfaces) and "missing" in owner_gaps((moved,), surfaces)[0]
|
||||
# skill:b pasted it; docstrings are not scanned for copies.
|
||||
assert copies((topic,), surfaces) == ["'widget' (owner skill:a) is also stated on ['skill:b']"]
|
||||
assert copies((topic._replace(shared_with=("skill:b",)),), surfaces) == []
|
||||
assert index_gaps((topic,), surfaces) == ["'widget': ['not in the index']"]
|
||||
|
||||
|
||||
# ── The skills are client-neutral (milestone 410 step 2) ────────────────
|
||||
#
|
||||
# Agent Skills is an open format read by dozens of clients (#4023), so the
|
||||
# skills folder is the part of every client package that is shared verbatim.
|
||||
# Anything only one client understands belongs in that client's adapter —
|
||||
# for Claude Code, the plugin's static context, hooks and commands — or the
|
||||
# skill reads as nonsense everywhere else. Each marker below names a thing
|
||||
# exactly one client has:
|
||||
# - "claude", "~/.claude", ".claude/" the client itself and its paths
|
||||
# - "claude.md", "memory.md", "auto-memory", "automemoryenabled"
|
||||
# Claude Code's local memory files/setting
|
||||
# - "/compact", "/scribe:" Claude Code slash commands
|
||||
# - "sessionstart", "userpromptsubmit", "pretooluse", "posttooluse"
|
||||
# Claude Code hook event names
|
||||
# - "write/edit" Claude Code's editor tool names
|
||||
# - "claude_plugin_root" the Claude Code plugin root variable
|
||||
# `Bash` is matched case-sensitively as a word: it is Claude Code's shell tool
|
||||
# name, while "bash" in prose is just the shell.
|
||||
CLIENT_SPECIFIC = (
|
||||
"claude", "~/.claude", ".claude/", "claude.md", "memory.md", "auto-memory",
|
||||
"automemoryenabled", "/compact", "/scribe:", "sessionstart", "userpromptsubmit",
|
||||
"pretooluse", "posttooluse", "write/edit", "claude_plugin_root",
|
||||
)
|
||||
CLIENT_TOOL_NAME = re.compile(r"\bBash\b")
|
||||
|
||||
|
||||
def client_specific_hits(text: str) -> list[str]:
|
||||
hits = [m for m in CLIENT_SPECIFIC if m in text.lower()]
|
||||
if CLIENT_TOOL_NAME.search(text):
|
||||
hits.append("Bash")
|
||||
return hits
|
||||
|
||||
|
||||
def test_the_skills_name_no_particular_client():
|
||||
offenders = {
|
||||
str(p.relative_to(ROOT)): client_specific_hits(p.read_text())
|
||||
for p in sorted((ROOT / "plugin/skills").glob("*/SKILL.md"))
|
||||
}
|
||||
offenders = {path: hits for path, hits in offenders.items() if hits}
|
||||
assert not offenders, (
|
||||
f"skills that name one client: {offenders}. Skills are shared by every "
|
||||
f"client package (decision #4027); say the universal thing in the skill "
|
||||
f"and put the client's own name for it in that client's adapter."
|
||||
)
|
||||
|
||||
|
||||
def test_the_client_guard_can_fail():
|
||||
assert client_specific_hits("keep a copy in CLAUDE.md") == ["claude", "claude.md"]
|
||||
assert client_specific_hits("edits made through Bash") == ["Bash"]
|
||||
assert client_specific_hits("edits made through a bash shell") == []
|
||||
|
||||
|
||||
# A skill is shipped verbatim inside every client's package, at whatever path
|
||||
# that client installs skills to. A reference from a skill to anything outside
|
||||
# its own folder — the adapter's hooks, a manifest, a relative path upward —
|
||||
# points at a file that exists in one package layout and nowhere else
|
||||
# (plugin/PACKAGING.md).
|
||||
OUTSIDE_THE_SKILL = re.compile(r"\.\./|plugin/|hooks/|commands/|\.claude-plugin|\bplugin\.json\b|\bhooks\.json\b")
|
||||
|
||||
|
||||
def test_the_skills_reference_nothing_outside_their_folder():
|
||||
offenders = {
|
||||
str(p.relative_to(ROOT)): sorted(set(OUTSIDE_THE_SKILL.findall(p.read_text())))
|
||||
for p in sorted((ROOT / "plugin/skills").glob("*/SKILL.md"))
|
||||
}
|
||||
offenders = {path: refs for path, refs in offenders.items() if refs}
|
||||
assert not offenders, (
|
||||
f"skills that reference files outside their own folder: {offenders}. A "
|
||||
f"skill ships verbatim in every client package; see plugin/PACKAGING.md."
|
||||
)
|
||||
|
||||
|
||||
def test_the_layout_guard_can_fail():
|
||||
assert OUTSIDE_THE_SKILL.findall("run ../hooks/sync.sh") == ["../", "hooks/"]
|
||||
assert OUTSIDE_THE_SKILL.findall("see plugin.json") == ["plugin.json"]
|
||||
assert OUTSIDE_THE_SKILL.findall("record the plugin's behaviour") == []
|
||||
@@ -1,39 +1,35 @@
|
||||
"""The instruction surfaces must agree on how a rule reaches a session.
|
||||
"""What the instruction surfaces must say that one owner per topic cannot enforce.
|
||||
|
||||
WHY THIS EXISTS
|
||||
WHY THIS FILE IS SMALLER THAN IT WAS
|
||||
|
||||
Rule #119 makes the instruction surfaces the SPECIFICATION for product
|
||||
behaviour — there is no other place the "load the operator's rules" obligation
|
||||
is written down, and no code path enforces it. So a surface that states it
|
||||
differently isn't a documentation slip; it is the product behaving differently.
|
||||
Until milestone 410 this file pinned REDUNDANCY: every session-start surface —
|
||||
the MCP `_INSTRUCTIONS`, the plugin's static context, the `using-scribe` skill —
|
||||
had to restate how a rule reaches a session, what an empty session means, the
|
||||
Systems reflex and the snippet-recording triggers. That was the design #2494
|
||||
chose after #2497 (two surfaces disagreeing about who loads the rules) and
|
||||
#2198 (every hook silently inert): say it everywhere, so no single failure
|
||||
loses it.
|
||||
|
||||
That happened (#2497). `_INSTRUCTIONS` said the SessionStart hook "is the
|
||||
bridge" for getting rules into a session, while the `using-scribe` skill said to
|
||||
pull them yourself and treat any push as a bonus. An agent weighting the first
|
||||
would reasonably skip the pull.
|
||||
Decision #4027 replaced that with ownership, and tests/test_guidance_ownership.py
|
||||
now enforces it: every topic stated in full by exactly one owner, with a
|
||||
one-line pointer in the server's index. The tests that required every surface
|
||||
to repeat itself were retired there, because they enforced the very duplication
|
||||
that drifted (#4022). What they protected is still protected:
|
||||
- "state how to ask for a rule" / "say an empty session is not an empty
|
||||
rulebook" → the rules topic: owned by using-scribe, indexed in
|
||||
`_INSTRUCTIONS` with both phrases;
|
||||
- the Systems reflex and the snippet-recording triggers → their topics'
|
||||
markers, required on their owners;
|
||||
- "never name the SessionStart push without stating the ask" (#2497's exact
|
||||
shape) → retired outright. Since milestone 394 the push carries no rules,
|
||||
so naming it can no longer imply rules were handled, and requiring every
|
||||
surface that mentions it to restate the ask would force a copy.
|
||||
|
||||
#2198 is the case where that is wrong: every plugin hook was silently inert for
|
||||
an extended period, and nothing announced it. An agent trusting the push would
|
||||
have run with no binding rules and no signal — while those rules govern branch,
|
||||
commit, push and other hard-to-reverse actions.
|
||||
WHAT STAYS HERE
|
||||
|
||||
The asymmetry is the whole argument, and it is what these tests pin: asking
|
||||
when a rule had already arrived costs one redundant call; not asking when
|
||||
nothing arrived costs the operator's rules entirely.
|
||||
|
||||
MILESTONE 394 SHARPENED IT RATHER THAN RETIRING IT. There is no longer a
|
||||
resident set to pull, so "no rule in front of me" went from a rare and
|
||||
suspicious state to the ordinary state of most turns. The instruction that
|
||||
used to be supplementary — go and ask — is now the only route a rule has, and
|
||||
the surfaces must additionally say what an EMPTY session means, or a session
|
||||
reads silence as permission on nearly every turn.
|
||||
|
||||
WHAT THIS DOES NOT DO
|
||||
|
||||
It cannot tell whether two surfaces contradict each other in prose generally —
|
||||
that needs a reader. It pins the instructions whose absence is known to be
|
||||
load-bearing, and the specific shape the #2497 defect took: naming the push
|
||||
without also stating how to ask.
|
||||
Properties of a CLAIM rather than of who owns a topic: the index fits the
|
||||
fold; a surface that says rules bind also names what does not; using-scribe
|
||||
still sends a session to the reporting-back skill.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
@@ -41,44 +37,6 @@ import pathlib
|
||||
|
||||
ROOT = pathlib.Path(__file__).resolve().parents[1]
|
||||
|
||||
# THE PULL IS NOW THE ASK (milestone 394). This was `list_always_on_rules`,
|
||||
# the call that fetched the resident set. There is no resident set and no such
|
||||
# call: a rule reaches a session by retrieval, and the only thing a session can
|
||||
# DO about a rule it has not been handed is go looking for one.
|
||||
#
|
||||
# So the two halves this file used to pin separately — "pull the resident set"
|
||||
# and "and retrieve the conditional ones too" — have collapsed into one
|
||||
# instruction, and it is the load-bearing one rather than the supplementary
|
||||
# one it used to be.
|
||||
ASK = 'content_type="rule"'
|
||||
|
||||
# A surface must also say what an EMPTY session means, which is the half that
|
||||
# is newly dangerous. Under residency, "no rule in front of me" was rare and
|
||||
# suspicious. Under retrieval it is the ordinary state of most turns, so a
|
||||
# session that reads it as "there is no rule" is wrong on nearly every turn
|
||||
# rather than occasionally — the #3720 defect at session scale.
|
||||
#
|
||||
# Claim phrases, not a single word, for the reason BINDING_CLAIMS gives below:
|
||||
# a bare "matched" or "silence" appears in prose that is not making this claim
|
||||
# at all. A surface passes by asserting the distinction however it words it.
|
||||
ABSENCE_CLAIMS = (
|
||||
"nothing matched",
|
||||
"is not the same as \"there is no rule",
|
||||
"never \"there is no rule",
|
||||
"silence is not absence",
|
||||
"not evidence there is none",
|
||||
)
|
||||
|
||||
# Surfaces a session loads before substantive work. Hand-written because
|
||||
# "is this a session-start surface?" is an editorial fact, not a derivable one —
|
||||
# but each entry is asserted to EXIST, so a move or rename fails loudly here
|
||||
# instead of quietly dropping that surface from the check.
|
||||
SESSION_START_SURFACES = (
|
||||
ROOT / "src" / "scribe" / "mcp" / "server.py",
|
||||
ROOT / "plugin" / "hooks" / "scribe_static_context.md",
|
||||
ROOT / "plugin" / "skills" / "using-scribe" / "SKILL.md",
|
||||
)
|
||||
|
||||
|
||||
def _all_surfaces() -> list[tuple[str, str]]:
|
||||
"""(label, text) for every file a SESSION loads as instructions.
|
||||
@@ -99,54 +57,6 @@ def _all_surfaces() -> list[tuple[str, str]]:
|
||||
return found
|
||||
|
||||
|
||||
def test_every_session_start_surface_states_the_ask():
|
||||
"""Retrieval is the only delivery, so asking is the only recourse."""
|
||||
missing = []
|
||||
for path in SESSION_START_SURFACES:
|
||||
assert path.exists(), (
|
||||
f"{path.relative_to(ROOT)} is gone — it was one of the surfaces "
|
||||
f"carrying the rules instruction. If it moved, update "
|
||||
f"SESSION_START_SURFACES; if it was retired, check the instruction "
|
||||
f"still lives somewhere a fresh session reads."
|
||||
)
|
||||
if ASK not in path.read_text():
|
||||
missing.append(str(path.relative_to(ROOT)))
|
||||
assert not missing, (
|
||||
f"these surfaces never tell the agent how to ask for a rule "
|
||||
f"({ASK}): {missing}. Nothing is pushed and nothing is resident, so a "
|
||||
f"surface that omits this leaves a session with no way to reach a rule "
|
||||
f"it was not handed — bound by nothing (#2198, #2497, milestone 394)."
|
||||
)
|
||||
|
||||
|
||||
def test_every_session_start_surface_says_an_empty_session_is_not_an_empty_rulebook():
|
||||
"""The half that got dangerous when residency went away.
|
||||
|
||||
Under the old model a session opened holding every applicable rule, so
|
||||
"nothing is in front of me" was a rare state and a suspicious one. Under
|
||||
retrieval it is the NORMAL state of most turns. A surface that describes
|
||||
where rules come from, without also saying what their absence means, leaves
|
||||
a session reading silence as permission — on nearly every turn rather than
|
||||
occasionally.
|
||||
|
||||
That is #3720's defect ("absence reads as non-existence") moved from a
|
||||
readout to the session itself, and this milestone is what makes every
|
||||
session start in the absent state.
|
||||
"""
|
||||
missing = []
|
||||
for path in SESSION_START_SURFACES:
|
||||
text = path.read_text().lower()
|
||||
if not any(c.lower() in text for c in ABSENCE_CLAIMS):
|
||||
missing.append(str(path.relative_to(ROOT)))
|
||||
assert not missing, (
|
||||
f"these surfaces say how a rule arrives but never what it means when "
|
||||
f"none does: {missing}. 'No rule arrived' means 'nothing matched', "
|
||||
f"never 'there is no rule' — and only one of those has been checked. "
|
||||
f"Say it however you like; one of {ABSENCE_CLAIMS} is what this looks "
|
||||
f"for."
|
||||
)
|
||||
|
||||
|
||||
def _instructions_text() -> str:
|
||||
"""The _INSTRUCTIONS literal from server.py, as the client would see it."""
|
||||
import re
|
||||
@@ -166,114 +76,11 @@ INSTRUCTIONS_BUDGET = 2000
|
||||
def test_instructions_fit_the_fold():
|
||||
text = _instructions_text()
|
||||
assert len(text) <= INSTRUCTIONS_BUDGET, (
|
||||
f"_INSTRUCTIONS is {len(text)} chars; the client injects only ~2,048 "
|
||||
f"and silently cuts the rest (#2562). This block is a MAP — move the "
|
||||
f"detail to the tool's docstring (delivered at reach-for time), the "
|
||||
f"plugin static context (always delivered), or a skill; see the "
|
||||
f"comment above _INSTRUCTIONS."
|
||||
)
|
||||
|
||||
|
||||
def test_floor_states_the_systems_reflex():
|
||||
"""Write-time tagging guidance must live on the surface that always arrives.
|
||||
|
||||
#2562's behavioral finding: with the guidance only in tool descriptions,
|
||||
sessions filed records untagged. The static context is the delivery floor,
|
||||
so the tag-as-you-write reflex has to be stated there.
|
||||
"""
|
||||
floor = (ROOT / "plugin" / "hooks" / "scribe_static_context.md").read_text()
|
||||
for needle in ("system_ids", "create_system"):
|
||||
assert needle in floor, (
|
||||
f"plugin/hooks/scribe_static_context.md no longer mentions "
|
||||
f"{needle} — the Systems tagging reflex must be stated on the "
|
||||
f"floor, not only in tool descriptions (#2562)."
|
||||
)
|
||||
|
||||
|
||||
def test_floor_names_the_snippet_recording_triggers():
|
||||
"""The floor must state the pattern-library recording model, by name.
|
||||
|
||||
#2664's behavioral finding: recording guidance as a trailing clause of the
|
||||
reuse bullet converted zero times outside snippet-minded sessions. The
|
||||
2026-08-16 ruling (decision #2686) then replaced the reactive model
|
||||
entirely: every shape is recorded at FIRST build — no "will it recur?"
|
||||
judgment — and second-copy consolidation is only the backstop. The floor
|
||||
is the delivery surface for that reflex, so all three elements must stay
|
||||
stated: the tool, the first-build trigger, and the backstop.
|
||||
"""
|
||||
floor = (ROOT / "plugin" / "hooks" / "scribe_static_context.md").read_text()
|
||||
for needle in ("create_snippet", "first build", "second copy"):
|
||||
assert needle in floor, (
|
||||
f"plugin/hooks/scribe_static_context.md no longer states the "
|
||||
f"snippet-recording model ({needle!r}) — record-every-shape-at-"
|
||||
f"first-build with second-copy consolidation as the backstop must "
|
||||
f"be stated on the floor (#2664, decision #2686)."
|
||||
)
|
||||
|
||||
|
||||
# Topics displaced from _INSTRUCTIONS when it was cut to fit the fold. Each
|
||||
# must remain stated on at least one DELIVERED surface: a tool docstring
|
||||
# (arrives with the tool schema), the plugin static context (always arrives),
|
||||
# or a bundled skill (arrives on trigger match). Keyed by a phrase distinctive
|
||||
# enough that its disappearance means the guidance is gone, not reworded —
|
||||
# update the phrase alongside a deliberate rewording.
|
||||
DISPLACED_TOPICS = {
|
||||
"supersedes": "supersedes",
|
||||
"trash is recoverable": "deleted_batch_id",
|
||||
"duplicate gate": "duplicate",
|
||||
"systems tag-as-you-write": "system_ids",
|
||||
"reference note vs dev-log": "reference note",
|
||||
"work-logs over body rewrites": "add_task_log",
|
||||
"rule homes / altitude": "create_project_rule",
|
||||
"rules vs other entities": "standing instruction",
|
||||
"shared records are suggestions": "shared",
|
||||
"processes run verbatim": "verbatim",
|
||||
"snippet reuse reflex": "when_to_use",
|
||||
"compaction at seams": "compact",
|
||||
"plans are milestones": "start_planning",
|
||||
"scope to the entered project": "cross-project",
|
||||
"project bootstrap needs confirmation": "never guessing a project",
|
||||
}
|
||||
|
||||
|
||||
def test_displaced_topics_live_on_a_delivered_surface():
|
||||
corpus = ""
|
||||
for p in (ROOT / "src" / "scribe" / "mcp" / "tools").glob("*.py"):
|
||||
corpus += p.read_text()
|
||||
corpus += (ROOT / "plugin" / "hooks" / "scribe_static_context.md").read_text()
|
||||
for p in (ROOT / "plugin" / "skills").rglob("SKILL.md"):
|
||||
corpus += p.read_text()
|
||||
corpus = corpus.lower()
|
||||
missing = [
|
||||
f"{topic} (phrase: {phrase!r})"
|
||||
for topic, phrase in DISPLACED_TOPICS.items()
|
||||
if phrase.lower() not in corpus
|
||||
]
|
||||
assert not missing, (
|
||||
f"guidance displaced from _INSTRUCTIONS has fallen off every delivered "
|
||||
f"surface (tool docstrings / static context / skills): {missing}. It "
|
||||
f"was cut from _INSTRUCTIONS deliberately (#2562) on the premise it "
|
||||
f"lives elsewhere — restore it somewhere that delivers."
|
||||
)
|
||||
|
||||
|
||||
def test_no_surface_names_the_push_without_stating_the_ask():
|
||||
"""The exact shape #2497 took.
|
||||
|
||||
Mentioning the SessionStart hook is fine and often useful. Mentioning it
|
||||
*instead of* the pull is the defect: it reads as "this is handled", and the
|
||||
surface that says so is the one an agent has least reason to doubt.
|
||||
"""
|
||||
offenders = [
|
||||
label for label, text in _all_surfaces()
|
||||
if "SessionStart" in text and ASK not in text
|
||||
]
|
||||
assert not offenders, (
|
||||
f"these surfaces describe the SessionStart push but never state how to "
|
||||
f"ask: {offenders}. The push is a delivery optimisation, not the "
|
||||
f"bridge — it can be absent without saying so, and since milestone 394 "
|
||||
f"it carries no rules at all. Name it if it helps, but say how to ask "
|
||||
f"({ASK}) regardless."
|
||||
f"_INSTRUCTIONS is {len(text)} chars; Claude Code injects only ~2,048 "
|
||||
f"and silently cuts the rest (#2562). This block is an INDEX — state "
|
||||
f"the topic in full on its owner (a skill or the tool's docstring, "
|
||||
f"decision #4027) and give it at most a line here; see the comment "
|
||||
f"above _INSTRUCTIONS."
|
||||
)
|
||||
|
||||
|
||||
@@ -322,3 +129,38 @@ def test_a_surface_claiming_rules_bind_also_names_what_does_not():
|
||||
f"'how the operator likes this done' as something it may not proceed "
|
||||
f"past. Name the other kind, however briefly."
|
||||
)
|
||||
|
||||
|
||||
# ── Reporting back (milestone 409 step 3) ──────────────────────────────
|
||||
#
|
||||
# The reporting-back skill carries the shapes, but a skill only helps if it
|
||||
# fires. The reflex that sends a session to it is stated in using-scribe; the
|
||||
# server's index carries a REPORT line pointing at `placement`, and the in-band
|
||||
# cue on update_task is the half that fires on its own in every client. The
|
||||
# static context was a second plugin-side copy until milestone 410.
|
||||
|
||||
REPORT_REFLEX = "report back in a shape the operator can read"
|
||||
REPORT_SURFACES = (
|
||||
ROOT / "plugin" / "skills" / "using-scribe" / "SKILL.md",
|
||||
)
|
||||
|
||||
|
||||
def test_the_reporting_reflex_is_stated_in_using_scribe():
|
||||
missing = []
|
||||
for path in REPORT_SURFACES:
|
||||
text = " ".join(path.read_text().split()).lower()
|
||||
if REPORT_REFLEX not in text or "reporting-back" not in text:
|
||||
missing.append(str(path.relative_to(ROOT)))
|
||||
assert not missing, (
|
||||
f"these surfaces no longer send a session to the reporting-back skill "
|
||||
f"({REPORT_REFLEX!r} plus the skill's name): {missing}. Without the "
|
||||
f"reflex the skill loads only when its description happens to match."
|
||||
)
|
||||
|
||||
|
||||
def test_the_reporting_reflex_takes_placement_from_the_record():
|
||||
"""The failure milestone 409 began from: a placement reconstructed from
|
||||
memory reads exactly like a real one when it is wrong."""
|
||||
for path in REPORT_SURFACES:
|
||||
text = " ".join(path.read_text().split())
|
||||
assert "`placement`" in text, f"{path.relative_to(ROOT)} no longer points at the placement block"
|
||||
|
||||
@@ -0,0 +1,96 @@
|
||||
"""Real-Postgres tests for task placement (milestone 409 step 1).
|
||||
|
||||
What a mock cannot show: that step order is creation order rather than update
|
||||
order, that position and progress are computed over the steps the CALLER may
|
||||
read, and that a collaborator holding one shared task is not shown the titles
|
||||
of steps they cannot open.
|
||||
"""
|
||||
import pytest
|
||||
import pytest_asyncio
|
||||
|
||||
from scribe.models import async_session
|
||||
from scribe.models.note import Note
|
||||
from scribe.models.project import Project
|
||||
from scribe.models.share import NoteShare, ProjectShare
|
||||
from scribe.services import notes as notes_svc
|
||||
from scribe.services.placement import task_placement
|
||||
from scribe.services.record_batch import BatchItem, BatchMilestone, create_batch
|
||||
from tests.helpers import ensure_user
|
||||
|
||||
pytestmark = [pytest.mark.integration, pytest.mark.usefixtures("_dispose_engine", "_no_embedding")]
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def plan():
|
||||
"""Owner, a collaborator, a project and a four-step milestone."""
|
||||
async with async_session() as s:
|
||||
owner = await ensure_user(s, "placement_owner")
|
||||
other = await ensure_user(s, "placement_collaborator")
|
||||
project = Project(user_id=owner.id, title="Placement target")
|
||||
s.add(project)
|
||||
await s.flush()
|
||||
ids = {"owner": owner.id, "other": other.id, "pid": project.id}
|
||||
await s.commit()
|
||||
ms, steps = await create_batch(ids["owner"], [
|
||||
BatchItem(title="Step 1", status="done"),
|
||||
BatchItem(title="Step 2", status="in_progress"),
|
||||
BatchItem(title="Step 3"),
|
||||
BatchItem(title="Step 4"),
|
||||
], project_id=ids["pid"], milestone=BatchMilestone(title="The plan", body="b"))
|
||||
ids.update({"mid": ms.id, "steps": steps})
|
||||
return ids
|
||||
|
||||
|
||||
async def test_the_owner_sees_the_whole_placement(plan):
|
||||
step2 = plan["steps"][1]
|
||||
out = await task_placement(plan["owner"], step2)
|
||||
assert out["project"] == {"id": plan["pid"], "title": "Placement target"}
|
||||
assert out["milestone"] == {"id": plan["mid"], "title": "The plan", "status": "active"}
|
||||
assert out["position"] == {"step": 2, "of": 4}
|
||||
assert out["progress"] == {"completed": 1, "total": 4, "pct": 25.0}
|
||||
assert out["next"] == {"id": plan["steps"][2].id, "title": "Step 3", "status": "todo"}
|
||||
|
||||
|
||||
async def test_step_order_is_creation_order_not_last_update(plan):
|
||||
"""Touching step 1 must not move it to the end of the list."""
|
||||
first = plan["steps"][0]
|
||||
updated = await notes_svc.update_note(plan["owner"], first.id, body="edited later")
|
||||
out = await task_placement(plan["owner"], updated)
|
||||
assert out["position"] == {"step": 1, "of": 4}
|
||||
|
||||
|
||||
async def test_the_last_step_points_back_at_the_earliest_open_one(plan):
|
||||
last = await notes_svc.update_note(plan["owner"], plan["steps"][3].id, status="done")
|
||||
out = await task_placement(plan["owner"], last)
|
||||
assert out["next"]["title"] == "Step 2"
|
||||
assert out["progress"]["completed"] == 2
|
||||
|
||||
|
||||
async def test_a_task_outside_any_milestone_gets_its_project_alone(plan):
|
||||
loose = await notes_svc.create_note(plan["owner"], title="Loose task", status="todo",
|
||||
project_id=plan["pid"])
|
||||
out = await task_placement(plan["owner"], loose)
|
||||
assert out == {"project": {"id": plan["pid"], "title": "Placement target"}}
|
||||
|
||||
|
||||
async def test_a_collaborator_holding_one_shared_task_sees_no_plan(plan):
|
||||
"""A note share opens one task. It does not open the project, so neither the
|
||||
milestone nor its other steps' titles may leak through the placement."""
|
||||
step3 = plan["steps"][2]
|
||||
async with async_session() as s:
|
||||
s.add(NoteShare(note_id=step3.id, shared_with_user_id=plan["other"],
|
||||
permission="editor", invited_by=plan["owner"]))
|
||||
await s.commit()
|
||||
assert await task_placement(plan["other"], step3) is None
|
||||
|
||||
|
||||
async def test_a_project_collaborator_sees_the_plan(plan):
|
||||
async with async_session() as s:
|
||||
s.add(ProjectShare(project_id=plan["pid"], shared_with_user_id=plan["other"],
|
||||
permission="viewer", invited_by=plan["owner"]))
|
||||
await s.commit()
|
||||
step3 = await s.get(Note, plan["steps"][2].id)
|
||||
out = await task_placement(plan["other"], step3)
|
||||
assert out["milestone"]["title"] == "The plan"
|
||||
assert out["position"] == {"step": 3, "of": 4}
|
||||
assert out["next"]["title"] == "Step 4"
|
||||
@@ -0,0 +1,113 @@
|
||||
"""Real-Postgres proof that a batch create cannot be interfered with (#4016).
|
||||
|
||||
What mocks cannot show: that the ids a placeholder is rewritten to are the ids
|
||||
the rows actually got, that two batches running at the same time each cite
|
||||
their OWN records, and that a failing batch leaves nothing behind. Those are
|
||||
properties of the sequence and the transaction, so they are measured against
|
||||
the real ones.
|
||||
"""
|
||||
import asyncio
|
||||
|
||||
import pytest
|
||||
import pytest_asyncio
|
||||
from sqlalchemy import func, select
|
||||
|
||||
from scribe.models import async_session
|
||||
from scribe.models.milestone import Milestone
|
||||
from scribe.models.note import Note
|
||||
from scribe.models.project import Project
|
||||
from scribe.services import record_refs
|
||||
from scribe.services.planning import start_planning
|
||||
from scribe.services.record_batch import BatchItem, create_batch
|
||||
from tests.helpers import ensure_user
|
||||
|
||||
pytestmark = [pytest.mark.integration, pytest.mark.usefixtures("_dispose_engine", "_no_embedding")]
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def seeded():
|
||||
async with async_session() as s:
|
||||
owner = await ensure_user(s, "record_batch_owner")
|
||||
project = Project(user_id=owner.id, title="Batch target")
|
||||
s.add(project)
|
||||
await s.flush()
|
||||
ids = {"owner": owner.id, "pid": project.id}
|
||||
await s.commit()
|
||||
return ids
|
||||
|
||||
|
||||
async def _count_in_project(pid: int) -> tuple[int, int]:
|
||||
async with async_session() as s:
|
||||
notes = (await s.execute(select(func.count(Note.id)).where(Note.project_id == pid))).scalar()
|
||||
milestones = (await s.execute(
|
||||
select(func.count(Milestone.id)).where(Milestone.project_id == pid)
|
||||
)).scalar()
|
||||
return notes, milestones
|
||||
|
||||
|
||||
async def test_a_plan_and_its_steps_cite_each_others_real_ids(seeded):
|
||||
owner, pid = seeded["owner"], seeded["pid"]
|
||||
out = await start_planning(
|
||||
owner, pid, "Ship it",
|
||||
body="First {{ref:1}}, then {{ref:2}}.",
|
||||
steps=[
|
||||
BatchItem(title="Step one", body="Part of {{ref:milestone}}; next is {{ref:2}}."),
|
||||
BatchItem(title="Step two", body="Follows {{ref:1}}."),
|
||||
],
|
||||
)
|
||||
ms = out["milestone"]
|
||||
one, two = out["steps"]
|
||||
assert ms["body"] == f'First #{one["id"]} "Step one", then #{two["id"]} "Step two".'
|
||||
assert one["body"] == f'Part of milestone {ms["id"]} "Ship it"; next is #{two["id"]} "Step two".'
|
||||
assert two["body"] == f'Follows #{one["id"]} "Step one".'
|
||||
assert one["milestone_id"] == two["milestone_id"] == ms["id"]
|
||||
assert one["project_id"] == pid
|
||||
|
||||
# What was returned is what was stored — not a rewrite applied to the
|
||||
# response alone.
|
||||
async with async_session() as s:
|
||||
stored = await s.get(Note, two["id"])
|
||||
assert stored.body == f'Follows #{one["id"]} "Step one".'
|
||||
|
||||
|
||||
async def test_concurrent_batches_each_cite_their_own_records(seeded):
|
||||
"""The operator's scenario: several sessions creating at once. Each batch
|
||||
must resolve {{ref:1}} to ITS first record, whatever the others take."""
|
||||
owner, pid = seeded["owner"], seeded["pid"]
|
||||
|
||||
async def one_batch(tag: str):
|
||||
return await create_batch(owner, [
|
||||
BatchItem(title=f"{tag} first"),
|
||||
BatchItem(title=f"{tag} second", body="points at {{ref:1}}"),
|
||||
], project_id=pid)
|
||||
|
||||
results = await asyncio.gather(*(one_batch(f"session-{i}") for i in range(6)))
|
||||
|
||||
all_ids = [n.id for _ms, notes in results for n in notes]
|
||||
assert len(all_ids) == len(set(all_ids)), "the sequence handed out an id twice"
|
||||
for _ms, (first, second) in results:
|
||||
assert second.body == f'points at #{first.id} "{first.title}"'
|
||||
async with async_session() as s:
|
||||
assert (await s.get(Note, first.id)).title == first.title
|
||||
|
||||
|
||||
async def test_a_failing_batch_writes_nothing(seeded):
|
||||
owner, pid = seeded["owner"], seeded["pid"]
|
||||
before = await _count_in_project(pid)
|
||||
with pytest.raises(ValueError, match="Invalid status"):
|
||||
await start_planning(owner, pid, "Doomed", steps=[
|
||||
BatchItem(title="fine"),
|
||||
BatchItem(title="broken", status="someday"),
|
||||
])
|
||||
assert await _count_in_project(pid) == before, "a partial batch (or its milestone) was left behind"
|
||||
|
||||
|
||||
async def test_a_guess_above_the_real_highest_id_is_refused(seeded):
|
||||
owner, pid = seeded["owner"], seeded["pid"]
|
||||
_ms, (made,) = await create_batch(owner, [BatchItem(title="anchor")], project_id=pid)
|
||||
# `made` is the newest row, so anything just past it can only be a guess —
|
||||
# unless a parallel test in the lane created more since, which only moves
|
||||
# the ceiling up and keeps made.id itself a real id.
|
||||
with pytest.raises(ValueError, match="does not exist yet"):
|
||||
await record_refs.refuse_guessed_ids(f"#{made.id + record_refs.GUESS_WINDOW}")
|
||||
await record_refs.refuse_guessed_ids(f"#{made.id}")
|
||||
@@ -0,0 +1,170 @@
|
||||
"""The batch create door, and the guessed-id refusal at every create/update door (#4016).
|
||||
|
||||
What the unit lane can pin: the door refuses a guessed id BEFORE the duplicate
|
||||
gate or any write, a batch is refused whole, and a malformed record fails with
|
||||
a message naming it. That the ids really come back resolved and atomic needs
|
||||
Postgres — see test_integration_record_batch.py.
|
||||
"""
|
||||
from unittest.mock import AsyncMock, MagicMock, patch
|
||||
|
||||
import pytest
|
||||
|
||||
from scribe.services import record_refs
|
||||
from scribe.services.record_batch import BatchItem, BatchMilestone, check_batch
|
||||
|
||||
pytestmark = pytest.mark.usefixtures("_bind_user")
|
||||
|
||||
|
||||
def _max(n: int):
|
||||
return patch.object(record_refs, "_max_note_id", AsyncMock(return_value=n))
|
||||
|
||||
|
||||
# ── the refusal reaches every door ─────────────────────────────────────────
|
||||
|
||||
|
||||
@pytest.mark.parametrize(("module", "name", "kwargs"), [
|
||||
("scribe.mcp.tools.tasks", "create_task", {"title": "t", "body": "after #101"}),
|
||||
("scribe.mcp.tools.tasks", "update_task", {"task_id": 5, "body": "see #101"}),
|
||||
("scribe.mcp.tools.notes", "create_note", {"title": "n", "body": "index: #101"}),
|
||||
("scribe.mcp.tools.notes", "update_note", {"note_id": 5, "body": "index: #101"}),
|
||||
("scribe.mcp.tools.milestones", "create_milestone", {"project_id": 1, "title": "m", "body": "#101"}),
|
||||
("scribe.mcp.tools.milestones", "update_milestone", {"project_id": 1, "milestone_id": 2, "body": "#101"}),
|
||||
("scribe.mcp.tools.tasks", "start_planning", {"project_id": 1, "title": "p", "body": "step #101"}),
|
||||
("scribe.mcp.tools.tasks", "create_records", {"records": [{"title": "a", "body": "#101"}]}),
|
||||
])
|
||||
async def test_every_writing_door_refuses_a_guessed_id_before_writing(module, name, kwargs):
|
||||
import importlib
|
||||
|
||||
mod = importlib.import_module(module)
|
||||
tool = getattr(mod, name)
|
||||
# Every service a door could reach is a MagicMock that must stay untouched:
|
||||
# the refusal has to land before the duplicate gate and before the write.
|
||||
with _max(100), \
|
||||
patch("scribe.services.dedup.find_duplicate_note", AsyncMock()) as dedup, \
|
||||
patch("scribe.services.notes.create_note", AsyncMock()) as create, \
|
||||
patch("scribe.services.notes.update_note", AsyncMock()) as update, \
|
||||
patch("scribe.services.milestones.create_milestone", AsyncMock()) as ms_create, \
|
||||
patch("scribe.services.milestones.update_milestone", AsyncMock()) as ms_update, \
|
||||
patch("scribe.services.record_batch.create_batch", AsyncMock()) as batch, \
|
||||
patch("scribe.services.planning.start_planning", AsyncMock()) as plan:
|
||||
with pytest.raises(ValueError, match="#101"):
|
||||
await tool(**kwargs)
|
||||
for mock in (dedup, create, update, ms_create, ms_update, batch, plan):
|
||||
mock.assert_not_awaited()
|
||||
|
||||
|
||||
async def test_a_real_id_passes_the_door():
|
||||
from scribe.mcp.tools.tasks import create_task
|
||||
|
||||
fake = MagicMock(id=90, project_id=None)
|
||||
fake.to_dict.return_value = {"id": 90}
|
||||
with _max(100), \
|
||||
patch("scribe.mcp.tools.tasks.dedup_svc.find_duplicate_note", AsyncMock(return_value=None)), \
|
||||
patch("scribe.mcp.tools.tasks.notes_svc.create_note", AsyncMock(return_value=fake)), \
|
||||
patch("scribe.mcp.tools.tasks.systems_tools.attach_systems", AsyncMock()):
|
||||
out = await create_task(title="t", body="follows #42")
|
||||
assert out["id"] == 90
|
||||
|
||||
|
||||
# ── create_records ─────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
async def test_unknown_record_fields_are_refused_not_dropped():
|
||||
from scribe.mcp.tools.tasks import create_records
|
||||
|
||||
with pytest.raises(ValueError, match=r"record 2 has unknown field\(s\) \['milestone'\]"):
|
||||
await create_records(records=[{"title": "a"}, {"title": "b", "milestone": 4}])
|
||||
|
||||
|
||||
async def test_a_record_type_must_be_task_or_note():
|
||||
from scribe.mcp.tools.tasks import create_records
|
||||
|
||||
with pytest.raises(ValueError, match="type must be 'task' or 'note'"):
|
||||
await create_records(records=[{"title": "a", "type": "rule"}])
|
||||
|
||||
|
||||
async def test_one_duplicate_blocks_the_whole_batch():
|
||||
from scribe.mcp.tools.tasks import create_records
|
||||
|
||||
dup = MagicMock(id=33, title="existing", similarity=1.0, reason="title")
|
||||
gate = AsyncMock(side_effect=[None, dup])
|
||||
with patch("scribe.mcp.tools.tasks.dedup_svc.find_duplicate_note", gate), \
|
||||
patch("scribe.mcp.tools.tasks.batch_svc.create_batch", AsyncMock()) as batch:
|
||||
out = await create_records(records=[{"title": "fresh"}, {"title": "existing"}], project_id=3)
|
||||
assert out["duplicate"] is True and out["record"] == 2 and out["existing_id"] == 33
|
||||
assert "Nothing in the batch was created" in out["message"]
|
||||
batch.assert_not_awaited()
|
||||
|
||||
|
||||
async def test_create_records_returns_ids_in_order():
|
||||
from scribe.mcp.tools.tasks import create_records
|
||||
|
||||
notes = []
|
||||
for nid in (51, 53):
|
||||
n = MagicMock(id=nid)
|
||||
n.to_dict.return_value = {"id": nid}
|
||||
notes.append(n)
|
||||
with patch("scribe.mcp.tools.tasks.dedup_svc.find_duplicate_note", AsyncMock(return_value=None)), \
|
||||
patch("scribe.mcp.tools.tasks.batch_svc.create_batch", AsyncMock(return_value=(None, notes))) as batch:
|
||||
out = await create_records(
|
||||
records=[{"title": "a"}, {"title": "b", "type": "note", "body": "after {{ref:1}}"}],
|
||||
milestone_id=8,
|
||||
)
|
||||
assert out["ids"] == [51, 53]
|
||||
items = batch.call_args.args[1]
|
||||
assert [i.is_task for i in items] == [True, False]
|
||||
assert batch.call_args.kwargs["milestone_id"] == 8
|
||||
|
||||
|
||||
async def test_start_planning_hands_its_steps_to_the_service():
|
||||
from scribe.mcp.tools.tasks import start_planning
|
||||
|
||||
with patch("scribe.mcp.tools.tasks.dedup_svc.find_duplicate_note", AsyncMock(return_value=None)), \
|
||||
patch("scribe.mcp.tools.tasks.planning_svc.start_planning",
|
||||
AsyncMock(return_value={"milestone": {"id": 1}})) as svc:
|
||||
await start_planning(project_id=3, title="Plan", body="see {{ref:1}}",
|
||||
steps=[{"title": "Step 1", "kind": "spike"}])
|
||||
kwargs = svc.call_args.kwargs
|
||||
assert kwargs["body"] == "see {{ref:1}}"
|
||||
assert [s.title for s in kwargs["steps"]] == ["Step 1"]
|
||||
assert kwargs["steps"][0].task_kind == "spike"
|
||||
|
||||
|
||||
# ── check_batch: refused whole, before any write ───────────────────────────
|
||||
|
||||
|
||||
def test_a_placeholder_past_the_end_is_refused():
|
||||
with pytest.raises(ValueError, match=r"\{\{ref:3\}\} names no record"):
|
||||
check_batch([BatchItem(title="a"), BatchItem(title="b", body="{{ref:3}}")])
|
||||
|
||||
|
||||
def test_the_milestone_placeholder_needs_a_milestone():
|
||||
with pytest.raises(ValueError, match=r"\{\{ref:milestone\}\}"):
|
||||
check_batch([BatchItem(title="a", body="in {{ref:milestone}}")])
|
||||
|
||||
|
||||
def test_the_milestone_placeholder_is_valid_with_one():
|
||||
check_batch([BatchItem(title="a", body="in {{ref:milestone}}")],
|
||||
BatchMilestone(title="m", body="first {{ref:1}}"))
|
||||
|
||||
|
||||
def test_every_record_needs_a_title():
|
||||
with pytest.raises(ValueError, match="record 2 has no title"):
|
||||
check_batch([BatchItem(title="a"), BatchItem(title=" ")])
|
||||
|
||||
|
||||
def test_a_batch_has_a_ceiling():
|
||||
from scribe.services.record_batch import MAX_BATCH
|
||||
|
||||
with pytest.raises(ValueError, match="at most"):
|
||||
check_batch([BatchItem(title=f"t{i}") for i in range(MAX_BATCH + 1)])
|
||||
|
||||
|
||||
async def test_start_planning_refuses_placeholders_with_no_steps():
|
||||
from scribe.services.planning import start_planning
|
||||
|
||||
with patch("scribe.services.planning.projects_svc.get_project", AsyncMock(return_value=MagicMock())), \
|
||||
patch("scribe.services.planning.milestones_svc.create_milestone", AsyncMock()) as create:
|
||||
with pytest.raises(ValueError, match="no steps were given"):
|
||||
await start_planning(user_id=7, project_id=3, title="p", body="see {{ref:1}}")
|
||||
create.assert_not_awaited()
|
||||
@@ -16,7 +16,11 @@ async def test_start_planning_tool_delegates_to_service():
|
||||
from scribe.mcp.tools.tasks import start_planning
|
||||
out = await start_planning(project_id=3, title="Plan it")
|
||||
assert out["milestone"]["id"] == 5
|
||||
assert mock.call_args.kwargs == {"user_id": 7, "project_id": 3, "title": "Plan it"}
|
||||
# No body and no steps reach the service as None, so it seeds the template
|
||||
# and takes the single-milestone path.
|
||||
assert mock.call_args.kwargs == {
|
||||
"user_id": 7, "project_id": 3, "title": "Plan it", "body": None, "steps": None,
|
||||
}
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
|
||||
@@ -0,0 +1,33 @@
|
||||
"""update_task closes a task with a one-line reminder of what the report needs.
|
||||
|
||||
The in-band half of milestone 409 step 3: the skill and static context only
|
||||
exist in the Claude Code plugin, and a tool response reaches every MCP client
|
||||
at the moment a piece of work closes. Pinned on the response, not the wording.
|
||||
"""
|
||||
from unittest.mock import AsyncMock, MagicMock, patch
|
||||
|
||||
import pytest
|
||||
|
||||
pytestmark = pytest.mark.usefixtures("_bind_user")
|
||||
|
||||
|
||||
async def _update(**kwargs):
|
||||
from scribe.mcp.tools.tasks import update_task
|
||||
|
||||
note = MagicMock(id=5, user_id=7, project_id=None)
|
||||
note.to_dict.return_value = {"id": 5}
|
||||
with patch("scribe.mcp.tools.tasks.notes_svc.update_note", AsyncMock(return_value=note)), \
|
||||
patch("scribe.mcp.tools.tasks.systems_tools.attach_systems", AsyncMock()), \
|
||||
patch("scribe.mcp.tools.tasks.placement_svc.attach_placement", AsyncMock()):
|
||||
return await update_task(task_id=5, **kwargs)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("status", ["done", "cancelled"])
|
||||
async def test_closing_a_task_carries_the_cue(status):
|
||||
out = await _update(status=status)
|
||||
assert "placement" in out["report_back"] and "needs them" in out["report_back"]
|
||||
|
||||
|
||||
@pytest.mark.parametrize("kwargs", [{"status": "in_progress"}, {"status": "todo"}, {"body": "more notes"}])
|
||||
async def test_other_updates_do_not(kwargs):
|
||||
assert "report_back" not in await _update(**kwargs)
|
||||
@@ -0,0 +1,55 @@
|
||||
"""The reporting-back skill keeps its shape (milestone 409 step 2).
|
||||
|
||||
WHY THIS EXISTS
|
||||
|
||||
The skill is what turns a reply written in the order the work happened into
|
||||
one the operator can read: where the work sits, what changed, what needs
|
||||
them, what is next. Its value is in its SECTIONS, and a later tidy-up that
|
||||
folds them into prose would leave a skill that still loads and no longer
|
||||
shapes anything.
|
||||
|
||||
WHAT THIS PINS, AND WHAT IT DOES NOT
|
||||
|
||||
Structure, never wording — the same reason test_create_tools_disambiguate
|
||||
gives: a test that punishes rewriting gets deleted. It pins that the
|
||||
completion report keeps its five sections, that placement is taken from the
|
||||
record rather than recalled, and that the shipped shapes stay domain-neutral.
|
||||
Whether the guidance is any good is milestone 409's last step, read against
|
||||
real replies, not something a test can see.
|
||||
"""
|
||||
import pathlib
|
||||
import re
|
||||
|
||||
SKILL = pathlib.Path(__file__).resolve().parents[1] / "plugin/skills/reporting-back/SKILL.md"
|
||||
|
||||
|
||||
def _text() -> str:
|
||||
return " ".join(SKILL.read_text().split())
|
||||
|
||||
|
||||
def test_the_skill_names_itself_as_its_directory():
|
||||
front = re.search(r"^---\s*\nname:\s*(\S+)", SKILL.read_text())
|
||||
assert front and front.group(1) == "reporting-back"
|
||||
|
||||
|
||||
def test_the_completion_report_keeps_its_sections():
|
||||
text = _text()
|
||||
for section in ("Where this sits", "What now works", "How / why", "Needs you", "Next"):
|
||||
assert section in text, f"the completion report lost its {section!r} section"
|
||||
|
||||
|
||||
def test_placement_comes_from_the_record():
|
||||
"""The failure this milestone started from: a placement written from memory
|
||||
reads exactly like a real one when it is wrong."""
|
||||
text = _text().lower()
|
||||
assert "placement" in text and "take the placement from the record" in text
|
||||
|
||||
|
||||
def test_the_shipped_shapes_assume_no_particular_domain():
|
||||
"""Scribe is domain-neutral: a home-infrastructure or writing project reads
|
||||
these too. Software-specific evidence belongs in an operator's own
|
||||
preferences, never in the product default."""
|
||||
text = _text()
|
||||
dev_only = [w for w in (r"\bCI\b", r"\bcommit", r"\bpull request", r"file:line", r"\bpytest\b")
|
||||
if re.search(w, text, re.IGNORECASE)]
|
||||
assert not dev_only, f"software-only vocabulary in a product-wide shape: {dev_only}"
|
||||
@@ -0,0 +1,48 @@
|
||||
"""Placement — the pure half: what "next" means, and that a lookup never breaks a write.
|
||||
|
||||
The query half (position over readable steps, what a collaborator may see)
|
||||
needs Postgres and lives in test_integration_placement.py.
|
||||
"""
|
||||
from types import SimpleNamespace
|
||||
from unittest.mock import AsyncMock, patch
|
||||
|
||||
from scribe.services import placement
|
||||
from scribe.services.placement import _next_open
|
||||
|
||||
|
||||
def _steps(*statuses):
|
||||
return [SimpleNamespace(id=i, title=f"step {i}", status=s) for i, s in enumerate(statuses, start=1)]
|
||||
|
||||
|
||||
def test_next_is_the_first_open_step_after_this_one():
|
||||
steps = _steps("done", "in_progress", "done", "todo", "todo")
|
||||
assert _next_open(steps, current_id=2)["id"] == 4
|
||||
|
||||
|
||||
def test_next_falls_back_to_an_earlier_open_step():
|
||||
"""The last step finishing does not mean the milestone is finished."""
|
||||
steps = _steps("todo", "done", "done")
|
||||
assert _next_open(steps, current_id=3) == {"id": 1, "title": "step 1", "status": "todo"}
|
||||
|
||||
|
||||
def test_nothing_open_means_no_next():
|
||||
assert _next_open(_steps("done", "cancelled", "done"), current_id=3) is None
|
||||
|
||||
|
||||
async def test_a_task_with_no_project_or_milestone_has_no_placement_and_no_query():
|
||||
boom = AsyncMock(side_effect=AssertionError("must not open a session"))
|
||||
with patch.object(placement, "async_session", boom):
|
||||
task = SimpleNamespace(id=1, project_id=None, milestone_id=None)
|
||||
assert await placement.task_placement(7, task) is None
|
||||
|
||||
|
||||
async def test_attach_omits_placement_rather_than_sending_it_empty():
|
||||
with patch.object(placement, "task_placement", AsyncMock(return_value=None)):
|
||||
data = await placement.attach_placement(7, {"id": 1}, SimpleNamespace(id=1))
|
||||
assert "placement" not in data
|
||||
|
||||
|
||||
async def test_a_failing_lookup_returns_the_payload_untouched():
|
||||
with patch.object(placement, "task_placement", AsyncMock(side_effect=RuntimeError("db down"))):
|
||||
data = await placement.attach_placement(7, {"id": 1}, SimpleNamespace(id=1))
|
||||
assert data == {"id": 1}
|
||||
@@ -114,6 +114,10 @@ async def test_build_session_context_includes_project_when_scoped():
|
||||
# No design system on the project -> no design block at all. An install with
|
||||
# none is the ordinary case, not a degraded one.
|
||||
assert "## Design system" not in out["context"]
|
||||
# Live state only (decision #4027): how to work with Scribe is the
|
||||
# using-scribe skill's to say. A restated reflex here is a copy that drifts.
|
||||
assert 'content_type="rule"' not in out["context"]
|
||||
assert "Reflex:" not in out["context"]
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
@@ -246,14 +250,16 @@ async def test_build_session_context_caps_length():
|
||||
|
||||
Patching the cap rather than manufacturing 9,000 characters keeps the test
|
||||
about the TRUNCATION PATH — that it cuts, and that it says it cut — which
|
||||
is the part a reader depends on.
|
||||
is the part a reader depends on. The unbound-repo hint supplies the text:
|
||||
since milestone 410 the block carries live state only, and a bare session
|
||||
is a one-liner.
|
||||
"""
|
||||
from scribe.services import plugin_context as pc
|
||||
|
||||
with patch.object(pc, "_MAX_CHARS", 120):
|
||||
out = await pc.build_session_context(user_id=7)
|
||||
with patch.object(pc, "_MAX_CHARS", 60):
|
||||
out = await pc.build_session_context(user_id=7, unbound_repo="host/owner/repo")
|
||||
|
||||
assert len(out["context"]) <= 120 + 60 # cap + truncation note
|
||||
assert len(out["context"]) <= 60 + 20 # cap + truncation note
|
||||
assert "truncated" in out["context"], (
|
||||
"the block was cut without saying so — a reader cannot tell a "
|
||||
"truncated context from a short one"
|
||||
|
||||
@@ -0,0 +1,92 @@
|
||||
"""Guessed record ids are refused; `{{ref:N}}` placeholders resolve (#4016).
|
||||
|
||||
The defect: a session predicts the ids its next creates will get and writes
|
||||
them into a body before the records exist; any other session's create takes
|
||||
those numbers. These pin the recogniser — what counts as a `#N`, where a guess
|
||||
can sit relative to the highest id — and the placeholder rewrite the batch
|
||||
create uses instead of guessing.
|
||||
"""
|
||||
from unittest.mock import AsyncMock, patch
|
||||
|
||||
import pytest
|
||||
|
||||
from scribe.services import record_refs
|
||||
from scribe.services.record_refs import (
|
||||
GUESS_WINDOW,
|
||||
cited_ids,
|
||||
placeholder_keys,
|
||||
resolve_placeholders,
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.parametrize(("text", "expected"), [
|
||||
("see #12 and (#34)", {12, 34}),
|
||||
("#7 at the start", {7}),
|
||||
("an entity { is not a reference", set()),
|
||||
("a fragment page#12 is not one", set()),
|
||||
("a colour #123abc is not one", set()),
|
||||
("path/#9 is not one", set()),
|
||||
("", set()),
|
||||
])
|
||||
def test_cited_ids_reads_standalone_hash_numbers_only(text, expected):
|
||||
assert cited_ids(text) == expected
|
||||
|
||||
|
||||
def test_cited_ids_ignores_none_among_texts():
|
||||
assert cited_ids(None, "#5", None) == {5}
|
||||
|
||||
|
||||
async def test_a_number_just_above_the_highest_id_is_a_guess():
|
||||
with patch.object(record_refs, "_max_note_id", AsyncMock(return_value=100)):
|
||||
assert await record_refs.guessed_ids("next is #101, then #102") == [101, 102]
|
||||
|
||||
|
||||
async def test_existing_ids_and_far_numbers_are_not_guesses():
|
||||
"""Below the max is a real record; far above it is not a Scribe id at all
|
||||
(a PR, a forge issue) — neither may be refused."""
|
||||
far = 100 + GUESS_WINDOW + 1
|
||||
with patch.object(record_refs, "_max_note_id", AsyncMock(return_value=100)):
|
||||
assert await record_refs.guessed_ids(f"#1 #100 #{far}") == []
|
||||
|
||||
|
||||
async def test_the_window_edge_is_still_a_guess():
|
||||
edge = 100 + GUESS_WINDOW
|
||||
with patch.object(record_refs, "_max_note_id", AsyncMock(return_value=100)):
|
||||
assert await record_refs.guessed_ids(f"#{edge}") == [edge]
|
||||
|
||||
|
||||
async def test_text_without_references_never_queries():
|
||||
"""Most writes cite nothing; they must not pay a query for it."""
|
||||
probe = AsyncMock(return_value=100)
|
||||
with patch.object(record_refs, "_max_note_id", probe):
|
||||
await record_refs.refuse_guessed_ids("no numbers here", None, "")
|
||||
probe.assert_not_awaited()
|
||||
|
||||
|
||||
async def test_refusal_names_the_guess_and_the_fix():
|
||||
with patch.object(record_refs, "_max_note_id", AsyncMock(return_value=100)):
|
||||
with pytest.raises(ValueError) as exc:
|
||||
await record_refs.refuse_guessed_ids("steps #101 and #103")
|
||||
message = str(exc.value)
|
||||
assert "#101, #103" in message
|
||||
assert "create_records" in message and "{{ref:N}}" in message
|
||||
|
||||
|
||||
def test_placeholder_keys_tolerates_spacing():
|
||||
assert placeholder_keys("{{ref:1}} {{ ref: 2 }} {{ref:milestone}}") == {"1", "2", "milestone"}
|
||||
|
||||
|
||||
def test_resolve_placeholders_rewrites_every_occurrence():
|
||||
refs = {"1": '#40 "first"', "milestone": 'milestone 9 "plan"'}
|
||||
out = resolve_placeholders("{{ref:1}} then {{ref:1}} in {{ref:milestone}}", refs)
|
||||
assert out == '#40 "first" then #40 "first" in milestone 9 "plan"'
|
||||
|
||||
|
||||
def test_resolve_placeholders_passes_empty_through():
|
||||
assert resolve_placeholders(None, {}) is None
|
||||
assert resolve_placeholders("", {}) == ""
|
||||
|
||||
|
||||
def test_an_unknown_key_is_a_caller_bug_not_a_silent_leftover():
|
||||
with pytest.raises(KeyError):
|
||||
resolve_placeholders("{{ref:3}}", {"1": "#1"})
|
||||
Reference in New Issue
Block a user