Plugin reply shape, preferences UI, cited-record status, and the lesson kind #166

Merged
bvandeusen merged 7 commits from dev into main 2026-09-18 16:33:57 -04:00
Owner

Seven commits. The reason for merging NOW is the plugin: the marketplace checkout tracks main, so the two plugin files on dev cannot reach a session until this lands. Everything else here was already testable on :dev.

What this unblocks

plugin/skills/reporting-back/SKILL.md and the manifest (2026.09.18.04362026.09.18.1606). The version moves, so the executing-cache refresh (#2209) actually fires rather than silently keeping the old copy.

  • #4153 — a reply's sections are chosen, not filled. Sections that answer a standing question ("needs you") are always answered; sections that explain earn their place only when they change what the operator does. Plus: write the shortest reply that carries the answer.
  • #4154 — a record you only mention is a record to read.

After merging, a /plugin refresh or a fresh session picks it up. Nothing else is required.

What rides along

#4154 — a record you merely CITE carries its status. next_step on every milestone summary (one flat query per batch, not per milestone — #2384's fan-out), and a task's status in the injected line's kind marker. This is the gap that let #4014 be reported as milestone 409's open step when it had been done for four days.

#3895 — preferences are writable, and their drift is visible. GET /rules/drift, a restore route carved out for preferences only, the preference chip, and PreferenceDriftPane. Milestone 323 refused one-click restore for rules because a binding instruction should not be revertible in one click; the carve-out turns on that reasoning being about the operator's own decision, where a preference's rewrite is the agent's. Nothing is erased — the restore goes through update_rule and snapshots.

#3729 / #3730 — the lesson kind, steps 2 and 3 of milestone 385. Half-built on purpose, and inert on arrival:

  • no create_lesson exists yet (step 4), and the MCP create_note tool hardcodes note_type="note", so no agent surface can write one;
  • include_global_kinds is wired only to the explicit MCP search, which returns nothing because the corpus holds no lessons;
  • the missing browse chip in KnowledgeView.vue (step 7, #3734) stays unreachable until a lesson can exist.

Merging before step 4 rather than after is deliberate for exactly that reason.

No migration in this batch — note_type carries no CHECK constraint, so rule 36 has nothing to expand.

Verification

CI 7023 green on 1361ed7, all six jobs including the integration lane against real Postgres. Runs 7016 and 7005 went red first and are fixed in 7017 and 7006; both were guards working as intended.

🤖 Generated with Claude Code

https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy

Seven commits. The reason for merging NOW is the plugin: the marketplace checkout tracks `main`, so the two plugin files on `dev` cannot reach a session until this lands. Everything else here was already testable on `:dev`. ## What this unblocks `plugin/skills/reporting-back/SKILL.md` and the manifest (`2026.09.18.0436` → `2026.09.18.1606`). The version moves, so the executing-cache refresh (#2209) actually fires rather than silently keeping the old copy. - **#4153** — a reply's sections are chosen, not filled. Sections that answer a standing question ("needs you") are always answered; sections that explain earn their place only when they change what the operator does. Plus: write the shortest reply that carries the answer. - **#4154** — a record you only mention is a record to read. After merging, a `/plugin` refresh or a fresh session picks it up. Nothing else is required. ## What rides along **#4154 — a record you merely CITE carries its status.** `next_step` on every milestone summary (one flat query per batch, not per milestone — #2384's fan-out), and a task's status in the injected line's kind marker. This is the gap that let #4014 be reported as milestone 409's open step when it had been done for four days. **#3895 — preferences are writable, and their drift is visible.** `GET /rules/drift`, a restore route carved out for preferences only, the preference chip, and `PreferenceDriftPane`. Milestone 323 refused one-click restore for rules because a binding instruction should not be revertible in one click; the carve-out turns on that reasoning being about the operator's own decision, where a preference's rewrite is the agent's. Nothing is erased — the restore goes through `update_rule` and snapshots. **#3729 / #3730 — the lesson kind, steps 2 and 3 of milestone 385.** Half-built on purpose, and inert on arrival: - no `create_lesson` exists yet (step 4), and the MCP `create_note` tool hardcodes `note_type="note"`, so no agent surface can write one; - `include_global_kinds` is wired only to the explicit MCP search, which returns nothing because the corpus holds no lessons; - the missing browse chip in `KnowledgeView.vue` (step 7, #3734) stays unreachable until a lesson can exist. Merging before step 4 rather than after is deliberate for exactly that reason. No migration in this batch — `note_type` carries no CHECK constraint, so rule 36 has nothing to expand. ## Verification CI 7023 green on `1361ed7`, all six jobs including the integration lane against real Postgres. Runs 7016 and 7005 went red first and are fixed in 7017 and 7006; both were guards working as intended. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
bvandeusen added 7 commits 2026-09-18 16:33:48 -04:00
feat(plugin): a reply's sections are chosen, not filled (#4153)
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / Python lint (push) Successful in 4s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 49s
CI & Build / Python tests (push) Successful in 1m30s
CI & Build / Build & push image (push) Successful in 16s
b5df9d6dca
Milestone 409 step 6 measured the scaffold on live sessions and found its
two halves disagreeing: adherence passed and the read test failed.
Completion replies carried every section the table asks for and were
still hard to read.

The cause was in the skill, not in compliance with it. It said to pick a
kind of reply "then fill its sections... keep them even when one is
short", which is an instruction to complete a form, and nothing anywhere
set a ceiling. A faithful reply and an unreadable one were the same
reply.

Four changes to the discipline around the scaffold. The categories and
their sections are untouched.

- Sections are what to consider including, not a form to complete. A
  section answering a standing question ("does anything need me?") is
  always answered, even with "nothing"; a section that explains earns its
  place only when it changes what the operator does. Otherwise it belongs
  in the record's log, where it is available and not in the way.
- Write the shortest reply that carries the answer, with named exceptions
  so this cannot be read as "always be terse".
- "Needs you" takes BOTH tests: theirs to decide, AND work is waiting on
  it. A question answerable by reading something or taking an available
  measurement is work not yet done, not a request — settle it, say which
  way you went, and leave them free to overrule.
- A decision already made gets acted on. Re-arguing a settled question
  reads as contradicting yourself rather than as being careful, and costs
  the operator the decision twice.

"Before sending" gains a second pass for what can go, since the existing
check asks what is MISSING, which a bloated reply passes.

Guards in tests/test_reply_discipline.py, three topics registered for
ownership. Every guard was falsified against the pre-change text before
committing (rule 167): all five fail on it and pass on the fix, and the
sixth deliberately passes both since it guards the scaffold against
collateral damage. No absence checks — the skill legitimately discusses
filling in order to warn against it, so asserting "fill" is absent would
false-alarm on the corrected text (snippet #3352).

Instance-agnostic per rule 115: the added text carries no record ids, no
software-specific terms and no verbatim quotes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
feat(placement): a record you only cite carries its status (#4154)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 8s
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Failing after 1m4s
CI & Build / Build & push image (push) Skipped
5c6175ad97
Step 1 made placement cheap for a task whose status CHANGES: create_task
and update_task return where it sits, and the report is written from
that. It did nothing for a task a reply merely cites.

This milestone's own step-6 review reported "#4014 is the open step of
milestone 409". #4014 had been done for four days; the open step was
#4015. The id did not come from a read — it came from a retrieval hint,
which carries an id, a kind and a title and says nothing about status,
while list_milestones said "8 of 9" and would not say which one. The
gap was there to be filled and the nearest-looking id filled it.

Two surfaces, one principle: the status arrives with the id.

1. get_project_milestone_summaries gains next_step — the earliest open
   step, {id, title, status} or None — carried through _BRIEF_FIELDS to
   enter_project, get_project and list_milestones. One extra flat query
   for the whole batch, so #2384's fan-out does not come back.

   OPEN_STEP_STATUSES moves to services/milestones.py and placement.py
   imports it; both surfaces now answer "what is next" and must not
   drift on what counts as open. Both step queries take the same
   readable_notes_clause (rule 78), so a row cannot name a step its own
   progress numbers exclude.

2. _record_kind renders a task's status: [task (done)], [issue (todo)].
   A finished step and an open one read identically before, which is
   exactly the line the misreport was taken from. Only tasks — is_task
   IS status-is-not-None on the model, so there is no fallback branch.

reporting-back gains the practice, owned and registered in the guidance
ownership table: a record you only mention is a record to read.

The guards are structural and each fails on the regression it names:
the query count is asserted rather than the payload shape, and the two
surfaces' agreement is pinned on the rendered ORDER BY, since a mocked
session hands back whatever order the test chose.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
fix(tests): the counts query selects four columns, not three (#4154)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / integration (push) Successful in 40s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m36s
CI & Build / Build & push image (push) Successful in 23s
94ecb633a0
The new fixtures fed (milestone_id, status, count) and the query also
carries max(updated_at), which last_touched_at is computed from. Six
tests in the new module died unpacking it; the product path was never
reached, and the rest of the suite was green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
feat(rules): preferences are writable, and their drift arrives (#3895)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / integration (push) Successful in 52s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m32s
CI & Build / Build & push image (push) Successful in 34s
7038e41ec7
Milestone 399 step 5. Steps 1-4 put preferences into the backend: a kind
column, an inverted write path, a third register in the injected block, a
delivery slot. Nothing the operator could touch. Rule 27 forbids leaving it
there, and here it matters more than usual, because the UI is the only
guard against the risk the milestone named up front — an agent misreads one
session, rewrites a preference, and follows the rewritten version forever
while the operator never sees the moment it changed.

Four things ship.

A preference is DISTINGUISHABLE. `kind` reaches the client (the server has
always sent it in rule_brief) and a preference carries a chip. Force is the
one thing a list of instructions must not leave the reader to infer, and a
row that renders identically to a rule teaches the opposite of both facts
about a preference: it does not bind, and a session may rewrite it.

A preference is WRITABLE. The editor gains the kind as a first-class choice
with the test beside it — what happens when someone does not do this — and
says plainly, when preference is chosen, that sessions rewrite these
without asking and every rewrite is kept.

DRIFT ARRIVES. `GET /api/rules/drift` returns one row per rewritten
preference carrying its latest rewrite: what it said, what it says now, and
the record named by `arose_from_id` that taught the change. Both texts ride
along so the list shows the diff without a call per row. The new pane sits
beside the staleness sweep, because drift belongs to no one rulebook, and
it answers a question the operator would not have thought to ask.

REVERSION IS ONE ACTION, and this is the carve-out worth arguing with.
Milestone 323 refused a one-click restore for rules — "a binding
instruction should not be revertible in one click", because a silent revert
erases the only record of why the rewrite happened. That reasoning turns on
the rewrite being the operator's own decision. A preference's is not: the
agent makes it mid-work without asking, so reverting is a veto over someone
else's edit rather than an undo of your own, and a veto costing more than a
shrug is not supervision. The route refuses anything but a preference (409),
and nothing is erased: the restore goes through update_rule, so it snapshots
too and the history GAINS the revert. An integration test pins that, because
it is the whole basis for the exception.

Tested against real Postgres — every claim is about which rows come back
and in what order, which a stand-in session cannot judge.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
feat(lessons): the lesson kind, and one join for every trigger title (#3729)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Python tests (push) Failing after 1m5s
CI & Build / Build & push image (push) Skipped
0ab15d7d80
Milestone 385 step 2, implementing decision #4157 from the step-1 spike.

The kind: `note_type='lesson'`, a note findable by WHEN IT APPLIES rather
than by what it is about. The trigger lives in `notes.data.when_to_apply`,
mirrored into the title and the head of the body — the shape snippets
already use, and the reason nothing re-embeds: chunk_document is untouched,
so CHUNKER_VERSION does not move.

NO MIGRATION, and the step assumed there would be one. `note_type` carries
no CHECK — only `task_kind` does (0056, 0065). Migration 0036 added it as
plain Text with a server default and nothing has gated it since, so rule 36
has no whitelist to expand and #3128's failure mode (a value the database
refuses) cannot arise for this column. The vocabulary that actually decides
what a reader can reach is services.knowledge._FACETS, which since #3161 is
one table feeding the door's validation, the counts and both dialects of
the type filter — so the kind lands there in a single edit.

ONE JOIN, not a fourth copy. `{subject} — {trigger}` had three
implementations: rule_document, snippets.compose_title, and this step
needed another. #3207 records what that costs, so the join moves to
embeddings.trigger_title beside embedding_text and all three delegate.
Behaviour is unchanged for rules and snippets; the guard calls each through
its own public name, so a re-implementation fails it.

The #3163 bill is stated in the service docstring rather than left to be
inferred: versions, supersession, trash, the share ACL, tags, project and
System tagging, chunked embeddings and the duplicate gate are all
inherited; status/task_kind/milestone_id and recurrence are not, and
verify_with/expires_when are available but outside the kind's contract.
The status cell is the one that matters — `is_task` IS `status is not
None`, so a lesson that acquired one would become a task.

The integration guard asserts the WRITE rather than the constraint: it
holds whether or not note_type is ever gated, and goes red only if it is
gated without this value.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
fix(tests): the browse vocabulary guard names the fourth kind (#3729)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / integration (push) Successful in 42s
CI & Build / TypeScript typecheck (push) Successful in 56s
CI & Build / Python tests (push) Successful in 1m37s
CI & Build / Build & push image (push) Successful in 26s
d127d48c14
CI 7016. `test_non_task_facets_are_the_note_types_and_only_those` pins the
non-task vocabulary as a literal set, so adding `lesson` to `_FACETS` turned
it red — the guard working, not breaking. It stays a literal: derived from
_FACETS it would assert nothing, and rule 167 wants a guard that can fail.

The representative corpus had no lesson row, so `lesson` was reaching only
the two tests that iterate FACET_TYPES and never the one that asserts each
facet selects EXACTLY its own rows. It has one now, which is what pins the
half that matters: a lesson is not picked up by the Notes facet despite
both being non-task records.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
feat(lessons): the document shape is the stored record, and it travels (#3730)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / integration (push) Successful in 51s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / Python tests (push) Successful in 1m31s
CI & Build / Build & push image (push) Successful in 25s
1361ed7200
Milestone 385 step 3 — the step where the kind either works or is cosmetic.

THE DOCUMENT, and why there is no `lesson_document()` beside `rule_document()`
in embeddings. The step expected one. The difference is where the sharp shape
LIVES. A rule keeps its trigger in a column and its title is a plain name, so
`{title} — {trigger}` has to be synthesised at embed time and exists nowhere
else. A snippet — the only sharp record in the corpus by #2485's measurement,
0.153 top-to-second against 0.010–0.023 — gets there the other way: its STORED
title is already the join and its stored body already opens with the trigger,
so the ordinary `title\nbody` join IS the sharp document. Step 1 chose the
snippet route and step 2 built it, so `lessons.lesson_document` composes what
is STORED and the generic chunker does the rest.

The consequence the step asked about: `chunk_document` is untouched, so
CHUNKER_VERSION does not move and NOTHING re-embeds. The step's "Re-embed"
section describes a change this design does not make.

THE NARRATIVE stays in the body, departing from the step's instruction to keep
it out. `rule_document` excludes `why` because long dated narrative made
sixteen dev-logs land on the centroid of "development" — but that finding
predates chunking (#280). A body over budget is now split, and every chunk is
prefixed with the title, which for a lesson carries the trigger. The story
occupies its own vectors instead of averaging itself into the trigger's, and
each of those is still anchored to when the lesson applies. A guard asserts
exactly that. Holding the story out would cost the reader the only part that
explains the insight, to buy a sharpness the chunker already provides.

GLOBAL IN THE SEARCH is the real new code: `GLOBAL_NOTE_TYPES` and
`include_global_kinds` on `semantic_search_notes`, widening the PROJECT filter
alone. Off by default, because two callers depend on that filter holding — the
near-duplicate gate compares a record only against its own project on purpose,
and a globally visible kind there would let a lesson block an unrelated note's
create on a project its author never touched. It composes with `note_type`
rather than overriding it, so narrowing to snippets does not quietly acquire
lessons, and it changes nothing about the ACL: `notes_visibility_clause` still
gates every row.

Wired into the explicit MCP search only — the operator asked, and there is no
budget to spend. The unasked-for injection arms are step 5's subject (#3732)
and the legibility of a lesson appearing on a foreign project is step 7's
(#3734), so neither is turned on here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
bvandeusen merged commit adeaec88c5 into main 2026-09-18 16:33:57 -04:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: bvandeusen/FabledScribe#166