feat(shapes): the agent judges what it wrote, at the end of the turn (milestone 439 steps 1-3)
CI & Build / Python lint (push) Successful in 5s
CI & Build / Plugin hooks (push) Successful in 18s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m38s
CI & Build / Build & push image (push) Successful in 33s

Recording used to be decided by machinery — the only "record it" prompt
fired when a same-named copy already existed (#2664), so a first instance of
a reusable piece was never asked about, and judgment arrived only through
audits. Now the question is asked where the knowledge is: the end of the
turn that wrote the code, of the agent that wrote it.

- Write hooks keep `<sid>.written.ids` (path, kind, name) for every
  definition a write names; a new file adds a `file` line for its stem — a
  candidate in any language without a framework rule (scribe_written_append).
- Stop hook scribe_shape_check.sh sends the ledger to GET
  /api/plugin/shape-check and blocks once, in the server's words, when
  anything is unjudged. Same discipline as the report check: never twice,
  never without a recorded check, another hook's loop left alone; the ledger
  is kept when the instance cannot be reached.
- shape_ledger.unjudged_shapes: no row, unclassified, scoped and hook stamps
  are unjudged; an agent/audit/import verdict is not. A snippet recorded at
  the shape answers for it until the refresh stamps it canonical.
- services/shape_check owns the reason text and records every outcome in
  app_logs (passed / blocked / judged_after_block / left_after_block).
- classify_shapes(repo=…) judges a shape the ledger has not synced yet via a
  provisional row under a bound repo; the sync confirms it, or vanishes and
  revives it with the verdict intact. An unbound repo is refused.

Plugin version minted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-10-01 08:39:36 -04:00
co-authored by Claude Opus 5.5
parent eb5cc6d3a7
commit f1fbdf746a
17 changed files with 854 additions and 9 deletions
+18
View File
@@ -122,6 +122,24 @@ async def bindings_for_project(user_id: int, project_id: int) -> list[RepoBindin
return list(rows.scalars().all())
async def is_bound(project_id: int, raw_repo: str) -> str:
"""The normalized key when ``raw_repo`` is bound to ``project_id`` by
anyone, else "". By anyone, not by the caller: a collaborator judging a
shared project's shapes holds no binding of their own, and the question
is whether the repo belongs to the project, not who bound it."""
key = normalize_repo_key(raw_repo or "")
if not key or not project_id:
return ""
async with async_session() as session:
found = await session.execute(
select(RepoBinding.id).where(
RepoBinding.project_id == project_id,
RepoBinding.repo_key == key,
).limit(1)
)
return key if found.first() is not None else ""
async def keys_for_project(user_id: int, project_id: int) -> list[str]:
"""Every repo key bound to a project — the snippet→forge join (#2691).
+147
View File
@@ -0,0 +1,147 @@
"""The end-of-turn shape check: what a turn wrote that nobody has judged, and what the agent is asked.
WHY THIS EXISTS (milestone 439)
The shape ledger used to be filled by machinery: extraction found the
definitions, framework rules decided which were one-offs, and the write-path
hook stamped "instance of #N" with nobody reading. The agent that WROTE the
code — the one participant who knows what it is — was asked only in an audit,
weeks later, if anyone ran one. A first instance of a reusable piece was never
asked about at all, so it was never recorded, and the next session rebuilt it.
So the question moves to the moment the knowledge exists. The client's write
hooks keep a ledger of the definitions each turn wrote; its Stop hook sends
them here at the end of the turn. Two jobs live on this side, as with the
report check (services/report_check.py):
- DECIDING what is unjudged, against the ledger and the project's recorded
snippets — something only the server can see.
- OWNING THE WORDS. A hook carries timing and transport (plugin/PACKAGING.md);
the instruction comes from here, so every client asks the same question
and the wording changes in one place.
And it RECORDS every outcome, so "how much did agents leave unjudged after
being asked" is a number rather than an impression. app_logs, for the reason
report_check gives.
"""
from __future__ import annotations
import json
from scribe.models import async_session
from scribe.models.app_log import AppLog
# check — the turn's first stop: block if anything is unjudged.
# after — the stop that follows a block: never block again, only record
# whether the agent answered.
PHASES = ("check", "after")
OUTCOMES = ("passed", "blocked", "judged_after_block", "left_after_block")
# How many shapes the reason names before "+N more". Enough to judge a normal
# turn in one call; a turn that wrote more than this generated something.
_LISTED = 25
def _evidence_note(evidence: dict) -> str:
"""One parenthesis of what the machinery holds — offered, never asserted."""
parts = []
looks = evidence.get("looks_like") or {}
if looks.get("snippet_id"):
score = looks.get("score")
at = f" {score:.2f}" if isinstance(score, (int, float)) else ""
parts.append(f"looks like #{looks['snippet_id']}{at}")
stamped = evidence.get("stamped") or {}
if stamped.get("snippet_id"):
parts.append(f"the hook guessed #{stamped['snippet_id']}")
if evidence.get("copies"):
parts.append("other copies of it exist")
if evidence.get("diverges_from"):
parts.append(f"#{evidence['diverges_from']} is canon in this directory")
return f" ({'; '.join(parts)})" if parts else ""
def block_reason(unjudged: list[dict], *, project_id: int, repo: str) -> str:
"""What the agent is told at the end of a turn that wrote unjudged shapes.
Phrased as the practice wanted (note #3565): what to decide and the one
call that records it, not a prohibition. The reusing-code skill owns the
longer version of what each verdict means.
"""
lines = []
for item in unjudged[:_LISTED]:
new = " — new" if item.get("status") == "new" else ""
lines.append(
f"- {item['path']} · {item['symbol']} ({item['kind']}){new}"
f"{_evidence_note(item.get('evidence') or {})}"
)
more = len(unjudged) - _LISTED
if more > 0:
lines.append(f"- … and {more} more (list_shapes(project_id={project_id}, path=…) shows them)")
repo_arg = f', repo="{repo}"' if repo else ""
return (
f"This turn wrote {len(unjudged)} definition{'s' if len(unjudged) != 1 else ''} "
"that nobody has judged yet. You built them, so you know what each one is — "
"record it now, while that is still true, so the next session starts from your "
"answer instead of rebuilding it:\n"
+ "\n".join(lines)
+ "\n\nFor each: something another part of the code should reuse → record it "
"with create_snippet (name, code, when to reach for it, location); built from a "
"recorded snippet → `instance` of it; a deliberate departure from one → "
"`variant` with the why; a genuine one-off → `exempt` with a reason "
"(reason_code one-off-handler, test-helper, scoped-css, pure-helper … indexes it). "
"One call records them all: "
f"classify_shapes(project_id={project_id}{repo_arg}, classifications=[{{path, "
"symbol, kind, status, snippet_id?, reason?, reason_code?}}, …]). "
"`repo` lets a verdict land on a shape the ledger has not synced yet."
)
async def record_shape_check(
user_id: int | None,
outcome: str,
*,
written: int,
unjudged: int,
project_id: int | None = None,
) -> None:
if outcome not in OUTCOMES:
raise ValueError(f"unknown shape-check outcome {outcome!r}")
details: dict = {"outcome": outcome, "written": written, "unjudged": unjudged}
if project_id:
details["project_id"] = project_id
async with async_session() as session:
session.add(AppLog(
category="plugin",
user_id=user_id,
action="shape_check",
details=json.dumps(details),
))
await session.commit()
def outcome_for(phase: str, unjudged: list[dict]) -> str:
"""The recorded outcome: a first stop blocks or passes; the stop after a
block records whether the agent answered, and never blocks."""
if phase == "after":
return "left_after_block" if unjudged else "judged_after_block"
return "blocked" if unjudged else "passed"
def parse_written(raw: str, *, cap: int = 200) -> list[tuple[str, str, str]]:
"""The hook's ledger lines, `path<TAB>kind<TAB>name`, as triples.
Kinds outside the ledger's vocabulary and malformed lines are dropped
rather than echoed into an instruction; duplicates collapse; capped."""
out: list[tuple[str, str, str]] = []
for line in (raw or "").splitlines():
parts = line.split("\t")
if len(parts) != 3:
continue
path, kind, name = (p.strip() for p in parts)
if kind not in ("css", "sym", "file") or not path or not name:
continue
if (path, kind, name) not in out:
out.append((path, kind, name))
if len(out) >= cap:
break
return out
+114 -1
View File
@@ -535,6 +535,7 @@ async def classify_shapes(
classifications: list[dict],
*,
via: str = "agent",
repo: str = "",
) -> dict:
"""Apply a batch of judgments to a project's live ledger rows.
@@ -544,7 +545,19 @@ async def classify_shapes(
item carries one — and a target no live row matches is reported in
``unmatched``, not an error: the tree may simply have moved since the
caller listed. Idempotent by construction.
``repo`` — a repo bound to the project — lets a judgment land on a shape
the ledger has not synced yet (milestone 439): the shape the agent wrote
this turn. Such an item becomes a PROVISIONAL row under that repo, the
same row the write-path hook has always created, with no seen commit;
the next sync confirms it, or marks it vanished until the code reaches
the bound ref, and either way the judgment rides along (a vanished row
that reappears keeps its status). Without ``repo`` an unknown shape is
`unmatched`, exactly as before. An unbound ``repo`` is refused rather
than ignored: a verdict silently dropped reads as one recorded.
"""
from scribe.services import repo_bindings as repo_bindings_svc
from scribe.services import access
from scribe.services import snippets as snippets_svc
@@ -569,9 +582,18 @@ async def classify_shapes(
for sid in sorted(target_ids):
if await snippets_svc.get_snippet(user_id, sid) is None:
raise ValueError(f"snippet {sid} not found (or not readable)")
repo_key = ""
if (repo or "").strip():
repo_key = await repo_bindings_svc.is_bound(project_id, repo)
if not repo_key:
raise ValueError(
f"repo {repo!r} is not bound to project {project_id} — "
"bind_repo it, or omit repo to judge synced shapes only"
)
now = datetime.now(timezone.utc)
classified = 0
provisional = 0
unmatched: list[dict] = []
async with async_session() as session:
rows = (
@@ -592,6 +614,17 @@ async def classify_shapes(
kind = (item.get("kind") or "").strip()
if kind:
matches = [r for r in matches if r.kind == kind]
if not matches and repo_key and item.get("status") != "unclassified":
path = (item.get("path") or "").strip()
symbol = (item.get("symbol") or "").strip()
row = CodeShape(
project_id=project_id, repo_key=repo_key, path=path,
symbol=symbol, kind=kind or snippet_kind(symbol, ""),
)
session.add(row)
by_key.setdefault((path, symbol), []).append(row)
matches = [row]
provisional += 1
if not matches:
unmatched.append({
"path": item.get("path"), "symbol": item.get("symbol"),
@@ -610,7 +643,10 @@ async def classify_shapes(
evidence=item.get("reason"))
classified += 1
await session.commit()
return {"classified": classified, "unmatched": unmatched}
out = {"classified": classified, "unmatched": unmatched}
if provisional:
out["provisional"] = provisional
return out
def rule_matches(row: CodeShape, *, path: str, pattern: str, kind: str) -> bool:
@@ -2550,6 +2586,83 @@ async def flag_divergence(project_id: int, *, since: datetime | None) -> int:
return flagged
# --- the end-of-turn question (milestone 439): what did this turn write that
# nobody has judged? ---------------------------------------------------------
#
# The write hooks keep a session ledger of the definitions each turn wrote; at
# the end of the turn the client asks this, and the agent that wrote them
# answers. "Judged" means a verdict SOMEONE READ: an agent's, an audit's, an
# import's. The sync's `scoped` stamp, `unclassified`, a hook stamp and no row
# at all are all unjudged — each is a machine's guess or nothing, and the
# point of the question is that the writer, who knows, says.
def _is_judged(row) -> bool:
return (
row.status not in _MECHANICAL_TODO
and row.classified_by not in _UNATTENDED_BY
)
async def unjudged_shapes(
project_id: int, written: list[tuple[str, str, str]], *, repo_key: str = "",
) -> list[dict]:
"""The (path, kind, name) shapes in ``written`` that carry no judgment.
Each comes back as {path, kind, symbol, status, evidence} — status is the
row's, or "new" when the ledger has none yet; evidence is what the
machinery holds that the writer may want to weigh: the proposer's
candidate canon, a hook's stamp, a divergence flag. Evidence, never a
verdict — the writer is asked because the machine's guesses were the
thing that poisoned two ledgers. Order follows ``written``; duplicates
collapse. The caller has already resolved the project and checked access.
"""
wanted: list[tuple[str, str, str]] = []
for path, kind, name in written:
key = ((path or "").strip(), (kind or "").strip(), (name or "").strip())
if all(key) and key not in wanted:
wanted.append(key)
if not project_id or not wanted:
return []
paths = sorted({p for p, _k, _n in wanted})
query = select(CodeShape).where(
CodeShape.project_id == project_id,
CodeShape.vanished_at.is_(None),
CodeShape.path.in_(paths),
)
if repo_key:
query = query.where(CodeShape.repo_key == repo_key)
async with async_session() as session:
rows = (await session.execute(query)).scalars().all()
by_key = {(r.path, r.kind, r.symbol): r for r in rows}
out: list[dict] = []
for path, kind, name in wanted:
row = by_key.get((path, kind, name))
if row is not None and _is_judged(row):
continue
evidence: dict = {}
if row is not None:
if row.proposed_snippet_id:
evidence["looks_like"] = {
"snippet_id": row.proposed_snippet_id,
"basis": row.proposal_basis,
"score": row.proposal_score,
}
if row.classified_by in _UNATTENDED_BY and row.snippet_id:
evidence["stamped"] = {
"snippet_id": row.snippet_id, "score": stamp_score(row.reason),
}
if row.proposal_group:
evidence["copies"] = row.proposal_group
if row.diverges_from:
evidence["diverges_from"] = row.diverges_from
out.append({
"path": path, "kind": kind, "symbol": name,
"status": row.status if row is not None else "new",
"evidence": evidence,
})
return out
# How many of a canon's rows a review listing shows before "…" — enough to
# judge from, not the whole table.
_REVIEW_ROWS_SHOWN = 12