fix(ledger): a canon may hold a class and the to_dict beside it (#4220)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m41s
CI & Build / Build & push image (push) Successful in 28s

The review surface shipped yesterday reported two canons on its first live
day and both were sound. Coherence was "do all judged rows share a form",
which #2844 failed at 37/62 = 0.597 for containing a model class and the
to_dict the canon's own text says the class must carry, and #2849 failed at
4/7 for pairing sync loop-starters with the async ticks they schedule. A
review surface whose whole output is noise is one that stops being read.

The obvious repair is a trap, and there is now a test standing in front of
it. Grouping by family and keeping a majority test makes the check BLIND:
before #2844 was cleaned by hand it held 37 classes and 56 callables, which
as families is 56/93 = 0.602 — a clean pass, and the 31 rows that had no
business being there (Vue functions, route handlers, a dozen tests) would
never have been reported at all. A looser bar in the same shape is worse
than the bug.

So the verdict is inverted. Instead of asking whether most rows agree, it
asks how many rows the canon CANNOT ACCOUNT FOR: a row in the majority
family is accounted for; a callable defined in a file that also holds a
majority-family `type` row is a method of a member, not a foreign body; and
strangers above a fifth of the readable rows make the canon incoherent. The
majority vote abstains those methods, so a class's own serialisers cannot
outvote the classes and turn the members into the strangers.

Measured on the real ledger before it was written, which is why it is this
rule and not a nudge to the share: clean #2844 has 0 strangers in 62,
#2849 has 0 in 7, and polluted #2844 had 31 in 93 — the same 31 withdrawn
by hand this morning, named exactly.

The entry now carries `families`, the majority `family`, `attached`,
`stranger_count`, `unattended`, and `strangers` — THE ROWS THAT DO NOT FIT,
replacing a sample of the first twelve members. The reader's question is
which rows are wrong, and a sample of the agreeing majority cannot answer
it. `unattended` is the discriminator between a check that is too strict and
a ledger full of junk: both canons flagged on day one were entirely
audit-judged, and nothing showed that without opening each one.

Scoped to the review surface. `canon_form` still answers at the precise form
level for stamping and divergence, where a sync helper beside an async canon
is a fair question; nothing here changes what the ledger writes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
2026-09-20 23:39:01 -04:00
co-authored by Claude Opus 5
parent e87bcfa48c
commit 0fe19a8440
3 changed files with 341 additions and 33 deletions
+15 -5
View File
@@ -267,13 +267,23 @@ async def stamps_to_review(project_id: int, top: int = 10) -> dict:
score, signature and derived form so you can judge rather than trust the score, signature and derived form so you can judge rather than trust the
threshold. threshold.
`incoherent` — canons whose judged rows do not agree on a form: the `incoherent` — canons whose membership no longer agrees what the canon
`forms` histogram, how many members were hook-stamped weakly, and a is. A canon asserts that some shapes are the same sort of thing; when its
sample. A canon asserts that some shapes are the same sort of thing; when members are a class, three getters and a dozen tests, that assertion has
its members are a class, three getters and a dozen tests, that assertion stopped being true and every base-rate reading built on it — the
has stopped being true and every base-rate reading built on it — the
divergence prompt included — is reading noise. divergence prompt included — is reading noise.
Disagreement is judged by FAMILY (a sync helper and an async one are one
family), and a method defined beside a member class counts as part of the
shape rather than against it — so a model convention that covers both the
class and its `to_dict` reads as coherent, which it is (#4220). Each
entry carries `families` and `forms`, the majority `family`, `attached`
(minority rows excused as methods of a member), `strangers` — THE ROWS
THAT DO NOT FIT, which is what you judge — and `unattended` / `weak_rows`,
the count written by the hook with nobody reading. An incoherence whose
rows are all judged is likelier this check being strict than a bad
ledger; one made of unattended rows is the real thing.
A row listed here is a question, not a verdict. Some will be correct. A row listed here is a question, not a verdict. Some will be correct.
Read-only; requires read access to the project. Read-only; requires read access to the project.
+166 -21
View File
@@ -918,6 +918,12 @@ FORM_UNKNOWN = ""
# a canon poisoned by loose stamping falls silent instead of flagging. # a canon poisoned by loose stamping falls silent instead of flagging.
_FORM_SHARE = 0.6 _FORM_SHARE = 0.6
# How much of a canon's membership may be shapes it cannot account for
# before the canon stops meaning anything. Not a majority test — see
# `canon_coherence` for why a majority test goes blind exactly when the
# ledger is worst.
_STRANGER_SHARE = 0.2
# Leading words that say nothing about the form of the thing being declared. # Leading words that say nothing about the form of the thing being declared.
_FORM_NOISE = ("export ", "default ", "public ", "private ", "static ", "final ") _FORM_NOISE = ("export ", "default ", "public ", "private ", "static ", "final ")
@@ -1069,6 +1075,124 @@ def canon_form(rows: Iterable, snippet_id: int) -> str:
return top if n / sum(forms.values()) >= _FORM_SHARE else FORM_UNKNOWN return top if n / sum(forms.values()) >= _FORM_SHARE else FORM_UNKNOWN
# A judgment nobody read before it was written. "hook" is the write-path
# stamp — a similarity score and no reader. Every other value ("agent",
# "audit", "import", "mechanical") came from something that examined the
# shape and said so, which is a different kind of evidence, not a stronger
# score.
_UNATTENDED_BY = ("hook",)
def canon_coherence(members: Iterable) -> dict:
"""Does a canon's membership still agree on what the canon IS?
The verdict is NOT "do all the rows share a form". That was the first
version, and on this surface's first live day both canons it reported
were sound (#4220) while the one genuinely polluted canon had already
been cleaned by hand. Two things were wrong with it, and the fix for the
first nearly broke the second:
1. FAMILY, NOT FORM. A sync loop-starter and the async tick it schedules
are one shape written two ways; `shape_family` already collapses `fn`
and `async-fn` into `callable` for exactly this reason on the
divergence side, and this side never asked. Snippet #2849 — four
starters, three ticks, every row judged by an audit — scored 4/7 and
was called incoherent when nothing about it is.
2. A METHOD OF A MEMBER IS NOT A STRANGER. Snippet #2844 is the
SQLAlchemy model convention and its own text covers the class AND the
`to_dict` the class must carry. Its 37 classes and 25 serialisers
scored 37/62 = 0.597, missing the bar by three thousandths for
containing exactly what it says it contains.
THE TRAP, and it is why this is not simply a family histogram: grouping
by family and keeping a majority test makes the check BLIND. Before #2844
was cleaned it held 37 classes and 56 callables; as families that is
56/93 = 0.602, a clean pass, and the 31 rows that had no business being
there — Vue functions, route handlers, a dozen tests — would never have
been reported. A looser bar in the same shape is not a fix.
So the verdict is inverted. Instead of asking whether most rows agree, it
asks how many rows the canon CANNOT ACCOUNT FOR:
* a row in the majority family is accounted for;
* a callable defined in a FILE that also holds a majority-family `type`
row is accounted for — it is a method of a member;
* everything else is a STRANGER, and strangers above `_STRANGER_SHARE`
of the readable rows make the canon incoherent.
The majority vote abstains those methods, so a class's own serialisers
cannot outvote the classes and turn the members into the strangers. On
the real ledger the three populations separate completely: clean #2844
has 0 strangers in 62, #2849 has 0 in 7, and polluted #2844 had 31 in 93.
Scoped to the REVIEW surface on purpose. `canon_form` still answers at
the precise form level for stamping and divergence, where a sync helper
beside an async canon is a fair question; nothing here changes what the
ledger writes.
Returns a dict, deliberately. Widening a tuple return is the #4204 break
exactly — a 4-tuple grew a fifth field and one consumer went on unpacking
four, and Python said nothing until that line ran.
"""
forms: dict[str, int] = {}
families: dict[str, int] = {}
per_row: list[tuple[object, str]] = []
has_type: set[str] = set()
for r in members:
f = shape_form(getattr(r, "signature", "") or "", r.kind)
if not f:
continue # unreadable: never counts either way
fam = shape_family(f)
forms[f] = forms.get(f, 0) + 1
families[fam] = families.get(fam, 0) + 1
per_row.append((r, fam))
if fam == "type":
has_type.add(r.path)
out = {
"readable": len(per_row),
"forms": dict(sorted(forms.items(), key=lambda kv: -kv[1])),
"families": dict(sorted(families.items(), key=lambda kv: -kv[1])),
"family": FORM_UNKNOWN,
"attached": 0,
"strangers": [],
"coheres": True,
}
if not per_row:
return out # nothing legible: say nothing, not "broken"
# The vote, with methods abstaining. A callable sitting in a file that
# defines a class is presumed to belong to it and does not get to argue
# that the canon is really about callables.
def _attachable(row, fam: str) -> bool:
return fam == "callable" and row.path in has_type
votes: dict[str, int] = {}
for r, fam in per_row:
if _attachable(r, fam):
continue
votes[fam] = votes.get(fam, 0) + 1
if not votes: # every readable row is a method
votes = dict(families)
# Ties go to `type`: a canon that defines a class is about the class.
family = max(votes, key=lambda k: (votes[k], k == "type"))
attached, strangers = 0, []
for r, fam in per_row:
if fam == family:
continue
if family == "type" and _attachable(r, fam):
attached += 1
else:
strangers.append(r)
out["family"] = family
out["attached"] = attached
out["strangers"] = strangers
out["coheres"] = len(strangers) / len(per_row) <= _STRANGER_SHARE
return out
def signature_in(code: str, symbol: str, kind: str) -> str: def signature_in(code: str, symbol: str, kind: str) -> str:
"""The line in ``code`` that DEFINES ``symbol``, or "" if none does. """The line in ``code`` that DEFINES ``symbol``, or "" if none does.
@@ -2270,12 +2394,26 @@ def stamps_to_review(rows: Iterable[CodeShape], *, top: int = 10) -> dict:
person or an audit judged are never listed, however old: a human judgment person or an audit judged are never listed, however old: a human judgment
is not weak evidence, it is a different kind of evidence. is not weak evidence, it is a different kind of evidence.
`incoherent` — canons whose judged rows do not agree on a form. A canon is `incoherent` — canons whose membership no longer agrees what the canon
a claim that some set of shapes are the same sort of thing; when its own IS. A canon is a claim that some set of shapes are the same sort of
members are a class, three getters and a dozen tests, that claim has thing; when its own members are a class, three getters and a dozen tests,
stopped being true, and every base-rate reading built on it is reading that claim has stopped being true, and every base-rate reading built on
noise. `canon_form` already makes such a canon fall silent — this is what it is reading noise. `canon_form` already makes such a canon fall silent
makes it VISIBLE, which is the half that was missing. — this is what makes it VISIBLE, which is the half that was missing.
What counts as disagreement is `canon_coherence`, and it is deliberately
looser than "one form": it compares FAMILIES, and it does not hold a
method against the class it is defined beside. Both corrections came from
this surface's first live day, when the only two canons it reported were
both sound (#4220). A review surface whose output is noise is one that
stops being read, which costs more than the check was ever worth.
`strangers` names the rows that do not fit, not the first dozen members:
the reader's question is which ones are wrong. `unattended` counts rows
written by the hook with nobody reading — an incoherence made entirely of
judged rows is far more likely to be this check being too strict than a
ledger full of junk, and the reader should be able to see that without
opening the canon.
""" """
live = [r for r in rows if r.vanished_at is None] live = [r for r in rows if r.vanished_at is None]
weak = [] weak = []
@@ -2300,32 +2438,39 @@ def stamps_to_review(rows: Iterable[CodeShape], *, top: int = 10) -> dict:
by_canon.setdefault(int(r.snippet_id), []).append(r) by_canon.setdefault(int(r.snippet_id), []).append(r)
incoherent = [] incoherent = []
for sid, members in by_canon.items(): for sid, members in by_canon.items():
forms: dict[str, int] = {} coh = canon_coherence(members)
for r in members: if coh["readable"] < _DENSITY_MIN_JUDGED:
f = shape_form(r.signature or "", r.kind)
if f:
forms[f] = forms.get(f, 0) + 1
readable = sum(forms.values())
if readable < _DENSITY_MIN_JUDGED:
continue # too few to say anything either way continue # too few to say anything either way
if canon_form(members, sid): if coh["coheres"]:
continue # a form holds the majority: coherent continue
unattended = sum(1 for r in members if r.classified_by in _UNATTENDED_BY)
incoherent.append({ incoherent.append({
"snippet_id": sid, "snippet_id": sid,
"judged": len(members), "judged": len(members),
"forms": dict(sorted(forms.items(), key=lambda kv: -kv[1])), "forms": coh["forms"],
"families": coh["families"],
"family": coh["family"],
"attached": coh["attached"],
"unattended": unattended,
# The listing below is capped; without this you cannot tell a
# canon with twelve strangers from one with three hundred.
"stranger_count": len(coh["strangers"]),
"weak_rows": sum( "weak_rows": sum(
1 for r in members 1 for r in members
if r.classified_by == "hook" if r.classified_by in _UNATTENDED_BY
and (stamp_score(r.reason) or 1.0) < _RESEMBLE_MIN and (stamp_score(r.reason) or 1.0) < _RESEMBLE_MIN
), ),
"sample": [ # The rows that do NOT fit — not the first dozen members. The
# reader's question is "which ones are wrong", and a sample of
# the majority cannot answer it.
"strangers": [
{"path": r.path, "symbol": r.symbol, {"path": r.path, "symbol": r.symbol,
"signature": r.signature or "", "by": r.classified_by} "signature": r.signature or "", "by": r.classified_by,
for r in members[:_REVIEW_ROWS_SHOWN] "form": shape_form(r.signature or "", r.kind)}
for r in coh["strangers"][:_REVIEW_ROWS_SHOWN]
], ],
}) })
incoherent.sort(key=lambda d: (-d["weak_rows"], -d["judged"])) incoherent.sort(key=lambda d: (-d["weak_rows"], -d["unattended"], -d["judged"]))
return { return {
"weak_count": len(weak), "weak_count": len(weak),
"weak": weak[:top], "weak": weak[:top],
+160 -7
View File
@@ -25,7 +25,8 @@ hook's own history is made of.
import pytest import pytest
from scribe.services.shape_ledger import ( from scribe.services.shape_ledger import (
_RESEMBLE_MIN, _RESEMBLE_REASON, stamp_score, stamps_to_review, _RESEMBLE_MIN, _RESEMBLE_REASON, _REVIEW_ROWS_SHOWN, canon_coherence,
stamp_score, stamps_to_review,
) )
@@ -135,17 +136,26 @@ def test_a_canon_whose_members_agree_is_not_listed() -> None:
def test_a_canon_whose_members_are_all_different_things_is_listed() -> None: def test_a_canon_whose_members_are_all_different_things_is_listed() -> None:
"""Portal's #3283, in miniature: a class, a getter, a test and a binding """Portal's #3283, in miniature: a class, a getter, a test and a binding
recorded as instances of one pattern.""" recorded as instances of one pattern — and in four different files, as
rows = [_member("class SessionAbsent(RuntimeError):"), they really were. The paths matter now: a callable sharing a file with a
_member("def build_channel() -> str:"), class is read as that class's method (#4220), so putting them all in one
_member("async def attach(self) -> None:"), file would test the excuse rather than the disagreement."""
_member("MAX = 10")] rows = [_member("class SessionAbsent(RuntimeError):", path="src/errors.py",
symbol="SessionAbsent"),
_member("def build_channel() -> str:", path="src/channel.py",
symbol="build_channel"),
_member("async def attach(self) -> None:", path="src/attach.py",
symbol="attach"),
_member("MAX = 10", path="src/limits.py", symbol="MAX")]
out = stamps_to_review(rows) out = stamps_to_review(rows)
assert out["incoherent_count"] == 1 assert out["incoherent_count"] == 1
entry = out["incoherent"][0] entry = out["incoherent"][0]
assert entry["snippet_id"] == 7 and entry["judged"] == 4 assert entry["snippet_id"] == 7 and entry["judged"] == 4
assert set(entry["forms"]) == {"type", "fn", "async-fn", "binding"} assert set(entry["forms"]) == {"type", "fn", "async-fn", "binding"}
assert len(entry["sample"]) == 4 # The two callables carry the vote; the class and the binding are what
# the canon cannot account for, and they are what the reader is shown.
assert entry["family"] == "callable"
assert {s["symbol"] for s in entry["strangers"]} == {"SessionAbsent", "MAX"}
def test_too_few_readable_members_says_nothing_either_way() -> None: def test_too_few_readable_members_says_nothing_either_way() -> None:
@@ -197,3 +207,146 @@ def test_the_service_carries_no_machinery_for_bulk_withdrawal() -> None:
body = open(ledger.__file__).read() body = open(ledger.__file__).read()
for banned in ("def retire_weak", "def auto_unclassify", "def bulk_withdraw"): for banned in ("def retire_weak", "def auto_unclassify", "def bulk_withdraw"):
assert banned not in body, banned assert banned not in body, banned
# ── coherence is about what a canon CANNOT account for (#4220) ────────────
#
# The first version of this check asked whether every member shared a form.
# On its first live day the only two canons it reported were both sound,
# while the genuinely polluted one had already been cleaned by hand — a
# surface whose whole output is noise is one that stops being read.
#
# The obvious repair — group by family, keep the majority test — is a trap,
# and `test_family_grouping_alone_would_have_gone_blind` is the guard that
# stops anyone walking back into it.
def _model(sym, path):
"""A SQLAlchemy model class, as #2844 holds one."""
return _member(f"class {sym}(Base, TimestampMixin):", path=path)
def _to_dict(path):
"""The serialiser that model is required to carry, beside it."""
return _member("def to_dict(self) -> dict:", path=path, symbol="to_dict")
def test_a_class_and_the_to_dict_beside_it_are_one_shape() -> None:
"""#2844 exactly: the SQLAlchemy model convention, whose own text is
about the class AND the to_dict it must carry. Live it scored
37/62 = 0.597 and was called incoherent for containing precisely what it
says it contains."""
rows = []
for i in range(6):
rows += [_model(f"M{i}", f"src/scribe/models/m{i}.py"),
_to_dict(f"src/scribe/models/m{i}.py")]
out = stamps_to_review(rows)
assert out["incoherent_count"] == 0
def test_a_sync_starter_and_the_async_tick_it_schedules_are_one_shape() -> None:
"""#2849: four loop-starters and three ticks. `shape_family` already
collapses fn and async-fn for the divergence gate; this surface was the
one place still asking at form level."""
rows = [_member("def start_notification_loop() -> None:", path="src/n.py"),
_member("def start_log_retention_loop() -> None:", path="src/l.py"),
_member("def start_auth_token_retention_loop() -> None:", path="src/a.py"),
_member("async def _notification_tick() -> None:", path="src/n.py"),
_member("async def _retention_tick() -> None:", path="src/l.py"),
_member("async def _auth_token_retention_tick() -> None:", path="src/a.py")]
assert stamps_to_review(rows)["incoherent_count"] == 0
def test_family_grouping_alone_would_have_gone_blind() -> None:
"""THE REGRESSION GUARD. Polluted #2844 as it stood this morning: 37
model classes, 25 to_dict methods beside them, and 31 rows that had no
business being there. Counted as families that is 56 callables to 37
types — 0.602, a clean pass under any majority test. Those 31 rows are
the ones the surface exists to find, so a change that lets this canon
read as coherent has broken the feature while keeping every other test
in this file green."""
rows = []
for i in range(37):
rows.append(_model(f"M{i}", f"src/scribe/models/m{i}.py"))
for i in range(25):
rows.append(_to_dict(f"src/scribe/models/m{i % 37}.py"))
strangers = (["src/scribe/routes/rulebooks.py"] * 2
+ ["src/scribe/services/embeddings.py"] * 7
+ ["frontend/src/stores/rulebooks.ts"] * 3
+ ["tests/test_services_rulebooks.py"] * 19)
for i, path in enumerate(strangers):
rows.append(_member(f"async def f{i}():", path=path, symbol=f"f{i}"))
out = stamps_to_review(rows)
assert out["incoherent_count"] == 1, "the polluted canon must still be caught"
entry = out["incoherent"][0]
assert entry["families"]["callable"] > entry["families"]["type"]
assert entry["family"] == "type", "methods must not outvote their classes"
assert entry["attached"] == 25
assert entry["stranger_count"] == len(strangers) == 31
# The listing is capped, the count is not — a reader must be able to
# tell twelve strangers from thirty-one.
assert len(entry["strangers"]) == _REVIEW_ROWS_SHOWN
assert all(not s["path"].startswith("src/scribe/models")
for s in entry["strangers"])
def test_a_callable_in_a_file_with_no_class_is_a_stranger() -> None:
"""The excuse is "method of a member", not "callable anywhere". Without
this, any function in the repo would be excused by the existence of a
class somewhere else in the canon."""
rows = [_model("A", "src/models/a.py"), _to_dict("src/models/a.py"),
_model("B", "src/models/b.py"), _to_dict("src/models/b.py"),
_member("def helper():", path="src/services/free.py", symbol="helper"),
_member("def other():", path="src/services/free.py", symbol="other")]
out = stamps_to_review(rows)
assert out["incoherent_count"] == 1
assert {s["symbol"] for s in out["incoherent"][0]["strangers"]} == {"helper", "other"}
def test_the_strangers_are_named_not_a_sample_of_the_majority() -> None:
"""The reader's question is which rows are wrong. A sample of the
agreeing majority cannot answer it, which is what the first version
returned."""
# Seven and two: 2 of 9 is over `_STRANGER_SHARE`, 2 of 11 is under it.
# The tolerance is real and the counts here sit deliberately on the far
# side of it — see the test below for the near side.
rows = [_member("async def a():", path=f"src/ok{i}.py", symbol=f"a{i}")
for i in range(7)]
rows.append(_member("class Odd:", path="src/odd.py", symbol="Odd"))
rows.append(_member("class Odder:", path="src/odder.py", symbol="Odder"))
out = stamps_to_review(rows)
assert out["incoherent_count"] == 1
assert {s["symbol"] for s in out["incoherent"][0]["strangers"]} == {"Odd", "Odder"}
def test_a_single_odd_row_in_a_large_canon_is_tolerated() -> None:
"""One stranger in twenty is a row to fix, not a canon that has stopped
meaning anything. The surface is for the second thing."""
rows = [_member("async def a():", path=f"src/ok{i}.py", symbol=f"a{i}")
for i in range(19)]
rows.append(_member("class Odd:", path="src/odd.py", symbol="Odd"))
assert stamps_to_review(rows)["incoherent_count"] == 0
def test_an_incoherence_made_of_judged_rows_says_so() -> None:
"""The discriminator that tells a too-strict check from a bad ledger.
Both canons flagged on the first live day were entirely audit-judged,
and a reader could not see that without opening each one."""
judged = [_member("async def a():", path=f"src/j{i}.py", symbol=f"a{i}")
for i in range(4)]
judged += [_member("class J:", path="src/jc.py", symbol="J"),
_member("class J2:", path="src/jc2.py", symbol="J2")]
out = stamps_to_review(judged)
assert out["incoherent_count"] == 1
assert out["incoherent"][0]["unattended"] == 0
assert out["incoherent"][0]["weak_rows"] == 0
def test_the_coherence_verdict_reads_rows_and_writes_nothing() -> None:
"""`canon_coherence` is called on live ORM rows; it must not touch them."""
rows = [_model("A", "src/models/a.py"), _to_dict("src/models/a.py"),
_member("def loose():", path="src/x.py", symbol="loose")]
before = [(r.path, r.symbol, r.status, r.snippet_id, r.signature) for r in rows]
canon_coherence(rows)
assert [(r.path, r.symbol, r.status, r.snippet_id, r.signature)
for r in rows] == before