fix(shapes): the semantic arm proposes at the write-path floor, not a private 0.8 (#4208)
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 56s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Build & push image (push) Successful in 22s

0.8 was sized for proposals nobody reads; every semantic proposal is read
by a judge before it is confirmed. Measured, it proposed 0 of 150 while
judged instances of a canon score 0.68-0.71 - the write-path hint's own
floor asks the same question of the same documents at 0.68. The arm now
uses that floor, scans 8 hits instead of 3 (a true instance ranked 4th
behind snippets of other language families), and the proposer version
bumps to 5 so rows examined under the old floor are read again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
2026-09-22 08:33:01 -04:00
co-authored by Claude Opus 5
parent a875a1b2ee
commit 312dc9f6f7
3 changed files with 46 additions and 22 deletions
+28 -18
View File
@@ -1516,24 +1516,32 @@ _DERIVE_MIN_NAME_CSS = 2
# Semantic checks per repo per refresh — an embedding each (local fastembed), # Semantic checks per repo per refresh — an embedding each (local fastembed),
# bounded so a 4,000-row ledger is worked through over refreshes, not in one. # bounded so a 4,000-row ledger is worked through over refreshes, not in one.
_SEMANTIC_CAP = 150 _SEMANTIC_CAP = 150
# The semantic basis's own floor — stricter than the write-path arm's, because # The semantic arm has NO floor of its own: it uses the write-path surface's,
# a code BODY against a prose-forward snippet document scores 0.680.75 for # because it asks the write-path hint's question — "which canon does this body
# "both are about migrations"; first live run paired every alembic # resemble?" — of the same snippet documents (#4208).
# upgrade()/downgrade() with an unrelated canon at exactly that band.
_SEMANTIC_FLOOR = 0.8
# How many above-floor hits the semantic arm asks for.
_SEMANTIC_LIMIT = 3
# A MISS FROM THIS ARM IS NOT EVIDENCE, and nothing may read it as one (#4208).
# #
# It had one, 0.8, set after the first live run paired alembic
# upgrade()/downgrade() with unrelated canons at 0.680.75. That floor was
# sized for a proposal nobody reads, and nobody-reads is not how proposals are
# consumed: a judge reviews each semantic one (confirm_proposals refuses an
# unnamed batch, and the tool says "semantic deserves a look"). Measured on
# 2026-09-22 it proposed 0 of 150, while judged instances of the service-unit
# canon score 0.680.71 against it. A wrong proposal costs the judge a look; a
# floor that yields nothing costs the signal. Noise in this band is expected —
# the score rides on the proposal so the judge can weigh it.
#
# How many hits the arm scans for an ALLOWED canon. Wider than the hint's
# budget because disallowed snippets (other language families) share the
# ranking: one measured true instance sat 4th, behind three it could not use.
_SEMANTIC_LIMIT = 8
# A MISS FROM THIS ARM IS NOT EVIDENCE, and nothing may read it as one (#4208).
# #4208 briefly stored "compared, nothing cleared the floor" and let it # #4208 briefly stored "compared, nothing cleared the floor" and let it
# withdraw divergence prompts. Measured live on 2026-09-22, with the arm's own # withdraw divergence prompts. It silenced real divergences exactly as it
# query shape: judged instances of the service-unit canon (#2860) score 0.68 # silenced helpers: most judged instances of a canon score below any floor
# 0.71 against it at best, and most fall below 0.66; helpers beside it score # that keeps noise out (0.66 and under, in the measurement above), and helpers
# 0.660.75 against snippets they have nothing to do with. Nothing reaches # score 0.660.75 against snippets they have nothing to do with. A code body
# 0.8, so the arm answers "no canon" for almost every body, true instance or # against a prose-forward snippet document measures the wrong field (#2518);
# not — a real divergence was silenced exactly as a helper was. A code body # no floor separates the two bands. The arm PROPOSES on a hit; its silence
# against a prose-forward snippet document measures the wrong field (#2518),
# and no floor separates the two bands. The arm PROPOSES on a hit; its silence
# says nothing. # says nothing.
@@ -1561,7 +1569,9 @@ def _semantic_priority(row) -> tuple:
# v4: the semantic arm recorded its misses too (#4208). Retired — a miss is # v4: the semantic arm recorded its misses too (#4208). Retired — a miss is
# not evidence (see _SEMANTIC_LIMIT) — but the proposals themselves did not # not evidence (see _SEMANTIC_LIMIT) — but the proposals themselves did not
# change, so no re-examination is owed; a stale miss basis is inert. # change, so no re-examination is owed; a stale miss basis is inert.
_PROPOSER_VERSION = 4 # v5: the semantic arm's floor dropped from 0.8 to the write-path floor (#4208),
# so rows it examined and found nothing for must be read once more.
_PROPOSER_VERSION = 5
# Signature resemblance floor, name blanked (difflib ratio) — and a length # Signature resemblance floor, name blanked (difflib ratio) — and a length
# floor, because `def NAME():` resembles `def NAME(x):` at 0.95 while saying # floor, because `def NAME():` resembles `def NAME(x):` at 0.95 while saying
# nothing; a family shape has parameters to resemble. # nothing; a family shape has parameters to resemble.
@@ -1810,7 +1820,7 @@ async def _semantic_canon(
query = concept_query(body) or body query = concept_query(body) or body
hits = await semantic_search_notes( hits = await semantic_search_notes(
user_id, query, limit=_SEMANTIC_LIMIT, user_id, query, limit=_SEMANTIC_LIMIT,
threshold=max(WRITEPATH_DEFAULT_THRESHOLD, _SEMANTIC_FLOOR), threshold=WRITEPATH_DEFAULT_THRESHOLD,
note_type="snippet", scope="browse", note_type="snippet", scope="browse",
) )
for score, note in hits: for score, note in hits:
+17 -3
View File
@@ -11,9 +11,11 @@ do. A signature does not carry a job.
the floor" and let it withdraw the prompt. Measured live on 2026-09-22 it the floor" and let it withdraw the prompt. Measured live on 2026-09-22 it
cannot discriminate. Judged instances of the service-unit canon score 0.68 cannot discriminate. Judged instances of the service-unit canon score 0.68
0.71 against it at best, most below 0.66; helpers score 0.660.75 against 0.71 against it at best, most below 0.66; helpers score 0.660.75 against
snippets unrelated to them. Nothing reaches the 0.8 floor, so the "miss" snippets unrelated to them. The then-floor of 0.8 was reached by nothing, so
fires for nearly every body — the real divergence silenced exactly as the the "miss" fired for nearly every body — the real divergence silenced exactly
helper was. The gate was removed; the prompt asks and the judge answers. as the helper was. The gate was removed; the prompt asks and the judge
answers. The same measurement moved the arm's floor down to the write-path
hint's: 0.8 was sized for proposals nobody reads, and it proposed 0 of 150.
These tests pin what stayed: the arm proposes on a hit and says nothing on a These tests pin what stayed: the arm proposes on a hit and says nothing on a
miss, the divergence check does not consult it, and the capped pass reads new miss, the divergence check does not consult it, and the capped pass reads new
@@ -92,6 +94,18 @@ async def test_no_allowed_canon_spends_no_search() -> None:
mock.assert_not_awaited() mock.assert_not_awaited()
async def test_the_arm_proposes_at_the_write_path_floor() -> None:
"""It asks the write-path hint's question of the same documents, so it
uses that surface's floor — not a stricter private one that yields no
proposals for the judge to weigh (#4208)."""
from scribe.services.plugin_context import WRITEPATH_DEFAULT_THRESHOLD
mock = _hits()
with _patch(mock):
await _semantic_canon(1, BODY, {CANON})
assert mock.await_args.kwargs["threshold"] == WRITEPATH_DEFAULT_THRESHOLD
# ── ...and its silence reaches nothing ─────────────────────────────────── # ── ...and its silence reaches nothing ───────────────────────────────────
+1 -1
View File
@@ -897,7 +897,7 @@ async def test_a_semantic_miss_does_not_silence_the_prompt(seeded):
`confirmDanger` beside an async confirm canon is structurally identical to `confirmDanger` beside an async confirm canon is structurally identical to
a registry helper beside an async service canon, and the arm's miss was a registry helper beside an async service canon, and the arm's miss was
meant to tell them apart. Measured live, true instances of a canon rarely meant to tell them apart. Measured live, true instances of a canon rarely
clear the arm's 0.8 floor either, so a miss fires on both — silencing on clear the arm's floor either, so a miss fires on both — silencing on
it silenced #2793's own case. Here the arm has read the body, found it silenced #2793's own case. Here the arm has read the body, found
nothing, and the prompt is still RAISED. nothing, and the prompt is still RAISED.