fix(shapes): the semantic arm proposes at the write-path floor, not a private 0.8 (#4208)
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 56s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Build & push image (push) Successful in 22s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 56s
CI & Build / Python tests (push) Successful in 1m34s
CI & Build / Build & push image (push) Successful in 22s
0.8 was sized for proposals nobody reads; every semantic proposal is read by a judge before it is confirmed. Measured, it proposed 0 of 150 while judged instances of a canon score 0.68-0.71 - the write-path hint's own floor asks the same question of the same documents at 0.68. The arm now uses that floor, scans 8 hits instead of 3 (a true instance ranked 4th behind snippets of other language families), and the proposer version bumps to 5 so rows examined under the old floor are read again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
@@ -1516,24 +1516,32 @@ _DERIVE_MIN_NAME_CSS = 2
|
|||||||
# Semantic checks per repo per refresh — an embedding each (local fastembed),
|
# Semantic checks per repo per refresh — an embedding each (local fastembed),
|
||||||
# bounded so a 4,000-row ledger is worked through over refreshes, not in one.
|
# bounded so a 4,000-row ledger is worked through over refreshes, not in one.
|
||||||
_SEMANTIC_CAP = 150
|
_SEMANTIC_CAP = 150
|
||||||
# The semantic basis's own floor — stricter than the write-path arm's, because
|
# The semantic arm has NO floor of its own: it uses the write-path surface's,
|
||||||
# a code BODY against a prose-forward snippet document scores 0.68–0.75 for
|
# because it asks the write-path hint's question — "which canon does this body
|
||||||
# "both are about migrations"; first live run paired every alembic
|
# resemble?" — of the same snippet documents (#4208).
|
||||||
# upgrade()/downgrade() with an unrelated canon at exactly that band.
|
|
||||||
_SEMANTIC_FLOOR = 0.8
|
|
||||||
# How many above-floor hits the semantic arm asks for.
|
|
||||||
_SEMANTIC_LIMIT = 3
|
|
||||||
# A MISS FROM THIS ARM IS NOT EVIDENCE, and nothing may read it as one (#4208).
|
|
||||||
#
|
#
|
||||||
|
# It had one, 0.8, set after the first live run paired alembic
|
||||||
|
# upgrade()/downgrade() with unrelated canons at 0.68–0.75. That floor was
|
||||||
|
# sized for a proposal nobody reads, and nobody-reads is not how proposals are
|
||||||
|
# consumed: a judge reviews each semantic one (confirm_proposals refuses an
|
||||||
|
# unnamed batch, and the tool says "semantic deserves a look"). Measured on
|
||||||
|
# 2026-09-22 it proposed 0 of 150, while judged instances of the service-unit
|
||||||
|
# canon score 0.68–0.71 against it. A wrong proposal costs the judge a look; a
|
||||||
|
# floor that yields nothing costs the signal. Noise in this band is expected —
|
||||||
|
# the score rides on the proposal so the judge can weigh it.
|
||||||
|
#
|
||||||
|
# How many hits the arm scans for an ALLOWED canon. Wider than the hint's
|
||||||
|
# budget because disallowed snippets (other language families) share the
|
||||||
|
# ranking: one measured true instance sat 4th, behind three it could not use.
|
||||||
|
_SEMANTIC_LIMIT = 8
|
||||||
|
# A MISS FROM THIS ARM IS NOT EVIDENCE, and nothing may read it as one (#4208).
|
||||||
# #4208 briefly stored "compared, nothing cleared the floor" and let it
|
# #4208 briefly stored "compared, nothing cleared the floor" and let it
|
||||||
# withdraw divergence prompts. Measured live on 2026-09-22, with the arm's own
|
# withdraw divergence prompts. It silenced real divergences exactly as it
|
||||||
# query shape: judged instances of the service-unit canon (#2860) score 0.68–
|
# silenced helpers: most judged instances of a canon score below any floor
|
||||||
# 0.71 against it at best, and most fall below 0.66; helpers beside it score
|
# that keeps noise out (0.66 and under, in the measurement above), and helpers
|
||||||
# 0.66–0.75 against snippets they have nothing to do with. Nothing reaches
|
# score 0.66–0.75 against snippets they have nothing to do with. A code body
|
||||||
# 0.8, so the arm answers "no canon" for almost every body, true instance or
|
# against a prose-forward snippet document measures the wrong field (#2518);
|
||||||
# not — a real divergence was silenced exactly as a helper was. A code body
|
# no floor separates the two bands. The arm PROPOSES on a hit; its silence
|
||||||
# against a prose-forward snippet document measures the wrong field (#2518),
|
|
||||||
# and no floor separates the two bands. The arm PROPOSES on a hit; its silence
|
|
||||||
# says nothing.
|
# says nothing.
|
||||||
|
|
||||||
|
|
||||||
@@ -1561,7 +1569,9 @@ def _semantic_priority(row) -> tuple:
|
|||||||
# v4: the semantic arm recorded its misses too (#4208). Retired — a miss is
|
# v4: the semantic arm recorded its misses too (#4208). Retired — a miss is
|
||||||
# not evidence (see _SEMANTIC_LIMIT) — but the proposals themselves did not
|
# not evidence (see _SEMANTIC_LIMIT) — but the proposals themselves did not
|
||||||
# change, so no re-examination is owed; a stale miss basis is inert.
|
# change, so no re-examination is owed; a stale miss basis is inert.
|
||||||
_PROPOSER_VERSION = 4
|
# v5: the semantic arm's floor dropped from 0.8 to the write-path floor (#4208),
|
||||||
|
# so rows it examined and found nothing for must be read once more.
|
||||||
|
_PROPOSER_VERSION = 5
|
||||||
# Signature resemblance floor, name blanked (difflib ratio) — and a length
|
# Signature resemblance floor, name blanked (difflib ratio) — and a length
|
||||||
# floor, because `def NAME():` resembles `def NAME(x):` at 0.95 while saying
|
# floor, because `def NAME():` resembles `def NAME(x):` at 0.95 while saying
|
||||||
# nothing; a family shape has parameters to resemble.
|
# nothing; a family shape has parameters to resemble.
|
||||||
@@ -1810,7 +1820,7 @@ async def _semantic_canon(
|
|||||||
query = concept_query(body) or body
|
query = concept_query(body) or body
|
||||||
hits = await semantic_search_notes(
|
hits = await semantic_search_notes(
|
||||||
user_id, query, limit=_SEMANTIC_LIMIT,
|
user_id, query, limit=_SEMANTIC_LIMIT,
|
||||||
threshold=max(WRITEPATH_DEFAULT_THRESHOLD, _SEMANTIC_FLOOR),
|
threshold=WRITEPATH_DEFAULT_THRESHOLD,
|
||||||
note_type="snippet", scope="browse",
|
note_type="snippet", scope="browse",
|
||||||
)
|
)
|
||||||
for score, note in hits:
|
for score, note in hits:
|
||||||
|
|||||||
@@ -11,9 +11,11 @@ do. A signature does not carry a job.
|
|||||||
the floor" and let it withdraw the prompt. Measured live on 2026-09-22 it
|
the floor" and let it withdraw the prompt. Measured live on 2026-09-22 it
|
||||||
cannot discriminate. Judged instances of the service-unit canon score 0.68–
|
cannot discriminate. Judged instances of the service-unit canon score 0.68–
|
||||||
0.71 against it at best, most below 0.66; helpers score 0.66–0.75 against
|
0.71 against it at best, most below 0.66; helpers score 0.66–0.75 against
|
||||||
snippets unrelated to them. Nothing reaches the 0.8 floor, so the "miss"
|
snippets unrelated to them. The then-floor of 0.8 was reached by nothing, so
|
||||||
fires for nearly every body — the real divergence silenced exactly as the
|
the "miss" fired for nearly every body — the real divergence silenced exactly
|
||||||
helper was. The gate was removed; the prompt asks and the judge answers.
|
as the helper was. The gate was removed; the prompt asks and the judge
|
||||||
|
answers. The same measurement moved the arm's floor down to the write-path
|
||||||
|
hint's: 0.8 was sized for proposals nobody reads, and it proposed 0 of 150.
|
||||||
|
|
||||||
These tests pin what stayed: the arm proposes on a hit and says nothing on a
|
These tests pin what stayed: the arm proposes on a hit and says nothing on a
|
||||||
miss, the divergence check does not consult it, and the capped pass reads new
|
miss, the divergence check does not consult it, and the capped pass reads new
|
||||||
@@ -92,6 +94,18 @@ async def test_no_allowed_canon_spends_no_search() -> None:
|
|||||||
mock.assert_not_awaited()
|
mock.assert_not_awaited()
|
||||||
|
|
||||||
|
|
||||||
|
async def test_the_arm_proposes_at_the_write_path_floor() -> None:
|
||||||
|
"""It asks the write-path hint's question of the same documents, so it
|
||||||
|
uses that surface's floor — not a stricter private one that yields no
|
||||||
|
proposals for the judge to weigh (#4208)."""
|
||||||
|
from scribe.services.plugin_context import WRITEPATH_DEFAULT_THRESHOLD
|
||||||
|
|
||||||
|
mock = _hits()
|
||||||
|
with _patch(mock):
|
||||||
|
await _semantic_canon(1, BODY, {CANON})
|
||||||
|
assert mock.await_args.kwargs["threshold"] == WRITEPATH_DEFAULT_THRESHOLD
|
||||||
|
|
||||||
|
|
||||||
# ── ...and its silence reaches nothing ───────────────────────────────────
|
# ── ...and its silence reaches nothing ───────────────────────────────────
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
@@ -897,7 +897,7 @@ async def test_a_semantic_miss_does_not_silence_the_prompt(seeded):
|
|||||||
`confirmDanger` beside an async confirm canon is structurally identical to
|
`confirmDanger` beside an async confirm canon is structurally identical to
|
||||||
a registry helper beside an async service canon, and the arm's miss was
|
a registry helper beside an async service canon, and the arm's miss was
|
||||||
meant to tell them apart. Measured live, true instances of a canon rarely
|
meant to tell them apart. Measured live, true instances of a canon rarely
|
||||||
clear the arm's 0.8 floor either, so a miss fires on both — silencing on
|
clear the arm's floor either, so a miss fires on both — silencing on
|
||||||
it silenced #2793's own case. Here the arm has read the body, found
|
it silenced #2793's own case. Here the arm has read the body, found
|
||||||
nothing, and the prompt is still RAISED.
|
nothing, and the prompt is still RAISED.
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user