fix(tests): two expectations that the new contract corrected (#4101)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / integration (push) Successful in 41s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 2m9s
CI & Build / Build & push image (push) Successful in 27s

Both were mine, and both show the change behaving as designed.

`test_session_dedup_marks_the_reuse_arms_rather_than_silencing_them` asserted
#12 stays out of the search's `exclude_ids`. It does not, and should not: the
place arm has just listed it, so the semantic arm must not list it again. That
is the same-call rule, not the ledger. The claim I meant — the LEDGER never
reaches the query — is asserted on its own against a record the call has not
otherwise rendered.

`test_a_pulled_snippet_already_seen_is_evidence_not_menu` expected #7 and #8.
It gets #7 alone, because #7 (0.91) now anchors the margin band at 0.81 and #8
scores 0.80. Withholding #7 used to promote a materially weaker hit into a slot
it had not earned, with nothing in the output saying so — the band measures
distance from the best answer, and letting the ledger decide which answer that
is was the axis confusion the rule band was written to avoid.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
2026-09-16 21:19:05 -04:00
co-authored by Claude Opus 5
parent 825491d859
commit 5d47342fbc
+12 -2
View File
@@ -345,7 +345,12 @@ async def test_session_dedup_marks_the_reuse_arms_rather_than_silencing_them():
assert "seen" in next( assert "seen" in next(
line for line in out["context"].splitlines() if "#12" in line line for line in out["context"].splitlines() if "#12" in line
) )
assert 12 not in (search.await_args.kwargs.get("exclude_ids") or set()) # #12 IS in this call's `exclude_ids`, and that is the other rule working:
# the place arm has just listed it, so the semantic arm must not list it
# again. What must never reach the query is the LEDGER, which is asserted
# on its own in tests/test_ledger_references_not_silence.py against a
# record this call has not otherwise rendered.
assert 12 in search.await_args.kwargs["exclude_ids"]
# --- attribution + telemetry ------------------------------------------------- # --- attribution + telemetry -------------------------------------------------
@@ -1313,9 +1318,14 @@ async def test_a_pulled_snippet_already_seen_is_evidence_not_menu():
assert 7 not in kw["exclude_ids"] assert 7 not in kw["exclude_ids"]
assert kw["limit"] == 3 assert kw["limit"] == 3
# ...and the menu carries it, marked. # ...and the menu carries it, marked.
assert out["note_ids"] == [7, 8]
seven = next(line for line in out["context"].splitlines() if "#7" in line) seven = next(line for line in out["context"].splitlines() if "#7" in line)
assert "seen" in seven assert "seen" in seven
# #8 (0.80) now falls OUTSIDE the margin band, because #7 (0.91) anchors it
# at 0.81 instead of being filtered out first. That is the point rather than
# a casualty: the band measures distance from the best answer, and the best
# answer here is #7 — suppressing it used to promote a materially weaker hit
# into a slot it had not earned, with nothing in the output saying so.
assert out["note_ids"] == [7]
# The stamp saw the pull and the resemblance score for this payload. # The stamp saw the pull and the resemblance score for this payload.
skw = stamp.call_args.kwargs skw = stamp.call_args.kwargs
assert skw["pulled"] == {7: skw["pulled"][7]} assert skw["pulled"] == {7: skw["pulled"][7]}