Two fixes about the same confusion in two places: treating "the reader already had this" as if it were "the ranker found nothing", and treating "was shown once" as "is known now".
#3739 — near_misses counted a repeat as a rejection
Caught on the first live read after yesterday's deploy. The readout contradicted itself:
pre_tool_rule top_score.min 0.7204 ← the lowest score ever RETURNED
near_misses.max 0.7457 ← "rejected", but scored higher
best_available_score is measured pre-threshold, which is right, but for the rule arms it is also pre-exclusion, which is not. The note arms pass exclude_ids into semantic_search_notes so their score is already post-exclusion and clean; semantic_search_rules takes no such parameter, so the rule arms filter in Python afterwards and a rule that cleared the bar and was dropped as a repeat still reported its score on a zero-result row.
That is #3497's distinction — a ranker decline versus a reader already ahead of it — reintroduced one level up, inside the field built to replace a tautology.
The population now also requires suppressed_count IS NULL OR = 0. The NULL arm is principled rather than permissive: null means the caller filtered inside the search, exactly the case where the score cannot be contaminated. Deliberately conservative — a call carrying both a repeat and a lower genuine miss is dropped whole. It undercounts; it cannot corrupt.
It also makes near_misses.max < threshold true by construction: an above-bar candidate nobody excluded would have been returned, so its call is not in the population at all.
The test did not catch this, and that is the part worth keeping. The assertion was already there with the right intent; it passed because the fixture contained no suppressed call. The guard held because the breaking shape was absent, not because the code was right — rule 167's stated failure mode, in a test written while citing rule 167. The fixture now builds that shape.
#3749 — the rule ledger outlived the context it described
The hooks record every rule id they have named and hand it back as exclude_rule_ids, so a rule is surfaced once per session and then goes quiet. Correct while the session still holds what it was told.
A compaction breaks that in the worst available way: it summarises the injections out of context and does not touch the filesystem. The rule ends up absent from context and still excluded — unreachable for the rest of the session. The compaction banner tells the model to re-pull its always-on rules, but an arm-surfaced rule is conditional and is not in that set. The rules most likely to be in that state are the ones that fire most often.
Cleared on the sources that destroy context and only those: compact and clear yes; resume no, because the context genuinely came back and re-surfacing everything after a restore that lost nothing is the same defect from the other side. fork keeps it, and that answer holds whichever way forks are keyed.
Only the rule ledger — the note arms' .ids / .sync.ids / .derive.ids are a different question with a different answer, and a rm glob would have decided it silently.
The guard pins the discrimination, not the deletion: the whole source table is one assertion, so a blanket delete (all False) and a no-op (all True) both fail.
Plugin version
2026.09.09.0408, minted. The manifest gate caught the missing bump and was right to: #3749 is entirely a hook change, and without the bump it would have reached the repo and stopped — the marketplace clone updates, the cache that executes does not (#2209, #1040, #2220).
So this one needs a plugin refresh as well as a deploy.#3749 does nothing until the cache picks up the new version.
Two fixes about the same confusion in two places: treating "the reader already had this" as if it were "the ranker found nothing", and treating "was shown once" as "is known now".
## #3739 — `near_misses` counted a repeat as a rejection
Caught on the first live read after yesterday's deploy. The readout contradicted itself:
```
pre_tool_rule top_score.min 0.7204 ← the lowest score ever RETURNED
near_misses.max 0.7457 ← "rejected", but scored higher
```
`best_available_score` is measured pre-threshold, which is right, but for the rule arms it is also **pre-exclusion**, which is not. The note arms pass `exclude_ids` into `semantic_search_notes` so their score is already post-exclusion and clean; `semantic_search_rules` takes no such parameter, so the rule arms filter in Python afterwards and a rule that cleared the bar and was dropped as a repeat still reported its score on a zero-result row.
That is #3497's distinction — a ranker decline versus a reader already ahead of it — reintroduced one level up, inside the field built to replace a tautology.
The population now also requires `suppressed_count IS NULL OR = 0`. The NULL arm is principled rather than permissive: null means the caller filtered *inside* the search, exactly the case where the score cannot be contaminated. Deliberately conservative — a call carrying both a repeat and a lower genuine miss is dropped whole. It undercounts; it cannot corrupt.
It also makes `near_misses.max < threshold` true **by construction**: an above-bar candidate nobody excluded would have been returned, so its call is not in the population at all.
**The test did not catch this, and that is the part worth keeping.** The assertion was already there with the right intent; it passed because the fixture contained no suppressed call. The guard held because the breaking shape was absent, not because the code was right — rule 167's stated failure mode, in a test written while citing rule 167. The fixture now builds that shape.
## #3749 — the rule ledger outlived the context it described
The hooks record every rule id they have named and hand it back as `exclude_rule_ids`, so a rule is surfaced once per session and then goes quiet. Correct while the session still holds what it was told.
A compaction breaks that in the worst available way: it summarises the injections out of context and does not touch the filesystem. The rule ends up absent from context **and** still excluded — unreachable for the rest of the session. The compaction banner tells the model to re-pull its *always-on* rules, but an arm-surfaced rule is conditional and is not in that set. The rules most likely to be in that state are the ones that fire most often.
Cleared on the sources that destroy context and only those: `compact` and `clear` yes; `resume` no, because the context genuinely came back and re-surfacing everything after a restore that lost nothing is the same defect from the other side. `fork` keeps it, and that answer holds whichever way forks are keyed.
Only the **rule** ledger — the note arms' `.ids` / `.sync.ids` / `.derive.ids` are a different question with a different answer, and a `rm` glob would have decided it silently.
The guard pins the **discrimination**, not the deletion: the whole source table is one assertion, so a blanket delete (all False) and a no-op (all True) both fail.
## Plugin version
`2026.09.09.0408`, minted. The manifest gate caught the missing bump and was right to: #3749 is entirely a hook change, and without the bump it would have reached the repo and stopped — the marketplace clone updates, the cache that executes does not (#2209, #1040, #2220).
**So this one needs a plugin refresh as well as a deploy.** #3749 does nothing until the cache picks up the new version.
CI green on `ab14f78` (run 6126), all six lanes.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_011cPyzNnegXHr5iRMzzy5KJ
Caught on the first live read after deploying #3670. The readout
contradicted itself:
pre_tool_rule top_score.min 0.7204 the lowest score ever RETURNED
near_misses.max 0.7457 "rejected", but scored higher
`best_available_score` is measured pre-threshold, which is right, but for
the rule arms it is also PRE-EXCLUSION, which is not. The note arms pass
`exclude_ids` into semantic_search_notes so their score is already
post-exclusion and clean; `semantic_search_rules` takes no such parameter,
so the rule arms filter in Python after the search and a rule that cleared
the bar and was dropped as a repeat still reported its score on a
zero-result row.
That is #3497's distinction — a ranker decline versus a reader already
ahead of it — reintroduced one level up, inside the field built to replace
a tautology.
The population now also requires `suppressed_count IS NULL OR = 0`. The
NULL arm is principled rather than permissive: null means the caller
filtered INSIDE the search, which is exactly the case where the reported
score cannot be contaminated.
Deliberately conservative — a call carrying both a repeat and a lower
genuine miss is dropped whole, losing that point. It undercounts; it
cannot corrupt, which is the right way round for a number read against a
bar.
It also makes `near_misses.max < threshold` true BY CONSTRUCTION rather
than by fixture: an above-bar candidate nobody excluded would have been
returned, so its call is not in the population at all.
THE TEST DID NOT CATCH THIS, and that is the part worth keeping. The
assertion `nm["max"] < 0.72` was already there, with exactly the right
intent. It passed because the fixture contained no suppressed call — the
guard held because the breaking shape was absent, not because the code was
right. Rule 167's stated failure mode, in a test written while citing rule
167. The fixture now builds that shape: a 0.9 hit dropped as a repeat,
which lands in the population and drags `max` above the threshold unless
the predicate excludes it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cPyzNnegXHr5iRMzzy5KJ
The prior-art and tool-rule hooks record every rule id they have named in
<state>/<sid>.rules.ids and hand it back as exclude_rule_ids, so a rule is
surfaced once per session and then goes quiet. That is correct while the
session still HOLDS what it was told.
A compaction breaks it in the worst available way: it summarizes the
earlier injections out of context and does not touch the filesystem. The
rule ends up absent from context AND still excluded — unreachable for the
rest of the session. The compaction banner this hook already prints tells
the model to re-pull its ALWAYS-ON rules, but a rule an arm surfaced is
conditional and is not in that set, so it has no other way back. The rules
most likely to be in that state are the ones that fire most often.
The stale ledger is genuinely found again rather than orphaned: the etag
marker further down this same hook is rewritten on `compact` and keyed by
session_id, which is only meaningful if the id survives a compaction.
CLEARED ON THE SOURCES THAT DESTROY CONTEXT, AND ONLY THOSE. `compact` and
`clear` destroy it while the file survives. `resume` does not — the
context came back intact, so clearing there would re-surface every rule
after a restore that lost nothing, which is the same defect from the other
side. `startup` is a no-op against a new session id. `fork` keeps it, and
the answer holds whichever way forks are keyed: a fork carries the
conversation, so an inherited id means an accurate ledger and a new id
means an empty file.
Only the RULE ledger. The same directory holds .ids / .sync.ids /
.derive.ids for the note arms; whether a surfaced note should return after
a compaction is a different question with a different answer, and a `rm`
glob would have decided it silently.
The guard pins the DISCRIMINATION, not the deletion: the whole source
table is asserted in one statement, so a blanket delete (all False) and a
no-op (all True) both fail, and neither can be made to pass by editing one
case. A second test pins the scope against that glob, and a third proves
an event with no session id clears nothing rather than falling back to a
wildcard.
Runs with no SCRIBE_URL/SCRIBE_TOKEN on purpose — the clear is local,
keyless and networkless, and must still happen against an unreachable
instance. That is also why it sits above the config read rather than
inside the dynamic tier.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cPyzNnegXHr5iRMzzy5KJ
The manifest gate caught this, which is what it is for:
FAIL plugin content changed but the version is still 2026.09.04.0140.
An install has two halves and only one self-updates. The marketplace clone
pulls on its own; the cache that actually EXECUTES refreshes only when this
string changes. So a hook edit shipped without a bump reaches the repo and
stops there — and the obvious debugging move, inspecting the clone, shows
the fix present while the broken copy keeps running. That is #2209, #1040
and #2220, and the only detector was the operator saying "I don't think it
updated".
Minted with scripts/mint_plugin_version.py rather than hand-edited: the
plugin ships straight from the repo with no build step, so there is no
moment at which CI could stamp a value, and the script is the path the
mint guard pins.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cPyzNnegXHr5iRMzzy5KJ
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Two fixes about the same confusion in two places: treating "the reader already had this" as if it were "the ranker found nothing", and treating "was shown once" as "is known now".
#3739 —
near_missescounted a repeat as a rejectionCaught on the first live read after yesterday's deploy. The readout contradicted itself:
best_available_scoreis measured pre-threshold, which is right, but for the rule arms it is also pre-exclusion, which is not. The note arms passexclude_idsintosemantic_search_notesso their score is already post-exclusion and clean;semantic_search_rulestakes no such parameter, so the rule arms filter in Python afterwards and a rule that cleared the bar and was dropped as a repeat still reported its score on a zero-result row.That is #3497's distinction — a ranker decline versus a reader already ahead of it — reintroduced one level up, inside the field built to replace a tautology.
The population now also requires
suppressed_count IS NULL OR = 0. The NULL arm is principled rather than permissive: null means the caller filtered inside the search, exactly the case where the score cannot be contaminated. Deliberately conservative — a call carrying both a repeat and a lower genuine miss is dropped whole. It undercounts; it cannot corrupt.It also makes
near_misses.max < thresholdtrue by construction: an above-bar candidate nobody excluded would have been returned, so its call is not in the population at all.The test did not catch this, and that is the part worth keeping. The assertion was already there with the right intent; it passed because the fixture contained no suppressed call. The guard held because the breaking shape was absent, not because the code was right — rule 167's stated failure mode, in a test written while citing rule 167. The fixture now builds that shape.
#3749 — the rule ledger outlived the context it described
The hooks record every rule id they have named and hand it back as
exclude_rule_ids, so a rule is surfaced once per session and then goes quiet. Correct while the session still holds what it was told.A compaction breaks that in the worst available way: it summarises the injections out of context and does not touch the filesystem. The rule ends up absent from context and still excluded — unreachable for the rest of the session. The compaction banner tells the model to re-pull its always-on rules, but an arm-surfaced rule is conditional and is not in that set. The rules most likely to be in that state are the ones that fire most often.
Cleared on the sources that destroy context and only those:
compactandclearyes;resumeno, because the context genuinely came back and re-surfacing everything after a restore that lost nothing is the same defect from the other side.forkkeeps it, and that answer holds whichever way forks are keyed.Only the rule ledger — the note arms'
.ids/.sync.ids/.derive.idsare a different question with a different answer, and armglob would have decided it silently.The guard pins the discrimination, not the deletion: the whole source table is one assertion, so a blanket delete (all False) and a no-op (all True) both fail.
Plugin version
2026.09.09.0408, minted. The manifest gate caught the missing bump and was right to: #3749 is entirely a hook change, and without the bump it would have reached the repo and stopped — the marketplace clone updates, the cache that executes does not (#2209, #1040, #2220).So this one needs a plugin refresh as well as a deploy. #3749 does nothing until the cache picks up the new version.
CI green on
ab14f78(run 6126), all six lanes.🤖 Generated with Claude Code
https://claude.ai/code/session_011cPyzNnegXHr5iRMzzy5KJ
The manifest gate caught this, which is what it is for: FAIL plugin content changed but the version is still 2026.09.04.0140. An install has two halves and only one self-updates. The marketplace clone pulls on its own; the cache that actually EXECUTES refreshes only when this string changes. So a hook edit shipped without a bump reaches the repo and stops there — and the obvious debugging move, inspecting the clone, shows the fix present while the broken copy keeps running. That is #2209, #1040 and #2220, and the only detector was the operator saying "I don't think it updated". Minted with scripts/mint_plugin_version.py rather than hand-edited: the plugin ships straight from the repo with no build step, so there is no moment at which CI could stamp a value, and the script is the path the mint guard pins. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011cPyzNnegXHr5iRMzzy5KJ