A repeat is not a rejection, and a compaction is not knowledge #145

Merged
bvandeusen merged 3 commits from dev into main 2026-09-09 00:13:02 -04:00
3 Commits
Author SHA1 Message Date
bvandeusenandClaude Opus 5 ab14f783e1 chore(plugin): mint 2026.09.09.0408 — the hook change has to reach the cache (#3749)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / TypeScript typecheck (push) Successful in 33s
CI & Build / integration (push) Successful in 36s
CI & Build / Python tests (push) Successful in 1m8s
CI & Build / Build & push image (push) Successful in 16s
The manifest gate caught this, which is what it is for:

    FAIL  plugin content changed but the version is still 2026.09.04.0140.

An install has two halves and only one self-updates. The marketplace clone
pulls on its own; the cache that actually EXECUTES refreshes only when this
string changes. So a hook edit shipped without a bump reaches the repo and
stops there — and the obvious debugging move, inspecting the clone, shows
the fix present while the broken copy keeps running. That is #2209, #1040
and #2220, and the only detector was the operator saying "I don't think it
updated".

Minted with scripts/mint_plugin_version.py rather than hand-edited: the
plugin ships straight from the repo with no build step, so there is no
moment at which CI could stamp a value, and the script is the path the
mint guard pins.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cPyzNnegXHr5iRMzzy5KJ
2026-09-09 00:08:55 -04:00
bvandeusenandClaude Opus 5 1cfbf43ccd fix(plugin): the rule ledger clears when the context it describes is destroyed (#3749)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Failing after 8s
CI & Build / TypeScript typecheck (push) Successful in 11s
CI & Build / integration (push) Successful in 36s
CI & Build / Python tests (push) Successful in 1m6s
CI & Build / Build & push image (push) Successful in 21s
The prior-art and tool-rule hooks record every rule id they have named in
<state>/<sid>.rules.ids and hand it back as exclude_rule_ids, so a rule is
surfaced once per session and then goes quiet. That is correct while the
session still HOLDS what it was told.

A compaction breaks it in the worst available way: it summarizes the
earlier injections out of context and does not touch the filesystem. The
rule ends up absent from context AND still excluded — unreachable for the
rest of the session. The compaction banner this hook already prints tells
the model to re-pull its ALWAYS-ON rules, but a rule an arm surfaced is
conditional and is not in that set, so it has no other way back. The rules
most likely to be in that state are the ones that fire most often.

The stale ledger is genuinely found again rather than orphaned: the etag
marker further down this same hook is rewritten on `compact` and keyed by
session_id, which is only meaningful if the id survives a compaction.

CLEARED ON THE SOURCES THAT DESTROY CONTEXT, AND ONLY THOSE. `compact` and
`clear` destroy it while the file survives. `resume` does not — the
context came back intact, so clearing there would re-surface every rule
after a restore that lost nothing, which is the same defect from the other
side. `startup` is a no-op against a new session id. `fork` keeps it, and
the answer holds whichever way forks are keyed: a fork carries the
conversation, so an inherited id means an accurate ledger and a new id
means an empty file.

Only the RULE ledger. The same directory holds .ids / .sync.ids /
.derive.ids for the note arms; whether a surfaced note should return after
a compaction is a different question with a different answer, and a `rm`
glob would have decided it silently.

The guard pins the DISCRIMINATION, not the deletion: the whole source
table is asserted in one statement, so a blanket delete (all False) and a
no-op (all True) both fail, and neither can be made to pass by editing one
case. A second test pins the scope against that glob, and a third proves
an event with no session id clears nothing rather than falling back to a
wildcard.

Runs with no SCRIBE_URL/SCRIBE_TOKEN on purpose — the clear is local,
keyless and networkless, and must still happen against an unreachable
instance. That is also why it sits above the config read rather than
inside the dynamic tier.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cPyzNnegXHr5iRMzzy5KJ
2026-09-09 00:04:33 -04:00
bvandeusenandClaude Opus 5 a165483b92 fix(telemetry): a repeat is not a rejection, and near_misses counted it as one (#3739)
CI & Build / Build & push image (push) Successful in 31s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / TypeScript typecheck (push) Successful in 35s
CI & Build / integration (push) Successful in 45s
CI & Build / Python tests (push) Successful in 1m27s
Caught on the first live read after deploying #3670. The readout
contradicted itself:

    pre_tool_rule   top_score.min    0.7204   the lowest score ever RETURNED
                    near_misses.max  0.7457   "rejected", but scored higher

`best_available_score` is measured pre-threshold, which is right, but for
the rule arms it is also PRE-EXCLUSION, which is not. The note arms pass
`exclude_ids` into semantic_search_notes so their score is already
post-exclusion and clean; `semantic_search_rules` takes no such parameter,
so the rule arms filter in Python after the search and a rule that cleared
the bar and was dropped as a repeat still reported its score on a
zero-result row.

That is #3497's distinction — a ranker decline versus a reader already
ahead of it — reintroduced one level up, inside the field built to replace
a tautology.

The population now also requires `suppressed_count IS NULL OR = 0`. The
NULL arm is principled rather than permissive: null means the caller
filtered INSIDE the search, which is exactly the case where the reported
score cannot be contaminated.

Deliberately conservative — a call carrying both a repeat and a lower
genuine miss is dropped whole, losing that point. It undercounts; it
cannot corrupt, which is the right way round for a number read against a
bar.

It also makes `near_misses.max < threshold` true BY CONSTRUCTION rather
than by fixture: an above-bar candidate nobody excluded would have been
returned, so its call is not in the population at all.

THE TEST DID NOT CATCH THIS, and that is the part worth keeping. The
assertion `nm["max"] < 0.72` was already there, with exactly the right
intent. It passed because the fixture contained no suppressed call — the
guard held because the breaking shape was absent, not because the code was
right. Rule 167's stated failure mode, in a test written while citing rule
167. The fixture now builds that shape: a 0.9 hit dropped as a repeat,
which lands in the population and drags `max` above the threshold unless
the predicate excludes it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cPyzNnegXHr5iRMzzy5KJ
2026-09-08 16:42:51 -04:00