feat(telemetry): every counter says when it started being recorded (#3712)
CI & Build / Python lint (push) Failing after 3s
CI & Build / Plugin hooks (push) Successful in 8s
CI & Build / integration (push) Failing after 17s
CI & Build / Python tests (push) Failing after 26s
CI & Build / TypeScript typecheck (push) Successful in 34s
CI & Build / Build & push image (push) Skipped
CI & Build / Python lint (push) Failing after 3s
CI & Build / Plugin hooks (push) Successful in 8s
CI & Build / integration (push) Failing after 17s
CI & Build / Python tests (push) Failing after 26s
CI & Build / TypeScript typecheck (push) Successful in 34s
CI & Build / Build & push image (push) Skipped
A window that opens before a counter existed reports that counter as though it had been measured throughout. The reader cannot tell "zero because nothing happened" from "zero because nobody was counting yet", and — worse — cannot tell a partial count from a complete one. That middle case yields a plausible FRACTION rather than an obvious zero, which is what makes it dangerous. It is not hypothetical. A 7-day window opened while the ranked rule surfacing recorders were four days old produced an apparent 64% write loss, which survived a code review, four ruled-out alternative causes and a five-step milestone before an identity check falsified it in one read. Every counter block now carries `complete_from` and `covers_window`. THE GRAIN IS THE SOURCE. retrieval_logs accumulates for months, so a per-table earliest row says months for every source it holds — including an arm added days ago whose counter means something else entirely. The old source would vouch for the young one, which is the exact reading this prevents. A SECTION TAKES ITS LATEST CONTRIBUTOR, NOT ITS EARLIEST. A figure summing several sources is complete only once every one of them was being written, so "*" is a max. Using min would reproduce the original error in miniature. `covers_window` is null, never false, when nothing was ever recorded: "no measurement" is not "partial measurement" — the null convention #3497 established for `suppression`, one level up. Also corrects a stale claim in the tool docstring: it still taught readers that write_path_rule "has never once declined to fire" (#3311). That was the arm writing its retrieval_logs row only on calls that found something; #3497 fixed it, and the arm declines the large majority of its calls. _complete_from takes the caller's session rather than opening its own, departing from the services canon (#2860) because it runs inside an existing block; to be recorded against the ledger once it ingests. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011cPyzNnegXHr5iRMzzy5KJ
This commit is contained in:
@@ -638,3 +638,148 @@ async def test_ambient_alone_reports_no_ratio(_dispose_engine):
|
||||
assert ru["pull_through"] is None
|
||||
finally:
|
||||
await cleanup()
|
||||
|
||||
|
||||
# ── Window coverage (#3712) ────────────────────────────────────────────
|
||||
#
|
||||
# A counter added last week, read over a 30-day window, reports a real count
|
||||
# against an imagined denominator. The result is a plausible FRACTION rather
|
||||
# than an obvious zero, which is what makes it dangerous — #379 spent five
|
||||
# planned steps on a defect that turned out to be a window opening before the
|
||||
# recording it was measuring existed.
|
||||
|
||||
|
||||
def test_coverage_says_nothing_rather_than_false_when_nothing_was_recorded():
|
||||
"""Null, never False. "No measurement" is not "partial measurement".
|
||||
|
||||
The same distinction `suppression`'s null carries (#3497): absent must not
|
||||
read as a verdict. A False here would assert the window is under-covered,
|
||||
which is a claim nobody is in a position to make.
|
||||
"""
|
||||
from datetime import datetime, timezone
|
||||
|
||||
from scribe.services.retrieval_telemetry import _coverage
|
||||
|
||||
since = datetime(2026, 9, 1, tzinfo=timezone.utc)
|
||||
assert _coverage(None, since) == {
|
||||
"complete_from": None, "covers_window": None,
|
||||
}
|
||||
|
||||
|
||||
def test_coverage_reads_a_start_before_the_window_as_covered():
|
||||
from datetime import datetime, timezone
|
||||
|
||||
from scribe.services.retrieval_telemetry import _coverage
|
||||
|
||||
since = datetime(2026, 9, 1, tzinfo=timezone.utc)
|
||||
older = datetime(2026, 8, 1, tzinfo=timezone.utc)
|
||||
newer = datetime(2026, 9, 5, tzinfo=timezone.utc)
|
||||
|
||||
assert _coverage(older, since)["covers_window"] is True
|
||||
assert _coverage(newer, since)["covers_window"] is False, (
|
||||
"a counter that started inside the window covers only part of it"
|
||||
)
|
||||
assert _coverage(newer, since)["complete_from"] == newer.isoformat()
|
||||
|
||||
|
||||
@pytest.mark.integration
|
||||
@pytest.mark.asyncio
|
||||
async def test_coverage_is_per_source_because_the_table_is_older_than_its_arms(
|
||||
_dispose_engine,
|
||||
):
|
||||
"""THE grain question, and the reason a per-table answer is useless.
|
||||
|
||||
`retrieval_logs` accumulates for months. A table-level "earliest row"
|
||||
therefore says months for every source it holds — including one added
|
||||
days ago whose counter means something quite different. The old source
|
||||
would vouch for the young one, which is exactly the reading this exists
|
||||
to prevent.
|
||||
"""
|
||||
from datetime import datetime, timedelta, timezone
|
||||
|
||||
from sqlalchemy import delete
|
||||
|
||||
from scribe.models import async_session
|
||||
from scribe.models.retrieval_log import RetrievalLog
|
||||
from scribe.services.retrieval_telemetry import retrieval_summary
|
||||
|
||||
UID = 990077
|
||||
now = datetime.now(timezone.utc)
|
||||
async with async_session() as s:
|
||||
# An old surface, recording since well before any window we ask for.
|
||||
s.add(RetrievalLog(
|
||||
user_id=UID, source="auto_inject", result_count=1,
|
||||
created_at=now - timedelta(days=90),
|
||||
))
|
||||
# A young arm, first written INSIDE the window below.
|
||||
s.add(RetrievalLog(
|
||||
user_id=UID, source="pre_tool_rule", result_count=1,
|
||||
created_at=now - timedelta(days=2),
|
||||
))
|
||||
await s.commit()
|
||||
|
||||
try:
|
||||
out = await retrieval_summary(UID, days=30)
|
||||
|
||||
assert out["sources"]["auto_inject"]["covers_window"] is True
|
||||
assert out["sources"]["pre_tool_rule"]["covers_window"] is False, (
|
||||
"the young arm was reported as covering a 30-day window — the "
|
||||
"table's age has been allowed to vouch for one of its sources"
|
||||
)
|
||||
finally:
|
||||
async with async_session() as s:
|
||||
await s.execute(delete(RetrievalLog).where(RetrievalLog.user_id == UID))
|
||||
await s.commit()
|
||||
|
||||
|
||||
@pytest.mark.integration
|
||||
@pytest.mark.asyncio
|
||||
async def test_a_section_is_complete_only_from_its_latest_contributor(
|
||||
_dispose_engine,
|
||||
):
|
||||
"""A sum is complete once EVERY contributor was being written — so the
|
||||
section takes the LATEST first-row, not the earliest.
|
||||
|
||||
Taking the earliest would be worse than reporting nothing: it would pick
|
||||
the oldest source in the table and use it to certify a total that a
|
||||
newer source is still only partly contributing to. That is the original
|
||||
error in miniature.
|
||||
"""
|
||||
from datetime import datetime, timedelta, timezone
|
||||
|
||||
from sqlalchemy import delete
|
||||
|
||||
from scribe.models import async_session
|
||||
from scribe.models.rule_usage import RuleUsageEvent
|
||||
from scribe.services.retrieval_telemetry import retrieval_summary
|
||||
|
||||
UID = 990078
|
||||
now = datetime.now(timezone.utc)
|
||||
old = now - timedelta(days=90)
|
||||
young = now - timedelta(days=2)
|
||||
async with async_session() as s:
|
||||
s.add_all([
|
||||
RuleUsageEvent(
|
||||
user_id=UID, rule_id=1, event="surfaced",
|
||||
source="list_always_on_rules", created_at=old,
|
||||
),
|
||||
RuleUsageEvent(
|
||||
user_id=UID, rule_id=2, event="surfaced",
|
||||
source="pre_tool_rule", created_at=young,
|
||||
),
|
||||
])
|
||||
await s.commit()
|
||||
|
||||
try:
|
||||
out = await retrieval_summary(UID, days=30)
|
||||
ru = out["rule_usage"]
|
||||
|
||||
assert ru["complete_from"] == young.isoformat(), (
|
||||
"the section reported completeness from its OLDEST source; a "
|
||||
"total is only as complete as its newest contributor"
|
||||
)
|
||||
assert ru["covers_window"] is False
|
||||
finally:
|
||||
async with async_session() as s:
|
||||
await s.execute(delete(RuleUsageEvent).where(RuleUsageEvent.user_id == UID))
|
||||
await s.commit()
|
||||
|
||||
Reference in New Issue
Block a user