feat(telemetry): rule_usage_events — the table, the service, and a restore that maps rule ids through the rule map (#3315)
CI & Build / Python lint (push) Failing after 3s
CI & Build / Plugin hooks (push) Successful in 8s
CI & Build / integration (push) Successful in 30s
CI & Build / TypeScript typecheck (push) Successful in 33s
CI & Build / Python tests (push) Successful in 1m5s
CI & Build / Build & push image (push) Skipped

Milestone 333 step 1. The write-path standing-rule arm is the only retrieval
surface in Scribe whose usefulness cannot be observed — and, not
coincidentally, the only one that has never declined to fire. 296 calls, zero
zero-result, 100% clearing its threshold, while every other surface declines
most of the time (#3311, and re-measured in note #3430). `retrieval_logs`
gives it scores; scores say what the ranker thought, never whether the hint
landed.

WHY A SIBLING TABLE AND NOT A COLUMN ON note_usage_events. The row carries no
note-specific field and the readout is the same shape, which is the strongest
case for sharing that note #3163 admits. What decides against it is identity at
RESTORE: the note importer maps note_id through note_id_map, so a rule id
parked in that column comes back attached to whatever note holds that number in
the target database. Not dropped — reattached. The restore reports success, the
counters are populated, and every one is about the wrong record, with no other
field to disagree with. rule_versions made the same call for the same reason;
this is the third rule-side sibling and it reads like the first two.

FK-free on rule_id and user_id, matching note_usage_events / retrieval_logs /
app_logs, and deliberately unlike rule_versions. A version belongs to a rule's
history and dies with it; telemetry outlives what it describes. Deleting a rule
must not erase the evidence that it was surfaced forty times and opened never,
because that evidence is the case for having deleted it.

The service uses `background.spawn` rather than a third copy of the
strong-reference dance — that module's own docstring says new callers should,
and a fourth copy is how one of them drifts. The AppLog canary #2663 demands is
kept, and since `rule_usage` needed exactly `note_usage`'s semantics, that
canary moved into `background.report_telemetry_failure` and note_usage now
calls it. `retrieval_telemetry` deliberately keeps its own: its canary is a
different shape (one process-wide flag, no AppLog row), so repointing it would
change behaviour rather than consolidate it.

No ambient bucket, and that is a decision. The note twin splits ranked from
ambient surfacings because enter_project and the skill sync deliver records
without choosing them (#2477). Rules have the same problem waiting —
list_always_on_rules loads them wholesale — but nothing emits here yet, so an
empty AMBIENT_SOURCES would be machinery pretending to a distinction the data
does not contain. `source` stays granular, so the split stays a readout-level
change needing no migration.

Backup carries it (v14). The round-trip test seeds a NOTE alongside the rule so
the target database has a note id to collide with — without that decoy, a
restore running rule ids through the wrong map would merely drop them and the
test would pass by absence, rather than failing on the populated-and-wrong
result that is the actual hazard.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TcCs1CcQ1ormdnzSshKqvN
This commit is contained in:
2026-09-02 16:55:22 -04:00
co-authored by Claude Opus 5
parent e029a7db64
commit 8826be7a91
10 changed files with 904 additions and 34 deletions
+118
View File
@@ -0,0 +1,118 @@
"""Rule usage telemetry — the parts that need no database (milestone 333 step 1).
The round trip lives in `test_integration_backup_rule_usage_roundtrip.py`.
What is here is the payload building and the zero shape: cheap, and the half
where a mistake is silent rather than loud.
"""
import pytest
from scribe.models.rule_usage import PULLED, SURFACED, RuleUsageEvent
from scribe.services import rule_usage
@pytest.fixture
def captured(monkeypatch):
"""Intercept the scheduler so the payload can be read without a loop.
Patching `_schedule` rather than `background.spawn` keeps the test on this
module's own seam: what is under test is which rows get built, not whether
the shared fire-and-forget machinery works — that has its own home.
"""
rows: list[list[dict]] = []
monkeypatch.setattr(rule_usage, "_schedule", rows.append)
return rows
def test_a_surfacing_records_one_row_per_rule(captured):
"""The arm shows a hint containing several rules at once; each needs its
own row, because the readout is per rule."""
rule_usage.record_rule_surfaced(
user_id=7, rule_ids=[156, 157], source="write_path_rule"
)
[batch] = captured
assert batch == [
{"user_id": 7, "rule_id": 156, "event": SURFACED, "source": "write_path_rule"},
{"user_id": 7, "rule_id": 157, "event": SURFACED, "source": "write_path_rule"},
]
def test_the_whole_hint_lands_as_one_batch(captured):
"""One scheduled insert for the hint, not one per rule. A hint is a single
decision and its rows should land together — a partial batch would read as
a hint that surfaced fewer rules than it did."""
rule_usage.record_rule_surfaced(
user_id=7, rule_ids=[1, 2, 3], source="write_path_rule"
)
assert len(captured) == 1
assert len(captured[0]) == 3
def test_a_pull_records_one_row(captured):
rule_usage.record_rule_pulled(user_id=7, rule_id=156, source="mcp_get_rule")
assert captured == [
[{"user_id": 7, "rule_id": 156, "event": PULLED, "source": "mcp_get_rule"}]
]
def test_an_actorless_event_is_still_recorded(captured):
"""The arm fires from a hook that may carry no authenticated user. Dropping
those would silently shrink the denominator the ratio divides by — the
surfacings would vanish while any later pull still counted."""
rule_usage.record_rule_surfaced(
user_id=None, rule_ids=[156], source="write_path_rule"
)
assert captured[0][0]["user_id"] is None
def test_an_empty_surfacing_builds_no_rows(captured):
"""The arm can rank everything out — `exclude_rule_ids` drops what the
session already holds. That is not a surfacing, and the empty batch is
where `_schedule` returns early rather than opening a session to insert
nothing."""
rule_usage.record_rule_surfaced(user_id=7, rule_ids=[], source="write_path_rule")
assert captured == [[]]
def test_the_real_scheduler_returns_early_on_an_empty_batch():
"""The guard itself, against the REAL `_schedule` the stub above replaces.
There is no running loop in a unit test, so `spawn` would be harmless
anyway — but it would build a coroutine only to close it, and the point is
that an empty batch never gets that far.
"""
rule_usage._schedule([]) # must not raise
def test_a_bad_rule_id_is_dropped_not_raised(captured):
"""Telemetry must never break the surface it observes. An unconvertible id
is a bug somewhere upstream, and the right response is to lose the row and
log it — not to take down the write-path hint."""
rule_usage.record_rule_pulled(
user_id=7, rule_id="not-an-int", source="mcp_get_rule" # type: ignore[arg-type]
)
assert captured == []
def test_the_zero_readout_names_every_key():
"""Callers render this shape unconditionally. Every rule in an existing
install predates the table, so for a while "no events" is the NORMAL state
— a missing key here would read as a broken readout on almost every row."""
assert rule_usage.empty_rule_usage() == {
"surfaced_count": 0,
"pull_count": 0,
"last_surfaced_at": None,
"last_pulled_at": None,
}
def test_the_model_serialises_the_fields_the_ratio_needs():
ev = RuleUsageEvent(
user_id=7, rule_id=156, event=SURFACED, source="write_path_rule"
)
row = ev.to_dict()
assert row["rule_id"] == 156
assert row["event"] == SURFACED
assert row["source"] == "write_path_rule"
# created_at is server-defaulted, so it is None until the row is flushed —
# `iso()` must tolerate that rather than raising on a fresh instance.
assert row["created_at"] is None