CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / TypeScript typecheck (push) Successful in 23s
CI & Build / integration (push) Successful in 31s
CI & Build / Python tests (push) Successful in 1m6s
CI & Build / Build & push image (push) Successful in 27s
The ranked rule arm became measurable in M333. The preload did not — and that is the surface whose value is actually in question. `list_always_on_rules`, the SessionStart block and every `rules_payload` caller handed rules over wholesale and emitted nothing, so the resident set's token cost was certain and its usefulness could not be tested even in principle. Bulk deliveries now record as AMBIENT, beside the ranked count and never inside pull-through. Folding them in would mean growing the always-on set depressed the arm's measured precision and trimming it flattered the arm, neither for any reason to do with the arm. `RANKED_SOURCES` inverts the note twin's `AMBIENT_SOURCES` deliberately: there is one ranked rule source and this change adds seven bulk ones, so naming the rare half makes a forgotten surface default to ambient — under-counting it — rather than padding the denominator with surfacings nobody chose. Two lookalike call sites are deliberately left silent, with a test to keep them that way: the write-path etag arm and `rules_etag_for` read the rules to build or compare a MARKER and show nobody anything. No migration — `event` and `source` are plain Text with no CHECK (rule 36 does not apply). Snippet #2858 updated to the new `rules_payload` contract. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011cPyzNnegXHr5iRMzzy5KJ
281 lines
12 KiB
Python
281 lines
12 KiB
Python
"""Rule usage telemetry — did a surfaced rule ever get read?
|
|
|
|
The sibling of `note_usage`, for the one retrieval surface in Scribe that
|
|
could not be measured at all.
|
|
|
|
Two event streams, deliberately independent:
|
|
|
|
- SURFACED: the write-path standing-rule arm put this rule in front of the
|
|
agent, unbidden, during a write.
|
|
- PULLED: someone then opened it in full (`get_rule`, or the REST detail
|
|
route).
|
|
|
|
WHY THIS ARM AND NOT ANOTHER. Every other surface declines most of the time —
|
|
`write_path` returns nothing on 78% of calls, `reuse_slot` on 79%, auto-inject
|
|
on 39%. The rule arm has never once returned nothing (#3311). That is either a
|
|
perfectly tuned surface or a bar it cannot fail to clear, and `retrieval_logs`
|
|
cannot tell the two apart: it records what the ranker scored, never whether the
|
|
hint was any use. The ratio these two streams produce is the missing half, and
|
|
without it any threshold change is a number picked off a histogram.
|
|
|
|
Design notes, mirroring `note_usage`:
|
|
- Writes are fire-and-forget through `background.spawn`, so telemetry never
|
|
adds latency to — or can break — the surface it observes. This module does
|
|
NOT carry its own copy of the strong-reference dance; `background` is the
|
|
one place that gets it right, and a fourth copy is how one of them drifts.
|
|
- Failures degrade, but never SILENTLY. `report_telemetry_failure` logs at
|
|
WARNING and drops one AppLog row per process per site. #2663 is the record
|
|
of this exact subsystem class running at zero for weeks — indistinguishable
|
|
from "nobody uses this" — because every failure went to `logger.debug`.
|
|
- Reads (`usage_for_rules`) are awaited and aggregated in one round-trip for
|
|
a whole page, never per row.
|
|
|
|
AMBIENT VS RANKED. The note twin splits ranked surfacings from ambient ones
|
|
because `enter_project` and the skill sync put records in front of the agent
|
|
without choosing them, and counting those as surfacings makes recency read as
|
|
popularity (#2477). Rules have exactly that shape: the SessionStart preload,
|
|
`list_always_on_rules`, and every `rules_payload` surface hand over the whole
|
|
applicable set at once, chosen by nobody.
|
|
|
|
Until 2026-09-03 those bulk surfaces emitted nothing, and this module said so —
|
|
"an empty `AMBIENT_SOURCES` would be machinery pretending to a distinction the
|
|
data does not yet contain". True as far as it went, but it had a consequence
|
|
worth naming, because it is the reason the bucket exists now: the always-on
|
|
set's token cost was certain and its usefulness was UNFALSIFIABLE, permanently
|
|
and by construction. The one surface whose value was actually in question was
|
|
the one surface exempt from the scoreboard that judges every other.
|
|
|
|
They emit now. The split is the readout-level change the old note promised — a
|
|
`case()`, no migration, because `event` and `source` are plain Text with no
|
|
CHECK constraint. `source` stays granular so a reader can still tell the
|
|
preload from `enter_project` from the ranked arm.
|
|
|
|
WHY THIS NAMES THE RANKED SOURCES AND THE TWIN NAMES THE AMBIENT ONES. A
|
|
deliberate divergence, on the failure mode rather than on symmetry. Both shapes
|
|
fail silently when someone adds a surface and forgets the list, so the question
|
|
is which list changes more often — and here it is emphatically the ambient one:
|
|
there is exactly ONE ranked rule source, and this change alone adds seven bulk
|
|
ones. Naming the rare, stable half means a newly-added bulk surface defaults to
|
|
`ambient`, which merely under-counts it, instead of defaulting to `ranked`,
|
|
which would quietly pad the pull-through denominator with surfacings nobody
|
|
chose and make the arm look imprecise. Same argument #3191 and #3430 make
|
|
against hand-kept lists: keep the list that must be remembered as short and as
|
|
slow-moving as possible.
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
import logging
|
|
|
|
from sqlalchemy import case, func, select
|
|
|
|
from scribe.models import async_session
|
|
from scribe.models.base import iso
|
|
from scribe.models.rule_usage import PULLED, SURFACED, RuleUsageEvent
|
|
from scribe.services.background import report_telemetry_failure, spawn
|
|
|
|
logger = logging.getLogger(__name__)
|
|
|
|
# The surfaces that CHOSE the rules they showed. Everything else is ambient —
|
|
# see the module docstring for why the rare half is the half that gets named.
|
|
#
|
|
# Membership is the whole definition of the pull-through denominator: a ranked
|
|
# surfacing is a claim ("this rule may apply to what you are doing") that a pull
|
|
# can confirm or refute, while an ambient one is a delivery nobody decided on.
|
|
# Add a source here only when a ranker picked it.
|
|
RANKED_SOURCES = ("write_path_rule",)
|
|
|
|
|
|
def is_ambient(source: str) -> bool:
|
|
"""Was this surfacing a bulk delivery rather than a ranked choice?
|
|
|
|
One definition, read by both the per-rule badge readout and the aggregate
|
|
in `retrieval_telemetry` — the two used to be able to disagree about what
|
|
"surfaced" counted, which is the class of drift #3246 found across the
|
|
rules system.
|
|
|
|
Sync and pure, per the service canon (#2860), but deliberately PUBLIC where
|
|
that canon says such helpers stay `_private`. The departure is the point:
|
|
a module-private copy in each caller is exactly the second definition this
|
|
exists to prevent.
|
|
"""
|
|
return source not in RANKED_SOURCES
|
|
|
|
|
|
async def _report_failure(site: str) -> None:
|
|
await report_telemetry_failure("rule_usage", site)
|
|
|
|
|
|
async def _insert_events(rows: list[dict]) -> None:
|
|
"""Persist usage rows. Best-effort: failures degrade, visibly."""
|
|
try:
|
|
async with async_session() as session:
|
|
session.add_all([RuleUsageEvent(**row) for row in rows])
|
|
await session.commit()
|
|
except Exception:
|
|
await _report_failure("write")
|
|
|
|
|
|
def _schedule(rows: list[dict]) -> None:
|
|
if not rows:
|
|
return
|
|
spawn(_insert_events(rows), site="rule_usage_write")
|
|
|
|
|
|
def record_rule_surfaced(
|
|
*, user_id: int | None, rule_ids: list[int] | set[int], source: str
|
|
) -> None:
|
|
"""Fire-and-forget: record that these rules were shown to the agent.
|
|
|
|
Takes the whole delivery at once — one insert per surfacing event, not per
|
|
rule — because a hint is a single decision and its rows should land
|
|
together.
|
|
|
|
Record what was actually SHOWN, never what was considered. For the ranked
|
|
arm that means the post-filter hits: it drops what the session already
|
|
holds (`exclude_rule_ids`) before it speaks, and a rule considered and not
|
|
shown was not surfaced. Counting those would inflate the denominator with
|
|
claims the agent never saw, which reads as a precision problem the arm does
|
|
not have.
|
|
|
|
Bulk surfaces pass their whole delivered set, which is the same rule read
|
|
from the other end — everything in a preload IS shown. `source` is what
|
|
separates the two afterwards (see `RANKED_SOURCES`); this function does not
|
|
care which kind it is recording.
|
|
"""
|
|
try:
|
|
rows = [
|
|
{
|
|
"user_id": user_id,
|
|
"rule_id": int(rid),
|
|
"event": SURFACED,
|
|
"source": source,
|
|
}
|
|
for rid in rule_ids
|
|
]
|
|
except Exception:
|
|
logger.debug("rule usage payload build failed", exc_info=True)
|
|
return
|
|
_schedule(rows)
|
|
|
|
|
|
def record_rule_pulled(*, user_id: int | None, rule_id: int, source: str) -> None:
|
|
"""Fire-and-forget: record that a rule was opened in full.
|
|
|
|
A PULL is somebody choosing to open one record. `list_always_on_rules` and
|
|
`enter_project` are NOT pulls — they are bulk resident loads that hand over
|
|
every applicable rule at once, and counting them would swamp the signal
|
|
with the very ambient delivery the ratio exists to distinguish from.
|
|
"""
|
|
try:
|
|
rows = [
|
|
{
|
|
"user_id": user_id,
|
|
"rule_id": int(rule_id),
|
|
"event": PULLED,
|
|
"source": source,
|
|
}
|
|
]
|
|
except Exception:
|
|
logger.debug("rule usage payload build failed", exc_info=True)
|
|
return
|
|
_schedule(rows)
|
|
|
|
|
|
def empty_rule_usage() -> dict:
|
|
"""The zero readout — what a rule with no recorded events looks like.
|
|
|
|
Callers render this shape unconditionally, so a rule predating the table
|
|
reads as "never surfaced, never pulled" rather than as a missing key. That
|
|
distinction matters more here than for notes: every rule in an install
|
|
predates this table, so for a while "no events" is the normal state and it
|
|
must not look like a broken readout.
|
|
|
|
`surfaced_count` is RANKED surfacings only; `ambient_count` is the bulk
|
|
deliveries (see `RANKED_SOURCES`). The split is what keeps the badge's
|
|
"shown often, opened never → dead weight" reading honest: every rule in an
|
|
always-on set is delivered every session, so an unsplit counter would rank
|
|
the resident set as the most-surfaced rules in the install purely for being
|
|
resident.
|
|
"""
|
|
return {
|
|
"surfaced_count": 0,
|
|
"ambient_count": 0,
|
|
"pull_count": 0,
|
|
"last_surfaced_at": None,
|
|
"last_pulled_at": None,
|
|
}
|
|
|
|
|
|
async def usage_for_rules(rule_ids: list[int]) -> dict[int, dict]:
|
|
"""Aggregate usage for a set of rules: {rule_id: {counts + timestamps}}.
|
|
|
|
One GROUP BY for the whole page rather than a query per row — this feeds a
|
|
list view, so the per-row shape would be N+1 by construction. Rules with no
|
|
events come back with `empty_rule_usage()`, so the caller never has to tell
|
|
"no events" from "not in the result".
|
|
"""
|
|
ids = [int(r) for r in rule_ids]
|
|
out: dict[int, dict] = {rid: empty_rule_usage() for rid in ids}
|
|
if not ids:
|
|
return out
|
|
|
|
# Classified in SQL so the group stays small: per rule we get at most
|
|
# (surfaced-ranked, surfaced-ambient, pulled) rather than a row per distinct
|
|
# source. ONE labelled expression, bound to a variable and reused in the
|
|
# GROUP BY — a second `case()` instance there renders its own expanding-IN
|
|
# bind names under asyncpg, so the database sees two DIFFERENT expressions
|
|
# and rejects the query with a GroupingError. The note twin carries the
|
|
# same warning for the same reason, and #2663 is what it cost: the
|
|
# rejection was swallowed and every counter read zero in production while
|
|
# the writes were landing fine.
|
|
ambient = case(
|
|
(RuleUsageEvent.source.notin_(RANKED_SOURCES), True),
|
|
else_=False,
|
|
).label("ambient")
|
|
try:
|
|
async with async_session() as session:
|
|
rows = (
|
|
await session.execute(
|
|
select(
|
|
RuleUsageEvent.rule_id,
|
|
RuleUsageEvent.event,
|
|
func.count().label("n"),
|
|
func.max(RuleUsageEvent.created_at).label("last_at"),
|
|
ambient,
|
|
)
|
|
.where(RuleUsageEvent.rule_id.in_(ids))
|
|
.group_by(
|
|
RuleUsageEvent.rule_id,
|
|
RuleUsageEvent.event,
|
|
ambient,
|
|
)
|
|
)
|
|
).all()
|
|
except Exception:
|
|
# A telemetry readout must not be able to break the list it decorates —
|
|
# but it must say it failed, or a broken readout is indistinguishable
|
|
# from a corpus nobody uses (#2663).
|
|
await _report_failure("readout")
|
|
return out
|
|
|
|
for rule_id, event, n, last_at, is_amb in rows:
|
|
slot = out.get(int(rule_id))
|
|
if slot is None:
|
|
continue
|
|
if event == SURFACED and is_amb:
|
|
slot["ambient_count"] = int(n)
|
|
elif event == SURFACED:
|
|
slot["surfaced_count"] = int(n)
|
|
slot["last_surfaced_at"] = iso(last_at)
|
|
elif event == PULLED:
|
|
# Pulls are pulls regardless of what surfaced the rule — "did
|
|
# anyone ever open this?" does not depend on how it was found. Both
|
|
# halves accumulate, so this ADDS rather than assigns: a rule can
|
|
# now be pulled after a ranked hint and after a preload, and the
|
|
# split arrives as two rows.
|
|
slot["pull_count"] = slot["pull_count"] + int(n)
|
|
latest = iso(last_at)
|
|
if latest and (slot["last_pulled_at"] or "") < latest:
|
|
slot["last_pulled_at"] = latest
|
|
return out
|