CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / TypeScript typecheck (push) Successful in 22s
CI & Build / integration (push) Successful in 33s
CI & Build / Python tests (push) Failing after 48s
CI & Build / Build & push image (push) Skipped
`write_path_rule` reported `zero_result_calls: 0` and `cleared_threshold: 133/133` — a perfect record no other surface comes near (`write_path` 421 zeroes of 613, `reuse_slot` 124/199, `auto_inject` 114/326). #3311 read that as a measurement and milestone 333 was scoped on it. It was an artifact. Both arms called `record_retrieval` inside a guard on having results — the write-path arm behind `if fresh:`, the pre-tool arm below `if not fresh: return out` — so a call that found nothing wrote no row. The statistic was a fact about the shape of the code, true at any threshold whatsoever. The call log moves out of the guard in both arms. The surfacing log stays in it: nothing was shown, so no surfacing occurred. `results=fresh` is kept deliberately — the note arms pass exclusions into `semantic_search_notes`, so what they log is already post-exclusion, and logging `hits` here would make this row mean something other than every other row in the same readout. The defect bites hardest on the pre-tool arm, which fires on every Bash call: with no rows at all, a ranker that declined is indistinguishable from a hook that never fired — the silent failure the arm exists to stop. Tests cover both arms behaviourally (found nothing; found only what the session already held; searched nothing at all, which must stay silent) plus a structural guard, because this was one level of indentation and it appeared independently in two places. #3311 and the `rule_usage` docstring corrected rather than quietly rewritten. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011cPyzNnegXHr5iRMzzy5KJ
296 lines
13 KiB
Python
296 lines
13 KiB
Python
"""Rule usage telemetry — did a surfaced rule ever get read?
|
|
|
|
The sibling of `note_usage`, for the one retrieval surface in Scribe that
|
|
could not be measured at all.
|
|
|
|
Two event streams, deliberately independent:
|
|
|
|
- SURFACED: the write-path standing-rule arm put this rule in front of the
|
|
agent, unbidden, during a write.
|
|
- PULLED: someone then opened it in full (`get_rule`, or the REST detail
|
|
route).
|
|
|
|
WHY THIS ARM AND NOT ANOTHER. Every other surface declines most of the time —
|
|
`write_path` returns nothing on 78% of calls, `reuse_slot` on 79%, auto-inject
|
|
on 39%. The rule arm APPEARED never to have returned nothing (#3311), and this
|
|
docstring used to put that forward as the puzzle worth measuring: "either a
|
|
perfectly tuned surface or a bar it cannot fail to clear".
|
|
|
|
It was neither, and the correction belongs here rather than being quietly
|
|
deleted. The arm wrote its `retrieval_logs` row only on calls that FOUND
|
|
something (#3497), so `zero_result_calls` sat at 0 and `cleared_threshold` at
|
|
`calls` because of the shape of the code — at any threshold whatsoever. A
|
|
statistic that could not vary was read as a finding about the corpus. It is the
|
|
#2663 failure mode one level up: there the broken readout was a zero, here it
|
|
was a hundred percent, which is far better camouflage.
|
|
|
|
The reason to measure this arm survives the correction, and is stronger for it.
|
|
`retrieval_logs` records what the ranker scored, never whether the hint was any
|
|
use, so even an honest clear-rate would not settle the question. The ratio these
|
|
two streams produce is the missing half, and without it any threshold change is
|
|
a number picked off a histogram.
|
|
|
|
Design notes, mirroring `note_usage`:
|
|
- Writes are fire-and-forget through `background.spawn`, so telemetry never
|
|
adds latency to — or can break — the surface it observes. This module does
|
|
NOT carry its own copy of the strong-reference dance; `background` is the
|
|
one place that gets it right, and a fourth copy is how one of them drifts.
|
|
- Failures degrade, but never SILENTLY. `report_telemetry_failure` logs at
|
|
WARNING and drops one AppLog row per process per site. #2663 is the record
|
|
of this exact subsystem class running at zero for weeks — indistinguishable
|
|
from "nobody uses this" — because every failure went to `logger.debug`.
|
|
- Reads (`usage_for_rules`) are awaited and aggregated in one round-trip for
|
|
a whole page, never per row.
|
|
|
|
AMBIENT VS RANKED. The note twin splits ranked surfacings from ambient ones
|
|
because `enter_project` and the skill sync put records in front of the agent
|
|
without choosing them, and counting those as surfacings makes recency read as
|
|
popularity (#2477). Rules have exactly that shape: the SessionStart preload,
|
|
`list_always_on_rules`, and every `rules_payload` surface hand over the whole
|
|
applicable set at once, chosen by nobody.
|
|
|
|
Until 2026-09-03 those bulk surfaces emitted nothing, and this module said so —
|
|
"an empty `AMBIENT_SOURCES` would be machinery pretending to a distinction the
|
|
data does not yet contain". True as far as it went, but it had a consequence
|
|
worth naming, because it is the reason the bucket exists now: the always-on
|
|
set's token cost was certain and its usefulness was UNFALSIFIABLE, permanently
|
|
and by construction. The one surface whose value was actually in question was
|
|
the one surface exempt from the scoreboard that judges every other.
|
|
|
|
They emit now. The split is the readout-level change the old note promised — a
|
|
`case()`, no migration, because `event` and `source` are plain Text with no
|
|
CHECK constraint. `source` stays granular so a reader can still tell the
|
|
preload from `enter_project` from the ranked arm.
|
|
|
|
WHY THIS NAMES THE RANKED SOURCES AND THE TWIN NAMES THE AMBIENT ONES. A
|
|
deliberate divergence, on the failure mode rather than on symmetry. Both shapes
|
|
fail silently when someone adds a surface and forgets the list, so the question
|
|
is which list changes more often — and here it is emphatically the ambient one:
|
|
there are TWO ranked rule sources (the write-path arm and the pre-tool arm)
|
|
against the seven bulk ones the preload alone contributes. Ranked sources are
|
|
added when somebody builds a ranker, which is rare and deliberate; bulk ones
|
|
appear whenever a surface hands rules over, which is most of them. Naming the
|
|
rare, slow-moving half means a newly-added bulk surface defaults to
|
|
`ambient`, which merely under-counts it, instead of defaulting to `ranked`,
|
|
which would quietly pad the pull-through denominator with surfacings nobody
|
|
chose and make the arm look imprecise. Same argument #3191 and #3430 make
|
|
against hand-kept lists: keep the list that must be remembered as short and as
|
|
slow-moving as possible.
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
import logging
|
|
|
|
from sqlalchemy import case, func, select
|
|
|
|
from scribe.models import async_session
|
|
from scribe.models.base import iso
|
|
from scribe.models.rule_usage import PULLED, SURFACED, RuleUsageEvent
|
|
from scribe.services.background import report_telemetry_failure, spawn
|
|
|
|
logger = logging.getLogger(__name__)
|
|
|
|
# The surfaces that CHOSE the rules they showed. Everything else is ambient —
|
|
# see the module docstring for why the rare half is the half that gets named.
|
|
#
|
|
# Membership is the whole definition of the pull-through denominator: a ranked
|
|
# surfacing is a claim ("this rule may apply to what you are doing") that a pull
|
|
# can confirm or refute, while an ambient one is a delivery nobody decided on.
|
|
# Add a source here only when a ranker picked it.
|
|
RANKED_SOURCES = ("write_path_rule", "pre_tool_rule")
|
|
|
|
|
|
def is_ambient(source: str) -> bool:
|
|
"""Was this surfacing a bulk delivery rather than a ranked choice?
|
|
|
|
One definition, read by both the per-rule badge readout and the aggregate
|
|
in `retrieval_telemetry` — the two used to be able to disagree about what
|
|
"surfaced" counted, which is the class of drift #3246 found across the
|
|
rules system.
|
|
|
|
Sync and pure, per the service canon (#2860), but deliberately PUBLIC where
|
|
that canon says such helpers stay `_private`. The departure is the point:
|
|
a module-private copy in each caller is exactly the second definition this
|
|
exists to prevent.
|
|
"""
|
|
return source not in RANKED_SOURCES
|
|
|
|
|
|
async def _report_failure(site: str) -> None:
|
|
await report_telemetry_failure("rule_usage", site)
|
|
|
|
|
|
async def _insert_events(rows: list[dict]) -> None:
|
|
"""Persist usage rows. Best-effort: failures degrade, visibly."""
|
|
try:
|
|
async with async_session() as session:
|
|
session.add_all([RuleUsageEvent(**row) for row in rows])
|
|
await session.commit()
|
|
except Exception:
|
|
await _report_failure("write")
|
|
|
|
|
|
def _schedule(rows: list[dict]) -> None:
|
|
if not rows:
|
|
return
|
|
spawn(_insert_events(rows), site="rule_usage_write")
|
|
|
|
|
|
def record_rule_surfaced(
|
|
*, user_id: int | None, rule_ids: list[int] | set[int], source: str
|
|
) -> None:
|
|
"""Fire-and-forget: record that these rules were shown to the agent.
|
|
|
|
Takes the whole delivery at once — one insert per surfacing event, not per
|
|
rule — because a hint is a single decision and its rows should land
|
|
together.
|
|
|
|
Record what was actually SHOWN, never what was considered. For the ranked
|
|
arm that means the post-filter hits: it drops what the session already
|
|
holds (`exclude_rule_ids`) before it speaks, and a rule considered and not
|
|
shown was not surfaced. Counting those would inflate the denominator with
|
|
claims the agent never saw, which reads as a precision problem the arm does
|
|
not have.
|
|
|
|
Bulk surfaces pass their whole delivered set, which is the same rule read
|
|
from the other end — everything in a preload IS shown. `source` is what
|
|
separates the two afterwards (see `RANKED_SOURCES`); this function does not
|
|
care which kind it is recording.
|
|
"""
|
|
try:
|
|
rows = [
|
|
{
|
|
"user_id": user_id,
|
|
"rule_id": int(rid),
|
|
"event": SURFACED,
|
|
"source": source,
|
|
}
|
|
for rid in rule_ids
|
|
]
|
|
except Exception:
|
|
logger.debug("rule usage payload build failed", exc_info=True)
|
|
return
|
|
_schedule(rows)
|
|
|
|
|
|
def record_rule_pulled(*, user_id: int | None, rule_id: int, source: str) -> None:
|
|
"""Fire-and-forget: record that a rule was opened in full.
|
|
|
|
A PULL is somebody choosing to open one record. `list_always_on_rules` and
|
|
`enter_project` are NOT pulls — they are bulk resident loads that hand over
|
|
every applicable rule at once, and counting them would swamp the signal
|
|
with the very ambient delivery the ratio exists to distinguish from.
|
|
"""
|
|
try:
|
|
rows = [
|
|
{
|
|
"user_id": user_id,
|
|
"rule_id": int(rule_id),
|
|
"event": PULLED,
|
|
"source": source,
|
|
}
|
|
]
|
|
except Exception:
|
|
logger.debug("rule usage payload build failed", exc_info=True)
|
|
return
|
|
_schedule(rows)
|
|
|
|
|
|
def empty_rule_usage() -> dict:
|
|
"""The zero readout — what a rule with no recorded events looks like.
|
|
|
|
Callers render this shape unconditionally, so a rule predating the table
|
|
reads as "never surfaced, never pulled" rather than as a missing key. That
|
|
distinction matters more here than for notes: every rule in an install
|
|
predates this table, so for a while "no events" is the normal state and it
|
|
must not look like a broken readout.
|
|
|
|
`surfaced_count` is RANKED surfacings only; `ambient_count` is the bulk
|
|
deliveries (see `RANKED_SOURCES`). The split is what keeps the badge's
|
|
"shown often, opened never → dead weight" reading honest: every rule in an
|
|
always-on set is delivered every session, so an unsplit counter would rank
|
|
the resident set as the most-surfaced rules in the install purely for being
|
|
resident.
|
|
"""
|
|
return {
|
|
"surfaced_count": 0,
|
|
"ambient_count": 0,
|
|
"pull_count": 0,
|
|
"last_surfaced_at": None,
|
|
"last_pulled_at": None,
|
|
}
|
|
|
|
|
|
async def usage_for_rules(rule_ids: list[int]) -> dict[int, dict]:
|
|
"""Aggregate usage for a set of rules: {rule_id: {counts + timestamps}}.
|
|
|
|
One GROUP BY for the whole page rather than a query per row — this feeds a
|
|
list view, so the per-row shape would be N+1 by construction. Rules with no
|
|
events come back with `empty_rule_usage()`, so the caller never has to tell
|
|
"no events" from "not in the result".
|
|
"""
|
|
ids = [int(r) for r in rule_ids]
|
|
out: dict[int, dict] = {rid: empty_rule_usage() for rid in ids}
|
|
if not ids:
|
|
return out
|
|
|
|
# Classified in SQL so the group stays small: per rule we get at most
|
|
# (surfaced-ranked, surfaced-ambient, pulled) rather than a row per distinct
|
|
# source. ONE labelled expression, bound to a variable and reused in the
|
|
# GROUP BY — a second `case()` instance there renders its own expanding-IN
|
|
# bind names under asyncpg, so the database sees two DIFFERENT expressions
|
|
# and rejects the query with a GroupingError. The note twin carries the
|
|
# same warning for the same reason, and #2663 is what it cost: the
|
|
# rejection was swallowed and every counter read zero in production while
|
|
# the writes were landing fine.
|
|
ambient = case(
|
|
(RuleUsageEvent.source.notin_(RANKED_SOURCES), True),
|
|
else_=False,
|
|
).label("ambient")
|
|
try:
|
|
async with async_session() as session:
|
|
rows = (
|
|
await session.execute(
|
|
select(
|
|
RuleUsageEvent.rule_id,
|
|
RuleUsageEvent.event,
|
|
func.count().label("n"),
|
|
func.max(RuleUsageEvent.created_at).label("last_at"),
|
|
ambient,
|
|
)
|
|
.where(RuleUsageEvent.rule_id.in_(ids))
|
|
.group_by(
|
|
RuleUsageEvent.rule_id,
|
|
RuleUsageEvent.event,
|
|
ambient,
|
|
)
|
|
)
|
|
).all()
|
|
except Exception:
|
|
# A telemetry readout must not be able to break the list it decorates —
|
|
# but it must say it failed, or a broken readout is indistinguishable
|
|
# from a corpus nobody uses (#2663).
|
|
await _report_failure("readout")
|
|
return out
|
|
|
|
for rule_id, event, n, last_at, is_amb in rows:
|
|
slot = out.get(int(rule_id))
|
|
if slot is None:
|
|
continue
|
|
if event == SURFACED and is_amb:
|
|
slot["ambient_count"] = int(n)
|
|
elif event == SURFACED:
|
|
slot["surfaced_count"] = int(n)
|
|
slot["last_surfaced_at"] = iso(last_at)
|
|
elif event == PULLED:
|
|
# Pulls are pulls regardless of what surfaced the rule — "did
|
|
# anyone ever open this?" does not depend on how it was found. Both
|
|
# halves accumulate, so this ADDS rather than assigns: a rule can
|
|
# now be pulled after a ranked hint and after a preload, and the
|
|
# split arrives as two rows.
|
|
slot["pull_count"] = slot["pull_count"] + int(n)
|
|
latest = iso(last_at)
|
|
if latest and (slot["last_pulled_at"] or "") < latest:
|
|
slot["last_pulled_at"] = latest
|
|
return out
|