feat(telemetry): retrieval_telemetry says what is wrong (#3431)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 8s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m30s
CI & Build / Build & push image (push) Successful in 23s

The tool returned distributions and left the reading to the caller, so every
readout was the same four checks done by hand — #3430's baseline, #3835's rule
near-misses, the #1038 rerank gate. Mechanical, and therefore forgettable.

Tonight's acceptance pass on #3898 was the case for doing this. Reading it by
hand meant catching that two surfaces had `covers_window: false`, that
`prompt_rule`'s floor had moved three times inside the window (which made the
readout self-contradictory: deliveries at 0.622 beside refusals at 0.7199),
and that 15 of 20 near-misses were one record against text no operator wrote.
Miss any of those and the obvious conclusion was "the bar is too tight" — a
floor change that would have injected one preference into every notification.

`warnings` is always present and empty when clean, so its emptiness is an
answer rather than a gap. Each entry carries the numbers that produced it:
"345 calls, 0 declined" is the analysis, "check write_path_rule" is an
instruction to redo it. Five codes — cannot_decline, band_hugs_floor,
no_duration, surfaced_never_pulled, unregistered_source.

cannot_decline has three guards, each a bug it would otherwise cause. Asked
surfaces are exempt (a search returning a list every time is working). An arm
not known to log unconditionally is exempt — that is #3497 exactly, where both
rule arms recorded only their hits, so a decline count of zero was a LOGGING
defect and this warning would have sent the reader to a threshold that was
never involved. Unregistered sources get numbers but no verdict.

`silent_surfaces` is the half the rows cannot show: an arm that emitted
nothing is invisible to every row-based check and looks exactly like an arm
that does not exist. It is driven by a new declared registry,
`retrieval_registry.POINTS` — deliberately NOT `retrieval_surfaces.SURFACES`,
which answers "what can be tuned" and excludes the reserved slots because a
budget of 1 is their feature. This answers "what can be measured", and the
reserved slots belong in it precisely because they are judgeable without being
tunable. A test asserts the two cannot drift apart.

The registry test derives sources from the call sites with `ast`, not grep,
and the difference is not theoretical: `wide_net` and `report_preference`
reach their recorder as `source=SOURCE` through a module constant, so a grep
for `source="` is blind to both — the narrowing #3191 warns about. Three sites
pass `source` as a variable and are declared in FAN_OUT_SITES; the test pins
those sites but not the values they can pass, which is why the
`unregistered_source` warning exists to catch the rest at first fire.

Thresholds are settings (rule 25) defaulted so a fresh install with almost no
data produces no warnings at all (rule 115) — a new user's first readout
naming five broken things would be describing the emptiness.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
2026-09-20 20:24:56 -04:00
co-authored by Claude Opus 5
parent 1f7ff7b215
commit e2c3a5c2b5
5 changed files with 1003 additions and 0 deletions
+223
View File
@@ -0,0 +1,223 @@
"""Every point that retrieves or surfaces a record, declared (#3431, #3432).
WHY A DECLARED LIST WHEN THE ROWS ALREADY NAME THEIR SOURCE. Because the rows
can only describe arms that FIRED. A surface that emits nothing writes no row,
so it is indistinguishable from a surface that does not exist — and the two
call for opposite responses. `retrieval_telemetry` could report perfect health
for a window in which half the arms never ran. #3430 hit exactly this: its
"silent surface" finding could only be made by a human who happened to know
which arms were supposed to be there.
This module is that knowledge, written down. It carries the two facts the rows
structurally cannot:
* **asked or unbidden.** `mcp_search` returning nothing is a search with no
matches, which is its job. `auto_inject` returning nothing on every call
for a week is an arm that has lost the ability to speak. Same row shape,
opposite meanings, and only the caller's intent separates them.
* **expected to emit, or deliberately quiet with a reason.** A justified
silence must not read as a gap (#2475). `rest_*` sources fire only when a
person opens the web UI; on an install driven entirely through MCP their
absence is correct and permanent, and reporting it every window would
train the reader to ignore the list that also carries the real ones.
NOT `retrieval_surfaces.SURFACES`, and the distinction is deliberate rather
than duplication. That registry answers "what can be TUNED" — it carries a
floor key, a budget key and the measurement stamp behind their defaults, and
it deliberately EXCLUDES the reserved slots (`preference_slot`, `reuse_slot`,
`lesson_slot`) because a budget of 1 is their feature and exposing it would
invite setting it to 0. This registry answers "what can be MEASURED", and the
reserved slots belong in it precisely because they are judgeable without being
tunable. One is a control panel; the other is an inventory. A test asserts
every tunable surface also appears here, so the two cannot drift apart.
ADDING AN ARM MEANS ADDING A ROW HERE. `tests/test_retrieval_registry.py`
walks the call sites with the ast module and fails on a source it cannot find
below — deliberately not a grep, because two of the sources in this file
(`wide_net`, `report_preference`) reach their recorder as `source=SOURCE`
through a module constant and a grep for `source="` misses both. That is the
narrowing #3191 warns about, caught here in the act.
"""
from __future__ import annotations
from dataclasses import dataclass
# ── How a point is reached ────────────────────────────────────────────────
#
# UNBIDDEN: fires on its own, against a query the agent did not write — a
# prompt, a file being written, a command about to run. It interrupts, so it
# must be able to stay quiet, and a stream of empty calls is a real signal.
#
# ASKED: someone called a search and wants a ranked list. Returning nothing is
# an answer, not a failure, so "cannot decline" must never fire on these.
#
# AMBIENT: records travel with a reply that was going to be sent anyway (a
# project handshake, a planning read). Nothing was ranked and nothing was
# chosen, so score-shaped warnings do not apply.
#
# PULL: a reader opened a record. The terminal event of the whole system, and
# the numerator of pull-through.
UNBIDDEN = "unbidden"
ASKED = "asked"
AMBIENT = "ambient"
PULL = "pull"
@dataclass(frozen=True)
class Point:
"""One place a record reaches a reader."""
source: str
"""The telemetry `source` value. MUST equal the string the call site passes."""
kind: str
"""UNBIDDEN / ASKED / AMBIENT / PULL — see above."""
what: str
"""One line, for an agent reading a warning that names this source."""
expects_traffic: bool = True
"""Whether silence over an active window is worth reporting.
False means "this can be quiet forever and that is correct" — and then
`quiet_because` must say why, so a reader meets a decision rather than a
hole (#2475).
"""
quiet_because: str = ""
logs_unconditionally: bool = True
"""Whether this arm writes a row even when it returns NOTHING.
THE #3497 GUARD, and the reason this field is not merely documentation.
Both rule arms once logged only their hits, so their zero-result count was
structurally 0 — and a "cannot decline" warning computed over that would
have fired on a LOGGING BUG while reporting a ranking problem, sending the
reader to move a floor that was never involved. An arm whose logging has
not been confirmed unconditional is not eligible for that warning; it is
excluded from the check rather than trusted.
"""
def _p(source, kind, what, **kw) -> tuple[str, Point]:
return source, Point(source=source, kind=kind, what=what, **kw)
POINTS: dict[str, Point] = dict([
# ── Unbidden push arms ───────────────────────────────────────────────
_p("auto_inject", UNBIDDEN, "the notes menu offered at the prompt boundary"),
_p("write_path", UNBIDDEN, "prior art offered when a file is about to be written"),
_p("write_path_rule", UNBIDDEN, "rules that may govern the file being written"),
_p("pre_tool_rule", UNBIDDEN, "rules that may govern a command about to run"),
_p("prompt_rule", UNBIDDEN, "rules that may govern what the operator just asked"),
_p("preference_slot", UNBIDDEN,
"the one line reserved for a preference at the prompt boundary"),
_p("reuse_slot", UNBIDDEN, "the one line reserved for a reusable snippet"),
_p("lesson_slot", UNBIDDEN, "the one line reserved for a lesson"),
_p("report_preference", UNBIDDEN,
"the fixed question asked when a task finishes: how should this report read"),
# ── Asked: a caller wanted a ranked list ─────────────────────────────
_p("mcp_search", ASKED, "an agent called search"),
_p("wide_net", ASKED, "an agent asked what might apply, with no bar"),
_p("rest_search", ASKED, "the web UI searched", expects_traffic=False,
quiet_because="only fires when a person uses the web UI; an install "
"driven entirely through MCP is correctly silent here"),
_p("browse_search", ASKED, "the browse listing searched", expects_traffic=False,
quiet_because="same as rest_search — a human-only entry point"),
# ── Ambient: carried along with a reply already being sent ───────────
_p("enter_project", AMBIENT, "the project handshake"),
_p("process_skill_sync", AMBIENT, "stored Processes synced into skills"),
_p("start_planning", AMBIENT, "the rules listed when a plan is opened"),
_p("get_task", AMBIENT, "the rules listed beside a task"),
_p("get_project", AMBIENT, "the rules listed beside a project"),
_p("get_milestone", AMBIENT, "the rules listed beside a milestone"),
# ── The write-path menu's own arms ───────────────────────────────────
#
# These are NOT in retrieval_logs. The place arm carries no score and so
# has no row there at all (#2085) — before it was split out, the arm
# firing on the strongest possible claim, "there is already a canonical
# helper in this exact file", was the one arm nobody could measure.
_p("write_path_place", AMBIENT, "a snippet already placed in this very file"),
_p("write_path_semantic", AMBIENT, "prior art matched by meaning"),
_p("write_path_sync", AMBIENT,
"records that should be updated alongside this edit"),
# ── Pull: a reader opened the record ─────────────────────────────────
_p("mcp_get_note", PULL, "an agent opened a note"),
_p("mcp_get_task", PULL, "an agent opened a task"),
_p("mcp_get_snippet", PULL, "an agent opened a snippet"),
_p("mcp_get_process", PULL, "an agent opened a process"),
_p("mcp_get_lesson", PULL, "an agent opened a lesson"),
_p("mcp_get_rule", PULL, "an agent opened a rule"),
_p("rest_note", PULL, "a person opened a note", expects_traffic=False,
quiet_because="web UI only"),
_p("rest_task", PULL, "a person opened a task", expects_traffic=False,
quiet_because="web UI only"),
_p("rest_snippet", PULL, "a person opened a snippet", expects_traffic=False,
quiet_because="web UI only"),
_p("rest_lesson", PULL, "a person opened a lesson", expects_traffic=False,
quiet_because="web UI only"),
_p("rest_rule", PULL, "a person opened a rule", expects_traffic=False,
quiet_because="web UI only"),
])
# Call sites that pass `source` as a VARIABLE, and what they can pass.
#
# WHY THESE ARE DECLARED RATHER THAN RESOLVED. The registry test reads the
# call sites with `ast`, which settles a literal and a module constant but not
# a value that arrives through a parameter or a dict key. Rather than let
# those sites go unchecked — the silent half of a guard that appears to cover
# everything — they are named here, and the test FAILS if it meets an
# unresolved site that is not in this list.
#
# THE HONEST LIMIT, stated so nobody over-trusts this: the test pins the SITE,
# not the values. Adding a fourth arm inside `plugin_context`'s `by_arm` fan
# out would not fail here. What catches that is `assert_registered`, which the
# recorders call as the row is written.
# Keyed by the EXACT string the extractor produces, matched by equality rather
# than by prefix. A prefix match was the first spelling and it silently matched
# nothing, because the paths are relative to `src/` and so begin `scribe/` —
# every site sailed through a check that looked like it was running.
FAN_OUT_SITES: dict[str, tuple[str, ...]] = {
"scribe/services/plugin_context.py::record_surfaced(source=arm)": (
"write_path_place", "write_path_semantic", "write_path_sync",
),
"scribe/services/rulebooks.py::record_rule_surfaced(source=source)": (
"enter_project", "start_planning", "get_task", "get_project",
"get_milestone",
),
}
def get_point(source: str) -> Point | None:
"""The registered point for a source, or None if it is not declared."""
return POINTS.get(source)
def sources_expected_to_emit() -> list[str]:
"""Every point whose silence over an ACTIVE window is worth reporting."""
return [s for s, p in POINTS.items() if p.expects_traffic]
def is_registered(source: str) -> bool:
"""Is this source declared?
Read by the TELEMETRY READOUT, not by the recorders, and that placement is
the design rather than an accident of where it was easy. Validating at
write time would put a dict lookup on every retrieval to catch a mistake
that is only ever made once, when an arm is added — and it could not
raise anyway, because telemetry that breaks its caller is the worse bug
(`record_retrieval` is fire-and-forget for exactly that reason). So the
check sits where the cost is already paid and the reader is already
looking: a source that appears in the rows and not in this file comes back
as an `unregistered_source` warning.
That is what covers FAN_OUT_SITES above. The static test pins those sites
but not the values they can pass, so a fourth arm added inside one would
slip past it — and then show up here the first time it fires.
"""
return source in POINTS