feat(telemetry): retrieval_telemetry says what is wrong (#3431)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 8s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m30s
CI & Build / Build & push image (push) Successful in 23s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 8s
CI & Build / integration (push) Successful in 48s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m30s
CI & Build / Build & push image (push) Successful in 23s
The tool returned distributions and left the reading to the caller, so every readout was the same four checks done by hand — #3430's baseline, #3835's rule near-misses, the #1038 rerank gate. Mechanical, and therefore forgettable. Tonight's acceptance pass on #3898 was the case for doing this. Reading it by hand meant catching that two surfaces had `covers_window: false`, that `prompt_rule`'s floor had moved three times inside the window (which made the readout self-contradictory: deliveries at 0.622 beside refusals at 0.7199), and that 15 of 20 near-misses were one record against text no operator wrote. Miss any of those and the obvious conclusion was "the bar is too tight" — a floor change that would have injected one preference into every notification. `warnings` is always present and empty when clean, so its emptiness is an answer rather than a gap. Each entry carries the numbers that produced it: "345 calls, 0 declined" is the analysis, "check write_path_rule" is an instruction to redo it. Five codes — cannot_decline, band_hugs_floor, no_duration, surfaced_never_pulled, unregistered_source. cannot_decline has three guards, each a bug it would otherwise cause. Asked surfaces are exempt (a search returning a list every time is working). An arm not known to log unconditionally is exempt — that is #3497 exactly, where both rule arms recorded only their hits, so a decline count of zero was a LOGGING defect and this warning would have sent the reader to a threshold that was never involved. Unregistered sources get numbers but no verdict. `silent_surfaces` is the half the rows cannot show: an arm that emitted nothing is invisible to every row-based check and looks exactly like an arm that does not exist. It is driven by a new declared registry, `retrieval_registry.POINTS` — deliberately NOT `retrieval_surfaces.SURFACES`, which answers "what can be tuned" and excludes the reserved slots because a budget of 1 is their feature. This answers "what can be measured", and the reserved slots belong in it precisely because they are judgeable without being tunable. A test asserts the two cannot drift apart. The registry test derives sources from the call sites with `ast`, not grep, and the difference is not theoretical: `wide_net` and `report_preference` reach their recorder as `source=SOURCE` through a module constant, so a grep for `source="` is blind to both — the narrowing #3191 warns about. Three sites pass `source` as a variable and are declared in FAN_OUT_SITES; the test pins those sites but not the values they can pass, which is why the `unregistered_source` warning exists to catch the rest at first fire. Thresholds are settings (rule 25) defaulted so a fresh install with almost no data produces no warnings at all (rule 115) — a new user's first readout naming five broken things would be describing the emptiness. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
@@ -0,0 +1,223 @@
|
||||
"""Every point that retrieves or surfaces a record, declared (#3431, #3432).
|
||||
|
||||
WHY A DECLARED LIST WHEN THE ROWS ALREADY NAME THEIR SOURCE. Because the rows
|
||||
can only describe arms that FIRED. A surface that emits nothing writes no row,
|
||||
so it is indistinguishable from a surface that does not exist — and the two
|
||||
call for opposite responses. `retrieval_telemetry` could report perfect health
|
||||
for a window in which half the arms never ran. #3430 hit exactly this: its
|
||||
"silent surface" finding could only be made by a human who happened to know
|
||||
which arms were supposed to be there.
|
||||
|
||||
This module is that knowledge, written down. It carries the two facts the rows
|
||||
structurally cannot:
|
||||
|
||||
* **asked or unbidden.** `mcp_search` returning nothing is a search with no
|
||||
matches, which is its job. `auto_inject` returning nothing on every call
|
||||
for a week is an arm that has lost the ability to speak. Same row shape,
|
||||
opposite meanings, and only the caller's intent separates them.
|
||||
* **expected to emit, or deliberately quiet with a reason.** A justified
|
||||
silence must not read as a gap (#2475). `rest_*` sources fire only when a
|
||||
person opens the web UI; on an install driven entirely through MCP their
|
||||
absence is correct and permanent, and reporting it every window would
|
||||
train the reader to ignore the list that also carries the real ones.
|
||||
|
||||
NOT `retrieval_surfaces.SURFACES`, and the distinction is deliberate rather
|
||||
than duplication. That registry answers "what can be TUNED" — it carries a
|
||||
floor key, a budget key and the measurement stamp behind their defaults, and
|
||||
it deliberately EXCLUDES the reserved slots (`preference_slot`, `reuse_slot`,
|
||||
`lesson_slot`) because a budget of 1 is their feature and exposing it would
|
||||
invite setting it to 0. This registry answers "what can be MEASURED", and the
|
||||
reserved slots belong in it precisely because they are judgeable without being
|
||||
tunable. One is a control panel; the other is an inventory. A test asserts
|
||||
every tunable surface also appears here, so the two cannot drift apart.
|
||||
|
||||
ADDING AN ARM MEANS ADDING A ROW HERE. `tests/test_retrieval_registry.py`
|
||||
walks the call sites with the ast module and fails on a source it cannot find
|
||||
below — deliberately not a grep, because two of the sources in this file
|
||||
(`wide_net`, `report_preference`) reach their recorder as `source=SOURCE`
|
||||
through a module constant and a grep for `source="` misses both. That is the
|
||||
narrowing #3191 warns about, caught here in the act.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass
|
||||
|
||||
# ── How a point is reached ────────────────────────────────────────────────
|
||||
#
|
||||
# UNBIDDEN: fires on its own, against a query the agent did not write — a
|
||||
# prompt, a file being written, a command about to run. It interrupts, so it
|
||||
# must be able to stay quiet, and a stream of empty calls is a real signal.
|
||||
#
|
||||
# ASKED: someone called a search and wants a ranked list. Returning nothing is
|
||||
# an answer, not a failure, so "cannot decline" must never fire on these.
|
||||
#
|
||||
# AMBIENT: records travel with a reply that was going to be sent anyway (a
|
||||
# project handshake, a planning read). Nothing was ranked and nothing was
|
||||
# chosen, so score-shaped warnings do not apply.
|
||||
#
|
||||
# PULL: a reader opened a record. The terminal event of the whole system, and
|
||||
# the numerator of pull-through.
|
||||
UNBIDDEN = "unbidden"
|
||||
ASKED = "asked"
|
||||
AMBIENT = "ambient"
|
||||
PULL = "pull"
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Point:
|
||||
"""One place a record reaches a reader."""
|
||||
|
||||
source: str
|
||||
"""The telemetry `source` value. MUST equal the string the call site passes."""
|
||||
|
||||
kind: str
|
||||
"""UNBIDDEN / ASKED / AMBIENT / PULL — see above."""
|
||||
|
||||
what: str
|
||||
"""One line, for an agent reading a warning that names this source."""
|
||||
|
||||
expects_traffic: bool = True
|
||||
"""Whether silence over an active window is worth reporting.
|
||||
|
||||
False means "this can be quiet forever and that is correct" — and then
|
||||
`quiet_because` must say why, so a reader meets a decision rather than a
|
||||
hole (#2475).
|
||||
"""
|
||||
|
||||
quiet_because: str = ""
|
||||
|
||||
logs_unconditionally: bool = True
|
||||
"""Whether this arm writes a row even when it returns NOTHING.
|
||||
|
||||
THE #3497 GUARD, and the reason this field is not merely documentation.
|
||||
Both rule arms once logged only their hits, so their zero-result count was
|
||||
structurally 0 — and a "cannot decline" warning computed over that would
|
||||
have fired on a LOGGING BUG while reporting a ranking problem, sending the
|
||||
reader to move a floor that was never involved. An arm whose logging has
|
||||
not been confirmed unconditional is not eligible for that warning; it is
|
||||
excluded from the check rather than trusted.
|
||||
"""
|
||||
|
||||
|
||||
def _p(source, kind, what, **kw) -> tuple[str, Point]:
|
||||
return source, Point(source=source, kind=kind, what=what, **kw)
|
||||
|
||||
|
||||
POINTS: dict[str, Point] = dict([
|
||||
# ── Unbidden push arms ───────────────────────────────────────────────
|
||||
_p("auto_inject", UNBIDDEN, "the notes menu offered at the prompt boundary"),
|
||||
_p("write_path", UNBIDDEN, "prior art offered when a file is about to be written"),
|
||||
_p("write_path_rule", UNBIDDEN, "rules that may govern the file being written"),
|
||||
_p("pre_tool_rule", UNBIDDEN, "rules that may govern a command about to run"),
|
||||
_p("prompt_rule", UNBIDDEN, "rules that may govern what the operator just asked"),
|
||||
_p("preference_slot", UNBIDDEN,
|
||||
"the one line reserved for a preference at the prompt boundary"),
|
||||
_p("reuse_slot", UNBIDDEN, "the one line reserved for a reusable snippet"),
|
||||
_p("lesson_slot", UNBIDDEN, "the one line reserved for a lesson"),
|
||||
_p("report_preference", UNBIDDEN,
|
||||
"the fixed question asked when a task finishes: how should this report read"),
|
||||
|
||||
# ── Asked: a caller wanted a ranked list ─────────────────────────────
|
||||
_p("mcp_search", ASKED, "an agent called search"),
|
||||
_p("wide_net", ASKED, "an agent asked what might apply, with no bar"),
|
||||
_p("rest_search", ASKED, "the web UI searched", expects_traffic=False,
|
||||
quiet_because="only fires when a person uses the web UI; an install "
|
||||
"driven entirely through MCP is correctly silent here"),
|
||||
_p("browse_search", ASKED, "the browse listing searched", expects_traffic=False,
|
||||
quiet_because="same as rest_search — a human-only entry point"),
|
||||
|
||||
# ── Ambient: carried along with a reply already being sent ───────────
|
||||
_p("enter_project", AMBIENT, "the project handshake"),
|
||||
_p("process_skill_sync", AMBIENT, "stored Processes synced into skills"),
|
||||
_p("start_planning", AMBIENT, "the rules listed when a plan is opened"),
|
||||
_p("get_task", AMBIENT, "the rules listed beside a task"),
|
||||
_p("get_project", AMBIENT, "the rules listed beside a project"),
|
||||
_p("get_milestone", AMBIENT, "the rules listed beside a milestone"),
|
||||
|
||||
# ── The write-path menu's own arms ───────────────────────────────────
|
||||
#
|
||||
# These are NOT in retrieval_logs. The place arm carries no score and so
|
||||
# has no row there at all (#2085) — before it was split out, the arm
|
||||
# firing on the strongest possible claim, "there is already a canonical
|
||||
# helper in this exact file", was the one arm nobody could measure.
|
||||
_p("write_path_place", AMBIENT, "a snippet already placed in this very file"),
|
||||
_p("write_path_semantic", AMBIENT, "prior art matched by meaning"),
|
||||
_p("write_path_sync", AMBIENT,
|
||||
"records that should be updated alongside this edit"),
|
||||
|
||||
# ── Pull: a reader opened the record ─────────────────────────────────
|
||||
_p("mcp_get_note", PULL, "an agent opened a note"),
|
||||
_p("mcp_get_task", PULL, "an agent opened a task"),
|
||||
_p("mcp_get_snippet", PULL, "an agent opened a snippet"),
|
||||
_p("mcp_get_process", PULL, "an agent opened a process"),
|
||||
_p("mcp_get_lesson", PULL, "an agent opened a lesson"),
|
||||
_p("mcp_get_rule", PULL, "an agent opened a rule"),
|
||||
_p("rest_note", PULL, "a person opened a note", expects_traffic=False,
|
||||
quiet_because="web UI only"),
|
||||
_p("rest_task", PULL, "a person opened a task", expects_traffic=False,
|
||||
quiet_because="web UI only"),
|
||||
_p("rest_snippet", PULL, "a person opened a snippet", expects_traffic=False,
|
||||
quiet_because="web UI only"),
|
||||
_p("rest_lesson", PULL, "a person opened a lesson", expects_traffic=False,
|
||||
quiet_because="web UI only"),
|
||||
_p("rest_rule", PULL, "a person opened a rule", expects_traffic=False,
|
||||
quiet_because="web UI only"),
|
||||
])
|
||||
|
||||
|
||||
# Call sites that pass `source` as a VARIABLE, and what they can pass.
|
||||
#
|
||||
# WHY THESE ARE DECLARED RATHER THAN RESOLVED. The registry test reads the
|
||||
# call sites with `ast`, which settles a literal and a module constant but not
|
||||
# a value that arrives through a parameter or a dict key. Rather than let
|
||||
# those sites go unchecked — the silent half of a guard that appears to cover
|
||||
# everything — they are named here, and the test FAILS if it meets an
|
||||
# unresolved site that is not in this list.
|
||||
#
|
||||
# THE HONEST LIMIT, stated so nobody over-trusts this: the test pins the SITE,
|
||||
# not the values. Adding a fourth arm inside `plugin_context`'s `by_arm` fan
|
||||
# out would not fail here. What catches that is `assert_registered`, which the
|
||||
# recorders call as the row is written.
|
||||
# Keyed by the EXACT string the extractor produces, matched by equality rather
|
||||
# than by prefix. A prefix match was the first spelling and it silently matched
|
||||
# nothing, because the paths are relative to `src/` and so begin `scribe/` —
|
||||
# every site sailed through a check that looked like it was running.
|
||||
FAN_OUT_SITES: dict[str, tuple[str, ...]] = {
|
||||
"scribe/services/plugin_context.py::record_surfaced(source=arm)": (
|
||||
"write_path_place", "write_path_semantic", "write_path_sync",
|
||||
),
|
||||
"scribe/services/rulebooks.py::record_rule_surfaced(source=source)": (
|
||||
"enter_project", "start_planning", "get_task", "get_project",
|
||||
"get_milestone",
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
def get_point(source: str) -> Point | None:
|
||||
"""The registered point for a source, or None if it is not declared."""
|
||||
return POINTS.get(source)
|
||||
|
||||
|
||||
def sources_expected_to_emit() -> list[str]:
|
||||
"""Every point whose silence over an ACTIVE window is worth reporting."""
|
||||
return [s for s, p in POINTS.items() if p.expects_traffic]
|
||||
|
||||
|
||||
def is_registered(source: str) -> bool:
|
||||
"""Is this source declared?
|
||||
|
||||
Read by the TELEMETRY READOUT, not by the recorders, and that placement is
|
||||
the design rather than an accident of where it was easy. Validating at
|
||||
write time would put a dict lookup on every retrieval to catch a mistake
|
||||
that is only ever made once, when an arm is added — and it could not
|
||||
raise anyway, because telemetry that breaks its caller is the worse bug
|
||||
(`record_retrieval` is fire-and-forget for exactly that reason). So the
|
||||
check sits where the cost is already paid and the reader is already
|
||||
looking: a source that appears in the rows and not in this file comes back
|
||||
as an `unregistered_source` warning.
|
||||
|
||||
That is what covers FAN_OUT_SITES above. The static test pins those sites
|
||||
but not the values they can pass, so a fourth arm added inside one would
|
||||
slip past it — and then show up here the first time it fires.
|
||||
"""
|
||||
return source in POINTS
|
||||
Reference in New Issue
Block a user