Files
FabledScribe/src/scribe/services/retrieval_registry.py
T
bvandeusenandClaude Opus 5.5 0a1bb68808
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 13s
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / Python tests (push) Successful in 1m44s
CI & Build / Build & push image (push) Successful in 36s
feat(rulings): system_usage_events is read back — per-System counts and a telemetry block (#4769)
#4769 "Rulings are counted where someone will read them": milestone 444
step 4 wrote system_usage_events and nothing read it.

- retrieval_telemetry gains a `system_usage` block: surfacings and opens by
  source, distinct counts, and `by_system` naming the areas most shown.
  There is deliberately no pull-through ratio, because rulings travel in full
  in the line and opens are the exception.
- usage_for_systems (one GROUP BY) adds `usage` to the REST Systems list and
  detail, and to MCP get_system. MCP list_systems is unchanged.
- The Systems UI shows a "rulings shown N×" chip.
- rulings_pre_tool, rulings_write_path and mcp_get_system are now declared
  registry points; the registry guard covers their recorders.
- The Systems store merges a PATCH reply instead of replacing the row.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 08:49:20 -04:00

282 lines
15 KiB
Python

"""Every point that retrieves or surfaces a record, declared (#3431, #3432).
WHY A DECLARED LIST WHEN THE ROWS ALREADY NAME THEIR SOURCE. Because the rows
can only describe arms that FIRED. A surface that emits nothing writes no row,
so it is indistinguishable from a surface that does not exist — and the two
call for opposite responses. `retrieval_telemetry` could report perfect health
for a window in which half the arms never ran. #3430 hit exactly this: its
"silent surface" finding could only be made by a human who happened to know
which arms were supposed to be there.
This module is that knowledge, written down. It carries the two facts the rows
structurally cannot:
* **asked or unbidden.** `mcp_search` returning nothing is a search with no
matches, which is its job. `auto_inject` returning nothing on every call
for a week is an arm that has lost the ability to speak. Same row shape,
opposite meanings, and only the caller's intent separates them.
* **expected to emit, or deliberately quiet with a reason.** A justified
silence must not read as a gap (#2475). `rest_*` sources fire only when a
person opens the web UI; on an install driven entirely through MCP their
absence is correct and permanent, and reporting it every window would
train the reader to ignore the list that also carries the real ones.
NOT `retrieval_surfaces.SURFACES`, and the distinction is deliberate rather
than duplication. That registry answers "what can be TUNED" — it carries a
floor key, a budget key and the measurement stamp behind their defaults, and
it deliberately EXCLUDES the reserved slots (`preference_slot`, `reuse_slot`,
`lesson_slot`) because a budget of 1 is their feature and exposing it would
invite setting it to 0. This registry answers "what can be MEASURED", and the
reserved slots belong in it precisely because they are judgeable without being
tunable. One is a control panel; the other is an inventory. A test asserts
every tunable surface also appears here, so the two cannot drift apart.
ADDING AN ARM MEANS ADDING A ROW HERE. `tests/test_retrieval_registry.py`
walks the call sites with the ast module and fails on a source it cannot find
below — deliberately not a grep, because two of the sources in this file
(`wide_net`, `report_preference`) reach their recorder as `source=SOURCE`
through a module constant and a grep for `source="` misses both. That is the
narrowing #3191 warns about, caught here in the act.
"""
from __future__ import annotations
from dataclasses import dataclass
# ── How a point is reached ────────────────────────────────────────────────
#
# UNBIDDEN: fires on its own, against a query the agent did not write — a
# prompt, a file being written, a command about to run. It interrupts, so it
# must be able to stay quiet, and a stream of empty calls is a real signal.
#
# ASKED: someone called a search and wants a ranked list. Returning nothing is
# an answer, not a failure, so "cannot decline" must never fire on these.
#
# AMBIENT: records travel with a reply that was going to be sent anyway (a
# project handshake, a planning read). Nothing was ranked and nothing was
# chosen, so score-shaped warnings do not apply.
#
# PULL: a reader opened a record. The terminal event of the whole system, and
# the numerator of pull-through.
UNBIDDEN = "unbidden"
ASKED = "asked"
AMBIENT = "ambient"
PULL = "pull"
@dataclass(frozen=True)
class Point:
"""One place a record reaches a reader."""
source: str
"""The telemetry `source` value. MUST equal the string the call site passes."""
kind: str
"""UNBIDDEN / ASKED / AMBIENT / PULL — see above."""
what: str
"""One line, for an agent reading a warning that names this source."""
expects_traffic: bool = True
"""Whether silence over an active window is worth reporting.
False means "this can be quiet forever and that is correct" — and then
`quiet_because` must say why, so a reader meets a decision rather than a
hole (#2475).
"""
quiet_because: str = ""
fixed_query: bool = False
"""Whether this arm always searches the SAME query string.
THE #3497 GUARD, ONE STEP OVER. `logs_unconditionally` below exists
because a warning computed over a LOGGING property read as a ranking
problem. This field exists because a warning computed over a QUERY-SHAPE
property does the same thing.
An arm with a fixed query scores against one constant. Its decline rate is
therefore 0% or 100% and nothing in between — which of the two depends
only on whether the bar sits below or above that single number. So "never
returned nothing" says nothing at all about whether a floor is applied,
and `cannot_decline` — whose whole remedy is "check that it applies its
floor" — is uninformative here and skips these arms.
What IS informative for them is the mirror image, and `reply_preferences`
names it in its own docstring: every call returning nothing means the bar
sits above the constant, no traffic will ever move it, and the arm is
dead. That has happened — 69 consecutive declines at 0.0006 under the bar
(see `retrieval_surfaces`) — so it gets its own warning rather than
inheriting one written for arms whose score can vary.
"""
logs_unconditionally: bool = True
"""Whether this arm writes a row even when it returns NOTHING.
THE #3497 GUARD, and the reason this field is not merely documentation.
Both rule arms once logged only their hits, so their zero-result count was
structurally 0 — and a "cannot decline" warning computed over that would
have fired on a LOGGING BUG while reporting a ranking problem, sending the
reader to move a floor that was never involved. An arm whose logging has
not been confirmed unconditional is not eligible for that warning; it is
excluded from the check rather than trusted.
"""
def _p(source, kind, what, **kw) -> tuple[str, Point]:
return source, Point(source=source, kind=kind, what=what, **kw)
POINTS: dict[str, Point] = dict([
# ── Unbidden push arms ───────────────────────────────────────────────
_p("auto_inject", UNBIDDEN, "the notes menu offered at the prompt boundary"),
_p("write_path", UNBIDDEN, "prior art offered when a file is about to be written"),
_p("write_path_rule", UNBIDDEN, "rules that may govern the file being written"),
_p("pre_tool_rule", UNBIDDEN, "rules that may govern a command about to run"),
_p("prompt_rule", UNBIDDEN, "rules that may govern what the operator just asked"),
_p("preference_slot", UNBIDDEN,
"the one line reserved for a preference at the prompt boundary"),
_p("reuse_slot", UNBIDDEN, "the one line reserved for a reusable snippet"),
_p("lesson_slot", UNBIDDEN, "the one line reserved for a lesson"),
_p("rule_via_lesson", UNBIDDEN,
"a rule reached through a lesson confirmed as an instance of it",
expects_traffic=False,
quiet_because="searches only once some lesson has a confirmed link to "
"a rule; an install where none has been judged is "
"correctly silent here"),
# `fixed_query`: COMPLETION_QUERY is a module constant in
# services/reply_preferences.py, so this arm's top score is the same number
# on every call — measured at 0.791 across 45 consecutive calls, with p10,
# p50, p90, min and max all identical. Five equal percentiles is the tell.
_p("report_preference", UNBIDDEN,
"the fixed question asked when a task finishes: how should this report read",
fixed_query=True),
# The rulings arm (milestone 444). A LOOKUP, not a ranker: a path falls
# under a System's patterns or it does not, so nothing here writes to
# retrieval_logs and no score-shaped warning can apply. Its rows are in
# `system_usage_events`, read back as `retrieval_telemetry`'s
# `system_usage` block (#4769).
_p("rulings_pre_tool", UNBIDDEN,
"an area's rulings, shown because a command touched its files",
expects_traffic=False,
quiet_because="speaks only once a System has both path_patterns and a "
"Rulings section; an install where none does is "
"correctly silent here"),
_p("rulings_write_path", UNBIDDEN,
"an area's rulings, shown because an edit touched its files",
expects_traffic=False,
quiet_because="same as rulings_pre_tool — no System with patterns "
"and rulings, nothing to show"),
# ── Asked: a caller wanted a ranked list ─────────────────────────────
_p("mcp_search", ASKED, "an agent called search"),
_p("wide_net", ASKED, "an agent asked what might apply, with no bar"),
_p("rest_search", ASKED, "the web UI searched", expects_traffic=False,
quiet_because="only fires when a person uses the web UI; an install "
"driven entirely through MCP is correctly silent here"),
_p("browse_search", ASKED, "the browse listing searched", expects_traffic=False,
quiet_because="same as rest_search — a human-only entry point"),
# ── Ambient: carried along with a reply already being sent ───────────
_p("enter_project", AMBIENT, "the project handshake"),
_p("process_skill_sync", AMBIENT, "stored Processes synced into skills"),
_p("start_planning", AMBIENT, "the rules listed when a plan is opened"),
_p("get_task", AMBIENT, "the rules listed beside a task"),
_p("get_project", AMBIENT, "the rules listed beside a project"),
_p("get_milestone", AMBIENT, "the rules listed beside a milestone"),
# ── The write-path menu's own arms ───────────────────────────────────
#
# These are NOT in retrieval_logs. The place arm carries no score and so
# has no row there at all (#2085) — before it was split out, the arm
# firing on the strongest possible claim, "there is already a canonical
# helper in this exact file", was the one arm nobody could measure.
_p("write_path_place", AMBIENT, "a snippet already placed in this very file"),
_p("write_path_semantic", AMBIENT, "prior art matched by meaning"),
_p("write_path_sync", AMBIENT,
"records that should be updated alongside this edit"),
# ── Pull: a reader opened the record ─────────────────────────────────
_p("mcp_get_note", PULL, "an agent opened a note"),
_p("mcp_get_task", PULL, "an agent opened a task"),
_p("mcp_get_snippet", PULL, "an agent opened a snippet"),
_p("mcp_get_process", PULL, "an agent opened a process"),
_p("mcp_get_lesson", PULL, "an agent opened a lesson"),
_p("mcp_get_rule", PULL, "an agent opened a rule"),
_p("mcp_get_system", PULL, "an agent opened a System", expects_traffic=False,
quiet_because="a System's rulings arrive in full in the rulings line, "
"so opening one is the exception rather than the "
"measure; a window without it is not a gap"),
_p("rest_note", PULL, "a person opened a note", expects_traffic=False,
quiet_because="web UI only"),
_p("rest_task", PULL, "a person opened a task", expects_traffic=False,
quiet_because="web UI only"),
_p("rest_snippet", PULL, "a person opened a snippet", expects_traffic=False,
quiet_because="web UI only"),
_p("rest_lesson", PULL, "a person opened a lesson", expects_traffic=False,
quiet_because="web UI only"),
_p("rest_rule", PULL, "a person opened a rule", expects_traffic=False,
quiet_because="web UI only"),
])
# Call sites that pass `source` as a VARIABLE, and what they can pass.
#
# WHY THESE ARE DECLARED RATHER THAN RESOLVED. The registry test reads the
# call sites with `ast`, which settles a literal and a module constant but not
# a value that arrives through a parameter or a dict key. Rather than let
# those sites go unchecked — the silent half of a guard that appears to cover
# everything — they are named here, and the test FAILS if it meets an
# unresolved site that is not in this list.
#
# THE HONEST LIMIT, stated so nobody over-trusts this: the test pins the SITE,
# not the values. Adding a fourth arm inside `plugin_context`'s `by_arm` fan
# out would not fail here. What catches that is `assert_registered`, which the
# recorders call as the row is written.
# Keyed by the EXACT string the extractor produces, matched by equality rather
# than by prefix. A prefix match was the first spelling and it silently matched
# nothing, because the paths are relative to `src/` and so begin `scribe/` —
# every site sailed through a check that looked like it was running.
FAN_OUT_SITES: dict[str, tuple[str, ...]] = {
"scribe/services/plugin_context.py::record_surfaced(source=arm)": (
"write_path_place", "write_path_semantic", "write_path_sync",
),
"scribe/services/system_rulings.py::record_system_surfaced(source=source)": (
"rulings_pre_tool", "rulings_write_path",
),
"scribe/services/rulebooks.py::record_rule_surfaced(source=source)": (
"enter_project", "start_planning", "get_task", "get_project",
"get_milestone",
),
}
def get_point(source: str) -> Point | None:
"""The registered point for a source, or None if it is not declared."""
return POINTS.get(source)
def sources_expected_to_emit() -> list[str]:
"""Every point whose silence over an ACTIVE window is worth reporting."""
return [s for s, p in POINTS.items() if p.expects_traffic]
def is_registered(source: str) -> bool:
"""Is this source declared?
Read by the TELEMETRY READOUT, not by the recorders, and that placement is
the design rather than an accident of where it was easy. Validating at
write time would put a dict lookup on every retrieval to catch a mistake
that is only ever made once, when an arm is added — and it could not
raise anyway, because telemetry that breaks its caller is the worse bug
(`record_retrieval` is fire-and-forget for exactly that reason). So the
check sits where the cost is already paid and the reader is already
looking: a source that appears in the rows and not in this file comes back
as an `unregistered_source` warning.
That is what covers FAN_OUT_SITES above. The static test pins those sites
but not the values they can pass, so a fourth arm added inside one would
slip past it — and then show up here the first time it fires.
"""
return source in POINTS