refactor(retrieval): the specs are the registry - SURFACES, the ranked POINTS rows and RANKED_SOURCES are read off the pipeline specs (milestone 456 step 5, #4907)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 2m1s
CI & Build / Build & push image (push) Successful in 27s

Each arm spec now carries its tuning pair (Surface, moved verbatim into
retrieval_pipeline) and a Declared block - what the telemetry readout
must know and cannot read off its rows. The two ranked stages that are
not arms (preference_slot, rule_via_lesson) are RankedSource specs, and
the note slots carry their own declaration.

- retrieval_surfaces.SURFACES = the TUNED_ARMS tuning, same order
- retrieval_registry.POINTS ranked rows = one per spec in RANKED; the
  lookups, asked, ambient and pull rows stay declared there
- rule_usage.RANKED_SOURCES = RULE_RANKED_SOURCES + moment_rule

Settings keys, defaults, prose and order are unchanged (checked field by
field against HEAD). tests/test_retrieval_specs.py pins the derivation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-10-05 20:11:20 -04:00
co-authored by Claude Opus 5.5
parent 2b8f41229d
commit 869046dda2
5 changed files with 360 additions and 197 deletions
+254 -1
View File
@@ -331,6 +331,110 @@ def _rule_hint_line(
)
# ── What a spec declares about itself (milestone 456 step 5) ─────────────
#
# A ranked surface used to be described in three hand-kept lists besides its
# code: its tuning pair in `retrieval_surfaces.SURFACES`, its row in
# `retrieval_registry.POINTS`, and — for a rule surface — its membership in
# `rule_usage.RANKED_SOURCES`. Tests existed to make the four agree. Now the
# spec carries all of it, and the three lists are read off the specs, so an
# arm added here is tunable, measured and counted by construction.
#
# THE CLASSES LIVE HERE, NOT BESIDE THE LISTS, for the import graph: the
# registry, the tuning table and the usage counter all read the specs, so the
# specs cannot import any of them back.
@dataclass(frozen=True)
class Surface:
"""One push arm's tunable pair, plus enough prose to tune it responsibly.
`asks` / `over` / `fires` are not documentation for this file — they are
rendered by the tuning tool and the Settings UI. A floor cannot be moved
sensibly by anyone, model or human, who does not know what the query is, what
corpus it runs against, or how often it costs something. Those three facts
are exactly what separates these arms from each other, and they were
previously recoverable only by reading `plugin_context.py`.
"""
name: str
"""The telemetry `source` value, and the join key.
MUST equal the string this arm passes to `record_retrieval`. Everything
useful about tuning depends on that identity: the tool that moves a floor
and the table that says what the floor did have to be talking about the same
arm. A test asserts it rather than a comment asking nicely.
"""
floor_key: str
floor_default: float
budget_key: str
budget_default: int
asks: str
over: str
fires: str
measured_model: str = "BAAI/bge-small-en-v1.5"
measured_shape: int = 1
"""What the SHIPPED defaults above were measured against (#4104).
A floor is a distance in one embedding model's geometry, over documents cut
one particular way. Either can change, and when one does every number in
this table describes something that no longer exists.
TWO FIELDS, NEVER ONE FUSED STRING (rule 149). A mismatch has to be able to
say WHICH half moved: a new embedding model and a re-cut document shape
invalidate the same numbers for different reasons and call for different
responses. `"<model>@<n>"` could only report that something changed, which
is the answer nobody can act on. Same reason `calibration_stamp()` returns
a dict and the event table gives each half its own column.
Recorded per surface rather than once for the module because they need not
move together: a surface retuned after a model change carries the new stamp
while its untouched siblings still carry the old one, and telling those
apart is the whole job.
LITERALS, deliberately, rather than an import of the live values — a stamp
says what was true when the number was chosen, so one that tracked the
current model would always agree with it and could never report staleness.
"""
budget_falls_back_to: str = ""
"""A budget key to inherit when this surface has none of its own set.
Only `write_path` uses it, and only because it USED to share auto-inject's
`top_k` outright. Giving it a key without this would silently reset the
budget of every install that had tuned the shared one — a behaviour change
delivered as a default, which is the shape of regression nobody reports
because nothing looks broken.
"""
@dataclass(frozen=True)
class Declared:
"""What the telemetry readout must know about a ranked source and cannot
read off its rows — `retrieval_registry.Point`'s fields, for the specs
that produce one. Every ranked source is UNBIDDEN: nobody asked for it."""
what: str
"""One line, for an agent reading a warning that names this source."""
quiet_because: str = ""
"""Set when silence over an active window is correct, saying why (#2475)."""
fixed_query: bool = False
"""The arm always searches the same string, so its decline rate is 0% or
100% and `cannot_decline` says nothing about it."""
@dataclass(frozen=True)
class RankedSource:
"""A ranked source that is a stage of an arm rather than an arm: it
records under its own name, and is measured but never tuned."""
source: str
declared: Declared
# ── The specs ────────────────────────────────────────────────────────────
@@ -364,6 +468,12 @@ class RuleArm:
recorded as surfaced — the rest were ranked, and the call row counts them,
but nobody was shown them."""
tuning: Surface | None = None
"""Its floor and budget, and the prose a tuner reads. Every arm has one;
`SURFACES` is read off them."""
declared: Declared | None = None
# Today's differences, reproduced exactly (milestone 456 step 2). Whether the
# prompt arm should band, and whether the act arms should reserve a
@@ -371,14 +481,47 @@ class RuleArm:
PROMPT_RULE = RuleArm(
"prompt_rule", band=False, compact_tail=False, checkpoint=False,
preference_slot=True,
tuning=Surface(
name="prompt_rule",
floor_key="kb_promptrule_threshold",
floor_default=0.72,
budget_key="kb_promptrule_top_k",
budget_default=3,
asks="the operator's message, against rule triggers",
over="global rules plus the bound project's own",
fires="once per operator turn",
),
declared=Declared("rules that may govern what the operator just asked"),
)
PRE_TOOL_RULE = RuleArm(
"pre_tool_rule", band=True, compact_tail=True, checkpoint=True,
preference_slot=False,
tuning=Surface(
name="pre_tool_rule",
floor_key="kb_toolrule_threshold",
floor_default=0.68,
budget_key="kb_toolrule_top_k",
budget_default=5,
asks="the command about to run, against rule triggers",
over="global rules plus the bound project's own",
fires="before every Bash call — the busiest arm there is",
),
declared=Declared("rules that may govern a command about to run"),
)
WRITE_PATH_RULE = RuleArm(
"write_path_rule", band=True, compact_tail=True, checkpoint=True,
preference_slot=False,
tuning=Surface(
name="write_path_rule",
floor_key="kb_rulehint_threshold",
floor_default=0.72,
budget_key="kb_rulehint_top_k",
budget_default=5,
asks="the code being written, against rule triggers",
over="global rules plus the bound project's own",
fires="before every Write and Edit",
),
declared=Declared("rules that may govern the file being written"),
)
# The completion report's preferences (milestone 409 step 4): a FIXED query,
# preferences only, read by update_task as records rather than as lines. Its
@@ -386,6 +529,30 @@ WRITE_PATH_RULE = RuleArm(
REPORT_PREFERENCE = RuleArm(
"report_preference", band=False, compact_tail=False, checkpoint=False,
preference_slot=False, kind="preference",
tuning=Surface(
name="report_preference",
floor_key="kb_reportpref_threshold",
floor_default=0.72,
budget_key="kb_reportpref_top_k",
budget_default=3,
# THE ONE FIXED QUERY, and the reason this arm behaves unlike the rest.
# The others score something that varies per call; this one scores a
# constant string, so its top score for a given corpus is also a
# constant. A floor a hair above that constant is not a quiet arm, it
# is a dead one, and no amount of traffic will ever reveal it — which
# is precisely how this arm spent 69 calls declining the same record.
asks="a fixed question about how to lay out a completion report",
over="preferences",
fires="when a task finishes",
),
# `fixed_query`: COMPLETION_QUERY is a module constant, so this arm's top
# score is the same number on every call — measured at 0.791 across 45
# consecutive calls, with p10, p50, p90, min and max all identical. Five
# equal percentiles is the tell.
declared=Declared(
"the fixed question asked when a task finishes: how should this report read",
fixed_query=True,
),
)
# The backstop for every arm that ran earlier in the turn and missed: the
# finished reply against every rule's trigger. Its floor IS its stop bar
@@ -393,12 +560,35 @@ REPORT_PREFERENCE = RuleArm(
REPLY_RULE = RuleArm(
"reply_rule", band=False, compact_tail=False, checkpoint=True,
preference_slot=False, stop_only=True,
# THE REPLY BACKSTOP (milestone 458, folded in from 456 step 8). Its floor
# is a STOP bar, not a hint bar: at the end of a turn nothing can be shown
# beside the reply, so a hit either holds the reply for one read or says
# nothing. Hence a default at the checkpoint's level and a budget of one —
# the call row's results are then exactly the rule that would hold.
tuning=Surface(
name="reply_rule",
floor_key="kb_replyrule_threshold",
floor_default=0.80,
budget_key="kb_replyrule_top_k",
budget_default=1,
asks="the reply that ends a turn, against rule triggers",
over="global rules plus the bound project's own",
fires="once per turn, when the reply is finished",
),
declared=Declared(
"a rule that holds the finished reply for one read — the backstop "
"for whatever the earlier arms missed"
),
)
RULE_ARMS: tuple[RuleArm, ...] = (
WRITE_PATH_RULE, PRE_TOOL_RULE, PROMPT_RULE, REPORT_PREFERENCE, REPLY_RULE,
)
PREFERENCE_SLOT_SOURCE = "preference_slot"
PREFERENCE_SLOT = RankedSource(
PREFERENCE_SLOT_SOURCE,
Declared("the one line reserved for a preference at the prompt boundary"),
)
# Every source the shared stages below can record under — what the registry
# declares for this module's fan-out sites, since `source` reaches the
@@ -861,15 +1051,21 @@ class NoteSlot:
"""The slot's line is recorded as surfaced under the slot's own source,
and so never again under the arm's."""
declared: Declared | None = None
# Order is load-bearing: reuse evicts the menu's weakest hit while the lesson
# slot extends, so running them the other way round would let a reserved
# lesson be the line reuse throws off — a slot another slot can silently undo
# is not a guarantee.
REUSE_SLOT = NoteSlot("reuse_slot", ("snippet", "process"), evicts=True)
REUSE_SLOT = NoteSlot(
"reuse_slot", ("snippet", "process"), evicts=True,
declared=Declared("the one line reserved for a reusable snippet"),
)
LESSON_SLOT = NoteSlot(
"lesson_slot", (LESSON_NOTE_TYPE,), evicts=False,
include_global_kinds=True, books_own=True,
declared=Declared("the one line reserved for a lesson"),
)
@@ -899,9 +1095,23 @@ class NoteArm:
kept out of the search (`exclude_ids`) — same-call duplication, which is a
different claim from the session ledger (#4101)."""
tuning: Surface | None = None
declared: Declared | None = None
AUTO_INJECT = NoteArm(
"auto_inject", surfaced_as="auto_inject", slots=(REUSE_SLOT, LESSON_SLOT),
tuning=Surface(
name="auto_inject",
floor_key="kb_autoinject_threshold",
floor_default=0.55,
budget_key="kb_autoinject_top_k",
budget_default=3,
asks="the operator's message, as they typed it",
over="notes, snippets, processes and issues",
fires="once per operator turn",
),
declared=Declared("the notes menu offered at the prompt boundary"),
)
# Snippets AND recorded experience (#2246): an issue saying "we tried this and
# it deadlocked" is prior art for the code about to be written. `task_kind`
@@ -914,6 +1124,18 @@ WRITE_PATH = NoteArm(
"write_path", surfaced_as="write_path_semantic",
note_type=("snippet", "note", LESSON_NOTE_TYPE), task_kind="issue",
withholds=True,
tuning=Surface(
name="write_path",
floor_key="kb_writepath_threshold",
floor_default=0.68,
budget_key="kb_writepath_top_k",
budget_default=3,
budget_falls_back_to="kb_autoinject_top_k",
asks="the code being written, rewritten as a concept query",
over="snippets and recorded issues",
fires="before every Write and Edit",
),
declared=Declared("prior art offered when a file is about to be written"),
)
NOTE_ARMS: tuple[NoteArm, ...] = (AUTO_INJECT, WRITE_PATH)
NOTE_SLOTS: tuple[NoteSlot, ...] = (REUSE_SLOT, LESSON_SLOT)
@@ -1231,6 +1453,15 @@ async def run_note_arm(
# line costs one line. Logged as its own source so it can be judged (#4636).
VIA_LESSON_SOURCE = "rule_via_lesson"
VIA_LESSON = RankedSource(
VIA_LESSON_SOURCE,
Declared(
"a rule reached through a lesson confirmed as an instance of it",
quiet_because="searches only once some lesson has a confirmed link to "
"a rule; an install where none has been judged is "
"correctly silent here",
),
)
VIA_LESSON_LIMIT = 1
# Lessons fetched before keeping only the linked ones. The search cannot be
# told "linked lessons only", so it overfetches and filters; on a corpus where
@@ -1340,6 +1571,28 @@ async def run_via_lesson_arm(
return RuleResult()
# ── The specs, read as lists (milestone 456 step 5) ──────────────────────
#
# What `retrieval_surfaces.SURFACES`, the ranked rows of
# `retrieval_registry.POINTS` and `rule_usage.RANKED_SOURCES` are read from.
# The order is the order the Settings page and the readouts list them in.
TUNED_ARMS: tuple = (*NOTE_ARMS, *RULE_ARMS)
"""Every arm with a floor and a budget: one per tunable surface."""
RANKED: tuple = (
*TUNED_ARMS, PREFERENCE_SLOT, *NOTE_SLOTS, VIA_LESSON,
)
"""Every source that RANKED what it showed — the arms, and the stages that
run a query of their own and record under their own name."""
RULE_RANKED_SOURCES: tuple[str, ...] = (
*(arm.source for arm in RULE_ARMS), PREFERENCE_SLOT_SOURCE, VIA_LESSON_SOURCE,
)
"""The ranked sources that surface RULES — a ranker chose each line, so a
pull can confirm or refute it (`rule_usage.RANKED_SOURCES`)."""
# ── The moment arm (milestone 458) ───────────────────────────────────────
#
# A LOOKUP beside the ranked arms, not one of them. A rule mounted on a moment