refactor(retrieval): the specs are the registry - SURFACES, the ranked POINTS rows and RANKED_SOURCES are read off the pipeline specs (milestone 456 step 5, #4907)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 2m1s
CI & Build / Build & push image (push) Successful in 27s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 2m1s
CI & Build / Build & push image (push) Successful in 27s
Each arm spec now carries its tuning pair (Surface, moved verbatim into retrieval_pipeline) and a Declared block - what the telemetry readout must know and cannot read off its rows. The two ranked stages that are not arms (preference_slot, rule_via_lesson) are RankedSource specs, and the note slots carry their own declaration. - retrieval_surfaces.SURFACES = the TUNED_ARMS tuning, same order - retrieval_registry.POINTS ranked rows = one per spec in RANKED; the lookups, asked, ambient and pull rows stay declared there - rule_usage.RANKED_SOURCES = RULE_RANKED_SOURCES + moment_rule Settings keys, defaults, prose and order are unchanged (checked field by field against HEAD). tests/test_retrieval_specs.py pins the derivation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -331,6 +331,110 @@ def _rule_hint_line(
|
||||
)
|
||||
|
||||
|
||||
# ── What a spec declares about itself (milestone 456 step 5) ─────────────
|
||||
#
|
||||
# A ranked surface used to be described in three hand-kept lists besides its
|
||||
# code: its tuning pair in `retrieval_surfaces.SURFACES`, its row in
|
||||
# `retrieval_registry.POINTS`, and — for a rule surface — its membership in
|
||||
# `rule_usage.RANKED_SOURCES`. Tests existed to make the four agree. Now the
|
||||
# spec carries all of it, and the three lists are read off the specs, so an
|
||||
# arm added here is tunable, measured and counted by construction.
|
||||
#
|
||||
# THE CLASSES LIVE HERE, NOT BESIDE THE LISTS, for the import graph: the
|
||||
# registry, the tuning table and the usage counter all read the specs, so the
|
||||
# specs cannot import any of them back.
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Surface:
|
||||
"""One push arm's tunable pair, plus enough prose to tune it responsibly.
|
||||
|
||||
`asks` / `over` / `fires` are not documentation for this file — they are
|
||||
rendered by the tuning tool and the Settings UI. A floor cannot be moved
|
||||
sensibly by anyone, model or human, who does not know what the query is, what
|
||||
corpus it runs against, or how often it costs something. Those three facts
|
||||
are exactly what separates these arms from each other, and they were
|
||||
previously recoverable only by reading `plugin_context.py`.
|
||||
"""
|
||||
|
||||
name: str
|
||||
"""The telemetry `source` value, and the join key.
|
||||
|
||||
MUST equal the string this arm passes to `record_retrieval`. Everything
|
||||
useful about tuning depends on that identity: the tool that moves a floor
|
||||
and the table that says what the floor did have to be talking about the same
|
||||
arm. A test asserts it rather than a comment asking nicely.
|
||||
"""
|
||||
|
||||
floor_key: str
|
||||
floor_default: float
|
||||
budget_key: str
|
||||
budget_default: int
|
||||
asks: str
|
||||
over: str
|
||||
fires: str
|
||||
measured_model: str = "BAAI/bge-small-en-v1.5"
|
||||
measured_shape: int = 1
|
||||
"""What the SHIPPED defaults above were measured against (#4104).
|
||||
|
||||
A floor is a distance in one embedding model's geometry, over documents cut
|
||||
one particular way. Either can change, and when one does every number in
|
||||
this table describes something that no longer exists.
|
||||
|
||||
TWO FIELDS, NEVER ONE FUSED STRING (rule 149). A mismatch has to be able to
|
||||
say WHICH half moved: a new embedding model and a re-cut document shape
|
||||
invalidate the same numbers for different reasons and call for different
|
||||
responses. `"<model>@<n>"` could only report that something changed, which
|
||||
is the answer nobody can act on. Same reason `calibration_stamp()` returns
|
||||
a dict and the event table gives each half its own column.
|
||||
|
||||
Recorded per surface rather than once for the module because they need not
|
||||
move together: a surface retuned after a model change carries the new stamp
|
||||
while its untouched siblings still carry the old one, and telling those
|
||||
apart is the whole job.
|
||||
|
||||
LITERALS, deliberately, rather than an import of the live values — a stamp
|
||||
says what was true when the number was chosen, so one that tracked the
|
||||
current model would always agree with it and could never report staleness.
|
||||
"""
|
||||
|
||||
budget_falls_back_to: str = ""
|
||||
"""A budget key to inherit when this surface has none of its own set.
|
||||
|
||||
Only `write_path` uses it, and only because it USED to share auto-inject's
|
||||
`top_k` outright. Giving it a key without this would silently reset the
|
||||
budget of every install that had tuned the shared one — a behaviour change
|
||||
delivered as a default, which is the shape of regression nobody reports
|
||||
because nothing looks broken.
|
||||
"""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Declared:
|
||||
"""What the telemetry readout must know about a ranked source and cannot
|
||||
read off its rows — `retrieval_registry.Point`'s fields, for the specs
|
||||
that produce one. Every ranked source is UNBIDDEN: nobody asked for it."""
|
||||
|
||||
what: str
|
||||
"""One line, for an agent reading a warning that names this source."""
|
||||
|
||||
quiet_because: str = ""
|
||||
"""Set when silence over an active window is correct, saying why (#2475)."""
|
||||
|
||||
fixed_query: bool = False
|
||||
"""The arm always searches the same string, so its decline rate is 0% or
|
||||
100% and `cannot_decline` says nothing about it."""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class RankedSource:
|
||||
"""A ranked source that is a stage of an arm rather than an arm: it
|
||||
records under its own name, and is measured but never tuned."""
|
||||
|
||||
source: str
|
||||
declared: Declared
|
||||
|
||||
|
||||
# ── The specs ────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
@@ -364,6 +468,12 @@ class RuleArm:
|
||||
recorded as surfaced — the rest were ranked, and the call row counts them,
|
||||
but nobody was shown them."""
|
||||
|
||||
tuning: Surface | None = None
|
||||
"""Its floor and budget, and the prose a tuner reads. Every arm has one;
|
||||
`SURFACES` is read off them."""
|
||||
|
||||
declared: Declared | None = None
|
||||
|
||||
|
||||
# Today's differences, reproduced exactly (milestone 456 step 2). Whether the
|
||||
# prompt arm should band, and whether the act arms should reserve a
|
||||
@@ -371,14 +481,47 @@ class RuleArm:
|
||||
PROMPT_RULE = RuleArm(
|
||||
"prompt_rule", band=False, compact_tail=False, checkpoint=False,
|
||||
preference_slot=True,
|
||||
tuning=Surface(
|
||||
name="prompt_rule",
|
||||
floor_key="kb_promptrule_threshold",
|
||||
floor_default=0.72,
|
||||
budget_key="kb_promptrule_top_k",
|
||||
budget_default=3,
|
||||
asks="the operator's message, against rule triggers",
|
||||
over="global rules plus the bound project's own",
|
||||
fires="once per operator turn",
|
||||
),
|
||||
declared=Declared("rules that may govern what the operator just asked"),
|
||||
)
|
||||
PRE_TOOL_RULE = RuleArm(
|
||||
"pre_tool_rule", band=True, compact_tail=True, checkpoint=True,
|
||||
preference_slot=False,
|
||||
tuning=Surface(
|
||||
name="pre_tool_rule",
|
||||
floor_key="kb_toolrule_threshold",
|
||||
floor_default=0.68,
|
||||
budget_key="kb_toolrule_top_k",
|
||||
budget_default=5,
|
||||
asks="the command about to run, against rule triggers",
|
||||
over="global rules plus the bound project's own",
|
||||
fires="before every Bash call — the busiest arm there is",
|
||||
),
|
||||
declared=Declared("rules that may govern a command about to run"),
|
||||
)
|
||||
WRITE_PATH_RULE = RuleArm(
|
||||
"write_path_rule", band=True, compact_tail=True, checkpoint=True,
|
||||
preference_slot=False,
|
||||
tuning=Surface(
|
||||
name="write_path_rule",
|
||||
floor_key="kb_rulehint_threshold",
|
||||
floor_default=0.72,
|
||||
budget_key="kb_rulehint_top_k",
|
||||
budget_default=5,
|
||||
asks="the code being written, against rule triggers",
|
||||
over="global rules plus the bound project's own",
|
||||
fires="before every Write and Edit",
|
||||
),
|
||||
declared=Declared("rules that may govern the file being written"),
|
||||
)
|
||||
# The completion report's preferences (milestone 409 step 4): a FIXED query,
|
||||
# preferences only, read by update_task as records rather than as lines. Its
|
||||
@@ -386,6 +529,30 @@ WRITE_PATH_RULE = RuleArm(
|
||||
REPORT_PREFERENCE = RuleArm(
|
||||
"report_preference", band=False, compact_tail=False, checkpoint=False,
|
||||
preference_slot=False, kind="preference",
|
||||
tuning=Surface(
|
||||
name="report_preference",
|
||||
floor_key="kb_reportpref_threshold",
|
||||
floor_default=0.72,
|
||||
budget_key="kb_reportpref_top_k",
|
||||
budget_default=3,
|
||||
# THE ONE FIXED QUERY, and the reason this arm behaves unlike the rest.
|
||||
# The others score something that varies per call; this one scores a
|
||||
# constant string, so its top score for a given corpus is also a
|
||||
# constant. A floor a hair above that constant is not a quiet arm, it
|
||||
# is a dead one, and no amount of traffic will ever reveal it — which
|
||||
# is precisely how this arm spent 69 calls declining the same record.
|
||||
asks="a fixed question about how to lay out a completion report",
|
||||
over="preferences",
|
||||
fires="when a task finishes",
|
||||
),
|
||||
# `fixed_query`: COMPLETION_QUERY is a module constant, so this arm's top
|
||||
# score is the same number on every call — measured at 0.791 across 45
|
||||
# consecutive calls, with p10, p50, p90, min and max all identical. Five
|
||||
# equal percentiles is the tell.
|
||||
declared=Declared(
|
||||
"the fixed question asked when a task finishes: how should this report read",
|
||||
fixed_query=True,
|
||||
),
|
||||
)
|
||||
# The backstop for every arm that ran earlier in the turn and missed: the
|
||||
# finished reply against every rule's trigger. Its floor IS its stop bar
|
||||
@@ -393,12 +560,35 @@ REPORT_PREFERENCE = RuleArm(
|
||||
REPLY_RULE = RuleArm(
|
||||
"reply_rule", band=False, compact_tail=False, checkpoint=True,
|
||||
preference_slot=False, stop_only=True,
|
||||
# THE REPLY BACKSTOP (milestone 458, folded in from 456 step 8). Its floor
|
||||
# is a STOP bar, not a hint bar: at the end of a turn nothing can be shown
|
||||
# beside the reply, so a hit either holds the reply for one read or says
|
||||
# nothing. Hence a default at the checkpoint's level and a budget of one —
|
||||
# the call row's results are then exactly the rule that would hold.
|
||||
tuning=Surface(
|
||||
name="reply_rule",
|
||||
floor_key="kb_replyrule_threshold",
|
||||
floor_default=0.80,
|
||||
budget_key="kb_replyrule_top_k",
|
||||
budget_default=1,
|
||||
asks="the reply that ends a turn, against rule triggers",
|
||||
over="global rules plus the bound project's own",
|
||||
fires="once per turn, when the reply is finished",
|
||||
),
|
||||
declared=Declared(
|
||||
"a rule that holds the finished reply for one read — the backstop "
|
||||
"for whatever the earlier arms missed"
|
||||
),
|
||||
)
|
||||
RULE_ARMS: tuple[RuleArm, ...] = (
|
||||
WRITE_PATH_RULE, PRE_TOOL_RULE, PROMPT_RULE, REPORT_PREFERENCE, REPLY_RULE,
|
||||
)
|
||||
|
||||
PREFERENCE_SLOT_SOURCE = "preference_slot"
|
||||
PREFERENCE_SLOT = RankedSource(
|
||||
PREFERENCE_SLOT_SOURCE,
|
||||
Declared("the one line reserved for a preference at the prompt boundary"),
|
||||
)
|
||||
|
||||
# Every source the shared stages below can record under — what the registry
|
||||
# declares for this module's fan-out sites, since `source` reaches the
|
||||
@@ -861,15 +1051,21 @@ class NoteSlot:
|
||||
"""The slot's line is recorded as surfaced under the slot's own source,
|
||||
and so never again under the arm's."""
|
||||
|
||||
declared: Declared | None = None
|
||||
|
||||
|
||||
# Order is load-bearing: reuse evicts the menu's weakest hit while the lesson
|
||||
# slot extends, so running them the other way round would let a reserved
|
||||
# lesson be the line reuse throws off — a slot another slot can silently undo
|
||||
# is not a guarantee.
|
||||
REUSE_SLOT = NoteSlot("reuse_slot", ("snippet", "process"), evicts=True)
|
||||
REUSE_SLOT = NoteSlot(
|
||||
"reuse_slot", ("snippet", "process"), evicts=True,
|
||||
declared=Declared("the one line reserved for a reusable snippet"),
|
||||
)
|
||||
LESSON_SLOT = NoteSlot(
|
||||
"lesson_slot", (LESSON_NOTE_TYPE,), evicts=False,
|
||||
include_global_kinds=True, books_own=True,
|
||||
declared=Declared("the one line reserved for a lesson"),
|
||||
)
|
||||
|
||||
|
||||
@@ -899,9 +1095,23 @@ class NoteArm:
|
||||
kept out of the search (`exclude_ids`) — same-call duplication, which is a
|
||||
different claim from the session ledger (#4101)."""
|
||||
|
||||
tuning: Surface | None = None
|
||||
declared: Declared | None = None
|
||||
|
||||
|
||||
AUTO_INJECT = NoteArm(
|
||||
"auto_inject", surfaced_as="auto_inject", slots=(REUSE_SLOT, LESSON_SLOT),
|
||||
tuning=Surface(
|
||||
name="auto_inject",
|
||||
floor_key="kb_autoinject_threshold",
|
||||
floor_default=0.55,
|
||||
budget_key="kb_autoinject_top_k",
|
||||
budget_default=3,
|
||||
asks="the operator's message, as they typed it",
|
||||
over="notes, snippets, processes and issues",
|
||||
fires="once per operator turn",
|
||||
),
|
||||
declared=Declared("the notes menu offered at the prompt boundary"),
|
||||
)
|
||||
# Snippets AND recorded experience (#2246): an issue saying "we tried this and
|
||||
# it deadlocked" is prior art for the code about to be written. `task_kind`
|
||||
@@ -914,6 +1124,18 @@ WRITE_PATH = NoteArm(
|
||||
"write_path", surfaced_as="write_path_semantic",
|
||||
note_type=("snippet", "note", LESSON_NOTE_TYPE), task_kind="issue",
|
||||
withholds=True,
|
||||
tuning=Surface(
|
||||
name="write_path",
|
||||
floor_key="kb_writepath_threshold",
|
||||
floor_default=0.68,
|
||||
budget_key="kb_writepath_top_k",
|
||||
budget_default=3,
|
||||
budget_falls_back_to="kb_autoinject_top_k",
|
||||
asks="the code being written, rewritten as a concept query",
|
||||
over="snippets and recorded issues",
|
||||
fires="before every Write and Edit",
|
||||
),
|
||||
declared=Declared("prior art offered when a file is about to be written"),
|
||||
)
|
||||
NOTE_ARMS: tuple[NoteArm, ...] = (AUTO_INJECT, WRITE_PATH)
|
||||
NOTE_SLOTS: tuple[NoteSlot, ...] = (REUSE_SLOT, LESSON_SLOT)
|
||||
@@ -1231,6 +1453,15 @@ async def run_note_arm(
|
||||
# line costs one line. Logged as its own source so it can be judged (#4636).
|
||||
|
||||
VIA_LESSON_SOURCE = "rule_via_lesson"
|
||||
VIA_LESSON = RankedSource(
|
||||
VIA_LESSON_SOURCE,
|
||||
Declared(
|
||||
"a rule reached through a lesson confirmed as an instance of it",
|
||||
quiet_because="searches only once some lesson has a confirmed link to "
|
||||
"a rule; an install where none has been judged is "
|
||||
"correctly silent here",
|
||||
),
|
||||
)
|
||||
VIA_LESSON_LIMIT = 1
|
||||
# Lessons fetched before keeping only the linked ones. The search cannot be
|
||||
# told "linked lessons only", so it overfetches and filters; on a corpus where
|
||||
@@ -1340,6 +1571,28 @@ async def run_via_lesson_arm(
|
||||
return RuleResult()
|
||||
|
||||
|
||||
# ── The specs, read as lists (milestone 456 step 5) ──────────────────────
|
||||
#
|
||||
# What `retrieval_surfaces.SURFACES`, the ranked rows of
|
||||
# `retrieval_registry.POINTS` and `rule_usage.RANKED_SOURCES` are read from.
|
||||
# The order is the order the Settings page and the readouts list them in.
|
||||
|
||||
TUNED_ARMS: tuple = (*NOTE_ARMS, *RULE_ARMS)
|
||||
"""Every arm with a floor and a budget: one per tunable surface."""
|
||||
|
||||
RANKED: tuple = (
|
||||
*TUNED_ARMS, PREFERENCE_SLOT, *NOTE_SLOTS, VIA_LESSON,
|
||||
)
|
||||
"""Every source that RANKED what it showed — the arms, and the stages that
|
||||
run a query of their own and record under their own name."""
|
||||
|
||||
RULE_RANKED_SOURCES: tuple[str, ...] = (
|
||||
*(arm.source for arm in RULE_ARMS), PREFERENCE_SLOT_SOURCE, VIA_LESSON_SOURCE,
|
||||
)
|
||||
"""The ranked sources that surface RULES — a ranker chose each line, so a
|
||||
pull can confirm or refute it (`rule_usage.RANKED_SOURCES`)."""
|
||||
|
||||
|
||||
# ── The moment arm (milestone 458) ───────────────────────────────────────
|
||||
#
|
||||
# A LOOKUP beside the ranked arms, not one of them. A rule mounted on a moment
|
||||
|
||||
Reference in New Issue
Block a user