Files
FabledScribe/src/scribe/services/retrieval_surfaces.py
T
bvandeusenandClaude Opus 5.5 f68b73f922
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m10s
CI & Build / Python tests (push) Successful in 1m54s
CI & Build / Build & push image (push) Successful in 31s
feat(500): retire the fixed-question preference arm - completion-report preferences ride the moments
report_preference searched one constant string when a task closed and handed
matches back as update_task's reply_preferences. Milestone 500 step 3 delivers
the reply shapes and the preferences mounted beside them on the moments, so
the arm goes, whole (rule 22):

- services/reply_preferences.py, the REPORT_PREFERENCE RuleArm, its re-scorer
  and corpus entries, and the REPORTPREF constants
- update_task's reply_preferences block and its cue; the docstring now points
  at moment_rules
- the fixed_query field and the fixed_query_never_clears warning: this arm was
  the only one that set it, so the concept and the cannot_decline exemption
  go with it
- the two Settings fields for its floor and budget
- migration 0121 deletes its rows (logs, judgments, rule-usage events, tuning
  history, settings keys), operator-approved 2026-10-09. Without it the
  readout would call the source unregistered and its surfacings would count
  as ambient. The unindexed rule_usage delete is bounded by created_at.

#5496 (step 4 of milestone 500).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-09 15:51:04 -04:00

171 lines
7.9 KiB
Python

"""One registry of the retrieval surfaces, and the two numbers each one has (#4102).
WHY THIS EXISTS
Six push arms each carried their own loose copy of the same shape: a settings
key, a default, and a limit that was usually a module constant nobody could
change. The read-and-clamp was written out separately in `plugin_context` (twice
over, for auto-inject and the write path), in three rule arms, and again in
the completion-report arm (since retired, milestone 500). That was survivable
while the numbers were shipped constants an operator occasionally edited.
It stops being survivable once the numbers are meant to MOVE. The operator's
decision for this step:
"the floor should be chosen and adjusted by the model using it. we've come
back to something either fails or has to be looked at by the user we need a
model consistent surface for the adjustment of these floor values. the user
should be able to touch it but the model should be the thing handling it 9
times out of 10."
A tuning surface cannot be "model consistent" if every arm spells its own
configuration differently. So the arms stop owning their numbers and read them
from here instead, and the tuning tool, the routes and the Settings UI all
enumerate THIS table rather than hard-coding six special cases.
WHAT A FLOOR IS NOW, AND WHAT IT IS NOT
It is not a relevance judgement. Relevance is decided by the reader, which is
the only participant that can read a trigger against a situation — that is the
milestone's whole argument, and the injected line has always said so out loud
("read it before deciding it does not apply").
A floor answers the cheaper question: *is this worth ranking at all*. `k` is
what binds, and `k` is a BUDGET — how much of this surface's attention a
candidate list may spend. An arm that fires before every Bash call cannot
afford what an arm that fires once a turn can.
WHY THE DEFAULTS BELOW ARE STARTING POINTS AND NOT ANSWERS
A cosine score is a distance in BAAI/bge-small-en-v1.5's vector space, measured
against THIS corpus. It cannot transfer to an install with different records,
and rule 115 forbids defending a shipped default from this instance's
telemetry — which is what made the old design unbuildable: every number was a
guess everywhere except here, and there was no mechanism that could ever
improve it.
The mechanism is the fix. Scribe ships a starting point and the means to
correct it, so the values below carry the measurement that motivated them (see
the long comments in `plugin_context.py`, which are kept where they are because
they record how each number was first arrived at) without claiming to be right
for anybody else.
HOW A FLOOR SHOULD ACTUALLY BE MOVED
By reading the records the floor refused — `retrieval_telemetry(
near_miss_samples=N)` returns them by id — and never by the percentile alone.
That is not a style preference; it is the one case where the two disagreed and
was checked. `report_preference` logged 69 consecutive declines with the
refused score 0.0006 under the bar, and every percentile said "lower it".
Reading the refused record showed it was rule 77 "Extract intent from loose
phrasing", a false positive, so lowering the bar would have delivered that rule
on every completion report ever written. The statistic and the correct action
pointed in opposite directions, and only opening the record could tell.
"""
from __future__ import annotations
from scribe.services.settings import bounded_float, get_setting
from scribe.services.retrieval_pipeline import ( # noqa: F401 - Surface re-exported
TUNED_ARMS, Surface,
)
# A budget nobody should be able to set past. Not a tuning value — a guard on
# the worst case, so a mistyped setting cannot turn a menu into a wall of text.
# Shared by every surface because it bounds the same thing everywhere: how many
# lines a single unsolicited injection may occupy.
MAX_BUDGET = 10
# THE TABLE IS READ OFF THE SPECS (milestone 456 step 5). Each ranked arm in
# `retrieval_pipeline` carries its own `Surface` — the floor and budget pair,
# its settings keys and defaults, and the prose a tuner reads — so an arm added
# there is tunable by construction, and nothing here can drift from it. The
# keys and defaults are the ones this table held before the move, unchanged.
SURFACES: dict[str, Surface] = {
arm.tuning.name: arm.tuning for arm in TUNED_ARMS if arm.tuning is not None
}
# Reserved slots are deliberately absent. `preference_slot`, `reuse_slot` and
# `lesson_slot` borrow their parent arm's floor and are hard-limited to one hit
# each, because their entire purpose is to guarantee a single line to a kind of
# record that keeps losing a general score contest (#2246, #3894) — or, for
# `lesson_slot`, to a kind whose loss is total rather than merely a lost
# convenience, since a lesson has no act arm to fall back on and nobody browses
# lessons looking for one (milestone 385). A budget of "1" is the feature;
# exposing it as tunable would invite setting it to 0 and silently removing the
# guarantee.
#
# They are still logged under their own `source`, which is what keeps them
# judgeable without being tunable: `best_available_id` names the record each
# bar refused, so a slot that never places, or one that places weak hits, shows
# up as evidence rather than as an argument.
def surface_names() -> list[str]:
"""Every tunable surface, in a stable order for menus and listings."""
return list(SURFACES)
def get_surface(name: str) -> Surface:
"""Look one up, refusing an unknown name loudly.
A typo'd surface must not be writable. Settings keys are free-form strings
in a generic table, so a tuning call naming `pretool_rule` would otherwise
write a key nothing ever reads — a change that appears to succeed, reports a
new value, and alters nothing.
"""
try:
return SURFACES[name]
except KeyError:
raise ValueError(
f"unknown retrieval surface {name!r}. Tunable surfaces are: "
+ ", ".join(surface_names())
) from None
def dial_for_key(key: str) -> tuple[str, str] | None:
"""Which `(surface, dial)` a settings key belongs to, or None.
The registry read backwards, and it exists for one caller: the generic
`/api/settings` endpoint, which accepts any key at all. Without this, a
floor written through that endpoint moves with no event recorded, and the
tuning history says nothing happened — a trail with holes in it, which is
worse than no trail because it reads as complete.
Derived rather than listed so a seventh surface is covered the moment it is
added here, which is the only way this stays true.
"""
for surface in SURFACES.values():
if key == surface.floor_key:
return (surface.name, "floor")
if key == surface.budget_key:
return (surface.name, "budget")
return None
async def floor_for(user_id: int, name: str) -> float:
"""This install's current floor for a surface, clamped to [0, 1]."""
s = get_surface(name)
raw = await get_setting(user_id, s.floor_key, str(s.floor_default))
return bounded_float(raw, s.floor_default)
async def budget_for(user_id: int, name: str) -> int:
"""This install's current budget for a surface, clamped to [1, MAX_BUDGET].
The lower clamp is 1, never 0: a surface turned off is turned off by its
`enabled` switch, which says so. A budget of zero would be an arm that runs
a search, logs a retrieval, and renders nothing — indistinguishable in the
telemetry from a bar nothing cleared, which is the exact confusion this
milestone exists to remove.
"""
s = get_surface(name)
raw = await get_setting(user_id, s.budget_key, "")
if not raw and s.budget_falls_back_to:
raw = await get_setting(user_id, s.budget_falls_back_to, "")
try:
value = int(float(raw)) if raw else s.budget_default
except (TypeError, ValueError):
value = s.budget_default
return min(MAX_BUDGET, max(1, value))