CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m10s
CI & Build / Python tests (push) Successful in 1m54s
CI & Build / Build & push image (push) Successful in 31s
report_preference searched one constant string when a task closed and handed matches back as update_task's reply_preferences. Milestone 500 step 3 delivers the reply shapes and the preferences mounted beside them on the moments, so the arm goes, whole (rule 22): - services/reply_preferences.py, the REPORT_PREFERENCE RuleArm, its re-scorer and corpus entries, and the REPORTPREF constants - update_task's reply_preferences block and its cue; the docstring now points at moment_rules - the fixed_query field and the fixed_query_never_clears warning: this arm was the only one that set it, so the concept and the cannot_decline exemption go with it - the two Settings fields for its floor and budget - migration 0121 deletes its rows (logs, judgments, rule-usage events, tuning history, settings keys), operator-approved 2026-10-09. Without it the readout would call the source unregistered and its surfacings would count as ambient. The unindexed rule_usage delete is bounded by created_at. #5496 (step 4 of milestone 500). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
171 lines
7.9 KiB
Python
171 lines
7.9 KiB
Python
"""One registry of the retrieval surfaces, and the two numbers each one has (#4102).
|
|
|
|
WHY THIS EXISTS
|
|
|
|
Six push arms each carried their own loose copy of the same shape: a settings
|
|
key, a default, and a limit that was usually a module constant nobody could
|
|
change. The read-and-clamp was written out separately in `plugin_context` (twice
|
|
over, for auto-inject and the write path), in three rule arms, and again in
|
|
the completion-report arm (since retired, milestone 500). That was survivable
|
|
while the numbers were shipped constants an operator occasionally edited.
|
|
|
|
It stops being survivable once the numbers are meant to MOVE. The operator's
|
|
decision for this step:
|
|
|
|
"the floor should be chosen and adjusted by the model using it. we've come
|
|
back to something either fails or has to be looked at by the user we need a
|
|
model consistent surface for the adjustment of these floor values. the user
|
|
should be able to touch it but the model should be the thing handling it 9
|
|
times out of 10."
|
|
|
|
A tuning surface cannot be "model consistent" if every arm spells its own
|
|
configuration differently. So the arms stop owning their numbers and read them
|
|
from here instead, and the tuning tool, the routes and the Settings UI all
|
|
enumerate THIS table rather than hard-coding six special cases.
|
|
|
|
WHAT A FLOOR IS NOW, AND WHAT IT IS NOT
|
|
|
|
It is not a relevance judgement. Relevance is decided by the reader, which is
|
|
the only participant that can read a trigger against a situation — that is the
|
|
milestone's whole argument, and the injected line has always said so out loud
|
|
("read it before deciding it does not apply").
|
|
|
|
A floor answers the cheaper question: *is this worth ranking at all*. `k` is
|
|
what binds, and `k` is a BUDGET — how much of this surface's attention a
|
|
candidate list may spend. An arm that fires before every Bash call cannot
|
|
afford what an arm that fires once a turn can.
|
|
|
|
WHY THE DEFAULTS BELOW ARE STARTING POINTS AND NOT ANSWERS
|
|
|
|
A cosine score is a distance in BAAI/bge-small-en-v1.5's vector space, measured
|
|
against THIS corpus. It cannot transfer to an install with different records,
|
|
and rule 115 forbids defending a shipped default from this instance's
|
|
telemetry — which is what made the old design unbuildable: every number was a
|
|
guess everywhere except here, and there was no mechanism that could ever
|
|
improve it.
|
|
|
|
The mechanism is the fix. Scribe ships a starting point and the means to
|
|
correct it, so the values below carry the measurement that motivated them (see
|
|
the long comments in `plugin_context.py`, which are kept where they are because
|
|
they record how each number was first arrived at) without claiming to be right
|
|
for anybody else.
|
|
|
|
HOW A FLOOR SHOULD ACTUALLY BE MOVED
|
|
|
|
By reading the records the floor refused — `retrieval_telemetry(
|
|
near_miss_samples=N)` returns them by id — and never by the percentile alone.
|
|
That is not a style preference; it is the one case where the two disagreed and
|
|
was checked. `report_preference` logged 69 consecutive declines with the
|
|
refused score 0.0006 under the bar, and every percentile said "lower it".
|
|
Reading the refused record showed it was rule 77 "Extract intent from loose
|
|
phrasing", a false positive, so lowering the bar would have delivered that rule
|
|
on every completion report ever written. The statistic and the correct action
|
|
pointed in opposite directions, and only opening the record could tell.
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
from scribe.services.settings import bounded_float, get_setting
|
|
from scribe.services.retrieval_pipeline import ( # noqa: F401 - Surface re-exported
|
|
TUNED_ARMS, Surface,
|
|
)
|
|
|
|
# A budget nobody should be able to set past. Not a tuning value — a guard on
|
|
# the worst case, so a mistyped setting cannot turn a menu into a wall of text.
|
|
# Shared by every surface because it bounds the same thing everywhere: how many
|
|
# lines a single unsolicited injection may occupy.
|
|
MAX_BUDGET = 10
|
|
|
|
|
|
# THE TABLE IS READ OFF THE SPECS (milestone 456 step 5). Each ranked arm in
|
|
# `retrieval_pipeline` carries its own `Surface` — the floor and budget pair,
|
|
# its settings keys and defaults, and the prose a tuner reads — so an arm added
|
|
# there is tunable by construction, and nothing here can drift from it. The
|
|
# keys and defaults are the ones this table held before the move, unchanged.
|
|
SURFACES: dict[str, Surface] = {
|
|
arm.tuning.name: arm.tuning for arm in TUNED_ARMS if arm.tuning is not None
|
|
}
|
|
|
|
# Reserved slots are deliberately absent. `preference_slot`, `reuse_slot` and
|
|
# `lesson_slot` borrow their parent arm's floor and are hard-limited to one hit
|
|
# each, because their entire purpose is to guarantee a single line to a kind of
|
|
# record that keeps losing a general score contest (#2246, #3894) — or, for
|
|
# `lesson_slot`, to a kind whose loss is total rather than merely a lost
|
|
# convenience, since a lesson has no act arm to fall back on and nobody browses
|
|
# lessons looking for one (milestone 385). A budget of "1" is the feature;
|
|
# exposing it as tunable would invite setting it to 0 and silently removing the
|
|
# guarantee.
|
|
#
|
|
# They are still logged under their own `source`, which is what keeps them
|
|
# judgeable without being tunable: `best_available_id` names the record each
|
|
# bar refused, so a slot that never places, or one that places weak hits, shows
|
|
# up as evidence rather than as an argument.
|
|
|
|
|
|
def surface_names() -> list[str]:
|
|
"""Every tunable surface, in a stable order for menus and listings."""
|
|
return list(SURFACES)
|
|
|
|
|
|
def get_surface(name: str) -> Surface:
|
|
"""Look one up, refusing an unknown name loudly.
|
|
|
|
A typo'd surface must not be writable. Settings keys are free-form strings
|
|
in a generic table, so a tuning call naming `pretool_rule` would otherwise
|
|
write a key nothing ever reads — a change that appears to succeed, reports a
|
|
new value, and alters nothing.
|
|
"""
|
|
try:
|
|
return SURFACES[name]
|
|
except KeyError:
|
|
raise ValueError(
|
|
f"unknown retrieval surface {name!r}. Tunable surfaces are: "
|
|
+ ", ".join(surface_names())
|
|
) from None
|
|
|
|
|
|
def dial_for_key(key: str) -> tuple[str, str] | None:
|
|
"""Which `(surface, dial)` a settings key belongs to, or None.
|
|
|
|
The registry read backwards, and it exists for one caller: the generic
|
|
`/api/settings` endpoint, which accepts any key at all. Without this, a
|
|
floor written through that endpoint moves with no event recorded, and the
|
|
tuning history says nothing happened — a trail with holes in it, which is
|
|
worse than no trail because it reads as complete.
|
|
|
|
Derived rather than listed so a seventh surface is covered the moment it is
|
|
added here, which is the only way this stays true.
|
|
"""
|
|
for surface in SURFACES.values():
|
|
if key == surface.floor_key:
|
|
return (surface.name, "floor")
|
|
if key == surface.budget_key:
|
|
return (surface.name, "budget")
|
|
return None
|
|
|
|
|
|
async def floor_for(user_id: int, name: str) -> float:
|
|
"""This install's current floor for a surface, clamped to [0, 1]."""
|
|
s = get_surface(name)
|
|
raw = await get_setting(user_id, s.floor_key, str(s.floor_default))
|
|
return bounded_float(raw, s.floor_default)
|
|
|
|
|
|
async def budget_for(user_id: int, name: str) -> int:
|
|
"""This install's current budget for a surface, clamped to [1, MAX_BUDGET].
|
|
|
|
The lower clamp is 1, never 0: a surface turned off is turned off by its
|
|
`enabled` switch, which says so. A budget of zero would be an arm that runs
|
|
a search, logs a retrieval, and renders nothing — indistinguishable in the
|
|
telemetry from a bar nothing cleared, which is the exact confusion this
|
|
milestone exists to remove.
|
|
"""
|
|
s = get_surface(name)
|
|
raw = await get_setting(user_id, s.budget_key, "")
|
|
if not raw and s.budget_falls_back_to:
|
|
raw = await get_setting(user_id, s.budget_falls_back_to, "")
|
|
try:
|
|
value = int(float(raw)) if raw else s.budget_default
|
|
except (TypeError, ValueError):
|
|
value = s.budget_default
|
|
return min(MAX_BUDGET, max(1, value))
|