Files
FabledScribe/src/scribe/services/retrieval_pipeline.py
T
bvandeusenandClaude Opus 5.5 869046dda2
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 2m1s
CI & Build / Build & push image (push) Successful in 27s
refactor(retrieval): the specs are the registry - SURFACES, the ranked POINTS rows and RANKED_SOURCES are read off the pipeline specs (milestone 456 step 5, #4907)
Each arm spec now carries its tuning pair (Surface, moved verbatim into
retrieval_pipeline) and a Declared block - what the telemetry readout
must know and cannot read off its rows. The two ranked stages that are
not arms (preference_slot, rule_via_lesson) are RankedSource specs, and
the note slots carry their own declaration.

- retrieval_surfaces.SURFACES = the TUNED_ARMS tuning, same order
- retrieval_registry.POINTS ranked rows = one per spec in RANKED; the
  lookups, asked, ambient and pull rows stay declared there
- rule_usage.RANKED_SOURCES = RULE_RANKED_SOURCES + moment_rule

Settings keys, defaults, prose and order are unchanged (checked field by
field against HEAD). tests/test_retrieval_specs.py pins the derivation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 20:11:20 -04:00

1691 lines
76 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""One retrieval pipeline: a ranked surface is a spec, not a hand-copied arm.
Milestone 456. Every ranked rule surface used to write out the same sequence
by hand — search, band, split fresh from repeats, log the call before any early
return, reserve a slot, render, record what was surfaced — once per arm, each
modelled on the nearest sibling. The copies drifted, and the defects that
mattered were fixed one copy at a time (#3497, #3750, #3752): "this arm shipped
with the same defect inherited from its sibling" was a comment in the code.
The invariants were held together by a parametrised test over three copies
rather than by there being one implementation.
Here the sequence is written ONCE (`run_rule_arm`), and an arm is a `RuleArm`:
which source it records under, and which of the stages it takes. The
differences between arms that used to be buried in their copies are now four
booleans on one screen — and the ones that are accidents rather than
decisions are milestone 456 step 7's to settle, one at a time, on evidence.
THE INVARIANTS, each written once:
- the call row is written before any early return (#3497), because a call
that found nothing is the only evidence a bar is too high;
- the call row's results and the surfacing rows are both the FRESH cut, so
the two tables describe the same delivery (#3668, #3752);
- the session ledger decides how a repeat RENDERS, never whether it is
retrieved (#3750, #4101);
- band before dedup (#3851), so "you were shown this" never changes what
counts as relevant;
- telemetry never costs the reader a line, and the arm as a whole fails
open: a recall aid may never break the prompt, the command or the write.
THE I/O IS PASSED IN (`RuleIO`), not imported. The pipeline is the policy;
which ranker it asks and which recorders it reports to belong to the caller.
The callers resolve them from their own module at call time.
"""
from __future__ import annotations
import logging
import time
from dataclasses import dataclass, field
from typing import Any, Callable
from scribe.services.lessons import LESSON_NOTE_TYPE, claim_line
logger = logging.getLogger(__name__)
# ── Shared rendering and ranking helpers (moved from plugin_context) ──────
# MEASURED, NOT REASONED — and the reasoning it replaced was wrong (#3851).
#
# The prediction was that rules would rank SHARPLY, because `rule_document()`
# shapes them the way snippets are shaped — trigger in the embedded title and
# again above the body — and note 2485 measured snippets separating their top
# hit by 0.153 while every other kind managed 0.010–0.023.
#
# They do not. Three probes against real act queries, scored by the same
# embedding the arms use:
#
# `git push origin dev` top 0.757, gap to second 0.022
# `docker compose up -d` top 0.685, gap to second 0.016
# a bare-owner-filter query top 0.656, gap to second 0.020
#
# That is dev-log territory, not snippet territory: rules arrive as a
# tightly-packed block. Shaping alone did not buy separation, which is worth
# recording because the opposite was the natural inference from 2485.
#
# So the band is narrow BECAUSE the corpus is flat. At 0.10 — the notes
# menu's value — every one of the top eight on the push probe falls inside,
# including a CI-registry rule and another project's branch policy. At 0.05
# it admits roughly three ranks, which is the span where the scores are still
# saying something. 2485's Finding 3 is the standing caveat: no band value
# fixes a tie, and if rules ever rank as flat as dev-logs did this control
# stops working and the answer is a reranker (#1038), not a smaller number.
_RULEHINT_BAND = 0.05
def _rule_band(hits: list) -> list:
"""The top hit, plus every hit within `_RULEHINT_BAND` of it (#3851).
The instrument that lets an act surface a SET without inventing one: a
fixed k fills its slots whether or not anything deserves them, while this
keeps only what the scores say is close, so one clearly-relevant rule
still shows one and four competing rules show four.
Sync and pure, and deliberately its own function rather than a comparison
written twice — the two act arms are the pair #3497 records drifting apart
by being modelled on each other instead of sharing.
Takes `(score, rule)` pairs already ordered best-first, as both
`semantic_search_*` helpers return them.
"""
if not hits:
return []
top = hits[0][0]
return [(s, r) for s, r in hits if s >= top - _RULEHINT_BAND]
def checkpoint_for(
kept: list, *, held: set[int], floor: float, where: str,
) -> dict:
"""The one rule, if any, that should STOP this act rather than annotate it.
Sync and pure so it can be read against a fixed list of hits without a
database — the two arms share it for the reason `_rule_band` is shared:
#3497 is the record of these two drifting apart by being modelled on each
other instead of sharing one function.
FOUR CONDITIONS, AND EACH IS A DIFFERENT KIND OF WRONG IT PREVENTS.
1. THE SCORE CLEARS `floor`. Not the arm's own floor — a much higher bar,
measured at `_CHECKPOINT_DEFAULT`. The hint arms keep nudging at their
floor; only a hit the corpus is confident about is allowed to stop
anything.
2. IT IS A RULE, NEVER A PREFERENCE. A preference says how something has
been done before and following it is what keeps work consistent; a rule
says what happens if you do not. Stopping an act over a preference
would assert a force the record explicitly does not claim, and
`_rule_hint_line` already keeps that distinction in the one word that
names it.
3. THE SESSION HAS NOT OPENED IT. `held` is observable — a PostToolUse
hook watches for the `get_rule` call (#4100) — so this is a recorded
event and not a model's self-report about its own context. A session
that read the rule has already had the thing the checkpoint exists to
produce, and stopping it again would be punishing the behaviour being
asked for.
4. IT IS THE TOP HIT. `kept` is a band, and a band's tail is there to let
an act surface a SET; the ranker's confidence claim attaches to its
first element only. A checkpoint raised on the fourth line of a band is
a stop justified by a score nobody claimed.
Returns a dict rather than a rule or a tuple. Widening a tuple is an
interface change to every unpack site that the compiler does not report
(#4207, learned the expensive way in this same milestone's first week), and
this value crosses a JSON boundary into a shell script where a missing
field is a silently empty variable.
"""
if not kept or floor <= 0:
return {}
score, rule = kept[0]
if score < floor:
return {}
if getattr(rule, "kind", "") == "preference":
return {}
if rule.id in held:
return {}
found = {
"rule_id": rule.id,
"title": rule.title,
"trigger": (rule.when_to_apply or "").strip(),
"score": round(float(score), 4),
"where": where,
}
# RENDERED HERE, at the one call site, rather than by each arm. Two arms
# that each remember to render it are two arms that can stop agreeing on
# what a stop says — which is #3497's history for this exact pair. The
# text stays a separate function so it can be read and tested without a
# rule object, but nothing outside this line decides whether to call it.
found["reason"] = checkpoint_reason(found)
return found
def checkpoint_reason(checkpoint: dict) -> str:
"""The text the agent reads INSTEAD of running the act.
Written as a practice rather than a prohibition (rule 165): it says what
to do and why it is worth doing, not what is forbidden. The act is not
wrong — nothing here knows whether it is — and saying so plainly is what
keeps the stop from reading as an accusation the system is in no position
to make.
It names the remedy as ONE call, because a stop whose remedy is vague
costs more than the miss it prevents. And it says the act may simply be
re-submitted afterwards, so a reader who finds the rule irrelevant is out
in two calls rather than negotiating with a hook.
"""
if not checkpoint:
return ""
trigger = checkpoint.get("trigger") or ""
return (
f"Held for one read. “{checkpoint['title']}” is a standing rule "
f"this session has not opened, and it scores {checkpoint['score']} "
f"against what you are about to do"
+ (f" ({trigger})" if trigger else "")
+ f". Read it with get_rule({checkpoint['rule_id']}), then go ahead — "
f"re-submit this call unchanged if the rule does not apply, which is a "
f"judgement only you can make. Nothing here has decided the act is "
f"wrong; the rule is being put in front of it rather than beside it, "
f"because a rule delivered alongside a result arrives after the "
f"decision it was meant to inform."
)
def _rule_hint_line(
rule, *, where: str, seen: bool, held: bool = False, compact: bool = False,
) -> str:
"""One rule hint line — both arms, both tails, both kinds (#3750, #3849).
THREE INDEPENDENT AXES SINCE #3851. `compact` joins `kind` and `seen`, and
like them it reads none of the others: it says how much ROOM this line
gets, which is a fact about its rank among today's hits rather than about
the rule. A compact line is still a full claim that the rule may apply —
it simply cites the rule instead of quoting its trigger. Crucially it
still carries the `seen` tail, so the three axes stay genuinely
independent: shortening a line must not decide what it says about whether
the session is holding the rule.
WHY THE LATER LINES ARE QUIETER. Measured at #3851: a full line runs ~143
tokens once the trigger is rendered, and #3855 tripled trigger lengths
across the corpus, so five full lines cost ~568 tokens before every Bash
call. Top-full-plus-references costs ~299 — about 2x the old single line,
for four more rules. The budget argument that once justified a single slot
was real; what it actually forbids is four voices at full volume, not four
voices.
The top hit keeps the full rendering because it is the one the ranker is
most confident about, and a reader who acts on exactly one line should
have acted on that one.
TWO INDEPENDENT AXES. `kind` decides the head, `seen` decides the tail,
and neither reads the other. A preference and a rule differ in force; a
repeat and a first surfacing differ in whether the session already holds
the line. Those are unrelated facts, and keeping them unrelated in the
code is what stopped the second kind from reopening the repeat question.
ONE FUNCTION BECAUSE THE TAILS MUST NOT DRIFT. The two arms phrase their
heads differently ("may apply here" vs "may apply to this Bash call") and
that difference is deliberate. Everything after it must not differ, and
#3497's history is that the pre-tool arm inherited a defect from its
sibling by being modelled on it rather than sharing with it. Two copies of
a two-branch string is how one branch gets fixed and the other does not.
WHY A REPEAT GETS A LINE AT ALL. Both arms used to drop a hit whose id was
already on the session's exclusion ledger and emit nothing. That is correct
only while the session still HOLDS what it was told, and a compaction
breaks exactly that: the earlier injection is summarized away while the id
stays on the ledger, so the rule is absent from context AND unreachable for
the rest of the session (#3749 closes the compaction half; this closes the
ordinary half, where a session simply stops holding a line it read an hour
ago).
Only ONE CLAUSE of the original line is false on a repeat — the claim that
the rule is not in the session's loaded set. So only that clause changes.
Title, trigger and pull pointer are identical either way, the statement is
never injected either way, and a repeat therefore costs the same ~40 tokens
as a first surfacing and no more.
DELIBERATELY NOT ASKING THE SESSION WHETHER IT HOLDS THE RULE. A model
asked "do you still hold rule 156?" will say yes, and the claim is
unverifiable self-report about its own context. The answer is also not
needed: the line is cheap enough to always emit and carries its own remedy
in both branches. Removing the question removes the fragility rather than
managing it.
"""
trigger = (rule.when_to_apply or "").strip()
preference = rule.kind == "preference"
# KIND CHANGES THE HEAD; `seen` CHANGES THE TAIL. The two axes are
# independent and stay that way, which is what lets the repeat logic above
# survive a second kind without being reasoned about again: whether a
# record is already on the ledger has nothing to do with how much force it
# carries, so the seen-branch is shared verbatim.
#
# The noun is the whole of the visual difference, and that is deliberate.
# A reader skimming an injected block gets one word to place the register \u2014
# so it is the SECOND word that moves, and it is the word naming force.
# Everything structural after it is identical, so the three kinds read as
# one set rather than three formats (milestone 385's step 5 writes the
# lesson voice against these two; they are one paragraph, not three).
noun = "Preference" if preference else "Standing rule"
# The only other place force is asserted. A rule's line tells the reader
# not to dismiss it unread, because dismissing a rule unread is how the
# thing it prevents happens. A preference makes no such claim: it says
# where to find how this has been done, and following it is what keeps
# things consistent rather than what keeps them correct.
reason = (
"for how this has been done before" if preference
else "before deciding it does not apply"
)
# THREE STATES, BECAUSE TWO OF THEM WERE BEING TOLD THE SAME LIE (#4100).
#
# `seen` means an arm NAMED this rule earlier. It does not mean the session
# read it — the line is a teaser, and a teaser skimmed past leaves nothing
# behind, least of all after a compaction summarises the turn it arrived
# in. "You saw it earlier this session" asserted something about the
# reader's context that the server had no way to know.
#
# `held` is the observable half: a PostToolUse hook watches for the
# `get_rule` call itself, so this is a recorded EVENT rather than a claim.
# That distinction is what keeps the non-goal above intact — the objection
# was to asking a model about its own context, not to noticing what it did.
#
# The middle state is the honest one and the one that was missing: named,
# not opened. It gets the full invitation, because a session that skipped
# the teaser is in almost the same position as one that never saw it.
if held:
tail = (
f"You opened it earlier this session; pull it with "
f"get_rule({rule.id}) again if you no longer hold it."
)
elif seen:
tail = (
f"Mentioned earlier this session but not opened — read it with "
f"get_rule({rule.id}) {reason}."
)
else:
tail = (
f"Read it with get_rule({rule.id}) {reason}; it is not in this "
"session's loaded set."
)
if compact:
# THE TRIGGER GOES; THE TAIL STAYS. Only one of the two is expensive \u2014
# a trigger runs 300-400 characters after #3855, the tail about 100 \u2014
# so dropping the trigger is nearly the whole saving and dropping the
# tail would be mostly sacrifice.
#
# It would also destroy the one thing #3750 exists to say. The tail is
# what tells a reader whether they were already told this and may no
# longer be holding it, and that is the entire difference they can act
# on; a reference with no tail reads as a first surfacing whether it is
# one or not. This branch shipped without it for one commit and
# test_a_rule_the_session_already_holds_is_referenced_not_re_offered
# caught it, which is the guard working exactly as #3750 intended.
return (
f"Also \u2014 {noun.lower()} \u201c{rule.title}\u201d. {tail}"
)
return (
f"{noun} that may apply {where} \u2014 \u201c{rule.title}\u201d"
+ (f" ({trigger})" if trigger else "")
+ f". {tail}"
)
# ── What a spec declares about itself (milestone 456 step 5) ─────────────
#
# A ranked surface used to be described in three hand-kept lists besides its
# code: its tuning pair in `retrieval_surfaces.SURFACES`, its row in
# `retrieval_registry.POINTS`, and — for a rule surface — its membership in
# `rule_usage.RANKED_SOURCES`. Tests existed to make the four agree. Now the
# spec carries all of it, and the three lists are read off the specs, so an
# arm added here is tunable, measured and counted by construction.
#
# THE CLASSES LIVE HERE, NOT BESIDE THE LISTS, for the import graph: the
# registry, the tuning table and the usage counter all read the specs, so the
# specs cannot import any of them back.
@dataclass(frozen=True)
class Surface:
"""One push arm's tunable pair, plus enough prose to tune it responsibly.
`asks` / `over` / `fires` are not documentation for this file — they are
rendered by the tuning tool and the Settings UI. A floor cannot be moved
sensibly by anyone, model or human, who does not know what the query is, what
corpus it runs against, or how often it costs something. Those three facts
are exactly what separates these arms from each other, and they were
previously recoverable only by reading `plugin_context.py`.
"""
name: str
"""The telemetry `source` value, and the join key.
MUST equal the string this arm passes to `record_retrieval`. Everything
useful about tuning depends on that identity: the tool that moves a floor
and the table that says what the floor did have to be talking about the same
arm. A test asserts it rather than a comment asking nicely.
"""
floor_key: str
floor_default: float
budget_key: str
budget_default: int
asks: str
over: str
fires: str
measured_model: str = "BAAI/bge-small-en-v1.5"
measured_shape: int = 1
"""What the SHIPPED defaults above were measured against (#4104).
A floor is a distance in one embedding model's geometry, over documents cut
one particular way. Either can change, and when one does every number in
this table describes something that no longer exists.
TWO FIELDS, NEVER ONE FUSED STRING (rule 149). A mismatch has to be able to
say WHICH half moved: a new embedding model and a re-cut document shape
invalidate the same numbers for different reasons and call for different
responses. `"<model>@<n>"` could only report that something changed, which
is the answer nobody can act on. Same reason `calibration_stamp()` returns
a dict and the event table gives each half its own column.
Recorded per surface rather than once for the module because they need not
move together: a surface retuned after a model change carries the new stamp
while its untouched siblings still carry the old one, and telling those
apart is the whole job.
LITERALS, deliberately, rather than an import of the live values — a stamp
says what was true when the number was chosen, so one that tracked the
current model would always agree with it and could never report staleness.
"""
budget_falls_back_to: str = ""
"""A budget key to inherit when this surface has none of its own set.
Only `write_path` uses it, and only because it USED to share auto-inject's
`top_k` outright. Giving it a key without this would silently reset the
budget of every install that had tuned the shared one — a behaviour change
delivered as a default, which is the shape of regression nobody reports
because nothing looks broken.
"""
@dataclass(frozen=True)
class Declared:
"""What the telemetry readout must know about a ranked source and cannot
read off its rows — `retrieval_registry.Point`'s fields, for the specs
that produce one. Every ranked source is UNBIDDEN: nobody asked for it."""
what: str
"""One line, for an agent reading a warning that names this source."""
quiet_because: str = ""
"""Set when silence over an active window is correct, saying why (#2475)."""
fixed_query: bool = False
"""The arm always searches the same string, so its decline rate is 0% or
100% and `cannot_decline` says nothing about it."""
@dataclass(frozen=True)
class RankedSource:
"""A ranked source that is a stage of an arm rather than an arm: it
records under its own name, and is measured but never tuned."""
source: str
declared: Declared
# ── The specs ────────────────────────────────────────────────────────────
@dataclass(frozen=True)
class RuleArm:
"""One ranked rule surface: where it records, and which stages it takes."""
source: str
"""The telemetry `source`. MUST equal the registry and SURFACES key."""
band: bool
"""Keep only the hits within `_RULEHINT_BAND` of the top one (#3851)."""
compact_tail: bool
"""Lines after the first cite the rule instead of quoting its trigger."""
checkpoint: bool
"""The top hit may hold the act for one read (`checkpoint_for`)."""
preference_slot: bool
"""Reserve one line for a preference that lost the ranking (#3894)."""
kind: str | None = None
"""Ask the ranker for one record kind only ("preference"), or every kind."""
stop_only: bool = False
"""The door shows nothing but the stop (milestone 458's reply moment).
At the end of a turn there is no context to annotate: a hit either holds
the reply for one read or reaches nobody. So only the checkpoint's rule is
recorded as surfaced — the rest were ranked, and the call row counts them,
but nobody was shown them."""
tuning: Surface | None = None
"""Its floor and budget, and the prose a tuner reads. Every arm has one;
`SURFACES` is read off them."""
declared: Declared | None = None
# Today's differences, reproduced exactly (milestone 456 step 2). Whether the
# prompt arm should band, and whether the act arms should reserve a
# preference, are step 7's questions — not settled by this refactor.
PROMPT_RULE = RuleArm(
"prompt_rule", band=False, compact_tail=False, checkpoint=False,
preference_slot=True,
tuning=Surface(
name="prompt_rule",
floor_key="kb_promptrule_threshold",
floor_default=0.72,
budget_key="kb_promptrule_top_k",
budget_default=3,
asks="the operator's message, against rule triggers",
over="global rules plus the bound project's own",
fires="once per operator turn",
),
declared=Declared("rules that may govern what the operator just asked"),
)
PRE_TOOL_RULE = RuleArm(
"pre_tool_rule", band=True, compact_tail=True, checkpoint=True,
preference_slot=False,
tuning=Surface(
name="pre_tool_rule",
floor_key="kb_toolrule_threshold",
floor_default=0.68,
budget_key="kb_toolrule_top_k",
budget_default=5,
asks="the command about to run, against rule triggers",
over="global rules plus the bound project's own",
fires="before every Bash call — the busiest arm there is",
),
declared=Declared("rules that may govern a command about to run"),
)
WRITE_PATH_RULE = RuleArm(
"write_path_rule", band=True, compact_tail=True, checkpoint=True,
preference_slot=False,
tuning=Surface(
name="write_path_rule",
floor_key="kb_rulehint_threshold",
floor_default=0.72,
budget_key="kb_rulehint_top_k",
budget_default=5,
asks="the code being written, against rule triggers",
over="global rules plus the bound project's own",
fires="before every Write and Edit",
),
declared=Declared("rules that may govern the file being written"),
)
# The completion report's preferences (milestone 409 step 4): a FIXED query,
# preferences only, read by update_task as records rather than as lines. Its
# query is `reply_preferences.COMPLETION_QUERY`.
REPORT_PREFERENCE = RuleArm(
"report_preference", band=False, compact_tail=False, checkpoint=False,
preference_slot=False, kind="preference",
tuning=Surface(
name="report_preference",
floor_key="kb_reportpref_threshold",
floor_default=0.72,
budget_key="kb_reportpref_top_k",
budget_default=3,
# THE ONE FIXED QUERY, and the reason this arm behaves unlike the rest.
# The others score something that varies per call; this one scores a
# constant string, so its top score for a given corpus is also a
# constant. A floor a hair above that constant is not a quiet arm, it
# is a dead one, and no amount of traffic will ever reveal it — which
# is precisely how this arm spent 69 calls declining the same record.
asks="a fixed question about how to lay out a completion report",
over="preferences",
fires="when a task finishes",
),
# `fixed_query`: COMPLETION_QUERY is a module constant, so this arm's top
# score is the same number on every call — measured at 0.791 across 45
# consecutive calls, with p10, p50, p90, min and max all identical. Five
# equal percentiles is the tell.
declared=Declared(
"the fixed question asked when a task finishes: how should this report read",
fixed_query=True,
),
)
# The backstop for every arm that ran earlier in the turn and missed: the
# finished reply against every rule's trigger. Its floor IS its stop bar
# (the reply_rule surface), so it is passed as both.
REPLY_RULE = RuleArm(
"reply_rule", band=False, compact_tail=False, checkpoint=True,
preference_slot=False, stop_only=True,
# THE REPLY BACKSTOP (milestone 458, folded in from 456 step 8). Its floor
# is a STOP bar, not a hint bar: at the end of a turn nothing can be shown
# beside the reply, so a hit either holds the reply for one read or says
# nothing. Hence a default at the checkpoint's level and a budget of one —
# the call row's results are then exactly the rule that would hold.
tuning=Surface(
name="reply_rule",
floor_key="kb_replyrule_threshold",
floor_default=0.80,
budget_key="kb_replyrule_top_k",
budget_default=1,
asks="the reply that ends a turn, against rule triggers",
over="global rules plus the bound project's own",
fires="once per turn, when the reply is finished",
),
declared=Declared(
"a rule that holds the finished reply for one read — the backstop "
"for whatever the earlier arms missed"
),
)
RULE_ARMS: tuple[RuleArm, ...] = (
WRITE_PATH_RULE, PRE_TOOL_RULE, PROMPT_RULE, REPORT_PREFERENCE, REPLY_RULE,
)
PREFERENCE_SLOT_SOURCE = "preference_slot"
PREFERENCE_SLOT = RankedSource(
PREFERENCE_SLOT_SOURCE,
Declared("the one line reserved for a preference at the prompt boundary"),
)
# Every source the shared stages below can record under — what the registry
# declares for this module's fan-out sites, since `source` reaches the
# recorders here as a value rather than a literal.
RULE_SOURCES: tuple[str, ...] = (
*(arm.source for arm in RULE_ARMS), PREFERENCE_SLOT_SOURCE,
)
@dataclass(frozen=True)
class RuleIO:
"""The ranker and the two recorders an arm reports to."""
search: Callable[..., Any]
"""`semantic_search_rules`, or a stand-in with its signature."""
record_retrieval: Callable[..., Any]
record_rule_surfaced: Callable[..., Any]
@dataclass(frozen=True)
class RuleMoment:
"""What is happening: the query an arm searches with, and the session."""
user_id: int
query: str
project_id: int | None
"""The bound project, 0 or None when unbound. Logged as given, searched as
`project_id or None` — global rules plus that project's own (milestone 414)."""
where: str = ""
"""How a line names the moment: "here", "to this Bash call", …"""
checkpoint_where: str = ""
exclude: frozenset[int] = frozenset()
"""Rules the session has been shown — rendered as repeats, not hidden."""
held: frozenset[int] = frozenset()
"""Rules the session has OPENED (#4100)."""
@dataclass
class RuleResult:
"""What an arm put in front of the reader, and what it counted."""
lines: list[str] = field(default_factory=list)
rule_ids: list[int] = field(default_factory=list)
"""The FRESH cut — telemetry's denominator, never a repeat."""
shown_rule_ids: list[int] = field(default_factory=list)
"""Every rule LINE, repeats and the reserved slot included."""
shown: list = field(default_factory=list)
"""The (score, rule) pairs behind those lines, best first — for a caller
that hands back records rather than lines (the completion report)."""
checkpoint: dict = field(default_factory=dict)
# ── The stages ───────────────────────────────────────────────────────────
def _telemetry_failed(source: str) -> None:
"""A recorder raised. Telemetry observes the surface and must never take it
down — or take a rendered line away from the reader, which an exception
reaching the arm's own fail-open would do. So it costs only its row.
The recorders are called DIRECTLY at each site, inside their own try,
rather than through a wrapper taking the recorder as an argument: the
registry test finds recorder call sites by their name, and a wrapper would
hide every site in this module from it (#3191's narrowing).
"""
logger.debug("retrieval telemetry failed for %s", source, exc_info=True)
async def _ranked(
io: RuleIO, moment: RuleMoment, *, source: str, floor: float, limit: int,
band: bool = False, kind: str | None = None,
) -> tuple[list, list, list]:
"""Search, band, split fresh from repeats, and log the call — every call.
Returns (hits, kept, fresh): what the ranker returned, what the band kept,
and the kept hits this session has not been shown. The row is written
HERE, before any caller can return early (#3497), with `results` the fresh
cut and `suppressed` counting both the band and the ledger, so a long
session's all-repeat zero reads apart from a bar that turned everything
away.
"""
t0 = time.perf_counter()
report: dict = {}
kwargs: dict = {"limit": limit, "threshold": floor, "report": report,
"project_id": moment.project_id or None}
if kind:
# KIND-FILTERED, so the query can only answer with what was asked
# for. Verifying the kind afterwards would be weaker: an unfiltered
# search that happened to return a rule would spend a preference's
# slot on it, indistinguishable from a line that earned its place.
kwargs["kind"] = kind
hits = await io.search(moment.user_id, moment.query, **kwargs)
if kind:
# And checked on the way out: a record of another kind that slipped
# through must not be logged, shown or counted under a source whose
# name claims the kind — a rule handed back as "how the operator likes
# this done" asserts a force the record does not have.
hits = [(score, rule) for score, rule in hits
if getattr(rule, "kind", None) == kind]
kept = _rule_band(hits) if band else hits
fresh = [(score, rule) for score, rule in kept if rule.id not in moment.exclude]
try:
io.record_retrieval(
user_id=moment.user_id, source=source, query=moment.query,
threshold=floor, limit=limit, project_id=moment.project_id,
is_task=None, results=fresh,
duration_ms=(time.perf_counter() - t0) * 1000.0,
best_available=report.get("best_available_score"),
best_available_id=report.get("best_available_id"),
searched=bool(report.get("searched", True)),
suppressed=len(hits) - len(fresh),
)
except Exception: # noqa: BLE001 - observation never breaks the observed
_telemetry_failed(source)
return hits, kept, fresh
async def _reserve_slot_for_preference(
io: RuleIO, moment: RuleMoment, hits: list, *, floor: float,
) -> tuple[list, int | None]:
"""Guarantee a preference one slot, if one clears the bar (#3894).
THE ASYMMETRY THIS EXISTS FOR. A rule and a preference are not equally
served by a shared score contest, because their losses are not equal:
- a RULE crowded out here can still fire at the act arm. A `git push`
reaches `pre_tool_rule`, a file write reaches `write_path_rule`. The
prompt hit is a preview of a second chance.
- a PREFERENCE about how to answer has no second chance. There is no
later act — the response IS the act — so crowded out here it is never
delivered at all.
A straight ranking therefore favours the record whose loss is recoverable
over the one whose loss is total, and it does so INVISIBLY: the rule that
won is a legitimate hit, the telemetry looks healthy, and the only symptom
is a preference that quietly never arrives. `reuse_slot` exists for the
same shape one corpus over (#2463), where snippets kept losing to project
records that merely resembled the query.
THE SLOT BUYS POSITION, NOT A LOWER BAR — same as `reuse_slot`, which also
reserves at `cfg["threshold"]`. A weak preference cannot buy the slot, so
silence stays the default and the reserved line is never worse than the
ones it sits beside. If `preference_slot` later shows a stream of
near-misses, `best_available_id` (#3807) names which preference was
refused and a separate bar becomes an argument with evidence behind it
rather than a knob added on a guess.
LEDGER REPEATS STILL COUNT AS REPRESENTED. A preference already on the
session's ledger occupies the slot rather than being skipped for a fresh
one: it is still rendered (#3750), just with the tail that says so, and a
preference is the kind of record where being reminded is the point.
Returns the possibly-extended hit list, and the id the slot spent — the
caller needs that to keep each source's surfaced set matching its own log
row (#3668), since the slot logs under its own name.
"""
if any(rule.kind == "preference" for _s, rule in hits):
return hits, None
# ITS OWN SOURCE, and both sides of the trade logged. #2463's own finding
# is the warning rather than the precedent here: the hit that slot pushed
# OUT was in retrieval_logs while the query that pushed it out was not, so
# the slot could never be judged against what it displaced.
found, _kept, _fresh = await _ranked(
io, moment, source=PREFERENCE_SLOT_SOURCE, floor=floor, limit=1,
kind="preference",
)
seen = {rule.id for _s, rule in hits}
slot = [(s, r) for s, r in found
if r.kind == "preference" and r.id not in seen][:1]
if not slot:
return hits, None
slot_id = int(slot[0][1].id)
if slot_id not in moment.exclude:
try:
io.record_rule_surfaced(
user_id=moment.user_id, rule_ids=[slot_id],
source=PREFERENCE_SLOT_SOURCE,
)
except Exception: # noqa: BLE001 - observation never breaks the observed
_telemetry_failed(PREFERENCE_SLOT_SOURCE)
# IT EXTENDS, IT NEVER DISPLACES — and here it parts company with
# `reuse_slot`, which evicts its menu's weakest hit. The reason is the
# ledger rather than taste. A displaced hit was RETURNED by the general
# search and is sitting in that call's `retrieval_logs` row, but would not
# have been shown — so `prompt_rule`'s surfaced set would stop matching
# its own log row, and #3668's identity would break for a reason nothing
# in the data explains. That identity is the cheapest true statement
# available about this pair of tables, and milestone #379 is what it costs
# to lose it: five steps planned against a gap that was two counters
# disagreeing, not a write path dropping rows.
#
# The price is one extra line, only when the general search already filled
# the limit AND a preference cleared the bar without placing. Cheap, and
# it buys a surface whose two tables can always be checked against each
# other.
return hits + slot, slot_id
# ── The pipeline ─────────────────────────────────────────────────────────
async def run_rule_arm(
arm: RuleArm, moment: RuleMoment, *, floor: float, budget: int, io: RuleIO,
checkpoint_floor: float = 0.0,
) -> RuleResult:
"""Run one rule arm through every stage it takes, in the one order.
search → band → fresh/repeat split → call row → reserved slot → render →
surfacing rows → checkpoint. Fails open to an empty result.
"""
try:
_hits, kept, fresh = await _ranked(
io, moment, source=arm.source, floor=floor, limit=budget,
band=arm.band, kind=arm.kind,
)
shown = kept
if arm.preference_slot:
# BEFORE the bail-out, and the ordering is load-bearing. An empty
# general result is not proof that no preference qualifies: the
# general search overfetches by distance and then collapses, so a
# preference ranked below that window is invisible to it while a
# kind-filtered query finds it at once.
shown, _slot_id = await _reserve_slot_for_preference(
io, moment, shown, floor=floor,
)
# `shown`, not `fresh` (#3750): a call whose only hit is a repeat
# still has something to say, it just says it differently.
if not shown:
return RuleResult()
lines = [
_rule_hint_line(
rule, where=moment.where,
seen=rule.id in moment.exclude,
held=rule.id in moment.held,
# Rank decides volume (#3851): the ranker's best guess gets
# the trigger, the rest get cited.
compact=arm.compact_tail and idx > 0,
)
for idx, (_score, rule) in enumerate(shown)
]
# FRESH-ONLY (#3752): a repeat is a rendering decision, not a
# retrieval outcome, and counting it would inflate the denominator
# pull_through is read from. It is also exactly what this call's own
# row recorded, so the two tables stay equal (#3668). RANKED, not
# ambient: the source is in `rule_usage.RANKED_SOURCES`.
rule_ids = [rule.id for _score, rule in fresh]
source = arm.source
checkpoint = (
checkpoint_for(
kept, held=set(moment.held), floor=checkpoint_floor,
where=moment.checkpoint_where,
)
if arm.checkpoint else {}
)
if arm.stop_only:
rule_ids = [rid for rid in rule_ids if rid == checkpoint.get("rule_id")]
if rule_ids:
try:
io.record_rule_surfaced(
user_id=moment.user_id, rule_ids=rule_ids, source=source,
)
except Exception: # noqa: BLE001 - observation never breaks the observed
_telemetry_failed(source)
return RuleResult(
lines=lines, rule_ids=rule_ids,
shown_rule_ids=[rule.id for _score, rule in shown],
checkpoint=checkpoint,
shown=list(shown),
)
except Exception: # noqa: BLE001 - a recall aid never breaks its act
logger.debug("%s arm failed", arm.source, exc_info=True)
return RuleResult()
# ── The notes corpus (milestone 456 step 4) ──────────────────────────────
#
# The same stages over the other corpus: search, withhold what this response
# already lists, split fresh from repeats, log the call before any early
# return, band, reserved slots, surfacing rows. Two arms take them — the
# prompt's menu (`auto_inject`) and the write path's match by meaning
# (`write_path`) — and they differed only in which kinds they ask for, whether
# a menu sits above them in the same response, and which slots they reserve.
#
# ONE DIFFERENCE FROM THE RULE ARMS IS KEPT, ON PURPOSE, FOR STEP 7. A notes
# arm logs the cut BEFORE its band; a rule arm logs after it. #2085 chose the
# first for notes (the log row is the candidate set a floor is tuned against,
# the surfacing rows are what the reader saw) and #3851 the second for rules.
# Both are recorded decisions, so this refactor reproduces both.
#
# What stays OUTSIDE: every lookup. The records named by number, the snippets
# recorded at a path, the rulings, the design system and the shape ledger
# have no score and no floor; the route builders compose them around what
# this returns.
# Margin gate: drop any hit more than this far below the top hit's score, so a
# single strong match doesn't drag in a wall of barely-passing neighbours.
# Twice the rule band, which is narrow because the rule corpus is flat (#3851).
_NOTE_BAND = 0.10
def _record_kind(note) -> str:
"""The kind marker for an injected menu line — and, for a task, its status.
The menu is drawn from every record that carries an embedding, so a snippet,
a stored process, an issue and a stray dev-log all arrive looking identical.
Recorded prior art only stands out if the line says what it is — and the kind
is also what tells the reader which tool opens it.
Task-ness wins over `note_type` because it's the more useful distinction at a
glance: "there's an open issue about this" beats "there's a note about this".
A TASK ALSO CARRIES ITS STATUS, because for that kind alone the line is
read as a claim about live work. A finished step and an open one rendered
identically is not a cosmetic gap: a done step was cited as a milestone's
open one on the strength of a line exactly like this, which carries an id,
a kind and a title and said nothing about where the work stood (#4154).
Only for tasks — a note or a snippet has no status to be wrong about.
"""
if note.is_task:
kind = "issue" if note.task_kind == "issue" else "task"
# No fallback for a missing status: `is_task` IS `status is not None`
# (models/note.py), so a branch for a task without one could never be
# taken, and a dead branch is a claim about the data that isn't true.
#
# Parenthesised rather than dot-joined: the write-path prior-art line
# joins its own fields with " · ", so a dotted status would read as
# another flag beside `seen` instead of as part of the kind.
return f"{kind} ({note.status})"
return note.note_type or "note"
# WHAT A MENU LINE CARRIES (#4364): the record's NAME, its kind and System,
# and the WHOLE passage that matched. Metadata plus the evidence, rather than a
# title asked to be both.
#
# The name, not the title. A snippet's or lesson's title is `name — when it
# applies` by construction (`embeddings.trigger_title`), because that join is
# what makes it rank on its situation. That is an EMBEDDING shape, and rendered
# as a menu line it ran to 1,500+ characters — the trigger paragraph spent
# again on every line, and again on every repeat. The trigger still arrives
# when it is what matched: it is in the chunk, and `_menu_passage` hands it
# over when the title was the whole match.
#
# The whole passage, not 200 characters of it. The search already chose the
# chunk that matched; the old cut kept its head and tail, and the head is the
# title every chunk is prefixed with — so the reader got the title twice and
# lost the middle, which is where the match was (lesson #4248). A chunk is at
# most ~1.4 KB (`embeddings._CHUNK_CHAR_BUDGET`), and it is shown once: a
# repeat is a one-line pointer (`_menu_seen_line`), not a second copy.
def _menu_name(title: str | None, note_type: str | None, data=None, body: str | None = "") -> str:
"""The record's name — its title without the trigger composed into it."""
title = (title or "(untitled)").replace("\n", " ").strip()
data = data if isinstance(data, dict) else {}
if note_type == "snippet":
from scribe.services.embeddings import TRIGGER_SEP
return (data.get("name") or title.partition(TRIGGER_SEP)[0]).strip() or title
if note_type == LESSON_NOTE_TYPE:
from types import SimpleNamespace
from scribe.services.embeddings import untrigger_title
from scribe.services.lessons import lesson_trigger
trigger = lesson_trigger(SimpleNamespace(data=data, body=body or ""))
# One line even when the stored name is a story (#4797).
return claim_line(untrigger_title(title, trigger).strip() or title)
return title
def _menu_passage(title: str | None, chunk_text: str | None, name: str = "") -> str:
"""The matched chunk on one line, without the title it was embedded under.
Every chunk is `title\nsection` (`embeddings.embedding_text`), and `title`
here must be the EMBEDDED one (`embeddings.document_title`) — for a snippet
or lesson that is `name — trigger`, not the stored name — so the prefix is
stripped exactly. A chunk that WAS only the title — a short
record, or the head chunk of one — matched on the title, and for a
trigger-keyed kind the part of it the name line no longer shows is the
trigger: that is returned, because it is precisely what matched.
One line, so the menu's blockquote survives it.
"""
title = (title or "").strip()
text = (chunk_text or "").strip()
if title and text.startswith(title):
text = text[len(title):]
text = " ".join(text.split())
if not text and name and title.startswith(name) and title != name:
text = " ".join(title[len(name):].lstrip(" —-").split())
return text
def _menu_label(kind: str, systems: list[str] | None) -> str:
"""`issue (done) · Plugin & hooks` — the kind, then where it belongs."""
return " · ".join([kind, *systems]) if systems else kind
def _menu_seen_line(note_id: int, kind: str, name: str) -> str:
"""A pointer to a record this session was already shown, not a copy of it."""
return f"> - #{note_id} [{kind} · seen] {name}"
def menu_entry(
note_id: int, *, kind: str, name: str, systems: list[str] | None = None,
seen: bool = False, stale: bool = False, score: float | None = None,
shared_by: str = "", under: str = "",
) -> list[str]:
"""One record on a notes menu: its line, and what sits under it.
The prompt menu's two blocks — records named by number and records ranked
by meaning — used to write this out twice, differing only in what they
pass: a ranked line carries its score, a named one does not, and each puts
its own text under the line (the matched passage, or the record's opening).
A repeat is a POINTER, not a copy (#4364). The record is in this session's
context already — the ledger is cleared at compaction, so "seen" stays true
— and re-rendering it spent its whole line again for nothing.
A superseded record is DEMOTED, not removed (#278), so one can still reach
a menu, and when it does the reader has to be told: an agent handed stale
material with nothing marking it acts on it with full confidence.
"""
if seen:
line = _menu_seen_line(note_id, kind, name)
return [line + (" — SUPERSEDED" if stale else "")]
line = f"> - #{note_id} [{_menu_label(kind, systems)}] \"{name}\""
if score is not None:
line += f" ({score:.2f})"
if stale:
line += " — SUPERSEDED, a later record covers this; check that first"
if shared_by:
line += f" — shared by {shared_by}"
return [line, f"> ↳ {under}"] if under else [line]
@dataclass(frozen=True)
class NoteSlot:
"""One line reserved for a kind the open ranking keeps losing."""
source: str
kinds: tuple[str, ...]
"""What the slot is FOR — asked for by the search and checked on the way
out, so a slot is never spent on a line indistinguishable from one that
earned its place on score."""
evicts: bool
"""Take the menu's last line when the menu is full, rather than adding one."""
include_global_kinds: bool = False
books_own: bool = False
"""The slot's line is recorded as surfaced under the slot's own source,
and so never again under the arm's."""
declared: Declared | None = None
# Order is load-bearing: reuse evicts the menu's weakest hit while the lesson
# slot extends, so running them the other way round would let a reserved
# lesson be the line reuse throws off — a slot another slot can silently undo
# is not a guarantee.
REUSE_SLOT = NoteSlot(
"reuse_slot", ("snippet", "process"), evicts=True,
declared=Declared("the one line reserved for a reusable snippet"),
)
LESSON_SLOT = NoteSlot(
"lesson_slot", (LESSON_NOTE_TYPE,), evicts=False,
include_global_kinds=True, books_own=True,
declared=Declared("the one line reserved for a lesson"),
)
@dataclass(frozen=True)
class NoteArm:
"""One ranked notes surface: what it searches, and which stages it takes."""
source: str
"""The telemetry `source` of its call row. MUST equal the SURFACES key."""
surfaced_as: str
"""The `source` its surfacing rows carry."""
note_type: tuple[str, ...] | None = None
task_kind: str | None = None
include_global_kinds: bool = True
"""Lessons are project-independent (#3730), so they join the candidate set
from wherever they were learned."""
scope: str = "browse"
"""Nobody asked for an injected line, so it takes the BROWSE scope: never
a record shared one-to-one with the operator."""
slots: tuple[NoteSlot, ...] = ()
withholds: bool = False
"""A menu sits above this arm in the same response, and what it lists is
kept out of the search (`exclude_ids`) — same-call duplication, which is a
different claim from the session ledger (#4101)."""
tuning: Surface | None = None
declared: Declared | None = None
AUTO_INJECT = NoteArm(
"auto_inject", surfaced_as="auto_inject", slots=(REUSE_SLOT, LESSON_SLOT),
tuning=Surface(
name="auto_inject",
floor_key="kb_autoinject_threshold",
floor_default=0.55,
budget_key="kb_autoinject_top_k",
budget_default=3,
asks="the operator's message, as they typed it",
over="notes, snippets, processes and issues",
fires="once per operator turn",
),
declared=Declared("the notes menu offered at the prompt boundary"),
)
# Snippets AND recorded experience (#2246): an issue saying "we tried this and
# it deadlocked" is prior art for the code about to be written. `task_kind`
# keeps the open to-do list out — a task resembles the code and answers
# nothing. Lessons too (milestone 385 step 5): the arm is kind-FILTERED, so a
# kind absent here is unreachable, not merely outranked (#3702). No reserved
# slot: this arm fires before every Write and Edit, and the field is already
# narrow enough that the 200:1 dilution a slot answers does not happen.
WRITE_PATH = NoteArm(
"write_path", surfaced_as="write_path_semantic",
note_type=("snippet", "note", LESSON_NOTE_TYPE), task_kind="issue",
withholds=True,
tuning=Surface(
name="write_path",
floor_key="kb_writepath_threshold",
floor_default=0.68,
budget_key="kb_writepath_top_k",
budget_default=3,
budget_falls_back_to="kb_autoinject_top_k",
asks="the code being written, rewritten as a concept query",
over="snippets and recorded issues",
fires="before every Write and Edit",
),
declared=Declared("prior art offered when a file is about to be written"),
)
NOTE_ARMS: tuple[NoteArm, ...] = (AUTO_INJECT, WRITE_PATH)
NOTE_SLOTS: tuple[NoteSlot, ...] = (REUSE_SLOT, LESSON_SLOT)
# What the notes stages record under, for the registry's fan-out sites.
NOTE_SOURCES: tuple[str, ...] = (
*(arm.source for arm in NOTE_ARMS), *(slot.source for slot in NOTE_SLOTS),
)
NOTE_SURFACED_SOURCES: tuple[str, ...] = (
*(arm.surfaced_as for arm in NOTE_ARMS),
*(slot.source for slot in NOTE_SLOTS if slot.books_own),
)
def note_search_filters(arm: NoteArm) -> dict:
"""The arm's constant search keywords — which kinds, whose records.
Read by `retrieval_review` too: a re-run that searched with different
visibility or kinds from the arm's would judge a menu nobody was shown.
"""
out: dict = {}
if arm.note_type:
out["note_type"] = arm.note_type
if arm.task_kind:
out["task_kind"] = arm.task_kind
if arm.include_global_kinds:
out["include_global_kinds"] = True
out["scope"] = arm.scope
return out
@dataclass(frozen=True)
class NoteIO:
"""The ranker and the two recorders a notes arm reports to."""
search: Callable[..., Any]
"""`semantic_search_notes`, or a stand-in with its signature."""
record_retrieval: Callable[..., Any]
record_surfaced: Callable[..., Any]
@dataclass(frozen=True)
class NoteMoment:
"""What a notes arm searches with, and what this response already holds."""
user_id: int
query: str
project_id: int | None
"""Searched and logged as `project_id or None`."""
seen: frozenset[int] = frozenset()
"""The session ledger: rendered as a pointer, never withheld (#4101)."""
named: frozenset[int] = frozenset()
"""Records this response shows by LOOKUP (named by number). Shown in their
own block and booked under `named_ref`, so this arm did not surface them."""
in_menu: frozenset[int] = frozenset()
"""Records listed earlier in this same response — withheld."""
still_scored: frozenset[int] = frozenset()
"""Of `in_menu`, those the search must still score (a pulled snippet's
resemblance to the payload is evidence for the shape ledger)."""
@dataclass
class NoteResult:
"""What a notes arm put on the menu, and what its search answered."""
menu: list = field(default_factory=list)
"""(score, note) best first: band, slots, and the named records removed."""
answered: list = field(default_factory=list)
"""Everything the search returned, before this response's own menu was
withheld — what `resembles` is read from."""
chunks: dict = field(default_factory=dict)
"""The passage each hit matched on, from the arm's OWN search — so a chunk
is only ever paired with the query that matched it. The slots' lines are
fetched by their own queries and so have none here."""
slot_ids: dict = field(default_factory=dict)
"""{slot source: the id it spent}."""
def _note_band(hits: list) -> list:
"""The top hit, plus every hit within `_NOTE_BAND` of it.
Computed over ALL hits, repeats included: the band measures distance from
the top SCORE, and letting the ledger move that cutoff would make "you
were shown this" change what counts as relevant (#3851's axis
independence).
"""
if not hits:
return []
top = hits[0][0]
return [(s, n) for s, n in hits if s >= top - _NOTE_BAND]
async def _reserve_note_slot(
io: NoteIO, arm: NoteArm, slot: NoteSlot, moment: NoteMoment, menu: list,
*, floor: float, budget: int, shown: set[int],
) -> tuple[list, int | None]:
"""Guarantee `slot.kinds` one line, if one clears the arm's own bar.
WHY A SLOT AT ALL (#2246, milestone 385 step 5). Ranking by raw cosine is
blind to what KIND of record answers what kind of ask, and the corpus
makes that fatal: Scribe's project records are about software work, so a
task about building a helper outranks the snippet that IS one. Snippets
were ~0.5% of the corpus when this was measured; no floor fixes 200:1. A
lesson crowded out is worse still — it exists only to be met at the moment
it applies, so the arm that surfaces it IS its delivery.
THE SLOT BUYS POSITION, NOT A LOWER BAR. It reserves at the menu's own
floor and is NOT held to the band (the top score is the very thing these
kinds lose to), so a weak record cannot buy the line.
ITS OWN SOURCE, from the first deploy. A guarantee has to be falsifiable,
and the hit a slot pushed out sits in the arm's row while the query that
pushed it out would otherwise be nowhere (#2463).
THE LEDGER IS NOT AN EXCLUSION HERE EITHER (#4101), but this call's own
menu is: a record already on it must not be shown twice, while one shown
in an EARLIER call is exactly what a slot may spend itself on — a snippet
relevant then and now is the reuse case, not a duplicate of it.
Returns the possibly-changed menu and the id the slot spent.
"""
if any(_record_kind(n) in slot.kinds for _s, n in menu):
return menu, None
t0 = time.perf_counter()
report: dict = {}
on_menu = {int(n.id) for _s, n in menu}
kwargs: dict = {
"limit": 1, "threshold": floor, "project_id": moment.project_id or None,
"exclude_ids": set(on_menu),
# KIND-FILTERED, so the slot can only be spent on what it is for.
"note_type": slot.kinds, "scope": arm.scope, "report": report,
}
if slot.include_global_kinds:
kwargs["include_global_kinds"] = True
found = await io.search(moment.user_id, moment.query, **kwargs)
fresh = [(s, n) for s, n in found if int(n.id) not in shown]
source = slot.source
try:
io.record_retrieval(
user_id=moment.user_id, source=source, query=moment.query,
threshold=floor, limit=1, project_id=moment.project_id or None,
is_task=None, results=fresh,
best_available=report.get("best_available_score"),
best_available_id=report.get("best_available_id"),
searched=bool(report.get("searched", True)),
suppressed=len(found) - len(fresh),
duration_ms=(time.perf_counter() - t0) * 1000.0,
)
except Exception: # noqa: BLE001 - observation never breaks the observed
_telemetry_failed(source)
# Verified, not trusted: the kind the query asked for, and not a line
# this menu already carries.
placed = [
(s, n) for s, n in found
if _record_kind(n) in slot.kinds and int(n.id) not in on_menu
][:1]
if not placed:
return menu, None
slot_id = int(placed[0][1].id)
# FRESH ONLY, matching the row above: a ledger repeat is rendered (#4101)
# but is not a new surfacing.
if slot.books_own and slot_id not in shown:
try:
io.record_surfaced(
user_id=moment.user_id, note_ids=[slot_id], source=source,
project_id=moment.project_id or None,
)
except Exception: # noqa: BLE001 - observation never breaks the observed
_telemetry_failed(source)
if not slot.evicts:
# IT EXTENDS, IT NEVER DISPLACES, siding with `preference_slot`: a
# displaced hit sits in the arm's row, and evicting it would make the
# two tables disagree about the same call (#3668). And a record that
# does not bind should not throw a better-scoring one off the menu.
return menu + placed, slot_id
# Take the LAST line, never the first: the strongest overall hit is still
# the best answer, and displacing it would trade one blindness for another.
if len(menu) >= budget:
return menu[:budget - 1] + placed, slot_id
return (menu + placed)[:budget], slot_id
async def run_note_arm(
arm: NoteArm, moment: NoteMoment, *, floor: float, budget: int, io: NoteIO,
) -> NoteResult:
"""Run one notes arm through every stage it takes, in the one order.
search → withhold this response's menu → fresh/repeat split → call row →
band → reserved slots → surfacing rows. Fails open to an empty result.
"""
try:
t0 = time.perf_counter()
report: dict = {}
scored = moment.in_menu & moment.still_scored
kwargs: dict = {
# The still-scored ids come back and are dropped below, so the
# limit has to cover them.
"limit": budget + len(scored), "threshold": floor,
"project_id": moment.project_id or None,
**note_search_filters(arm), "report": report,
}
if arm.withholds:
kwargs["exclude_ids"] = set(moment.in_menu - scored)
answered = await io.search(moment.user_id, moment.query, **kwargs)
hits = [(s, n) for s, n in answered if int(n.id) not in moment.in_menu]
# WHAT THIS ARM WITHHELD AFTER THE SEARCH ANSWERED (#3739). The score
# the search reports is measured BEFORE that drop, so a record already
# listed in this response could be logged as one the BAR turned away
# — live proof once read a "rejection" at 0.822 against a lowest
# acceptance of 0.6857. Whenever this removed anything, the honest
# best-available is null: "not measured on this call".
withheld = len(answered) - len(hits)
hits = hits[:budget]
# `results=fresh` and `suppressed` together (#3752): a rendered repeat
# is not a new surfacing, so it is COUNTED rather than reported, and an
# all-repeat zero reads apart from a bar nothing cleared. A record
# NAMED in this response is counted the same way — shown, but by the
# lookup and booked under `named_ref`.
fresh = [
(s, n) for s, n in hits
if int(n.id) not in moment.seen and int(n.id) not in moment.named
]
source = arm.source
try:
io.record_retrieval(
user_id=moment.user_id, source=source, query=moment.query,
threshold=floor, limit=budget,
project_id=moment.project_id or None, is_task=None,
results=fresh,
best_available=(
None if withheld else report.get("best_available_score")
),
# Withheld on the SAME condition: a surviving id beside a null
# score would name a record without saying what it scored.
best_available_id=(
None if withheld else report.get("best_available_id")
),
searched=bool(report.get("searched", True)),
suppressed=len(hits) - len(fresh),
duration_ms=(time.perf_counter() - t0) * 1000.0,
)
except Exception: # noqa: BLE001 - observation never breaks the observed
_telemetry_failed(source)
result = NoteResult(answered=answered, chunks=report.get("best_chunk") or {})
if not hits:
return result
menu = _note_band(hits)
# The slots are told about the named records as if already shown, so
# a slot that picks one does not book a surfacing the lookup booked.
shown = set(moment.seen | moment.named)
for slot in arm.slots:
menu, slot_id = await _reserve_note_slot(
io, arm, slot, moment, menu, floor=floor, budget=budget,
shown=shown,
)
if slot_id is not None:
result.slot_ids[slot.source] = slot_id
# A named record is shown once, in its own block, never again as a match.
menu = [(s, n) for s, n in menu if int(n.id) not in moment.named]
result.menu = menu
# What SURVIVED the band — the menu the reader saw — where the call
# row holds the candidate set a floor is tuned against (#2085). FRESH
# ONLY, the row's own cut (#4101, #3668). A slot that books its own
# line is not this arm's surfacing: counting it twice would leave the
# slot's two tables describing different numbers of one event.
own = {
sid for src, sid in result.slot_ids.items()
if any(slot.books_own and slot.source == src for slot in arm.slots)
}
ids = [
int(n.id) for _s, n in menu
if int(n.id) not in moment.seen and int(n.id) not in own
]
if ids:
source = arm.surfaced_as
try:
io.record_surfaced(
user_id=moment.user_id, note_ids=ids, source=source,
project_id=moment.project_id,
)
except Exception: # noqa: BLE001 - observation never breaks the observed
_telemetry_failed(source)
return result
except Exception: # noqa: BLE001 - a recall aid never breaks its act
logger.debug("%s arm failed", arm.source, exc_info=True)
return NoteResult()
# ── A rule reached through its lessons (milestone 440, #4633) ────────────
#
# A rule's own document is written in the rule's words, which are general by
# design; the situations that keep proving it are often closer to what a
# session is actually doing. A lesson JUDGED to be an instance of a rule (a
# CONFIRMED link) carries that situation, so a lesson matching the moment
# brings its rule along — in rule voice, naming the lesson that reached it.
# It searches the NOTES corpus and answers with rules, which is why it sits
# between the two halves of this module.
#
# ITS OWN SLOT, NOT A RULE SLOT, as a stated default. There was nothing to
# measure until links are confirmed (#4632), so the asymmetry decides it: a
# via-lesson line that took a rule slot could push out a rule the ranker
# matched DIRECTLY, a stronger claim displaced by a weaker one, while an extra
# line costs one line. Logged as its own source so it can be judged (#4636).
VIA_LESSON_SOURCE = "rule_via_lesson"
VIA_LESSON = RankedSource(
VIA_LESSON_SOURCE,
Declared(
"a rule reached through a lesson confirmed as an instance of it",
quiet_because="searches only once some lesson has a confirmed link to "
"a rule; an install where none has been judged is "
"correctly silent here",
),
)
VIA_LESSON_LIMIT = 1
# Lessons fetched before keeping only the linked ones. The search cannot be
# told "linked lessons only", so it overfetches and filters; on a corpus where
# most lessons are unlinked, a fetch of one would almost always be spent on a
# lesson that carries nothing.
_VIA_LESSON_OVERFETCH = 10
@dataclass(frozen=True)
class ViaLessonIO:
"""The lesson search, the link lookups, and the recorders."""
search: Callable[..., Any]
"""`semantic_search_notes`, or a stand-in with its signature."""
linked: Callable[..., Any]
"""`lesson_rules.confirmed_lessons`: the lessons with a confirmed link."""
rules_for: Callable[..., Any]
"""`lesson_rules.confirmed_rules_in_scope`: their rules, in scope."""
floor: Callable[..., Any]
"""The notes menu's own bar — a lesson too weak to be shown cannot carry a
rule in. Awaited only once a linked lesson exists."""
record_retrieval: Callable[..., Any]
record_rule_surfaced: Callable[..., Any]
async def run_via_lesson_arm(
moment: RuleMoment, *, skip: frozenset[int], io: ViaLessonIO,
) -> RuleResult:
"""Rule lines reached through a matching lesson's CONFIRMED links.
`skip` is every rule this response already names plus the session ledger:
suppression applies to the RULE, whichever lesson reached it. Fails open.
"""
try:
confirmed = await io.linked(moment.user_id)
if not confirmed:
return RuleResult()
bar = await io.floor()
t0 = time.perf_counter()
report: dict = {}
found = await io.search(
moment.user_id, moment.query, limit=_VIA_LESSON_OVERFETCH,
threshold=bar, project_id=moment.project_id,
note_type=(LESSON_NOTE_TYPE,), include_global_kinds=True,
scope="browse", report=report,
)
matched = [(s, n) for s, n in found if int(n.id) in confirmed]
by_lesson = await io.rules_for(
moment.user_id, [int(n.id) for _s, n in matched], moment.project_id,
)
candidates = [
(s, rule, n) for s, n in matched for rule in by_lesson.get(int(n.id), [])
]
chosen: list = []
taken: set[int] = set(skip)
for s, rule, n in candidates:
if rule.id in taken:
continue
taken.add(rule.id)
chosen.append((s, rule, n))
if len(chosen) >= VIA_LESSON_LIMIT:
break
# Logged whenever a search ran, results or not — the #3497 guard. The
# score is the LESSON's, because the lesson is what was matched; the
# best-available id is left off for the same reason, since the row's
# results are rules and an id beside them would read as a rule id.
try:
io.record_retrieval(
user_id=moment.user_id, source=VIA_LESSON_SOURCE,
query=moment.query, threshold=bar, limit=VIA_LESSON_LIMIT,
project_id=moment.project_id, is_task=None,
results=[(s, rule) for s, rule, _n in chosen],
duration_ms=(time.perf_counter() - t0) * 1000.0,
best_available=report.get("best_available_score"),
searched=bool(report.get("searched", True)),
suppressed=len({r.id for _s, r, _n in candidates}) - len(chosen),
)
except Exception: # noqa: BLE001 - observation never breaks the observed
_telemetry_failed(VIA_LESSON_SOURCE)
if not chosen:
return RuleResult()
rule_ids = [rule.id for _s, rule, _n in chosen]
try:
io.record_rule_surfaced(
user_id=moment.user_id, rule_ids=rule_ids, source=VIA_LESSON_SOURCE,
)
except Exception: # noqa: BLE001 - observation never breaks the observed
_telemetry_failed(VIA_LESSON_SOURCE)
lines = [
_rule_hint_line(rule, where=moment.where, seen=False,
held=rule.id in moment.held)
+ f" Reached through lesson #{n.id} "
+ f"“{_menu_name(n.title, n.note_type, n.data, n.body)}”, "
+ "a recorded instance of it."
for _s, rule, n in chosen
]
return RuleResult(
lines=lines, rule_ids=rule_ids, shown_rule_ids=list(rule_ids),
shown=[(s, rule) for s, rule, _n in chosen],
)
except Exception: # noqa: BLE001 - a recall aid never breaks its act
logger.debug("%s arm failed", VIA_LESSON_SOURCE, exc_info=True)
return RuleResult()
# ── The specs, read as lists (milestone 456 step 5) ──────────────────────
#
# What `retrieval_surfaces.SURFACES`, the ranked rows of
# `retrieval_registry.POINTS` and `rule_usage.RANKED_SOURCES` are read from.
# The order is the order the Settings page and the readouts list them in.
TUNED_ARMS: tuple = (*NOTE_ARMS, *RULE_ARMS)
"""Every arm with a floor and a budget: one per tunable surface."""
RANKED: tuple = (
*TUNED_ARMS, PREFERENCE_SLOT, *NOTE_SLOTS, VIA_LESSON,
)
"""Every source that RANKED what it showed — the arms, and the stages that
run a query of their own and record under their own name."""
RULE_RANKED_SOURCES: tuple[str, ...] = (
*(arm.source for arm in RULE_ARMS), PREFERENCE_SLOT_SOURCE, VIA_LESSON_SOURCE,
)
"""The ranked sources that surface RULES — a ranker chose each line, so a
pull can confirm or refute it (`rule_usage.RANKED_SOURCES`)."""
# ── The moment arm (milestone 458) ───────────────────────────────────────
#
# A LOOKUP beside the ranked arms, not one of them. A rule mounted on a moment
# arrives because somebody said it belongs there, not because the work's words
# resemble it — which is the whole point: a rule about WHEN work is finished
# never scores against what is said while finishing it. So there is no score,
# no floor and no band, and it writes no retrieval_logs row (the rulings arm's
# precedent, #4769): every score-shaped warning would misread a lookup. What it
# records is the surfacing, with the moment beside each row, so a mount's
# pull-through can be read per moment.
#
# It shares everything else with the ranked arms: the line renderer, the
# session ledger (a repeat renders as a citation, never hidden — #3750) and
# the fresh-only surfacing rows (#3752).
MOMENT_RULE_SOURCE = "moment_rule"
# A cap against a misconfiguration, not a ranking budget. Mounts are
# deliberate, so the ordinary case is a handful; a rulebook that mounted
# dozens of rules on one moment would otherwise bury the act it annotates.
# The remedy for a moment that hits this is fewer mounts (unmount through
# update_rule), which is why it is not a tuning dial.
MOMENT_RULE_LIMIT = 8
@dataclass(frozen=True)
class MomentIO:
"""The mounted-rule lookup and the recorder the moment arm reports to."""
lookup: Callable[..., Any]
"""`rulebooks.rules_on_moments`, or a stand-in with its signature."""
record_rule_surfaced: Callable[..., Any]
def _reached_by(hit: dict) -> str:
"""How the line names what reached a moment: the action, as mapped."""
return f"`{hit.get('match') or hit.get('tool') or '?'}`"
async def run_moment_arm(
reached: list[dict], moment: RuleMoment, *, io: MomentIO,
limit: int = MOMENT_RULE_LIMIT,
) -> RuleResult:
"""Deliver the rules mounted on the moments this act reached.
`reached` is `moment_actions.resolve`'s answer — each moment with the
action that reached it — and every line says both ("at work.deliver,
reached by `git push`"), so a moment that fired on the wrong act is
visible in the line itself and can be corrected in the session with
`unmap_action` (step 2's ruling). Fresh rules come first and carry their
trigger; a rule the session was already shown is cited, not repeated in
full. Fails open to an empty result.
"""
try:
names = [hit["moment"] for hit in reached]
if not names:
return RuleResult()
pairs = await io.lookup(moment.user_id, names, moment.project_id or None)
if not pairs:
return RuleResult()
by_moment = {hit["moment"]: hit for hit in reached}
# Stable: catalog order within each half, fresh half first.
shown = sorted(pairs, key=lambda pair: pair[0].id in moment.exclude)[:limit]
lines = []
for rule, at in shown:
seen = rule.id in moment.exclude
lines.append(_rule_hint_line(
rule,
where=f"at {at}, reached by {_reached_by(by_moment.get(at, {}))}",
seen=seen, held=rule.id in moment.held,
# A repeat is cited rather than quoted: the session was told
# which moment brought it the first time.
compact=seen,
))
fresh = [(rule, at) for rule, at in shown if rule.id not in moment.exclude]
rule_ids = [rule.id for rule, _at in fresh]
if rule_ids:
try:
io.record_rule_surfaced(
user_id=moment.user_id, rule_ids=rule_ids,
source=MOMENT_RULE_SOURCE,
detail={rule.id: at for rule, at in fresh},
)
except Exception: # noqa: BLE001 - observation never breaks the observed
_telemetry_failed(MOMENT_RULE_SOURCE)
return RuleResult(
lines=lines, rule_ids=rule_ids,
shown_rule_ids=[rule.id for rule, _at in shown],
shown=[(None, rule) for rule, _at in shown],
)
except Exception: # noqa: BLE001 - a recall aid never breaks its act
logger.debug("%s arm failed", MOMENT_RULE_SOURCE, exc_info=True)
return RuleResult()