refactor(retrieval): the notes arms run on the one pipeline - auto_inject, its reuse and lesson slots, the write path by meaning, and rule_via_lesson are specs (milestone 456 step 4, #4906)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m2s
CI & Build / Python tests (push) Successful in 1m52s
CI & Build / Build & push image (push) Successful in 35s

retrieval_pipeline gains the notes half: NoteArm / NoteSlot / NoteMoment /
NoteIO / NoteResult and run_note_arm, which writes once the stages both
notes arms copied: search, withhold this response's own menu (#3739),
fresh/repeat split, the call row before any return (#3497, #3752), the
band, the reserved slots in their order (reuse evicts, lesson extends),
and the surfacing rows. The note renderer (_record_kind, _menu_name,
_menu_passage, the seen pointer, menu_entry) moves with it, and
run_via_lesson_arm takes rule_via_lesson.

Behaviour-preserving, with flags for today's differences: notes still log
BEFORE the band and rules after it (step 7's question). One deliberate
change: a notes arm now fails open like the rule arms, so a failing
search costs its lines and no longer the whole hook response.

The I/O is resolved from plugin_context at call time (_note_io), so the
existing patches keep working. The review re-run reads the AUTO_INJECT
spec instead of restating it, and its guard now compares the two live
searches. The registry declares the pipeline's notes fan-out sites.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-10-05 20:02:37 -04:00
co-authored by Claude Opus 5.5
parent 4b1060ae8b
commit 2b8f41229d
10 changed files with 1172 additions and 763 deletions
+659
View File
@@ -39,6 +39,8 @@ import time
from dataclasses import dataclass, field
from typing import Any, Callable
from scribe.services.lessons import LESSON_NOTE_TYPE, claim_line
logger = logging.getLogger(__name__)
# ── Shared rendering and ranking helpers (moved from plugin_context) ──────
@@ -681,6 +683,663 @@ async def run_rule_arm(
return RuleResult()
# ── The notes corpus (milestone 456 step 4) ──────────────────────────────
#
# The same stages over the other corpus: search, withhold what this response
# already lists, split fresh from repeats, log the call before any early
# return, band, reserved slots, surfacing rows. Two arms take them — the
# prompt's menu (`auto_inject`) and the write path's match by meaning
# (`write_path`) — and they differed only in which kinds they ask for, whether
# a menu sits above them in the same response, and which slots they reserve.
#
# ONE DIFFERENCE FROM THE RULE ARMS IS KEPT, ON PURPOSE, FOR STEP 7. A notes
# arm logs the cut BEFORE its band; a rule arm logs after it. #2085 chose the
# first for notes (the log row is the candidate set a floor is tuned against,
# the surfacing rows are what the reader saw) and #3851 the second for rules.
# Both are recorded decisions, so this refactor reproduces both.
#
# What stays OUTSIDE: every lookup. The records named by number, the snippets
# recorded at a path, the rulings, the design system and the shape ledger
# have no score and no floor; the route builders compose them around what
# this returns.
# Margin gate: drop any hit more than this far below the top hit's score, so a
# single strong match doesn't drag in a wall of barely-passing neighbours.
# Twice the rule band, which is narrow because the rule corpus is flat (#3851).
_NOTE_BAND = 0.10
def _record_kind(note) -> str:
"""The kind marker for an injected menu line — and, for a task, its status.
The menu is drawn from every record that carries an embedding, so a snippet,
a stored process, an issue and a stray dev-log all arrive looking identical.
Recorded prior art only stands out if the line says what it is — and the kind
is also what tells the reader which tool opens it.
Task-ness wins over `note_type` because it's the more useful distinction at a
glance: "there's an open issue about this" beats "there's a note about this".
A TASK ALSO CARRIES ITS STATUS, because for that kind alone the line is
read as a claim about live work. A finished step and an open one rendered
identically is not a cosmetic gap: a done step was cited as a milestone's
open one on the strength of a line exactly like this, which carries an id,
a kind and a title and said nothing about where the work stood (#4154).
Only for tasks — a note or a snippet has no status to be wrong about.
"""
if note.is_task:
kind = "issue" if note.task_kind == "issue" else "task"
# No fallback for a missing status: `is_task` IS `status is not None`
# (models/note.py), so a branch for a task without one could never be
# taken, and a dead branch is a claim about the data that isn't true.
#
# Parenthesised rather than dot-joined: the write-path prior-art line
# joins its own fields with " · ", so a dotted status would read as
# another flag beside `seen` instead of as part of the kind.
return f"{kind} ({note.status})"
return note.note_type or "note"
# WHAT A MENU LINE CARRIES (#4364): the record's NAME, its kind and System,
# and the WHOLE passage that matched. Metadata plus the evidence, rather than a
# title asked to be both.
#
# The name, not the title. A snippet's or lesson's title is `name — when it
# applies` by construction (`embeddings.trigger_title`), because that join is
# what makes it rank on its situation. That is an EMBEDDING shape, and rendered
# as a menu line it ran to 1,500+ characters — the trigger paragraph spent
# again on every line, and again on every repeat. The trigger still arrives
# when it is what matched: it is in the chunk, and `_menu_passage` hands it
# over when the title was the whole match.
#
# The whole passage, not 200 characters of it. The search already chose the
# chunk that matched; the old cut kept its head and tail, and the head is the
# title every chunk is prefixed with — so the reader got the title twice and
# lost the middle, which is where the match was (lesson #4248). A chunk is at
# most ~1.4 KB (`embeddings._CHUNK_CHAR_BUDGET`), and it is shown once: a
# repeat is a one-line pointer (`_menu_seen_line`), not a second copy.
def _menu_name(title: str | None, note_type: str | None, data=None, body: str | None = "") -> str:
"""The record's name — its title without the trigger composed into it."""
title = (title or "(untitled)").replace("\n", " ").strip()
data = data if isinstance(data, dict) else {}
if note_type == "snippet":
from scribe.services.embeddings import TRIGGER_SEP
return (data.get("name") or title.partition(TRIGGER_SEP)[0]).strip() or title
if note_type == LESSON_NOTE_TYPE:
from types import SimpleNamespace
from scribe.services.embeddings import untrigger_title
from scribe.services.lessons import lesson_trigger
trigger = lesson_trigger(SimpleNamespace(data=data, body=body or ""))
# One line even when the stored name is a story (#4797).
return claim_line(untrigger_title(title, trigger).strip() or title)
return title
def _menu_passage(title: str | None, chunk_text: str | None, name: str = "") -> str:
"""The matched chunk on one line, without the title it was embedded under.
Every chunk is `title\nsection` (`embeddings.embedding_text`), and `title`
here must be the EMBEDDED one (`embeddings.document_title`) — for a snippet
or lesson that is `name — trigger`, not the stored name — so the prefix is
stripped exactly. A chunk that WAS only the title — a short
record, or the head chunk of one — matched on the title, and for a
trigger-keyed kind the part of it the name line no longer shows is the
trigger: that is returned, because it is precisely what matched.
One line, so the menu's blockquote survives it.
"""
title = (title or "").strip()
text = (chunk_text or "").strip()
if title and text.startswith(title):
text = text[len(title):]
text = " ".join(text.split())
if not text and name and title.startswith(name) and title != name:
text = " ".join(title[len(name):].lstrip(" —-").split())
return text
def _menu_label(kind: str, systems: list[str] | None) -> str:
"""`issue (done) · Plugin & hooks` — the kind, then where it belongs."""
return " · ".join([kind, *systems]) if systems else kind
def _menu_seen_line(note_id: int, kind: str, name: str) -> str:
"""A pointer to a record this session was already shown, not a copy of it."""
return f"> - #{note_id} [{kind} · seen] {name}"
def menu_entry(
note_id: int, *, kind: str, name: str, systems: list[str] | None = None,
seen: bool = False, stale: bool = False, score: float | None = None,
shared_by: str = "", under: str = "",
) -> list[str]:
"""One record on a notes menu: its line, and what sits under it.
The prompt menu's two blocks — records named by number and records ranked
by meaning — used to write this out twice, differing only in what they
pass: a ranked line carries its score, a named one does not, and each puts
its own text under the line (the matched passage, or the record's opening).
A repeat is a POINTER, not a copy (#4364). The record is in this session's
context already — the ledger is cleared at compaction, so "seen" stays true
— and re-rendering it spent its whole line again for nothing.
A superseded record is DEMOTED, not removed (#278), so one can still reach
a menu, and when it does the reader has to be told: an agent handed stale
material with nothing marking it acts on it with full confidence.
"""
if seen:
line = _menu_seen_line(note_id, kind, name)
return [line + (" — SUPERSEDED" if stale else "")]
line = f"> - #{note_id} [{_menu_label(kind, systems)}] \"{name}\""
if score is not None:
line += f" ({score:.2f})"
if stale:
line += " — SUPERSEDED, a later record covers this; check that first"
if shared_by:
line += f" — shared by {shared_by}"
return [line, f"> ↳ {under}"] if under else [line]
@dataclass(frozen=True)
class NoteSlot:
"""One line reserved for a kind the open ranking keeps losing."""
source: str
kinds: tuple[str, ...]
"""What the slot is FOR — asked for by the search and checked on the way
out, so a slot is never spent on a line indistinguishable from one that
earned its place on score."""
evicts: bool
"""Take the menu's last line when the menu is full, rather than adding one."""
include_global_kinds: bool = False
books_own: bool = False
"""The slot's line is recorded as surfaced under the slot's own source,
and so never again under the arm's."""
# Order is load-bearing: reuse evicts the menu's weakest hit while the lesson
# slot extends, so running them the other way round would let a reserved
# lesson be the line reuse throws off — a slot another slot can silently undo
# is not a guarantee.
REUSE_SLOT = NoteSlot("reuse_slot", ("snippet", "process"), evicts=True)
LESSON_SLOT = NoteSlot(
"lesson_slot", (LESSON_NOTE_TYPE,), evicts=False,
include_global_kinds=True, books_own=True,
)
@dataclass(frozen=True)
class NoteArm:
"""One ranked notes surface: what it searches, and which stages it takes."""
source: str
"""The telemetry `source` of its call row. MUST equal the SURFACES key."""
surfaced_as: str
"""The `source` its surfacing rows carry."""
note_type: tuple[str, ...] | None = None
task_kind: str | None = None
include_global_kinds: bool = True
"""Lessons are project-independent (#3730), so they join the candidate set
from wherever they were learned."""
scope: str = "browse"
"""Nobody asked for an injected line, so it takes the BROWSE scope: never
a record shared one-to-one with the operator."""
slots: tuple[NoteSlot, ...] = ()
withholds: bool = False
"""A menu sits above this arm in the same response, and what it lists is
kept out of the search (`exclude_ids`) — same-call duplication, which is a
different claim from the session ledger (#4101)."""
AUTO_INJECT = NoteArm(
"auto_inject", surfaced_as="auto_inject", slots=(REUSE_SLOT, LESSON_SLOT),
)
# Snippets AND recorded experience (#2246): an issue saying "we tried this and
# it deadlocked" is prior art for the code about to be written. `task_kind`
# keeps the open to-do list out — a task resembles the code and answers
# nothing. Lessons too (milestone 385 step 5): the arm is kind-FILTERED, so a
# kind absent here is unreachable, not merely outranked (#3702). No reserved
# slot: this arm fires before every Write and Edit, and the field is already
# narrow enough that the 200:1 dilution a slot answers does not happen.
WRITE_PATH = NoteArm(
"write_path", surfaced_as="write_path_semantic",
note_type=("snippet", "note", LESSON_NOTE_TYPE), task_kind="issue",
withholds=True,
)
NOTE_ARMS: tuple[NoteArm, ...] = (AUTO_INJECT, WRITE_PATH)
NOTE_SLOTS: tuple[NoteSlot, ...] = (REUSE_SLOT, LESSON_SLOT)
# What the notes stages record under, for the registry's fan-out sites.
NOTE_SOURCES: tuple[str, ...] = (
*(arm.source for arm in NOTE_ARMS), *(slot.source for slot in NOTE_SLOTS),
)
NOTE_SURFACED_SOURCES: tuple[str, ...] = (
*(arm.surfaced_as for arm in NOTE_ARMS),
*(slot.source for slot in NOTE_SLOTS if slot.books_own),
)
def note_search_filters(arm: NoteArm) -> dict:
"""The arm's constant search keywords — which kinds, whose records.
Read by `retrieval_review` too: a re-run that searched with different
visibility or kinds from the arm's would judge a menu nobody was shown.
"""
out: dict = {}
if arm.note_type:
out["note_type"] = arm.note_type
if arm.task_kind:
out["task_kind"] = arm.task_kind
if arm.include_global_kinds:
out["include_global_kinds"] = True
out["scope"] = arm.scope
return out
@dataclass(frozen=True)
class NoteIO:
"""The ranker and the two recorders a notes arm reports to."""
search: Callable[..., Any]
"""`semantic_search_notes`, or a stand-in with its signature."""
record_retrieval: Callable[..., Any]
record_surfaced: Callable[..., Any]
@dataclass(frozen=True)
class NoteMoment:
"""What a notes arm searches with, and what this response already holds."""
user_id: int
query: str
project_id: int | None
"""Searched and logged as `project_id or None`."""
seen: frozenset[int] = frozenset()
"""The session ledger: rendered as a pointer, never withheld (#4101)."""
named: frozenset[int] = frozenset()
"""Records this response shows by LOOKUP (named by number). Shown in their
own block and booked under `named_ref`, so this arm did not surface them."""
in_menu: frozenset[int] = frozenset()
"""Records listed earlier in this same response — withheld."""
still_scored: frozenset[int] = frozenset()
"""Of `in_menu`, those the search must still score (a pulled snippet's
resemblance to the payload is evidence for the shape ledger)."""
@dataclass
class NoteResult:
"""What a notes arm put on the menu, and what its search answered."""
menu: list = field(default_factory=list)
"""(score, note) best first: band, slots, and the named records removed."""
answered: list = field(default_factory=list)
"""Everything the search returned, before this response's own menu was
withheld — what `resembles` is read from."""
chunks: dict = field(default_factory=dict)
"""The passage each hit matched on, from the arm's OWN search — so a chunk
is only ever paired with the query that matched it. The slots' lines are
fetched by their own queries and so have none here."""
slot_ids: dict = field(default_factory=dict)
"""{slot source: the id it spent}."""
def _note_band(hits: list) -> list:
"""The top hit, plus every hit within `_NOTE_BAND` of it.
Computed over ALL hits, repeats included: the band measures distance from
the top SCORE, and letting the ledger move that cutoff would make "you
were shown this" change what counts as relevant (#3851's axis
independence).
"""
if not hits:
return []
top = hits[0][0]
return [(s, n) for s, n in hits if s >= top - _NOTE_BAND]
async def _reserve_note_slot(
io: NoteIO, arm: NoteArm, slot: NoteSlot, moment: NoteMoment, menu: list,
*, floor: float, budget: int, shown: set[int],
) -> tuple[list, int | None]:
"""Guarantee `slot.kinds` one line, if one clears the arm's own bar.
WHY A SLOT AT ALL (#2246, milestone 385 step 5). Ranking by raw cosine is
blind to what KIND of record answers what kind of ask, and the corpus
makes that fatal: Scribe's project records are about software work, so a
task about building a helper outranks the snippet that IS one. Snippets
were ~0.5% of the corpus when this was measured; no floor fixes 200:1. A
lesson crowded out is worse still — it exists only to be met at the moment
it applies, so the arm that surfaces it IS its delivery.
THE SLOT BUYS POSITION, NOT A LOWER BAR. It reserves at the menu's own
floor and is NOT held to the band (the top score is the very thing these
kinds lose to), so a weak record cannot buy the line.
ITS OWN SOURCE, from the first deploy. A guarantee has to be falsifiable,
and the hit a slot pushed out sits in the arm's row while the query that
pushed it out would otherwise be nowhere (#2463).
THE LEDGER IS NOT AN EXCLUSION HERE EITHER (#4101), but this call's own
menu is: a record already on it must not be shown twice, while one shown
in an EARLIER call is exactly what a slot may spend itself on — a snippet
relevant then and now is the reuse case, not a duplicate of it.
Returns the possibly-changed menu and the id the slot spent.
"""
if any(_record_kind(n) in slot.kinds for _s, n in menu):
return menu, None
t0 = time.perf_counter()
report: dict = {}
on_menu = {int(n.id) for _s, n in menu}
kwargs: dict = {
"limit": 1, "threshold": floor, "project_id": moment.project_id or None,
"exclude_ids": set(on_menu),
# KIND-FILTERED, so the slot can only be spent on what it is for.
"note_type": slot.kinds, "scope": arm.scope, "report": report,
}
if slot.include_global_kinds:
kwargs["include_global_kinds"] = True
found = await io.search(moment.user_id, moment.query, **kwargs)
fresh = [(s, n) for s, n in found if int(n.id) not in shown]
source = slot.source
try:
io.record_retrieval(
user_id=moment.user_id, source=source, query=moment.query,
threshold=floor, limit=1, project_id=moment.project_id or None,
is_task=None, results=fresh,
best_available=report.get("best_available_score"),
best_available_id=report.get("best_available_id"),
searched=bool(report.get("searched", True)),
suppressed=len(found) - len(fresh),
duration_ms=(time.perf_counter() - t0) * 1000.0,
)
except Exception: # noqa: BLE001 - observation never breaks the observed
_telemetry_failed(source)
# Verified, not trusted: the kind the query asked for, and not a line
# this menu already carries.
placed = [
(s, n) for s, n in found
if _record_kind(n) in slot.kinds and int(n.id) not in on_menu
][:1]
if not placed:
return menu, None
slot_id = int(placed[0][1].id)
# FRESH ONLY, matching the row above: a ledger repeat is rendered (#4101)
# but is not a new surfacing.
if slot.books_own and slot_id not in shown:
try:
io.record_surfaced(
user_id=moment.user_id, note_ids=[slot_id], source=source,
project_id=moment.project_id or None,
)
except Exception: # noqa: BLE001 - observation never breaks the observed
_telemetry_failed(source)
if not slot.evicts:
# IT EXTENDS, IT NEVER DISPLACES, siding with `preference_slot`: a
# displaced hit sits in the arm's row, and evicting it would make the
# two tables disagree about the same call (#3668). And a record that
# does not bind should not throw a better-scoring one off the menu.
return menu + placed, slot_id
# Take the LAST line, never the first: the strongest overall hit is still
# the best answer, and displacing it would trade one blindness for another.
if len(menu) >= budget:
return menu[:budget - 1] + placed, slot_id
return (menu + placed)[:budget], slot_id
async def run_note_arm(
arm: NoteArm, moment: NoteMoment, *, floor: float, budget: int, io: NoteIO,
) -> NoteResult:
"""Run one notes arm through every stage it takes, in the one order.
search → withhold this response's menu → fresh/repeat split → call row →
band → reserved slots → surfacing rows. Fails open to an empty result.
"""
try:
t0 = time.perf_counter()
report: dict = {}
scored = moment.in_menu & moment.still_scored
kwargs: dict = {
# The still-scored ids come back and are dropped below, so the
# limit has to cover them.
"limit": budget + len(scored), "threshold": floor,
"project_id": moment.project_id or None,
**note_search_filters(arm), "report": report,
}
if arm.withholds:
kwargs["exclude_ids"] = set(moment.in_menu - scored)
answered = await io.search(moment.user_id, moment.query, **kwargs)
hits = [(s, n) for s, n in answered if int(n.id) not in moment.in_menu]
# WHAT THIS ARM WITHHELD AFTER THE SEARCH ANSWERED (#3739). The score
# the search reports is measured BEFORE that drop, so a record already
# listed in this response could be logged as one the BAR turned away
# — live proof once read a "rejection" at 0.822 against a lowest
# acceptance of 0.6857. Whenever this removed anything, the honest
# best-available is null: "not measured on this call".
withheld = len(answered) - len(hits)
hits = hits[:budget]
# `results=fresh` and `suppressed` together (#3752): a rendered repeat
# is not a new surfacing, so it is COUNTED rather than reported, and an
# all-repeat zero reads apart from a bar nothing cleared. A record
# NAMED in this response is counted the same way — shown, but by the
# lookup and booked under `named_ref`.
fresh = [
(s, n) for s, n in hits
if int(n.id) not in moment.seen and int(n.id) not in moment.named
]
source = arm.source
try:
io.record_retrieval(
user_id=moment.user_id, source=source, query=moment.query,
threshold=floor, limit=budget,
project_id=moment.project_id or None, is_task=None,
results=fresh,
best_available=(
None if withheld else report.get("best_available_score")
),
# Withheld on the SAME condition: a surviving id beside a null
# score would name a record without saying what it scored.
best_available_id=(
None if withheld else report.get("best_available_id")
),
searched=bool(report.get("searched", True)),
suppressed=len(hits) - len(fresh),
duration_ms=(time.perf_counter() - t0) * 1000.0,
)
except Exception: # noqa: BLE001 - observation never breaks the observed
_telemetry_failed(source)
result = NoteResult(answered=answered, chunks=report.get("best_chunk") or {})
if not hits:
return result
menu = _note_band(hits)
# The slots are told about the named records as if already shown, so
# a slot that picks one does not book a surfacing the lookup booked.
shown = set(moment.seen | moment.named)
for slot in arm.slots:
menu, slot_id = await _reserve_note_slot(
io, arm, slot, moment, menu, floor=floor, budget=budget,
shown=shown,
)
if slot_id is not None:
result.slot_ids[slot.source] = slot_id
# A named record is shown once, in its own block, never again as a match.
menu = [(s, n) for s, n in menu if int(n.id) not in moment.named]
result.menu = menu
# What SURVIVED the band — the menu the reader saw — where the call
# row holds the candidate set a floor is tuned against (#2085). FRESH
# ONLY, the row's own cut (#4101, #3668). A slot that books its own
# line is not this arm's surfacing: counting it twice would leave the
# slot's two tables describing different numbers of one event.
own = {
sid for src, sid in result.slot_ids.items()
if any(slot.books_own and slot.source == src for slot in arm.slots)
}
ids = [
int(n.id) for _s, n in menu
if int(n.id) not in moment.seen and int(n.id) not in own
]
if ids:
source = arm.surfaced_as
try:
io.record_surfaced(
user_id=moment.user_id, note_ids=ids, source=source,
project_id=moment.project_id,
)
except Exception: # noqa: BLE001 - observation never breaks the observed
_telemetry_failed(source)
return result
except Exception: # noqa: BLE001 - a recall aid never breaks its act
logger.debug("%s arm failed", arm.source, exc_info=True)
return NoteResult()
# ── A rule reached through its lessons (milestone 440, #4633) ────────────
#
# A rule's own document is written in the rule's words, which are general by
# design; the situations that keep proving it are often closer to what a
# session is actually doing. A lesson JUDGED to be an instance of a rule (a
# CONFIRMED link) carries that situation, so a lesson matching the moment
# brings its rule along — in rule voice, naming the lesson that reached it.
# It searches the NOTES corpus and answers with rules, which is why it sits
# between the two halves of this module.
#
# ITS OWN SLOT, NOT A RULE SLOT, as a stated default. There was nothing to
# measure until links are confirmed (#4632), so the asymmetry decides it: a
# via-lesson line that took a rule slot could push out a rule the ranker
# matched DIRECTLY, a stronger claim displaced by a weaker one, while an extra
# line costs one line. Logged as its own source so it can be judged (#4636).
VIA_LESSON_SOURCE = "rule_via_lesson"
VIA_LESSON_LIMIT = 1
# Lessons fetched before keeping only the linked ones. The search cannot be
# told "linked lessons only", so it overfetches and filters; on a corpus where
# most lessons are unlinked, a fetch of one would almost always be spent on a
# lesson that carries nothing.
_VIA_LESSON_OVERFETCH = 10
@dataclass(frozen=True)
class ViaLessonIO:
"""The lesson search, the link lookups, and the recorders."""
search: Callable[..., Any]
"""`semantic_search_notes`, or a stand-in with its signature."""
linked: Callable[..., Any]
"""`lesson_rules.confirmed_lessons`: the lessons with a confirmed link."""
rules_for: Callable[..., Any]
"""`lesson_rules.confirmed_rules_in_scope`: their rules, in scope."""
floor: Callable[..., Any]
"""The notes menu's own bar — a lesson too weak to be shown cannot carry a
rule in. Awaited only once a linked lesson exists."""
record_retrieval: Callable[..., Any]
record_rule_surfaced: Callable[..., Any]
async def run_via_lesson_arm(
moment: RuleMoment, *, skip: frozenset[int], io: ViaLessonIO,
) -> RuleResult:
"""Rule lines reached through a matching lesson's CONFIRMED links.
`skip` is every rule this response already names plus the session ledger:
suppression applies to the RULE, whichever lesson reached it. Fails open.
"""
try:
confirmed = await io.linked(moment.user_id)
if not confirmed:
return RuleResult()
bar = await io.floor()
t0 = time.perf_counter()
report: dict = {}
found = await io.search(
moment.user_id, moment.query, limit=_VIA_LESSON_OVERFETCH,
threshold=bar, project_id=moment.project_id,
note_type=(LESSON_NOTE_TYPE,), include_global_kinds=True,
scope="browse", report=report,
)
matched = [(s, n) for s, n in found if int(n.id) in confirmed]
by_lesson = await io.rules_for(
moment.user_id, [int(n.id) for _s, n in matched], moment.project_id,
)
candidates = [
(s, rule, n) for s, n in matched for rule in by_lesson.get(int(n.id), [])
]
chosen: list = []
taken: set[int] = set(skip)
for s, rule, n in candidates:
if rule.id in taken:
continue
taken.add(rule.id)
chosen.append((s, rule, n))
if len(chosen) >= VIA_LESSON_LIMIT:
break
# Logged whenever a search ran, results or not — the #3497 guard. The
# score is the LESSON's, because the lesson is what was matched; the
# best-available id is left off for the same reason, since the row's
# results are rules and an id beside them would read as a rule id.
try:
io.record_retrieval(
user_id=moment.user_id, source=VIA_LESSON_SOURCE,
query=moment.query, threshold=bar, limit=VIA_LESSON_LIMIT,
project_id=moment.project_id, is_task=None,
results=[(s, rule) for s, rule, _n in chosen],
duration_ms=(time.perf_counter() - t0) * 1000.0,
best_available=report.get("best_available_score"),
searched=bool(report.get("searched", True)),
suppressed=len({r.id for _s, r, _n in candidates}) - len(chosen),
)
except Exception: # noqa: BLE001 - observation never breaks the observed
_telemetry_failed(VIA_LESSON_SOURCE)
if not chosen:
return RuleResult()
rule_ids = [rule.id for _s, rule, _n in chosen]
try:
io.record_rule_surfaced(
user_id=moment.user_id, rule_ids=rule_ids, source=VIA_LESSON_SOURCE,
)
except Exception: # noqa: BLE001 - observation never breaks the observed
_telemetry_failed(VIA_LESSON_SOURCE)
lines = [
_rule_hint_line(rule, where=moment.where, seen=False,
held=rule.id in moment.held)
+ f" Reached through lesson #{n.id} "
+ f"“{_menu_name(n.title, n.note_type, n.data, n.body)}”, "
+ "a recorded instance of it."
for _s, rule, n in chosen
]
return RuleResult(
lines=lines, rule_ids=rule_ids, shown_rule_ids=list(rule_ids),
shown=[(s, rule) for s, rule, _n in chosen],
)
except Exception: # noqa: BLE001 - a recall aid never breaks its act
logger.debug("%s arm failed", VIA_LESSON_SOURCE, exc_info=True)
return RuleResult()
# ── The moment arm (milestone 458) ───────────────────────────────────────
#
# A LOOKUP beside the ranked arms, not one of them. A rule mounted on a moment