refactor(retrieval): the notes arms run on the one pipeline - auto_inject, its reuse and lesson slots, the write path by meaning, and rule_via_lesson are specs (milestone 456 step 4, #4906)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m2s
CI & Build / Python tests (push) Successful in 1m52s
CI & Build / Build & push image (push) Successful in 35s

retrieval_pipeline gains the notes half: NoteArm / NoteSlot / NoteMoment /
NoteIO / NoteResult and run_note_arm, which writes once the stages both
notes arms copied: search, withhold this response's own menu (#3739),
fresh/repeat split, the call row before any return (#3497, #3752), the
band, the reserved slots in their order (reuse evicts, lesson extends),
and the surfacing rows. The note renderer (_record_kind, _menu_name,
_menu_passage, the seen pointer, menu_entry) moves with it, and
run_via_lesson_arm takes rule_via_lesson.

Behaviour-preserving, with flags for today's differences: notes still log
BEFORE the band and rules after it (step 7's question). One deliberate
change: a notes arm now fails open like the rule arms, so a failing
search costs its lines and no longer the whole hook response.

The I/O is resolved from plugin_context at call time (_note_io), so the
existing patches keep working. The review re-run reads the AUTO_INJECT
spec instead of restating it, and its guard now compares the two live
searches. The registry declares the pipeline's notes fan-out sites.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-10-05 20:02:37 -04:00
co-authored by Claude Opus 5.5
parent 4b1060ae8b
commit 2b8f41229d
10 changed files with 1172 additions and 763 deletions
+149 -711
View File
@@ -17,7 +17,6 @@ from __future__ import annotations
import logging
import re
import textwrap
import time
from scribe.services import design_systems as design_systems_svc
@@ -35,7 +34,7 @@ from scribe.services.embeddings import (
semantic_search_rules,
)
from scribe.services import lesson_rules as lesson_rules_svc
from scribe.services.lessons import LESSON_NOTE_TYPE, claim_line
from scribe.services.lessons import LESSON_NOTE_TYPE
from scribe.services.note_usage import record_surfaced
from scribe.services.rule_usage import record_rule_surfaced
from scribe.services.supersession import superseded_ids
@@ -48,6 +47,8 @@ from scribe.services.retrieval_surfaces import (
from scribe.services import retrieval_pipeline as rp
from scribe.services.retrieval_pipeline import ( # noqa: F401 - re-exported
_RULEHINT_BAND, _rule_band, _rule_hint_line, checkpoint_for, checkpoint_reason,
VIA_LESSON_LIMIT, _menu_label, _menu_name, _menu_passage, _menu_seen_line,
_record_kind, menu_entry,
)
from scribe.services.retrieval_telemetry import record_retrieval
from scribe.services.settings import get_setting
@@ -60,74 +61,9 @@ logger = logging.getLogger(__name__)
# Defensive cap below Claude Code's 10k additionalContext limit.
_MAX_CHARS = 9000
# WHAT A MENU LINE CARRIES (#4364): the record's NAME, its kind and System,
# and the WHOLE passage that matched. Metadata plus the evidence, rather than a
# title asked to be both.
#
# The name, not the title. A snippet's or lesson's title is `name — when it
# applies` by construction (`embeddings.trigger_title`), because that join is
# what makes it rank on its situation. That is an EMBEDDING shape, and rendered
# as a menu line it ran to 1,500+ characters — the trigger paragraph spent
# again on every line, and again on every repeat. The trigger still arrives
# when it is what matched: it is in the chunk, and `_menu_passage` hands it
# over when the title was the whole match.
#
# The whole passage, not 200 characters of it. The search already chose the
# chunk that matched; the old cut kept its head and tail, and the head is the
# title every chunk is prefixed with — so the reader got the title twice and
# lost the middle, which is where the match was (lesson #4248). A chunk is at
# most ~1.4 KB (`embeddings._CHUNK_CHAR_BUDGET`), and it is shown once: a
# repeat is a one-line pointer (`_menu_seen_line`), not a second copy.
def _menu_name(title: str | None, note_type: str | None, data=None, body: str | None = "") -> str:
"""The record's name — its title without the trigger composed into it."""
title = (title or "(untitled)").replace("\n", " ").strip()
data = data if isinstance(data, dict) else {}
if note_type == "snippet":
from scribe.services.embeddings import TRIGGER_SEP
return (data.get("name") or title.partition(TRIGGER_SEP)[0]).strip() or title
if note_type == LESSON_NOTE_TYPE:
from types import SimpleNamespace
from scribe.services.embeddings import untrigger_title
from scribe.services.lessons import lesson_trigger
trigger = lesson_trigger(SimpleNamespace(data=data, body=body or ""))
# One line even when the stored name is a story (#4797).
return claim_line(untrigger_title(title, trigger).strip() or title)
return title
def _menu_passage(title: str | None, chunk_text: str | None, name: str = "") -> str:
"""The matched chunk on one line, without the title it was embedded under.
Every chunk is `title\nsection` (`embeddings.embedding_text`), and `title`
here must be the EMBEDDED one (`embeddings.document_title`) — for a snippet
or lesson that is `name — trigger`, not the stored name — so the prefix is
stripped exactly. A chunk that WAS only the title — a short
record, or the head chunk of one — matched on the title, and for a
trigger-keyed kind the part of it the name line no longer shows is the
trigger: that is returned, because it is precisely what matched.
One line, so the menu's blockquote survives it.
"""
title = (title or "").strip()
text = (chunk_text or "").strip()
if title and text.startswith(title):
text = text[len(title):]
text = " ".join(text.split())
if not text and name and title.startswith(name) and title != name:
text = " ".join(title[len(name):].lstrip(" —-").split())
return text
def _menu_label(kind: str, systems: list[str] | None) -> str:
"""`issue (done) · Plugin & hooks` — the kind, then where it belongs."""
return " · ".join([kind, *systems]) if systems else kind
def _menu_seen_line(note_id: int, kind: str, name: str) -> str:
"""A pointer to a record this session was already shown, not a copy of it."""
return f"> - #{note_id} [{kind} · seen] {name}"
# The notes renderer — a record's name, kind, System and matched passage —
# lives in `retrieval_pipeline` with the rest of the notes stages (milestone
# 456 step 4) and is re-exported above.
# How much of a named record's body sits under its line. The opening, not an
@@ -201,7 +137,7 @@ AUTOINJECT_DEFAULT_TOP_K = SURFACES["auto_inject"].budget_default
# Any two Python-shaped payloads share keywords, indentation and structure, so
# the floor for "some code" is ~0.55-0.63 — auto-inject's 0.55 lands INSIDE that
# noise band, and 6 of 8 probe payloads produced a nudge (4 of them noise). The
# margin gate can't rescue it either: _AUTOINJECT_BAND is relative to the top
# margin gate can't rescue it either: the notes band (retrieval_pipeline._NOTE_BAND) is relative to the top
# hit, so with a single hit it never engages.
#
# 0.68 clears every measured false positive with margin and still sits 0.05
@@ -591,9 +527,6 @@ _CONCEPT_MAX_DECLS = 4
# the code itself (e.g. all we found was `f()`), so we keep the raw payload.
_CONCEPT_MIN_CHARS = 16
# Margin gate: drop any hit more than this far below the top hit's score, so a
# single strong match doesn't drag in a wall of barely-passing neighbours.
_AUTOINJECT_BAND = 0.10
# Hard ceiling on top-k regardless of the user's setting — this is an
# awareness menu (titles only), never a content dump.
# The budget ceiling, now shared by every surface rather than owned by this
@@ -814,243 +747,6 @@ async def get_autoinject_config(user_id: int) -> dict:
}
def _record_kind(note) -> str:
"""The kind marker for an injected menu line — and, for a task, its status.
The menu is drawn from every record that carries an embedding, so a snippet,
a stored process, an issue and a stray dev-log all arrive looking identical.
Recorded prior art only stands out if the line says what it is — and the kind
is also what tells the reader which tool opens it.
Task-ness wins over `note_type` because it's the more useful distinction at a
glance: "there's an open issue about this" beats "there's a note about this".
A TASK ALSO CARRIES ITS STATUS, because for that kind alone the line is
read as a claim about live work. A finished step and an open one rendered
identically is not a cosmetic gap: a done step was cited as a milestone's
open one on the strength of a line exactly like this, which carries an id,
a kind and a title and said nothing about where the work stood (#4154).
Only for tasks — a note or a snippet has no status to be wrong about.
"""
if note.is_task:
kind = "issue" if note.task_kind == "issue" else "task"
# No fallback for a missing status: `is_task` IS `status is not None`
# (models/note.py), so a branch for a task without one could never be
# taken, and a dead branch is a claim about the data that isn't true.
#
# Parenthesised rather than dot-joined: the write-path prior-art line
# joins its own fields with " · ", so a dotted status would read as
# another flag beside `seen` instead of as part of the kind.
return f"{kind} ({note.status})"
return note.note_type or "note"
_REUSE_KINDS = ("snippet", "process")
async def _reserve_slot_for_reuse(
user_id: int,
query: str,
kept: list,
cfg: dict,
*,
project_id: int | None,
already: set[int],
) -> list:
"""Guarantee the reuse-shaped kinds one slot, if one clears threshold (#2246).
Ranking by raw cosine is blind to what KIND of record answers what kind of
ask, and the corpus makes that fatal rather than merely imperfect: Scribe's
project records are *about software work*, so a task titled "surface snippets
before the agent writes code" is a near-perfect lexical match for "write a
function…" while being useless as an answer to it. Measured live, a prompt
asking for a helper returned three records about BUILDING the retrieval
system and zero snippets.
The bias is structural and gets WORSE as the project record grows — which is
the direction Scribe is supposed to grow. Snippets are ~0.5% of the corpus
here; no threshold tuning fixes a 200:1 ratio.
So the reserved hit is deliberately NOT held to the margin band. The band
measures distance from the top overall score, and that top score is the very
thing snippets lose to. It still has to clear the configured threshold, so a
weak snippet cannot buy the slot — silence stays the default.
"""
if any(_record_kind(n) in _REUSE_KINDS for _s, n in kept):
return kept # reuse already represented; nothing to do
top_k = cfg["top_k"]
_t0 = time.perf_counter()
_rep: dict = {}
# THE LEDGER IS NOT AN EXCLUSION HERE EITHER (#4101) — but what is already
# in `kept` still is, and the two are different claims. A record sitting in
# this call's own menu must not be shown twice in it; a record shown in an
# EARLIER call is exactly what the slot should be allowed to spend itself
# on, because a snippet that was relevant then and is relevant now is the
# reuse case rather than a duplicate of it.
reuse = await semantic_search_notes(
user_id, query,
limit=1,
threshold=cfg["threshold"],
project_id=project_id,
exclude_ids={int(n.id) for _s, n in kept},
note_type=_REUSE_KINDS,
scope="browse",
report=_rep,
)
fresh_reuse = [(s, n) for s, n in reuse if int(n.id) not in already]
# A real semantic query competing for a menu slot — logged like the scored
# arm it displaces. Before this, the hit it PUSHED OUT was in
# retrieval_logs and the query that pushed it out was not, so the slot
# could never be evaluated against what it replaced (#2463; #1038 and
# #2085 are gated on this ledger being complete).
record_retrieval(
user_id=user_id, source="reuse_slot", query=query,
threshold=cfg["threshold"], limit=1, project_id=project_id,
is_task=None, results=fresh_reuse,
best_available=_rep.get("best_available_score"),
best_available_id=_rep.get("best_available_id"),
searched=bool(_rep.get("searched", True)),
suppressed=len(reuse) - len(fresh_reuse),
duration_ms=(time.perf_counter() - _t0) * 1000.0,
)
# Verify the kind rather than trusting the query that asked for it, and
# dedup against this call's own menu. This slot exists FOR reuse kinds — a
# slot silently spent on something else is worse than no slot, because the
# line is indistinguishable from one that earned its place on score.
kept_ids = {int(n.id) for _s, n in kept}
fresh = [
(s, n) for s, n in reuse
if _record_kind(n) in _REUSE_KINDS and int(n.id) not in kept_ids
][:1]
if not fresh:
return kept
# Take the LAST slot, never the first: the strongest overall hit is still the
# best answer to the prompt, and displacing it would trade one blindness for
# another.
if len(kept) >= top_k:
return kept[:top_k - 1] + fresh
return (kept + fresh)[:top_k]
async def _reserve_slot_for_lesson(
user_id: int,
query: str,
kept: list,
cfg: dict,
*,
project_id: int | None,
already: set[int],
) -> tuple[list, int | None]:
"""Guarantee a lesson one slot, if one clears the bar (milestone 385 step 5).
WHY A SLOT, AND WHAT IT ACTUALLY DISPLACES
The step's own framing was that "a slot spent on a lesson is a slot not
spent on a rule that binds". That is not what happens here, and the
correction matters for judging the cost: the notes menu and the rule hints
are separate functions with separate budgets, composed by the caller
(`build_prompt_rule_hint` says why). A line reserved in THIS menu displaces
a note, a snippet or an issue — never a rule.
THE ASYMMETRY, which is `preference_slot`'s argument on a different corpus:
- a NOTE crowded out of this menu is a lost convenience. It stays
searchable, and the operator can ask for it.
- a RULE crowded out still fires at an act arm. The prompt hit is a
preview of a second chance.
- a LESSON crowded out is the feature failing. A lesson exists only to be
met at the moment it applies — nobody browses lessons looking for one —
so the arm that surfaces it IS its delivery, and the loss is total and
silent. Silent delivery failure is the exact shape #3727 recorded: an
insight with no home arrived as a rule proposal instead.
And the ratio only moves one way. Lessons are by design rare and hard-won
while project records grow with the work, which is the 200:1 problem
`reuse_slot` was built for (#2246), before it has had a chance to be
measured here.
THE SLOT BUYS POSITION, NOT A LOWER BAR. It reserves at the menu's own
threshold, so a weak lesson cannot buy the line and silence stays the
default — the discipline both existing slots keep.
IT EXTENDS, IT NEVER DISPLACES, siding with `preference_slot` over
`reuse_slot`. Two reasons, and the second is the one that would be hard to
recover later: a displaced hit was returned by the general search and sits
in that call's `retrieval_logs` row, so evicting it makes the two tables
disagree about the same call for a reason nothing in the data explains
(#3668, and milestone #379 is what that costs). The first is voice — a
lesson does not bind, and a record that does not bind should not be able
to throw a better-scoring one off the menu.
Returns the possibly-extended list and the id the slot spent. The caller
needs that id to keep each source's surfaced set matching its own log row:
this slot records its own surfacing under its own name, so counting it
again under `auto_inject` would double it.
"""
if any(_record_kind(n) == LESSON_NOTE_TYPE for _s, n in kept):
return kept, None # a lesson already placed on score
_t0 = time.perf_counter()
_rep: dict = {}
# KIND-FILTERED, so the slot can only ever be spent on what it is for —
# `preference_slot`'s reasoning: verifying the kind after an open search
# would let a stray note buy the line, and that line would be
# indistinguishable from one that earned its place.
#
# `include_global_kinds` is the half that makes a lesson reachable at all
# from a project it was not written on, which is this kind's whole claim
# (#3730). Without it the slot would be a guarantee that silently only
# applies to lessons learned here.
found = await semantic_search_notes(
user_id, query,
limit=1,
threshold=cfg["threshold"],
project_id=project_id,
exclude_ids={int(n.id) for _s, n in kept},
note_type=(LESSON_NOTE_TYPE,),
include_global_kinds=True,
scope="browse",
report=_rep,
)
fresh = [(s, n) for s, n in found if int(n.id) not in already]
# ITS OWN SOURCE, from the first deploy. This slot is a claim that a kind
# deserves a guaranteed line, and a claim like that has to be falsifiable:
# `best_available_id` (#3807) names the lesson a bar refused, and the
# result count says how often the guarantee was actually spent. Without
# this row the question "does the lesson slot earn its line?" would have no
# data behind it in either direction — which is #2463's finding, recorded
# about the slot that shipped without one.
record_retrieval(
user_id=user_id, source="lesson_slot", query=query,
threshold=cfg["threshold"], limit=1, project_id=project_id,
is_task=None, results=fresh,
best_available=_rep.get("best_available_score"),
best_available_id=_rep.get("best_available_id"),
searched=bool(_rep.get("searched", True)),
suppressed=len(found) - len(fresh),
duration_ms=(time.perf_counter() - _t0) * 1000.0,
)
kept_ids = {int(n.id) for _s, n in kept}
slot = [
(s, n) for s, n in found
if _record_kind(n) == LESSON_NOTE_TYPE and int(n.id) not in kept_ids
][:1]
if not slot:
return kept, None
slot_id = int(slot[0][1].id)
# FRESH ONLY, matching the row above: a ledger repeat is rendered (#4101)
# but is not a new surfacing, so this source's two tables stay identical.
if slot_id not in already:
record_surfaced(
user_id=user_id, note_ids=[slot_id], source="lesson_slot",
project_id=project_id,
)
return kept + slot, slot_id
async def build_autoinject_hint(
user_id: int,
query: str,
@@ -1062,7 +758,7 @@ async def build_autoinject_hint(
The four anti-bloat gates (see the module + milestone-93 design):
1. high-confidence threshold (stricter than pull) — set per-user;
2. margin gate — keep only hits within _AUTOINJECT_BAND of the top score;
2. margin gate — keep only hits within retrieval_pipeline._NOTE_BAND of the top score;
3. session marking — caller passes already-injected ids as `exclude_ids`
and they are rendered again with `[seen]`, never withheld (#4101);
4. title-first payload — id + kind + title + score only, never bodies.
@@ -1107,83 +803,23 @@ async def build_autoinject_hint(
# before. A note line is a title and a score — the repeat costs about as
# much as the comma in this sentence — so there is nothing here to save.
already = {int(i) for i in (exclude_ids or [])}
t0 = time.perf_counter()
_rep_ai: dict = {}
hits = await semantic_search_notes(
user_id, q,
limit=cfg["top_k"],
threshold=cfg["threshold"],
project_id=(project_id or None),
# LESSONS ARE PROJECT-INDEPENDENT (#3730), so they join this menu's
# candidate set from wherever they were learned. Widening it is what
# makes the reserved slot below falsifiable rather than decorative: if
# the slot were the only path a lesson had, the general contest would
# be permanently closed to the kind and "the slot earns its line" would
# be true by construction. The switch adds nothing else — it ORs in
# `GLOBAL_NOTE_TYPES` and no other kind is in it.
include_global_kinds=True,
# Injection is the one retrieval nobody asked for, so it takes the BROWSE
# scope: never a record shared one-to-one with the operator. What can
# still appear is a collaborator's note inside a shared project — legible
# only because the line below names its owner.
scope="browse",
report=_rep_ai,
# The ranked half is `retrieval_pipeline.run_note_arm` (milestone 456):
# search, the fresh/repeat split, the call row, the band, the reuse and
# lesson slots in that order, and the surfacing rows are written there.
# What is decided HERE is what makes this moment this moment — the query
# the operator's words became, and the records they named by number.
result = await rp.run_note_arm(
rp.AUTO_INJECT,
rp.NoteMoment(
user_id=user_id, query=q, project_id=project_id,
seen=frozenset(already), named=frozenset(named_ids),
),
floor=cfg["threshold"], budget=cfg["top_k"], io=_note_io(),
)
# `results=fresh` and `suppressed` together, on the rule arms' contract
# (#3752): a rendered repeat is not a new surfacing, so it stays out of the
# row's result set and is COUNTED instead. That keeps this source's surfaced
# set identical to its own log row (#3668) while making a zero-result call
# readable — `result_count == 0` with `suppressed_count > 0` is "everything
# that matched, this session has already seen", which is a different fact
# about the bar from "nothing cleared it" and used to be unreportable here.
# A record the operator NAMED is counted the same way: it is shown, but by
# the lookup above and booked under `named_ref`, so this arm did not
# surface it and must not claim it.
fresh = [
(s, n) for s, n in hits
if int(n.id) not in already and int(n.id) not in named_ids
]
record_retrieval(
user_id=user_id, source="auto_inject", query=q,
threshold=cfg["threshold"], limit=cfg["top_k"],
project_id=(project_id or None), is_task=None, results=fresh,
best_available=_rep_ai.get("best_available_score"),
best_available_id=_rep_ai.get("best_available_id"),
searched=bool(_rep_ai.get("searched", True)),
suppressed=len(hits) - len(fresh),
duration_ms=(time.perf_counter() - t0) * 1000.0,
)
if not hits and not named:
kept = result.menu
if not kept and not named:
return empty
kept: list = []
lesson_slot_id = None
if hits:
# Margin gate: keep only hits close to the strongest one. Computed over
# ALL hits, repeats included — the band measures distance from the top
# SCORE, and letting the ledger move that cutoff would make "you were
# shown this" change what counts as relevant, which is the axis
# independence the rule band keeps for the same reason.
top_score = hits[0][0]
kept = [(s, n) for s, n in hits if s >= top_score - _AUTOINJECT_BAND]
# The slots are told about the named records as if already shown, so a
# slot that picks one does not book a surfacing the lookup booked too.
shown = already | set(named_ids)
kept = await _reserve_slot_for_reuse(
user_id, q, kept, cfg, project_id=(project_id or None),
already=shown,
)
# AFTER the reuse slot, because that one evicts the menu's weakest hit
# while this one extends: running them the other way round would let a
# reserved lesson be the line reuse throws off, and a slot that another
# slot can silently undo is not a guarantee.
kept, lesson_slot_id = await _reserve_slot_for_lesson(
user_id, q, kept, cfg, project_id=(project_id or None),
already=shown,
)
# A named record is shown once, in its own block, never again as a match.
kept = [(s, n) for s, n in kept if int(n.id) not in named_ids]
# A collaborator's note can reach this menu via a shared project, and the
# operator never asked for it — so say whose it is. Unattributed, it reads as
# something they wrote and settled.
@@ -1192,16 +828,19 @@ async def build_autoinject_hint(
if n.user_id != user_id
})
# A superseded record is DEMOTED, not removed (#278) — so one can still reach
# this menu, and when it does the reader has to be told. An agent handed
# stale material with nothing marking it acts on it with full confidence,
# which is worse than never having surfaced it. One query for the whole menu.
# A superseded record is DEMOTED, not removed (#278) — `menu_entry` says
# so on the line. One query for the whole menu.
stale = await superseded_ids([*named_ids, *(int(n.id) for _s, n in kept)])
systems = await system_names_for(
{i for i in named_ids if i not in already}
| {int(n.id) for _s, n in kept if int(n.id) not in already}
)
def _shared_by(note) -> str:
if note.user_id == user_id:
return ""
return owners.get(int(note.user_id)) or "another user"
lines: list[str] = []
note_ids: list[int] = []
# FIRST, because the operator said which record they meant and everything
@@ -1217,24 +856,15 @@ async def build_autoinject_hint(
for note in named:
nid = int(note.id)
note_ids.append(nid)
kind = _record_kind(note)
name = _menu_name(note.title, note.note_type, note.data, note.body)
if nid in already:
line = _menu_seen_line(nid, kind, name)
if nid in stale:
line += " — SUPERSEDED"
lines.append(line)
continue
line = f"> - #{nid} [{_menu_label(kind, systems.get(nid))}] \"{name}\""
if nid in stale:
line += " — SUPERSEDED, a later record covers this; check that first"
if note.user_id != user_id:
who = owners.get(int(note.user_id)) or "another user"
line += f" — shared by {who}"
lines.append(line)
opening = _named_opening(note.body)
if opening:
lines.append(f"> ↳ {opening}")
# The OPENING under the line, not a passage: nothing matched here, so
# there is no passage to keep, and what the reader needs is what the
# record IS — which is where a record says it.
lines.extend(menu_entry(
nid, kind=_record_kind(note),
name=_menu_name(note.title, note.note_type, note.data, note.body),
systems=systems.get(nid), seen=nid in already, stale=nid in stale,
shared_by=_shared_by(note), under=_named_opening(note.body),
))
# "records", not "notes" — the menu can hold snippets, processes and tasks
# too, and the kind marker on each line is only legible if the header doesn't
@@ -1275,78 +905,35 @@ async def build_autoinject_hint(
"what you are doing and use your judgement — a lesson is not a "
"rule and binds nothing."
)
# From THIS arm's own search (`_rep_ai`), so a chunk is only ever paired
# with the query that actually matched it.
menu_chunks = _rep_ai.get("best_chunk") or {}
menu_ids: list[int] = []
for score, note in kept:
nid = int(note.id)
note_ids.append(nid)
menu_ids.append(nid)
kind = _record_kind(note)
# The NAME, not the title (#4364): a snippet's or lesson's title is its
# embedding shape, trigger and all, and ran past 1,500 characters here.
name = _menu_name(note.title, note.note_type, note.data, note.body)
if nid in already:
# A POINTER, not a copy (#4364). The record is in this session's
# context already — the ledger is cleared at compaction, so "seen"
# stays true — and re-rendering it spent its whole line again for
# nothing. What the reader needs is the reminder that it matched
# again, and the id to open it if it has scrolled out of mind.
line = _menu_seen_line(nid, kind, name)
if nid in stale:
line += " — SUPERSEDED"
lines.append(line)
continue
line = f"> - #{nid} [{_menu_label(kind, systems.get(nid))}] \"{name}\" ({score:.2f})"
if nid in stale:
line += " — SUPERSEDED, a later record covers this; check that first"
if note.user_id != user_id:
who = owners.get(int(note.user_id)) or "another user"
line += f" — shared by {who}, treat as a suggestion"
lines.append(line)
# The passage that earned the line, WHOLE, indented under it (#4364).
# Absent when the record has no stored chunk — an un-embedded row, or
# the reserved lesson and reuse slots, which are fetched by their own
# queries and so are not in this search's report. No fallback to the
# body's opening: on a menu that would be a line of preamble dressed as
# a reason, and a reader cannot tell the two apart once indented alike.
passage = _menu_passage(
passage = "" if nid in already else _menu_passage(
document_title(note.title, note.note_type, note.data, note.body),
(menu_chunks.get(nid) or {}).get("text"), name,
(result.chunks.get(nid) or {}).get("text"), name,
)
if passage:
lines.append(f"> ↳ {passage}")
who = _shared_by(note)
lines.extend(menu_entry(
nid, kind=_record_kind(note), name=name, systems=systems.get(nid),
seen=nid in already, stale=nid in stale, score=score,
shared_by=f"{who}, treat as a suggestion" if who else "",
under=passage,
))
# Records what SURVIVED the margin gate, not what the ranker returned — the
# menu the agent actually saw. retrieval_logs already holds the full
# candidate set for threshold tuning; conflating the two would make
# "surfaced" mean two different things depending on the surface (#2085).
#
# FRESH ONLY, which is the same cut the log row above takes (#4101). A
# repeat is rendered but is not a new surfacing, and counting it again would
# make this table disagree with `retrieval_logs` about the same call —
# #3668's identity, which is the cheapest true statement available about
# this pair of tables and is not worth a marker's convenience.
# THE RESERVED LESSON IS NOT THIS ARM'S SURFACING. It was fetched by its own
# query and already recorded under `lesson_slot`, so counting it here would
# book one delivery twice and leave `lesson_slot`'s two tables describing
# different numbers of the same event — #3668's identity, which is the
# cheapest true statement available about this pair of tables. It stays in
# `note_ids`, which is the session LEDGER and must list every line rendered.
record_surfaced(
user_id=user_id,
note_ids=[
i for i in menu_ids
if i not in already and i != lesson_slot_id
],
source="auto_inject",
project_id=project_id,
)
# A lookup, not a ranking, so it writes no retrieval_logs row — there is no
# score or bar to tune — and its surfacings are its own source, which is
# what lets pull-through say whether a named record gets opened.
# what lets pull-through say whether a named record gets opened. The
# ranked menu's own rows were written by the pipeline.
named_fresh = [i for i in named_ids if i not in already]
if named_fresh:
record_surfaced(
@@ -1369,97 +956,46 @@ async def build_autoinject_hint(
}
# ── A rule reached through its lessons (milestone 440, #4633) ──────────────
#
# A rule's own document is written in the rule's words, which are general by
# design; the situations that keep proving it are often closer to what a
# session is actually doing. A lesson JUDGED to be an instance of a rule (a
# CONFIRMED link) carries that situation, so a lesson matching the moment
# brings its rule along — in rule voice, naming the lesson that reached it.
#
# ITS OWN SLOT, NOT A RULE SLOT, as a stated default. The plan asked for this
# to be settled by measurement, and there is nothing to measure until links
# are confirmed (#4632). Until then the asymmetry decides it: a via-lesson
# line that took a rule slot could push out a rule the ranker matched
# DIRECTLY, a stronger claim displaced by a weaker one, while an extra line
# costs one line. `rule_via_lesson` is logged as its own source so the
# question can be answered from data later (#4636).
VIA_LESSON_LIMIT = 1
# Lessons fetched before keeping only the linked ones. The search cannot be
# told "linked lessons only", so it overfetches and filters; on a corpus where
# most lessons are unlinked, a fetch of one would almost always be spent on a
# lesson that carries nothing.
_VIA_LESSON_OVERFETCH = 10
def _note_io() -> rp.NoteIO:
"""The ranker and recorders the notes arms report to, read at call time —
`_rule_io`'s reason: whatever this module's names are when the arm runs."""
return rp.NoteIO(
search=semantic_search_notes,
record_retrieval=record_retrieval,
record_surfaced=record_surfaced,
)
async def _rules_via_lessons(
user_id: int, query: str, *, project_id: int | None, skip: set[int],
held: set[int], where: str,
) -> tuple[list[str], list[int]]:
"""Rule lines reached through a matching lesson's CONFIRMED links.
"""Rule lines reached through a matching lesson's CONFIRMED links — the
pipeline's via-lesson arm (`retrieval_pipeline.run_via_lesson_arm`, where
the design is written), with its I/O read from this module at call time.
The lesson bar is the notes menu's own threshold — the bar a lesson has to
clear to be shown at all — so a lesson too weak to surface cannot carry a
rule in. `skip` is every rule this response already names plus the
session ledger: suppression applies to the RULE, whichever lesson reached
it. Returns (lines, rule ids shown). Fails open, like every arm.
The lesson bar is the notes menu's own threshold. Returns (lines, rule ids
shown).
"""
try:
confirmed = await lesson_rules_svc.confirmed_lessons(user_id)
if not confirmed:
return [], []
bar = (await get_autoinject_config(user_id))["threshold"]
t0 = time.perf_counter()
_rep: dict = {}
found = await semantic_search_notes(
user_id, query, limit=_VIA_LESSON_OVERFETCH, threshold=bar,
project_id=project_id, note_type=(LESSON_NOTE_TYPE,),
include_global_kinds=True, scope="browse", report=_rep,
)
matched = [(s, n) for s, n in found if int(n.id) in confirmed]
by_lesson = await lesson_rules_svc.confirmed_rules_in_scope(
user_id, [int(n.id) for _s, n in matched], project_id,
)
candidates = [
(s, rule, n) for s, n in matched for rule in by_lesson.get(int(n.id), [])
]
chosen: list = []
taken: set[int] = set(skip)
for s, rule, n in candidates:
if rule.id in taken:
continue
taken.add(rule.id)
chosen.append((s, rule, n))
if len(chosen) >= VIA_LESSON_LIMIT:
break
# Logged whenever a search ran, results or not — the #3497 guard. The
# score is the LESSON's, because the lesson is what was matched; the
# best-available id is left off for the same reason, since the row's
# results are rules and an id beside them would read as a rule id.
record_retrieval(
user_id=user_id, source="rule_via_lesson", query=query,
threshold=bar, limit=VIA_LESSON_LIMIT, project_id=project_id,
is_task=None, results=[(s, rule) for s, rule, _n in chosen],
duration_ms=(time.perf_counter() - t0) * 1000.0,
best_available=_rep.get("best_available_score"),
searched=bool(_rep.get("searched", True)),
suppressed=len({r.id for _s, r, _n in candidates}) - len(chosen),
)
if not chosen:
return [], []
rule_ids = [rule.id for _s, rule, _n in chosen]
record_rule_surfaced(user_id=user_id, rule_ids=rule_ids, source="rule_via_lesson")
lines = [
_rule_hint_line(rule, where=where, seen=False, held=rule.id in held)
+ f" Reached through lesson #{n.id} "
+ f"\u201c{_menu_name(n.title, n.note_type, n.data, n.body)}\u201d, "
+ "a recorded instance of it."
for _s, rule, n in chosen
]
return lines, rule_ids
except Exception:
logger.debug("rule-via-lesson arm failed", exc_info=True)
return [], []
async def bar() -> float:
return (await get_autoinject_config(user_id))["threshold"]
result = await rp.run_via_lesson_arm(
rp.RuleMoment(
user_id=user_id, query=query, project_id=project_id, where=where,
held=frozenset(held),
),
skip=frozenset(skip),
io=rp.ViaLessonIO(
search=semantic_search_notes,
linked=lesson_rules_svc.confirmed_lessons,
rules_for=lesson_rules_svc.confirmed_rules_in_scope,
floor=bar,
record_retrieval=record_retrieval,
record_rule_surfaced=record_rule_surfaced,
),
)
return result.lines, result.rule_ids
async def _add_rules_via_lessons(
@@ -2026,171 +1562,73 @@ async def build_write_path_hint(
# empty mapping means every line falls back to its title alone.
wp_chunks: dict[int, dict] = {}
if remaining > 0 and query:
t0 = time.perf_counter()
# Pulled-and-already-listed ids stay in the query (as evidence for
# `resembles`) but never in the menu, so the limit has to cover them.
# `seen` is this call's own menu now, which is the only thing left that
# is a reason to withhold (#4101).
in_menu = seen
pulled_in_menu = in_menu & set(pulled)
_rep_wp: dict = {}
hits = await semantic_search_notes(
user_id, query,
limit=remaining + len(pulled_in_menu),
threshold=cfg["threshold"],
project_id=scope_project,
exclude_ids=in_menu - set(pulled),
# Snippets AND recorded experience (#2246). This arm was
# snippets-only, which is auto-inject's mistake inverted: an issue
# saying "we tried this and it deadlocked", or a dev-log recording
# how a problem was solved, is prior art for the code about to be
# written — arguably better prior art than a resembling helper,
# because it says what NOT to do.
#
# `task_kind="issue"` keeps the open to-do list out. A task titled
# "add debouncing to the search box" resembles the code being
# written and answers nothing; an ISSUE is corrective work with a
# root cause in it, and a non-task note is durable knowledge. Both
# earned their place; a todo did not.
#
# AND LESSONS (milestone 385 step 5). This arm is kind-FILTERED, so
# a kind absent from this tuple is not merely outranked here — it
# is unreachable, and nothing reports an arm that never had the
# candidate (#3702). The founding example of the kind is a lesson
# about a code shape ("two absolutely-positioned siblings"), which
# is the moment this arm fires and no other.
#
# NO RESERVED SLOT HERE, unlike the prompt menu. This arm fires
# before EVERY Write and Edit, where a guaranteed extra line is a
# guaranteed extra interruption per keystroke-batch — the same
# argument that keeps the act arms' budgets tight. And the contest
# is already fair: the field is snippets, issues and lessons rather
# than the whole corpus, so the 200:1 dilution a slot answers is
# not what happens here. Lessons surfaced by this arm are
# identifiable in the telemetry by their kind.
note_type=("snippet", "note", LESSON_NOTE_TYPE),
task_kind="issue",
# Project-independent, for the reason the prompt menu passes it:
# a lesson's claim is that it transfers, and an arm scoped to the
# project it was written on cannot test that claim (#3730).
include_global_kinds=True,
# Same reasoning as auto-inject: nobody asked for this, so it takes
# the browse scope and never surfaces a one-to-one direct share.
scope="browse",
report=_rep_wp,
# The ranked half is `retrieval_pipeline.run_note_arm` with the
# WRITE_PATH spec (milestone 456): which kinds it asks for and why,
# the withheld-menu accounting (#3739), the fresh/repeat split, the
# call row and the surfacing rows are written there. What is decided
# HERE is the menu above it — what place and sync already listed is
# withheld, and a PULLED snippet among those is still scored, because
# its resemblance to this payload is the stamping feed's evidence.
result = await rp.run_note_arm(
rp.WRITE_PATH,
rp.NoteMoment(
user_id=user_id, query=query, project_id=project_id,
seen=frozenset(excluded), in_menu=frozenset(seen),
still_scored=frozenset(pulled),
),
floor=cfg["threshold"], budget=remaining, io=_note_io(),
)
wp_chunks = _rep_wp.get("best_chunk") or {}
wp_chunks = result.chunks
resembles = {
int(note.id): float(score) for score, note in hits
int(note.id): float(score) for score, note in result.answered
if int(note.id) in pulled
}
shown = [(s, n) for s, n in hits if int(n.id) not in in_menu]
# WHAT THIS ARM WITHHELD AFTER THE SEARCH ANSWERED, and the reason
# `best_available_score` cannot always be reported here (#3739 again,
# from the side its fix did not reach).
#
# This arm is the one note arm that filters TWICE. `exclude_ids` takes
# the already-listed ids into the search, but the PULLED ones among them
# stay in the query deliberately — `resembles` above needs them — and
# are dropped in the line above instead. So the score the search
# reported is PRE that drop while the row's `result_count` is POST it,
# and a record already listed in this same menu could be logged as
# something the BAR turned away. Live proof on the first read after
# #3739 shipped: write_path's near-miss max was 0.822 while the lowest
# score it ever RETURNED was 0.6857 — a "rejection" that beat every
# acceptance.
#
# So the honest answer is null — "not measured on this call" — whenever
# this filter removed anything, because then the bar is not the only
# thing that turned something away and the reported score may belong to
# a record we withheld ourselves. Calls where nothing was dropped keep
# reporting it, which is most of them.
#
# THE LEDGER IS NO LONGER PART OF THIS (#4101), and that is why the
# suppression column can now be filled in where it could not before.
# The old objection was that a count here would be PARTIAL — covering
# the drops made in this function but not the ones `exclude_ids` made
# inside the search — and a partial number under a name that reads as
# complete is the substitution this milestone exists to stop. That was
# right while the ledger was one of the things `exclude_ids` carried.
# It no longer is: every ledger repeat comes back from the search and is
# rendered, so `suppressed` counts all of them and none are hidden
# inside the query. What `exclude_ids` still removes is this call's own
# menu, which is not suppression at all — those records ARE being shown,
# one block further up.
withheld_here = len(hits) - len(shown)
hits = shown[:remaining]
fresh = [(s, n) for s, n in hits if int(n.id) not in excluded]
record_retrieval(
user_id=user_id, source="write_path", query=query,
threshold=cfg["threshold"], limit=remaining,
# is_task is None, not False: this arm now returns issues too, and
# recording it as a notes-only retrieval would misdescribe the
# candidate set the threshold is being tuned against.
project_id=scope_project, is_task=None, results=fresh,
best_available=(
None if withheld_here else _rep_wp.get("best_available_score")
),
# Withheld on the SAME condition as the score. A surviving id
# beside a null score would name a record without saying what it
# scored, which is the pair disagreeing in the other direction.
best_available_id=(
None if withheld_here else _rep_wp.get("best_available_id")
),
searched=bool(_rep_wp.get("searched", True)),
suppressed=len(hits) - len(fresh),
duration_ms=(time.perf_counter() - t0) * 1000.0,
)
if hits:
top_score = hits[0][0]
for score, note in hits:
if score < top_score - _AUTOINJECT_BAND:
continue
# Name the kind unless it's a snippet — the menu's default and
# the header's default reading. An issue or a dev-log offered
# here is a different KIND of claim ("this was already tried")
# and an unlabelled line would be read as "here is code to
# reuse", which is the opposite of what it says.
kind = _record_kind(note)
marker = (
f"similar {score:.2f}" if kind == "snippet"
else f"similar {score:.2f} · {kind}"
)
# The repeat marker rides the same dotted list as the kind, so a
# reference costs four characters and needs no second line
# (#4101). Same word as the auto-inject menu deliberately: a
# reader meeting `seen` on two different surfaces should not
# have to work out whether they mean the same thing.
if int(note.id) in excluded:
marker += " · seen"
scored.append((
marker,
{
"id": int(note.id), "title": note.title, "user_id": note.user_id,
# The name the line shows, and whether this session has
# it already — carried as data for the reason `kind` is
# (#4364): the line is built from facts, not from
# re-reading its own marker.
"name": _menu_name(note.title, note.note_type, note.data, note.body),
# What its chunks are prefixed with, for stripping.
"doc_title": document_title(
note.title, note.note_type, note.data, note.body,
),
"seen": int(note.id) in excluded,
# Carried, not re-read off the rendered marker. The
# marker is prose assembled for a human and it already
# varies by kind, language and the `seen` flag — a
# header that decided what to say by matching substrings
# in it would break the next time a marker is reworded,
# silently and in the direction of saying nothing.
"kind": kind,
# Carried so the line can disclose a cross-language hit
# (#2244). The semantic arm is where these actually arise —
# a snippet recorded at the path you're editing is almost
# never in another language, but a concept match easily is.
"language": (note.data or {}).get("language") if note.data else None,
},
))
for score, note in result.menu:
# Name the kind unless it's a snippet — the menu's default and
# the header's default reading. An issue or a dev-log offered
# here is a different KIND of claim ("this was already tried")
# and an unlabelled line would be read as "here is code to
# reuse", which is the opposite of what it says.
kind = _record_kind(note)
marker = (
f"similar {score:.2f}" if kind == "snippet"
else f"similar {score:.2f} · {kind}"
)
# The repeat marker rides the same dotted list as the kind, so a
# reference costs four characters and needs no second line
# (#4101). Same word as the auto-inject menu deliberately: a
# reader meeting `seen` on two different surfaces should not
# have to work out whether they mean the same thing.
if int(note.id) in excluded:
marker += " · seen"
scored.append((
marker,
{
"id": int(note.id), "title": note.title, "user_id": note.user_id,
# The name the line shows, and whether this session has
# it already — carried as data for the reason `kind` is
# (#4364): the line is built from facts, not from
# re-reading its own marker.
"name": _menu_name(note.title, note.note_type, note.data, note.body),
# What its chunks are prefixed with, for stripping.
"doc_title": document_title(
note.title, note.note_type, note.data, note.body,
),
"seen": int(note.id) in excluded,
# Carried, not re-read off the rendered marker. The
# marker is prose assembled for a human and it already
# varies by kind, language and the `seen` flag — a
# header that decided what to say by matching substrings
# in it would break the next time a marker is reworded,
# silently and in the direction of saying nothing.
"kind": kind,
# Carried so the line can disclose a cross-language hit
# (#2244). The semantic arm is where these actually arise —
# a snippet recorded at the path you're editing is almost
# never in another language, but a concept match easily is.
"language": (note.data or {}).get("language") if note.data else None,
},
))
menu = (placed + scored)[:max(0, top_k - len(synced))]
@@ -2403,12 +1841,12 @@ async def build_write_path_hint(
for marker, item in menu:
# A rendered repeat is not a new surfacing (#4101) — same cut the log
# row takes, so this table and `retrieval_logs` keep agreeing about the
# same call (#3668).
if int(item["id"]) in excluded:
# same call (#3668). The semantic arm's rows (`write_path_semantic`)
# were written by the pipeline beside its call row, so only the
# lookups are booked here.
if int(item["id"]) in excluded or not marker.startswith("nearby"):
continue
arm = ("write_path_place" if marker.startswith("nearby")
else "write_path_semantic")
by_arm.setdefault(arm, []).append(int(item["id"]))
by_arm.setdefault("write_path_place", []).append(int(item["id"]))
for arm, ids in by_arm.items():
record_surfaced(
user_id=user_id, note_ids=ids, source=arm, project_id=project_id,