feat(rules): rules retrieve against the operator's message (#3852)
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m35s
CI & Build / Build & push image (push) Successful in 38s

The third rule arm, and the one the other two cannot reach. `write_path_rule`
is keyed on code, `pre_tool_rule` on a command — both things the session is
about to DO. A rule that governs what to SAY has no such trigger: extract
intent from loose phrasing, raise a conflict before acting, hand off an
action with its reason, end a finding with an offer all bind on a RESPONSE,
and no tool call precedes one.

The operator's message is the only query that exists before a response is
composed. That hook searched notes alone, so no rule had ever been retrieved
against a thing the operator actually said — and residency was the only
surface those rules had, which is what milestone 394 removes.

A SEPARATE FUNCTION, not a branch in build_autoinject_hint, because of its
early returns. That arm bails when auto-inject is disabled, when the query is
blank, when nothing clears the note bar — every one a statement about NOTES.
Folded in, an operator who turned the awareness menu off would silently lose
their rules, a coupling with no symptom since both look like a quiet hook.
Two functions, two sets of gates, composed in the route. Guarded as "the rule
arm never asks the notes arm's config", which is the structural fact.

Joins _ARMS rather than getting its own test file. #3497's history is that
the pre-tool arm inherited a defect from its sibling by being MODELLED on it
instead of sharing with it, and a third arm modelled on two is two chances to
repeat that. Repeat rendering, fresh-only counting, log-before-bailout, the
kind register and the two-recorders identity are properties of every arm or
of none.

The bar is INHERITED and says so. 0.72 was tuned against code and commands;
prose is a different query shape against the same documents, and triggers are
written in the vocabulary of the moment — which for most rules is act
vocabulary. Starting at the only number with evidence behind it and logging
every call from the first deploy is what makes it settleable; guessing lower
would put an unmeasured bar in front of a corpus that binds.

k=3, anchored on this hook's own budget rather than the act arms'.
RULEHINT_LIMIT is 1 because that arm fires before every Bash call; this one
fires once per turn, beside a notes menu already spending three slots. And a
prompt genuinely contains more than one act — "merge to main and then start
on X" is two — where a command is one thing.

`prompt_rule` added to RANKED_SOURCES: a ranker picked it, and a ranked
source missing from that tuple is silently counted as bulk delivery and drops
out of the pull-through denominator.

The hook reads and writes the SHARED rule ledger under scribe-priorart, not a
private one — one session keeps one list, aged (#3751), so a rule named here
is not re-announced before the next Bash call.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011cPyzNnegXHr5iRMzzy5KJ
This commit is contained in:
2026-09-11 07:56:19 -04:00
co-authored by Claude Opus 5
parent 8406871085
commit 44e0b0541f
6 changed files with 376 additions and 10 deletions
+24
View File
@@ -91,13 +91,37 @@ async def autoinject_retrieve():
project_id (opt) — explicit project scope override (ad-hoc/testing).
exclude_ids (opt) — comma-separated note ids already injected this
session; skipped so each note injects at most once.
exclude_rule_ids — comma-separated rule ids already surfaced this
(opt) session. SHARED with /prior-art and /tool-rules on
purpose: one session keeps ONE rule ledger, so a
rule named by any arm is not re-announced by
another. Ages out (#3751), so salience decays.
TWO ARMS, TWO SETS OF GATES. Rules ride the same hook and the same query
but nothing else: the notes menu can be disabled, thresholded and top-k'd
by the operator without touching whether a rule reaches them. Composed
here rather than inside either builder so neither one's early return can
silently suppress the other.
Rules come FIRST in the payload. A rule or preference governing the answer
is more consequential than a menu of things that might be worth reading,
and a reader who stops after the first block should have stopped after
the right one.
"""
q = (request.args.get("q") or "").strip()
project_id, _repo, _unbound = await _project_scope()
exclude_ids = _int_list(request.args.get("exclude_ids"))
exclude_rule_ids = _int_list(request.args.get("exclude_rule_ids"))
rules = await plugin_ctx_svc.build_prompt_rule_hint(
g.user.id, q, project_id=project_id, exclude_rule_ids=exclude_rule_ids
)
result = await plugin_ctx_svc.build_autoinject_hint(
g.user.id, q, project_id=project_id, exclude_ids=exclude_ids
)
blocks = [b for b in (rules["context"], result["context"]) if b]
result["context"] = "\n\n".join(blocks)
result["rule_ids"] = rules["rule_ids"]
return jsonify(result)
+145
View File
@@ -314,6 +314,42 @@ _AUTOINJECT_BAND = 0.10
# awareness menu (titles only), never a content dump.
_AUTOINJECT_MAX_TOP_K = 10
# --- the prompt-boundary rule arm (#3852) ------------------------------------
#
# Both existing rule arms are keyed on something the session is about to DO —
# a file write, a command. A rule that governs what to SAY has no such moment.
# Extract intent from loose phrasing, raise a conflict before acting, hand off
# an action with its reason, end a finding with an offer: every one binds on a
# RESPONSE, and no tool call precedes a response.
#
# The operator's message is the only query that exists before one is composed,
# and this arm is what runs against it. Until now that hook searched notes
# alone, so no rule had ever been retrieved against a thing the operator said.
PROMPTRULE_THRESHOLD_KEY = "kb_promptrule_threshold"
# INHERITED FROM THE ACT ARMS, AND NOT YET EARNED HERE. 0.72 was tuned against
# code and shell commands. An operator's prose is a different query shape
# against the same documents, and nothing yet says the two distributions line
# up — triggers are written in the vocabulary of the MOMENT, which for most
# rules is act vocabulary, so prose may well score lower across the board.
#
# Starting at the act arms' number anyway is deliberate: it is the only value
# with evidence behind it, and guessing lower would put an unmeasured bar in
# front of a corpus that binds. Every call is logged under `prompt_rule` from
# the first deploy, so a few days of real traffic settles it — read
# `near_miss_samples` (#3807) before moving this, not the percentile alone.
PROMPTRULE_DEFAULT_THRESHOLD = 0.72
# MORE THAN THE ACT ARMS' SINGLE SLOT, anchored on this hook's budget rather
# than theirs. RULEHINT_LIMIT is 1 because that arm fires before EVERY Bash
# call, where a second line is a second interruption per command. This arm
# fires once per TURN, on the same hook whose notes menu already spends
# AUTOINJECT_DEFAULT_TOP_K slots — so that is the comparable budget.
#
# And a prompt genuinely contains more than one act. "Merge to main and then
# start on X" is two, governed by different rules; k=1 cannot serve that case
# at all, where the act arms never face it because a command is one thing.
PROMPTRULE_LIMIT = 3
def _slugify(text: str) -> str:
"""kebab-case slug for a skill directory name (a-z0-9 + single hyphens)."""
@@ -642,6 +678,115 @@ async def build_autoinject_hint(
return {"context": "\n".join(lines), "note_ids": note_ids, "config": cfg}
async def build_prompt_rule_hint(
user_id: int,
query: str,
*,
project_id: int = 0,
exclude_rule_ids: list[int] | None = None,
) -> dict:
"""Rules and preferences that may apply to what the operator just asked.
The third rule arm, and the one that closes a gap the other two cannot
reach. `write_path_rule` is keyed on code, `pre_tool_rule` on a command —
both are things the session is about to DO. A rule that governs what to
SAY has no such trigger, and residency was the only surface it ever had.
Removing residency (milestone 394) without this would drop that half of
the corpus on the floor.
A SEPARATE FUNCTION, not a branch inside build_autoinject_hint, and the
reason is its early returns. That arm bails when auto-inject is disabled,
when the query is blank, when nothing clears the note bar — and every one
of those is a statement about NOTES. Folded in, a user who turned the
notes menu off would silently lose their rules too, which is the kind of
coupling nothing downstream could see. Two functions, two sets of gates,
composed by the caller.
THE OUTPUT IS DELIBERATELY NOT QUOTED, where the notes menu is. The task
asked whether the two share a header; the answer is that neither needs
one. A note line is a bare title and needs the menu's header to say what
it is doing there, while a rule line names itself in its opening words
("Standing rule that may apply…" / "Preference that may apply…"). Leaving
rules unquoted separates the two claims visually with no extra prose, and
matches how a rule line already renders on both act arms.
Fails open and returns empty context on any error, like its siblings: a
recall aid may never break the operator's prompt.
"""
out: dict = {"context": "", "rule_ids": []}
q = (query or "").strip()
if not q:
return out
try:
try:
threshold = float(await get_setting(
user_id, PROMPTRULE_THRESHOLD_KEY,
str(PROMPTRULE_DEFAULT_THRESHOLD)))
except (TypeError, ValueError):
threshold = PROMPTRULE_DEFAULT_THRESHOLD
threshold = min(1.0, max(0.0, threshold))
t0 = time.perf_counter()
_rep: dict = {}
# NOT scoped to the project, and that is the corpus's own decision
# rather than an omission here — semantic_search_rules is scoped by
# OWNERSHIP on purpose, because "is there a rule about this" is asked
# across a whole rulebook. `project_id` below reaches the log row and
# nothing else.
hits = await semantic_search_rules(
user_id, q, limit=PROMPTRULE_LIMIT, threshold=threshold,
report=_rep,
)
duration_ms = (time.perf_counter() - t0) * 1000.0
already = set(exclude_rule_ids or [])
fresh = [(score, rule) for score, rule in hits if rule.id not in already]
# BEFORE the early return, for the reason both sibling arms spell out
# at length: a call that found nothing is the only evidence a bar is
# too high, and an arm that logs only the calls it liked reports a
# flawless clear-rate however badly it is tuned. This bar is inherited
# and unverified for this corpus, so the zero rows are the point.
record_retrieval(
user_id=user_id, source="prompt_rule", query=q,
threshold=threshold, limit=PROMPTRULE_LIMIT,
project_id=project_id,
is_task=None, results=fresh, duration_ms=duration_ms,
best_available=_rep.get("best_available_score"),
best_available_id=_rep.get("best_available_id"),
searched=bool(_rep.get("searched", True)),
suppressed=len(hits) - len(fresh),
)
# `hits`, not `fresh` (#3750): a call whose only hit is a repeat still
# has something to say, it just says it differently.
if not hits:
return out
lines = [
_rule_hint_line(rule, where="to this request", seen=rule.id in already)
for _score, rule in hits
]
# FRESH-ONLY (#3752). A reference is a rendering decision, not a
# retrieval outcome, and counting one here would inflate the
# denominator pull_through is read from.
rule_ids = [rule.id for _score, rule in fresh]
# RANKED, not ambient: this arm chose what it showed. The name is also
# in `rule_usage.RANKED_SOURCES`, and it has to be — a ranked source
# missing from that tuple is counted as a bulk delivery nobody decided
# on, which silently moves it out of the pull-through denominator.
if rule_ids:
record_rule_surfaced(
user_id=user_id, rule_ids=rule_ids, source="prompt_rule",
)
out["context"] = "\n".join(lines)
out["rule_ids"] = rule_ids
except Exception:
logger.debug("prompt rule arm failed", exc_info=True)
return out
# --- Write-path trigger (#2082): prior art at the moment code is written ------
# Auto-inject above fires on the operator's prompt. The moment reuse is actually
# lost is later — when the AGENT decides mid-task to write a helper — and nothing
+1 -1
View File
@@ -97,7 +97,7 @@ logger = logging.getLogger(__name__)
# surfacing is a claim ("this rule may apply to what you are doing") that a pull
# can confirm or refute, while an ambient one is a delivery nobody decided on.
# Add a source here only when a ranker picked it.
RANKED_SOURCES = ("write_path_rule", "pre_tool_rule")
RANKED_SOURCES = ("write_path_rule", "pre_tool_rule", "prompt_rule")
def is_ambient(source: str) -> bool: