feat(rules): rules retrieve against the operator's message (#3852)
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m35s
CI & Build / Build & push image (push) Successful in 38s
CI & Build / Plugin hooks (push) Successful in 9s
CI & Build / Python lint (push) Successful in 3s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m35s
CI & Build / Build & push image (push) Successful in 38s
The third rule arm, and the one the other two cannot reach. `write_path_rule` is keyed on code, `pre_tool_rule` on a command — both things the session is about to DO. A rule that governs what to SAY has no such trigger: extract intent from loose phrasing, raise a conflict before acting, hand off an action with its reason, end a finding with an offer all bind on a RESPONSE, and no tool call precedes one. The operator's message is the only query that exists before a response is composed. That hook searched notes alone, so no rule had ever been retrieved against a thing the operator actually said — and residency was the only surface those rules had, which is what milestone 394 removes. A SEPARATE FUNCTION, not a branch in build_autoinject_hint, because of its early returns. That arm bails when auto-inject is disabled, when the query is blank, when nothing clears the note bar — every one a statement about NOTES. Folded in, an operator who turned the awareness menu off would silently lose their rules, a coupling with no symptom since both look like a quiet hook. Two functions, two sets of gates, composed in the route. Guarded as "the rule arm never asks the notes arm's config", which is the structural fact. Joins _ARMS rather than getting its own test file. #3497's history is that the pre-tool arm inherited a defect from its sibling by being MODELLED on it instead of sharing with it, and a third arm modelled on two is two chances to repeat that. Repeat rendering, fresh-only counting, log-before-bailout, the kind register and the two-recorders identity are properties of every arm or of none. The bar is INHERITED and says so. 0.72 was tuned against code and commands; prose is a different query shape against the same documents, and triggers are written in the vocabulary of the moment — which for most rules is act vocabulary. Starting at the only number with evidence behind it and logging every call from the first deploy is what makes it settleable; guessing lower would put an unmeasured bar in front of a corpus that binds. k=3, anchored on this hook's own budget rather than the act arms'. RULEHINT_LIMIT is 1 because that arm fires before every Bash call; this one fires once per turn, beside a notes menu already spending three slots. And a prompt genuinely contains more than one act — "merge to main and then start on X" is two — where a command is one thing. `prompt_rule` added to RANKED_SOURCES: a ranker picked it, and a ranked source missing from that tuple is silently counted as bulk delivery and drops out of the pull-through denominator. The hook reads and writes the SHARED rule ledger under scribe-priorart, not a private one — one session keeps one list, aged (#3751), so a rule named here is not re-announced before the next Bash call. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011cPyzNnegXHr5iRMzzy5KJ
This commit is contained in:
@@ -91,13 +91,37 @@ async def autoinject_retrieve():
|
||||
project_id (opt) — explicit project scope override (ad-hoc/testing).
|
||||
exclude_ids (opt) — comma-separated note ids already injected this
|
||||
session; skipped so each note injects at most once.
|
||||
exclude_rule_ids — comma-separated rule ids already surfaced this
|
||||
(opt) session. SHARED with /prior-art and /tool-rules on
|
||||
purpose: one session keeps ONE rule ledger, so a
|
||||
rule named by any arm is not re-announced by
|
||||
another. Ages out (#3751), so salience decays.
|
||||
|
||||
TWO ARMS, TWO SETS OF GATES. Rules ride the same hook and the same query
|
||||
but nothing else: the notes menu can be disabled, thresholded and top-k'd
|
||||
by the operator without touching whether a rule reaches them. Composed
|
||||
here rather than inside either builder so neither one's early return can
|
||||
silently suppress the other.
|
||||
|
||||
Rules come FIRST in the payload. A rule or preference governing the answer
|
||||
is more consequential than a menu of things that might be worth reading,
|
||||
and a reader who stops after the first block should have stopped after
|
||||
the right one.
|
||||
"""
|
||||
q = (request.args.get("q") or "").strip()
|
||||
project_id, _repo, _unbound = await _project_scope()
|
||||
exclude_ids = _int_list(request.args.get("exclude_ids"))
|
||||
exclude_rule_ids = _int_list(request.args.get("exclude_rule_ids"))
|
||||
|
||||
rules = await plugin_ctx_svc.build_prompt_rule_hint(
|
||||
g.user.id, q, project_id=project_id, exclude_rule_ids=exclude_rule_ids
|
||||
)
|
||||
result = await plugin_ctx_svc.build_autoinject_hint(
|
||||
g.user.id, q, project_id=project_id, exclude_ids=exclude_ids
|
||||
)
|
||||
blocks = [b for b in (rules["context"], result["context"]) if b]
|
||||
result["context"] = "\n\n".join(blocks)
|
||||
result["rule_ids"] = rules["rule_ids"]
|
||||
return jsonify(result)
|
||||
|
||||
|
||||
|
||||
@@ -314,6 +314,42 @@ _AUTOINJECT_BAND = 0.10
|
||||
# awareness menu (titles only), never a content dump.
|
||||
_AUTOINJECT_MAX_TOP_K = 10
|
||||
|
||||
# --- the prompt-boundary rule arm (#3852) ------------------------------------
|
||||
#
|
||||
# Both existing rule arms are keyed on something the session is about to DO —
|
||||
# a file write, a command. A rule that governs what to SAY has no such moment.
|
||||
# Extract intent from loose phrasing, raise a conflict before acting, hand off
|
||||
# an action with its reason, end a finding with an offer: every one binds on a
|
||||
# RESPONSE, and no tool call precedes a response.
|
||||
#
|
||||
# The operator's message is the only query that exists before one is composed,
|
||||
# and this arm is what runs against it. Until now that hook searched notes
|
||||
# alone, so no rule had ever been retrieved against a thing the operator said.
|
||||
PROMPTRULE_THRESHOLD_KEY = "kb_promptrule_threshold"
|
||||
# INHERITED FROM THE ACT ARMS, AND NOT YET EARNED HERE. 0.72 was tuned against
|
||||
# code and shell commands. An operator's prose is a different query shape
|
||||
# against the same documents, and nothing yet says the two distributions line
|
||||
# up — triggers are written in the vocabulary of the MOMENT, which for most
|
||||
# rules is act vocabulary, so prose may well score lower across the board.
|
||||
#
|
||||
# Starting at the act arms' number anyway is deliberate: it is the only value
|
||||
# with evidence behind it, and guessing lower would put an unmeasured bar in
|
||||
# front of a corpus that binds. Every call is logged under `prompt_rule` from
|
||||
# the first deploy, so a few days of real traffic settles it — read
|
||||
# `near_miss_samples` (#3807) before moving this, not the percentile alone.
|
||||
PROMPTRULE_DEFAULT_THRESHOLD = 0.72
|
||||
|
||||
# MORE THAN THE ACT ARMS' SINGLE SLOT, anchored on this hook's budget rather
|
||||
# than theirs. RULEHINT_LIMIT is 1 because that arm fires before EVERY Bash
|
||||
# call, where a second line is a second interruption per command. This arm
|
||||
# fires once per TURN, on the same hook whose notes menu already spends
|
||||
# AUTOINJECT_DEFAULT_TOP_K slots — so that is the comparable budget.
|
||||
#
|
||||
# And a prompt genuinely contains more than one act. "Merge to main and then
|
||||
# start on X" is two, governed by different rules; k=1 cannot serve that case
|
||||
# at all, where the act arms never face it because a command is one thing.
|
||||
PROMPTRULE_LIMIT = 3
|
||||
|
||||
|
||||
def _slugify(text: str) -> str:
|
||||
"""kebab-case slug for a skill directory name (a-z0-9 + single hyphens)."""
|
||||
@@ -642,6 +678,115 @@ async def build_autoinject_hint(
|
||||
return {"context": "\n".join(lines), "note_ids": note_ids, "config": cfg}
|
||||
|
||||
|
||||
async def build_prompt_rule_hint(
|
||||
user_id: int,
|
||||
query: str,
|
||||
*,
|
||||
project_id: int = 0,
|
||||
exclude_rule_ids: list[int] | None = None,
|
||||
) -> dict:
|
||||
"""Rules and preferences that may apply to what the operator just asked.
|
||||
|
||||
The third rule arm, and the one that closes a gap the other two cannot
|
||||
reach. `write_path_rule` is keyed on code, `pre_tool_rule` on a command —
|
||||
both are things the session is about to DO. A rule that governs what to
|
||||
SAY has no such trigger, and residency was the only surface it ever had.
|
||||
Removing residency (milestone 394) without this would drop that half of
|
||||
the corpus on the floor.
|
||||
|
||||
A SEPARATE FUNCTION, not a branch inside build_autoinject_hint, and the
|
||||
reason is its early returns. That arm bails when auto-inject is disabled,
|
||||
when the query is blank, when nothing clears the note bar — and every one
|
||||
of those is a statement about NOTES. Folded in, a user who turned the
|
||||
notes menu off would silently lose their rules too, which is the kind of
|
||||
coupling nothing downstream could see. Two functions, two sets of gates,
|
||||
composed by the caller.
|
||||
|
||||
THE OUTPUT IS DELIBERATELY NOT QUOTED, where the notes menu is. The task
|
||||
asked whether the two share a header; the answer is that neither needs
|
||||
one. A note line is a bare title and needs the menu's header to say what
|
||||
it is doing there, while a rule line names itself in its opening words
|
||||
("Standing rule that may apply…" / "Preference that may apply…"). Leaving
|
||||
rules unquoted separates the two claims visually with no extra prose, and
|
||||
matches how a rule line already renders on both act arms.
|
||||
|
||||
Fails open and returns empty context on any error, like its siblings: a
|
||||
recall aid may never break the operator's prompt.
|
||||
"""
|
||||
out: dict = {"context": "", "rule_ids": []}
|
||||
q = (query or "").strip()
|
||||
if not q:
|
||||
return out
|
||||
|
||||
try:
|
||||
try:
|
||||
threshold = float(await get_setting(
|
||||
user_id, PROMPTRULE_THRESHOLD_KEY,
|
||||
str(PROMPTRULE_DEFAULT_THRESHOLD)))
|
||||
except (TypeError, ValueError):
|
||||
threshold = PROMPTRULE_DEFAULT_THRESHOLD
|
||||
threshold = min(1.0, max(0.0, threshold))
|
||||
|
||||
t0 = time.perf_counter()
|
||||
_rep: dict = {}
|
||||
# NOT scoped to the project, and that is the corpus's own decision
|
||||
# rather than an omission here — semantic_search_rules is scoped by
|
||||
# OWNERSHIP on purpose, because "is there a rule about this" is asked
|
||||
# across a whole rulebook. `project_id` below reaches the log row and
|
||||
# nothing else.
|
||||
hits = await semantic_search_rules(
|
||||
user_id, q, limit=PROMPTRULE_LIMIT, threshold=threshold,
|
||||
report=_rep,
|
||||
)
|
||||
duration_ms = (time.perf_counter() - t0) * 1000.0
|
||||
|
||||
already = set(exclude_rule_ids or [])
|
||||
fresh = [(score, rule) for score, rule in hits if rule.id not in already]
|
||||
|
||||
# BEFORE the early return, for the reason both sibling arms spell out
|
||||
# at length: a call that found nothing is the only evidence a bar is
|
||||
# too high, and an arm that logs only the calls it liked reports a
|
||||
# flawless clear-rate however badly it is tuned. This bar is inherited
|
||||
# and unverified for this corpus, so the zero rows are the point.
|
||||
record_retrieval(
|
||||
user_id=user_id, source="prompt_rule", query=q,
|
||||
threshold=threshold, limit=PROMPTRULE_LIMIT,
|
||||
project_id=project_id,
|
||||
is_task=None, results=fresh, duration_ms=duration_ms,
|
||||
best_available=_rep.get("best_available_score"),
|
||||
best_available_id=_rep.get("best_available_id"),
|
||||
searched=bool(_rep.get("searched", True)),
|
||||
suppressed=len(hits) - len(fresh),
|
||||
)
|
||||
# `hits`, not `fresh` (#3750): a call whose only hit is a repeat still
|
||||
# has something to say, it just says it differently.
|
||||
if not hits:
|
||||
return out
|
||||
|
||||
lines = [
|
||||
_rule_hint_line(rule, where="to this request", seen=rule.id in already)
|
||||
for _score, rule in hits
|
||||
]
|
||||
# FRESH-ONLY (#3752). A reference is a rendering decision, not a
|
||||
# retrieval outcome, and counting one here would inflate the
|
||||
# denominator pull_through is read from.
|
||||
rule_ids = [rule.id for _score, rule in fresh]
|
||||
|
||||
# RANKED, not ambient: this arm chose what it showed. The name is also
|
||||
# in `rule_usage.RANKED_SOURCES`, and it has to be — a ranked source
|
||||
# missing from that tuple is counted as a bulk delivery nobody decided
|
||||
# on, which silently moves it out of the pull-through denominator.
|
||||
if rule_ids:
|
||||
record_rule_surfaced(
|
||||
user_id=user_id, rule_ids=rule_ids, source="prompt_rule",
|
||||
)
|
||||
out["context"] = "\n".join(lines)
|
||||
out["rule_ids"] = rule_ids
|
||||
except Exception:
|
||||
logger.debug("prompt rule arm failed", exc_info=True)
|
||||
return out
|
||||
|
||||
|
||||
# --- Write-path trigger (#2082): prior art at the moment code is written ------
|
||||
# Auto-inject above fires on the operator's prompt. The moment reuse is actually
|
||||
# lost is later — when the AGENT decides mid-task to write a helper — and nothing
|
||||
|
||||
@@ -97,7 +97,7 @@ logger = logging.getLogger(__name__)
|
||||
# surfacing is a claim ("this rule may apply to what you are doing") that a pull
|
||||
# can confirm or refute, while an ambient one is a delivery nobody decided on.
|
||||
# Add a source here only when a ranker picked it.
|
||||
RANKED_SOURCES = ("write_path_rule", "pre_tool_rule")
|
||||
RANKED_SOURCES = ("write_path_rule", "pre_tool_rule", "prompt_rule")
|
||||
|
||||
|
||||
def is_ambient(source: str) -> bool:
|
||||
|
||||
Reference in New Issue
Block a user