feat(lessons): soft links — a lesson and a rule arriving together in distinct situations are proposed as a link (milestone 440 step 3, #4637)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Build & push image (push) Successful in 30s
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m45s
CI & Build / Build & push image (push) Successful in 30s
When one hook response puts a lesson and a rule in front of the reader, the pair is recorded as evidence on a SUGGESTED lesson_rule_links row; once the pair has arrived together in PROPOSE_SITUATIONS (3) distinct situations, the next co-arrival carries one line asking the reader to judge it with judge_lesson_link. Nothing about surfacing changes: a suggested link carries no rule anywhere (that is #4633, confirmed links only). - lesson_rules: co_surfaced (fail-open; only pairs the reader could confirm — a lesson they may write, a rule they own; judged pairs gather nothing; one proposal per response; PROPOSE_COOLDOWN 6h between asks), plus the pure counting rules: situation_key, add_evidence, proposal_due, evidence_summary. A situation is the prompt on /retrieve (word tokens, sorted and de-duplicated, so trivial rewordings count once) and the FILE on /prior-art (every edit to one file is one situation). - rules_for_lessons shows a suggested link's evidence counts. - plugin_context: build_autoinject_hint returns lesson_ids, build_prompt_rule_hint returns shown_rule_ids, build_write_path_hint records its own pair; routes/plugin /retrieve records the prompt pair. Shown lines, repeats included: relevance makes a co-arrival, not the session ledger. - Evidence lives in the existing evidence column — no migration, and backup already carries it. - Tests: counting rules (unit), the recorder against Postgres (bar, repeat, judged pairs, ownership, cooldown, evidence kept on confirm); conftest stubs co_surfaced for unit tests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -12,6 +12,7 @@ from quart import Blueprint, g, jsonify, request
|
||||
|
||||
from scribe.auth import admin_required, get_current_user_id, login_required
|
||||
from scribe.config import Config
|
||||
from scribe.services import lesson_rules as lesson_rules_svc
|
||||
from scribe.services import plugin_context as plugin_ctx_svc
|
||||
from scribe.services import repo_bindings as repo_bindings_svc
|
||||
from scribe.services import report_check as report_check_svc
|
||||
@@ -147,7 +148,14 @@ async def autoinject_retrieve():
|
||||
result = await plugin_ctx_svc.build_autoinject_hint(
|
||||
g.user.id, q, project_id=project_id, exclude_ids=exclude_ids, context=ctx,
|
||||
)
|
||||
blocks = [b for b in (rules["context"], result["context"]) if b]
|
||||
# The one place both arms' lines meet, so the one place a lesson and a
|
||||
# rule arriving TOGETHER can be counted (#4637). Keyed on the operator's
|
||||
# prompt — the situation — and fails open to "".
|
||||
proposal = await lesson_rules_svc.co_surfaced(
|
||||
g.user.id, result.get("lesson_ids") or [], rules.get("shown_rule_ids") or [],
|
||||
arm="p", situation=q, project_id=project_id or None,
|
||||
)
|
||||
blocks = [b for b in (rules["context"], result["context"], proposal) if b]
|
||||
result["context"] = "\n\n".join(blocks)
|
||||
result["rule_ids"] = rules["rule_ids"]
|
||||
return jsonify(result)
|
||||
|
||||
@@ -17,6 +17,14 @@ UNJUDGED, and `list_unjudged` lists those. The three are kept apart because a
|
||||
lesson that stands alone is the raw material for a rule nobody has written yet
|
||||
(#4634), and that is unreadable if "nobody looked" means the same thing.
|
||||
|
||||
THE SOFT LINK (#4637). When one hook request puts a lesson and a rule in
|
||||
front of the reader together, `co_surfaced` records the pair as evidence on a
|
||||
SUGGESTED link, and once that evidence crosses a bar the next co-arrival asks
|
||||
the reader to judge it. The evidence counts distinct SITUATIONS, not
|
||||
arrivals: two records that resemble each other arrive together on every
|
||||
similar prompt, and counting each time would measure their similarity over
|
||||
and over rather than any relation between them.
|
||||
|
||||
ACL (rule 78). Linking changes what a lesson says about itself, so it needs
|
||||
WRITE on the lesson (`access.can_write_note`, share-aware). It names a rule, so
|
||||
it needs the rule to be one the caller may read — rules are owner-scoped, and
|
||||
@@ -26,10 +34,13 @@ someone must not hand them the titles of its owner's private rules.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import logging
|
||||
from datetime import datetime, timezone
|
||||
import re
|
||||
from datetime import datetime, timedelta, timezone
|
||||
|
||||
from sqlalchemy import and_, delete, exists, func, not_, select
|
||||
from sqlalchemy.exc import IntegrityError
|
||||
|
||||
from scribe.models import async_session
|
||||
from scribe.models.lesson_rule_link import (
|
||||
@@ -55,6 +66,25 @@ _NO_RULE_NOTE = "no rule fits: "
|
||||
# What `rule_judgment` reads on a lesson payload — the three answers.
|
||||
LINKED, NO_RULE, UNJUDGED = "linked", "no_rule", "unjudged"
|
||||
|
||||
# THE SOFT-LINK EVIDENCE BAR (#4637) — defaults, stated as defaults (rules 32,
|
||||
# 115): nothing about this install chose them.
|
||||
#
|
||||
# Three distinct situations before a pair is put to the reader. One is
|
||||
# coincidence and two is a pattern only by courtesy; three different prompts
|
||||
# or files that each brought the same lesson and rule up together is the
|
||||
# smallest count that says the pairing is not a single resemblance repeated.
|
||||
PROPOSE_SITUATIONS = 3
|
||||
# A proposal the reader passed over is asked again only after this long, so
|
||||
# one session's run of related prompts does not repeat the same question on
|
||||
# every turn.
|
||||
PROPOSE_COOLDOWN = timedelta(hours=6)
|
||||
# Fingerprints kept per link. The bar needs three; the rest is for the reader
|
||||
# judging later, and a list that grew with every situation would bloat a row
|
||||
# whose answer stopped changing long ago.
|
||||
_SITUATION_CAP = 50
|
||||
# A token shorter than this is grammar, not situation — "a", "is", "to".
|
||||
_MIN_TOKEN = 3
|
||||
|
||||
# How many rules a new lesson is offered to judge against. A handful, not the
|
||||
# fifty `what_might_apply` returns: this is read in the create response, at
|
||||
# the moment of writing, and a long tail there buries the one that fits.
|
||||
@@ -257,10 +287,15 @@ async def rules_for_lessons(user_id: int, lesson_ids) -> dict[int, list[dict]]:
|
||||
.where(_owned_rules_clause(user_id))
|
||||
)).all()
|
||||
for link, rule in sorted(rows, key=lambda r: (order.get(r[0].state, 9), r[1].id)):
|
||||
out[link.lesson_id].append({
|
||||
item = {
|
||||
"id": rule.id, "title": rule.title, "kind": rule.kind,
|
||||
"state": link.state, "note": link.note or "",
|
||||
})
|
||||
}
|
||||
if link.state == SUGGESTED:
|
||||
# What the suggestion rests on, so a reader can judge it from
|
||||
# here rather than having to go and count.
|
||||
item["evidence"] = evidence_summary(link.evidence)
|
||||
out[link.lesson_id].append(item)
|
||||
return out
|
||||
|
||||
|
||||
@@ -445,3 +480,173 @@ async def attach_rule_lessons(user_id: int, data: dict, rule_id: int) -> None:
|
||||
if lessons:
|
||||
data["lessons"] = lessons
|
||||
|
||||
|
||||
|
||||
def situation_key(arm: str, text: str) -> str:
|
||||
"""A fingerprint of one situation, so a repeat counts once.
|
||||
|
||||
`arm` is part of the key — "p" for a prompt, "w" for a file being written
|
||||
— because the two name different kinds of situation and must not collide.
|
||||
On the prompt arm the text is the prompt: lowercased, split into word
|
||||
tokens of _MIN_TOKEN or more, de-duplicated and sorted, so the same ask
|
||||
re-sent with different spacing, punctuation, case or word order is one
|
||||
situation. On the write arm the text is the PATH, not the code: every
|
||||
edit to one file is one situation, however much the code differs between
|
||||
them. Empty when there is nothing to key on.
|
||||
"""
|
||||
if arm == "w":
|
||||
basis = (text or "").strip()
|
||||
else:
|
||||
tokens = sorted({t for t in re.findall(r"[a-z0-9]+", (text or "").lower())
|
||||
if len(t) >= _MIN_TOKEN})
|
||||
basis = " ".join(tokens)
|
||||
if not basis:
|
||||
return ""
|
||||
return f"{arm}:" + hashlib.sha1(basis.encode()).hexdigest()[:16]
|
||||
|
||||
|
||||
def evidence_summary(evidence) -> dict:
|
||||
"""The counts a reader judges a suggested link by."""
|
||||
ev = evidence if isinstance(evidence, dict) else {}
|
||||
return {
|
||||
"situations": len(ev.get("situations") or []),
|
||||
"projects": len(ev.get("projects") or []),
|
||||
"co_surfaced": int(ev.get("co_surfaced") or 0),
|
||||
}
|
||||
|
||||
|
||||
def add_evidence(evidence, key: str, project_id: int | None, now: datetime) -> dict:
|
||||
"""A NEW evidence dict with this co-arrival counted. Pure, so the counting
|
||||
rules are testable without a database; a new dict, so SQLAlchemy sees the
|
||||
JSONB column change."""
|
||||
ev = dict(evidence) if isinstance(evidence, dict) else {}
|
||||
ev["co_surfaced"] = int(ev.get("co_surfaced") or 0) + 1
|
||||
situations = list(ev.get("situations") or [])
|
||||
if key and key not in situations and len(situations) < _SITUATION_CAP:
|
||||
situations.append(key)
|
||||
ev["situations"] = situations
|
||||
projects = list(ev.get("projects") or [])
|
||||
if project_id and int(project_id) not in projects:
|
||||
projects.append(int(project_id))
|
||||
ev["projects"] = projects
|
||||
ev.setdefault("first_at", now.isoformat())
|
||||
ev["last_at"] = now.isoformat()
|
||||
return ev
|
||||
|
||||
|
||||
def proposal_due(evidence, now: datetime) -> bool:
|
||||
"""At the bar, and not asked within the cooldown."""
|
||||
ev = evidence if isinstance(evidence, dict) else {}
|
||||
if len(ev.get("situations") or []) < PROPOSE_SITUATIONS:
|
||||
return False
|
||||
last = ev.get("proposed_at")
|
||||
if not last:
|
||||
return True
|
||||
try:
|
||||
return now - datetime.fromisoformat(last) >= PROPOSE_COOLDOWN
|
||||
except (TypeError, ValueError):
|
||||
return True
|
||||
|
||||
|
||||
def _proposal_line(lesson_id: int, lesson_title: str, rule_id: int, rule_title: str,
|
||||
kind: str, evidence: dict) -> str:
|
||||
counts = evidence_summary(evidence)
|
||||
where = (f" across {counts['projects']} projects" if counts["projects"] > 1 else "")
|
||||
return (
|
||||
f"> Lesson #{lesson_id} \"{lesson_title}\" and {kind} #{rule_id} "
|
||||
f"\"{rule_title}\" keep arriving together — {counts['situations']} "
|
||||
f"distinct situations{where}. If the lesson is an instance of the "
|
||||
f"{kind}, `judge_lesson_link({lesson_id}, {rule_id}, \"confirm\", "
|
||||
f"note=\"why\")` links them and the {kind} becomes reachable through "
|
||||
f"the lesson's situation; if it is not, `\"reject\"` with the why "
|
||||
f"stops the pair being asked about again."
|
||||
)
|
||||
|
||||
|
||||
async def co_surfaced(
|
||||
user_id: int, lesson_ids, rule_ids, *, arm: str, situation: str,
|
||||
project_id: int | None = None,
|
||||
) -> str:
|
||||
"""Record that these lessons and rules arrived together in one request,
|
||||
and return a proposal line when a pair has earned one ("" otherwise).
|
||||
|
||||
Called by the hook doors that put both kinds in front of the reader in
|
||||
the SAME response — `/retrieve` on a prompt, `/prior-art` on a write — so
|
||||
the pairing is exact: nothing is joined across requests, and no session
|
||||
identity is needed.
|
||||
|
||||
Only pairs the reader could confirm are recorded: a lesson they may write
|
||||
and a rule they own, the two checks `judge_link` makes. A pair already
|
||||
judged (confirmed or rejected) gathers nothing more — a confirmation needs
|
||||
no evidence, and a rejection is what stops the asking.
|
||||
|
||||
One proposal per request at most, so a menu never becomes a queue of
|
||||
questions. Fails open: this rides a hook, and a recall aid must never
|
||||
break the prompt or the write it decorates.
|
||||
"""
|
||||
try:
|
||||
return await _co_surfaced(user_id, lesson_ids, rule_ids, arm, situation, project_id)
|
||||
except Exception:
|
||||
logger.warning("lesson/rule co-surfacing could not be recorded", exc_info=True)
|
||||
return ""
|
||||
|
||||
|
||||
async def _co_surfaced(user_id, lesson_ids, rule_ids, arm, situation, project_id) -> str:
|
||||
from scribe.services.access import can_write_note
|
||||
from scribe.services.rulebooks import _owned_rules_clause
|
||||
|
||||
lessons = _ids(lesson_ids)
|
||||
rules = _ids(rule_ids)
|
||||
key = situation_key(arm, situation)
|
||||
if not lessons or not rules or not key:
|
||||
return ""
|
||||
lessons = [lid for lid in lessons if await can_write_note(user_id, lid)]
|
||||
if not lessons:
|
||||
return ""
|
||||
now = datetime.now(timezone.utc)
|
||||
proposal = None
|
||||
async with async_session() as session:
|
||||
owned = {
|
||||
r.id: r for r in (await session.execute(
|
||||
select(Rule).where(Rule.id.in_(rules)).where(_owned_rules_clause(user_id))
|
||||
)).scalars().all()
|
||||
}
|
||||
if not owned:
|
||||
return ""
|
||||
existing = {
|
||||
(row.lesson_id, row.rule_id): row for row in (await session.execute(
|
||||
select(LessonRuleLink).where(
|
||||
LessonRuleLink.lesson_id.in_(lessons),
|
||||
LessonRuleLink.rule_id.in_(list(owned)),
|
||||
)
|
||||
)).scalars().all()
|
||||
}
|
||||
for lid in lessons:
|
||||
for rid in owned:
|
||||
row = existing.get((lid, rid))
|
||||
if row is not None and row.state != SUGGESTED:
|
||||
continue
|
||||
if row is None:
|
||||
row = LessonRuleLink(lesson_id=lid, rule_id=rid, state=SUGGESTED)
|
||||
session.add(row)
|
||||
ev = add_evidence(row.evidence, key, project_id, now)
|
||||
if proposal is None and proposal_due(ev, now):
|
||||
ev["proposed_at"] = now.isoformat()
|
||||
ev["proposed_count"] = int(ev.get("proposed_count") or 0) + 1
|
||||
proposal = (lid, rid, ev)
|
||||
row.evidence = ev
|
||||
try:
|
||||
await session.commit()
|
||||
except IntegrityError:
|
||||
# Two hook requests created the same pair at once; the other
|
||||
# one's row stands, and this co-arrival is simply not counted.
|
||||
await session.rollback()
|
||||
return ""
|
||||
if proposal is None:
|
||||
return ""
|
||||
lid, rid, ev = proposal
|
||||
lesson_title = (await session.execute(
|
||||
select(Note.title).where(Note.id == lid)
|
||||
)).scalar_one_or_none() or ""
|
||||
rule = owned[rid]
|
||||
return _proposal_line(lid, lesson_title, rid, rule.title, rule.kind or "rule", ev)
|
||||
|
||||
@@ -33,6 +33,7 @@ from scribe.services.embeddings import (
|
||||
semantic_search_notes,
|
||||
semantic_search_rules,
|
||||
)
|
||||
from scribe.services import lesson_rules as lesson_rules_svc
|
||||
from scribe.services.lessons import LESSON_NOTE_TYPE
|
||||
from scribe.services.note_usage import record_surfaced
|
||||
from scribe.services.rule_usage import record_rule_surfaced
|
||||
@@ -1275,7 +1276,15 @@ async def build_autoinject_hint(
|
||||
project_id=project_id,
|
||||
)
|
||||
|
||||
return {"context": "\n".join(lines), "note_ids": note_ids, "config": cfg}
|
||||
# The lessons this menu put in front of the reader, repeats included — for
|
||||
# the soft-link recorder (#4637), which pairs them with the rule arm's
|
||||
# lines in the same response. Relevance, not the session ledger, is what
|
||||
# makes a co-arrival, so a `seen` lesson counts.
|
||||
lesson_ids = [int(n.id) for _s, n in kept if _record_kind(n) == LESSON_NOTE_TYPE]
|
||||
return {
|
||||
"context": "\n".join(lines), "note_ids": note_ids, "config": cfg,
|
||||
"lesson_ids": lesson_ids,
|
||||
}
|
||||
|
||||
|
||||
async def _reserve_slot_for_preference(
|
||||
@@ -1498,6 +1507,10 @@ async def build_prompt_rule_hint(
|
||||
# (#3668). The slot's hit is surfaced under `preference_slot` by the
|
||||
# helper, against that source's own row.
|
||||
rule_ids = [rule.id for _score, rule in fresh]
|
||||
# Every rule LINE, repeats and the reserved slot included — what the
|
||||
# reader actually had in front of them, for the soft-link recorder
|
||||
# (#4637). Distinct from `rule_ids`, which is telemetry's fresh-only cut.
|
||||
out["shown_rule_ids"] = [rule.id for _score, rule in hits]
|
||||
|
||||
# RANKED, not ambient: this arm chose what it showed. The name is also
|
||||
# in `rule_usage.RANKED_SOURCES`, and it has to be — a ranked source
|
||||
@@ -2583,6 +2596,7 @@ async def build_write_path_hint(
|
||||
#
|
||||
# Fails open like every other arm: a rule hint must never break a write.
|
||||
rule_ids: list[int] = []
|
||||
shown_rule_ids: list[int] = []
|
||||
checkpoint: dict = {}
|
||||
try:
|
||||
already = set(exclude_rule_ids or [])
|
||||
@@ -2623,6 +2637,9 @@ async def build_write_path_hint(
|
||||
# would inflate pull_through's denominator with a choice this arm never
|
||||
# made. A reference is a RENDERING decision, not a retrieval outcome.
|
||||
rule_ids.extend(rule.id for _score, rule in fresh)
|
||||
# Every rule line, repeats included — what the soft-link recorder
|
||||
# pairs with this response's lessons (#4637).
|
||||
shown_rule_ids = [rule.id for _score, rule in kept]
|
||||
# The stop, beside the lines rather than instead of them — see
|
||||
# `checkpoint_for`. `kept` is passed, not `fresh`: whether a rule
|
||||
# was named earlier this session says nothing about whether this
|
||||
@@ -2691,6 +2708,19 @@ async def build_write_path_hint(
|
||||
except Exception:
|
||||
logger.debug("write-path rule arm failed", exc_info=True)
|
||||
|
||||
# A lesson and a rule on this one response (#4637): the soft-link
|
||||
# recorder counts the pair, keyed on the FILE — every edit to one file is
|
||||
# one situation — and asks about it once the evidence holds. Fails open.
|
||||
shown_lessons = [
|
||||
int(item["id"]) for _m, item in menu if item.get("kind") == LESSON_NOTE_TYPE
|
||||
]
|
||||
proposal = await lesson_rules_svc.co_surfaced(
|
||||
user_id, shown_lessons, shown_rule_ids, arm="w", situation=path,
|
||||
project_id=project_id or None,
|
||||
)
|
||||
if proposal:
|
||||
lines.append(proposal)
|
||||
|
||||
return {
|
||||
"context": "\n".join(lines),
|
||||
"note_ids": note_ids,
|
||||
|
||||
Reference in New Issue
Block a user