Merge dev: scope-then-rank searches (#4958, #4961), CI gate on integration, milestone 456 steps 4-6 #205

Merged
bvandeusen merged 9 commits from dev into main 2026-10-05 22:23:48 -04:00
Owner

Nine commits, all CI-green on dev (latest: run 8226 on ccbccb0).

  • #4958: the rule search scopes before it ranks (983fd2c), and the scope test says why a search came back empty (79cd234).
  • Rule 177: the image build now waits on the integration lane (17f4348). The gate was watched rejecting a deliberately red run (9a76eb0), then that probe was reverted (4b1060a).
  • Milestone 456, steps 4-6:
  • #4961 (ccbccb0): every semantic search (notes, rules, milestones, systems) ranks exactly over the caller's in-scope chunks through _rank_scoped, so an in-scope record is no longer dropped behind ~40 nearer out-of-scope ones.

After deploy: compare search latency against the baseline in #4801.

🤖 Generated with Claude Code

Nine commits, all CI-green on dev (latest: run 8226 on ccbccb0). - **#4958:** the rule search scopes before it ranks (983fd2c), and the scope test says why a search came back empty (79cd234). - **Rule 177:** the image build now waits on the integration lane (17f4348). The gate was watched rejecting a deliberately red run (9a76eb0), then that probe was reverted (4b1060a). - **Milestone 456, steps 4-6:** - the notes arms run on the one retrieval pipeline (2b8f412, #4906); - the specs are the registry (869046d, #4907); - `_rule_moment` composes the rule builders (d2fa723, #4908). - **#4961 (ccbccb0):** every semantic search (notes, rules, milestones, systems) ranks exactly over the caller's in-scope chunks through `_rank_scoped`, so an in-scope record is no longer dropped behind ~40 nearer out-of-scope ones. After deploy: compare search latency against the baseline in #4801. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
bvandeusen added 9 commits 2026-10-05 22:23:43 -04:00
test(rules): the scope test says why a rule search returned nothing - an error or an empty scan (#4958)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 57s
CI & Build / integration (push) Successful in 1m6s
CI & Build / Python tests (push) Successful in 1m55s
CI & Build / Build & push image (push) Successful in 18s
79cd2341a6
semantic_search_rules fails open, so both failures on run 8209 and main
run 8212 read as set() with no trace. Assert report["searched"] and carry
the swallowed traceback into the failure.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
fix(retrieval): scope the rule search before ranking it - an in-scope rule is no longer lost behind other owners' nearer rules (#4958)
CI & Build / Python lint (push) Successful in 5s
CI & Build / Plugin hooks (push) Successful in 20s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / integration (push) Successful in 1m11s
CI & Build / Python tests (push) Successful in 2m0s
CI & Build / Build & push image (push) Successful in 29s
983fd2c4d1
Ordered straight off rule_embeddings, the planner walked the HNSW index,
which returns about hnsw.ef_search (40) nearest chunks across every owner
and project and filters by home only afterwards. A reader's own rule ranked
past the 40th chunk overall was silently dropped - on a shared install,
other users' rules fill those 40. This is the likely cause of the
intermittent test_integration_rule_scope failures, whose axis-vector
fixtures sit far from every real embedding in the graph.

The in-scope chunks now go through a MATERIALIZED CTE, which the index
cannot order, so the ranking over them is exact. A rulebook is hundreds of
chunks; exact is cheap. New test: 80 nearer rules belonging to someone else
no longer hide the reader's one rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ci: the image build waits on the integration lane (rule 177, #4958)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 52s
CI & Build / integration (push) Successful in 56s
CI & Build / Python tests (push) Successful in 1m53s
CI & Build / Build & push image (push) Successful in 18s
17f4348711
The integration job was added after the build gate (c6211e5) and never
joined its needs, so main run 8212 published :latest with integration red.
Both jobs carry the same ref condition, so the gate cannot be satisfied by
a skipped integration run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
test(ci): TEMPORARY deliberate integration failure, to watch the publish gate reject a red run (rule 177, #4958)
CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 11s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Failing after 1m1s
CI & Build / Python tests (push) Successful in 1m51s
CI & Build / Build & push image (push) Skipped
9a76eb0a34
Reverted next. Expect: integration red, Build & push skipped, :dev unmoved,
no image tagged with this commit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Revert the deliberate integration failure - the gate was watched rejecting run 8218 (rule 177, #4958)
CI & Build / Python lint (push) Successful in 4s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / integration (push) Successful in 53s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / Python tests (push) Successful in 1m53s
CI & Build / Build & push image (push) Successful in 19s
4b1060ae8b
Run 8218: integration failure, Build & push skipped, :dev still package
version 11098 (23:36:29Z, run 8217), and no image tagged 9a76eb0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
refactor(retrieval): the notes arms run on the one pipeline - auto_inject, its reuse and lesson slots, the write path by meaning, and rule_via_lesson are specs (milestone 456 step 4, #4906)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m2s
CI & Build / Python tests (push) Successful in 1m52s
CI & Build / Build & push image (push) Successful in 35s
2b8f41229d
retrieval_pipeline gains the notes half: NoteArm / NoteSlot / NoteMoment /
NoteIO / NoteResult and run_note_arm, which writes once the stages both
notes arms copied: search, withhold this response's own menu (#3739),
fresh/repeat split, the call row before any return (#3497, #3752), the
band, the reserved slots in their order (reuse evicts, lesson extends),
and the surfacing rows. The note renderer (_record_kind, _menu_name,
_menu_passage, the seen pointer, menu_entry) moves with it, and
run_via_lesson_arm takes rule_via_lesson.

Behaviour-preserving, with flags for today's differences: notes still log
BEFORE the band and rules after it (step 7's question). One deliberate
change: a notes arm now fails open like the rule arms, so a failing
search costs its lines and no longer the whole hook response.

The I/O is resolved from plugin_context at call time (_note_io), so the
existing patches keep working. The review re-run reads the AUTO_INJECT
spec instead of restating it, and its guard now compares the two live
searches. The registry declares the pipeline's notes fan-out sites.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
refactor(retrieval): the specs are the registry - SURFACES, the ranked POINTS rows and RANKED_SOURCES are read off the pipeline specs (milestone 456 step 5, #4907)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 59s
CI & Build / Python tests (push) Successful in 2m1s
CI & Build / Build & push image (push) Successful in 27s
869046dda2
Each arm spec now carries its tuning pair (Surface, moved verbatim into
retrieval_pipeline) and a Declared block - what the telemetry readout
must know and cannot read off its rows. The two ranked stages that are
not arms (preference_slot, rule_via_lesson) are RankedSource specs, and
the note slots carry their own declaration.

- retrieval_surfaces.SURFACES = the TUNED_ARMS tuning, same order
- retrieval_registry.POINTS ranked rows = one per spec in RANKED; the
  lookups, asked, ambient and pull rows stay declared there
- rule_usage.RANKED_SOURCES = RULE_RANKED_SOURCES + moment_rule

Settings keys, defaults, prose and order are unchanged (checked field by
field against HEAD). tests/test_retrieval_specs.py pins the derivation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
refactor(retrieval): the rule builders compose one moment - _rule_moment runs the arm and its via-lesson step for all three (milestone 456 step 6, #4908)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 14s
CI & Build / TypeScript typecheck (push) Successful in 53s
CI & Build / integration (push) Successful in 1m1s
CI & Build / Python tests (push) Successful in 1m54s
CI & Build / Build & push image (push) Successful in 28s
d2fa723373
build_prompt_rule_hint, build_tool_rule_hint and build_write_path_hint
each ran the rule arm and then the via-lesson step by hand, the first two
passing the query between them through a _via_query key on the payload.
They now call one composer, _rule_moment, and the split helpers
(_prompt_rule_hint, _tool_rule_hint, _add_rules_via_lessons) and the side
channel are deleted. Output shapes are unchanged: the prompt builder
returns no checkpoint key, the tool builder always does, and
shown_rule_ids stays the direct band.

A structural test pins that only _rule_moment runs a rule arm or the
via-lesson step, and that the three builders are its only callers.

plugin_context.py: 2,418 -> 2,361 lines this step; 3,558 when the
milestone began.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
fix(retrieval): every semantic search scopes first, then ranks - the note, milestone and system searches join the rule search on one shared shape, _rank_scoped (#4961)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 17s
CI & Build / TypeScript typecheck (push) Successful in 54s
CI & Build / integration (push) Successful in 1m0s
CI & Build / Python tests (push) Successful in 1m54s
CI & Build / Build & push image (push) Successful in 27s
ccbccb025c
Ordered straight off a *_embeddings table, the planner walks the HNSW index,
takes ~ef_search (40) nearest chunks across every owner and project, and only
then applies the scope: an in-scope record behind 40 nearer ones the caller
cannot see was silently dropped. #4958 fixed the rule search alone; the note,
milestone and system searches kept the fault.

_scoped_chunks builds the in-scope chunks with their distance; _rank_scoped
materializes them as a CTE and ranks exactly. All four searches go through it,
and the row shape is unchanged, so callers and mocks are untouched.

Tests: a structural guard that every semantic_search_* ranks through
_rank_scoped and orders nothing itself (with a replay of the old shape), a
compiled-SQL check of the MATERIALIZED CTE, and integration crowd tests - 80
nearer out-of-scope records - for notes (another user; the reader own other
project), milestones and systems.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
bvandeusen merged commit d4de0b9b54 into main 2026-10-05 22:23:48 -04:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: bvandeusen/FabledScribe#205