CI & Build / Python lint (push) Successful in 5s
CI & Build / Plugin hooks (push) Successful in 12s
CI & Build / TypeScript typecheck (push) Successful in 56s
CI & Build / integration (push) Successful in 1m6s
CI & Build / Python tests (push) Successful in 1m40s
CI & Build / Build & push image (push) Successful in 35s
Milestone 419 step 2. Step 1 made an outcome recordable; this makes it readable. `retrieval_summary`'s rule block gains `applied`, `departed` and `distinct_rules_acted`, and `_compute_warnings` gains two codes. TWO CODES, NOT ONE WITH A ZERO IN IT. `read_and_unacted` reports rules that were opened and left no outcome, against the ones that did. It only fires once outcomes exist anywhere in the window, because a window with none cannot tell "every rule was ignored" from "nothing calls `rule_outcome` yet" — and on every install the day this ships, the truth is the second. Claiming the first there would be #3311's failure exactly: a statistic that could not vary being read as a fact about the corpus. The cold case gets its own code, `outcomes_never_recorded`, whose prose says in as many words that it does NOT mean the rules were ignored. `applied` AND `departed` ARE NOT SUMMED. A departure carries the reason the agent gave and is evidence about the RULE; an application is evidence about the agent. Folded together they would say only "an outcome exists", which is true of both and useful about neither. `distinct_rules_acted` counts either, because for the unacted arithmetic the distinction does not matter. An outcome is not a pull. The fold branches on OUTCOMES first and never routes an outcome through the surfaced/ambient split: `source` on an outcome row names the door the outcome came through, not a ranker, so the ambient distinction has nothing to say about it. An integration test holds that line — if an outcome leaked into the pull counters the silently-unchanged rule would vanish into a compliant-looking total, which is the confusion #4212 was opened to end. Verified by lifting the shipped `_compute_warnings` out of source with `ast` and exercising it against the six populations the new tests assert: cold instrument, warm instrument, departures-only, full compliance, nothing opened, and a failed read. The integration tests for the new counts run against real Postgres in CI — count(distinct) with an IN over an unconstrained column is a SQL shape a mock would agree with whatever it did, which is what #2663 was. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
352 lines
15 KiB
Python
352 lines
15 KiB
Python
"""`retrieval_telemetry` says what is wrong, rather than implying it (#3431).
|
|
|
|
WHY THIS IS TESTED AS ARITHMETIC AND NOT THROUGH THE DATABASE. The warnings
|
|
are a pure function of the blocks the readout already built — that is the
|
|
design, so a verdict can never disagree with the numbers printed beside it —
|
|
and the risk in them is not whether rows load. It is whether a rule fires on
|
|
the wrong shape. Every case below is a shape that once produced, or would
|
|
produce, a wrong reading:
|
|
|
|
* an arm that never declines, which is either a missing floor or a missing
|
|
LOG (#3497 — the rule arms recorded only their hits, so their decline
|
|
count was structurally zero and the obvious warning would have sent a
|
|
reader to move a threshold that was never involved);
|
|
* a search that never declines, which is a search working correctly;
|
|
* a quiet window, where firing every check would describe an empty database
|
|
rather than a broken one (rule 115);
|
|
* a floor the band is sitting on, which is the case where tuning changes
|
|
volume while looking like it changes quality.
|
|
|
|
The ε and N boundaries are tested from BOTH sides. A threshold asserted only
|
|
where it fires is half-tested: the expensive failure here is a false positive,
|
|
because a readout that cries wolf is one nobody reads.
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
import pytest
|
|
|
|
from scribe.services.retrieval_telemetry import (
|
|
WARN_FLOOR_EPSILON_DEFAULT, WARN_MIN_CALLS_DEFAULT,
|
|
_compute_warnings, _num, _silent_surfaces,
|
|
)
|
|
|
|
N = WARN_MIN_CALLS_DEFAULT
|
|
EPS = WARN_FLOOR_EPSILON_DEFAULT
|
|
|
|
|
|
def src(**kw) -> dict:
|
|
"""One source block, shaped like the readout builds it."""
|
|
b = {
|
|
"calls": kw.pop("calls", 0),
|
|
"zero_result_calls": kw.pop("zero_result_calls", 0),
|
|
"top_score": {"p10": kw.pop("p10", None), "p50": None, "p90": None,
|
|
"min": None, "max": None},
|
|
"avg_result_count": None,
|
|
"p90_duration_ms": kw.pop("p90_duration_ms", 12.0),
|
|
}
|
|
b.update(kw)
|
|
return b
|
|
|
|
|
|
def warn(sources, usage=None, rule_usage=None, floors=None,
|
|
min_calls=N, epsilon=EPS) -> list[dict]:
|
|
return _compute_warnings(
|
|
sources, usage or {}, rule_usage or {}, floors or {}, min_calls, epsilon,
|
|
)
|
|
|
|
|
|
def codes(ws, source=None) -> set[str]:
|
|
return {w["code"] for w in ws if source is None or w["source"] == source}
|
|
|
|
|
|
# ── cannot_decline ────────────────────────────────────────────────────────
|
|
|
|
def test_an_unbidden_arm_that_never_declines_is_flagged() -> None:
|
|
ws = warn({"auto_inject": src(calls=300, zero_result_calls=0)})
|
|
assert "cannot_decline" in codes(ws, "auto_inject")
|
|
|
|
|
|
def test_the_warning_carries_the_numbers_that_produced_it() -> None:
|
|
"""Not "check auto_inject" — that is an instruction to redo the analysis."""
|
|
w = next(w for w in warn({"auto_inject": src(calls=300)})
|
|
if w["code"] == "cannot_decline")
|
|
assert w["numbers"] == {"calls": 300, "zero_result_calls": 0}
|
|
assert "300" in w["detail"]
|
|
|
|
|
|
def test_a_quiet_window_flags_nothing() -> None:
|
|
"""Three calls returning nothing is an afternoon, not a defect (rule 115)."""
|
|
assert codes(warn({"auto_inject": src(calls=3, zero_result_calls=0)})) == set()
|
|
|
|
|
|
def test_the_call_count_boundary_holds_on_both_sides() -> None:
|
|
assert "cannot_decline" not in codes(warn({"auto_inject": src(calls=N - 1)}))
|
|
assert "cannot_decline" in codes(warn({"auto_inject": src(calls=N)}))
|
|
|
|
|
|
def test_an_asked_surface_is_never_flagged_for_not_declining() -> None:
|
|
"""A search returning a list every time is a search doing its job."""
|
|
for asked in ("mcp_search", "wide_net", "rest_search", "browse_search"):
|
|
assert "cannot_decline" not in codes(
|
|
warn({asked: src(calls=500, zero_result_calls=0)}), asked)
|
|
|
|
|
|
def test_an_arm_that_does_decline_is_not_flagged() -> None:
|
|
ws = warn({"auto_inject": src(calls=300, zero_result_calls=1)})
|
|
assert "cannot_decline" not in codes(ws)
|
|
|
|
|
|
def test_an_arm_not_known_to_log_unconditionally_is_exempt(monkeypatch) -> None:
|
|
"""#3497: a structurally-zero decline count is a logging bug, not a floor.
|
|
|
|
Flagging it would report a ranking problem and send the reader to a
|
|
threshold that was never involved.
|
|
"""
|
|
from dataclasses import replace
|
|
|
|
from scribe.services import retrieval_registry as reg
|
|
|
|
patched = dict(reg.POINTS)
|
|
patched["auto_inject"] = replace(patched["auto_inject"],
|
|
logs_unconditionally=False)
|
|
monkeypatch.setattr(reg, "POINTS", patched)
|
|
assert "cannot_decline" not in codes(
|
|
warn({"auto_inject": src(calls=300, zero_result_calls=0)}))
|
|
|
|
|
|
# ── band_hugs_floor ───────────────────────────────────────────────────────
|
|
|
|
def test_a_band_sitting_on_its_floor_is_flagged() -> None:
|
|
ws = warn({"auto_inject": src(calls=100, zero_result_calls=5, p10=0.705)},
|
|
floors={"auto_inject": 0.70})
|
|
assert "band_hugs_floor" in codes(ws, "auto_inject")
|
|
|
|
|
|
def test_a_band_clear_of_its_floor_is_not() -> None:
|
|
ws = warn({"auto_inject": src(calls=100, zero_result_calls=5, p10=0.80)},
|
|
floors={"auto_inject": 0.70})
|
|
assert "band_hugs_floor" not in codes(ws)
|
|
|
|
|
|
@pytest.mark.parametrize("gap,flagged", [
|
|
(EPS / 2, True), # inside
|
|
(EPS, False), # exactly at the boundary is NOT "hugging"
|
|
(EPS * 2, False), # clear
|
|
])
|
|
def test_the_epsilon_boundary_holds_on_both_sides(gap, flagged) -> None:
|
|
ws = warn({"auto_inject": src(calls=100, zero_result_calls=5,
|
|
p10=round(0.70 + gap, 6))},
|
|
floors={"auto_inject": 0.70})
|
|
assert ("band_hugs_floor" in codes(ws)) is flagged
|
|
|
|
|
|
def test_an_arm_with_no_readable_floor_is_not_guessed_at() -> None:
|
|
"""Reserved slots borrow a parent's floor; naming it here would attribute
|
|
the parent's setting to a child that has no dial of its own."""
|
|
ws = warn({"preference_slot": src(calls=100, zero_result_calls=5, p10=0.7001)},
|
|
floors={})
|
|
assert "band_hugs_floor" not in codes(ws)
|
|
|
|
|
|
# ── no_duration ───────────────────────────────────────────────────────────
|
|
|
|
def test_rows_without_timings_are_a_logging_gap() -> None:
|
|
ws = warn({"auto_inject": src(calls=1, zero_result_calls=1,
|
|
p90_duration_ms=None)})
|
|
assert "no_duration" in codes(ws)
|
|
|
|
|
|
def test_one_untimed_call_is_already_the_defect() -> None:
|
|
"""No minimum on this one — unlike the others, it is not about volume."""
|
|
ws = warn({"auto_inject": src(calls=1, p90_duration_ms=None)})
|
|
assert "no_duration" in codes(ws)
|
|
assert "cannot_decline" not in codes(ws), "volume rules still need volume"
|
|
|
|
|
|
def test_a_source_with_no_calls_reports_no_timing_gap() -> None:
|
|
assert "no_duration" not in codes(warn({"auto_inject": src(calls=0)}))
|
|
|
|
|
|
# ── surfaced_never_pulled ─────────────────────────────────────────────────
|
|
|
|
def test_records_shown_and_never_opened_are_reported_per_corpus() -> None:
|
|
ws = warn(
|
|
{},
|
|
usage={"distinct_notes_surfaced": 171, "distinct_notes_pulled": 40},
|
|
rule_usage={"distinct_rules_surfaced": 69, "distinct_rules_pulled": 11},
|
|
)
|
|
found = {w["numbers"]["corpus"]: w["numbers"]
|
|
for w in ws if w["code"] == "surfaced_never_pulled"}
|
|
assert found["notes"]["never_pulled"] == 131
|
|
assert found["rules"]["never_pulled"] == 58
|
|
|
|
|
|
def test_everything_opened_reports_nothing() -> None:
|
|
ws = warn({}, usage={"distinct_notes_surfaced": 5, "distinct_notes_pulled": 5})
|
|
assert "surfaced_never_pulled" not in codes(ws)
|
|
|
|
|
|
def test_an_empty_corpus_reports_nothing_rather_than_zero() -> None:
|
|
ws = warn({}, usage={"distinct_notes_surfaced": 0, "distinct_notes_pulled": 0})
|
|
assert "surfaced_never_pulled" not in codes(ws)
|
|
|
|
|
|
# ── unregistered_source ───────────────────────────────────────────────────
|
|
|
|
def test_a_source_missing_from_the_registry_is_reported() -> None:
|
|
"""The check that covers what the static test cannot reach — a source
|
|
arriving through a parameter or a dict key."""
|
|
ws = warn({"a_fourth_write_path_arm": src(calls=50)})
|
|
assert codes(ws, "a_fourth_write_path_arm") == {"unregistered_source"}
|
|
|
|
|
|
def test_an_unregistered_source_gets_no_other_verdict() -> None:
|
|
"""Its numbers are real, but nothing says whether it was asked."""
|
|
ws = warn({"mystery": src(calls=500, zero_result_calls=0)})
|
|
assert "cannot_decline" not in codes(ws)
|
|
|
|
|
|
def test_every_registered_source_stays_quiet_when_healthy() -> None:
|
|
"""The negative control for the whole suite (rule 167).
|
|
|
|
A guard that cannot pass cleanly is not a guard — if a healthy window
|
|
produced warnings, every assertion above would be meaningless.
|
|
"""
|
|
from scribe.services.retrieval_registry import POINTS
|
|
|
|
healthy = {s: src(calls=100, zero_result_calls=40, p10=0.90)
|
|
for s in POINTS}
|
|
assert warn(healthy, floors={s: 0.70 for s in POINTS}) == []
|
|
|
|
|
|
# ── silent surfaces ───────────────────────────────────────────────────────
|
|
|
|
def test_a_registered_point_with_no_rows_is_reported() -> None:
|
|
quiet = _silent_surfaces({"auto_inject": src(calls=100)}, {}, {}, active=True)
|
|
assert "prompt_rule" in {p["source"] for p in quiet}
|
|
|
|
|
|
def test_a_point_that_emitted_is_not_reported() -> None:
|
|
quiet = _silent_surfaces({"auto_inject": src(calls=100)}, {}, {}, active=True)
|
|
assert "auto_inject" not in {p["source"] for p in quiet}
|
|
|
|
|
|
def test_a_deliberately_quiet_point_is_never_reported() -> None:
|
|
"""A justified silence must not read as a gap (#2475)."""
|
|
quiet = {p["source"] for p in _silent_surfaces({}, {}, {}, active=True)}
|
|
for web_only in ("rest_search", "browse_search", "rest_note", "rest_rule"):
|
|
assert web_only not in quiet
|
|
|
|
|
|
def test_an_inactive_window_reports_no_silence_at_all() -> None:
|
|
"""On a fresh install every point is silent; the list would be the
|
|
registry printed back (rule 115)."""
|
|
assert _silent_surfaces({}, {}, {}, active=False) == []
|
|
|
|
|
|
def test_a_point_seen_only_in_usage_counts_as_having_emitted() -> None:
|
|
"""The write-path arms and the pull sources never reach `sources` — they
|
|
live in note_usage_events — so reading only retrieval_logs would report
|
|
every one of them as silent."""
|
|
quiet = {p["source"] for p in _silent_surfaces(
|
|
{}, {"by_source": {"write_path_place": {}}},
|
|
{"by_source": {"enter_project": {}}}, active=True)}
|
|
assert "write_path_place" not in quiet
|
|
assert "enter_project" not in quiet
|
|
|
|
|
|
# ── settings parsing ──────────────────────────────────────────────────────
|
|
|
|
@pytest.mark.parametrize("raw,fallback,want", [
|
|
("50", 30, 50),
|
|
("0.05", 0.02, 0.05),
|
|
("", 30, 30),
|
|
("banana", 30, 30), # a malformed setting must not take the readout down
|
|
(None, 0.02, 0.02),
|
|
])
|
|
def test_a_setting_falls_back_rather_than_raising(raw, fallback, want) -> None:
|
|
assert _num(raw, fallback) == want
|
|
|
|
|
|
# ── read_and_unacted / outcomes_never_recorded (#4213, milestone 419) ─────
|
|
#
|
|
# The pair exists because ZERO OUTCOMES IS AMBIGUOUS, and getting that wrong
|
|
# would have been this milestone's own failure mode in miniature: a window
|
|
# with no outcome rows cannot tell "every rule was ignored" from "nothing
|
|
# reports outcomes yet". Reporting the first when the truth is the second
|
|
# manufactures a finding out of an unwired feature — #3311, where a statistic
|
|
# that could not vary was read as a fact about the corpus.
|
|
|
|
def ru(**kw) -> dict:
|
|
base = {
|
|
"distinct_rules_surfaced": 0, "distinct_rules_pulled": 0,
|
|
"distinct_rules_acted": 0, "applied": 0, "departed": 0,
|
|
}
|
|
base.update(kw)
|
|
return base
|
|
|
|
|
|
def test_rules_opened_with_no_outcome_machinery_running_says_so() -> None:
|
|
"""The cold-instrument case, which is what an install looks like the day
|
|
this ships. It must NOT read as "47 rules ignored"."""
|
|
ws = warn({}, rule_usage=ru(distinct_rules_pulled=47))
|
|
assert "outcomes_never_recorded" in codes(ws)
|
|
assert "read_and_unacted" not in codes(ws)
|
|
[w] = [w for w in ws if w["code"] == "outcomes_never_recorded"]
|
|
assert w["numbers"]["opened"] == 47
|
|
# The distinction is in the prose, because the prose is what gets read.
|
|
assert "does NOT mean they were ignored" in w["detail"]
|
|
|
|
|
|
def test_once_outcomes_exist_the_unacted_rules_are_named() -> None:
|
|
"""The instrument is live — some rules recorded an outcome — so the ones
|
|
that did not are a real finding rather than an artefact."""
|
|
ws = warn({}, rule_usage=ru(
|
|
distinct_rules_pulled=20, distinct_rules_acted=6, applied=5, departed=2,
|
|
))
|
|
assert "read_and_unacted" in codes(ws)
|
|
assert "outcomes_never_recorded" not in codes(ws)
|
|
[w] = [w for w in ws if w["code"] == "read_and_unacted"]
|
|
assert w["numbers"]["unacted"] == 14
|
|
assert w["numbers"]["opened"] == 20 and w["numbers"]["acted"] == 6
|
|
assert w["numbers"]["applied"] == 5 and w["numbers"]["departed"] == 2
|
|
|
|
|
|
def test_a_departure_alone_is_enough_to_warm_the_instrument() -> None:
|
|
"""Departures count as outcomes. An install whose every recorded outcome
|
|
is a departure is saying something loudly, and must not be mistaken for
|
|
one that records nothing."""
|
|
ws = warn({}, rule_usage=ru(
|
|
distinct_rules_pulled=9, distinct_rules_acted=2, departed=3,
|
|
))
|
|
assert "read_and_unacted" in codes(ws)
|
|
assert "outcomes_never_recorded" not in codes(ws)
|
|
|
|
|
|
def test_every_opened_rule_acted_on_reports_nothing() -> None:
|
|
ws = warn({}, rule_usage=ru(
|
|
distinct_rules_pulled=4, distinct_rules_acted=4, applied=4,
|
|
))
|
|
assert "read_and_unacted" not in codes(ws)
|
|
assert "outcomes_never_recorded" not in codes(ws)
|
|
|
|
|
|
def test_no_rules_opened_at_all_reports_neither() -> None:
|
|
"""Silence is not a finding. A window where nothing was opened has nothing
|
|
to say about outcomes, and saying it anyway would put a warning on every
|
|
fresh install (rule 115)."""
|
|
ws = warn({}, rule_usage=ru(distinct_rules_surfaced=12))
|
|
assert "read_and_unacted" not in codes(ws)
|
|
assert "outcomes_never_recorded" not in codes(ws)
|
|
|
|
|
|
def test_an_absent_rule_usage_block_is_not_a_finding() -> None:
|
|
"""A failed rule-usage read leaves the keys missing or zero. Neither may
|
|
become a warning, because a warning computed over rows that could not be
|
|
loaded describes the outage, not the corpus (#2663)."""
|
|
assert "read_and_unacted" not in codes(warn({}, rule_usage={}))
|
|
assert "outcomes_never_recorded" not in codes(warn({}, rule_usage={}))
|
|
assert "outcomes_never_recorded" not in codes(
|
|
warn({}, rule_usage={"rule_usage_failed": True})
|
|
)
|