CI & Build / Python lint (push) Successful in 2s
CI & Build / Plugin hooks (push) Successful in 8s
CI & Build / integration (push) Successful in 43s
CI & Build / TypeScript typecheck (push) Successful in 56s
CI & Build / Python tests (push) Failing after 1m5s
CI & Build / Build & push image (push) Skipped
Milestone 416 step 4's write half. `retrieval_surfaces.py` made the six
push arms describe their `{floor, budget}` the same way; this adds the
three MCP tools that let the model READ that and change it, and the
backup sections that carry the reasons.
The operator's decision, which this implements:
"the floor should be chosen and adjusted by the model using it… the
user should be able to touch it but the model should be the thing
handling it 9 times out of 10."
WHY A REASON IS REQUIRED, AND WHY THE TOOL ARGUES AGAINST PERCENTILES
The milestone originally listed self-tuning as a non-goal on one
measured case, and that case is now the tool's docstring rather than a
prohibition: `report_preference` logged 69 consecutive declines with
the refused record 0.0006 under the bar, and every percentile said
"lower it". The refused record was rule 77 "Extract intent from loose
phrasing" matched against a query about report layout — a false
positive. Lowering would have attached that rule to every completion
report ever written.
What separated the statistic from the correct action was OPENING the
record. So `tune_retrieval` refuses a blank or perfunctory reason,
tells the caller to read `retrieval_telemetry(near_miss_samples=5)`
and the record ids it names, and carries that 69-decline example — an
abstract warning loses to a number. The non-goal that survives is
*statistical* auto-tuning; nothing here reads a percentile and picks a
value.
BACKUP (v16), which is what CI caught
`retrieval_tuning_events` was neither backed up nor excluded, and
#2293's guard said so. It is backed up: `settings` already carried the
numbers, so dropping this would restore an install with six moved
dials and no argument for any of them — precisely the state the table
exists to prevent, and worse now that the model is the one moving
them. One `_retrieval_tuning_event_rows` builder called from both
exporters (snippet #2851); `user_id` travels because a restore has to
remap it, which is why the row builder is not the model's `to_dict()`.
`surface` is a registry name rather than a foreign key, so the history
survives a restore into an install whose ids all differ.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
444 lines
20 KiB
Python
444 lines
20 KiB
Python
"""Unit tests for the backup export contract.
|
|
|
|
This is the no-database lane, so these cover the parts that need none: the
|
|
version/coverage constants, the pure row helpers, and the export dict shape
|
|
(via a mocked session). The full FK-remapping round-trip needs real Postgres
|
|
and belongs in a `@pytest.mark.integration` module — it is not written yet,
|
|
which is why every row helper here is a plain function that can be tested
|
|
without a session.
|
|
"""
|
|
from datetime import datetime, timezone
|
|
from types import SimpleNamespace
|
|
from unittest.mock import patch
|
|
|
|
import pytest
|
|
|
|
from scribe.services import backup
|
|
|
|
|
|
def test_backup_version_is_current():
|
|
"""The bump is the point of the test — a payload section added without
|
|
moving the version produces backups that are structurally different and
|
|
indistinguishable by inspection.
|
|
|
|
(Named for the number it asserted until v10, which is exactly the drift a
|
|
name-carrying-a-value invites; it now says what it checks.)"""
|
|
assert backup.BACKUP_VERSION == 16
|
|
|
|
|
|
def _exportable_note(**over):
|
|
"""A Note-shaped stand-in for the pure row helper. SimpleNamespace, not a
|
|
MagicMock: `_note_rows` calls .isoformat() on the timestamps, and a mock
|
|
would happily return another mock instead of failing."""
|
|
base = dict(
|
|
id=1, user_id=7, title="t", body="b", description=None, tags=["x"],
|
|
parent_id=None, arose_from_id=None,
|
|
project_id=None, milestone_id=None, status=None, priority=None,
|
|
due_date=None, created_at=datetime(2026, 1, 1, tzinfo=timezone.utc),
|
|
updated_at=datetime(2026, 1, 2, tzinfo=timezone.utc),
|
|
note_type="note", task_kind="work", data=None,
|
|
started_at=None, completed_at=None,
|
|
recurrence_rule=None, recurrence_next_spawn_at=None,
|
|
verify_with=None, expires_when=None, verified_at=None,
|
|
)
|
|
base.update(over)
|
|
return SimpleNamespace(**base)
|
|
|
|
|
|
def test_note_rows_carry_the_verification_trio():
|
|
"""Operator judgment — somebody checked this fact, and this is when —
|
|
which nothing downstream can recompute (milestone 317)."""
|
|
[row] = backup._note_rows([_exportable_note(
|
|
verify_with="curl the AMO docs",
|
|
expires_when="AMO starts allowing re-signing",
|
|
verified_at=datetime(2026, 8, 28, tzinfo=timezone.utc),
|
|
)])
|
|
assert row["verify_with"] == "curl the AMO docs"
|
|
assert row["expires_when"] == "AMO starts allowing re-signing"
|
|
assert row["verified_at"] == "2026-08-28T00:00:00+00:00"
|
|
|
|
|
|
def test_a_never_checked_note_exports_a_null_stamp_and_restores_as_one():
|
|
"""The round trip that matters. NULL `verified_at` means nobody has ever
|
|
looked, and it is what sorts FIRST in the sweep. Restoring it as now() —
|
|
which is what `_dt` would do — silently converts the sweep's top result
|
|
into its bottom one."""
|
|
[row] = backup._note_rows([_exportable_note(verify_with="check the runner shell")])
|
|
assert row["verified_at"] is None
|
|
assert backup._dt_or_none(row["verified_at"]) is None
|
|
# ...and the helper that must NOT be used here, for contrast.
|
|
assert backup._dt(row["verified_at"]) is not None
|
|
|
|
|
|
def test_the_record_type_and_kind_survive_the_export():
|
|
"""#3182's headline. Without these two columns a restore reported success
|
|
and handed back a corpus where all 90 snippets and 3 processes were plain
|
|
notes and all 435 issues and the spike were `work` — the entire vocabulary
|
|
milestone 312 and #3128 were about, gone, with nothing to notice it by."""
|
|
[snippet] = backup._note_rows([_exportable_note(note_type="snippet")])
|
|
[issue] = backup._note_rows([_exportable_note(task_kind="issue", status="done")])
|
|
assert snippet["note_type"] == "snippet"
|
|
assert issue["task_kind"] == "issue"
|
|
|
|
|
|
def test_provenance_and_lifecycle_travel():
|
|
started = datetime(2026, 3, 1, tzinfo=timezone.utc)
|
|
[row] = backup._note_rows([_exportable_note(
|
|
arose_from_id=42,
|
|
description="one-liner",
|
|
started_at=started,
|
|
recurrence_rule={"freq": "weekly"},
|
|
)])
|
|
assert row["arose_from_id"] == 42
|
|
assert row["description"] == "one-liner"
|
|
assert row["started_at"] == started.isoformat()
|
|
assert row["recurrence_rule"] == {"freq": "weekly"}
|
|
# Absent lifecycle stamps stay absent — a note that never started must not
|
|
# restore as one that started at restore time.
|
|
assert row["completed_at"] is None
|
|
|
|
|
|
def test_a_milestone_carries_its_plan():
|
|
"""A milestone IS the plan (0066); `body` is its design and intent and
|
|
`description` is only the one-line summary. Dropping it restored every
|
|
plan as a title with no reasoning behind it (#3182)."""
|
|
m = SimpleNamespace(
|
|
id=1, user_id=7, project_id=2, title="t", description="d",
|
|
body="## Goal\n\nthe actual plan", status="active", order_index=0,
|
|
created_at=datetime(2026, 1, 1, tzinfo=timezone.utc),
|
|
updated_at=datetime(2026, 1, 2, tzinfo=timezone.utc),
|
|
)
|
|
[row] = backup._milestone_rows([m])
|
|
assert row["body"] == "## Goal\n\nthe actual plan"
|
|
|
|
|
|
def test_a_repo_binding_carries_the_branch_its_ledger_follows():
|
|
"""#2873. Without `ref` a restored binding silently falls back to the
|
|
default branch and the shape ledger starts accounting for a different
|
|
tree — a wrong answer that looks like a working one."""
|
|
b = SimpleNamespace(user_id=7, project_id=2, repo_key="Scribe", ref="dev")
|
|
[row] = backup._repo_binding_rows([b])
|
|
assert row["ref"] == "dev"
|
|
|
|
|
|
# The table -> (model, row helper) registry the column guard walks. Kept here
|
|
# rather than in the service because it exists only to be introspected: the
|
|
# product code already knows these pairings by calling them.
|
|
def _column_guard_targets():
|
|
from scribe.models.canonical_system import CanonicalSystem
|
|
from scribe.models.code_shape import CodeShape, CodeShapeEvent, CodeShapeUse
|
|
from scribe.models.design_system import DesignSystem, DesignToken
|
|
from scribe.models.milestone import Milestone
|
|
from scribe.models.note import Note
|
|
from scribe.models.note_draft import NoteDraft
|
|
from scribe.models.note_supersession import NoteSupersession
|
|
from scribe.models.note_usage import NoteUsageEvent
|
|
from scribe.models.rule_usage import RuleUsageEvent
|
|
from scribe.models.retrieval_tuning import RetrievalTuningEvent
|
|
from scribe.models.note_version import NoteVersion
|
|
from scribe.models.rule_version import RuleVersion
|
|
from scribe.models.project import Project
|
|
from scribe.models.repo_binding import RepoBinding
|
|
from scribe.models.rulebook import Rule, Rulebook, RulebookTopic, RuleRelation
|
|
from scribe.models.setting import Setting
|
|
from scribe.models.system import RecordSystem, System
|
|
from scribe.models.task_log import TaskLog
|
|
from scribe.models.user import User
|
|
|
|
return {
|
|
"users": (User, backup._user_rows),
|
|
"projects": (Project, backup._project_rows),
|
|
"milestones": (Milestone, backup._milestone_rows),
|
|
"notes": (Note, backup._note_rows),
|
|
"task_logs": (TaskLog, backup._task_log_rows),
|
|
"note_drafts": (NoteDraft, backup._note_draft_rows),
|
|
"note_versions": (NoteVersion, backup._note_version_rows),
|
|
"settings": (Setting, backup._setting_rows),
|
|
"rulebooks": (Rulebook, backup._rulebook_rows),
|
|
"rulebook_topics": (RulebookTopic, backup._topic_rows),
|
|
"rules": (Rule, backup._rule_rows),
|
|
"rule_versions": (RuleVersion, backup._rule_version_rows),
|
|
"systems": (System, lambda rows: backup._system_rows(rows, {})),
|
|
"canonical_systems": (CanonicalSystem, backup._canonical_system_rows),
|
|
"record_systems": (RecordSystem, backup._record_system_rows),
|
|
"note_supersessions": (NoteSupersession, backup._note_supersession_rows),
|
|
"rule_relations": (RuleRelation, backup._rule_relation_rows),
|
|
"note_usage_events": (NoteUsageEvent, backup._usage_event_rows),
|
|
"rule_usage_events": (RuleUsageEvent, backup._rule_usage_event_rows),
|
|
"retrieval_tuning_events": (
|
|
RetrievalTuningEvent, backup._retrieval_tuning_event_rows,
|
|
),
|
|
"design_systems": (DesignSystem, backup._design_system_rows),
|
|
"design_tokens": (DesignToken, backup._design_token_rows),
|
|
"repo_bindings": (RepoBinding, backup._repo_binding_rows),
|
|
"code_shapes": (CodeShape, backup._code_shape_rows),
|
|
"code_shape_events": (CodeShapeEvent, backup._code_shape_event_rows),
|
|
"code_shape_uses": (CodeShapeUse, backup._code_shape_use_rows),
|
|
}
|
|
|
|
|
|
def _stand_in(model):
|
|
"""A real instance of `model` with every column set to a value of roughly
|
|
the right type, so the serialiser runs and we can read which KEYS it
|
|
produced. Values are meaningless; only the shape of the output dict is
|
|
under test.
|
|
|
|
A real instance rather than a MagicMock because several helpers delegate to
|
|
the model's own `to_dict()`, and a mock would return another mock instead
|
|
of a dict. Typed rather than a bare instance because the helpers call
|
|
`.isoformat()` on the timestamps, which `None` does not have.
|
|
"""
|
|
import sqlalchemy as sa
|
|
|
|
row = model()
|
|
for column in model.__table__.columns:
|
|
t = column.type
|
|
if isinstance(t, sa.DateTime):
|
|
value = datetime(2026, 1, 1, tzinfo=timezone.utc)
|
|
elif isinstance(t, sa.Date):
|
|
value = datetime(2026, 1, 1).date()
|
|
elif isinstance(t, sa.Boolean):
|
|
value = False
|
|
elif isinstance(t, sa.Integer):
|
|
value = 1
|
|
elif isinstance(t, sa.ARRAY) or isinstance(getattr(t, "impl", None), sa.ARRAY):
|
|
value = []
|
|
elif isinstance(t, (sa.Text, sa.String)):
|
|
value = "x"
|
|
else:
|
|
# JSON/JSONB and anything exotic. None is what these actually hold
|
|
# most of the time, and no serialiser calls a method on one.
|
|
value = None
|
|
setattr(row, column.name, value)
|
|
return row
|
|
|
|
|
|
@pytest.mark.parametrize("table", sorted(_column_guard_targets()))
|
|
def test_every_column_is_exported_or_declared_excluded(table):
|
|
"""THE COLUMN GUARD (#3182) — _NOT_INCLUDED's shape, one level down.
|
|
|
|
The table guard below catches a whole table going missing. It cannot catch
|
|
a COLUMN going missing from a table it already considers covered, which is
|
|
how nine of them vanished from `notes` alone: note_type and task_kind, so
|
|
every snippet and process restored as a plain note and every issue and
|
|
spike as `work`; arose_from_id, so every provenance edge went; the
|
|
recurrence pair, so recurring tasks stopped recurring. Plus milestones.body
|
|
— which IS the plan — and repo_bindings.ref.
|
|
|
|
Each arrived the same way: added to the model and the migration, both of
|
|
which fail loudly, and never to the serialiser, which fails silently.
|
|
|
|
A new column must now be exported or named in _COLUMN_EXCLUSIONS with a
|
|
reason. Forgetting is no longer expressible.
|
|
"""
|
|
model, helper = _column_guard_targets()[table]
|
|
[row] = helper([_stand_in(model)])
|
|
|
|
columns = {c.name for c in model.__table__.columns}
|
|
missing = columns - set(row)
|
|
declared = backup._COLUMN_EXCLUSIONS[table]
|
|
|
|
assert missing == declared, (
|
|
f"{table}: exported columns and _COLUMN_EXCLUSIONS disagree.\n"
|
|
f" dropped but not declared: {sorted(missing - declared)}\n"
|
|
f" declared but exported anyway: {sorted(declared - missing)}"
|
|
)
|
|
|
|
|
|
def test_the_column_guard_covers_every_table_with_a_row_helper():
|
|
"""The guard is only as good as its registry — a table added to _BACKED_UP
|
|
with a new helper, and not to the registry, would be unguarded and look
|
|
guarded. Join tables have no model class and carry both their columns by
|
|
construction, so they are the only permitted absences."""
|
|
# REAL table names, as _BACKED_UP holds them — not the shorter keys the
|
|
# payload uses for the same sections. Getting this wrong is what the guard
|
|
# caught on its own first run.
|
|
join_tables = {"rule_systems"}
|
|
covered = set(_column_guard_targets()) | join_tables
|
|
assert set(backup._BACKED_UP) - covered == set()
|
|
# And no stale entries: every declaration must name a real target.
|
|
assert set(backup._COLUMN_EXCLUSIONS) == set(_column_guard_targets())
|
|
|
|
|
|
def test_not_included_lists_the_known_gaps():
|
|
# The deferred tables must be surfaced explicitly, not silently dropped.
|
|
# forge_connections is excluded as CREDENTIALS (api_keys reasoning): a
|
|
# backup that carries forge tokens is a token-exfiltration file (#2778).
|
|
for table in ("groups", "project_shares", "note_shares", "api_keys",
|
|
"note_embeddings", "retrieval_logs", "forge_connections"):
|
|
assert table in backup._NOT_INCLUDED
|
|
|
|
|
|
def test_every_table_is_either_backed_up_or_explicitly_excluded():
|
|
"""THE GUARD (#2293), and the only shape of test that catches an ABSENCE.
|
|
|
|
A new table gets a model and a migration — both fail loudly if wrong — and
|
|
then silently never gets a backup section. No error, no warning, and a
|
|
restore that reports success. That is how `systems`, `record_systems`,
|
|
`note_usage_events`, `design_systems`, `design_tokens` and `repo_bindings`
|
|
all went missing, over five migrations, with nothing to notice.
|
|
|
|
Extending the export fixes today. THIS fixes the next one: adding a table
|
|
now fails here until someone either backs it up or states in
|
|
`_NOT_INCLUDED` that it shouldn't be. Either is fine; silence is not.
|
|
"""
|
|
from scribe.models import Base
|
|
|
|
schema = set(Base.metadata.tables)
|
|
accounted = set(backup._BACKED_UP) | set(backup._NOT_INCLUDED)
|
|
|
|
unaccounted = schema - accounted
|
|
assert not unaccounted, (
|
|
f"{len(unaccounted)} table(s) are neither backed up nor explicitly "
|
|
f"excluded: {sorted(unaccounted)}. Add each to backup._BACKED_UP (and "
|
|
f"give it an export + restore section) or to backup._NOT_INCLUDED with "
|
|
f"a reason in the comment above it."
|
|
)
|
|
|
|
# And the reverse: a name in either list that no longer exists is a lie the
|
|
# guard would otherwise keep telling. This half is what caught "embeddings",
|
|
# "invitations" and "password_resets" — three entries that named nothing.
|
|
phantom = accounted - schema
|
|
assert not phantom, (
|
|
f"backup lists table(s) that are not in the schema: {sorted(phantom)}. "
|
|
f"Renamed or dropped — fix the list rather than leaving it to read as "
|
|
f"coverage."
|
|
)
|
|
|
|
|
|
class _Result:
|
|
def scalars(self):
|
|
return self
|
|
|
|
def all(self):
|
|
return []
|
|
|
|
|
|
class _Session:
|
|
async def execute(self, *a, **k):
|
|
return _Result()
|
|
|
|
async def get(self, *a, **k):
|
|
return None
|
|
|
|
|
|
class _CM:
|
|
async def __aenter__(self):
|
|
return _Session()
|
|
|
|
async def __aexit__(self, *a):
|
|
return False
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_export_full_backup_contains_every_declared_section():
|
|
with patch("scribe.services.backup.async_session", lambda: _CM()):
|
|
out = await backup.export_full_backup()
|
|
|
|
assert out["version"] == backup.BACKUP_VERSION
|
|
assert out["scope"] == "full"
|
|
assert "api_keys" in out["_not_included"]
|
|
# The sections v2 silently dropped, the six v5 added, v6's
|
|
# note_supersessions, and v7's code_shapes (all empty here).
|
|
for key in ("rulebooks", "rulebook_topics", "rules",
|
|
"systems", "record_systems", "design_systems",
|
|
"design_tokens", "note_usage_events", "repo_bindings",
|
|
"note_supersessions", "code_shapes", "code_shape_events",
|
|
"code_shape_uses",
|
|
# v16: the reasons beside the settings they explain.
|
|
"retrieval_tuning_events"):
|
|
assert key in out, f"missing export section: {key}"
|
|
assert out[key] == []
|
|
|
|
|
|
def test_supersession_rows_serialise_the_pair():
|
|
"""The row builder is a plain function precisely so it can be tested with
|
|
no database — same reason as the other v5/v6 builders."""
|
|
class _Row:
|
|
def __init__(self, a, b):
|
|
self.superseder_id, self.superseded_id = a, b
|
|
|
|
assert backup._note_supersession_rows([_Row(9, 4), _Row(9, 5)]) == [
|
|
{"superseder_id": 9, "superseded_id": 4},
|
|
{"superseder_id": 9, "superseded_id": 5},
|
|
]
|
|
|
|
|
|
def test_rule_rows_carry_the_verification_fields():
|
|
"""A rule's check must survive a backup.
|
|
|
|
`verify_with`/`expires_when`/`verified_at` (milestone 312) say whether a
|
|
rule is a fact that can go false and when it was last confirmed. A backup
|
|
that drops them restores a rulebook that has forgotten which of its rules
|
|
can rot — the exact blindness the fields were added to end.
|
|
|
|
Column additions do not bump BACKUP_VERSION; only new SECTIONS do. Same
|
|
call made for when_to_apply/arose_from_id in 0088 (commit 6ddb8bf).
|
|
"""
|
|
checked = datetime(2026, 8, 27, 12, 0, tzinfo=timezone.utc)
|
|
row = SimpleNamespace(
|
|
id=1, topic_id=2, project_id=None, title="t", statement="s",
|
|
why="w", how_to_apply="h", order_index=0,
|
|
when_to_apply="when", kind="rule",
|
|
verify_with="cat some/file", expires_when="the file grows a shell",
|
|
verified_at=checked, arose_from_id=99,
|
|
created_at=checked, updated_at=checked,
|
|
)
|
|
out = backup._rule_rows([row])[0]
|
|
|
|
assert out["verify_with"] == "cat some/file"
|
|
assert out["expires_when"] == "the file grows a shell"
|
|
assert out["verified_at"] == checked.isoformat()
|
|
# Provenance was exported from 0088 onward but silently dropped on the way
|
|
# back IN until milestone 312. Export side asserted here; the restore side
|
|
# remaps it through note_id_map.
|
|
assert out["arose_from_id"] == 99
|
|
|
|
|
|
def test_rule_rows_keep_an_unverified_rule_unverified():
|
|
"""NULL verified_at means never checked, and it must round-trip as null.
|
|
|
|
_dt substitutes now() so created_at/updated_at are never null. Reusing it
|
|
here would restore a rule nobody ever checked as though it had just been
|
|
checked — dropping it to the BOTTOM of the sweep it should top. That is
|
|
why _dt_or_none exists.
|
|
"""
|
|
row = SimpleNamespace(
|
|
id=1, topic_id=2, project_id=None, title="t", statement="s",
|
|
why=None, how_to_apply=None, order_index=0,
|
|
when_to_apply=None, kind="rule",
|
|
verify_with=None, expires_when=None, verified_at=None,
|
|
arose_from_id=None,
|
|
created_at=datetime(2026, 8, 27, tzinfo=timezone.utc),
|
|
updated_at=datetime(2026, 8, 27, tzinfo=timezone.utc),
|
|
)
|
|
assert backup._rule_rows([row])[0]["verified_at"] is None
|
|
assert backup._dt_or_none(None) is None
|
|
assert backup._dt_or_none("2026-08-27T12:00:00+00:00") == datetime(
|
|
2026, 8, 27, 12, 0, tzinfo=timezone.utc
|
|
)
|
|
|
|
|
|
def test_rule_rows_carry_the_kind_so_a_preference_does_not_restore_as_a_rule():
|
|
"""FORCE has to survive a backup, and the failure would be silent.
|
|
|
|
A preference that comes back as a rule is not a missing field anyone would
|
|
notice — the rule reads fine, it simply binds when it was only ever meant
|
|
to be how the operator prefers things done. Nothing in the restored
|
|
rulebook says it used to be softer.
|
|
|
|
Asserted with `preference` rather than `rule` on purpose: a fixture
|
|
carrying the DEFAULT would pass just as happily against a `_rule_rows`
|
|
that dropped the field entirely and let the importer's `or "rule"` fill
|
|
the hole back in, which is precisely the bug this guards.
|
|
"""
|
|
stamp = datetime(2026, 9, 10, tzinfo=timezone.utc)
|
|
row = SimpleNamespace(
|
|
id=1, topic_id=2, project_id=None, title="t", statement="s",
|
|
why=None, how_to_apply=None, order_index=0,
|
|
when_to_apply="when", kind="preference",
|
|
verify_with=None, expires_when=None, verified_at=None,
|
|
arose_from_id=None, created_at=stamp, updated_at=stamp,
|
|
)
|
|
assert backup._rule_rows([row])[0]["kind"] == "preference"
|