CI & Build / Python lint (push) Failing after 3s
CI & Build / Plugin hooks (push) Successful in 8s
CI & Build / integration (push) Successful in 30s
CI & Build / TypeScript typecheck (push) Successful in 33s
CI & Build / Python tests (push) Successful in 1m5s
CI & Build / Build & push image (push) Skipped
Milestone 333 step 1. The write-path standing-rule arm is the only retrieval surface in Scribe whose usefulness cannot be observed — and, not coincidentally, the only one that has never declined to fire. 296 calls, zero zero-result, 100% clearing its threshold, while every other surface declines most of the time (#3311, and re-measured in note #3430). `retrieval_logs` gives it scores; scores say what the ranker thought, never whether the hint landed. WHY A SIBLING TABLE AND NOT A COLUMN ON note_usage_events. The row carries no note-specific field and the readout is the same shape, which is the strongest case for sharing that note #3163 admits. What decides against it is identity at RESTORE: the note importer maps note_id through note_id_map, so a rule id parked in that column comes back attached to whatever note holds that number in the target database. Not dropped — reattached. The restore reports success, the counters are populated, and every one is about the wrong record, with no other field to disagree with. rule_versions made the same call for the same reason; this is the third rule-side sibling and it reads like the first two. FK-free on rule_id and user_id, matching note_usage_events / retrieval_logs / app_logs, and deliberately unlike rule_versions. A version belongs to a rule's history and dies with it; telemetry outlives what it describes. Deleting a rule must not erase the evidence that it was surfaced forty times and opened never, because that evidence is the case for having deleted it. The service uses `background.spawn` rather than a third copy of the strong-reference dance — that module's own docstring says new callers should, and a fourth copy is how one of them drifts. The AppLog canary #2663 demands is kept, and since `rule_usage` needed exactly `note_usage`'s semantics, that canary moved into `background.report_telemetry_failure` and note_usage now calls it. `retrieval_telemetry` deliberately keeps its own: its canary is a different shape (one process-wide flag, no AppLog row), so repointing it would change behaviour rather than consolidate it. No ambient bucket, and that is a decision. The note twin splits ranked from ambient surfacings because enter_project and the skill sync deliver records without choosing them (#2477). Rules have the same problem waiting — list_always_on_rules loads them wholesale — but nothing emits here yet, so an empty AMBIENT_SOURCES would be machinery pretending to a distinction the data does not contain. `source` stays granular, so the split stays a readout-level change needing no migration. Backup carries it (v14). The round-trip test seeds a NOTE alongside the rule so the target database has a note id to collide with — without that decoy, a restore running rule ids through the wrong map would merely drop them and the test would pass by absence, rather than failing on the populated-and-wrong result that is the actual hazard. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TcCs1CcQ1ormdnzSshKqvN
429 lines
19 KiB
Python
429 lines
19 KiB
Python
"""Unit tests for the backup export contract.
|
|
|
|
This is the no-database lane, so these cover the parts that need none: the
|
|
version/coverage constants, the pure row helpers, and the export dict shape
|
|
(via a mocked session). The full FK-remapping round-trip needs real Postgres
|
|
and belongs in a `@pytest.mark.integration` module — it is not written yet,
|
|
which is why every row helper here is a plain function that can be tested
|
|
without a session.
|
|
"""
|
|
from datetime import datetime, timezone
|
|
from types import SimpleNamespace
|
|
from unittest.mock import patch
|
|
|
|
import pytest
|
|
|
|
from scribe.services import backup
|
|
|
|
|
|
def test_backup_version_is_current():
|
|
"""The bump is the point of the test — a payload section added without
|
|
moving the version produces backups that are structurally different and
|
|
indistinguishable by inspection.
|
|
|
|
(Named for the number it asserted until v10, which is exactly the drift a
|
|
name-carrying-a-value invites; it now says what it checks.)"""
|
|
assert backup.BACKUP_VERSION == 14
|
|
|
|
|
|
def _exportable_note(**over):
|
|
"""A Note-shaped stand-in for the pure row helper. SimpleNamespace, not a
|
|
MagicMock: `_note_rows` calls .isoformat() on the timestamps, and a mock
|
|
would happily return another mock instead of failing."""
|
|
base = dict(
|
|
id=1, user_id=7, title="t", body="b", description=None, tags=["x"],
|
|
parent_id=None, arose_from_id=None,
|
|
project_id=None, milestone_id=None, status=None, priority=None,
|
|
due_date=None, created_at=datetime(2026, 1, 1, tzinfo=timezone.utc),
|
|
updated_at=datetime(2026, 1, 2, tzinfo=timezone.utc),
|
|
note_type="note", task_kind="work", data=None,
|
|
started_at=None, completed_at=None,
|
|
recurrence_rule=None, recurrence_next_spawn_at=None,
|
|
verify_with=None, expires_when=None, verified_at=None,
|
|
)
|
|
base.update(over)
|
|
return SimpleNamespace(**base)
|
|
|
|
|
|
def test_note_rows_carry_the_verification_trio():
|
|
"""Operator judgment — somebody checked this fact, and this is when —
|
|
which nothing downstream can recompute (milestone 317)."""
|
|
[row] = backup._note_rows([_exportable_note(
|
|
verify_with="curl the AMO docs",
|
|
expires_when="AMO starts allowing re-signing",
|
|
verified_at=datetime(2026, 8, 28, tzinfo=timezone.utc),
|
|
)])
|
|
assert row["verify_with"] == "curl the AMO docs"
|
|
assert row["expires_when"] == "AMO starts allowing re-signing"
|
|
assert row["verified_at"] == "2026-08-28T00:00:00+00:00"
|
|
|
|
|
|
def test_a_never_checked_note_exports_a_null_stamp_and_restores_as_one():
|
|
"""The round trip that matters. NULL `verified_at` means nobody has ever
|
|
looked, and it is what sorts FIRST in the sweep. Restoring it as now() —
|
|
which is what `_dt` would do — silently converts the sweep's top result
|
|
into its bottom one."""
|
|
[row] = backup._note_rows([_exportable_note(verify_with="check the runner shell")])
|
|
assert row["verified_at"] is None
|
|
assert backup._dt_or_none(row["verified_at"]) is None
|
|
# ...and the helper that must NOT be used here, for contrast.
|
|
assert backup._dt(row["verified_at"]) is not None
|
|
|
|
|
|
def test_the_record_type_and_kind_survive_the_export():
|
|
"""#3182's headline. Without these two columns a restore reported success
|
|
and handed back a corpus where all 90 snippets and 3 processes were plain
|
|
notes and all 435 issues and the spike were `work` — the entire vocabulary
|
|
milestone 312 and #3128 were about, gone, with nothing to notice it by."""
|
|
[snippet] = backup._note_rows([_exportable_note(note_type="snippet")])
|
|
[issue] = backup._note_rows([_exportable_note(task_kind="issue", status="done")])
|
|
assert snippet["note_type"] == "snippet"
|
|
assert issue["task_kind"] == "issue"
|
|
|
|
|
|
def test_provenance_and_lifecycle_travel():
|
|
started = datetime(2026, 3, 1, tzinfo=timezone.utc)
|
|
[row] = backup._note_rows([_exportable_note(
|
|
arose_from_id=42,
|
|
description="one-liner",
|
|
started_at=started,
|
|
recurrence_rule={"freq": "weekly"},
|
|
)])
|
|
assert row["arose_from_id"] == 42
|
|
assert row["description"] == "one-liner"
|
|
assert row["started_at"] == started.isoformat()
|
|
assert row["recurrence_rule"] == {"freq": "weekly"}
|
|
# Absent lifecycle stamps stay absent — a note that never started must not
|
|
# restore as one that started at restore time.
|
|
assert row["completed_at"] is None
|
|
|
|
|
|
def test_a_milestone_carries_its_plan():
|
|
"""A milestone IS the plan (0066); `body` is its design and intent and
|
|
`description` is only the one-line summary. Dropping it restored every
|
|
plan as a title with no reasoning behind it (#3182)."""
|
|
m = SimpleNamespace(
|
|
id=1, user_id=7, project_id=2, title="t", description="d",
|
|
body="## Goal\n\nthe actual plan", status="active", order_index=0,
|
|
created_at=datetime(2026, 1, 1, tzinfo=timezone.utc),
|
|
updated_at=datetime(2026, 1, 2, tzinfo=timezone.utc),
|
|
)
|
|
[row] = backup._milestone_rows([m])
|
|
assert row["body"] == "## Goal\n\nthe actual plan"
|
|
|
|
|
|
def test_a_repo_binding_carries_the_branch_its_ledger_follows():
|
|
"""#2873. Without `ref` a restored binding silently falls back to the
|
|
default branch and the shape ledger starts accounting for a different
|
|
tree — a wrong answer that looks like a working one."""
|
|
b = SimpleNamespace(user_id=7, project_id=2, repo_key="Scribe", ref="dev")
|
|
[row] = backup._repo_binding_rows([b])
|
|
assert row["ref"] == "dev"
|
|
|
|
|
|
# The table -> (model, row helper) registry the column guard walks. Kept here
|
|
# rather than in the service because it exists only to be introspected: the
|
|
# product code already knows these pairings by calling them.
|
|
def _column_guard_targets():
|
|
from scribe.models.canonical_system import CanonicalSystem
|
|
from scribe.models.code_shape import CodeShape, CodeShapeEvent, CodeShapeUse
|
|
from scribe.models.design_system import DesignSystem, DesignToken
|
|
from scribe.models.milestone import Milestone
|
|
from scribe.models.note import Note
|
|
from scribe.models.note_draft import NoteDraft
|
|
from scribe.models.note_supersession import NoteSupersession
|
|
from scribe.models.note_usage import NoteUsageEvent
|
|
from scribe.models.rule_usage import RuleUsageEvent
|
|
from scribe.models.note_version import NoteVersion
|
|
from scribe.models.rule_version import RuleVersion
|
|
from scribe.models.project import Project
|
|
from scribe.models.repo_binding import RepoBinding
|
|
from scribe.models.rulebook import Rule, Rulebook, RulebookTopic, RuleRelation
|
|
from scribe.models.setting import Setting
|
|
from scribe.models.system import RecordSystem, System
|
|
from scribe.models.task_log import TaskLog
|
|
from scribe.models.user import User
|
|
|
|
return {
|
|
"users": (User, backup._user_rows),
|
|
"projects": (Project, backup._project_rows),
|
|
"milestones": (Milestone, backup._milestone_rows),
|
|
"notes": (Note, backup._note_rows),
|
|
"task_logs": (TaskLog, backup._task_log_rows),
|
|
"note_drafts": (NoteDraft, backup._note_draft_rows),
|
|
"note_versions": (NoteVersion, backup._note_version_rows),
|
|
"settings": (Setting, backup._setting_rows),
|
|
"rulebooks": (Rulebook, backup._rulebook_rows),
|
|
"rulebook_topics": (RulebookTopic, backup._topic_rows),
|
|
"rules": (Rule, backup._rule_rows),
|
|
"rule_versions": (RuleVersion, backup._rule_version_rows),
|
|
"systems": (System, lambda rows: backup._system_rows(rows, {})),
|
|
"canonical_systems": (CanonicalSystem, backup._canonical_system_rows),
|
|
"record_systems": (RecordSystem, backup._record_system_rows),
|
|
"note_supersessions": (NoteSupersession, backup._note_supersession_rows),
|
|
"rule_relations": (RuleRelation, backup._rule_relation_rows),
|
|
"note_usage_events": (NoteUsageEvent, backup._usage_event_rows),
|
|
"rule_usage_events": (RuleUsageEvent, backup._rule_usage_event_rows),
|
|
"design_systems": (DesignSystem, backup._design_system_rows),
|
|
"design_tokens": (DesignToken, backup._design_token_rows),
|
|
"repo_bindings": (RepoBinding, backup._repo_binding_rows),
|
|
"code_shapes": (CodeShape, backup._code_shape_rows),
|
|
"code_shape_events": (CodeShapeEvent, backup._code_shape_event_rows),
|
|
"code_shape_uses": (CodeShapeUse, backup._code_shape_use_rows),
|
|
}
|
|
|
|
|
|
def _stand_in(model):
|
|
"""A real instance of `model` with every column set to a value of roughly
|
|
the right type, so the serialiser runs and we can read which KEYS it
|
|
produced. Values are meaningless; only the shape of the output dict is
|
|
under test.
|
|
|
|
A real instance rather than a MagicMock because several helpers delegate to
|
|
the model's own `to_dict()`, and a mock would return another mock instead
|
|
of a dict. Typed rather than a bare instance because the helpers call
|
|
`.isoformat()` on the timestamps, which `None` does not have.
|
|
"""
|
|
import sqlalchemy as sa
|
|
|
|
row = model()
|
|
for column in model.__table__.columns:
|
|
t = column.type
|
|
if isinstance(t, sa.DateTime):
|
|
value = datetime(2026, 1, 1, tzinfo=timezone.utc)
|
|
elif isinstance(t, sa.Date):
|
|
value = datetime(2026, 1, 1).date()
|
|
elif isinstance(t, sa.Boolean):
|
|
value = False
|
|
elif isinstance(t, sa.Integer):
|
|
value = 1
|
|
elif isinstance(t, sa.ARRAY) or isinstance(getattr(t, "impl", None), sa.ARRAY):
|
|
value = []
|
|
elif isinstance(t, (sa.Text, sa.String)):
|
|
value = "x"
|
|
else:
|
|
# JSON/JSONB and anything exotic. None is what these actually hold
|
|
# most of the time, and no serialiser calls a method on one.
|
|
value = None
|
|
setattr(row, column.name, value)
|
|
return row
|
|
|
|
|
|
@pytest.mark.parametrize("table", sorted(_column_guard_targets()))
|
|
def test_every_column_is_exported_or_declared_excluded(table):
|
|
"""THE COLUMN GUARD (#3182) — _NOT_INCLUDED's shape, one level down.
|
|
|
|
The table guard below catches a whole table going missing. It cannot catch
|
|
a COLUMN going missing from a table it already considers covered, which is
|
|
how nine of them vanished from `notes` alone: note_type and task_kind, so
|
|
every snippet and process restored as a plain note and every issue and
|
|
spike as `work`; arose_from_id, so every provenance edge went; the
|
|
recurrence pair, so recurring tasks stopped recurring. Plus milestones.body
|
|
— which IS the plan — and repo_bindings.ref.
|
|
|
|
Each arrived the same way: added to the model and the migration, both of
|
|
which fail loudly, and never to the serialiser, which fails silently.
|
|
|
|
A new column must now be exported or named in _COLUMN_EXCLUSIONS with a
|
|
reason. Forgetting is no longer expressible.
|
|
"""
|
|
model, helper = _column_guard_targets()[table]
|
|
[row] = helper([_stand_in(model)])
|
|
|
|
columns = {c.name for c in model.__table__.columns}
|
|
missing = columns - set(row)
|
|
declared = backup._COLUMN_EXCLUSIONS[table]
|
|
|
|
assert missing == declared, (
|
|
f"{table}: exported columns and _COLUMN_EXCLUSIONS disagree.\n"
|
|
f" dropped but not declared: {sorted(missing - declared)}\n"
|
|
f" declared but exported anyway: {sorted(declared - missing)}"
|
|
)
|
|
|
|
|
|
def test_the_column_guard_covers_every_table_with_a_row_helper():
|
|
"""The guard is only as good as its registry — a table added to _BACKED_UP
|
|
with a new helper, and not to the registry, would be unguarded and look
|
|
guarded. Join tables have no model class and carry both their columns by
|
|
construction, so they are the only permitted absences."""
|
|
# REAL table names, as _BACKED_UP holds them — not the shorter keys the
|
|
# payload uses for the same sections. Getting this wrong is what the guard
|
|
# caught on its own first run.
|
|
join_tables = {
|
|
"project_rulebook_subscriptions", "project_rule_suppressions",
|
|
"project_topic_suppressions", "project_rulebook_exclusions",
|
|
"rule_systems",
|
|
}
|
|
covered = set(_column_guard_targets()) | join_tables
|
|
assert set(backup._BACKED_UP) - covered == set()
|
|
# And no stale entries: every declaration must name a real target.
|
|
assert set(backup._COLUMN_EXCLUSIONS) == set(_column_guard_targets())
|
|
|
|
|
|
def test_not_included_lists_the_known_gaps():
|
|
# The deferred tables must be surfaced explicitly, not silently dropped.
|
|
# forge_connections is excluded as CREDENTIALS (api_keys reasoning): a
|
|
# backup that carries forge tokens is a token-exfiltration file (#2778).
|
|
for table in ("groups", "project_shares", "note_shares", "api_keys",
|
|
"note_embeddings", "retrieval_logs", "forge_connections"):
|
|
assert table in backup._NOT_INCLUDED
|
|
|
|
|
|
def test_every_table_is_either_backed_up_or_explicitly_excluded():
|
|
"""THE GUARD (#2293), and the only shape of test that catches an ABSENCE.
|
|
|
|
A new table gets a model and a migration — both fail loudly if wrong — and
|
|
then silently never gets a backup section. No error, no warning, and a
|
|
restore that reports success. That is how `systems`, `record_systems`,
|
|
`note_usage_events`, `design_systems`, `design_tokens` and `repo_bindings`
|
|
all went missing, over five migrations, with nothing to notice.
|
|
|
|
Extending the export fixes today. THIS fixes the next one: adding a table
|
|
now fails here until someone either backs it up or states in
|
|
`_NOT_INCLUDED` that it shouldn't be. Either is fine; silence is not.
|
|
"""
|
|
from scribe.models import Base
|
|
|
|
schema = set(Base.metadata.tables)
|
|
accounted = set(backup._BACKED_UP) | set(backup._NOT_INCLUDED)
|
|
|
|
unaccounted = schema - accounted
|
|
assert not unaccounted, (
|
|
f"{len(unaccounted)} table(s) are neither backed up nor explicitly "
|
|
f"excluded: {sorted(unaccounted)}. Add each to backup._BACKED_UP (and "
|
|
f"give it an export + restore section) or to backup._NOT_INCLUDED with "
|
|
f"a reason in the comment above it."
|
|
)
|
|
|
|
# And the reverse: a name in either list that no longer exists is a lie the
|
|
# guard would otherwise keep telling. This half is what caught "embeddings",
|
|
# "invitations" and "password_resets" — three entries that named nothing.
|
|
phantom = accounted - schema
|
|
assert not phantom, (
|
|
f"backup lists table(s) that are not in the schema: {sorted(phantom)}. "
|
|
f"Renamed or dropped — fix the list rather than leaving it to read as "
|
|
f"coverage."
|
|
)
|
|
|
|
|
|
def test_join_table_row_helpers_are_pure():
|
|
subs = [SimpleNamespace(project_id=1, rulebook_id=2)]
|
|
rsup = [SimpleNamespace(project_id=1, rule_id=9)]
|
|
tsup = [SimpleNamespace(project_id=1, topic_id=7)]
|
|
assert backup._subscription_rows(subs) == [{"project_id": 1, "rulebook_id": 2}]
|
|
assert backup._rule_suppression_rows(rsup) == [{"project_id": 1, "rule_id": 9}]
|
|
assert backup._topic_suppression_rows(tsup) == [{"project_id": 1, "topic_id": 7}]
|
|
|
|
|
|
class _Result:
|
|
def scalars(self):
|
|
return self
|
|
|
|
def all(self):
|
|
return []
|
|
|
|
|
|
class _Session:
|
|
async def execute(self, *a, **k):
|
|
return _Result()
|
|
|
|
async def get(self, *a, **k):
|
|
return None
|
|
|
|
|
|
class _CM:
|
|
async def __aenter__(self):
|
|
return _Session()
|
|
|
|
async def __aexit__(self, *a):
|
|
return False
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_export_full_backup_contains_every_declared_section():
|
|
with patch("scribe.services.backup.async_session", lambda: _CM()):
|
|
out = await backup.export_full_backup()
|
|
|
|
assert out["version"] == backup.BACKUP_VERSION
|
|
assert out["scope"] == "full"
|
|
assert "api_keys" in out["_not_included"]
|
|
# The sections v2 silently dropped, the six v5 added, v6's
|
|
# note_supersessions, and v7's code_shapes (all empty here).
|
|
for key in ("rulebooks", "rulebook_topics", "rules",
|
|
"rulebook_subscriptions", "rule_suppressions",
|
|
"topic_suppressions",
|
|
"systems", "record_systems", "design_systems",
|
|
"design_tokens", "note_usage_events", "repo_bindings",
|
|
"note_supersessions", "code_shapes", "code_shape_events",
|
|
"code_shape_uses", "rulebook_exclusions"):
|
|
assert key in out, f"missing export section: {key}"
|
|
assert out[key] == []
|
|
|
|
|
|
def test_supersession_rows_serialise_the_pair():
|
|
"""The row builder is a plain function precisely so it can be tested with
|
|
no database — same reason as the other v5/v6 builders."""
|
|
class _Row:
|
|
def __init__(self, a, b):
|
|
self.superseder_id, self.superseded_id = a, b
|
|
|
|
assert backup._note_supersession_rows([_Row(9, 4), _Row(9, 5)]) == [
|
|
{"superseder_id": 9, "superseded_id": 4},
|
|
{"superseder_id": 9, "superseded_id": 5},
|
|
]
|
|
|
|
|
|
def test_rule_rows_carry_the_verification_fields():
|
|
"""A rule's check must survive a backup.
|
|
|
|
`verify_with`/`expires_when`/`verified_at` (milestone 312) say whether a
|
|
rule is a fact that can go false and when it was last confirmed. A backup
|
|
that drops them restores a rulebook that has forgotten which of its rules
|
|
can rot — the exact blindness the fields were added to end.
|
|
|
|
Column additions do not bump BACKUP_VERSION; only new SECTIONS do. Same
|
|
call made for when_to_apply/tier/arose_from_id in 0088 (commit 6ddb8bf).
|
|
"""
|
|
checked = datetime(2026, 8, 27, 12, 0, tzinfo=timezone.utc)
|
|
row = SimpleNamespace(
|
|
id=1, topic_id=2, project_id=None, title="t", statement="s",
|
|
why="w", how_to_apply="h", order_index=0,
|
|
when_to_apply="when", tier="conditional",
|
|
verify_with="cat some/file", expires_when="the file grows a shell",
|
|
verified_at=checked, arose_from_id=99,
|
|
created_at=checked, updated_at=checked,
|
|
)
|
|
out = backup._rule_rows([row])[0]
|
|
|
|
assert out["verify_with"] == "cat some/file"
|
|
assert out["expires_when"] == "the file grows a shell"
|
|
assert out["verified_at"] == checked.isoformat()
|
|
# Provenance was exported from 0088 onward but silently dropped on the way
|
|
# back IN until milestone 312. Export side asserted here; the restore side
|
|
# remaps it through note_id_map.
|
|
assert out["arose_from_id"] == 99
|
|
|
|
|
|
def test_rule_rows_keep_an_unverified_rule_unverified():
|
|
"""NULL verified_at means never checked, and it must round-trip as null.
|
|
|
|
_dt substitutes now() so created_at/updated_at are never null. Reusing it
|
|
here would restore a rule nobody ever checked as though it had just been
|
|
checked — dropping it to the BOTTOM of the sweep it should top. That is
|
|
why _dt_or_none exists.
|
|
"""
|
|
row = SimpleNamespace(
|
|
id=1, topic_id=2, project_id=None, title="t", statement="s",
|
|
why=None, how_to_apply=None, order_index=0,
|
|
when_to_apply=None, tier="always_on",
|
|
verify_with=None, expires_when=None, verified_at=None,
|
|
arose_from_id=None,
|
|
created_at=datetime(2026, 8, 27, tzinfo=timezone.utc),
|
|
updated_at=datetime(2026, 8, 27, tzinfo=timezone.utc),
|
|
)
|
|
assert backup._rule_rows([row])[0]["verified_at"] is None
|
|
assert backup._dt_or_none(None) is None
|
|
assert backup._dt_or_none("2026-08-27T12:00:00+00:00") == datetime(
|
|
2026, 8, 27, 12, 0, tzinfo=timezone.utc
|
|
)
|