db: collapse alembic 0001..0089 into one baseline (#3266)
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
CI / extension-version (push) Successful in 3s
Build images / build-agent (push) Successful in 8s
CI / frontend-build (push) Successful in 17s
CI / backend-lint-and-test (push) Successful in 30s
Build images / build-web (push) Successful in 2m12s
Build images / build-ml (push) Successful in 2m52s
CI / integration (push) Failing after 3m42s

89 files and 6,300 lines become one file of 807. Nothing about the
resulting schema changes; what goes away is the requirement that a new
installation replay our development history to arrive at it.

revision = "0089", down_revision = None. That pairing IS the migration
strategy for existing installs, not a detail of it: a deployed database
already has alembic_version = '0089' from running the real 0089, so
alembic reads the version table, sees head reached, and does nothing. No
stamp is required — which matters, because `alembic stamp` writes a
version string without validating anything about the schema it is writing
it against, and a wrong stamp is indistinguishable from a right one until
the next migration fails. An empty database runs the file and records
0089. Both paths converge. The next migration is 0090, as it would have
been; the numbering is continuous across the collapse on purpose.

Autogenerate produced nearly all of this unaided, which was NOT true of
the first attempt — that one was reverted because the generator silently
dropped eleven indexes and three uniqueness guarantees. #3275 put those
on the models first, so the HNSW index with its opclass, the COALESCE
expression index, the partial uniques, 107 server_defaults and the enum
CHECKs are all emitted now. Doing the reconciliation before the squash,
rather than after, is what made this work.

Hand-added, because none of it can live in a model:

  * CREATE EXTENSION vector / tsm_system_rows (0001, 0004) — database
    objects, not table metadata.
  * The pgvector import. Autogenerate writes qualified
    pgvector.sqlalchemy.vector.VECTOR references without importing the
    package, so its own output cannot run (run 4988).
  * THE TWO SEED ROWS. 0002 and 0003 did not only build schema — each
    inserted a settings singleton, and nothing in the app ever creates
    them: ImportSettings.load() and MLSettings.load() are
    select(...).scalar_one(), which RAISES NoResultFound rather than
    returning None. A models-only baseline would leave both tables empty
    and crash a fresh install on first settings access, while
    baseline.yml reported a perfect schema match. Only running the app
    against a new database finds that.

Not carried over: 0023's DELETE FROM tag and 0047's series deletes, which
are historical cleanups operating on rows an empty database lacks.

downgrade() raises. A baseline's downgrade is "drop every table", which
is a data-loss event wearing a migration as a disguise; offering it as
one invites someone to run it. Restore from a backup.

Also removed, per the plan: the 10 test_migration_*.py files (they assert
intermediate states and backfills that no longer exist — a test that a
column exists is already the model tests' job) and
backend/app/utils/artist_backfill.py, whose only importer was 0008.
Verified no other consumer anywhere in backend/ or tests/.

baseline.yml changes with it. chain_ref now DEFAULTS to 725bf15, since
the tree no longer carries a chain to compare against — that pinned
commit is the last one that does.

And the CheckConstraint repair is removed, because it never fired. I
added it claiming autogenerate re-doubles a constraint name on the round
trip and asserted in b979062 that it was "still correct and still
needed". It is not: autogenerate wraps names in op.f(), which marks them
already-formatted and blocks the convention from re-applying. Tested
against the real candidate line — the regex matches nothing. What
actually fixed the mismatch was 0088's renames alone. The doubling is a
real hazard, but of hand-writing a pre-prefixed name, not of the
generator; the comment asserting otherwise was worse than the dead code
under it.

The header comment is rewritten for the same reason — it described 87
revisions, and claimed the HNSW index could not be expressed in a model,
which #3275 disproved. It now also states plainly what this check CANNOT
see: it compares schema, so a green run means the schema is right, not
that the baseline is.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017QHszn9H8VBvx5Ke8x1hvw
This commit is contained in:
2026-09-01 01:14:37 -04:00
co-authored by Claude Opus 5
parent bc4eba636d
commit 973db73221
102 changed files with 838 additions and 6890 deletions
-58
View File
@@ -1,58 +0,0 @@
"""Smoke test for migration 0002: confirms model classes import and the
tag-kind uniqueness rule shape is correct.
"""
from backend.app.models import (
Base,
ImportBatch,
ImportSettings,
ImportTask,
Tag,
TagKind,
)
def test_new_tables_registered():
expected = {"import_batch", "import_task", "import_settings"}
assert expected.issubset(Base.metadata.tables.keys())
def test_tag_has_kind_and_fandom_id():
cols = {c.name for c in Tag.__table__.columns}
assert "kind" in cols
assert "fandom_id" in cols
assert "namespace" not in cols
def test_tag_kind_enum_values():
# Current TagKind enum after alembic 0023 dropped meta + rating
# (operator-retired 2026-05-26). `artist` is still in the enum
# for backward-compat with historical rows, though new artist
# tags don't get created (Artist row is canonical per FC-2d-vii-c).
expected = {
"artist",
"character",
"fandom",
"general",
"series",
"archive",
"post",
}
assert {k.value for k in TagKind} == expected
def test_image_record_has_integrity_status():
from backend.app.models import ImageRecord
cols = {c.name for c in ImageRecord.__table__.columns}
assert "integrity_status" in cols
def test_import_task_has_state_columns():
cols = {c.name for c in ImportTask.__table__.columns}
for required in ("batch_id", "source_path", "task_type", "status", "result_image_id"):
assert required in cols
def test_import_settings_singleton_constraint():
constraints = {c.name for c in ImportSettings.__table__.constraints}
assert "ck_import_settings_singleton" in constraints
-46
View File
@@ -1,46 +0,0 @@
"""Smoke test for migration 0003: model classes import, schema shape correct."""
from backend.app.models import (
Base,
ImageRecord,
MLSettings,
TagAlias,
TagSuggestionRejection,
)
def test_new_tables_registered():
expected = {
"tag_suggestion_rejection",
"tag_alias",
"ml_settings",
}
assert expected.issubset(Base.metadata.tables.keys())
def test_image_record_columns_renamed():
cols = {c.name for c in ImageRecord.__table__.columns}
# Legacy tagger columns are all gone: tagger_predictions/wd14_* dropped in
# 0046, tagger_model_version + centroid_scores dropped in 0068 (#1199, Camie
# retirement). The SigLIP embedding columns are the live ML fields.
assert "siglip_embedding" in cols
assert "siglip_model_version" in cols
assert "tagger_model_version" not in cols
assert "centroid_scores" not in cols
assert "tagger_predictions" not in cols
assert "wd14_predictions" not in cols
def test_tag_alias_composite_pk():
pk_cols = {c.name for c in TagAlias.__table__.primary_key.columns}
assert pk_cols == {"alias_string", "alias_category"}
def test_ml_settings_singleton_constraint():
names = {c.name for c in MLSettings.__table__.constraints}
assert "ck_ml_settings_singleton" in names
def test_tag_suggestion_rejection_pk():
pk_cols = {c.name for c in TagSuggestionRejection.__table__.primary_key.columns}
assert pk_cols == {"image_record_id", "tag_id"}
-27
View File
@@ -1,27 +0,0 @@
"""Integration: the tsm_system_rows extension is installed by migration 0004.
Needs a real Postgres (CI does not provision one), so integration-marked.
"""
import pytest
from sqlalchemy import text
pytestmark = pytest.mark.integration
@pytest.mark.asyncio
async def test_tsm_system_rows_extension_present(db):
row = (
await db.execute(
text("SELECT 1 FROM pg_extension WHERE extname = 'tsm_system_rows'")
)
).first()
assert row is not None
@pytest.mark.asyncio
async def test_system_rows_sampling_is_usable(db):
# Should parse and execute even on an empty table.
await db.execute(
text("SELECT * FROM image_record TABLESAMPLE SYSTEM_ROWS(1)")
)
-48
View File
@@ -1,48 +0,0 @@
"""FC-2d-iv: post.description + post.attachment_count round-trip."""
from datetime import UTC, datetime
import pytest
from backend.app.models import Artist, Post, Source
pytestmark = pytest.mark.integration
async def _post(db, **post_kwargs):
artist = Artist(name="Nadia", slug="nadia")
db.add(artist)
await db.flush()
src = Source(artist_id=artist.id, platform="web", url="http://x")
db.add(src)
await db.flush()
post = Post(
source_id=src.id, artist_id=artist.id, external_post_id="p1",
post_date=datetime(2026, 3, 1, tzinfo=UTC),
**post_kwargs,
)
db.add(post)
await db.flush()
return post.id
def test_post_has_new_columns():
cols = {c.name for c in Post.__table__.columns}
assert "description" in cols
assert "attachment_count" in cols
@pytest.mark.asyncio
async def test_description_and_attachment_count_round_trip(db):
pid = await _post(db, description="<p>hi</p>", attachment_count=3)
row = await db.get(Post, pid)
assert row.description == "<p>hi</p>"
assert row.attachment_count == 3
@pytest.mark.asyncio
async def test_new_fields_default_null(db):
pid = await _post(db)
row = await db.get(Post, pid)
assert row.description is None
assert row.attachment_count is None
-137
View File
@@ -1,137 +0,0 @@
"""FC-2d-vii-c: image_record.artist_id + backfill + artist-tag delete."""
from datetime import UTC, datetime, timedelta
import pytest
from sqlalchemy import func, select, text
from backend.app.models import (
Artist,
ImageProvenance,
ImageRecord,
Post,
Source,
Tag,
TagKind,
)
from backend.app.models.tag import image_tag
from backend.app.utils.artist_backfill import (
BACKFILL_PRIMARY_SQL,
BACKFILL_PROVENANCE_SQL,
BACKFILL_TAG_SQL,
DELETE_ARTIST_TAGS_SQL,
)
pytestmark = pytest.mark.integration
def test_image_record_has_artist_id_column():
assert "artist_id" in {c.name for c in ImageRecord.__table__.columns}
async def _img(db, n):
rec = ImageRecord(
path=f"/images/bf/{n}.jpg", sha256=f"bf{n:062d}",
size_bytes=1, mime="image/jpeg", width=1, height=1,
origin="imported_filesystem", integrity_status="unknown",
)
rec.created_at = datetime.now(UTC) - timedelta(minutes=n)
db.add(rec)
await db.flush()
return rec
async def _artist_source(db, name, slug):
a = Artist(name=name, slug=slug)
db.add(a)
await db.flush()
s = Source(artist_id=a.id, platform="patreon",
url=f"https://p.test/{slug}")
db.add(s)
await db.flush()
return a, s
async def _run_backfill(db):
await db.execute(text(BACKFILL_PRIMARY_SQL))
await db.execute(text(BACKFILL_PROVENANCE_SQL))
await db.execute(text(BACKFILL_TAG_SQL))
@pytest.mark.asyncio
async def test_backfill_primary_post(db):
rec = await _img(db, 1)
a, s = await _artist_source(db, "Alice", "alice")
post = Post(source_id=s.id, artist_id=a.id, external_post_id="1")
db.add(post)
await db.flush()
rec.primary_post_id = post.id
await db.flush()
await _run_backfill(db)
got = await db.scalar(
select(ImageRecord.artist_id).where(ImageRecord.id == rec.id)
)
assert got == a.id
@pytest.mark.asyncio
async def test_backfill_provenance_fallback(db):
rec = await _img(db, 1)
a, s = await _artist_source(db, "Bob", "bob")
post = Post(source_id=s.id, artist_id=a.id, external_post_id="2")
db.add(post)
await db.flush()
db.add(ImageProvenance(image_record_id=rec.id, post_id=post.id,
source_id=s.id))
await db.flush()
await _run_backfill(db)
got = await db.scalar(
select(ImageRecord.artist_id).where(ImageRecord.id == rec.id)
)
assert got == a.id
@pytest.mark.asyncio
async def test_backfill_artist_tag_by_name(db):
rec = await _img(db, 1)
a = Artist(name="Carol", slug="carol")
db.add(a)
await db.flush()
tag = Tag(name="Carol", kind=TagKind.artist)
db.add(tag)
await db.flush()
await db.execute(image_tag.insert().values(
image_record_id=rec.id, tag_id=tag.id, source="auto"))
await db.flush()
await _run_backfill(db)
got = await db.scalar(
select(ImageRecord.artist_id).where(ImageRecord.id == rec.id)
)
assert got == a.id
@pytest.mark.asyncio
async def test_no_signal_stays_null(db):
rec = await _img(db, 1)
await _run_backfill(db)
got = await db.scalar(
select(ImageRecord.artist_id).where(ImageRecord.id == rec.id)
)
assert got is None
@pytest.mark.asyncio
async def test_delete_removes_only_artist_tags(db):
artist_tag = Tag(name="Dave", kind=TagKind.artist)
general_tag = Tag(name="forest", kind=TagKind.general)
db.add_all([artist_tag, general_tag])
await db.flush()
await db.execute(text(DELETE_ARTIST_TAGS_SQL))
remaining = await db.scalar(
select(func.count()).select_from(Tag).where(Tag.kind == TagKind.artist)
)
assert remaining == 0
survived = await db.scalar(
select(func.count()).select_from(Tag).where(Tag.kind == TagKind.general)
)
assert survived >= 1
-37
View File
@@ -1,37 +0,0 @@
"""FC-2d-iii: post_attachment table + import_batch.attachments column."""
import pytest
from backend.app.models import ImportBatch, PostAttachment
pytestmark = pytest.mark.integration
def test_post_attachment_columns():
cols = {c.name for c in PostAttachment.__table__.columns}
assert {
"id", "post_id", "artist_id", "sha256", "path",
"original_filename", "ext", "mime", "size_bytes", "captured_at",
} <= cols
def test_import_batch_has_attachments_counter():
assert "attachments" in {c.name for c in ImportBatch.__table__.columns}
@pytest.mark.asyncio
async def test_post_attachment_roundtrip(db):
from backend.app.models import Artist
a = Artist(name="Zed", slug="zed")
db.add(a)
await db.flush()
att = PostAttachment(
post_id=None, artist_id=a.id, sha256="z" + "0" * 63,
path="/images/attachments/z00/z.zip", original_filename="pack.zip",
ext=".zip", mime="application/zip", size_bytes=123,
)
db.add(att)
await db.flush()
got = await db.get(PostAttachment, att.id)
assert got.original_filename == "pack.zip" and got.post_id is None
-35
View File
@@ -1,35 +0,0 @@
import pytest
from sqlalchemy.exc import IntegrityError
from backend.app.models import Artist, Source
pytestmark = pytest.mark.integration
@pytest.mark.asyncio
async def test_duplicate_artist_platform_url_rejected(db):
artist = Artist(name="Alice", slug="alice")
db.add(artist)
await db.flush()
db.add(Source(
artist_id=artist.id, platform="patreon",
url="https://patreon.com/alice", enabled=True,
))
await db.flush()
db.add(Source(
artist_id=artist.id, platform="patreon",
url="https://patreon.com/alice", enabled=True,
))
with pytest.raises(IntegrityError):
await db.flush()
@pytest.mark.asyncio
async def test_same_url_under_different_artist_ok(db):
a = Artist(name="A", slug="a")
b = Artist(name="B", slug="b")
db.add_all([a, b])
await db.flush()
db.add(Source(artist_id=a.id, platform="patreon", url="https://x/y", enabled=True))
db.add(Source(artist_id=b.id, platform="patreon", url="https://x/y", enabled=True))
await db.flush() # must NOT raise
-32
View File
@@ -1,32 +0,0 @@
import pytest
from sqlalchemy import inspect, text
pytestmark = pytest.mark.integration
@pytest.mark.asyncio
async def test_credential_has_credential_type_not_kind(db):
cols = (await db.run_sync(
lambda sync_session: [c["name"] for c in inspect(sync_session.bind).get_columns("credential")]
))
assert "credential_type" in cols
assert "kind" not in cols
assert "status" not in cols
assert "last_verified" in cols
@pytest.mark.asyncio
async def test_credential_round_trip(db):
from backend.app.models import Credential
db.add(Credential(
platform="patreon",
credential_type="cookies",
encrypted_blob=b"\x00\x01\x02",
))
await db.flush()
row = (await db.execute(
text("SELECT credential_type, last_verified FROM credential WHERE platform='patreon'")
)).one()
assert row.credential_type == "cookies"
assert row.last_verified is None
-32
View File
@@ -1,32 +0,0 @@
import pytest
from sqlalchemy import select
from backend.app.models import AppSetting
pytestmark = pytest.mark.integration
@pytest.mark.asyncio
async def test_app_setting_table_round_trip(db):
db.add(AppSetting(key="extension_api_key", value="abc123"))
await db.flush()
row = (await db.execute(
select(AppSetting).where(AppSetting.key == "extension_api_key")
)).scalar_one()
assert row.value == "abc123"
assert row.updated_at is not None
@pytest.mark.asyncio
async def test_app_setting_upsert(db):
db.add(AppSetting(key="k", value="v1"))
await db.flush()
row = (await db.execute(
select(AppSetting).where(AppSetting.key == "k")
)).scalar_one()
row.value = "v2"
await db.flush()
again = (await db.execute(
select(AppSetting.value).where(AppSetting.key == "k")
)).scalar_one()
assert again == "v2"
-31
View File
@@ -1,31 +0,0 @@
import pytest
from sqlalchemy import inspect, select
from backend.app.models import ImportSettings
pytestmark = pytest.mark.integration
@pytest.mark.asyncio
async def test_download_event_has_metadata(db):
cols = await db.run_sync(
lambda s: {c["name"]: c for c in inspect(s.bind).get_columns("download_event")}
)
assert "metadata" in cols
assert cols["metadata"]["nullable"] is False
@pytest.mark.asyncio
async def test_import_settings_has_downloader_fields(db):
cols = await db.run_sync(
lambda s: {c["name"]: c for c in inspect(s.bind).get_columns("import_settings")}
)
assert "download_rate_limit_seconds" in cols
assert "download_validate_files" in cols
@pytest.mark.asyncio
async def test_import_settings_defaults(db):
row = (await db.execute(select(ImportSettings).where(ImportSettings.id == 1))).scalar_one()
assert row.download_rate_limit_seconds == 3.0
assert row.download_validate_files is True