db: collapse alembic 0001..0087 into one baseline (milestone 328 step 1)
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 5s
Build images / build-agent (push) Successful in 9s
CI / frontend-build (push) Successful in 24s
Build images / build-ml (push) Successful in 42s
CI / backend-lint-and-test (push) Successful in 53s
Build images / build-web (push) Successful in 33s
CI / integration (push) Failing after 3m47s

87 revisions narrating this project's build-out become one file that
creates the schema in a single step. They cost nothing at runtime — all
86 upgrade steps ran in 0.2s (note #3260) — so this is a presentation
change, not a performance one: a new installer should not inherit our
development history to stand up a database.

Deleted: 87 revisions (6,052 lines), the 10 tests/test_migration_*.py
files (483 lines) that asserted intermediate states and backfills which
no longer exist, and backend/app/utils/artist_backfill.py — the only
live module a migration imported, with no other consumer anywhere. That
last one satisfies the operator's separate request to inline it into
0008 and delete the module; the squash removes both outright.

THE REVISION ID IS "0087", NOT "0001", ON PURPOSE. It is the id of the
last revision collapsed, so an existing database is already at head and
`alembic upgrade head` does nothing. The alternative is `alembic stamp`
against live data, and stamp validates NOTHING — it writes a version
string whether or not the schema matches, so a wrong baseline surfaces
later, via the next real migration, with no clean way back. This removes
that operation rather than making it safe. Future revisions run from
0088.

Four things are hand-written because SQLAlchemy metadata does not carry
them, and none fail at generation time:

  1. CREATE EXTENSION vector          — the VECTOR columns cannot be
     created without it, so it is ordered first in upgrade().
  2. CREATE EXTENSION tsm_system_rows — surfaces only when the random
     sample query runs.
  3. the HNSW index on image_record.siglip_embedding, raw SQL because
     create_index cannot express USING hnsw (... vector_cosine_ops).
     The quietest of the four: everything works, similarity search just
     stops using an index.
  4. import pgvector.sqlalchemy.vector — autogenerate EMITS
     pgvector.sqlalchemy.vector.VECTOR references without importing it,
     so the generated file dies with NameError on first run.

The candidate came out of CI (run 4967) as checksummed base64 rather
than a plain cat, because run 4964's cat was truncated mid-line inside a
column definition with the step still green — 29 tables instead of 42,
and it looked entirely plausible. Verified here: 56,582 bytes,
sha256 471acfca69c0…, 42 tables, 66 indexes, 42 drops.

NOT YET PROVEN against the old chain. baseline.yml does that, and it is
step 2's gate; this commit does not claim the schemas match.
This commit is contained in:
2026-08-30 13:45:35 -04:00
parent 8f1ac0c96a
commit 2529b516e6
99 changed files with 872 additions and 6579 deletions
-44
View File
@@ -1,44 +0,0 @@
"""Literal SQL for the FC-2d-vii-c artist backfill / artist-tag delete.
Intentionally pure string constants — NO model/slug imports, NO logic —
so migration 0008 and its test share one drift-proof source of truth.
Backfill steps are ordered primary -> provenance -> artist-tag and each
only touches rows still NULL (idempotent, first match wins). The
artist-tag step matches Artist.name = Tag.name: the importer always
created both from the same artist_name string.
"""
BACKFILL_PRIMARY_SQL = """
UPDATE image_record AS ir
SET artist_id = s.artist_id
FROM post p
JOIN source s ON s.id = p.source_id
WHERE ir.primary_post_id = p.id
AND ir.artist_id IS NULL
"""
BACKFILL_PROVENANCE_SQL = """
UPDATE image_record AS ir
SET artist_id = s.artist_id
FROM (
SELECT DISTINCT ON (ip.image_record_id)
ip.image_record_id, src.artist_id
FROM image_provenance ip
JOIN source src ON src.id = ip.source_id
ORDER BY ip.image_record_id, ip.id
) AS s
WHERE ir.id = s.image_record_id
AND ir.artist_id IS NULL
"""
BACKFILL_TAG_SQL = """
UPDATE image_record AS ir
SET artist_id = a.id
FROM image_tag it
JOIN tag t ON t.id = it.tag_id AND t.kind = 'artist'
JOIN artist a ON a.name = t.name
WHERE it.image_record_id = ir.id
AND ir.artist_id IS NULL
"""
DELETE_ARTIST_TAGS_SQL = "DELETE FROM tag WHERE kind = 'artist'"