Item 1 of #3072. Both sweeps issued a single-row pg_insert(image_tag) from inside their per-image loop. Steady state that is nothing; a first sweep over a back-catalogue is one round-trip per applied tag, tens of thousands of them. Each chunk now collects its rows and writes them in one statement. The ticket suggested one insert per chunk PER TAG. A single multi-row VALUES carries every tag at once, so it is one statement per chunk full stop — and the sweeps already accumulate across all heads before they commit, so nothing had to be restructured to allow it. Not a new helper: wip_title.apply_wip_image_tags was already doing the chunked ON CONFLICT DO NOTHING insert, so that shape is extracted to services/image_tag_apply.insert_image_tags and all three writers share it. The extraction deliberately leaves wip_title's pre-SELECT behind rather than pulling it into the shared function — the sweeps don't need it (their `skip` sets already exclude applied and rejected images) and it exists only to produce an accurate count, which the sweeps also compute themselves. So the shared primitive returns nothing: psycopg reports rowcount -1 for a multi-row ON CONFLICT DO NOTHING insert, and a count taken from the statement would be a lie rather than an approximation. Ordering note for the system-tag sweep: tag rows are now written after that chunk's PresentationReview rows rather than interleaved before them. Safe — PresentationReview FKs to image_record and tag, not to image_tag. Chunk size stays 5000: 5000 rows x 3 bound params = 15000, inside Postgres' 65535-parameter ceiling with room to spare. tests/test_image_tag_apply.py covers the primitive directly, since it is now the single place three writers can be wrong at once — most importantly that a re-run never restamps a hand-applied tag's source, which would silently poison head training (it excludes the auto sources). Left alone: _insert_presentation_review is still per-row, and the retract path still deletes per-row. Both operate on sets that are small by construction, unlike the apply path. Refs #3072
FabledCurator
Self-hosted media curation — gallery, ML tagging, and subscription-driven downloading in one app. Part of the FabledSword family.
Combines what was ImageRepo (gallery, ML, importer) and GallerySubscriber (gallery-dl wrapper, subscriptions, credential capture) into a single product.
Status
In production. main is continuously deployed — every merge to main builds
and publishes :latest images, so whatever is on main is what is running.
Day-to-day work happens on dev, which publishes :dev images.
What's in here
Five deployable pieces, built by .forgejo/workflows/build.yml:
| Piece | Built from | Image | Role |
|---|---|---|---|
| Web / workers | Dockerfile |
fabledcurator |
Quart API + the built Vue SPA in one image. entrypoint.sh picks the role: web, worker, scheduler. The maintenance-long service is a second worker pinned to the long-running maintenance queue. |
| ML worker | Dockerfile.ml |
fabledcurator-ml |
Same app, plus requirements-ml.txt — tagging and embedding models that run in-container. |
| GPU agent | agent/Dockerfile |
fabledcurator-agent |
Optional desktop-GPU worker (agent/). Leases jobs over HTTP only — never touches the database or Redis. Run it for a burst, stop it to reclaim the card. See agent/README.md. |
| Firefox extension | extension/ |
signed XPI | MV3 extension: pushes platform session cookies into FC and adds a creator as a Source in one click. AMO-signed on main only, then bundled into the web image and served from Settings → Maintenance. See extension/README.md. |
| Data | — | pgvector/pgvector:pg16, redis:7-alpine |
Postgres with pgvector for embeddings; Redis as the Celery broker. |
Quick start
For local development and testing, just:
docker compose up -d
# UI: http://localhost:8080
That uses sane dev defaults baked into docker-compose.yml and the dev
override (docker-compose.override.yml, auto-merged) — local builds, DEBUG
logging, exposed Postgres + Redis ports on the host. No .env required.
For a production-like deployment, override the dev defaults via shell env
or a .env file (see .env.example for the variable names) and use:
docker compose -f docker-compose.yml up -d
# (skips the override so containers pull registry images)
The GPU agent is deployed separately, on the machine with the card —
agent/docker-compose.yml, not this stack.
Deployment posture
FabledCurator is designed to run inside a self-hosted homelab environment over plain HTTP. If you want TLS, terminate it at your reverse proxy. The app does not generate certificates, redirect to HTTPS, or set HSTS.
CI / Forgejo setup
Three workflows: ci.yml (lint, extension-version guard, backend unit tests,
frontend build, integration), extension.yml (extension lint, vitest, XPI
content verification), and build.yml (sign + publish).
The toolchain each job runs in is its container.image, not its runs-on
label. runs-on: python-ci only schedules the job onto a runner; every job
then names the image it actually wants. ci-requirements.md is the current,
authoritative list of images and per-job installs — read that rather than a
copy here, so the two can't drift.
The repo expects one secret:
-
RELEASE_TOKEN— a Forgejo PAT with:write:package+read:package— fordocker pushtogit.fabledsword.comwrite:release— for theext-<version>releases that cache the signed XPIwrite:issue— for issue-management automation
Generate at https://git.fabledsword.com/user/settings/applications. The injected
GITHUB_TOKENcannot be used because it lackswrite:package.
AMO signing additionally needs MOZILLA_AMO_JWT_KEY / MOZILLA_AMO_JWT_SECRET; it runs on
main only and is cached per version, since AMO rejects a re-signed version.
License
Personal project; use at your own discretion.