Files
FabledCurator/docker-compose.single.yml
T
bvandeusenandClaude Opus 5 b74a4c964b
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 6s
extension / lint (push) Successful in 19s
CI / frontend-build (push) Successful in 22s
CI / backend-lint-and-test (push) Successful in 33s
CI / integration (push) Successful in 2m15s
Build images / build-web (push) Successful in 2m52s
Build images / smoke-web (push) Successful in 1m16s
Build images / promote (push) Skipped
fix: the image brings its own PID 1 instead of asking for init: true (4295)
Operator, 2026-09-23: *"it is out of the norm to require this init call we
need to fix this."* Correct, and it is the same mistake as declaring the
healthcheck per service — the image needing a deployment to remember a flag
before it behaves correctly.

PID 1 carries a duty no other process has: every orphaned process in the
container reparents to it and must be reaped or it stays a zombie holding a
PID slot. This app makes orphans in normal operation — six service modules
shell out (gallery_dl, thumbnailer, backup_service, external_fetch,
download_service, download_backends) and celery's prefork pool forks children
that spawn them.

Whatever the role, something that is not an init ends up as PID 1:
supervisord for `all`, hypercorn for `web`, celery for a worker. `init: true`
covered that, and cost correctness the moment it was forgotten or silently
dropped — an older Swarm, a plain `docker run`, a compose file someone
copied. No signal either way.

So tini goes in the image and is the ENTRYPOINT. `docker run <image>` is
correct on its own now, `init: true` comes out of docker-compose.single.yml,
and nothing downstream has to know. The smoke asserts /proc/1/comm is tini,
read from /proc because the runtime stage installs no `ps`.

## A correction to what I told the operator

I justified `init: true` by saying supervisord "has no idea about orphans it
never started". That is very likely wrong: supervisord's reaper calls
waitpid(-1) and logs "reaped unknown pid" for children it did not spawn, so
it does reap orphans. I asserted the mechanism without checking it, and could
not check it here — `supervisor` is not installed in this environment.

It does not change this commit. tini is correct whichever way that lands, and
it covers the single-role containers too, where celery or hypercorn is PID 1
and the subprocess-spawning is heaviest. But the reason I gave was not a
verified one and should not have been stated as fact.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-23 10:18:47 -04:00

105 lines
4.3 KiB
YAML

# FabledCurator in three containers — the install path.
#
# docker compose -f docker-compose.single.yml up -d
#
# Milestone 422 step 5. FabledCurator runs web and every worker lane inside
# ONE container, with Postgres and Redis beside it. How much work each lane
# does is then a dial in the web UI (Settings -> Activity -> Worker lanes),
# live, with no compose edit and no restart.
#
# THE MULTI-SERVICE STACK IS NOT REPLACED. `docker-compose.yml` still runs the
# five app services separately and is the right shape for a Swarm deployment
# spread across hosts, where per-service rolling rollback and placement
# constraints matter. This file is the adopter path: one box, one command.
#
# What consolidating costs, stated here rather than discovered later:
# - Everything shares one host, so there is no spreading work across nodes.
# - Rollback is all-or-nothing; there is no rolling back `web` alone.
# - One stop timeout for the whole container, sized to the slowest lane.
#
# NOT a cost, recorded so it is not rediscovered and raised again: the
# multi-service stack mounts /images:ro on ml-worker and one container cannot
# mount one path two ways. Operator ruled that a non-issue (2026-09-22) — it
# is the same codebase either way.
#
# FabledCurator has no authentication. Whatever can reach ${PORT} is an
# administrator, including over the stored platform session cookies. Do not
# publish this port beyond a network you trust — see "Before you expose it"
# in README.md.
services:
redis:
image: redis:7-alpine
volumes:
- redis_data:/data
healthcheck:
test: ["CMD", "redis-cli", "ping"]
interval: 10s
timeout: 5s
retries: 5
restart: unless-stopped
postgres:
image: pgvector/pgvector:pg16
environment:
POSTGRES_USER: ${DB_USER:-curator}
POSTGRES_PASSWORD: ${DB_PASSWORD:-postgres}
POSTGRES_DB: ${DB_NAME:-curator}
volumes:
- postgres_data:/var/lib/postgresql/data
# pgvector index builds and the gallery's TABLESAMPLE reads both want more
# shared memory than docker's 64MB default.
shm_size: 512m
healthcheck:
test: ["CMD-SHELL", "pg_isready -U ${DB_USER:-curator} -d ${DB_NAME:-curator}"]
interval: 10s
timeout: 5s
retries: 5
restart: unless-stopped
fabledcurator:
image: git.fabledsword.com/bvandeusen/fabledcurator:latest
# No `command:`. Everything — hypercorn plus one celery process per lane
# under supervisord — is what the image does by default, and supervisord's
# config is generated from the application's own lane table so the two
# cannot disagree. `command: ["all"]` still works and means the same thing.
# Sized to the SLOWEST lane, not the average. maintenance_long runs DB
# backups, library audits and translation backfill, and gets 180s to
# finish a chunk; the lanes stop in parallel, so this covers the max
# rather than their sum. Below this, a routine restart becomes a SIGKILL
# mid-backup — which is recoverable (the work is chunked and idempotent)
# but wastes however long it had run.
stop_grace_period: 200s
# No healthcheck here either. The image declares one that reads the role
# it is running, and for this one that means BOTH halves: hypercorn
# answers AND every lane is answering the broker. A web-only check would
# report a healthy container while every lane inside it had crashed —
# the failure mode consolidation creates, since docker can no longer see
# the lanes as separate services.
environment:
DB_USER: ${DB_USER:-curator}
DB_PASSWORD: ${DB_PASSWORD:-postgres}
DB_HOST: postgres
DB_PORT: "5432"
DB_NAME: ${DB_NAME:-curator}
CELERY_BROKER_URL: redis://redis:6379/0
CELERY_RESULT_BACKEND: redis://redis:6379/0
SECRET_KEY: ${SECRET_KEY:-change-me-before-you-expose-this}
EXTENSION_API_KEY: ${EXTENSION_API_KEY:-}
LOG_LEVEL: ${LOG_LEVEL:-INFO}
ports:
- "${PORT:-8080}:8080"
volumes:
- ${IMAGES_DIR:-./images}:/images
# Read-only. The filesystem scan copies out of here and never writes to
# it, so a mistake cannot reach the source library.
- ${IMPORT_DIR:-./import}:/import:ro
depends_on:
postgres: { condition: service_healthy }
redis: { condition: service_healthy }
restart: unless-stopped
volumes:
redis_data:
postgres_data: