CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 6s
extension / lint (push) Successful in 19s
CI / frontend-build (push) Successful in 22s
CI / backend-lint-and-test (push) Successful in 33s
CI / integration (push) Successful in 2m15s
Build images / build-web (push) Successful in 2m52s
Build images / smoke-web (push) Successful in 1m16s
Build images / promote (push) Skipped
Operator, 2026-09-23: *"it is out of the norm to require this init call we need to fix this."* Correct, and it is the same mistake as declaring the healthcheck per service — the image needing a deployment to remember a flag before it behaves correctly. PID 1 carries a duty no other process has: every orphaned process in the container reparents to it and must be reaped or it stays a zombie holding a PID slot. This app makes orphans in normal operation — six service modules shell out (gallery_dl, thumbnailer, backup_service, external_fetch, download_service, download_backends) and celery's prefork pool forks children that spawn them. Whatever the role, something that is not an init ends up as PID 1: supervisord for `all`, hypercorn for `web`, celery for a worker. `init: true` covered that, and cost correctness the moment it was forgotten or silently dropped — an older Swarm, a plain `docker run`, a compose file someone copied. No signal either way. So tini goes in the image and is the ENTRYPOINT. `docker run <image>` is correct on its own now, `init: true` comes out of docker-compose.single.yml, and nothing downstream has to know. The smoke asserts /proc/1/comm is tini, read from /proc because the runtime stage installs no `ps`. ## A correction to what I told the operator I justified `init: true` by saying supervisord "has no idea about orphans it never started". That is very likely wrong: supervisord's reaper calls waitpid(-1) and logs "reaped unknown pid" for children it did not spawn, so it does reap orphans. I asserted the mechanism without checking it, and could not check it here — `supervisor` is not installed in this environment. It does not change this commit. tini is correct whichever way that lands, and it covers the single-role containers too, where celery or hypercorn is PID 1 and the subprocess-spawning is heaviest. But the reason I gave was not a verified one and should not have been stated as fact. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
105 lines
4.3 KiB
YAML
105 lines
4.3 KiB
YAML
# FabledCurator in three containers — the install path.
|
|
#
|
|
# docker compose -f docker-compose.single.yml up -d
|
|
#
|
|
# Milestone 422 step 5. FabledCurator runs web and every worker lane inside
|
|
# ONE container, with Postgres and Redis beside it. How much work each lane
|
|
# does is then a dial in the web UI (Settings -> Activity -> Worker lanes),
|
|
# live, with no compose edit and no restart.
|
|
#
|
|
# THE MULTI-SERVICE STACK IS NOT REPLACED. `docker-compose.yml` still runs the
|
|
# five app services separately and is the right shape for a Swarm deployment
|
|
# spread across hosts, where per-service rolling rollback and placement
|
|
# constraints matter. This file is the adopter path: one box, one command.
|
|
#
|
|
# What consolidating costs, stated here rather than discovered later:
|
|
# - Everything shares one host, so there is no spreading work across nodes.
|
|
# - Rollback is all-or-nothing; there is no rolling back `web` alone.
|
|
# - One stop timeout for the whole container, sized to the slowest lane.
|
|
#
|
|
# NOT a cost, recorded so it is not rediscovered and raised again: the
|
|
# multi-service stack mounts /images:ro on ml-worker and one container cannot
|
|
# mount one path two ways. Operator ruled that a non-issue (2026-09-22) — it
|
|
# is the same codebase either way.
|
|
#
|
|
# FabledCurator has no authentication. Whatever can reach ${PORT} is an
|
|
# administrator, including over the stored platform session cookies. Do not
|
|
# publish this port beyond a network you trust — see "Before you expose it"
|
|
# in README.md.
|
|
|
|
services:
|
|
redis:
|
|
image: redis:7-alpine
|
|
volumes:
|
|
- redis_data:/data
|
|
healthcheck:
|
|
test: ["CMD", "redis-cli", "ping"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 5
|
|
restart: unless-stopped
|
|
|
|
postgres:
|
|
image: pgvector/pgvector:pg16
|
|
environment:
|
|
POSTGRES_USER: ${DB_USER:-curator}
|
|
POSTGRES_PASSWORD: ${DB_PASSWORD:-postgres}
|
|
POSTGRES_DB: ${DB_NAME:-curator}
|
|
volumes:
|
|
- postgres_data:/var/lib/postgresql/data
|
|
# pgvector index builds and the gallery's TABLESAMPLE reads both want more
|
|
# shared memory than docker's 64MB default.
|
|
shm_size: 512m
|
|
healthcheck:
|
|
test: ["CMD-SHELL", "pg_isready -U ${DB_USER:-curator} -d ${DB_NAME:-curator}"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 5
|
|
restart: unless-stopped
|
|
|
|
fabledcurator:
|
|
image: git.fabledsword.com/bvandeusen/fabledcurator:latest
|
|
# No `command:`. Everything — hypercorn plus one celery process per lane
|
|
# under supervisord — is what the image does by default, and supervisord's
|
|
# config is generated from the application's own lane table so the two
|
|
# cannot disagree. `command: ["all"]` still works and means the same thing.
|
|
# Sized to the SLOWEST lane, not the average. maintenance_long runs DB
|
|
# backups, library audits and translation backfill, and gets 180s to
|
|
# finish a chunk; the lanes stop in parallel, so this covers the max
|
|
# rather than their sum. Below this, a routine restart becomes a SIGKILL
|
|
# mid-backup — which is recoverable (the work is chunked and idempotent)
|
|
# but wastes however long it had run.
|
|
stop_grace_period: 200s
|
|
# No healthcheck here either. The image declares one that reads the role
|
|
# it is running, and for this one that means BOTH halves: hypercorn
|
|
# answers AND every lane is answering the broker. A web-only check would
|
|
# report a healthy container while every lane inside it had crashed —
|
|
# the failure mode consolidation creates, since docker can no longer see
|
|
# the lanes as separate services.
|
|
environment:
|
|
DB_USER: ${DB_USER:-curator}
|
|
DB_PASSWORD: ${DB_PASSWORD:-postgres}
|
|
DB_HOST: postgres
|
|
DB_PORT: "5432"
|
|
DB_NAME: ${DB_NAME:-curator}
|
|
CELERY_BROKER_URL: redis://redis:6379/0
|
|
CELERY_RESULT_BACKEND: redis://redis:6379/0
|
|
SECRET_KEY: ${SECRET_KEY:-change-me-before-you-expose-this}
|
|
EXTENSION_API_KEY: ${EXTENSION_API_KEY:-}
|
|
LOG_LEVEL: ${LOG_LEVEL:-INFO}
|
|
ports:
|
|
- "${PORT:-8080}:8080"
|
|
volumes:
|
|
- ${IMAGES_DIR:-./images}:/images
|
|
# Read-only. The filesystem scan copies out of here and never writes to
|
|
# it, so a mistake cannot reach the source library.
|
|
- ${IMPORT_DIR:-./import}:/import:ro
|
|
depends_on:
|
|
postgres: { condition: service_healthy }
|
|
redis: { condition: service_healthy }
|
|
restart: unless-stopped
|
|
|
|
volumes:
|
|
redis_data:
|
|
postgres_data:
|