CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 36s
Build images / build-web (push) Successful in 1m9s
Build images / smoke-web (push) Skipped
Build images / build-ml (push) Successful in 2m8s
Build images / promote (push) Skipped
CI / integration (push) Successful in 2m37s
Milestone 422 step 5. `docker compose -f docker-compose.single.yml up -d` gives three containers — FabledCurator, Postgres, Redis — where the stack previously needed seven. THE MULTI-SERVICE STACK IS KEPT. docker-compose.yml still runs the five app services separately and remains the right shape for a Swarm deployment spread across hosts, where per-service rolling rollback and placement constraints matter. This adds a compose file; it deletes none. `entrypoint.sh all` GENERATES the supervisord config from worker_lanes.LANES and execs it as PID 1. Generated rather than checked in because a static .conf would spell out each lane's -Q list, making a FIFTH hand-kept copy of the queue names — after celery_app.task_routes and the three collapsed in steps 1, 2 and 4. Every one of those had already drifted when found. Generating gives a stronger guarantee than "they match today": a lane added to LANES gets a process, and a queue cannot end up with no consumer because someone missed a file. supervisord over s6-overlay: one pip dependency on an image already Python, with per-program stop timeouts and stopasgroup. The process-group part is not a detail — celery's prefork pool forks children, and a TERM reaching only the parent leaves them orphaned holding tasks. s6's advantage (PID-1 signal and zombie handling) comes from `init: true` instead. Nothing in FC talks to the supervisor, so the choice is reversible without touching product code. FOUR LANES, NOT FIVE. The ml lane is skipped: torch and the ML requirements live only in Dockerfile.ml until step 6 merges the images, so an `ml` program here would fail to import on every restart forever. `--with-ml` is the flag step 6 turns on. THREE BUGS FOUND BY READING IT BACK, none of which the first tests caught: 1. `environment=CELERY_QUEUES=default,import,thumbnail,download` — supervisord parses that key as a COMMA-separated list, so it reads as CELERY_QUEUES=default plus three malformed entries and the worker lane would have consumed only `default`. Silent: the worker starts, reports healthy, never picks up an import. Now quoted, and the test asserts the quoted form rather than the bare substring, which passed either way. 2. The generator emitted `entrypoint.sh <lane.name>`, but `maintenance_long` is not a role — compose runs it as the plain `worker` role with different queues. Lane now carries `entrypoint_role`, and a test reads entrypoint.sh to assert every role a lane names actually exists. 3. The `scheduler` role hardcoded --concurrency=1, ignoring CELERY_CONCURRENCY. Harmless while only compose started it and set none; with a generated value being passed, the lane would have sat at 1 until the reconcile noticed, with nothing saying why. The healthcheck asserts BOTH halves — hypercorn answers and every configured lane is answering the broker. That is the failure mode consolidation creates: docker can no longer see the lanes as separate services, so a web-only check would report a healthy container with every lane inside it dead. It deliberately ignores the `enabled` flag: a disabled lane still has a running process with its consumers cancelled, and marking the container unhealthy for turning tagging off would be wrong. stop_grace_period 200s, sized to the slowest lane (maintenance_long at 180s) rather than the average, with a test asserting no program's stopwaitsecs can exceed what compose allows. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
114 lines
4.6 KiB
YAML
114 lines
4.6 KiB
YAML
# FabledCurator in three containers — the install path.
|
|
#
|
|
# docker compose -f docker-compose.single.yml up -d
|
|
#
|
|
# Milestone 422 step 5. FabledCurator runs web and every worker lane inside
|
|
# ONE container, with Postgres and Redis beside it. How much work each lane
|
|
# does is then a dial in the web UI (Settings -> Activity -> Worker lanes),
|
|
# live, with no compose edit and no restart.
|
|
#
|
|
# THE MULTI-SERVICE STACK IS NOT REPLACED. `docker-compose.yml` still runs the
|
|
# five app services separately and is the right shape for a Swarm deployment
|
|
# spread across hosts, where per-service rolling rollback and placement
|
|
# constraints matter. This file is the adopter path: one box, one command.
|
|
#
|
|
# What consolidating costs, stated here rather than discovered later:
|
|
# - Everything shares one host, so there is no spreading work across nodes.
|
|
# - Rollback is all-or-nothing; there is no rolling back `web` alone.
|
|
# - The ML lane cannot write to the library read-only any more — in the
|
|
# multi-service stack ml-worker mounts /images:ro, and one container
|
|
# cannot mount one path two ways.
|
|
# - One stop timeout for the whole container, sized to the slowest lane.
|
|
#
|
|
# FabledCurator has no authentication. Whatever can reach ${PORT} is an
|
|
# administrator, including over the stored platform session cookies. Do not
|
|
# publish this port beyond a network you trust — see "Before you expose it"
|
|
# in README.md.
|
|
|
|
services:
|
|
redis:
|
|
image: redis:7-alpine
|
|
volumes:
|
|
- redis_data:/data
|
|
healthcheck:
|
|
test: ["CMD", "redis-cli", "ping"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 5
|
|
restart: unless-stopped
|
|
|
|
postgres:
|
|
image: pgvector/pgvector:pg16
|
|
environment:
|
|
POSTGRES_USER: ${DB_USER:-curator}
|
|
POSTGRES_PASSWORD: ${DB_PASSWORD:-postgres}
|
|
POSTGRES_DB: ${DB_NAME:-curator}
|
|
volumes:
|
|
- postgres_data:/var/lib/postgresql/data
|
|
# pgvector index builds and the gallery's TABLESAMPLE reads both want more
|
|
# shared memory than docker's 64MB default.
|
|
shm_size: 512m
|
|
healthcheck:
|
|
test: ["CMD-SHELL", "pg_isready -U ${DB_USER:-curator} -d ${DB_NAME:-curator}"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 5
|
|
restart: unless-stopped
|
|
|
|
fabledcurator:
|
|
image: git.fabledsword.com/bvandeusen/fabledcurator:latest
|
|
# Everything: hypercorn plus one celery process per lane, under
|
|
# supervisord, whose config is generated from the application's own lane
|
|
# table so the two cannot disagree.
|
|
command: ["all"]
|
|
# tini as PID 1, in front of supervisord. supervisord reaps its own
|
|
# children, but a container's PID 1 also inherits orphans from anywhere
|
|
# below — celery's prefork pool and gallery-dl's subprocesses both make
|
|
# them. Without this they accumulate as zombies for the life of the
|
|
# container.
|
|
init: true
|
|
# Sized to the SLOWEST lane, not the average. maintenance_long runs DB
|
|
# backups, library audits and translation backfill, and gets 180s to
|
|
# finish a chunk; the lanes stop in parallel, so this covers the max
|
|
# rather than their sum. Below this, a routine restart becomes a SIGKILL
|
|
# mid-backup — which is recoverable (the work is chunked and idempotent)
|
|
# but wastes however long it had run.
|
|
stop_grace_period: 200s
|
|
# BOTH halves: hypercorn answers AND every configured lane is answering
|
|
# the broker. A web-only check would report a healthy container while
|
|
# every lane inside it had crashed — the failure mode consolidation
|
|
# creates, since docker can no longer see the lanes as separate services.
|
|
healthcheck:
|
|
test: ["CMD", "python", "-m", "backend.app.scripts.healthcheck_all"]
|
|
interval: 30s
|
|
timeout: 15s
|
|
retries: 3
|
|
# Covers alembic + hypercorn boot + four celery workers registering.
|
|
start_period: 90s
|
|
environment:
|
|
DB_USER: ${DB_USER:-curator}
|
|
DB_PASSWORD: ${DB_PASSWORD:-postgres}
|
|
DB_HOST: postgres
|
|
DB_PORT: "5432"
|
|
DB_NAME: ${DB_NAME:-curator}
|
|
CELERY_BROKER_URL: redis://redis:6379/0
|
|
CELERY_RESULT_BACKEND: redis://redis:6379/0
|
|
SECRET_KEY: ${SECRET_KEY:-change-me-before-you-expose-this}
|
|
EXTENSION_API_KEY: ${EXTENSION_API_KEY:-}
|
|
LOG_LEVEL: ${LOG_LEVEL:-INFO}
|
|
ports:
|
|
- "${PORT:-8080}:8080"
|
|
volumes:
|
|
- ${IMAGES_DIR:-./images}:/images
|
|
# Read-only. The filesystem scan copies out of here and never writes to
|
|
# it, so a mistake cannot reach the source library.
|
|
- ${IMPORT_DIR:-./import}:/import:ro
|
|
depends_on:
|
|
postgres: { condition: service_healthy }
|
|
redis: { condition: service_healthy }
|
|
restart: unless-stopped
|
|
|
|
volumes:
|
|
redis_data:
|
|
postgres_data:
|