CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 5s
Build images / build-agent (push) Successful in 7s
extension / lint (push) Successful in 19s
CI / frontend-build (push) Successful in 24s
CI / backend-lint-and-test (push) Successful in 32s
Build images / build-web (push) Successful in 1m41s
CI / integration (push) Successful in 2m7s
Build images / smoke-web (push) Failing after 12m36s
Build images / promote (push) Skipped
Operator, 2026-09-23: *"I also want to see that we remove the need for the command line of the configuration in the consolidated version."* `CMD` was `web`, so the single-container layout only worked if you knew to ask for it by name. A compose file that forgot `command: ["all"]` got a web server with nothing processing its queues — a gallery that loads, accepts an import, and never finishes one. Nothing errors; it just never progresses. Now `docker run fabledcurator` with no command starts hypercorn plus every lane under supervisord. `entrypoint.sh`'s own default moves with it, since the two are doors to the same decision and a disagreement would only show up as `--entrypoint` behaving differently from a plain run. `docker-compose.single.yml` drops its `command:` line; `["all"]` still works and still means the same thing. The multi-service stack is untouched — every service there names its role explicitly, which is what makes it the multi-service stack. ## And CI now actually boots it This is the gap I should have named when I reported milestone 422 at 7/7 and did not. Measured, not inferred: the smoke booted role `web` only (build.yml:1454), nothing in CI ran `all`, `docker-compose.single.yml` was read as TEXT by one test checking stop_grace_period and never run, and test_gen_supervisord asserts the generated config against the lane table without ever handing it to supervisord. So the shape this milestone is NAMED for had started nowhere. Steps 5-7 were marked done on evidence that did not cover it, and the operator is about to collapse their production stack onto exactly that. The smoke now boots the image with NO command — checking the Dockerfile CMD, the entrypoint default and the role together, the way an adopter gets it — and asserts `healthcheck_all`, which was itself never executed. That check passes only when hypercorn answers AND every lane in the table answers the broker; a web-only check goes green with every worker dead, which is the failure mode consolidation creates. It then prints `supervisorctl status`, so a lane that is merely restart-looping is visible rather than inferred. Cheap because ml ships at 0 slots and disabled: nothing loads a model, and the lane answers `inspect` with its consumers cancelled, which is what healthy means for a disabled lane. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
116 lines
4.8 KiB
YAML
116 lines
4.8 KiB
YAML
# FabledCurator in three containers — the install path.
|
|
#
|
|
# docker compose -f docker-compose.single.yml up -d
|
|
#
|
|
# Milestone 422 step 5. FabledCurator runs web and every worker lane inside
|
|
# ONE container, with Postgres and Redis beside it. How much work each lane
|
|
# does is then a dial in the web UI (Settings -> Activity -> Worker lanes),
|
|
# live, with no compose edit and no restart.
|
|
#
|
|
# THE MULTI-SERVICE STACK IS NOT REPLACED. `docker-compose.yml` still runs the
|
|
# five app services separately and is the right shape for a Swarm deployment
|
|
# spread across hosts, where per-service rolling rollback and placement
|
|
# constraints matter. This file is the adopter path: one box, one command.
|
|
#
|
|
# What consolidating costs, stated here rather than discovered later:
|
|
# - Everything shares one host, so there is no spreading work across nodes.
|
|
# - Rollback is all-or-nothing; there is no rolling back `web` alone.
|
|
# - One stop timeout for the whole container, sized to the slowest lane.
|
|
#
|
|
# NOT a cost, recorded so it is not rediscovered and raised again: the
|
|
# multi-service stack mounts /images:ro on ml-worker and one container cannot
|
|
# mount one path two ways. Operator ruled that a non-issue (2026-09-22) — it
|
|
# is the same codebase either way.
|
|
#
|
|
# FabledCurator has no authentication. Whatever can reach ${PORT} is an
|
|
# administrator, including over the stored platform session cookies. Do not
|
|
# publish this port beyond a network you trust — see "Before you expose it"
|
|
# in README.md.
|
|
|
|
services:
|
|
redis:
|
|
image: redis:7-alpine
|
|
volumes:
|
|
- redis_data:/data
|
|
healthcheck:
|
|
test: ["CMD", "redis-cli", "ping"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 5
|
|
restart: unless-stopped
|
|
|
|
postgres:
|
|
image: pgvector/pgvector:pg16
|
|
environment:
|
|
POSTGRES_USER: ${DB_USER:-curator}
|
|
POSTGRES_PASSWORD: ${DB_PASSWORD:-postgres}
|
|
POSTGRES_DB: ${DB_NAME:-curator}
|
|
volumes:
|
|
- postgres_data:/var/lib/postgresql/data
|
|
# pgvector index builds and the gallery's TABLESAMPLE reads both want more
|
|
# shared memory than docker's 64MB default.
|
|
shm_size: 512m
|
|
healthcheck:
|
|
test: ["CMD-SHELL", "pg_isready -U ${DB_USER:-curator} -d ${DB_NAME:-curator}"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 5
|
|
restart: unless-stopped
|
|
|
|
fabledcurator:
|
|
image: git.fabledsword.com/bvandeusen/fabledcurator:latest
|
|
# No `command:`. Everything — hypercorn plus one celery process per lane
|
|
# under supervisord — is what the image does by default, and supervisord's
|
|
# config is generated from the application's own lane table so the two
|
|
# cannot disagree. `command: ["all"]` still works and means the same thing.
|
|
# tini as PID 1, in front of supervisord. supervisord reaps its own
|
|
# children, but a container's PID 1 also inherits orphans from anywhere
|
|
# below — celery's prefork pool and gallery-dl's subprocesses both make
|
|
# them. Without this they accumulate as zombies for the life of the
|
|
# container.
|
|
init: true
|
|
# Sized to the SLOWEST lane, not the average. maintenance_long runs DB
|
|
# backups, library audits and translation backfill, and gets 180s to
|
|
# finish a chunk; the lanes stop in parallel, so this covers the max
|
|
# rather than their sum. Below this, a routine restart becomes a SIGKILL
|
|
# mid-backup — which is recoverable (the work is chunked and idempotent)
|
|
# but wastes however long it had run.
|
|
stop_grace_period: 200s
|
|
# BOTH halves: hypercorn answers AND every configured lane is answering
|
|
# the broker. A web-only check would report a healthy container while
|
|
# every lane inside it had crashed — the failure mode consolidation
|
|
# creates, since docker can no longer see the lanes as separate services.
|
|
healthcheck:
|
|
test: ["CMD", "python", "-m", "backend.app.scripts.healthcheck_all"]
|
|
interval: 30s
|
|
timeout: 15s
|
|
retries: 3
|
|
# Covers alembic + hypercorn boot + four celery workers registering.
|
|
start_period: 90s
|
|
environment:
|
|
DB_USER: ${DB_USER:-curator}
|
|
DB_PASSWORD: ${DB_PASSWORD:-postgres}
|
|
DB_HOST: postgres
|
|
DB_PORT: "5432"
|
|
DB_NAME: ${DB_NAME:-curator}
|
|
CELERY_BROKER_URL: redis://redis:6379/0
|
|
CELERY_RESULT_BACKEND: redis://redis:6379/0
|
|
SECRET_KEY: ${SECRET_KEY:-change-me-before-you-expose-this}
|
|
EXTENSION_API_KEY: ${EXTENSION_API_KEY:-}
|
|
LOG_LEVEL: ${LOG_LEVEL:-INFO}
|
|
ports:
|
|
- "${PORT:-8080}:8080"
|
|
volumes:
|
|
- ${IMAGES_DIR:-./images}:/images
|
|
# Read-only. The filesystem scan copies out of here and never writes to
|
|
# it, so a mistake cannot reach the source library.
|
|
- ${IMPORT_DIR:-./import}:/import:ro
|
|
depends_on:
|
|
postgres: { condition: service_healthy }
|
|
redis: { condition: service_healthy }
|
|
restart: unless-stopped
|
|
|
|
volumes:
|
|
redis_data:
|
|
postgres_data:
|