Files
FabledCurator/docker-compose.yml
T
bvandeusenandClaude Opus 5 bc4eba636d
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 2s
Build images / build-agent (push) Successful in 7s
Build images / build-ml (push) Successful in 9s
Build images / build-web (push) Successful in 7s
CI / frontend-build (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 31s
CI / integration (push) Successful in 3m44s
fix: the documented install path pulled :dev, not :latest (#3270)
docker-compose.yml pinned :dev on all five app services — web, worker,
scheduler, maintenance-long, ml-worker. The README documents
`docker compose -f docker-compose.yml up -d` as the production path, and
-f means "use only this file", skipping the override and its build:
directives. So Compose pulled image:, and image: was the rolling
development channel. The documented way to install this product shipped
development builds.

It went unnoticed for a structural reason rather than a careless one:
nobody who works on the project takes that path. The operator deploys
from a swarm stack file; contributors get docker-compose.override.yml,
which sets build: for all five services, and build: wins over image:. The
broken path is reachable only by a stranger following the README — which
is exactly the audience that did not exist until now.

:latest, per rule 147: main IS production. It is also what the agent
stack (agent/docker-compose.yml) already pinned, so this makes the two
stacks agree rather than introducing a new convention.

Both paths verified with `docker compose config`, which merges and prints
without starting anything:

  dev path        — build: present on all five, image: not pulled
  -f production   — 0 build: directives, five :latest images resolved

Also checked the base file for anything a stranger could not satisfy:
no host-absolute volume paths, no operator-specific port bindings, no
device mappings. The tag was the only defect in the consumer path.

The comment on web.image is deliberately long (rule 32). A line reading
:latest inside a file a developer is debugging with is exactly the line
someone flips back to :dev to test something and then commits, and the
consequence — strangers silently installing bleeding edge — is invisible
to everyone who works here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017QHszn9H8VBvx5Ke8x1hvw
2026-08-31 16:55:43 -04:00

223 lines
9.3 KiB
YAML

# Base compose stack. Uses ${VAR:-default} interpolation throughout so the
# stack boots with zero config — sane dev defaults baked in. For production
# deployments, override the defaults via shell env vars or a .env file:
#
# DB_PASSWORD=...real... SECRET_KEY=...real... docker compose up
#
# The dev override (docker-compose.override.yml) is auto-merged when you
# run `docker compose up` from this directory and switches images to
# local builds + DEBUG logging.
# Rolling-deploy safety (Swarm / `docker stack deploy`): update one task at a
# time, START the new task before stopping the old (zero-downtime via the ingress
# mesh), and if the new task doesn't reach a healthy state within `monitor`, roll
# back to the previous image automatically. `monitor` is sized above web's
# healthcheck start_period so a broken image that never goes healthy is caught.
# Plain `docker compose up` ignores `deploy:` (it warns + skips), so the dev
# override is unaffected. Referenced by each long-lived service below.
x-deploy-policy: &deploy_policy
update_config:
order: start-first
failure_action: rollback
monitor: 90s
rollback_config:
order: start-first
restart_policy:
condition: any
delay: 10s
# Worker liveness: ping THIS container's celery node over the broker. Lenient
# (60s interval, 3 retries, 60s start_period) so a transient broker blip never
# false-flags a worker into a rollback. `$$HOSTNAME` → `$HOSTNAME` for the shell;
# celery's default node name is celery@<hostname> (the container id).
x-celery-healthcheck: &celery_healthcheck
test: ["CMD-SHELL", "celery -A backend.app.celery_app:celery inspect ping -d celery@$$HOSTNAME --timeout 10 >/dev/null 2>&1"]
interval: 60s
timeout: 15s
retries: 3
start_period: 60s
services:
redis:
image: redis:7-alpine
volumes:
- redis_data:/data
healthcheck:
test: ["CMD", "redis-cli", "ping"]
interval: 10s
timeout: 5s
retries: 5
postgres:
image: pgvector/pgvector:pg16
environment:
POSTGRES_USER: ${DB_USER:-fabledcurator}
POSTGRES_PASSWORD: ${DB_PASSWORD:-fabledcurator_dev}
POSTGRES_DB: ${DB_NAME:-fabledcurator}
volumes:
- postgres_data:/var/lib/postgresql/data
# Docker's default /dev/shm is 64MB; VACUUM (ANALYZE) + parallel queries
# allocate a POSIX dynamic-shared-memory segment in /dev/shm and fail with
# "could not resize shared memory segment ... No space left on device"
# (operator-flagged 2026-06-07, vacuum_analyze on import_task needed 67MB).
# NOTE: a `shm_size:` key is SILENTLY IGNORED under Docker Swarm
# (`docker stack deploy` — the prod deployment here), so size /dev/shm via
# a tmpfs mount instead — honored by both Swarm and plain Compose.
- type: tmpfs
target: /dev/shm
tmpfs:
size: 536870912 # 512 MB
healthcheck:
test: ["CMD-SHELL", "pg_isready -U ${DB_USER:-fabledcurator}"]
interval: 10s
timeout: 5s
retries: 5
web:
# :latest, NOT :dev — this file IS the install path.
#
# `docker compose up -d` merges docker-compose.override.yml, which sets
# build: for all five app services, and a build: wins over image:. So a
# contributor never pulls this tag and is unaffected by what it says.
#
# The tag is consulted only on `docker compose -f docker-compose.yml up -d`
# — the documented production path, which skips the override. That is a
# stranger installing the product, and they must land on the stable channel.
#
# :latest is main, which IS production (rule 147). :dev is the rolling
# bleeding-edge channel we work out of, republished several times a day with
# no stability promise. This file pinned :dev on all five services until
# 2026-08-31 (#3270), so the documented install shipped development builds.
# It went unnoticed because nobody who works on the project takes this path:
# the operator deploys from a swarm stack file, contributors get the
# override. Do not "fix" this back to :dev while debugging — use the
# override, or -f with an explicit tag on the command line.
image: git.fabledsword.com/bvandeusen/fabledcurator:latest
command: ["web"]
# Graceful shutdown: give the container time to drain in-flight work on a
# deploy (docker SIGTERMs, then SIGKILLs after this window — default is only
# 10s, far too short for real jobs). Hypercorn/Celery both warm-shut-down on
# SIGTERM; per-lane values sized to typical task length. Anything that still
# outruns the window is re-queued (task_reject_on_worker_lost) and re-driven
# by the 5-min recovery sweeps, so a kill never corrupts. web = short HTTP
# requests + the occasional file download.
stop_grace_period: 30s
# Liveness for rolling deploys: /api/health is a no-DB 200 (just proves the
# app booted + serves HTTP after `alembic upgrade head`). start_period covers
# the migration + boot so a slow start isn't mis-flagged.
healthcheck:
test: ["CMD-SHELL", "python -c \"import urllib.request,sys; sys.exit(0 if urllib.request.urlopen('http://localhost:8080/api/health', timeout=5).status==200 else 1)\""]
interval: 15s
timeout: 6s
retries: 3
start_period: 40s
deploy: *deploy_policy
ports:
- "${PORT:-8080}:8080"
environment: &app_env
DB_USER: ${DB_USER:-fabledcurator}
DB_PASSWORD: ${DB_PASSWORD:-fabledcurator_dev}
DB_HOST: postgres
DB_PORT: "5432"
DB_NAME: ${DB_NAME:-fabledcurator}
CELERY_BROKER_URL: redis://redis:6379/0
CELERY_RESULT_BACKEND: redis://redis:6379/0
SECRET_KEY: ${SECRET_KEY:-dev_secret_key_not_for_production_change_me}
EXTENSION_API_KEY: ${EXTENSION_API_KEY:-}
LOG_LEVEL: ${LOG_LEVEL:-INFO}
volumes:
- ./images:/images
- ./import:/import
# FC-5 legacy migration: bind-mount the host's ImageRepo images dir
# under /import (FC's existing filesystem scan picks them up). Read-only
# is sufficient — FC copies into /images during the scan. The worker +
# scheduler services see the same /import via their own mounts below
# because of /import volume reuse. Edit the host path to match your
# install before running Settings → Maintenance → Legacy migration.
# - /var/lib/imagerepo/images:/import/imagerepo:ro
depends_on:
postgres: { condition: service_healthy }
redis: { condition: service_healthy }
worker:
image: git.fabledsword.com/bvandeusen/fabledcurator:latest
command: ["worker"]
# Drain in-flight import/thumbnail/download tasks before SIGKILL on deploy.
stop_grace_period: 90s
healthcheck: *celery_healthcheck
deploy: *deploy_policy
environment:
<<: *app_env
CELERY_QUEUES: default,import,thumbnail,download
CELERY_CONCURRENCY: "2"
# /downloads dropped — nothing in the app references it (operator-flagged
# 2026-06-07: it wasn't mapped in prod and everything worked).
volumes:
- ./images:/images
- ./import:/import
depends_on:
postgres: { condition: service_healthy }
redis: { condition: service_healthy }
scheduler:
image: git.fabledsword.com/bvandeusen/fabledcurator:latest
command: ["scheduler"]
# Quick maintenance/scan lane + beat — short tasks, modest drain window.
stop_grace_period: 60s
healthcheck: *celery_healthcheck
deploy: *deploy_policy
environment:
<<: *app_env
CELERY_QUEUES: maintenance,scan
volumes:
- ./images:/images
- ./import:/import
depends_on:
postgres: { condition: service_healthy }
redis: { condition: service_healthy }
# Dedicated lane for long one-shot maintenance (DB backups, library audits,
# admin maintenance). Kept off the scheduler's quick `maintenance` lane so a
# 30-min backup or a multi-chunk audit can never starve the 5-min recovery
# sweeps / vacuum (operator-flagged 2026-06-07). One slot — these are heavy.
maintenance-long:
image: git.fabledsword.com/bvandeusen/fabledcurator:latest
command: ["worker"]
# Longest lane (DB backups, library audits, translation backfill) — give it
# the most room to finish a chunk gracefully. Chunked + idempotent, so a job
# that still outruns this resumes cleanly next run rather than corrupting.
stop_grace_period: 180s
healthcheck: *celery_healthcheck
deploy: *deploy_policy
environment:
<<: *app_env
CELERY_QUEUES: maintenance_long
CELERY_CONCURRENCY: "1"
# Only /images: backups write to /images/_backups, audits read /images, and
# the admin tasks (re-extract/cascade-delete/normalize) operate on /images.
volumes:
- ./images:/images
depends_on:
postgres: { condition: service_healthy }
redis: { condition: service_healthy }
ml-worker:
image: git.fabledsword.com/bvandeusen/fabledcurator-ml:latest
command: ["ml-worker"]
# A single GPU inference pass can run tens of seconds — let it finish.
stop_grace_period: 120s
healthcheck: *celery_healthcheck
deploy: *deploy_policy
environment:
<<: *app_env
volumes:
- ./images:/images:ro
- ./models:/models
depends_on:
postgres: { condition: service_healthy }
redis: { condition: service_healthy }
volumes:
redis_data:
postgres_data: