CI and images / lint (push) Successful in 3s
CI and images / extension-version (push) Successful in 3s
extension / lint (push) Successful in 16s
CI and images / frontend-build (push) Successful in 19s
CI and images / backend-lint-and-test (push) Successful in 31s
CI and images / integration (push) Successful in 2m22s
CI and images / sign-extension (push) Successful in 3s
CI and images / build-web (push) Successful in 3m7s
CI and images / smoke-web (push) Successful in 59s
CI and images / build-agent (push) Successful in 6m41s
CI and images / promote (push) Successful in 2s
Agent: - The image ran PyPI's CUDA-13 torch 2.14 and onnxruntime-gpu 1.30 on a CUDA 12.9 cudnn-runtime base. requirements.txt had silently replaced the Dockerfile's torch 2.6+cu124, because ultralytics pulls torchvision, which pulls its own torch. That left ~3 GB of base libraries and a ~3 GB torch nothing loaded: 10 GB compressed. - Now: an nvidia/cuda 13.0.3 `base` image, with torch and torchvision installed together from cu130. CUDA and cuDNN come from the nvidia-* pip packages; onnxruntime-gpu declares its [cuda,cudnn] extras. - fc_agent/accel.py preloads those libraries for onnxruntime. It then logs, and reports in /status, whether torch and the ONNX CUDA provider actually got the GPU, since both fall back to the CPU silently. Web image: - Drop opencv-python-headless and onnxruntime, plus the opencv-only apt libs. Both have been listed since the scaffold and nothing in backend/ imports them. - torch/torchvision move to 2.14/0.29, and the unexplained caps are lifted (rule 154). Redis: 8-alpine in both compose files and both CI service containers. That gives an AGPLv3 licence option, where 7.4 was RSAL/SSPL only. The client moves to >=8.1. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
105 lines
4.3 KiB
YAML
105 lines
4.3 KiB
YAML
# FabledCurator in three containers — the install path.
|
|
#
|
|
# docker compose -f docker-compose.single.yml up -d
|
|
#
|
|
# Milestone 422 step 5. FabledCurator runs web and every worker lane inside
|
|
# ONE container, with Postgres and Redis beside it. How much work each lane
|
|
# does is then a dial in the web UI (Settings -> Activity -> Worker lanes),
|
|
# live, with no compose edit and no restart.
|
|
#
|
|
# THE MULTI-SERVICE STACK IS NOT REPLACED. `docker-compose.yml` still runs the
|
|
# five app services separately and is the right shape for a Swarm deployment
|
|
# spread across hosts, where per-service rolling rollback and placement
|
|
# constraints matter. This file is the adopter path: one box, one command.
|
|
#
|
|
# What consolidating costs, stated here rather than discovered later:
|
|
# - Everything shares one host, so there is no spreading work across nodes.
|
|
# - Rollback is all-or-nothing; there is no rolling back `web` alone.
|
|
# - One stop timeout for the whole container, sized to the slowest lane.
|
|
#
|
|
# NOT a cost, recorded so it is not rediscovered and raised again: the
|
|
# multi-service stack mounts /images:ro on ml-worker and one container cannot
|
|
# mount one path two ways. Operator ruled that a non-issue (2026-09-22) — it
|
|
# is the same codebase either way.
|
|
#
|
|
# FabledCurator has no authentication. Whatever can reach ${PORT} is an
|
|
# administrator, including over the stored platform session cookies. Do not
|
|
# publish this port beyond a network you trust — see "Before you expose it"
|
|
# in README.md.
|
|
|
|
services:
|
|
redis:
|
|
image: redis:8-alpine
|
|
volumes:
|
|
- redis_data:/data
|
|
healthcheck:
|
|
test: ["CMD", "redis-cli", "ping"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 5
|
|
restart: unless-stopped
|
|
|
|
postgres:
|
|
image: pgvector/pgvector:pg16
|
|
environment:
|
|
POSTGRES_USER: ${DB_USER:-curator}
|
|
POSTGRES_PASSWORD: ${DB_PASSWORD:-postgres}
|
|
POSTGRES_DB: ${DB_NAME:-curator}
|
|
volumes:
|
|
- postgres_data:/var/lib/postgresql/data
|
|
# pgvector index builds and the gallery's TABLESAMPLE reads both want more
|
|
# shared memory than docker's 64MB default.
|
|
shm_size: 512m
|
|
healthcheck:
|
|
test: ["CMD-SHELL", "pg_isready -U ${DB_USER:-curator} -d ${DB_NAME:-curator}"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 5
|
|
restart: unless-stopped
|
|
|
|
fabledcurator:
|
|
image: git.fabledsword.com/bvandeusen/fabledcurator:latest
|
|
# No `command:`. Everything — hypercorn plus one celery process per lane
|
|
# under supervisord — is what the image does by default, and supervisord's
|
|
# config is generated from the application's own lane table so the two
|
|
# cannot disagree. `command: ["all"]` still works and means the same thing.
|
|
# Sized to the SLOWEST lane, not the average. maintenance_long runs DB
|
|
# backups, library audits and translation backfill, and gets 180s to
|
|
# finish a chunk; the lanes stop in parallel, so this covers the max
|
|
# rather than their sum. Below this, a routine restart becomes a SIGKILL
|
|
# mid-backup — which is recoverable (the work is chunked and idempotent)
|
|
# but wastes however long it had run.
|
|
stop_grace_period: 200s
|
|
# No healthcheck here either. The image declares one that reads the role
|
|
# it is running, and for this one that means BOTH halves: hypercorn
|
|
# answers AND every lane is answering the broker. A web-only check would
|
|
# report a healthy container while every lane inside it had crashed —
|
|
# the failure mode consolidation creates, since docker can no longer see
|
|
# the lanes as separate services.
|
|
environment:
|
|
DB_USER: ${DB_USER:-curator}
|
|
DB_PASSWORD: ${DB_PASSWORD:-postgres}
|
|
DB_HOST: postgres
|
|
DB_PORT: "5432"
|
|
DB_NAME: ${DB_NAME:-curator}
|
|
CELERY_BROKER_URL: redis://redis:6379/0
|
|
CELERY_RESULT_BACKEND: redis://redis:6379/0
|
|
SECRET_KEY: ${SECRET_KEY:-change-me-before-you-expose-this}
|
|
EXTENSION_API_KEY: ${EXTENSION_API_KEY:-}
|
|
LOG_LEVEL: ${LOG_LEVEL:-INFO}
|
|
ports:
|
|
- "${PORT:-8080}:8080"
|
|
volumes:
|
|
- ${IMAGES_DIR:-./images}:/images
|
|
# Read-only. The filesystem scan copies out of here and never writes to
|
|
# it, so a mistake cannot reach the source library.
|
|
- ${IMPORT_DIR:-./import}:/import:ro
|
|
depends_on:
|
|
postgres: { condition: service_healthy }
|
|
redis: { condition: service_healthy }
|
|
restart: unless-stopped
|
|
|
|
volumes:
|
|
redis_data:
|
|
postgres_data:
|