CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 5s
Build images / build-agent (push) Successful in 7s
extension / lint (push) Successful in 19s
CI / frontend-build (push) Successful in 24s
CI / backend-lint-and-test (push) Successful in 32s
Build images / build-web (push) Successful in 1m41s
CI / integration (push) Successful in 2m7s
Build images / smoke-web (push) Failing after 12m36s
Build images / promote (push) Skipped
Operator, 2026-09-23: *"I also want to see that we remove the need for the command line of the configuration in the consolidated version."* `CMD` was `web`, so the single-container layout only worked if you knew to ask for it by name. A compose file that forgot `command: ["all"]` got a web server with nothing processing its queues — a gallery that loads, accepts an import, and never finishes one. Nothing errors; it just never progresses. Now `docker run fabledcurator` with no command starts hypercorn plus every lane under supervisord. `entrypoint.sh`'s own default moves with it, since the two are doors to the same decision and a disagreement would only show up as `--entrypoint` behaving differently from a plain run. `docker-compose.single.yml` drops its `command:` line; `["all"]` still works and still means the same thing. The multi-service stack is untouched — every service there names its role explicitly, which is what makes it the multi-service stack. ## And CI now actually boots it This is the gap I should have named when I reported milestone 422 at 7/7 and did not. Measured, not inferred: the smoke booted role `web` only (build.yml:1454), nothing in CI ran `all`, `docker-compose.single.yml` was read as TEXT by one test checking stop_grace_period and never run, and test_gen_supervisord asserts the generated config against the lane table without ever handing it to supervisord. So the shape this milestone is NAMED for had started nowhere. Steps 5-7 were marked done on evidence that did not cover it, and the operator is about to collapse their production stack onto exactly that. The smoke now boots the image with NO command — checking the Dockerfile CMD, the entrypoint default and the role together, the way an adopter gets it — and asserts `healthcheck_all`, which was itself never executed. That check passes only when hypercorn answers AND every lane in the table answers the broker; a web-only check goes green with every worker dead, which is the failure mode consolidation creates. It then prints `supervisorctl status`, so a lane that is merely restart-looping is visible rather than inferred. Cheap because ml ships at 0 slots and disabled: nothing loads a model, and the lane answers `inspect` with its consumers cancelled, which is what healthy means for a disabled lane. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
142 lines
5.9 KiB
Docker
142 lines
5.9 KiB
Docker
# syntax=docker/dockerfile:1.25
|
|
|
|
FROM node:24-alpine AS frontend-builder
|
|
WORKDIR /build
|
|
COPY frontend/package.json frontend/package-lock.json* ./
|
|
# No package-lock.json is tracked yet (we don't run npm locally per
|
|
# feedback-no-local-runs), so `npm install` instead of `npm ci`. Flip to
|
|
# `npm ci` once a lockfile is committed.
|
|
RUN npm install --no-audit --no-fund
|
|
COPY frontend/ ./
|
|
RUN npm run build
|
|
|
|
FROM python:3.14-slim AS runtime
|
|
ENV PYTHONUNBUFFERED=1 \
|
|
PYTHONDONTWRITEBYTECODE=1 \
|
|
PIP_NO_CACHE_DIR=1 \
|
|
PIP_DISABLE_PIP_VERSION_CHECK=1
|
|
|
|
# System deps: ffmpeg (transcode + thumbnails, FC-2), unar (archives, FC-2),
|
|
# libpq for psycopg, postgresql-client + zstd for FC-5 backup/restore
|
|
# (pg_dump + tar --zstd), image libs, megatools (mega.nz public-link downloads
|
|
# for off-platform file-host links, #830 — `megatools dl`; Debian-native, no
|
|
# external MEGA apt repo needed).
|
|
RUN apt-get update && apt-get install -y --no-install-recommends \
|
|
ffmpeg \
|
|
unar \
|
|
libpq5 \
|
|
postgresql-client \
|
|
zstd \
|
|
megatools \
|
|
libjpeg62-turbo \
|
|
libwebp7 \
|
|
libpng16-16 \
|
|
ca-certificates \
|
|
# opencv-python-headless (via requirements-ml.txt) links these even in its
|
|
# headless build. Came from Dockerfile.ml when the images merged
|
|
# (milestone 422 step 6).
|
|
libgl1 \
|
|
libglib2.0-0 \
|
|
&& rm -rf /var/lib/apt/lists/*
|
|
|
|
WORKDIR /app
|
|
|
|
COPY requirements.txt requirements-ml.txt ./
|
|
RUN pip install -r requirements.txt
|
|
|
|
# --- ML, merged from Dockerfile.ml (milestone 422 step 6) --------------------
|
|
#
|
|
# ONE image now serves every lane. It was two because the ML lane ran in its
|
|
# own container; with the single-container layout (step 5) running every lane
|
|
# in one process tree, a second image would mean the `ml` lane could never be
|
|
# enabled from the UI — there would be no worker in this container to enable.
|
|
#
|
|
# THE COST, MEASURED from run 7273 rather than guessed — and it is far
|
|
# smaller than the estimate this comment first carried, which said "everyone
|
|
# pulls ~4GB":
|
|
#
|
|
# torch 2.12.1+cpu wheel 192.3 MB
|
|
# torchvision 0.27.1+cpu 1.8 MB
|
|
# transformers / onnxruntime / opencv / sklearn and friends
|
|
# 62.0, 35.3, 23.6, 16.7, 12.3, 9.2, 6.9 MB
|
|
# largest newly-pushed layer 222.07 MB
|
|
#
|
|
# So the ML code adds a few hundred MB to the pull, not gigabytes. The CPU
|
|
# index is what makes that true: the default PyPI torch wheel bundles the
|
|
# NVIDIA CUDA runtime and is ~2GB on its own.
|
|
#
|
|
# The GIGABYTES are in the MODEL — ~3.5GB of SigLIP weights — and those are
|
|
# NOT in this image. They arrive only when the operator enables the lane,
|
|
# which is what lets rule 164 permit a runtime fetch at all ("optional and
|
|
# clearly off"). That also settles the trade this step was asked to weigh:
|
|
# baking the weights in would add ~3.5GB to every pull for a feature many
|
|
# adopters never enable, against ~350MB for the code that makes the switch
|
|
# available. Off-by-default wins by an order of magnitude, which was NOT
|
|
# obvious before measuring — the estimate had the two costs within 15% of
|
|
# each other.
|
|
#
|
|
# `--index-url`, not `--extra-index-url`: the latter would let pip resolve a
|
|
# +cu wheel anyway, and the whole saving above depends on it not doing that.
|
|
#
|
|
# CPU-only torch from the PyTorch CPU index. Nothing here uses a GPU — the
|
|
# GPU agent is a separate service with its own image.
|
|
RUN pip install --index-url https://download.pytorch.org/whl/cpu \
|
|
"torch>=2.12,<3.0" "torchvision>=0.27,<0.28"
|
|
RUN pip install -r requirements-ml.txt
|
|
|
|
# Where the model lands. Deliberately NOT a VOLUME instruction: that mints an
|
|
# anonymous volume when nobody mounts one, which survives `docker rm` and
|
|
# accumulates 3.5GB copies nobody can find. The compose files mount it
|
|
# explicitly instead, so an unmounted run simply re-downloads — visible, and
|
|
# recoverable.
|
|
ENV HF_HOME=/models/.huggingface \
|
|
TRANSFORMERS_CACHE=/models/.huggingface \
|
|
ML_MODEL_DIR=/models
|
|
|
|
COPY backend/ ./backend/
|
|
COPY alembic/ ./alembic/
|
|
COPY alembic.ini ./
|
|
COPY entrypoint.sh ./
|
|
RUN chmod +x entrypoint.sh
|
|
|
|
COPY --from=frontend-builder /build/dist ./frontend/dist
|
|
|
|
# Which channel this image belongs to — `dev` or `main` (milestone 271 step 7).
|
|
# build.yml passes it; /api/extension/manifest reports it beside the version so
|
|
# an operator can tell which channel an install came from without the channel
|
|
# ever touching the version string.
|
|
#
|
|
# Empty by default, deliberately: a locally-built image then reports NO channel
|
|
# rather than claiming to be one, and the manifest omits the field entirely —
|
|
# indistinguishable from an image built before the field existed, which is
|
|
# exactly the shape every reader already has to handle.
|
|
#
|
|
# Declared LAST on purpose. An ARG/ENV invalidates every layer below it, and
|
|
# these are the values that differ between builds of otherwise identical
|
|
# source — put them any earlier and the two channels could never share a
|
|
# cached pip install.
|
|
#
|
|
# FC_VERSION is what the instance reports about itself in the UI. Since
|
|
# milestone 318 stopped publishing version image tags, that self-report is
|
|
# the only answer to "which build is this?" — nothing else names it.
|
|
ARG FC_CHANNEL=""
|
|
ENV FC_CHANNEL=${FC_CHANNEL}
|
|
ARG FC_VERSION=""
|
|
ENV FC_VERSION=${FC_VERSION}
|
|
|
|
EXPOSE 8080
|
|
|
|
ENTRYPOINT ["./entrypoint.sh"]
|
|
# The DEFAULT is the whole application, not one lane of it.
|
|
#
|
|
# `docker run fabledcurator` with no command starts hypercorn plus every
|
|
# worker lane under supervisord — the shape an adopter wants and the shape the
|
|
# consolidated stack runs. It was `web`, which meant the single-container
|
|
# layout only worked if you knew to ask for it by name, and a compose file
|
|
# that forgot `command:` got a web server with nothing processing its queues:
|
|
# a gallery that loads, accepts an import, and never finishes one.
|
|
#
|
|
# The multi-service stack is unaffected — every service there names its role
|
|
# explicitly, which is exactly what makes it the multi-service stack.
|
|
CMD ["all"]
|