feat: one image for every lane, with the model fetch gated on enabling (4296)
CI / lint (push) Successful in 2s
CI / extension-version (push) Successful in 2s
Build images / sign-extension (push) Successful in 3s
Build images / build-agent (push) Successful in 6s
CI / frontend-build (push) Successful in 20s
extension / lint (push) Successful in 23s
CI / backend-lint-and-test (push) Failing after 31s
CI / integration (push) Successful in 2m16s
Build images / build-ml (push) Successful in 3m8s
Build images / build-web (push) Successful in 3m16s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / lint (push) Successful in 2s
CI / extension-version (push) Successful in 2s
Build images / sign-extension (push) Successful in 3s
Build images / build-agent (push) Successful in 6s
CI / frontend-build (push) Successful in 20s
extension / lint (push) Successful in 23s
CI / backend-lint-and-test (push) Failing after 31s
CI / integration (push) Successful in 2m16s
Build images / build-ml (push) Successful in 3m8s
Build images / build-web (push) Successful in 3m16s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
Milestone 422 step 6. Dockerfile.ml is gone; the main image carries torch,
torchvision, transformers, onnxruntime and opencv, and serves every lane.
WHY IT HAD TO MERGE: step 5 runs every lane in one process tree, so a second
image would mean the `ml` lane could never be enabled from the UI — there
would be no worker in that container to enable. The switch needs something to
switch.
THE MODEL NO LONGER DOWNLOADS AT BOOT. `entrypoint.sh`'s ml-worker role ran
download_models before celery started, so every boot of that role reached
HuggingFace for ~3.5GB — a startup dependency on a third party for a feature
the operator may never use. Rule 164 permits a runtime fetch only for
something "optional and clearly off", so the fetch is now a TASK, enqueued
the moment the lane is ENABLED.
Being a task is what makes it visible: it gets a TaskRun row, so the download
shows in Activity with a duration and a status, and a failure is something an
operator can see and retry rather than a container that quietly never became
useful. Idempotent, so re-enabling a provisioned lane costs one no-op.
Enqueued only when the lane actually came ON (`enabled is True`, not the
resolved value) so re-saving slots does not re-fetch, and only when the
consumer change landed — a task queued onto a queue nothing consumes would
sit pending with no explanation.
`fabledcurator-ml` KEEPS PUBLISHING, from the merged Dockerfile. The
operator's Swarm stack references that name and lives outside this repo;
dropping it would not break their deploy, it would freeze it silently at the
last publish — the exact failure class this milestone keeps finding. Retiring
the NAME is its own task, gated on that stack moving. Same two-phase shape
#406 used for pixiv.
THREE LIVE BREAKAGES from deleting the file, found by grepping for it rather
than assuming the build was the only consumer:
- `docker-compose.override.yml` built the ml service from it (contributor
path would have failed at `docker compose build`).
- `tests/test_artifact_paths.py` pins the ml path set.
- `scripts/artifacts.sh` ML_PATHS named it. A path set naming a deleted file
silently stops contributing to the derived revision — which the reuse check
and the version string both read. That is #3202's recorded shape.
The `--with-ml` flag is gone from the generator and the healthcheck rather
than left defaulting to true. One image carries every lane now, so a flag
that can only be passed one way is a branch pretending to be a choice.
The advisory shipped in ecbd325 is what makes this honest to an adopter: the
lane says it is optional, names the model, and gives its download and
per-slot RAM before the switch is thrown.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
This commit is contained in:
+39
-1
@@ -32,13 +32,51 @@ RUN apt-get update && apt-get install -y --no-install-recommends \
|
||||
libwebp7 \
|
||||
libpng16-16 \
|
||||
ca-certificates \
|
||||
# opencv-python-headless (via requirements-ml.txt) links these even in its
|
||||
# headless build. Came from Dockerfile.ml when the images merged
|
||||
# (milestone 422 step 6).
|
||||
libgl1 \
|
||||
libglib2.0-0 \
|
||||
&& rm -rf /var/lib/apt/lists/*
|
||||
|
||||
WORKDIR /app
|
||||
|
||||
COPY requirements.txt ./
|
||||
COPY requirements.txt requirements-ml.txt ./
|
||||
RUN pip install -r requirements.txt
|
||||
|
||||
# --- ML, merged from Dockerfile.ml (milestone 422 step 6) --------------------
|
||||
#
|
||||
# ONE image now serves every lane. It was two because the ML lane ran in its
|
||||
# own container; with the single-container layout (step 5) running every lane
|
||||
# in one process tree, a second image would mean the `ml` lane could never be
|
||||
# enabled from the UI — there would be no worker in this container to enable.
|
||||
#
|
||||
# The COST, stated because it is real and falls on every adopter: this adds
|
||||
# torch, torchvision, transformers, onnxruntime and opencv to an image that
|
||||
# previously carried none of them. Everyone pulls it, including the many who
|
||||
# will never turn tagging on. That is the trade the milestone accepted for
|
||||
# being able to offer the lane as a switch rather than a second deployment.
|
||||
# What it buys back is that nothing downloads a MODEL until the switch is
|
||||
# thrown — the weights are not baked in, and rule 164 permits that only
|
||||
# because the feature is optional and clearly off.
|
||||
#
|
||||
# CPU-only torch from the PyTorch CPU index. The default PyPI wheel bundles
|
||||
# the NVIDIA CUDA runtime (~5.6GB of layer) and nothing here uses a GPU — the
|
||||
# GPU agent is a separate service with its own image. `--index-url`, not
|
||||
# `--extra-index-url`: the latter would let pip resolve a +cu wheel anyway.
|
||||
RUN pip install --index-url https://download.pytorch.org/whl/cpu \
|
||||
"torch>=2.12,<3.0" "torchvision>=0.27,<0.28"
|
||||
RUN pip install -r requirements-ml.txt
|
||||
|
||||
# Where the model lands. Deliberately NOT a VOLUME instruction: that mints an
|
||||
# anonymous volume when nobody mounts one, which survives `docker rm` and
|
||||
# accumulates 3.5GB copies nobody can find. The compose files mount it
|
||||
# explicitly instead, so an unmounted run simply re-downloads — visible, and
|
||||
# recoverable.
|
||||
ENV HF_HOME=/models/.huggingface \
|
||||
TRANSFORMERS_CACHE=/models/.huggingface \
|
||||
ML_MODEL_DIR=/models
|
||||
|
||||
COPY backend/ ./backend/
|
||||
COPY alembic/ ./alembic/
|
||||
COPY alembic.ini ./
|
||||
|
||||
Reference in New Issue
Block a user