feat: one image for every lane, with the model fetch gated on enabling (4296)
CI / lint (push) Successful in 2s
CI / extension-version (push) Successful in 2s
Build images / sign-extension (push) Successful in 3s
Build images / build-agent (push) Successful in 6s
CI / frontend-build (push) Successful in 20s
extension / lint (push) Successful in 23s
CI / backend-lint-and-test (push) Failing after 31s
CI / integration (push) Successful in 2m16s
Build images / build-ml (push) Successful in 3m8s
Build images / build-web (push) Successful in 3m16s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped

Milestone 422 step 6. Dockerfile.ml is gone; the main image carries torch,
torchvision, transformers, onnxruntime and opencv, and serves every lane.

WHY IT HAD TO MERGE: step 5 runs every lane in one process tree, so a second
image would mean the `ml` lane could never be enabled from the UI — there
would be no worker in that container to enable. The switch needs something to
switch.

THE MODEL NO LONGER DOWNLOADS AT BOOT. `entrypoint.sh`'s ml-worker role ran
download_models before celery started, so every boot of that role reached
HuggingFace for ~3.5GB — a startup dependency on a third party for a feature
the operator may never use. Rule 164 permits a runtime fetch only for
something "optional and clearly off", so the fetch is now a TASK, enqueued
the moment the lane is ENABLED.

Being a task is what makes it visible: it gets a TaskRun row, so the download
shows in Activity with a duration and a status, and a failure is something an
operator can see and retry rather than a container that quietly never became
useful. Idempotent, so re-enabling a provisioned lane costs one no-op.

Enqueued only when the lane actually came ON (`enabled is True`, not the
resolved value) so re-saving slots does not re-fetch, and only when the
consumer change landed — a task queued onto a queue nothing consumes would
sit pending with no explanation.

`fabledcurator-ml` KEEPS PUBLISHING, from the merged Dockerfile. The
operator's Swarm stack references that name and lives outside this repo;
dropping it would not break their deploy, it would freeze it silently at the
last publish — the exact failure class this milestone keeps finding. Retiring
the NAME is its own task, gated on that stack moving. Same two-phase shape
#406 used for pixiv.

THREE LIVE BREAKAGES from deleting the file, found by grepping for it rather
than assuming the build was the only consumer:

- `docker-compose.override.yml` built the ml service from it (contributor
  path would have failed at `docker compose build`).
- `tests/test_artifact_paths.py` pins the ml path set.
- `scripts/artifacts.sh` ML_PATHS named it. A path set naming a deleted file
  silently stops contributing to the derived revision — which the reuse check
  and the version string both read. That is #3202's recorded shape.

The `--with-ml` flag is gone from the generator and the healthcheck rather
than left defaulting to true. One image carries every lane now, so a flag
that can only be passed one way is a branch pretending to be a choice.

The advisory shipped in ecbd325 is what makes this honest to an adopter: the
lane says it is optional, names the model, and gives its download and
per-slot RAM before the switch is thrown.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
This commit is contained in:
2026-09-22 08:51:44 -04:00
co-authored by Claude Opus 5
parent ecbd325437
commit ffcd13096a
15 changed files with 183 additions and 102 deletions
+1 -1
View File
@@ -31,7 +31,7 @@ ROOT = Path(__file__).resolve().parent.parent
# artifact -> (dockerfile, build context relative to the repo root)
ARTIFACTS = {
"web": ("Dockerfile", ""),
"ml": ("Dockerfile.ml", ""),
"ml": ("Dockerfile", ""),
"agent": ("agent/Dockerfile", "agent"),
}
+17 -16
View File
@@ -37,20 +37,21 @@ def test_it_is_valid_ini_with_a_supervisord_section():
assert cp.get("supervisord", "nodaemon") == "true"
def test_web_and_every_non_ml_lane_get_a_program():
def test_every_lane_gets_a_program():
"""One image carries every lane since step 6, so nothing is conditional.
A lane in LANES with no program is a queue with no consumer."""
cp = _parse()
expected = {"program:web"} | {
f"program:{lane.name}" for lane in LANES if lane.name != "ml"
}
expected = {"program:web"} | {f"program:{lane.name}" for lane in LANES}
assert set(cp.sections()) - {"supervisord"} == expected
def test_the_ml_lane_is_absent_until_its_deps_are_in_the_image():
"""The web image has no torch. An `ml` program here would fail to import
on every restart, forever — startretries would give up and the lane would
be permanently dead while the container reported healthy."""
assert not _parse().has_section("program:ml")
assert _parse(with_ml=True).has_section("program:ml")
def test_the_ml_lane_runs_even_though_it_ships_disabled():
"""It holds a PROCESS and no model. `add_consumer` needs a running worker
to reach, so without this the UI switch would have nothing to switch —
and nothing is downloaded by starting it, which is what lets rule 164
permit the fetch at all."""
assert _parse().has_section("program:ml")
assert LANES_BY_NAME["ml"].default_enabled is False
# --- the coupling this generator exists to guarantee -------------------------
@@ -60,7 +61,7 @@ def test_each_program_serves_exactly_its_lane_s_queues():
"""The whole point: the container's processes and the application's lane
table are one list. A queue in LANES with no program means work that
queues forever with nothing consuming it."""
cp = _parse(with_ml=True)
cp = _parse()
for lane in LANES:
env = cp.get(f"program:{lane.name}", "environment")
# The QUOTED form. supervisord splits `environment` on commas, so an
@@ -75,7 +76,7 @@ def test_each_program_invokes_the_lane_s_entrypoint_role_not_its_name():
queue — exactly as docker-compose starts it today. Invoking
`entrypoint.sh maintenance_long` would hit the unknown-role branch and
exit 1 on every restart."""
cp = _parse(with_ml=True)
cp = _parse()
for lane in LANES:
command = cp.get(f"program:{lane.name}", "command")
assert f"entrypoint.sh {lane.entrypoint_role}" in command
@@ -104,7 +105,7 @@ def test_every_program_signals_its_whole_process_group():
parent leaves them running and holding tasks — a 'graceful' shutdown that
orphans workers. The `sh -c … | sed` wrapper makes this doubly necessary:
without it the signal reaches the shell holding the pipeline, not celery."""
cp = _parse(with_ml=True)
cp = _parse()
for section in cp.sections():
if not section.startswith("program:"):
continue
@@ -133,7 +134,7 @@ def test_no_program_waits_longer_than_the_compose_stop_grace_period():
assert m, "docker-compose.single.yml has no stop_grace_period"
grace = int(m.group(1))
cp = _parse(with_ml=True)
cp = _parse()
for section in cp.sections():
if section.startswith("program:"):
assert cp.getint(section, "stopwaitsecs") <= grace, section
@@ -146,7 +147,7 @@ def test_a_zero_slot_lane_still_gets_a_running_process():
"""ML ships at 0 slots and disabled — but `add_consumer` needs something to
reach. With no process there would be nothing for the UI switch to switch,
and enabling tagging could not work at all."""
cp = _parse(with_ml=True)
cp = _parse()
assert LANES_BY_NAME["ml"].default_slots == 0
env = cp.get("program:ml", "environment")
assert "CELERY_CONCURRENCY=1" in env
@@ -174,7 +175,7 @@ def test_every_program_writes_to_the_container_stdout_with_its_lane_named():
"""Four celery workers and hypercorn on one stream are indistinguishable
without this. Unbuffered (`maxbytes 0`) so `docker logs` is live rather
than arriving in rotated chunks."""
cp = _parse(with_ml=True)
cp = _parse()
for section in cp.sections():
if not section.startswith("program:"):
continue