Files
FabledCurator/tests/test_gen_supervisord.py
T
bvandeusenandClaude Opus 5 ffcd13096a
CI / lint (push) Successful in 2s
CI / extension-version (push) Successful in 2s
Build images / sign-extension (push) Successful in 3s
Build images / build-agent (push) Successful in 6s
CI / frontend-build (push) Successful in 20s
extension / lint (push) Successful in 23s
CI / backend-lint-and-test (push) Failing after 31s
CI / integration (push) Successful in 2m16s
Build images / build-ml (push) Successful in 3m8s
Build images / build-web (push) Successful in 3m16s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
feat: one image for every lane, with the model fetch gated on enabling (4296)
Milestone 422 step 6. Dockerfile.ml is gone; the main image carries torch,
torchvision, transformers, onnxruntime and opencv, and serves every lane.

WHY IT HAD TO MERGE: step 5 runs every lane in one process tree, so a second
image would mean the `ml` lane could never be enabled from the UI — there
would be no worker in that container to enable. The switch needs something to
switch.

THE MODEL NO LONGER DOWNLOADS AT BOOT. `entrypoint.sh`'s ml-worker role ran
download_models before celery started, so every boot of that role reached
HuggingFace for ~3.5GB — a startup dependency on a third party for a feature
the operator may never use. Rule 164 permits a runtime fetch only for
something "optional and clearly off", so the fetch is now a TASK, enqueued
the moment the lane is ENABLED.

Being a task is what makes it visible: it gets a TaskRun row, so the download
shows in Activity with a duration and a status, and a failure is something an
operator can see and retry rather than a container that quietly never became
useful. Idempotent, so re-enabling a provisioned lane costs one no-op.

Enqueued only when the lane actually came ON (`enabled is True`, not the
resolved value) so re-saving slots does not re-fetch, and only when the
consumer change landed — a task queued onto a queue nothing consumes would
sit pending with no explanation.

`fabledcurator-ml` KEEPS PUBLISHING, from the merged Dockerfile. The
operator's Swarm stack references that name and lives outside this repo;
dropping it would not break their deploy, it would freeze it silently at the
last publish — the exact failure class this milestone keeps finding. Retiring
the NAME is its own task, gated on that stack moving. Same two-phase shape
#406 used for pixiv.

THREE LIVE BREAKAGES from deleting the file, found by grepping for it rather
than assuming the build was the only consumer:

- `docker-compose.override.yml` built the ml service from it (contributor
  path would have failed at `docker compose build`).
- `tests/test_artifact_paths.py` pins the ml path set.
- `scripts/artifacts.sh` ML_PATHS named it. A path set naming a deleted file
  silently stops contributing to the derived revision — which the reuse check
  and the version string both read. That is #3202's recorded shape.

The `--with-ml` flag is gone from the generator and the healthcheck rather
than left defaulting to true. One image carries every lane now, so a flag
that can only be passed one way is a branch pretending to be a choice.

The advisory shipped in ecbd325 is what makes this honest to an adopter: the
lane says it is optional, names the model, and gives its download and
per-slot RAM before the switch is thrown.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-22 08:51:44 -04:00

187 lines
7.8 KiB
Python

"""The generated supervisord config (milestone 422 step 5).
Asserts the config against the LANE TABLE rather than against a fixture of
expected text. A fixture would have to be updated whenever a lane changes,
which is the same hand-kept coupling generating the config exists to remove —
and it would pass while describing a container that does not match the
application's own idea of what it runs.
"""
from __future__ import annotations
import configparser
from backend.app.scripts import gen_supervisord as gen
from backend.app.services.worker_lanes import LANES, LANES_BY_NAME
# --- structure ---------------------------------------------------------------
def _parse(**kwargs) -> configparser.ConfigParser:
"""supervisord's config is ini, so parse it rather than grepping strings.
A substring assertion passes on a line that is present but malformed —
inside a comment, in the wrong section, or with a typo'd key that
supervisord silently ignores.
"""
cp = configparser.ConfigParser()
cp.read_string(gen.render(**kwargs))
return cp
def test_it_is_valid_ini_with_a_supervisord_section():
cp = _parse()
assert cp.has_section("supervisord")
# PID 1 in a container: daemonising would exit immediately and take the
# container with it.
assert cp.get("supervisord", "nodaemon") == "true"
def test_every_lane_gets_a_program():
"""One image carries every lane since step 6, so nothing is conditional.
A lane in LANES with no program is a queue with no consumer."""
cp = _parse()
expected = {"program:web"} | {f"program:{lane.name}" for lane in LANES}
assert set(cp.sections()) - {"supervisord"} == expected
def test_the_ml_lane_runs_even_though_it_ships_disabled():
"""It holds a PROCESS and no model. `add_consumer` needs a running worker
to reach, so without this the UI switch would have nothing to switch —
and nothing is downloaded by starting it, which is what lets rule 164
permit the fetch at all."""
assert _parse().has_section("program:ml")
assert LANES_BY_NAME["ml"].default_enabled is False
# --- the coupling this generator exists to guarantee -------------------------
def test_each_program_serves_exactly_its_lane_s_queues():
"""The whole point: the container's processes and the application's lane
table are one list. A queue in LANES with no program means work that
queues forever with nothing consuming it."""
cp = _parse()
for lane in LANES:
env = cp.get(f"program:{lane.name}", "environment")
# The QUOTED form. supervisord splits `environment` on commas, so an
# unquoted multi-queue value silently degrades to its first queue —
# and an assertion on the bare string passes either way, which is how
# that would have shipped.
assert f'CELERY_QUEUES="{",".join(lane.queues)}"' in env
def test_each_program_invokes_the_lane_s_entrypoint_role_not_its_name():
"""`maintenance_long` is the plain `worker` role pointed at a different
queue — exactly as docker-compose starts it today. Invoking
`entrypoint.sh maintenance_long` would hit the unknown-role branch and
exit 1 on every restart."""
cp = _parse()
for lane in LANES:
command = cp.get(f"program:{lane.name}", "command")
assert f"entrypoint.sh {lane.entrypoint_role}" in command
def test_every_entrypoint_role_a_lane_names_actually_exists():
"""Reads entrypoint.sh itself. The generator can only emit a role name;
whether the script handles it is a separate fact, and getting it wrong
fails at container start rather than here."""
from pathlib import Path
script = Path(__file__).resolve().parents[1] / "entrypoint.sh"
text = script.read_text()
for lane in LANES:
# Roles are `case` arms: ` worker)` possibly in an alternation.
assert f" {lane.entrypoint_role})" in text or \
f"|{lane.entrypoint_role})" in text, \
f"{lane.name} names entrypoint role {lane.entrypoint_role!r}, which does not exist"
# --- shutdown ----------------------------------------------------------------
def test_every_program_signals_its_whole_process_group():
"""Celery's prefork pool forks children. A TERM delivered only to the
parent leaves them running and holding tasks — a 'graceful' shutdown that
orphans workers. The `sh -c … | sed` wrapper makes this doubly necessary:
without it the signal reaches the shell holding the pipeline, not celery."""
cp = _parse()
for section in cp.sections():
if not section.startswith("program:"):
continue
assert cp.get(section, "stopasgroup") == "true", section
assert cp.get(section, "killasgroup") == "true", section
def test_the_long_maintenance_lane_keeps_its_180s_drain():
"""The per-service stop_grace_period values from the multi-service stack
are preserved per program. maintenance_long runs DB backups and library
audits; cutting its drain turns a restart into a SIGKILL mid-backup."""
cp = _parse()
assert cp.getint("program:maintenance_long", "stopwaitsecs") == 180
def test_no_program_waits_longer_than_the_compose_stop_grace_period():
"""The container gets ONE timeout and the programs stop in parallel, so it
must cover the slowest. If a lane's stopwaitsecs ever exceeds what
docker-compose.single.yml allows, docker kills the container while that
lane still believes it has time to drain."""
import re
from pathlib import Path
compose = (Path(__file__).resolve().parents[1] / "docker-compose.single.yml").read_text()
m = re.search(r"stop_grace_period:\s*(\d+)s", compose)
assert m, "docker-compose.single.yml has no stop_grace_period"
grace = int(m.group(1))
cp = _parse()
for section in cp.sections():
if section.startswith("program:"):
assert cp.getint(section, "stopwaitsecs") <= grace, section
# --- what runs, and how much ------------------------------------------------
def test_a_zero_slot_lane_still_gets_a_running_process():
"""ML ships at 0 slots and disabled — but `add_consumer` needs something to
reach. With no process there would be nothing for the UI switch to switch,
and enabling tagging could not work at all."""
cp = _parse()
assert LANES_BY_NAME["ml"].default_slots == 0
env = cp.get("program:ml", "environment")
assert "CELERY_CONCURRENCY=1" in env
def test_programs_restart_but_back_off_rather_than_looping():
"""A lane that dies instantly and repeatedly is a broken image, not a
transient fault. Unbounded restarts would burn a core forever and bury the
original error under its own noise."""
cp = _parse()
for section in cp.sections():
if section.startswith("program:"):
assert cp.get(section, "autorestart") == "true", section
assert cp.getint(section, "startretries") >= 1, section
def test_web_starts_first_because_it_runs_the_migration():
"""A worker booting against an un-migrated schema fails in a way that
looks like application breakage rather than an ordering problem."""
cp = _parse()
assert cp.getint("program:web", "priority") == 1
def test_every_program_writes_to_the_container_stdout_with_its_lane_named():
"""Four celery workers and hypercorn on one stream are indistinguishable
without this. Unbuffered (`maxbytes 0`) so `docker logs` is live rather
than arriving in rotated chunks."""
cp = _parse()
for section in cp.sections():
if not section.startswith("program:"):
continue
name = section.split(":", 1)[1]
assert cp.get(section, "stdout_logfile") == "/dev/fd/1", section
assert cp.getint(section, "stdout_logfile_maxbytes") == 0, section
assert cp.get(section, "redirect_stderr") == "true", section
assert f"[{name}] " in cp.get(section, "command"), section