CI / lint (push) Successful in 2s
CI / extension-version (push) Successful in 2s
Build images / sign-extension (push) Successful in 3s
Build images / build-agent (push) Successful in 6s
CI / frontend-build (push) Successful in 20s
extension / lint (push) Successful in 23s
CI / backend-lint-and-test (push) Failing after 31s
CI / integration (push) Successful in 2m16s
Build images / build-ml (push) Successful in 3m8s
Build images / build-web (push) Successful in 3m16s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
Milestone 422 step 6. Dockerfile.ml is gone; the main image carries torch,
torchvision, transformers, onnxruntime and opencv, and serves every lane.
WHY IT HAD TO MERGE: step 5 runs every lane in one process tree, so a second
image would mean the `ml` lane could never be enabled from the UI — there
would be no worker in that container to enable. The switch needs something to
switch.
THE MODEL NO LONGER DOWNLOADS AT BOOT. `entrypoint.sh`'s ml-worker role ran
download_models before celery started, so every boot of that role reached
HuggingFace for ~3.5GB — a startup dependency on a third party for a feature
the operator may never use. Rule 164 permits a runtime fetch only for
something "optional and clearly off", so the fetch is now a TASK, enqueued
the moment the lane is ENABLED.
Being a task is what makes it visible: it gets a TaskRun row, so the download
shows in Activity with a duration and a status, and a failure is something an
operator can see and retry rather than a container that quietly never became
useful. Idempotent, so re-enabling a provisioned lane costs one no-op.
Enqueued only when the lane actually came ON (`enabled is True`, not the
resolved value) so re-saving slots does not re-fetch, and only when the
consumer change landed — a task queued onto a queue nothing consumes would
sit pending with no explanation.
`fabledcurator-ml` KEEPS PUBLISHING, from the merged Dockerfile. The
operator's Swarm stack references that name and lives outside this repo;
dropping it would not break their deploy, it would freeze it silently at the
last publish — the exact failure class this milestone keeps finding. Retiring
the NAME is its own task, gated on that stack moving. Same two-phase shape
#406 used for pixiv.
THREE LIVE BREAKAGES from deleting the file, found by grepping for it rather
than assuming the build was the only consumer:
- `docker-compose.override.yml` built the ml service from it (contributor
path would have failed at `docker compose build`).
- `tests/test_artifact_paths.py` pins the ml path set.
- `scripts/artifacts.sh` ML_PATHS named it. A path set naming a deleted file
silently stops contributing to the derived revision — which the reuse check
and the version string both read. That is #3202's recorded shape.
The `--with-ml` flag is gone from the generator and the healthcheck rather
than left defaulting to true. One image carries every lane now, so a flag
that can only be passed one way is a branch pretending to be a choice.
The advisory shipped in ecbd325 is what makes this honest to an adopter: the
lane says it is optional, names the model, and gives its download and
per-slot RAM before the switch is thrown.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
187 lines
7.8 KiB
Python
187 lines
7.8 KiB
Python
"""The generated supervisord config (milestone 422 step 5).
|
|
|
|
Asserts the config against the LANE TABLE rather than against a fixture of
|
|
expected text. A fixture would have to be updated whenever a lane changes,
|
|
which is the same hand-kept coupling generating the config exists to remove —
|
|
and it would pass while describing a container that does not match the
|
|
application's own idea of what it runs.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import configparser
|
|
|
|
from backend.app.scripts import gen_supervisord as gen
|
|
from backend.app.services.worker_lanes import LANES, LANES_BY_NAME
|
|
|
|
# --- structure ---------------------------------------------------------------
|
|
|
|
|
|
def _parse(**kwargs) -> configparser.ConfigParser:
|
|
"""supervisord's config is ini, so parse it rather than grepping strings.
|
|
|
|
A substring assertion passes on a line that is present but malformed —
|
|
inside a comment, in the wrong section, or with a typo'd key that
|
|
supervisord silently ignores.
|
|
"""
|
|
cp = configparser.ConfigParser()
|
|
cp.read_string(gen.render(**kwargs))
|
|
return cp
|
|
|
|
|
|
def test_it_is_valid_ini_with_a_supervisord_section():
|
|
cp = _parse()
|
|
assert cp.has_section("supervisord")
|
|
# PID 1 in a container: daemonising would exit immediately and take the
|
|
# container with it.
|
|
assert cp.get("supervisord", "nodaemon") == "true"
|
|
|
|
|
|
def test_every_lane_gets_a_program():
|
|
"""One image carries every lane since step 6, so nothing is conditional.
|
|
A lane in LANES with no program is a queue with no consumer."""
|
|
cp = _parse()
|
|
expected = {"program:web"} | {f"program:{lane.name}" for lane in LANES}
|
|
assert set(cp.sections()) - {"supervisord"} == expected
|
|
|
|
|
|
def test_the_ml_lane_runs_even_though_it_ships_disabled():
|
|
"""It holds a PROCESS and no model. `add_consumer` needs a running worker
|
|
to reach, so without this the UI switch would have nothing to switch —
|
|
and nothing is downloaded by starting it, which is what lets rule 164
|
|
permit the fetch at all."""
|
|
assert _parse().has_section("program:ml")
|
|
assert LANES_BY_NAME["ml"].default_enabled is False
|
|
|
|
|
|
# --- the coupling this generator exists to guarantee -------------------------
|
|
|
|
|
|
def test_each_program_serves_exactly_its_lane_s_queues():
|
|
"""The whole point: the container's processes and the application's lane
|
|
table are one list. A queue in LANES with no program means work that
|
|
queues forever with nothing consuming it."""
|
|
cp = _parse()
|
|
for lane in LANES:
|
|
env = cp.get(f"program:{lane.name}", "environment")
|
|
# The QUOTED form. supervisord splits `environment` on commas, so an
|
|
# unquoted multi-queue value silently degrades to its first queue —
|
|
# and an assertion on the bare string passes either way, which is how
|
|
# that would have shipped.
|
|
assert f'CELERY_QUEUES="{",".join(lane.queues)}"' in env
|
|
|
|
|
|
def test_each_program_invokes_the_lane_s_entrypoint_role_not_its_name():
|
|
"""`maintenance_long` is the plain `worker` role pointed at a different
|
|
queue — exactly as docker-compose starts it today. Invoking
|
|
`entrypoint.sh maintenance_long` would hit the unknown-role branch and
|
|
exit 1 on every restart."""
|
|
cp = _parse()
|
|
for lane in LANES:
|
|
command = cp.get(f"program:{lane.name}", "command")
|
|
assert f"entrypoint.sh {lane.entrypoint_role}" in command
|
|
|
|
|
|
def test_every_entrypoint_role_a_lane_names_actually_exists():
|
|
"""Reads entrypoint.sh itself. The generator can only emit a role name;
|
|
whether the script handles it is a separate fact, and getting it wrong
|
|
fails at container start rather than here."""
|
|
from pathlib import Path
|
|
|
|
script = Path(__file__).resolve().parents[1] / "entrypoint.sh"
|
|
text = script.read_text()
|
|
for lane in LANES:
|
|
# Roles are `case` arms: ` worker)` possibly in an alternation.
|
|
assert f" {lane.entrypoint_role})" in text or \
|
|
f"|{lane.entrypoint_role})" in text, \
|
|
f"{lane.name} names entrypoint role {lane.entrypoint_role!r}, which does not exist"
|
|
|
|
|
|
# --- shutdown ----------------------------------------------------------------
|
|
|
|
|
|
def test_every_program_signals_its_whole_process_group():
|
|
"""Celery's prefork pool forks children. A TERM delivered only to the
|
|
parent leaves them running and holding tasks — a 'graceful' shutdown that
|
|
orphans workers. The `sh -c … | sed` wrapper makes this doubly necessary:
|
|
without it the signal reaches the shell holding the pipeline, not celery."""
|
|
cp = _parse()
|
|
for section in cp.sections():
|
|
if not section.startswith("program:"):
|
|
continue
|
|
assert cp.get(section, "stopasgroup") == "true", section
|
|
assert cp.get(section, "killasgroup") == "true", section
|
|
|
|
|
|
def test_the_long_maintenance_lane_keeps_its_180s_drain():
|
|
"""The per-service stop_grace_period values from the multi-service stack
|
|
are preserved per program. maintenance_long runs DB backups and library
|
|
audits; cutting its drain turns a restart into a SIGKILL mid-backup."""
|
|
cp = _parse()
|
|
assert cp.getint("program:maintenance_long", "stopwaitsecs") == 180
|
|
|
|
|
|
def test_no_program_waits_longer_than_the_compose_stop_grace_period():
|
|
"""The container gets ONE timeout and the programs stop in parallel, so it
|
|
must cover the slowest. If a lane's stopwaitsecs ever exceeds what
|
|
docker-compose.single.yml allows, docker kills the container while that
|
|
lane still believes it has time to drain."""
|
|
import re
|
|
from pathlib import Path
|
|
|
|
compose = (Path(__file__).resolve().parents[1] / "docker-compose.single.yml").read_text()
|
|
m = re.search(r"stop_grace_period:\s*(\d+)s", compose)
|
|
assert m, "docker-compose.single.yml has no stop_grace_period"
|
|
grace = int(m.group(1))
|
|
|
|
cp = _parse()
|
|
for section in cp.sections():
|
|
if section.startswith("program:"):
|
|
assert cp.getint(section, "stopwaitsecs") <= grace, section
|
|
|
|
|
|
# --- what runs, and how much ------------------------------------------------
|
|
|
|
|
|
def test_a_zero_slot_lane_still_gets_a_running_process():
|
|
"""ML ships at 0 slots and disabled — but `add_consumer` needs something to
|
|
reach. With no process there would be nothing for the UI switch to switch,
|
|
and enabling tagging could not work at all."""
|
|
cp = _parse()
|
|
assert LANES_BY_NAME["ml"].default_slots == 0
|
|
env = cp.get("program:ml", "environment")
|
|
assert "CELERY_CONCURRENCY=1" in env
|
|
|
|
|
|
def test_programs_restart_but_back_off_rather_than_looping():
|
|
"""A lane that dies instantly and repeatedly is a broken image, not a
|
|
transient fault. Unbounded restarts would burn a core forever and bury the
|
|
original error under its own noise."""
|
|
cp = _parse()
|
|
for section in cp.sections():
|
|
if section.startswith("program:"):
|
|
assert cp.get(section, "autorestart") == "true", section
|
|
assert cp.getint(section, "startretries") >= 1, section
|
|
|
|
|
|
def test_web_starts_first_because_it_runs_the_migration():
|
|
"""A worker booting against an un-migrated schema fails in a way that
|
|
looks like application breakage rather than an ordering problem."""
|
|
cp = _parse()
|
|
assert cp.getint("program:web", "priority") == 1
|
|
|
|
|
|
def test_every_program_writes_to_the_container_stdout_with_its_lane_named():
|
|
"""Four celery workers and hypercorn on one stream are indistinguishable
|
|
without this. Unbuffered (`maxbytes 0`) so `docker logs` is live rather
|
|
than arriving in rotated chunks."""
|
|
cp = _parse()
|
|
for section in cp.sections():
|
|
if not section.startswith("program:"):
|
|
continue
|
|
name = section.split(":", 1)[1]
|
|
assert cp.get(section, "stdout_logfile") == "/dev/fd/1", section
|
|
assert cp.getint(section, "stdout_logfile_maxbytes") == 0, section
|
|
assert cp.get(section, "redirect_stderr") == "true", section
|
|
assert f"[{name}] " in cp.get(section, "command"), section
|