CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 36s
Build images / build-web (push) Successful in 1m9s
Build images / smoke-web (push) Skipped
Build images / build-ml (push) Successful in 2m8s
Build images / promote (push) Skipped
CI / integration (push) Successful in 2m37s
Milestone 422 step 5. `docker compose -f docker-compose.single.yml up -d` gives three containers — FabledCurator, Postgres, Redis — where the stack previously needed seven. THE MULTI-SERVICE STACK IS KEPT. docker-compose.yml still runs the five app services separately and remains the right shape for a Swarm deployment spread across hosts, where per-service rolling rollback and placement constraints matter. This adds a compose file; it deletes none. `entrypoint.sh all` GENERATES the supervisord config from worker_lanes.LANES and execs it as PID 1. Generated rather than checked in because a static .conf would spell out each lane's -Q list, making a FIFTH hand-kept copy of the queue names — after celery_app.task_routes and the three collapsed in steps 1, 2 and 4. Every one of those had already drifted when found. Generating gives a stronger guarantee than "they match today": a lane added to LANES gets a process, and a queue cannot end up with no consumer because someone missed a file. supervisord over s6-overlay: one pip dependency on an image already Python, with per-program stop timeouts and stopasgroup. The process-group part is not a detail — celery's prefork pool forks children, and a TERM reaching only the parent leaves them orphaned holding tasks. s6's advantage (PID-1 signal and zombie handling) comes from `init: true` instead. Nothing in FC talks to the supervisor, so the choice is reversible without touching product code. FOUR LANES, NOT FIVE. The ml lane is skipped: torch and the ML requirements live only in Dockerfile.ml until step 6 merges the images, so an `ml` program here would fail to import on every restart forever. `--with-ml` is the flag step 6 turns on. THREE BUGS FOUND BY READING IT BACK, none of which the first tests caught: 1. `environment=CELERY_QUEUES=default,import,thumbnail,download` — supervisord parses that key as a COMMA-separated list, so it reads as CELERY_QUEUES=default plus three malformed entries and the worker lane would have consumed only `default`. Silent: the worker starts, reports healthy, never picks up an import. Now quoted, and the test asserts the quoted form rather than the bare substring, which passed either way. 2. The generator emitted `entrypoint.sh <lane.name>`, but `maintenance_long` is not a role — compose runs it as the plain `worker` role with different queues. Lane now carries `entrypoint_role`, and a test reads entrypoint.sh to assert every role a lane names actually exists. 3. The `scheduler` role hardcoded --concurrency=1, ignoring CELERY_CONCURRENCY. Harmless while only compose started it and set none; with a generated value being passed, the lane would have sat at 1 until the reconcile noticed, with nothing saying why. The healthcheck asserts BOTH halves — hypercorn answers and every configured lane is answering the broker. That is the failure mode consolidation creates: docker can no longer see the lanes as separate services, so a web-only check would report a healthy container with every lane inside it dead. It deliberately ignores the `enabled` flag: a disabled lane still has a running process with its consumers cancelled, and marking the container unhealthy for turning tagging off would be wrong. stop_grace_period 200s, sized to the slowest lane (maintenance_long at 180s) rather than the average, with a test asserting no program's stopwaitsecs can exceed what compose allows. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
186 lines
7.7 KiB
Python
186 lines
7.7 KiB
Python
"""The generated supervisord config (milestone 422 step 5).
|
|
|
|
Asserts the config against the LANE TABLE rather than against a fixture of
|
|
expected text. A fixture would have to be updated whenever a lane changes,
|
|
which is the same hand-kept coupling generating the config exists to remove —
|
|
and it would pass while describing a container that does not match the
|
|
application's own idea of what it runs.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import configparser
|
|
|
|
from backend.app.scripts import gen_supervisord as gen
|
|
from backend.app.services.worker_lanes import LANES, LANES_BY_NAME
|
|
|
|
# --- structure ---------------------------------------------------------------
|
|
|
|
|
|
def _parse(**kwargs) -> configparser.ConfigParser:
|
|
"""supervisord's config is ini, so parse it rather than grepping strings.
|
|
|
|
A substring assertion passes on a line that is present but malformed —
|
|
inside a comment, in the wrong section, or with a typo'd key that
|
|
supervisord silently ignores.
|
|
"""
|
|
cp = configparser.ConfigParser()
|
|
cp.read_string(gen.render(**kwargs))
|
|
return cp
|
|
|
|
|
|
def test_it_is_valid_ini_with_a_supervisord_section():
|
|
cp = _parse()
|
|
assert cp.has_section("supervisord")
|
|
# PID 1 in a container: daemonising would exit immediately and take the
|
|
# container with it.
|
|
assert cp.get("supervisord", "nodaemon") == "true"
|
|
|
|
|
|
def test_web_and_every_non_ml_lane_get_a_program():
|
|
cp = _parse()
|
|
expected = {"program:web"} | {
|
|
f"program:{lane.name}" for lane in LANES if lane.name != "ml"
|
|
}
|
|
assert set(cp.sections()) - {"supervisord"} == expected
|
|
|
|
|
|
def test_the_ml_lane_is_absent_until_its_deps_are_in_the_image():
|
|
"""The web image has no torch. An `ml` program here would fail to import
|
|
on every restart, forever — startretries would give up and the lane would
|
|
be permanently dead while the container reported healthy."""
|
|
assert not _parse().has_section("program:ml")
|
|
assert _parse(with_ml=True).has_section("program:ml")
|
|
|
|
|
|
# --- the coupling this generator exists to guarantee -------------------------
|
|
|
|
|
|
def test_each_program_serves_exactly_its_lane_s_queues():
|
|
"""The whole point: the container's processes and the application's lane
|
|
table are one list. A queue in LANES with no program means work that
|
|
queues forever with nothing consuming it."""
|
|
cp = _parse(with_ml=True)
|
|
for lane in LANES:
|
|
env = cp.get(f"program:{lane.name}", "environment")
|
|
# The QUOTED form. supervisord splits `environment` on commas, so an
|
|
# unquoted multi-queue value silently degrades to its first queue —
|
|
# and an assertion on the bare string passes either way, which is how
|
|
# that would have shipped.
|
|
assert f'CELERY_QUEUES="{",".join(lane.queues)}"' in env
|
|
|
|
|
|
def test_each_program_invokes_the_lane_s_entrypoint_role_not_its_name():
|
|
"""`maintenance_long` is the plain `worker` role pointed at a different
|
|
queue — exactly as docker-compose starts it today. Invoking
|
|
`entrypoint.sh maintenance_long` would hit the unknown-role branch and
|
|
exit 1 on every restart."""
|
|
cp = _parse(with_ml=True)
|
|
for lane in LANES:
|
|
command = cp.get(f"program:{lane.name}", "command")
|
|
assert f"entrypoint.sh {lane.entrypoint_role}" in command
|
|
|
|
|
|
def test_every_entrypoint_role_a_lane_names_actually_exists():
|
|
"""Reads entrypoint.sh itself. The generator can only emit a role name;
|
|
whether the script handles it is a separate fact, and getting it wrong
|
|
fails at container start rather than here."""
|
|
from pathlib import Path
|
|
|
|
script = Path(__file__).resolve().parents[1] / "entrypoint.sh"
|
|
text = script.read_text()
|
|
for lane in LANES:
|
|
# Roles are `case` arms: ` worker)` possibly in an alternation.
|
|
assert f" {lane.entrypoint_role})" in text or \
|
|
f"|{lane.entrypoint_role})" in text, \
|
|
f"{lane.name} names entrypoint role {lane.entrypoint_role!r}, which does not exist"
|
|
|
|
|
|
# --- shutdown ----------------------------------------------------------------
|
|
|
|
|
|
def test_every_program_signals_its_whole_process_group():
|
|
"""Celery's prefork pool forks children. A TERM delivered only to the
|
|
parent leaves them running and holding tasks — a 'graceful' shutdown that
|
|
orphans workers. The `sh -c … | sed` wrapper makes this doubly necessary:
|
|
without it the signal reaches the shell holding the pipeline, not celery."""
|
|
cp = _parse(with_ml=True)
|
|
for section in cp.sections():
|
|
if not section.startswith("program:"):
|
|
continue
|
|
assert cp.get(section, "stopasgroup") == "true", section
|
|
assert cp.get(section, "killasgroup") == "true", section
|
|
|
|
|
|
def test_the_long_maintenance_lane_keeps_its_180s_drain():
|
|
"""The per-service stop_grace_period values from the multi-service stack
|
|
are preserved per program. maintenance_long runs DB backups and library
|
|
audits; cutting its drain turns a restart into a SIGKILL mid-backup."""
|
|
cp = _parse()
|
|
assert cp.getint("program:maintenance_long", "stopwaitsecs") == 180
|
|
|
|
|
|
def test_no_program_waits_longer_than_the_compose_stop_grace_period():
|
|
"""The container gets ONE timeout and the programs stop in parallel, so it
|
|
must cover the slowest. If a lane's stopwaitsecs ever exceeds what
|
|
docker-compose.single.yml allows, docker kills the container while that
|
|
lane still believes it has time to drain."""
|
|
import re
|
|
from pathlib import Path
|
|
|
|
compose = (Path(__file__).resolve().parents[1] / "docker-compose.single.yml").read_text()
|
|
m = re.search(r"stop_grace_period:\s*(\d+)s", compose)
|
|
assert m, "docker-compose.single.yml has no stop_grace_period"
|
|
grace = int(m.group(1))
|
|
|
|
cp = _parse(with_ml=True)
|
|
for section in cp.sections():
|
|
if section.startswith("program:"):
|
|
assert cp.getint(section, "stopwaitsecs") <= grace, section
|
|
|
|
|
|
# --- what runs, and how much ------------------------------------------------
|
|
|
|
|
|
def test_a_zero_slot_lane_still_gets_a_running_process():
|
|
"""ML ships at 0 slots and disabled — but `add_consumer` needs something to
|
|
reach. With no process there would be nothing for the UI switch to switch,
|
|
and enabling tagging could not work at all."""
|
|
cp = _parse(with_ml=True)
|
|
assert LANES_BY_NAME["ml"].default_slots == 0
|
|
env = cp.get("program:ml", "environment")
|
|
assert "CELERY_CONCURRENCY=1" in env
|
|
|
|
|
|
def test_programs_restart_but_back_off_rather_than_looping():
|
|
"""A lane that dies instantly and repeatedly is a broken image, not a
|
|
transient fault. Unbounded restarts would burn a core forever and bury the
|
|
original error under its own noise."""
|
|
cp = _parse()
|
|
for section in cp.sections():
|
|
if section.startswith("program:"):
|
|
assert cp.get(section, "autorestart") == "true", section
|
|
assert cp.getint(section, "startretries") >= 1, section
|
|
|
|
|
|
def test_web_starts_first_because_it_runs_the_migration():
|
|
"""A worker booting against an un-migrated schema fails in a way that
|
|
looks like application breakage rather than an ordering problem."""
|
|
cp = _parse()
|
|
assert cp.getint("program:web", "priority") == 1
|
|
|
|
|
|
def test_every_program_writes_to_the_container_stdout_with_its_lane_named():
|
|
"""Four celery workers and hypercorn on one stream are indistinguishable
|
|
without this. Unbuffered (`maxbytes 0`) so `docker logs` is live rather
|
|
than arriving in rotated chunks."""
|
|
cp = _parse(with_ml=True)
|
|
for section in cp.sections():
|
|
if not section.startswith("program:"):
|
|
continue
|
|
name = section.split(":", 1)[1]
|
|
assert cp.get(section, "stdout_logfile") == "/dev/fd/1", section
|
|
assert cp.getint(section, "stdout_logfile_maxbytes") == 0, section
|
|
assert cp.get(section, "redirect_stderr") == "true", section
|
|
assert f"[{name}] " in cp.get(section, "command"), section
|