Files
FabledCurator/tests/test_gen_supervisord.py
T
bvandeusenandClaude Opus 5 445164c852
CI and images / lint (push) Successful in 4s
CI and images / extension-version (push) Successful in 4s
CI and images / frontend-build (push) Successful in 24s
CI and images / integration (push) Failing after 24s
CI and images / backend-lint-and-test (push) Failing after 34s
CI and images / sign-extension (push) Skipped
CI and images / build-web (push) Skipped
CI and images / smoke-web (push) Skipped
CI and images / promote (push) Skipped
CI and images / build-agent (push) Skipped
feat: one number per lane — the cap — and the autoscaler is the mechanism (4295)
Operator, 2026-09-23: *"auto should be always on, not a setting, so that idle
instances quiet down when not running. the number that is visible and
something the user can tweak and manage should be the cap itself the number of
running workers is handled by the autoscaling function which is always on."*

They are right, and the reason it was not built this way is worth stating: the
manual dial came first (steps 2-4) and the autoscaler came last (step 7), as
an opt-in BESIDE a control that already existed. Nothing ever asked whether
the dial should still exist once something could move it automatically. Each
step was defensible; the result was three operator settings over one number.

## `slots`, `enabled` and `autoscale` are gone

`slots` was a MEASUREMENT wearing a preference's clothes. How many workers a
lane runs is read live and moved every minute; storing it meant the operator
had to keep two numbers in agreement and the autoscaler had to be told it was
allowed to touch one of them.

`autoscale` gated the mechanism behind a choice, so a lane nobody opted in
never gave its workers back — which is why an idle instance never quieted
down.

`enabled` is derived: a cap of zero means no consumers. "Off" and "may use no
workers" were two spellings of one fact, stored separately, free to disagree.

## Two sweeps become one

`reconcile_lanes_sync` drove the pool to the stored `slots`; `autoscale_lanes_
sync` moved it away from that same number; and most of step 7's hardest
reasoning — a stored value that is a FLOOR, a target of `max(stored, current)`
— existed only to stop them fighting. Delete the stored number and the problem
is not solved, it is absent.

`size_lanes_sync` runs every minute and owns both consumers and pool size. It
also subsumes what the reconcile was for: a worker restarted at its ENV
concurrency is corrected on the next tick rather than after five.

Growth is immediate, shrink is one worker per tick. Deliberately asymmetric —
"always on" is only pleasant if the ramp keeps up, and +1/minute would take
four minutes to answer a burst. Being one worker too large for a minute costs
a sleeping process; being too small costs work not happening. For ML the
asymmetry matters most: every new slot reloads a multi-GB model, so the slow
shrink is what stops a quiet patch from paying that cost again a minute later.

## The caps ship at one, and zero for ML

Per the operator. Conservative on purpose — and a conservative default nobody
knows how to raise is just a slow product, which is the other half of what
they asked for:

    "there needs to be something that tells the user to bump those numbers to
     improve processing rate or they'd never know the controls exist."

So a lane running everything its cap allows while work piles up says so, in
its own row, with the headroom named: *"4,060 waiting and all 1 worker busy.
Raise the cap to run more at once — this machine allows up to 7."*

It fires only when raising the cap would actually help. Not when the lane is
keeping up, not when the sizing pass has room it has not taken, and not at the
machine ceiling — where "raise the cap" is advice nobody can take.

## Migration 0105 rewrites the caps rather than carrying them

The old defaults (4/2/2/1) bounded a manual control and were loose because
moving within them was the ordinary act. The number now means "the most
workers this lane may use", which is a different promise; carrying the old
figure over would quadruple the worker lane on every existing install at the
moment this deploys.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-23 12:48:22 -04:00

303 lines
13 KiB
Python

"""The generated supervisord config (milestone 422 step 5).
Asserts the config against the LANE TABLE rather than against a fixture of
expected text. A fixture would have to be updated whenever a lane changes,
which is the same hand-kept coupling generating the config exists to remove —
and it would pass while describing a container that does not match the
application's own idea of what it runs.
"""
from __future__ import annotations
import configparser
from backend.app.scripts import gen_supervisord as gen
from backend.app.services.worker_lanes import LANES, LANES_BY_NAME
# --- structure ---------------------------------------------------------------
def _parse(**kwargs) -> configparser.ConfigParser:
"""supervisord's config is ini, so parse it rather than grepping strings.
A substring assertion passes on a line that is present but malformed —
inside a comment, in the wrong section, or with a typo'd key that
supervisord silently ignores.
"""
cp = configparser.ConfigParser()
cp.read_string(gen.render(**kwargs))
return cp
def test_it_is_valid_ini_with_a_supervisord_section():
cp = _parse()
assert cp.has_section("supervisord")
# PID 1 in a container: daemonising would exit immediately and take the
# container with it.
assert cp.get("supervisord", "nodaemon") == "true"
def test_every_lane_gets_a_program():
"""One image carries every lane since step 6, so nothing is conditional.
A lane in LANES with no program is a queue with no consumer.
Compared over the `program:` sections ONLY, both directions: no lane
without a program, and no program that is not a lane. It used to compare
every section minus `[supervisord]`, which made it fail the moment the
config grew non-program plumbing — the control socket did exactly that,
and the property it exists for had not moved at all. A guard that fires on
a correct change is one people learn to edit rather than read.
"""
cp = _parse()
programs = {s for s in cp.sections() if s.startswith("program:")}
expected = {"program:web"} | {f"program:{lane.name}" for lane in LANES}
assert programs == expected
def test_the_ml_lane_runs_even_though_it_ships_off():
"""It holds a PROCESS and no model. `add_consumer` needs a running worker
to reach, so without this raising the cap would have nothing to reach —
and nothing is downloaded by starting it, which is what lets rule 164
permit the fetch at all."""
assert _parse().has_section("program:ml")
assert LANES_BY_NAME["ml"].default_slots_cap == 0
# --- the coupling this generator exists to guarantee -------------------------
def test_each_program_serves_exactly_its_lane_s_queues():
"""The whole point: the container's processes and the application's lane
table are one list. A queue in LANES with no program means work that
queues forever with nothing consuming it."""
cp = _parse()
for lane in LANES:
env = cp.get(f"program:{lane.name}", "environment")
# The QUOTED form. supervisord splits `environment` on commas, so an
# unquoted multi-queue value silently degrades to its first queue —
# and an assertion on the bare string passes either way, which is how
# that would have shipped.
assert f'CELERY_QUEUES="{",".join(lane.queues)}"' in env
def test_each_program_invokes_the_lane_s_entrypoint_role_not_its_name():
"""`maintenance_long` is the plain `worker` role pointed at a different
queue — exactly as docker-compose starts it today. Invoking
`entrypoint.sh maintenance_long` would hit the unknown-role branch and
exit 1 on every restart."""
cp = _parse()
for lane in LANES:
command = cp.get(f"program:{lane.name}", "command")
assert f"entrypoint.sh {lane.entrypoint_role}" in command
def test_every_entrypoint_role_a_lane_names_actually_exists():
"""Reads entrypoint.sh itself. The generator can only emit a role name;
whether the script handles it is a separate fact, and getting it wrong
fails at container start rather than here."""
from pathlib import Path
script = Path(__file__).resolve().parents[1] / "entrypoint.sh"
text = script.read_text()
for lane in LANES:
# Roles are `case` arms: ` worker)` possibly in an alternation.
assert f" {lane.entrypoint_role})" in text or \
f"|{lane.entrypoint_role})" in text, \
f"{lane.name} names entrypoint role {lane.entrypoint_role!r}, which does not exist"
# --- shutdown ----------------------------------------------------------------
def test_every_program_signals_its_whole_process_group():
"""Celery's prefork pool forks children. A TERM delivered only to the
parent leaves them running and holding tasks — a 'graceful' shutdown that
orphans workers. The `sh -c … | sed` wrapper makes this doubly necessary:
without it the signal reaches the shell holding the pipeline, not celery."""
cp = _parse()
for section in cp.sections():
if not section.startswith("program:"):
continue
assert cp.get(section, "stopasgroup") == "true", section
assert cp.get(section, "killasgroup") == "true", section
def test_the_long_maintenance_lane_keeps_its_180s_drain():
"""The per-service stop_grace_period values from the multi-service stack
are preserved per program. maintenance_long runs DB backups and library
audits; cutting its drain turns a restart into a SIGKILL mid-backup."""
cp = _parse()
assert cp.getint("program:maintenance_long", "stopwaitsecs") == 180
def test_no_program_waits_longer_than_the_compose_stop_grace_period():
"""The container gets ONE timeout and the programs stop in parallel, so it
must cover the slowest. If a lane's stopwaitsecs ever exceeds what
docker-compose.single.yml allows, docker kills the container while that
lane still believes it has time to drain."""
import re
from pathlib import Path
compose = (Path(__file__).resolve().parents[1] / "docker-compose.single.yml").read_text()
m = re.search(r"stop_grace_period:\s*(\d+)s", compose)
assert m, "docker-compose.single.yml has no stop_grace_period"
grace = int(m.group(1))
cp = _parse()
for section in cp.sections():
if section.startswith("program:"):
assert cp.getint(section, "stopwaitsecs") <= grace, section
# --- what runs, and how much ------------------------------------------------
def test_a_lane_capped_at_zero_still_gets_a_running_process():
"""ML ships at cap 0 — but `add_consumer` needs something to reach. With
no process there would be nothing for the cap to switch back on, and
enabling tagging could not work at all."""
cp = _parse()
assert LANES_BY_NAME["ml"].default_slots_cap == 0
env = cp.get("program:ml", "environment")
assert "CELERY_CONCURRENCY=1" in env
def test_programs_restart_but_back_off_rather_than_looping():
"""A lane that dies instantly and repeatedly is a broken image, not a
transient fault. Unbounded restarts would burn a core forever and bury the
original error under its own noise."""
cp = _parse()
for section in cp.sections():
if section.startswith("program:"):
assert cp.get(section, "autorestart") == "true", section
assert cp.getint(section, "startretries") >= 1, section
def test_web_starts_first_because_it_runs_the_migration():
"""A worker booting against an un-migrated schema fails in a way that
looks like application breakage rather than an ordering problem."""
cp = _parse()
assert cp.getint("program:web", "priority") == 1
def test_every_program_writes_to_the_container_stdout_with_its_lane_named():
"""Four celery workers and hypercorn on one stream are indistinguishable
without this. Unbuffered (`maxbytes 0`) so `docker logs` is live rather
than arriving in rotated chunks."""
cp = _parse()
for section in cp.sections():
if not section.startswith("program:"):
continue
name = section.split(":", 1)[1]
assert cp.get(section, "stdout_logfile") == "/dev/fd/1", section
assert cp.getint(section, "stdout_logfile_maxbytes") == 0, section
assert cp.get(section, "redirect_stderr") == "true", section
assert f"[{name}] " in cp.get(section, "command"), section
def test_every_lane_gets_its_own_celery_node_name():
"""The bug that made the single-container layout unusable, found the first
time CI actually booted it (run 7319).
Celery's default node name is `celery@<hostname>`, and these processes
share one hostname — so all four lanes registered as the SAME node:
DuplicateNodenameWarning: Received multiple replies from node name:
celery@72adc5b706a7
`inspect` then collapses four replies into one dict key, last one wins.
Three lanes read as absent, and WHICH three varied between calls:
lanes not answering: maintenance_long, ml, worker
lanes not answering: maintenance_long, scheduler, worker
Fatal twice: the composite healthcheck can never pass, so the container is
permanently unhealthy and Swarm restart-loops it; and `pool_grow` takes a
`destination` of node names, so the UI dial and the autoscaler would have
resized whichever lane happened to answer rather than the one asked for.
Asserted as DISTINCTNESS across the whole table rather than as a fixed
string per lane — the property is that no two lanes collide, which is what
actually broke, and it keeps holding when a lane is added.
"""
cp = _parse()
names = {}
for section in cp.sections():
if not section.startswith("program:"):
continue
lane = section.split(":", 1)[1]
if lane == "web":
continue # hypercorn, not a celery node
env = cp.get(section, "environment")
pairs = dict(
p.split("=", 1) for p in env.split(",") if "=" in p and not p.startswith('"')
)
assert "CELERY_NODENAME" in pairs, f"{lane} has no CELERY_NODENAME: {env}"
names[lane] = pairs["CELERY_NODENAME"]
assert len(set(names.values())) == len(names), (
f"two lanes share a celery node name, which is the collapse itself: {names}"
)
for lane, node in names.items():
assert node == lane, f"{lane} announces itself as {node!r}"
def test_supervisorctl_can_reach_supervisord():
"""Consolidation took away `docker ps` as the way to see the lanes, and
this is what replaces it.
Without the three control-socket sections supervisord runs perfectly and
`supervisorctl` cannot talk to it at all:
Error: .ini file does not include supervisorctl section
That is the first thing anyone reaches for when a lane misbehaves inside
the one container — `supervisorctl status` to see which processes are up,
`restart ml` to bounce one without taking the application down. Shipping
without it leaves an operator with five processes and no way to ask about
any of them.
Asserted as the three sections AGREEING on one socket path, not merely as
three sections existing: a serverurl pointing somewhere the server does
not listen fails exactly the same way, and reads as configured.
"""
cp = _parse()
for section in ("unix_http_server", "supervisorctl", "rpcinterface:supervisor"):
assert cp.has_section(section), f"no [{section}] — supervisorctl is blind"
listening = cp.get("unix_http_server", "file")
talking = cp.get("supervisorctl", "serverurl")
assert talking == f"unix://{listening}", (
f"supervisorctl talks to {talking}, supervisord listens on {listening}"
)
def test_each_program_starts_at_the_smallest_pool_the_control_path_allows():
"""The two ends of the same floor, asserted together.
`gen_supervisord` starts every lane at one process because billiard will
not run a pool of zero. The sizing pass has the same floor for the
opposite reason: it cannot SHRINK to zero either —
[ml] pidbox command error:
ValueError("Can't shrink pool. All processes busy!")
Live, 2026-09-23. ML started at one process and stored zero, so the
reconcile tried 1 -> 0 on every tick, billiard refused, and
`set_lane_slots_sync` — which returns True on SENDING the message —
reported the lane changed forever (lesson #4183, on the default
configuration of every install).
One constant now, read from the same module by both, rather than two that
must agree. Asserted through it rather than against a literal 1, so
raising the floor moves both ends at once.
"""
from backend.app.services.worker_lanes import MIN_POOL_SLOTS
cp = _parse()
for lane in LANES:
env = cp.get(f"program:{lane.name}", "environment")
assert f"CELERY_CONCURRENCY={MIN_POOL_SLOTS}," in env, (
f"{lane.name} starts at a size the control path cannot reach"
)