Files
FabledCurator/backend/app/scripts/gen_supervisord.py
T
bvandeusenandClaude Opus 5 ffcd13096a
CI / lint (push) Successful in 2s
CI / extension-version (push) Successful in 2s
Build images / sign-extension (push) Successful in 3s
Build images / build-agent (push) Successful in 6s
CI / frontend-build (push) Successful in 20s
extension / lint (push) Successful in 23s
CI / backend-lint-and-test (push) Failing after 31s
CI / integration (push) Successful in 2m16s
Build images / build-ml (push) Successful in 3m8s
Build images / build-web (push) Successful in 3m16s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
feat: one image for every lane, with the model fetch gated on enabling (4296)
Milestone 422 step 6. Dockerfile.ml is gone; the main image carries torch,
torchvision, transformers, onnxruntime and opencv, and serves every lane.

WHY IT HAD TO MERGE: step 5 runs every lane in one process tree, so a second
image would mean the `ml` lane could never be enabled from the UI — there
would be no worker in that container to enable. The switch needs something to
switch.

THE MODEL NO LONGER DOWNLOADS AT BOOT. `entrypoint.sh`'s ml-worker role ran
download_models before celery started, so every boot of that role reached
HuggingFace for ~3.5GB — a startup dependency on a third party for a feature
the operator may never use. Rule 164 permits a runtime fetch only for
something "optional and clearly off", so the fetch is now a TASK, enqueued
the moment the lane is ENABLED.

Being a task is what makes it visible: it gets a TaskRun row, so the download
shows in Activity with a duration and a status, and a failure is something an
operator can see and retry rather than a container that quietly never became
useful. Idempotent, so re-enabling a provisioned lane costs one no-op.

Enqueued only when the lane actually came ON (`enabled is True`, not the
resolved value) so re-saving slots does not re-fetch, and only when the
consumer change landed — a task queued onto a queue nothing consumes would
sit pending with no explanation.

`fabledcurator-ml` KEEPS PUBLISHING, from the merged Dockerfile. The
operator's Swarm stack references that name and lives outside this repo;
dropping it would not break their deploy, it would freeze it silently at the
last publish — the exact failure class this milestone keeps finding. Retiring
the NAME is its own task, gated on that stack moving. Same two-phase shape
#406 used for pixiv.

THREE LIVE BREAKAGES from deleting the file, found by grepping for it rather
than assuming the build was the only consumer:

- `docker-compose.override.yml` built the ml service from it (contributor
  path would have failed at `docker compose build`).
- `tests/test_artifact_paths.py` pins the ml path set.
- `scripts/artifacts.sh` ML_PATHS named it. A path set naming a deleted file
  silently stops contributing to the derived revision — which the reuse check
  and the version string both read. That is #3202's recorded shape.

The `--with-ml` flag is gone from the generator and the healthcheck rather
than left defaulting to true. One image carries every lane now, so a flag
that can only be passed one way is a branch pretending to be a choice.

The advisory shipped in ecbd325 is what makes this honest to an adopter: the
lane says it is optional, names the model, and gives its download and
per-slot RAM before the switch is thrown.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-22 08:51:44 -04:00

179 lines
7.3 KiB
Python

"""Emit a supervisord config for the single-container layout.
Milestone 422 step 5. Writes to stdout; `entrypoint.sh all` redirects it to a
file and execs supervisord against it.
## Why this is generated and not a checked-in .conf
A static config would spell out each lane's `-Q` list, and that would be a
FIFTH hand-kept copy of the queue names — after `celery_app.task_routes`, and
the three collapsed in steps 1, 2 and 4 (`service_roster.ROLE_NAMES`,
`system_activity._QUEUE_NAMES`, and the Activity filter). Every one of those
had already drifted by the time it was found.
Generating from `worker_lanes.LANES` makes a stronger guarantee than "they
match today": the processes this container runs and the lanes the application
believes in are the same list, so a lane added to `LANES` gets a process
without anyone remembering to add one, and a queue can never end up with no
consumer because a config file was missed.
## Why supervisord
It is one pip dependency on an image that is already Python, and it does the
four things this needs without being clever: restart a program that exits,
give each one its OWN stop timeout, signal the process GROUP rather than the
leader, and put every program's output on one stdout.
The process-group part is not a detail. Celery's prefork pool forks children,
and a TERM delivered only to the parent leaves them running — which is how a
"graceful" shutdown turns into orphaned workers holding tasks. `stopasgroup`
and `killasgroup` are both set for every program.
s6-overlay is the other standard answer and would work; it needs a build-time
download and a second mental model, and its advantage (correct PID-1 signal
and zombie handling) is available here from `init: true` in compose, which
puts tini in front of supervisord. Neither choice reaches the application —
nothing in FC talks to the supervisor — so this is reversible without touching
a line of product code.
## Every lane, including ml
Step 6 merged the images, so this one carries torch and the ML requirements
and the `ml` lane gets a program like any other. It starts at one slot with
its consumers CANCELLED — `enabled=false` in the seeded settings — so it
holds a process and no model. That matters: `add_consumer` needs a running
worker to reach, and without one the UI switch would have nothing to switch.
Nothing is downloaded by starting it. The model fetch is enqueued when the
lane is enabled, which is what lets rule 164 permit a runtime fetch at all.
"""
from __future__ import annotations
import argparse
import shlex
import sys
from ..services.worker_lanes import LANES, Lane
# One number for the whole container, and it must cover the SLOWEST lane —
# docker gives the container a single stop timeout, where compose today gives
# each service its own (90/60/180/120s). `maintenance_long` is the 180s one:
# DB backups, library audits and translation backfill. Anything less turns a
# routine restart into a SIGKILL mid-backup.
#
# Per-program values below are the old per-service ones, preserved: supervisord
# waits `stopwaitsecs` for each, and they stop in parallel, so the container's
# own timeout needs to cover the max rather than the sum.
STOP_WAIT_SECONDS: dict[str, int] = {
"worker": 90,
"scheduler": 60,
"maintenance_long": 180,
"ml": 120,
}
DEFAULT_STOP_WAIT = 60
def _program(lane: Lane, *, slots: int) -> str:
"""One [program:x] block.
`stdout_logfile=/dev/fd/1` with maxbytes 0 puts the lane's output straight
on the container's stdout unbuffered, so `docker logs` shows every lane
interleaved rather than supervisord swallowing them into rotated files.
The output is prefixed through `sed` so a line can be attributed to a lane
— four celery workers and hypercorn on one stream are otherwise
indistinguishable. The shell that the pipe requires is exactly why
`stopasgroup` matters: the signal has to reach the celery process, not the
`sh` holding the pipeline.
"""
inner = f"./entrypoint.sh {lane.entrypoint_role}"
prefixed = f"{inner} 2>&1 | sed -u 's/^/[{lane.name}] /'"
stop_wait = STOP_WAIT_SECONDS.get(lane.name, DEFAULT_STOP_WAIT)
return "\n".join([
f"[program:{lane.name}]",
f"command=sh -c {shlex.quote(prefixed)}",
# QUOTED, and that is load-bearing. supervisord parses `environment`
# as a COMMA-separated KEY=VALUE list, so an unquoted queue list reads
# as CELERY_QUEUES=default followed by three malformed entries — and
# the lane would consume only its first queue. Silent: the worker
# starts, reports healthy, and simply never picks up `import`.
f'environment=CELERY_QUEUES="{",".join(lane.queues)}",'
f"CELERY_CONCURRENCY={slots}",
"autostart=true",
"autorestart=true",
# A lane that dies instantly and repeatedly is a broken image, not a
# transient fault. Backing off stops it burning a core in a restart
# loop while still recovering from a one-off crash.
"startretries=3",
"startsecs=5",
f"stopwaitsecs={stop_wait}",
"stopasgroup=true",
"killasgroup=true",
"stdout_logfile=/dev/fd/1",
"stdout_logfile_maxbytes=0",
"redirect_stderr=true",
"",
])
def _web_program() -> str:
"""hypercorn. Started FIRST (priority) because its role runs
`alembic upgrade head`, and a worker that boots against an un-migrated
schema fails in a way that looks like application breakage."""
prefixed = "./entrypoint.sh web 2>&1 | sed -u 's/^/[web] /'"
return "\n".join([
"[program:web]",
f"command=sh -c {shlex.quote(prefixed)}",
"priority=1",
"autostart=true",
"autorestart=true",
"startretries=3",
"startsecs=5",
# Short: HTTP requests and the occasional file download. Matches the
# 30s the operator's production stack gives the web service.
"stopwaitsecs=30",
"stopasgroup=true",
"killasgroup=true",
"stdout_logfile=/dev/fd/1",
"stdout_logfile_maxbytes=0",
"redirect_stderr=true",
"",
])
def render() -> str:
parts = [
"\n".join([
"[supervisord]",
# PID 1 in the container, so it must not daemonise.
"nodaemon=true",
# supervisord's OWN log. /dev/fd/1 keeps it on the container's
# stdout beside the programs rather than in a file nobody reads.
"logfile=/dev/fd/1",
"logfile_maxbytes=0",
"loglevel=info",
"",
]),
_web_program(),
]
# Lanes after web, in LANES order, so the log reads in a stable sequence.
for lane in LANES:
# A lane configured at zero slots still gets a PROCESS, at one slot
# with its consumers cancelled by the reconcile. Without a running
# worker there is nothing for `add_consumer` to reach, so enabling the
# lane from the UI could not work at all — the process has to exist for
# the switch to have something to switch.
parts.append(_program(lane, slots=max(1, lane.default_slots)))
return "\n".join(parts)
def main(argv: list[str] | None = None) -> int:
ap = argparse.ArgumentParser(description=__doc__)
ap.parse_args(argv)
sys.stdout.write(render())
return 0
if __name__ == "__main__":
raise SystemExit(main())