CI / lint (push) Successful in 2s
CI / extension-version (push) Successful in 2s
Build images / sign-extension (push) Successful in 3s
Build images / build-agent (push) Successful in 6s
CI / frontend-build (push) Successful in 20s
extension / lint (push) Successful in 23s
CI / backend-lint-and-test (push) Failing after 31s
CI / integration (push) Successful in 2m16s
Build images / build-ml (push) Successful in 3m8s
Build images / build-web (push) Successful in 3m16s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
Milestone 422 step 6. Dockerfile.ml is gone; the main image carries torch,
torchvision, transformers, onnxruntime and opencv, and serves every lane.
WHY IT HAD TO MERGE: step 5 runs every lane in one process tree, so a second
image would mean the `ml` lane could never be enabled from the UI — there
would be no worker in that container to enable. The switch needs something to
switch.
THE MODEL NO LONGER DOWNLOADS AT BOOT. `entrypoint.sh`'s ml-worker role ran
download_models before celery started, so every boot of that role reached
HuggingFace for ~3.5GB — a startup dependency on a third party for a feature
the operator may never use. Rule 164 permits a runtime fetch only for
something "optional and clearly off", so the fetch is now a TASK, enqueued
the moment the lane is ENABLED.
Being a task is what makes it visible: it gets a TaskRun row, so the download
shows in Activity with a duration and a status, and a failure is something an
operator can see and retry rather than a container that quietly never became
useful. Idempotent, so re-enabling a provisioned lane costs one no-op.
Enqueued only when the lane actually came ON (`enabled is True`, not the
resolved value) so re-saving slots does not re-fetch, and only when the
consumer change landed — a task queued onto a queue nothing consumes would
sit pending with no explanation.
`fabledcurator-ml` KEEPS PUBLISHING, from the merged Dockerfile. The
operator's Swarm stack references that name and lives outside this repo;
dropping it would not break their deploy, it would freeze it silently at the
last publish — the exact failure class this milestone keeps finding. Retiring
the NAME is its own task, gated on that stack moving. Same two-phase shape
#406 used for pixiv.
THREE LIVE BREAKAGES from deleting the file, found by grepping for it rather
than assuming the build was the only consumer:
- `docker-compose.override.yml` built the ml service from it (contributor
path would have failed at `docker compose build`).
- `tests/test_artifact_paths.py` pins the ml path set.
- `scripts/artifacts.sh` ML_PATHS named it. A path set naming a deleted file
silently stops contributing to the derived revision — which the reuse check
and the version string both read. That is #3202's recorded shape.
The `--with-ml` flag is gone from the generator and the healthcheck rather
than left defaulting to true. One image carries every lane now, so a flag
that can only be passed one way is a branch pretending to be a choice.
The advisory shipped in ecbd325 is what makes this honest to an adopter: the
lane says it is optional, names the model, and gives its download and
per-slot RAM before the switch is thrown.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
82 lines
2.8 KiB
Python
82 lines
2.8 KiB
Python
"""Container healthcheck for the single-container layout.
|
|
|
|
Milestone 422 step 5. Exit 0 healthy, non-zero unhealthy.
|
|
|
|
## Why this is not just "does :8080 answer"
|
|
|
|
In the multi-service stack every service has its OWN healthcheck, so a dead
|
|
worker turns that service unhealthy while web stays green — docker knows which
|
|
part failed. Collapsing them into one container collapses that too: a web-only
|
|
check would report a perfectly healthy container while every lane inside it
|
|
had crashed and been abandoned by supervisord after its retries.
|
|
|
|
So this asserts both halves: hypercorn answers, AND every lane this container
|
|
was configured to run is answering the broker.
|
|
|
|
## What it deliberately does NOT do
|
|
|
|
It does not read the database, and it does not consult the `enabled` flag. A
|
|
DISABLED lane still has a running process with its consumers cancelled (see
|
|
the config generator), so it answers `inspect` and is healthy. Health is
|
|
"is the process alive", and whether it should be consuming is a settings
|
|
question the reconcile owns — conflating them would make turning a lane off
|
|
in the UI mark the container unhealthy.
|
|
|
|
It also cannot distinguish "the broker is down" from "every lane is down",
|
|
and reports unhealthy either way. That is correct: a container that cannot
|
|
reach its broker is not serving, whichever half is at fault.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import sys
|
|
import urllib.error
|
|
import urllib.request
|
|
|
|
WEB_URL = "http://localhost:8080/api/health"
|
|
WEB_TIMEOUT = 5.0
|
|
|
|
|
|
def _web_ok() -> tuple[bool, str]:
|
|
try:
|
|
with urllib.request.urlopen(WEB_URL, timeout=WEB_TIMEOUT) as resp:
|
|
if resp.status == 200:
|
|
return True, ""
|
|
return False, f"web returned {resp.status}"
|
|
except (urllib.error.URLError, OSError) as exc:
|
|
return False, f"web unreachable: {exc}"
|
|
|
|
|
|
def _lanes_ok() -> tuple[bool, str]:
|
|
from ..services.worker_control import inspect_lanes_sync
|
|
from ..services.worker_lanes import LANES
|
|
|
|
# Every lane, ml included: one image carries them all since step 6, and a
|
|
# disabled lane still runs a process (consumers cancelled), so it answers
|
|
# inspect and is healthy. Health is "is the process alive"; whether it
|
|
# should be consuming is the reconcile's business.
|
|
expected = {lane.name for lane in LANES}
|
|
live = inspect_lanes_sync()
|
|
missing = sorted(n for n in expected if not live[n].present)
|
|
if missing:
|
|
return False, "lanes not answering: " + ", ".join(missing)
|
|
return True, ""
|
|
|
|
|
|
def main(argv: list[str] | None = None) -> int:
|
|
ok, detail = _web_ok()
|
|
if not ok:
|
|
print(detail, file=sys.stderr)
|
|
return 1
|
|
|
|
ok, detail = _lanes_ok()
|
|
if not ok:
|
|
print(detail, file=sys.stderr)
|
|
return 1
|
|
|
|
return 0
|
|
|
|
|
|
if __name__ == "__main__":
|
|
raise SystemExit(main())
|