CI / lint (push) Successful in 2s
CI / extension-version (push) Successful in 2s
Build images / sign-extension (push) Successful in 3s
Build images / build-agent (push) Successful in 6s
CI / frontend-build (push) Successful in 20s
extension / lint (push) Successful in 23s
CI / backend-lint-and-test (push) Failing after 31s
CI / integration (push) Successful in 2m16s
Build images / build-ml (push) Successful in 3m8s
Build images / build-web (push) Successful in 3m16s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
Milestone 422 step 6. Dockerfile.ml is gone; the main image carries torch,
torchvision, transformers, onnxruntime and opencv, and serves every lane.
WHY IT HAD TO MERGE: step 5 runs every lane in one process tree, so a second
image would mean the `ml` lane could never be enabled from the UI — there
would be no worker in that container to enable. The switch needs something to
switch.
THE MODEL NO LONGER DOWNLOADS AT BOOT. `entrypoint.sh`'s ml-worker role ran
download_models before celery started, so every boot of that role reached
HuggingFace for ~3.5GB — a startup dependency on a third party for a feature
the operator may never use. Rule 164 permits a runtime fetch only for
something "optional and clearly off", so the fetch is now a TASK, enqueued
the moment the lane is ENABLED.
Being a task is what makes it visible: it gets a TaskRun row, so the download
shows in Activity with a duration and a status, and a failure is something an
operator can see and retry rather than a container that quietly never became
useful. Idempotent, so re-enabling a provisioned lane costs one no-op.
Enqueued only when the lane actually came ON (`enabled is True`, not the
resolved value) so re-saving slots does not re-fetch, and only when the
consumer change landed — a task queued onto a queue nothing consumes would
sit pending with no explanation.
`fabledcurator-ml` KEEPS PUBLISHING, from the merged Dockerfile. The
operator's Swarm stack references that name and lives outside this repo;
dropping it would not break their deploy, it would freeze it silently at the
last publish — the exact failure class this milestone keeps finding. Retiring
the NAME is its own task, gated on that stack moving. Same two-phase shape
#406 used for pixiv.
THREE LIVE BREAKAGES from deleting the file, found by grepping for it rather
than assuming the build was the only consumer:
- `docker-compose.override.yml` built the ml service from it (contributor
path would have failed at `docker compose build`).
- `tests/test_artifact_paths.py` pins the ml path set.
- `scripts/artifacts.sh` ML_PATHS named it. A path set naming a deleted file
silently stops contributing to the derived revision — which the reuse check
and the version string both read. That is #3202's recorded shape.
The `--with-ml` flag is gone from the generator and the healthcheck rather
than left defaulting to true. One image carries every lane now, so a flag
that can only be passed one way is a branch pretending to be a choice.
The advisory shipped in ecbd325 is what makes this honest to an adopter: the
lane says it is optional, names the model, and gives its download and
per-slot RAM before the switch is thrown.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
107 lines
4.3 KiB
Bash
Executable File
107 lines
4.3 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
set -euo pipefail
|
|
|
|
ROLE="${1:-web}"
|
|
shift || true
|
|
|
|
case "$ROLE" in
|
|
web)
|
|
echo "[entrypoint] Running alembic upgrade head"
|
|
alembic upgrade head
|
|
echo "[entrypoint] Starting hypercorn on :8080"
|
|
# create_app is a factory — the `()` tells hypercorn to call it once
|
|
# and serve the returned Quart (ASGI) app, rather than treating the
|
|
# function itself as the application (which it then mis-invokes as WSGI).
|
|
# Default 4 workers (was 2): each worker is one asyncio loop, and a large
|
|
# file download occupies its worker for the transfer — 2 was too few once the
|
|
# GPU agent + the browser's thumbnail grid hit /images concurrently (they
|
|
# queued behind each other). Env-tunable via HYPERCORN_WORKERS.
|
|
exec hypercorn \
|
|
--bind 0.0.0.0:8080 \
|
|
--workers "${HYPERCORN_WORKERS:-4}" \
|
|
--access-logfile - \
|
|
"backend.app:create_app()"
|
|
;;
|
|
|
|
worker)
|
|
QUEUES="${CELERY_QUEUES:-default,import,thumbnail}"
|
|
CONCURRENCY="${CELERY_CONCURRENCY:-2}"
|
|
echo "[entrypoint] Starting Celery worker queues=$QUEUES concurrency=$CONCURRENCY"
|
|
exec celery -A backend.app.celery_app:celery worker \
|
|
--loglevel=info \
|
|
-Q "$QUEUES" \
|
|
--concurrency="$CONCURRENCY"
|
|
;;
|
|
|
|
scheduler)
|
|
QUEUES="${CELERY_QUEUES:-maintenance,scan}"
|
|
# Honours CELERY_CONCURRENCY like the `worker` role does. It was hardcoded
|
|
# to 1, which was harmless while only compose started this lane and set no
|
|
# concurrency for it — but the generated supervisord config (milestone 422
|
|
# step 5) passes one, and a value silently ignored at boot would leave the
|
|
# lane at 1 until the reconcile sweep noticed, with nothing saying why.
|
|
CONCURRENCY="${CELERY_CONCURRENCY:-1}"
|
|
echo "[entrypoint] Starting Celery beat+worker queues=$QUEUES concurrency=$CONCURRENCY"
|
|
exec celery -A backend.app.celery_app:celery worker \
|
|
--beat \
|
|
--loglevel=info \
|
|
-Q "$QUEUES" \
|
|
--concurrency="$CONCURRENCY"
|
|
;;
|
|
|
|
ml-worker)
|
|
# NO MODEL DOWNLOAD HERE (milestone 422 step 6). This used to run
|
|
# download_models before celery started, which made every boot of this
|
|
# role reach HuggingFace for ~3.5GB. Rule 164 permits a runtime fetch only
|
|
# for a feature that is "optional and clearly off" — so the fetch moved to
|
|
# the moment the operator ENABLES the lane, where it is visible, retryable
|
|
# and attributable, instead of being a silent precondition of starting.
|
|
#
|
|
# The worker therefore starts with no model present, which is correct: it
|
|
# is not consuming the ml queue until the lane is enabled, and enabling it
|
|
# is what enqueues ensure_models.
|
|
QUEUES="${CELERY_QUEUES:-ml}"
|
|
CONCURRENCY="${CELERY_CONCURRENCY:-1}"
|
|
echo "[entrypoint] Starting ML Celery worker queues=$QUEUES concurrency=$CONCURRENCY"
|
|
exec celery -A backend.app.celery_app:celery worker \
|
|
--loglevel=info \
|
|
-Q "$QUEUES" \
|
|
--concurrency="$CONCURRENCY"
|
|
;;
|
|
|
|
all)
|
|
# The single-container layout (milestone 422 step 5): hypercorn plus one
|
|
# celery process per lane, under supervisord, in one container beside
|
|
# Postgres and Redis.
|
|
#
|
|
# The config is GENERATED from services/worker_lanes.LANES rather than
|
|
# checked in, so the processes this container runs and the lanes the
|
|
# application believes in cannot disagree — see the generator's docstring
|
|
# for why a static .conf would have been a fifth copy of the queue names.
|
|
#
|
|
# supervisord is PID 1 here and never reads the database. Every lane boots
|
|
# at its LANES default; the reconcile sweep raises it to whatever the
|
|
# operator stored, within one tick. That ordering is deliberate: settings
|
|
# adjust a baseline that already works, and can never prevent a boot.
|
|
CONF="${SUPERVISOR_CONF:-/tmp/supervisord.conf}"
|
|
echo "[entrypoint] Generating $CONF from the lane table"
|
|
python -m backend.app.scripts.gen_supervisord > "$CONF"
|
|
echo "[entrypoint] Starting supervisord (web + worker lanes)"
|
|
exec supervisord -c "$CONF"
|
|
;;
|
|
|
|
shell|bash)
|
|
exec /bin/bash "$@"
|
|
;;
|
|
|
|
alembic)
|
|
exec alembic "$@"
|
|
;;
|
|
|
|
*)
|
|
echo "[entrypoint] Unknown role: $ROLE" >&2
|
|
echo "[entrypoint] Valid roles: all | web | worker | scheduler | maintenance_long | ml | ml-worker | shell | alembic" >&2
|
|
exit 1
|
|
;;
|
|
esac
|