Files
FabledCurator/entrypoint.sh
T
bvandeusenandClaude Opus 5 828c6a5ae3
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 2s
CI / frontend-build (push) Successful in 19s
CI / backend-lint-and-test (push) Successful in 31s
CI / integration (push) Successful in 2m6s
Build images / sign-extension (push) Successful in 3s
Build images / build-agent (push) Successful in 6s
Build images / build-web (push) Successful in 1m52s
Build images / smoke-web (push) Successful in 56s
Build images / promote (push) Skipped
fix: each lane needs its own celery node name, or three of four vanish (4295)
Found by the all-role smoke on its very first execution (run 7319), which is
the whole argument for having added it one commit ago.

Celery's default node name is `celery@<hostname>`. In the single-container
layout all four lanes share one hostname, so all four registered as the SAME
node. Celery says so itself:

    DuplicateNodenameWarning: Received multiple replies from node name:
    celery@72adc5b706a7

`inspect` collapses four replies into one dict key and the last one wins, so
three lanes read as absent — and WHICH three varies between calls:

    lanes not answering: maintenance_long, ml, worker
    lanes not answering: maintenance_long, scheduler, worker

Fatal twice over:

  * The composite healthcheck can never pass. In Swarm that is a container
    that never goes healthy — restart loop, then an automatic rollback of a
    deploy whose image was fine.
  * `pool_grow`/`pool_shrink` take a `destination` of node names. The UI dial
    and the autoscaler would have resized whichever lane happened to answer
    rather than the one asked for — silently, and differently each time.

Every celery role now starts with `-n "${CELERY_NODENAME:-celery}@%h"`, and
the generated supervisord config sets that per lane. The lanes become
worker@<cid>, scheduler@<cid>, maintenance_long@<cid>, ml@<cid> — distinct,
so inspect keeps four entries and `destination` addresses what it names.
`inspect_lanes_sync` maps hostname to lane by QUEUES, so nothing there
changes; it just stops having three of its four entries overwritten.

Unset, it falls back to `celery` — exactly celery's own default — so every
service in the multi-service stack is byte-identical to before, including the
`celery@$HOSTNAME` healthcheck in docker-compose.yml and in the operator's
Swarm stack.

The test asserts DISTINCTNESS across the whole lane table rather than a fixed
string per lane. The property that broke is that no two collide, and stating
it that way keeps holding when a lane is added.

This is the bug I said a live deploy was needed to find, found in CI instead
for the price of one `docker run` — and it would have met the operator as a
rollback loop on their first consolidated deploy.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-23 08:52:56 -04:00

139 lines
5.9 KiB
Bash
Executable File

#!/usr/bin/env bash
set -euo pipefail
# Defaults to the whole application (see the Dockerfile's CMD). Kept in step
# with that CMD deliberately: they are two doors to the same decision, and a
# disagreement between them would only show up as `docker run --entrypoint`
# behaving differently from `docker run`.
ROLE="${1:-all}"
# CELERY NODE NAME. Every celery role below starts with `-n $CELERY_NODENAME@%h`.
#
# Celery's default node name is `celery@<hostname>`, and in the single-
# container layout all four lanes share one hostname — so all four registered
# as the SAME node. celery's own words for it, observed on run 7319:
#
# DuplicateNodenameWarning: Received multiple replies from node name:
# celery@72adc5b706a7
#
# `inspect` then collapses four replies into one dict key and the last one
# wins, so three lanes read as absent and WHICH three varies per call:
#
# lanes not answering: maintenance_long, ml, worker
# lanes not answering: maintenance_long, scheduler, worker
#
# That is fatal twice over. The composite healthcheck can never pass, so the
# container is permanently unhealthy; and `pool_grow(destination=[hostname])`
# addresses a lane BY that name, so the UI dial and the autoscaler would have
# resized whichever lane happened to answer rather than the one asked for.
#
# The generated supervisord config sets this per lane. Unset — which is every
# service in the multi-service stack — it falls back to `celery`, exactly
# celery's own default, so `celery@$HOSTNAME` healthchecks there still work.
: "${CELERY_NODENAME:=celery}"
shift || true
case "$ROLE" in
web)
echo "[entrypoint] Running alembic upgrade head"
alembic upgrade head
echo "[entrypoint] Starting hypercorn on :8080"
# create_app is a factory — the `()` tells hypercorn to call it once
# and serve the returned Quart (ASGI) app, rather than treating the
# function itself as the application (which it then mis-invokes as WSGI).
# Default 4 workers (was 2): each worker is one asyncio loop, and a large
# file download occupies its worker for the transfer — 2 was too few once the
# GPU agent + the browser's thumbnail grid hit /images concurrently (they
# queued behind each other). Env-tunable via HYPERCORN_WORKERS.
exec hypercorn \
--bind 0.0.0.0:8080 \
--workers "${HYPERCORN_WORKERS:-4}" \
--access-logfile - \
"backend.app:create_app()"
;;
worker)
QUEUES="${CELERY_QUEUES:-default,import,thumbnail}"
CONCURRENCY="${CELERY_CONCURRENCY:-2}"
echo "[entrypoint] Starting Celery worker queues=$QUEUES concurrency=$CONCURRENCY"
exec celery -A backend.app.celery_app:celery worker \
-n "${CELERY_NODENAME:-celery}@%h" \
--loglevel=info \
-Q "$QUEUES" \
--concurrency="$CONCURRENCY"
;;
scheduler)
QUEUES="${CELERY_QUEUES:-maintenance,scan}"
# Honours CELERY_CONCURRENCY like the `worker` role does. It was hardcoded
# to 1, which was harmless while only compose started this lane and set no
# concurrency for it — but the generated supervisord config (milestone 422
# step 5) passes one, and a value silently ignored at boot would leave the
# lane at 1 until the reconcile sweep noticed, with nothing saying why.
CONCURRENCY="${CELERY_CONCURRENCY:-1}"
echo "[entrypoint] Starting Celery beat+worker queues=$QUEUES concurrency=$CONCURRENCY"
exec celery -A backend.app.celery_app:celery worker \
-n "${CELERY_NODENAME:-celery}@%h" \
--beat \
--loglevel=info \
-Q "$QUEUES" \
--concurrency="$CONCURRENCY"
;;
ml-worker)
# NO MODEL DOWNLOAD HERE (milestone 422 step 6). This used to run
# download_models before celery started, which made every boot of this
# role reach HuggingFace for ~3.5GB. Rule 164 permits a runtime fetch only
# for a feature that is "optional and clearly off" — so the fetch moved to
# the moment the operator ENABLES the lane, where it is visible, retryable
# and attributable, instead of being a silent precondition of starting.
#
# The worker therefore starts with no model present, which is correct: it
# is not consuming the ml queue until the lane is enabled, and enabling it
# is what enqueues ensure_models.
QUEUES="${CELERY_QUEUES:-ml}"
CONCURRENCY="${CELERY_CONCURRENCY:-1}"
echo "[entrypoint] Starting ML Celery worker queues=$QUEUES concurrency=$CONCURRENCY"
exec celery -A backend.app.celery_app:celery worker \
-n "${CELERY_NODENAME:-celery}@%h" \
--loglevel=info \
-Q "$QUEUES" \
--concurrency="$CONCURRENCY"
;;
all)
# The single-container layout (milestone 422 step 5): hypercorn plus one
# celery process per lane, under supervisord, in one container beside
# Postgres and Redis.
#
# The config is GENERATED from services/worker_lanes.LANES rather than
# checked in, so the processes this container runs and the lanes the
# application believes in cannot disagree — see the generator's docstring
# for why a static .conf would have been a fifth copy of the queue names.
#
# supervisord is PID 1 here and never reads the database. Every lane boots
# at its LANES default; the reconcile sweep raises it to whatever the
# operator stored, within one tick. That ordering is deliberate: settings
# adjust a baseline that already works, and can never prevent a boot.
CONF="${SUPERVISOR_CONF:-/tmp/supervisord.conf}"
echo "[entrypoint] Generating $CONF from the lane table"
python -m backend.app.scripts.gen_supervisord > "$CONF"
echo "[entrypoint] Starting supervisord (web + worker lanes)"
exec supervisord -c "$CONF"
;;
shell|bash)
exec /bin/bash "$@"
;;
alembic)
exec alembic "$@"
;;
*)
echo "[entrypoint] Unknown role: $ROLE" >&2
echo "[entrypoint] Valid roles: all | web | worker | scheduler | maintenance_long | ml | ml-worker | shell | alembic" >&2
exit 1
;;
esac