Files
FabledCurator/tests/test_wait_for_deps.py
T
bvandeusenandClaude Opus 5 b2da3acce9
CI / lint (push) Successful in 2s
CI / extension-version (push) Successful in 2s
Build images / sign-extension (push) Successful in 3s
Build images / build-agent (push) Successful in 6s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Failing after 32s
Build images / build-web (push) Successful in 1m43s
CI / integration (push) Successful in 2m12s
Build images / smoke-web (push) Successful in 57s
Build images / promote (push) Skipped
feat: wait for Postgres and Redis before starting work (4295)
Operator, 2026-09-23: *"it's a single container that need to connect
successfully to redis and postgres before starting work shouldn't that simply
be a check (with retries) at the start of the container."*

Yes, and the consolidated layout makes it necessary rather than tidy.

Swarm has no ordering primitive — it ignores `depends_on` outright — so every
service in a stack starts at once and this container has always raced its own
database on a cold deploy. The multi-service stack hid how sharp that is: a
`web` task that failed `alembic upgrade head` against a still-initialising
Postgres simply died, and Swarm restarted it until it worked. Nobody ever saw
a problem worth naming.

Consolidation removes that safety net. Each supervisord program gets
`startretries=3`, so three quick failures put the program in FATAL and leave
it there — supervisord keeps running, the container keeps running, and the
application never starts. It would present as a permanently unhealthy
container whose image was fine and whose database merely took twenty seconds
to initialise, which is a miserable thing to debug on a first deploy.

A TCP connect, not a query: the same probe ci.yml's integration lane and the
build smoke already use. It answers the question actually being asked — is
something listening — and cannot fail for a reason that retrying will never
fix. A real query would be a stronger readiness signal and a worse gate,
since a wrong password or a missing database is not transient, and a loop
waiting for one to heal turns a five-second misconfiguration into a
two-minute timeout with a misleading message. Those belong to alembic, which
runs seconds later and says exactly what is wrong.

Targets are derived from the same env the application reads, so the wait
cannot drift from what the app will actually connect to — a gate checking a
different host than the app uses is worse than no gate.

Bounded at 120s (rule 156), reporting every few attempts so `docker logs` on
a waiting container says what it is waiting for. Skipped for `shell`, which
exists precisely for when something else is broken and you want a prompt
rather than a gate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-23 09:17:58 -04:00

92 lines
3.1 KiB
Python

"""The startup gate that stands in for Swarm's missing `depends_on`.
Swarm ignores `depends_on`, so the container races its own database on every
cold deploy. Under supervisord a program that fails three times goes FATAL and
stays there, so losing that race does not self-heal the way it did when Swarm
restarted a failed task — it leaves a running container with a dead
application inside it.
"""
from __future__ import annotations
import pytest
from backend.app.scripts import wait_for_deps as w
def test_it_waits_for_both_postgres_and_redis(monkeypatch):
monkeypatch.setenv("DB_HOST", "postgres")
monkeypatch.setenv("DB_PORT", "5432")
monkeypatch.setenv("CELERY_BROKER_URL", "redis://redis:6379/0")
assert w.targets() == [
("postgres", ("postgres", 5432)),
("redis", ("redis", 6379)),
]
def test_the_targets_come_from_the_same_env_the_app_reads(monkeypatch):
"""A gate that checks a different host than the application connects to is
worse than no gate — it would pass while the app still cannot reach its
database."""
monkeypatch.setenv("DB_HOST", "10.0.0.5")
monkeypatch.setenv("DB_PORT", "6543")
monkeypatch.setenv("CELERY_BROKER_URL", "redis://broker.internal:6380/2")
assert w.targets() == [
("postgres", ("10.0.0.5", 6543)),
("redis", ("broker.internal", 6380)),
]
def test_a_broker_url_without_a_port_falls_back_to_the_default(monkeypatch):
monkeypatch.delenv("DB_HOST", raising=False)
monkeypatch.setenv("CELERY_BROKER_URL", "redis://redis/0")
assert w.targets() == [("redis", ("redis", 6379))]
def test_nothing_configured_is_not_an_error(monkeypatch):
"""`shell` and one-off runs are legitimate. Refusing to start would make
this gate the reason a debugging container will not boot."""
monkeypatch.delenv("DB_HOST", raising=False)
monkeypatch.delenv("CELERY_BROKER_URL", raising=False)
assert w.targets() == []
assert w.main([]) == 0
def test_it_returns_as_soon_as_the_port_accepts(monkeypatch):
attempts = {"n": 0}
def accepts(host, port):
attempts["n"] += 1
return attempts["n"] >= 3
monkeypatch.setattr(w, "_accepts", accepts)
monkeypatch.setattr(w.time, "sleep", lambda _: None)
assert w.wait("postgres", "h", 5432, deadline=float("inf")) is True
assert attempts["n"] == 3
def test_it_gives_up_at_the_deadline_rather_than_hanging(monkeypatch):
"""A wait with no deadline is a bug (rule 156). A container that never
starts and never says why is the shape this must not have."""
monkeypatch.setattr(w, "_accepts", lambda h, p: False)
monkeypatch.setattr(w.time, "sleep", lambda _: None)
clock = iter([0.0, 0.0, 99.0])
assert w.wait(
"postgres", "h", 5432, deadline=10.0, now=lambda: next(clock),
) is False
def test_a_dependency_that_never_comes_up_fails_the_boot(monkeypatch):
monkeypatch.setenv("DB_HOST", "postgres")
monkeypatch.delenv("CELERY_BROKER_URL", raising=False)
monkeypatch.setattr(w, "_accepts", lambda h, p: False)
monkeypatch.setattr(w.time, "sleep", lambda _: None)
assert w.main(["--timeout", "0"]) == 1