On 2026-09-19 a host storage stall made one Postgres checkpoint of 14
buffers take 281 seconds against a 1.3-second baseline. The app restarted
into the tail of it, `get_maintenance_hour()` — the first DB read in
`before_serving` — hung with no deadline, Hypercorn killed the worker at
its 60-second lifespan timeout, and nothing retries a failed lifespan. A
five-minute disk hiccup became a three-hour outage that only a human
restart could clear. Every MCP call returned 405, which reads like a
routing fault and was nothing of the kind: nothing was serving.
Three changes, none of which prevent a stall — they stop a transient one
becoming a permanent one.
1. THE STARTUP READ IS BOUNDED (rule 156). `get_maintenance_hour` already
answered `_DEFAULT_HOUR` for a value it could not parse; a database
that will not answer in three seconds is the same class of "no usable
value here". The failure is now a WARNING naming the symptom — the
breadcrumb whose absence meant this was only diagnosable from
Postgres's own log — and a default run-hour, instead of the app.
2. THE BACKFILL NO LONGER RACES STARTUP. Its comment said it "never
blocks the server from accepting requests": true of requests, false of
startup, because the task began while `before_serving` was still
running and competed for the same pool. Both of the incident's
cancelled statements were in flight together. It now waits on a flag
released on the hook's way out — in a `finally`, never after the work
(rule 157), because an undeadlined wait is only safe when the wake-up
cannot be missed.
3. THE ENGINE CANNOT WAIT FOREVER TO CONNECT. asyncpg's default is 60s,
the whole lifespan budget spent before a query is sent. `command_timeout`
is deliberately NOT set alongside it and the comment says why: it would
apply to every statement, and this app runs long ones on purpose.
tests/test_startup_survives_a_slow_database.py asserts the shape rather
than the stall: a read that never returns still yields an hour, the
warning names the symptom, a healthy read is unaffected, the backfill
does no work before release, the flag is released even when startup
raises, and the engine's connect args carry a deadline but no blanket
statement timeout.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy