fix(startup): a slow database costs seconds, not the instance (#4181)
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m31s
CI & Build / Build & push image (push) Successful in 26s
CI & Build / Python lint (push) Successful in 3s
CI & Build / Plugin hooks (push) Successful in 10s
CI & Build / integration (push) Successful in 50s
CI & Build / TypeScript typecheck (push) Successful in 55s
CI & Build / Python tests (push) Successful in 1m31s
CI & Build / Build & push image (push) Successful in 26s
On 2026-09-19 a host storage stall made one Postgres checkpoint of 14 buffers take 281 seconds against a 1.3-second baseline. The app restarted into the tail of it, `get_maintenance_hour()` — the first DB read in `before_serving` — hung with no deadline, Hypercorn killed the worker at its 60-second lifespan timeout, and nothing retries a failed lifespan. A five-minute disk hiccup became a three-hour outage that only a human restart could clear. Every MCP call returned 405, which reads like a routing fault and was nothing of the kind: nothing was serving. Three changes, none of which prevent a stall — they stop a transient one becoming a permanent one. 1. THE STARTUP READ IS BOUNDED (rule 156). `get_maintenance_hour` already answered `_DEFAULT_HOUR` for a value it could not parse; a database that will not answer in three seconds is the same class of "no usable value here". The failure is now a WARNING naming the symptom — the breadcrumb whose absence meant this was only diagnosable from Postgres's own log — and a default run-hour, instead of the app. 2. THE BACKFILL NO LONGER RACES STARTUP. Its comment said it "never blocks the server from accepting requests": true of requests, false of startup, because the task began while `before_serving` was still running and competed for the same pool. Both of the incident's cancelled statements were in flight together. It now waits on a flag released on the hook's way out — in a `finally`, never after the work (rule 157), because an undeadlined wait is only safe when the wake-up cannot be missed. 3. THE ENGINE CANNOT WAIT FOREVER TO CONNECT. asyncpg's default is 60s, the whole lifespan budget spent before a query is sent. `command_timeout` is deliberately NOT set alongside it and the comment says why: it would apply to every statement, and this app runs long ones on purpose. tests/test_startup_survives_a_slow_database.py asserts the shape rather than the stall: a read that never returns still yields an hour, the warning names the symptom, a healthy read is unaffected, the backfill does no work before release, the flag is released even when startup raises, and the engine's connect args carry a deadline but no blanket statement timeout. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01821k5B3Ysecp9fNYs92Kuy
This commit is contained in:
@@ -3,11 +3,39 @@ from sqlalchemy.orm import DeclarativeBase
|
||||
|
||||
from scribe.config import Config
|
||||
|
||||
# Named rather than inlined so the deadline below is READABLE. SQLAlchemy
|
||||
# captures `connect_args` in a closure and merges it at connect time, so an
|
||||
# inline dict cannot be recovered from the engine — and a guard that cannot
|
||||
# read the value it guards is a guard that passes forever.
|
||||
_CONNECT_ARGS: dict = {
|
||||
# A DEADLINE ON ESTABLISHING A CONNECTION (#4181, rule 156).
|
||||
#
|
||||
# asyncpg's `timeout` bounds the CONNECT — the TCP handshake plus session
|
||||
# setup — and nothing else. Against a host whose storage has wedged, that
|
||||
# handshake does not fail, it waits, and without this the wait is
|
||||
# asyncpg's own 60-second default: the entire lifespan budget spent before
|
||||
# a single query is even sent. Ten seconds is far longer than a healthy
|
||||
# local connect (single-digit milliseconds) and short enough to leave room
|
||||
# to fail usefully rather than be killed.
|
||||
#
|
||||
# WHAT THIS DOES NOT COVER, said plainly so the next reader doesn't assume
|
||||
# it does: a query on an already-open connection, which includes
|
||||
# `pool_pre_ping`'s liveness check. Bounding those is `command_timeout`,
|
||||
# and that is deliberately NOT set here — it would apply to every
|
||||
# statement, and this app legitimately runs long ones (the embedding
|
||||
# backfills, VACUUM ANALYZE). A blanket statement deadline would trade
|
||||
# this failure mode for a worse one. Callers that must not hang — the
|
||||
# lifespan hook above all — bound their own await instead; see
|
||||
# `_STARTUP_READ_TIMEOUT` in services/db_maintenance_scheduler.py.
|
||||
"timeout": 10,
|
||||
}
|
||||
|
||||
engine = create_async_engine(
|
||||
Config.DATABASE_URL,
|
||||
echo=False,
|
||||
pool_pre_ping=True,
|
||||
pool_recycle=1800,
|
||||
connect_args=_CONNECT_ARGS,
|
||||
)
|
||||
async_session = async_sessionmaker(engine, class_=AsyncSession, expire_on_commit=False)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user