CI and images / lint (push) Failing after 3s
CI and images / extension-version (push) Successful in 4s
CI and images / frontend-build (push) Successful in 31s
CI and images / backend-lint-and-test (push) Successful in 35s
CI and images / integration (push) Successful in 2m44s
CI and images / sign-extension (push) Skipped
CI and images / build-web (push) Skipped
CI and images / smoke-web (push) Skipped
CI and images / promote (push) Skipped
CI and images / build-agent (push) Skipped
Operator: "there is a repull every time this page loads is there a reason
this info isn't being tracked in the background and stored in some way?"
There was a reason and it had expired, and underneath it there was plain
waste.
The expired one: /api/system/workers was deliberately uncached because an
operator dragging the stepper must not be shown a pre-change value. That
stopped being true at 1353d34, when the UI began patching its row from the
write's reply instead of refetching.
The waste: size_worker_lanes already inspected the broker on a timer to
decide pool sizes — computing the pool, active, reserved and queue depth
the page shows, using them, and discarding them. The browser then asked
the broker for the same numbers four times a minute, per open tab.
So one inspect now feeds three things: the sizing decision, a stored
sample (worker_lane_sample, alembic 0107), and the celery roster. No
request path touches the broker at all — the roster refresh comes off
/api/system/health too, where it had been rate-limited to 20s and so made
worker liveness a function of whether anyone had a browser open.
Consequences, stated rather than hidden:
- The live figures are up to one sweep old. measured_at travels with each
lane and the page says how old, because a stale number presented as
current is how someone watches a queue "not move" that is moving.
- The sweep is the roster's only writer now, so its period and the
staleness thresholds are in a relationship. 60s against a 90s stale
threshold left one missed tick between normal and all-yellow — the
shape of lesson #4355 — so the period is 30s, named once in
worker_lanes, and system_health asserts its headroom at import with a
test stating the same thing in prose.
- An idle lane therefore also gives a worker back twice as fast. That is
the direction asked for: "idle instances quiet down when not running".
Also bounds the inspect in push_lane_cap, which was an await with no
deadline (rule 156) — harmless while it ran on a request, less so now
that it runs in a background task where a hang would be silent.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
204 lines
8.2 KiB
Python
204 lines
8.2 KiB
Python
"""Is every part of FabledCurator running? One verdict, one endpoint.
|
|
|
|
Milestone 365. The nav indicator and the System page both read this and
|
|
nothing else — composing a verdict is this module's job, not the UI's.
|
|
|
|
## Two kinds of part, answered two different ways
|
|
|
|
**Learned** — celery roles and the GPU agent, from `service_seen`. The
|
|
question is "how long since it checked in", and these are the parts that can
|
|
be ABSENT, which is the whole point: `celery inspect` alone reports presence,
|
|
so a dead worker is a shorter list rather than a red light.
|
|
|
|
**Probed live** — Postgres and Redis. Always expected, never learned, and a
|
|
last-seen for them would be actively misleading: that Redis answered thirty
|
|
seconds ago says nothing about now.
|
|
|
|
## This endpoint must never fail because something it checks has failed
|
|
|
|
The inversion is easy to write by accident and it destroys the feature exactly
|
|
when it is needed — a 500 when Redis is down, instead of `redis: down`. Every
|
|
probe is wrapped, every wait has a deadline (rule 156), and the roster refresh
|
|
swallows its own errors. The worst case is a part reported `unknown`, which is
|
|
a true statement.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import asyncio
|
|
import time
|
|
from datetime import UTC, datetime
|
|
|
|
from quart import Blueprint, jsonify
|
|
from sqlalchemy import select, text
|
|
|
|
from ..config import get_config
|
|
from ..extensions import get_session
|
|
from ..models import ServiceSeen
|
|
from ..services.worker_lanes import SWEEP_PERIOD_SECONDS
|
|
|
|
system_health_bp = Blueprint("system_health", __name__, url_prefix="/api/system")
|
|
|
|
# How long a learned part may go quiet before it is doubted, then disbelieved.
|
|
#
|
|
# These are deliberately generous, and the reason is a deploy rather than a
|
|
# worker: `docker compose up -d` rolls start-first, so a role is briefly served
|
|
# by two containers and then by neither while the old one drains. Thresholds
|
|
# tight enough to catch a crash in seconds would paint the page red every time
|
|
# the stack is updated, and an alarm that cries wolf on every deploy is one
|
|
# nobody reads. Tune down only after watching a real deploy pass through.
|
|
STALE_AFTER_SECONDS = 90
|
|
DOWN_AFTER_SECONDS = 300
|
|
|
|
# The celery roster is written by `size_worker_lanes` and by nothing else, so
|
|
# these thresholds are only meaningful against ITS cadence. Asserted at import
|
|
# rather than left to a reader, because this is precisely the comparison that
|
|
# was never made for the GPU agent: its lease poll backed off to 900s while
|
|
# the roster called it stopped at 300s, and both numbers were individually
|
|
# correct, in different directions, in different files (lesson #4355).
|
|
#
|
|
# Two clear sweeps before a part is even called STALE. One missed tick is
|
|
# routine — the sweep rides the maintenance queue and does an inspect that can
|
|
# take eleven seconds — and must not turn the page yellow.
|
|
_SWEEPS_BEFORE_STALE = 2
|
|
assert STALE_AFTER_SECONDS >= SWEEP_PERIOD_SECONDS * _SWEEPS_BEFORE_STALE, (
|
|
f"a {SWEEP_PERIOD_SECONDS}s sweep cannot keep a roster fresh against a "
|
|
f"{STALE_AFTER_SECONDS}s stale threshold: raise the threshold or shorten "
|
|
f"the sweep"
|
|
)
|
|
|
|
# Probes cross a process boundary, so they carry deadlines. A hung Postgres
|
|
# must make this endpoint say "postgres: down", not hang alongside it.
|
|
PROBE_TIMEOUT_SECONDS = 2.0
|
|
|
|
_OK, _STALE, _DOWN, _UNKNOWN = "ok", "stale", "down", "unknown"
|
|
|
|
# Worst-first, so an overall verdict is just the max.
|
|
_SEVERITY = {_OK: 0, _UNKNOWN: 1, _STALE: 2, _DOWN: 3}
|
|
|
|
|
|
def _age_state(age_seconds: float) -> str:
|
|
if age_seconds >= DOWN_AFTER_SECONDS:
|
|
return _DOWN
|
|
if age_seconds >= STALE_AFTER_SECONDS:
|
|
return _STALE
|
|
return _OK
|
|
|
|
|
|
def _describe_learned(name: str, state: str, age: float, details: dict) -> str:
|
|
"""Say what the state MEANS. A red chip tells an operator less than a
|
|
sentence does at the moment they are deciding whether to go and look."""
|
|
if state == _OK:
|
|
replicas = details.get("replicas")
|
|
if replicas and replicas > 1:
|
|
return f"{name} is running ({replicas} replicas)"
|
|
return f"{name} is running"
|
|
mins = int(age // 60)
|
|
ago = f"{mins} min" if mins else f"{int(age)}s"
|
|
if state == _STALE:
|
|
return f"{name} has not checked in for {ago}"
|
|
return f"{name} has not checked in for {ago} — treat it as stopped"
|
|
|
|
|
|
async def _probe_postgres(session) -> dict:
|
|
started = time.monotonic()
|
|
try:
|
|
await asyncio.wait_for(
|
|
session.execute(text("SELECT 1")), timeout=PROBE_TIMEOUT_SECONDS
|
|
)
|
|
except Exception as exc: # noqa: BLE001 — a probe reports, it never raises
|
|
return {
|
|
"key": "postgres", "kind": "datastore", "name": "PostgreSQL",
|
|
"state": _DOWN, "detail": f"not answering: {type(exc).__name__}",
|
|
}
|
|
return {
|
|
"key": "postgres", "kind": "datastore", "name": "PostgreSQL", "state": _OK,
|
|
"detail": "answering", "latency_ms": round((time.monotonic() - started) * 1000, 1),
|
|
}
|
|
|
|
|
|
def _ping_redis_sync() -> None:
|
|
import redis # local import; mirrors system_activity's pattern
|
|
|
|
client = redis.Redis.from_url(
|
|
get_config().celery_broker_url,
|
|
socket_connect_timeout=PROBE_TIMEOUT_SECONDS,
|
|
socket_timeout=PROBE_TIMEOUT_SECONDS,
|
|
)
|
|
client.ping()
|
|
|
|
|
|
async def _probe_redis() -> dict:
|
|
started = time.monotonic()
|
|
try:
|
|
await asyncio.wait_for(
|
|
asyncio.to_thread(_ping_redis_sync), timeout=PROBE_TIMEOUT_SECONDS * 2
|
|
)
|
|
except Exception as exc: # noqa: BLE001
|
|
return {
|
|
"key": "redis", "kind": "datastore", "name": "Redis",
|
|
"state": _DOWN,
|
|
"detail": f"not answering: {type(exc).__name__} — queues and workers "
|
|
f"cannot be reached either",
|
|
}
|
|
return {
|
|
"key": "redis", "kind": "datastore", "name": "Redis", "state": _OK,
|
|
"detail": "answering", "latency_ms": round((time.monotonic() - started) * 1000, 1),
|
|
}
|
|
|
|
|
|
@system_health_bp.route("/health", methods=["GET"])
|
|
async def system_health():
|
|
"""Every part, its state, and one overall verdict.
|
|
|
|
Response: {overall, parts: [{key, kind, name, state, detail, last_seen_at,
|
|
…}], checked_at}
|
|
"""
|
|
parts: list[dict] = []
|
|
now = datetime.now(UTC)
|
|
|
|
async with get_session() as session:
|
|
# Postgres first, and if it is unreachable nothing else can be read —
|
|
# say so rather than failing, because "the database is down" is the
|
|
# single most useful thing this endpoint can ever report.
|
|
pg = await _probe_postgres(session)
|
|
parts.append(pg)
|
|
|
|
if pg["state"] == _OK:
|
|
# A PURE READ since 2026-09-23. This used to refresh the celery
|
|
# roster here, rate-limited to once per 20s — so the roster only
|
|
# advanced while somebody had a browser open, and a broadcast rode
|
|
# on a request. `size_worker_lanes` writes it now, on a timer, and
|
|
# the assertion below is what keeps that cadence honest.
|
|
rows = (
|
|
await session.execute(select(ServiceSeen).order_by(ServiceSeen.display_name))
|
|
).scalars().all()
|
|
for row in rows:
|
|
age = (now - row.last_seen_at).total_seconds()
|
|
state = _age_state(age)
|
|
parts.append({
|
|
"key": row.key,
|
|
"kind": row.kind,
|
|
"name": row.display_name,
|
|
"state": state,
|
|
"detail": _describe_learned(row.display_name, state, age, row.details or {}),
|
|
"last_seen_at": row.last_seen_at.isoformat(),
|
|
"first_seen_at": row.first_seen_at.isoformat(),
|
|
**{k: v for k, v in (row.details or {}).items() if k != "agent_id"},
|
|
})
|
|
|
|
parts.append(await _probe_redis())
|
|
|
|
overall = max((p["state"] for p in parts), key=lambda s: _SEVERITY[s], default=_UNKNOWN)
|
|
return jsonify({
|
|
"overall": overall,
|
|
"parts": sorted(parts, key=lambda p: (-_SEVERITY[p["state"]], p["name"])),
|
|
"checked_at": now.isoformat(),
|
|
# So the UI can explain a `stale` without hard-coding the same numbers
|
|
# in a second place.
|
|
"thresholds": {
|
|
"stale_after_seconds": STALE_AFTER_SECONDS,
|
|
"down_after_seconds": DOWN_AFTER_SECONDS,
|
|
},
|
|
})
|