feat: the System tab reads a stored sample instead of inspecting per load (4295)
CI and images / lint (push) Failing after 3s
CI and images / extension-version (push) Successful in 4s
CI and images / frontend-build (push) Successful in 31s
CI and images / backend-lint-and-test (push) Successful in 35s
CI and images / integration (push) Successful in 2m44s
CI and images / sign-extension (push) Skipped
CI and images / build-web (push) Skipped
CI and images / smoke-web (push) Skipped
CI and images / promote (push) Skipped
CI and images / build-agent (push) Skipped

Operator: "there is a repull every time this page loads is there a reason
this info isn't being tracked in the background and stored in some way?"

There was a reason and it had expired, and underneath it there was plain
waste.

The expired one: /api/system/workers was deliberately uncached because an
operator dragging the stepper must not be shown a pre-change value. That
stopped being true at 1353d34, when the UI began patching its row from the
write's reply instead of refetching.

The waste: size_worker_lanes already inspected the broker on a timer to
decide pool sizes — computing the pool, active, reserved and queue depth
the page shows, using them, and discarding them. The browser then asked
the broker for the same numbers four times a minute, per open tab.

So one inspect now feeds three things: the sizing decision, a stored
sample (worker_lane_sample, alembic 0107), and the celery roster. No
request path touches the broker at all — the roster refresh comes off
/api/system/health too, where it had been rate-limited to 20s and so made
worker liveness a function of whether anyone had a browser open.

Consequences, stated rather than hidden:

- The live figures are up to one sweep old. measured_at travels with each
  lane and the page says how old, because a stale number presented as
  current is how someone watches a queue "not move" that is moving.
- The sweep is the roster's only writer now, so its period and the
  staleness thresholds are in a relationship. 60s against a 90s stale
  threshold left one missed tick between normal and all-yellow — the
  shape of lesson #4355 — so the period is 30s, named once in
  worker_lanes, and system_health asserts its headroom at import with a
  test stating the same thing in prose.
- An idle lane therefore also gives a worker back twice as fast. That is
  the direction asked for: "idle instances quiet down when not running".

Also bounds the inspect in push_lane_cap, which was an await with no
deadline (rule 156) — harmless while it ran on a request, less so now
that it runs in a background task where a hang would be silent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
This commit is contained in:
2026-09-23 18:52:08 -04:00
co-authored by Claude Opus 5
parent 7f1693a40d
commit 45bb7044f7
17 changed files with 871 additions and 133 deletions
+116 -1
View File
@@ -13,7 +13,8 @@ from sqlalchemy import select
from backend.app.api import workers as workers_api
from backend.app.models import WorkerLane
from backend.app.services import worker_control as wc
from backend.app.services.worker_lanes import LANES
from backend.app.services.worker_lanes import LANES, SWEEP_PERIOD_SECONDS
from tests.doubles import RecordingSession
pytestmark = pytest.mark.integration
@@ -319,6 +320,120 @@ async def _refreshed(db, name: str, expected: int) -> None:
assert row.slots_cap == expected, "a refused write must store nothing"
# --- the page does not ask the broker anything -------------------------------
@pytest.mark.asyncio
async def test_the_lane_read_makes_no_broker_call_at_all(client, monkeypatch):
"""Operator, 2026-09-23: *"there is a repull every time this page loads —
is there a reason this info isn't being tracked in the background and
stored in some way?"*
There was: this endpoint inspected the broker on every call — four
broadcasts on an eleven-second budget — while `size_worker_lanes` was
already inspecting on a timer and throwing the same numbers away. The
sweep stores them now and this reads the table.
Asserted by making the inspect RAISE, because a version that inspected and
was merely quick about it would pass a call-count test on a fast CI box.
"""
monkeypatch.setattr(wc, "inspect_lanes_sync", _never_called)
monkeypatch.setattr(wc, "_queue_depths_sync", _never_called)
resp = await client.get("/api/system/workers")
assert resp.status_code == 200
assert len((await resp.get_json())["lanes"]) == len(LANES)
@pytest.mark.asyncio
async def test_a_lane_with_no_sample_yet_reads_as_unmeasured(client, monkeypatch):
"""A fresh install inside its first sweep period. `measured_at` is null and
`present` is false — and those are DIFFERENT facts: nothing has asked yet,
versus something asked and nothing answered. The UI says different things
about them, so the payload must keep them apart."""
monkeypatch.setattr(wc, "inspect_lanes_sync", _never_called)
body = await (await client.get("/api/system/workers")).get_json()
for lane in body["lanes"]:
assert lane["measured_at"] is None, lane["name"]
assert lane["live"]["present"] is False, lane["name"]
@pytest.mark.asyncio
async def test_the_stored_sample_is_what_the_page_shows(client, db, monkeypatch):
"""The whole point of the table: the sweep writes, the endpoint reads."""
monkeypatch.setattr(wc, "inspect_lanes_sync", _never_called)
live = {lane.name: wc.LaneLiveState() for lane in LANES}
live["worker"] = wc.LaneLiveState(
present=True, replicas=1, active=2, reserved=3, pools={"worker@a": 4},
)
await _store_sample(db, live, {"default": 7, "import": 0,
"thumbnail": 0, "download": 0})
body = await (await client.get("/api/system/workers")).get_json()
worker = next(l for l in body["lanes"] if l["name"] == "worker")
assert worker["live"] == {
"present": True, "replicas": 1, "pool": 4, "active": 2, "reserved": 3,
}
assert worker["queue_depth"] == 7
# depth PLUS reserved — celery prefetches, so LLEN alone under-reports.
assert worker["pending"] == 10
assert worker["measured_at"] is not None
@pytest.mark.asyncio
async def test_a_lane_that_stopped_answering_overwrites_its_old_reading(
client, db, monkeypatch,
):
"""The sweep writes EVERY lane, including the ones that did not answer.
Skipping them would leave the previous sample in place, and the page would
go on showing a pool that is no longer there — a stale row read as a
current one (lesson #4202: the row is the thing that has to change).
"""
monkeypatch.setattr(wc, "inspect_lanes_sync", _never_called)
up = {lane.name: wc.LaneLiveState() for lane in LANES}
up["worker"] = wc.LaneLiveState(present=True, replicas=1, pools={"worker@a": 4})
await _store_sample(db, up, {})
await _store_sample(db, {lane.name: wc.LaneLiveState() for lane in LANES}, {})
body = await (await client.get("/api/system/workers")).get_json()
worker = next(l for l in body["lanes"] if l["name"] == "worker")
assert worker["live"]["present"] is False
assert worker["live"]["pool"] is None
@pytest.mark.asyncio
async def test_the_payload_says_how_often_it_is_measured(client, monkeypatch):
"""So the UI can explain the age of the numbers without keeping its own
copy of the cadence, which would be free to drift from the schedule."""
monkeypatch.setattr(wc, "inspect_lanes_sync", _never_called)
body = await (await client.get("/api/system/workers")).get_json()
assert body["sweep_period_seconds"] == SWEEP_PERIOD_SECONDS
def _never_called(*args, **kwargs):
raise AssertionError("the lane read talked to the broker")
async def _store_sample(db, live, depths) -> None:
"""Write a sweep's worth of samples through the real storer.
The production path is sync (a celery task owns a sync session); this
suite is async, so the statements are replayed on the async session rather
than reimplemented — the thing under test must be the shipped writer.
"""
rec = RecordingSession()
wc.store_lane_samples_sync(rec, live, depths)
for stmt in rec.stmts:
await db.execute(stmt)
await db.commit()
# --- no database connection is held across a broker round trip ---------------