feat: the System tab reads a stored sample instead of inspecting per load (4295)
CI and images / lint (push) Failing after 3s
CI and images / extension-version (push) Successful in 4s
CI and images / frontend-build (push) Successful in 31s
CI and images / backend-lint-and-test (push) Successful in 35s
CI and images / integration (push) Successful in 2m44s
CI and images / sign-extension (push) Skipped
CI and images / build-web (push) Skipped
CI and images / smoke-web (push) Skipped
CI and images / promote (push) Skipped
CI and images / build-agent (push) Skipped

Operator: "there is a repull every time this page loads is there a reason
this info isn't being tracked in the background and stored in some way?"

There was a reason and it had expired, and underneath it there was plain
waste.

The expired one: /api/system/workers was deliberately uncached because an
operator dragging the stepper must not be shown a pre-change value. That
stopped being true at 1353d34, when the UI began patching its row from the
write's reply instead of refetching.

The waste: size_worker_lanes already inspected the broker on a timer to
decide pool sizes — computing the pool, active, reserved and queue depth
the page shows, using them, and discarding them. The browser then asked
the broker for the same numbers four times a minute, per open tab.

So one inspect now feeds three things: the sizing decision, a stored
sample (worker_lane_sample, alembic 0107), and the celery roster. No
request path touches the broker at all — the roster refresh comes off
/api/system/health too, where it had been rate-limited to 20s and so made
worker liveness a function of whether anyone had a browser open.

Consequences, stated rather than hidden:

- The live figures are up to one sweep old. measured_at travels with each
  lane and the page says how old, because a stale number presented as
  current is how someone watches a queue "not move" that is moving.
- The sweep is the roster's only writer now, so its period and the
  staleness thresholds are in a relationship. 60s against a 90s stale
  threshold left one missed tick between normal and all-yellow — the
  shape of lesson #4355 — so the period is 30s, named once in
  worker_lanes, and system_health asserts its headroom at import with a
  test stating the same thing in prose.
- An idle lane therefore also gives a worker back twice as fast. That is
  the direction asked for: "idle instances quiet down when not running".

Also bounds the inspect in push_lane_cap, which was an await with no
deadline (rule 156) — harmless while it ran on a request, less so now
that it runs in a background task where a hang would be silent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
This commit is contained in:
2026-09-23 18:52:08 -04:00
co-authored by Claude Opus 5
parent 7f1693a40d
commit 45bb7044f7
17 changed files with 871 additions and 133 deletions
+62
View File
@@ -8,6 +8,7 @@ work drifting apart, which is invisible in each one read on its own.
from __future__ import annotations
from backend.app.services import service_roster as sr
from tests.doubles import RecordingSession
def test_the_wrapper_budget_exceeds_the_work_it_waits_for():
@@ -129,3 +130,64 @@ def test_a_worker_this_build_did_not_name_is_still_grouped_by_its_queues(
_stub_inspect(monkeypatch, {"celery@xyz": [{"name": "odd"}]})
assert list(sr._inspect_celery_sync()) == [("odd",)]
# --- the sweep writes the roster, not the page -------------------------------
#
# It used to be refreshed on the /api/system/health request path, rate-limited
# to once per 20s. So the roster only advanced while someone had a browser
# open: the liveness of the workers was a function of whether anyone was
# looking at them. Operator, 2026-09-23: *"is there a reason this info isn't
# being tracked in the background and stored in some way?"*
def test_the_sync_refresh_writes_a_row_for_everything_that_answered(monkeypatch):
monkeypatch.setattr(
sr, "_inspect_celery_sync",
lambda: {("ml",): {"hostnames": ["ml@a"], "active": 2},
("scan",): {"hostnames": ["scheduler@a"], "active": 0}},
)
session = RecordingSession()
sr.refresh_celery_roster_sync(session)
assert len(session.stmts) == 2
assert session.commits == 1
def test_a_broker_that_will_not_answer_does_not_kill_the_sweep(monkeypatch):
"""The sizing pass runs on a timer and does three things; a roster refresh
that raised would take the other two with it. Rows going stale IS the
correct report about a broker nobody can reach."""
def boom():
raise RuntimeError("no broker")
monkeypatch.setattr(sr, "_inspect_celery_sync", boom)
session = RecordingSession()
sr.refresh_celery_roster_sync(session)
assert session.stmts == []
assert session.commits == 0
def test_both_refreshes_build_the_same_row(monkeypatch):
"""The async path (an agent lease over the API) and the sync one (the
sweep) must not drift. Asserted on the shared mapping rather than by
running both, because the thing that could drift is what a roster row IS —
not which kind of session writes it."""
grouped = {("ml",): {"hostnames": ["ml@a", "ml@b"], "active": 3}}
rows = sr._roster_rows(grouped)
assert rows == [{
"key": "celery:ml",
"kind": "celery",
"display_name": sr.role_display_name(("ml",)),
"details": {
"queues": ["ml"],
"hostnames": ["ml@a", "ml@b"],
"replicas": 2,
"active": 3,
},
}]