feat: the System tab reads a stored sample instead of inspecting per load (4295)
CI and images / lint (push) Failing after 3s
CI and images / extension-version (push) Successful in 4s
CI and images / frontend-build (push) Successful in 31s
CI and images / backend-lint-and-test (push) Successful in 35s
CI and images / integration (push) Successful in 2m44s
CI and images / sign-extension (push) Skipped
CI and images / build-web (push) Skipped
CI and images / smoke-web (push) Skipped
CI and images / promote (push) Skipped
CI and images / build-agent (push) Skipped
CI and images / lint (push) Failing after 3s
CI and images / extension-version (push) Successful in 4s
CI and images / frontend-build (push) Successful in 31s
CI and images / backend-lint-and-test (push) Successful in 35s
CI and images / integration (push) Successful in 2m44s
CI and images / sign-extension (push) Skipped
CI and images / build-web (push) Skipped
CI and images / smoke-web (push) Skipped
CI and images / promote (push) Skipped
CI and images / build-agent (push) Skipped
Operator: "there is a repull every time this page loads is there a reason
this info isn't being tracked in the background and stored in some way?"
There was a reason and it had expired, and underneath it there was plain
waste.
The expired one: /api/system/workers was deliberately uncached because an
operator dragging the stepper must not be shown a pre-change value. That
stopped being true at 1353d34, when the UI began patching its row from the
write's reply instead of refetching.
The waste: size_worker_lanes already inspected the broker on a timer to
decide pool sizes — computing the pool, active, reserved and queue depth
the page shows, using them, and discarding them. The browser then asked
the broker for the same numbers four times a minute, per open tab.
So one inspect now feeds three things: the sizing decision, a stored
sample (worker_lane_sample, alembic 0107), and the celery roster. No
request path touches the broker at all — the roster refresh comes off
/api/system/health too, where it had been rate-limited to 20s and so made
worker liveness a function of whether anyone had a browser open.
Consequences, stated rather than hidden:
- The live figures are up to one sweep old. measured_at travels with each
lane and the page says how old, because a stale number presented as
current is how someone watches a queue "not move" that is moving.
- The sweep is the roster's only writer now, so its period and the
staleness thresholds are in a relationship. 60s against a 90s stale
threshold left one missed tick between normal and all-yellow — the
shape of lesson #4355 — so the period is 30s, named once in
worker_lanes, and system_health asserts its headroom at import with a
test stating the same thing in prose.
- An idle lane therefore also gives a worker back twice as fast. That is
the direction asked for: "idle instances quiet down when not running".
Also bounds the inspect in push_lane_cap, which was an await with no
deadline (rule 156) — harmless while it ran on a request, less so now
that it runs in a background task where a hang would be silent.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
This commit is contained in:
@@ -14,6 +14,7 @@ Queues:
|
||||
from celery import Celery
|
||||
|
||||
from .config import get_config
|
||||
from .services.worker_lanes import SWEEP_PERIOD_SECONDS
|
||||
|
||||
|
||||
def make_celery() -> Celery:
|
||||
@@ -113,7 +114,14 @@ def make_celery() -> Celery:
|
||||
},
|
||||
"size-worker-lanes": {
|
||||
"task": "backend.app.tasks.maintenance.size_worker_lanes",
|
||||
"schedule": 60.0, # every minute.
|
||||
"schedule": SWEEP_PERIOD_SECONDS,
|
||||
#
|
||||
# The number lives in `services/worker_lanes` because three
|
||||
# places must agree on it: this schedule, the freshness of the
|
||||
# sample the System tab reads, and the roster staleness
|
||||
# thresholds in `api/system_health` — which now depend on this
|
||||
# sweep rather than on a browser being open, and assert their
|
||||
# headroom over it at import.
|
||||
#
|
||||
# ONE entry, replacing `autoscale-worker-lanes` (60s) and
|
||||
# `reconcile-worker-lanes` (300s) on 2026-09-23. They were two
|
||||
@@ -121,14 +129,16 @@ def make_celery() -> Celery:
|
||||
# existed to stop the reconcile undoing its work; with the
|
||||
# stored `slots` gone there is nothing to disagree about.
|
||||
#
|
||||
# A minute because it reacts to a BACKLOG, and a five-minute
|
||||
# reaction to a queue filling up is no reaction. It also now
|
||||
# carries what the reconcile was for — a worker restarted at
|
||||
# its ENV concurrency is corrected on the next tick rather
|
||||
# than after five.
|
||||
# Fast enough to react to a BACKLOG — a five-minute reaction to
|
||||
# a queue filling up is no reaction. It also carries what the
|
||||
# reconcile was for: a worker restarted at its ENV concurrency
|
||||
# is corrected on the next tick rather than after five.
|
||||
#
|
||||
# Cheap when settled: one inspect plus one LLEN sweep, and no
|
||||
# control messages at all once every lane matches.
|
||||
# control messages at all once every lane matches. It is also
|
||||
# now the ONLY thing that inspects — nothing on a request path
|
||||
# does — so this is the whole broker cost of the System tab,
|
||||
# whether nobody or ten tabs are watching.
|
||||
},
|
||||
"cleanup-old-tasks": {
|
||||
"task": "backend.app.tasks.maintenance.cleanup_old_tasks",
|
||||
|
||||
Reference in New Issue
Block a user