CI / lint (push) Successful in 2s
CI / extension-version (push) Successful in 2s
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 6s
CI / frontend-build (push) Successful in 25s
CI / backend-lint-and-test (push) Successful in 31s
Build images / build-web (push) Successful in 55s
Build images / smoke-web (push) Skipped
Build images / build-ml (push) Successful in 1m41s
Build images / promote (push) Skipped
CI / integration (push) Successful in 2m8s
Milestone 422 step 2. `GET /api/system/workers` reports every lane joined to its live pool; `POST /api/system/workers/<name>` changes it. NO DOCKER SOCKET. Milestone 365 deferred "acting on the state" because restarting a dead worker needs a socket the web container deliberately does not have. That holds for restarting a CONTAINER; it does not hold for changing how much work a RUNNING worker does. celery's pool_grow / pool_shrink / add_consumer / cancel_consumer send a message over the Redis the app already uses, and the worker resizes itself. No new privilege, no new surface, and the security question that deferred this is never raised. PERSIST AND PUSH, in one call, in that order. pool_grow is not durable — a restart drops every lane to its env concurrency — so a UI that only pushed would lose the setting on the next deploy with nothing to show for it (lesson #4202). Storing alone would describe nothing until something restarted. A failed PUSH is not a failed setting: 200 with `applied: false` and a reason, so the UI says "saved, not yet live" rather than "that didn't work". Step 3's reconcile carries it when the lane answers again. PER-REPLICA DELTAS. `pool_grow(n, destination=[...])` adds n to EACH destination, so while `worker` runs `replicas: 2` a single delta from an aggregate is wrong for both. `slots` therefore means what CELERY_CONCURRENCY means — one process's pool — and each replica is driven to it from its OWN current size, so replicas that drifted apart converge rather than moving in lockstep. I wrote this wrong first: the docstring claimed per-replica while the code computed one delta from the max across replicas. LaneLiveState now carries `pools` per hostname and exposes `pool` as a property. A replica already at the target is sent nothing at all — the reachable fixed point step 3's periodic reconcile needs, or it re-issues a grow of zero every tick forever (lesson #4183). A replica that answered inspect but not stats is NAMED in the error rather than skipped silently, since otherwise it would run at a size the UI claims it does not. `present=False` is not "zero slots", it is "nothing answered" — kept distinct throughout, because step 3 skips an absent lane rather than correcting it. /workers now also reports pool size (from `insp.stats()`) and RESERVED count. Celery prefetches, so tasks that have left the Redis list but not started are invisible to LLEN: a lane can read depth 0 with thirty tasks held in worker memory. `pending` is depth + reserved. The UI is misleading without this and step 7's autoscaler would be simply wrong. Also kills the THIRD copy of the queue list: system_activity's _QUEUE_NAMES, whose own comment admitted the coupling ("must match celery_app.task_routes") and which sat alongside task_routes and the ROLE_NAMES copy step 1 collapsed. Now derived from LANES. The rendered order changes to lane grouping, which is the better shape for a lane-oriented UI. Separate blueprint rather than folding into system_activity, which states in its first line that it is read-only and answers a different question — its /workers is keyed on celery HOSTNAME and reports which nodes answered. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
104 lines
4.0 KiB
Python
104 lines
4.0 KiB
Python
"""Worker lanes: what each is doing, and the dial that changes it.
|
|
|
|
Milestone 422 step 2. The write half of a surface `api/system_activity.py`
|
|
only reads.
|
|
|
|
## Why this is a separate blueprint
|
|
|
|
`system_activity` says in its own first line that it is read-only, and it
|
|
answers a different question: its `/workers` is keyed on celery HOSTNAME and
|
|
reports which nodes answered. That stays as it is — the existing
|
|
SystemActivityTab consumes it.
|
|
|
|
This is keyed on LANE, joins the stored settings to the live pool, and
|
|
accepts writes. Two endpoints answering "which celery processes exist" and
|
|
"how much work is each lane allowed to do" are not the same endpoint, and
|
|
folding the second into the first would make a read-only module a write one.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
from datetime import UTC, datetime
|
|
|
|
from quart import Blueprint, jsonify, request
|
|
|
|
from ..extensions import get_session
|
|
from ..services.worker_control import LaneUpdateRefused, lane_view, set_lane
|
|
from ..services.worker_lanes import LANES_BY_NAME
|
|
from ._responses import error_response as _bad
|
|
|
|
workers_bp = Blueprint("workers", __name__, url_prefix="/api/system/workers")
|
|
|
|
|
|
@workers_bp.route("", methods=["GET"])
|
|
async def list_lanes():
|
|
"""Every lane: configured slots, the cap, the ceiling, and live state.
|
|
|
|
Response: {lanes: [...], fetched_at: iso8601}
|
|
|
|
Deliberately NOT cached, unlike system_activity's 2s/5s caches. This is
|
|
the surface an operator watches while dragging a stepper, and a cached
|
|
reply would show them the value from before their own change and read as
|
|
the control having failed.
|
|
"""
|
|
async with get_session() as session:
|
|
lanes = await lane_view(session)
|
|
return jsonify({
|
|
"lanes": lanes,
|
|
"fetched_at": datetime.now(UTC).isoformat(),
|
|
})
|
|
|
|
|
|
@workers_bp.route("/<name>", methods=["POST"])
|
|
async def update_lane(name: str):
|
|
"""Set a lane's slots, cap and/or enabled flag. Stores, then pushes live.
|
|
|
|
Partial: only the keys present are changed, so the UI's stepper can send
|
|
`{"slots": 3}` without restating the cap it did not touch.
|
|
|
|
Two failure kinds, deliberately different statuses:
|
|
|
|
* **400** — the value is not allowed (above the cap, above the ceiling,
|
|
negative). Nothing was stored. The body carries `detail`, which is the
|
|
sentence the UI shows; a refused control with no reason reads as a bug.
|
|
* **200 with `applied: false`** — the value WAS stored but could not be
|
|
pushed, because the lane is not currently answering. That is not an
|
|
error: step 3's reconcile carries it when the lane comes back, and the
|
|
UI should say "saved, not yet live" rather than "that didn't work".
|
|
"""
|
|
lane = LANES_BY_NAME.get(name)
|
|
if lane is None:
|
|
return _bad("unknown_lane", detail=name, known=sorted(LANES_BY_NAME))
|
|
|
|
body = await request.get_json()
|
|
if not isinstance(body, dict):
|
|
return _bad("invalid_body", detail="body must be a JSON object")
|
|
|
|
fields: dict = {}
|
|
for key in ("slots", "slots_cap"):
|
|
if key in body:
|
|
value = body[key]
|
|
# Rejected rather than coerced: `True` is an int in Python, and
|
|
# silently reading it as 1 slot would be a control that appears to
|
|
# work and sets something nobody asked for.
|
|
if not isinstance(value, int) or isinstance(value, bool):
|
|
return _bad("invalid_body", detail=f"{key} must be an integer")
|
|
fields[key] = value
|
|
if "enabled" in body:
|
|
if not isinstance(body["enabled"], bool):
|
|
return _bad("invalid_body", detail="enabled must be a boolean")
|
|
fields["enabled"] = body["enabled"]
|
|
|
|
if not fields:
|
|
return _bad(
|
|
"invalid_body",
|
|
detail="give at least one of slots, slots_cap, enabled",
|
|
)
|
|
|
|
async with get_session() as session:
|
|
try:
|
|
result = await set_lane(session, lane, **fields)
|
|
except LaneUpdateRefused as exc:
|
|
return _bad("refused", detail=str(exc))
|
|
return jsonify(result)
|