CI and images / lint (push) Successful in 4s
CI and images / extension-version (push) Successful in 4s
CI and images / frontend-build (push) Successful in 24s
CI and images / integration (push) Failing after 24s
CI and images / backend-lint-and-test (push) Failing after 34s
CI and images / sign-extension (push) Skipped
CI and images / build-web (push) Skipped
CI and images / smoke-web (push) Skipped
CI and images / promote (push) Skipped
CI and images / build-agent (push) Skipped
Operator, 2026-09-23: *"auto should be always on, not a setting, so that idle
instances quiet down when not running. the number that is visible and
something the user can tweak and manage should be the cap itself the number of
running workers is handled by the autoscaling function which is always on."*
They are right, and the reason it was not built this way is worth stating: the
manual dial came first (steps 2-4) and the autoscaler came last (step 7), as
an opt-in BESIDE a control that already existed. Nothing ever asked whether
the dial should still exist once something could move it automatically. Each
step was defensible; the result was three operator settings over one number.
## `slots`, `enabled` and `autoscale` are gone
`slots` was a MEASUREMENT wearing a preference's clothes. How many workers a
lane runs is read live and moved every minute; storing it meant the operator
had to keep two numbers in agreement and the autoscaler had to be told it was
allowed to touch one of them.
`autoscale` gated the mechanism behind a choice, so a lane nobody opted in
never gave its workers back — which is why an idle instance never quieted
down.
`enabled` is derived: a cap of zero means no consumers. "Off" and "may use no
workers" were two spellings of one fact, stored separately, free to disagree.
## Two sweeps become one
`reconcile_lanes_sync` drove the pool to the stored `slots`; `autoscale_lanes_
sync` moved it away from that same number; and most of step 7's hardest
reasoning — a stored value that is a FLOOR, a target of `max(stored, current)`
— existed only to stop them fighting. Delete the stored number and the problem
is not solved, it is absent.
`size_lanes_sync` runs every minute and owns both consumers and pool size. It
also subsumes what the reconcile was for: a worker restarted at its ENV
concurrency is corrected on the next tick rather than after five.
Growth is immediate, shrink is one worker per tick. Deliberately asymmetric —
"always on" is only pleasant if the ramp keeps up, and +1/minute would take
four minutes to answer a burst. Being one worker too large for a minute costs
a sleeping process; being too small costs work not happening. For ML the
asymmetry matters most: every new slot reloads a multi-GB model, so the slow
shrink is what stops a quiet patch from paying that cost again a minute later.
## The caps ship at one, and zero for ML
Per the operator. Conservative on purpose — and a conservative default nobody
knows how to raise is just a slow product, which is the other half of what
they asked for:
"there needs to be something that tells the user to bump those numbers to
improve processing rate or they'd never know the controls exist."
So a lane running everything its cap allows while work piles up says so, in
its own row, with the headroom named: *"4,060 waiting and all 1 worker busy.
Raise the cap to run more at once — this machine allows up to 7."*
It fires only when raising the cap would actually help. Not when the lane is
keeping up, not when the sizing pass has room it has not taken, and not at the
machine ceiling — where "raise the cap" is advice nobody can take.
## Migration 0105 rewrites the caps rather than carrying them
The old defaults (4/2/2/1) bounded a manual control and were loose because
moving within them was the ordinary act. The number now means "the most
workers this lane may use", which is a different promise; carrying the old
figure over would quadruple the worker lane on every existing install at the
moment this deploys.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
367 lines
15 KiB
Python
367 lines
15 KiB
Python
"""The worker lanes: what they are, and how many slots each may be given.
|
|
|
|
Milestone 422 step 1. This module is the ONE place that knows the lane set;
|
|
`models/worker_lane.py` holds only what the operator can change about them.
|
|
|
|
## Why the queues are here and not in the table
|
|
|
|
A lane's queue set is not a preference — it is decided by `celery_app.py`'s
|
|
`task_routes`, which is what puts a backup on `maintenance_long` and a
|
|
thumbnail on `thumbnail`. An operator cannot move a task to another lane, so
|
|
storing the queues as settings would create a row that can disagree with the
|
|
routing table, and nothing would notice until a queue had no consumer.
|
|
|
|
So: queues and display names are code, slots and caps are data. The table
|
|
stores three numbers and a flag, and nothing that could contradict celery.
|
|
|
|
This also collapses a duplicate rather than adding one.
|
|
`service_roster.ROLE_NAMES` was a second copy of "queue set -> the name an
|
|
operator recognises", and it had already drifted: `maintenance_long` is a
|
|
live lane with four task routes pointing at it, and the roster did not know
|
|
its name, so the System tab rendered it as `Worker (maintenance_long)`. That
|
|
map is now derived from `LANES` below, so a lane added here is named
|
|
everywhere at once.
|
|
|
|
## Why the ceiling is derived rather than configured
|
|
|
|
Consolidating the stack into one container (step 5) widens the OOM blast
|
|
radius: today an ml-worker that exhausts memory is killed by Docker on its
|
|
own, and web keeps serving. In one container the kernel picks a victim from
|
|
the whole cgroup, and it may pick hypercorn — so a tagging task can take the
|
|
UI down with it, on exactly the modest hardware least able to spare the
|
|
memory.
|
|
|
|
Operator, 2026-09-22: *"ram isn't an issue for me but some users might run
|
|
this on weaker hardware and I don't want it to kill their servers."*
|
|
|
|
So the maximum is computed from what the container actually has, and the
|
|
operator's own `slots_cap` must fit under it. Three numbers, not two, and the
|
|
ordering is the point:
|
|
|
|
slots <= slots_cap <= derived_ceiling
|
|
(live) (operator) (this module)
|
|
|
|
The operator can always lower their cap. They cannot raise it past what the
|
|
box can hold. The derived ceiling is never stored — a row that outlived a
|
|
change in container limits must not carry a stale one.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import logging
|
|
import os
|
|
from dataclasses import dataclass
|
|
from pathlib import Path
|
|
|
|
log = logging.getLogger(__name__)
|
|
|
|
# --- the lanes ---------------------------------------------------------------
|
|
|
|
|
|
GIB = 1024 ** 3
|
|
|
|
|
|
@dataclass(frozen=True)
|
|
class ModelRequirement:
|
|
"""A model a lane must download before it can do anything.
|
|
|
|
Surfaced to the UI so the operator is told WHICH model, how big, and what
|
|
it costs to hold — before they turn the lane on, not after a multi-GB
|
|
download has already started. The lane is optional and its cost is not
|
|
obvious from its name, which is the whole reason this is structured data
|
|
rather than a sentence in a component.
|
|
|
|
`measured=False` means the numbers are ESTIMATES and the UI must say so.
|
|
They come from the checkpoint's parameter count and dtype, not from a
|
|
build — and a number presented as fact decides whether someone's server
|
|
survives, so it is labelled rather than rounded confidently.
|
|
"""
|
|
|
|
# The Hugging Face repo id, which is the honest answer to "which model".
|
|
repo: str
|
|
# Roughly what the download costs, for the operator's bandwidth and disk.
|
|
approx_download_bytes: int
|
|
# Roughly what ONE slot holds while running. Prefork forks a child per
|
|
# slot and each loads its own copy, so this multiplies.
|
|
approx_resident_bytes: int
|
|
measured: bool = False
|
|
|
|
|
|
# SigLIP so400m — the only model FabledCurator itself downloads.
|
|
#
|
|
# What it is for, which is NOT obvious from the lane's name: it produces the
|
|
# image embeddings that back similarity search, duplicate grouping and the
|
|
# tag heads. WD14 tagging is the GPU AGENT's job, not this lane's — the
|
|
# comment in celery_app.py naming both is stale since B3 (#1238), when the
|
|
# agent took over and this lane was left as the CPU embed fallback for stacks
|
|
# running no agent at all (see MLSettings.cpu_embed_enabled).
|
|
#
|
|
# Both numbers are ESTIMATES, derived from the checkpoint rather than from a
|
|
# build: ~877M parameters at fp32 is ~3.5GB of weights, and holding them plus
|
|
# activations and the torch runtime is what the resident figure covers. They
|
|
# err high. Replace them with measurements — download the repo and read its
|
|
# size; run one embed and read the worker child's VmHWM — and set
|
|
# `measured=True` when you do.
|
|
SIGLIP_MODEL = ModelRequirement(
|
|
repo="google/siglip-so400m-patch14-384",
|
|
approx_download_bytes=3_500_000_000,
|
|
approx_resident_bytes=4 * GIB,
|
|
measured=False,
|
|
)
|
|
|
|
|
|
@dataclass(frozen=True)
|
|
class Lane:
|
|
"""A worker lane. `name` is the stable key the settings row is keyed on.
|
|
|
|
Keyed on a lane NAME rather than a container hostname for the reason
|
|
`models/service_seen.py` gives at length: celery's worker names here are
|
|
`celery@<container id>` and are minted fresh on every deploy, so anything
|
|
keyed on them records a death and a birth every time the stack updates.
|
|
"""
|
|
|
|
name: str
|
|
display_name: str
|
|
queues: tuple[str, ...]
|
|
# Which `entrypoint.sh` role starts this lane. NOT always the lane name:
|
|
# `maintenance_long` is the plain `worker` role pointed at a different
|
|
# queue, exactly as docker-compose starts it today (`command: ["worker"]`
|
|
# with CELERY_QUEUES=maintenance_long). Recorded here so the generated
|
|
# supervisord config and the compose file cannot disagree about it.
|
|
entrypoint_role: str
|
|
# THE cap a lane starts with — and, since 2026-09-23, the only number an
|
|
# operator sets for it. How many workers actually run is the autoscaler's
|
|
# job; this is the most it may use. Zero means the lane is off.
|
|
#
|
|
# One, and zero for ML. Deliberately far below the operator's own
|
|
# production numbers, which are tuned for their hardware and are not a
|
|
# sane first boot for a stranger — and low enough that a busy instance
|
|
# tells them to raise it rather than quietly consuming the machine.
|
|
default_slots_cap: int
|
|
# True when a slot costs a copy of the ML model rather than just a process.
|
|
# The only lane whose ceiling is decided by memory instead of by cores.
|
|
memory_bound: bool = False
|
|
# Models this lane downloads the first time it is enabled. Empty for every
|
|
# lane that needs none, which is how the UI knows whether to warn at all.
|
|
models: tuple[ModelRequirement, ...] = ()
|
|
# An optional lane is one the product works without. Shown as such, so
|
|
# nobody turns on a multi-GB download believing it is required.
|
|
optional: bool = False
|
|
|
|
@property
|
|
def queue_key(self) -> tuple[str, ...]:
|
|
"""The sorted queue set, which is how `service_seen` identifies a
|
|
running worker. The join between what is configured here and what
|
|
`celery inspect` reports."""
|
|
return tuple(sorted(self.queues))
|
|
|
|
|
|
# ONE CAP PER LANE, and that is the whole of what an operator sets.
|
|
#
|
|
# Operator, 2026-09-23: *"auto should be always on, not a setting, so that
|
|
# idle instances quiet down when not running. the number that is visible and
|
|
# something the user can tweak and manage should be the cap itself the number
|
|
# of running workers is handled by the autoscaling function which is always
|
|
# on."*
|
|
#
|
|
# Until then a lane had THREE operator values — `slots`, `slots_cap` and
|
|
# `autoscale` — because the manual dial was built first (steps 2-4) and the
|
|
# autoscaler arrived last (step 7) as an opt-in beside a control that already
|
|
# existed. Nothing ever asked whether the dial should still exist once
|
|
# something could move it automatically. It should not: "how many are running
|
|
# right now" is a measurement, not a preference.
|
|
#
|
|
# One of each, and ML at zero. ML at zero is also rule 164's carve-out: a cap
|
|
# of zero means no consumers, so a fresh install never loads a model or
|
|
# reaches HuggingFace, and raising the cap is what triggers the fetch.
|
|
#
|
|
# These are far below the operator's own production numbers, and deliberately
|
|
# so — they are what a stranger's first boot should do, not what a tuned
|
|
# machine can. The UI is what closes that gap: a lane sitting at its cap with
|
|
# a backlog says so, and says raising the cap is the fix. Without that a
|
|
# conservative default is just a slow instance nobody knows how to speed up.
|
|
LANES: tuple[Lane, ...] = (
|
|
Lane(
|
|
name="worker",
|
|
display_name="Worker",
|
|
queues=("default", "import", "thumbnail", "download"),
|
|
entrypoint_role="worker",
|
|
default_slots_cap=1,
|
|
),
|
|
Lane(
|
|
name="scheduler",
|
|
display_name="Scheduler",
|
|
queues=("maintenance", "scan"),
|
|
entrypoint_role="scheduler",
|
|
default_slots_cap=1,
|
|
),
|
|
Lane(
|
|
name="maintenance_long",
|
|
display_name="Long maintenance",
|
|
queues=("maintenance_long",),
|
|
entrypoint_role="worker",
|
|
default_slots_cap=1,
|
|
),
|
|
Lane(
|
|
name="ml",
|
|
display_name="ML tagging",
|
|
queues=("ml",),
|
|
entrypoint_role="ml-worker",
|
|
default_slots_cap=0,
|
|
memory_bound=True,
|
|
models=(SIGLIP_MODEL,),
|
|
optional=True,
|
|
),
|
|
)
|
|
|
|
LANES_BY_NAME: dict[str, Lane] = {lane.name: lane for lane in LANES}
|
|
LANES_BY_QUEUE_KEY: dict[tuple[str, ...], Lane] = {
|
|
lane.queue_key: lane for lane in LANES
|
|
}
|
|
|
|
|
|
# --- what the container actually has -----------------------------------------
|
|
|
|
# cgroup v2 first, then v1. A container started without an explicit memory
|
|
# limit reports "max" on v2 and a sentinel near 2**63 on v1; both mean "no
|
|
# limit", and the answer then is the host's RAM.
|
|
_CGROUP_V2_MEMORY = Path("/sys/fs/cgroup/memory.max")
|
|
_CGROUP_V1_MEMORY = Path("/sys/fs/cgroup/memory/memory.limit_in_bytes")
|
|
_CGROUP_V2_CPU = Path("/sys/fs/cgroup/cpu.max")
|
|
_CGROUP_V1_CPU_QUOTA = Path("/sys/fs/cgroup/cpu/cpu.cfs_quota_us")
|
|
_CGROUP_V1_CPU_PERIOD = Path("/sys/fs/cgroup/cpu/cpu.cfs_period_us")
|
|
|
|
# A v1 "unlimited" is PAGE_SIZE-aligned LONG_MAX, not a round number, so it is
|
|
# recognised by magnitude rather than by equality. Anything claiming more than
|
|
# a petabyte is a sentinel, not a machine.
|
|
_UNLIMITED_ABOVE = 1 << 50
|
|
|
|
# DERIVED from the model requirement above, never restated. The ceiling and
|
|
# the number shown to the operator before they enable the lane have to be the
|
|
# same figure, or the UI promises something the cap will then refuse.
|
|
ML_BYTES_PER_SLOT = SIGLIP_MODEL.approx_resident_bytes
|
|
|
|
# Held back for hypercorn and the non-ML lanes before any ML slot is offered.
|
|
# In the consolidated container these share one cgroup with ML, and they are
|
|
# the processes an OOM kill must not take (see the module docstring).
|
|
RESERVED_BYTES = 2 * GIB
|
|
|
|
# The smallest pool a lane can actually run: ONE process, never zero.
|
|
#
|
|
# billiard refuses to remove the last worker in a pool, so a lane asked to
|
|
# shrink to nothing gets `ValueError("Can't shrink pool. All processes
|
|
# busy!")` and the sizing pass re-sends the doomed message forever. Found on
|
|
# the operator's live deploy, 2026-09-23.
|
|
#
|
|
# It is also what makes "off" expressible: a lane at cap 0 keeps this one
|
|
# parked process with its consumers cancelled, so it still answers `inspect`
|
|
# (and so reads as present rather than crashed), and `add_consumer` has
|
|
# something to reach when the cap goes back up.
|
|
#
|
|
# Lives HERE rather than in `worker_control` because `gen_supervisord` needs
|
|
# it at container boot and must not import the models package to get it.
|
|
MIN_POOL_SLOTS = 1
|
|
|
|
# The floor a cores-derived ceiling never goes below. A single-core box still
|
|
# needs to be able to run its lanes; the ceiling exists to stop absurd values,
|
|
# not to make a small machine unusable.
|
|
MIN_CEILING = 1
|
|
|
|
# What an unreadable limit yields. Low rather than unlimited, on purpose: not
|
|
# knowing how much memory there is must never read as "plenty". An unswept
|
|
# absence is not a verdict.
|
|
UNKNOWN_CEILING = 1
|
|
|
|
|
|
def _read_int(path: Path) -> int | None:
|
|
try:
|
|
raw = path.read_text().strip()
|
|
except OSError:
|
|
return None
|
|
if raw == "max":
|
|
return None
|
|
try:
|
|
return int(raw)
|
|
except ValueError:
|
|
return None
|
|
|
|
|
|
def container_memory_bytes() -> int | None:
|
|
"""The memory this container may use, or None when it cannot be read.
|
|
|
|
None means UNKNOWN, never UNLIMITED. Every caller must treat it as the
|
|
conservative case — the whole point of the ceiling is to protect a machine
|
|
whose size we are unsure of.
|
|
"""
|
|
for path in (_CGROUP_V2_MEMORY, _CGROUP_V1_MEMORY):
|
|
value = _read_int(path)
|
|
if value is not None and value < _UNLIMITED_ABOVE:
|
|
return value
|
|
if value is not None:
|
|
# A sentinel: the cgroup exists but sets no limit, so the real
|
|
# bound is the host's.
|
|
break
|
|
try:
|
|
return os.sysconf("SC_PHYS_PAGES") * os.sysconf("SC_PAGE_SIZE")
|
|
except (ValueError, OSError, AttributeError):
|
|
return None
|
|
|
|
|
|
def container_cpu_count() -> int | None:
|
|
"""Effective cores, honouring a cgroup CPU quota.
|
|
|
|
`os.cpu_count()` reports the HOST's cores from inside a container, so a
|
|
quota of 2.0 on a 32-core host would otherwise offer 32 slots. The
|
|
operator's own stack sets `cpus: '4.0'` on ml-worker, so this is a real
|
|
configuration here and not a hypothetical.
|
|
"""
|
|
quota: float | None = None
|
|
try:
|
|
raw = _CGROUP_V2_CPU.read_text().strip().split()
|
|
if raw and raw[0] != "max":
|
|
quota = int(raw[0]) / int(raw[1])
|
|
except (OSError, ValueError, IndexError, ZeroDivisionError):
|
|
pass
|
|
if quota is None:
|
|
q = _read_int(_CGROUP_V1_CPU_QUOTA)
|
|
p = _read_int(_CGROUP_V1_CPU_PERIOD)
|
|
if q is not None and p and q > 0:
|
|
quota = q / p
|
|
if quota is not None and quota > 0:
|
|
return max(1, int(quota))
|
|
return os.cpu_count()
|
|
|
|
|
|
def derived_ceiling(lane: Lane) -> int:
|
|
"""The most slots `lane` may be given on this container.
|
|
|
|
Never stored. Recomputed on every read so a container whose limits changed
|
|
is bounded by what it has NOW rather than by what it had when its row was
|
|
written.
|
|
"""
|
|
if lane.memory_bound:
|
|
total = container_memory_bytes()
|
|
if total is None:
|
|
log.warning(
|
|
"worker_lanes: cannot read a memory limit; capping %s at %d",
|
|
lane.name, UNKNOWN_CEILING,
|
|
)
|
|
return UNKNOWN_CEILING
|
|
usable = total - RESERVED_BYTES
|
|
if usable < ML_BYTES_PER_SLOT:
|
|
# Honestly zero. A box that cannot hold one model alongside the web
|
|
# process must be told it cannot run tagging, not sold a slot that
|
|
# will OOM the container the first time it is used.
|
|
return 0
|
|
return int(usable // ML_BYTES_PER_SLOT)
|
|
|
|
cores = container_cpu_count()
|
|
if cores is None:
|
|
return UNKNOWN_CEILING
|
|
return max(MIN_CEILING, cores)
|
|
|
|
|
|
def ceilings() -> dict[str, int]:
|
|
"""Every lane's ceiling, for the settings API and the UI."""
|
|
return {lane.name: derived_ceiling(lane) for lane in LANES}
|