fix: the ML dial offered slots the machine had no cores to feed (4295)
CI and images / lint (push) Successful in 2s
CI and images / extension-version (push) Successful in 3s
CI and images / frontend-build (push) Successful in 22s
CI and images / backend-lint-and-test (push) Successful in 32s
CI and images / integration (push) Successful in 2m21s
CI and images / sign-extension (push) Successful in 3s
CI and images / build-agent (push) Successful in 5s
CI and images / build-web (push) Successful in 1m42s
CI and images / smoke-web (push) Successful in 1m7s
CI and images / promote (push) Skipped
CI and images / lint (push) Successful in 2s
CI and images / extension-version (push) Successful in 3s
CI and images / frontend-build (push) Successful in 22s
CI and images / backend-lint-and-test (push) Successful in 32s
CI and images / integration (push) Successful in 2m21s
CI and images / sign-extension (push) Successful in 3s
CI and images / build-agent (push) Successful in 5s
CI and images / build-web (push) Successful in 1m42s
CI and images / smoke-web (push) Successful in 1m7s
CI and images / promote (push) Skipped
Operator's 2026-09-23 log: embed_image taking 107-246s each, ~49 slots in flight by Little's law, and the daily CCIP sweep dying on its 1800s soft limit in a numpy matmul. The billiard/pool.py frame in that traceback is the soft-timeout signal handler, not a pool fault. Two causes, both mine. 1. `derived_ceiling` computed the ML lane from MEMORY ALONE. Meanwhile `embedder.py` carried `_INTRA_OP_THREADS = 4` beside a comment reading "keep N_replicas x this within the cores allotted to ML" — a constraint stated where nothing could act on it. A large-memory host offered ~49 slots, the operator took what the dial offered, and the lane asked the box for ~200 torch threads. The number moves onto the lane as `threads_per_slot`, the embedder reads it rather than restating it, and the ceiling is now the smaller of the two bounds. They fail differently on purpose: too little memory is honestly zero, because the first task would OOM the container; too few cores is merely slow, so it floors at one rather than making the lane unreachable on a small box. 2. `scheduled_ccip_auto_apply` scored one image per matmul, over every image in the library, on every daily run — ~119k products each too small to pay for its own BLAS setup. `char_maxima` does the same arithmetic in blocks bounded by elements, so its memory stays flat as either axis grows. Batching changes no arithmetic: a character's score for an image is a max over that image's figures AND that character's prototypes, and max does not care how it is grouped. Pinned against the old loop written out longhand, and against itself with the blocking forced to split every row. The UI copy said the ML ceiling came from memory; it says cores or memory, whichever runs out first. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
This commit is contained in:
@@ -57,3 +57,88 @@ def test_applied_or_rejected_unions_applied_any_source_and_rejected(db_sync):
|
||||
assert skip[b.id] == {imgs[3].id}
|
||||
assert imgs[4].id not in skip[a.id]
|
||||
assert imgs[4].id not in skip[b.id]
|
||||
|
||||
|
||||
# --- the CCIP auto-apply sweep's scorer ---------------------------------------
|
||||
#
|
||||
# `scheduled_ccip_auto_apply` scored one image per matmul, over every image in
|
||||
# the library, on every daily run — and on 2026-09-23 it hit its 1800s soft
|
||||
# limit on the operator's instance. `char_maxima` does the same arithmetic in
|
||||
# blocks. These pin THAT: same answer, whatever the blocking.
|
||||
|
||||
|
||||
def _score_fixture(np):
|
||||
"""Four images with 1-3 figures each, three characters with 2/5/1
|
||||
prototypes. Deliberately ragged — equal group sizes would let a wrong
|
||||
`reduceat` offset pass."""
|
||||
from backend.app.services.ml.training_data import _l2norm
|
||||
|
||||
rng = np.random.default_rng(7)
|
||||
dim = 16
|
||||
q_by_image = [
|
||||
_l2norm(rng.standard_normal((n, dim)).astype(np.float32), np)
|
||||
for n in (1, 3, 2, 1)
|
||||
]
|
||||
mats = [
|
||||
_l2norm(rng.standard_normal((k, dim)).astype(np.float32), np)
|
||||
for k in (2, 5, 1)
|
||||
]
|
||||
allref = np.vstack(mats)
|
||||
seg = np.cumsum([0] + [len(m) for m in mats])[:-1]
|
||||
return q_by_image, allref, seg
|
||||
|
||||
|
||||
def _naive(q_by_image, allref, seg, np):
|
||||
"""The loop as it was written before batching, kept longhand. The point of
|
||||
comparing against this rather than against a stored array is that it is
|
||||
the OLD CODE — if the batched form ever diverges, this says so in the
|
||||
terms the change was justified in."""
|
||||
return np.vstack([
|
||||
np.maximum.reduceat((q @ allref.T).max(axis=0), seg) for q in q_by_image
|
||||
])
|
||||
|
||||
|
||||
def test_char_maxima_matches_the_per_image_loop():
|
||||
import numpy as np
|
||||
|
||||
from backend.app.services.ml.ccip import char_maxima
|
||||
|
||||
q_by_image, allref, seg = _score_fixture(np)
|
||||
got = char_maxima(q_by_image, allref, seg, np)
|
||||
|
||||
assert got.shape == (len(q_by_image), len(seg))
|
||||
np.testing.assert_allclose(
|
||||
got, _naive(q_by_image, allref, seg, np), rtol=1e-6, atol=1e-6,
|
||||
)
|
||||
|
||||
|
||||
def test_the_answer_does_not_depend_on_where_the_blocks_fall():
|
||||
"""The one thing batching could get wrong. Rows are reduced over the
|
||||
PROTOTYPE axis inside a block and over the FIGURE axis afterwards, so a
|
||||
block boundary may fall in the middle of an image's figures — which is
|
||||
safe only because max does not care how it is grouped. `max_elems=1`
|
||||
forces a boundary between every single row."""
|
||||
import numpy as np
|
||||
|
||||
from backend.app.services.ml.ccip import char_maxima
|
||||
|
||||
q_by_image, allref, seg = _score_fixture(np)
|
||||
whole = char_maxima(q_by_image, allref, seg, np, max_elems=10_000_000)
|
||||
split = char_maxima(q_by_image, allref, seg, np, max_elems=1)
|
||||
|
||||
np.testing.assert_allclose(whole, split, rtol=1e-6, atol=1e-6)
|
||||
|
||||
|
||||
def test_one_character_and_one_figure_still_reduces():
|
||||
"""The degenerate shape `reduceat` is easiest to get wrong: a single
|
||||
segment starting at 0, and a single row."""
|
||||
import numpy as np
|
||||
|
||||
from backend.app.services.ml.ccip import char_maxima
|
||||
|
||||
q = np.array([[1.0, 0.0]], dtype=np.float32)
|
||||
allref = np.array([[1.0, 0.0], [0.0, 1.0]], dtype=np.float32)
|
||||
got = char_maxima([q], allref, np.array([0]), np)
|
||||
|
||||
assert got.shape == (1, 1)
|
||||
assert got[0][0] == pytest.approx(1.0)
|
||||
|
||||
@@ -153,10 +153,60 @@ def test_v1_sentinel_is_recognised_as_unlimited(monkeypatch, tmp_path):
|
||||
|
||||
def test_ml_ceiling_is_memory_divided_by_per_slot_cost(monkeypatch, tmp_path):
|
||||
# 2 GiB reserved for web and the other lanes, then 4 GiB per model copy.
|
||||
# Cores pinned high so the memory bound is the one being read here.
|
||||
_point_memory_at(monkeypatch, tmp_path, str(14 * wl.GIB))
|
||||
monkeypatch.setattr(wl, "container_cpu_count", lambda: 32)
|
||||
assert wl.derived_ceiling(wl.LANES_BY_NAME["ml"]) == 3
|
||||
|
||||
|
||||
def test_ml_is_bounded_by_cores_as_well_as_memory(monkeypatch, tmp_path):
|
||||
"""The bug of 2026-09-23, in one assertion.
|
||||
|
||||
The ML ceiling was memory ALONE. On the operator's large-memory host that
|
||||
offered ~49 slots; they took them, and each slot asks torch for
|
||||
`threads_per_slot` cores — so the lane ran ~200 threads over a box that
|
||||
has nowhere near that many. Embeds that should be seconds took 107-246s,
|
||||
and the daily CCIP sweep sharing that pool died on its 1800s soft limit.
|
||||
|
||||
Memory here says 49. The cores say 8 / 4 = 2, and the smaller bound is the
|
||||
only honest one: a control must not offer a number the machine cannot
|
||||
feed.
|
||||
"""
|
||||
_point_memory_at(monkeypatch, tmp_path, str(200 * wl.GIB))
|
||||
monkeypatch.setattr(wl, "container_cpu_count", lambda: 8)
|
||||
|
||||
ml = wl.LANES_BY_NAME["ml"]
|
||||
assert ml.threads_per_slot == 4
|
||||
assert wl.derived_ceiling(ml) == 2
|
||||
|
||||
|
||||
def test_a_few_cores_still_offer_one_ml_slot_rather_than_none(
|
||||
monkeypatch, tmp_path,
|
||||
):
|
||||
"""The two bounds fail differently, deliberately. Too little MEMORY is
|
||||
honestly zero — the first task would OOM the container. Too few CORES is
|
||||
merely slow, which is a trade an operator may want, so it floors at one
|
||||
rather than making the lane unreachable on a small box."""
|
||||
_point_memory_at(monkeypatch, tmp_path, str(200 * wl.GIB))
|
||||
monkeypatch.setattr(wl, "container_cpu_count", lambda: 1)
|
||||
assert wl.derived_ceiling(wl.LANES_BY_NAME["ml"]) == 1
|
||||
|
||||
|
||||
def test_the_embedder_asks_for_exactly_what_the_ceiling_budgeted(monkeypatch):
|
||||
"""The two halves of the same number, tied.
|
||||
|
||||
`threads_per_slot` is only meaningful because `embedder.load()` calls
|
||||
`torch.set_num_threads` with it. It lived in the embedder as a private 4
|
||||
beside a comment saying "keep N_replicas x this within the cores allotted
|
||||
to ML" — a constraint stated where nothing could enforce it, and nothing
|
||||
did. If these two ever drift, the ceiling is budgeting cores for a demand
|
||||
the worker does not make, and nothing else would notice.
|
||||
"""
|
||||
from backend.app.services.ml import embedder
|
||||
|
||||
assert embedder._INTRA_OP_THREADS == wl.LANES_BY_NAME["ml"].threads_per_slot
|
||||
|
||||
|
||||
def test_a_small_box_is_told_it_cannot_run_tagging(monkeypatch, tmp_path):
|
||||
"""Honestly zero rather than a floor of one. A 4GB box cannot hold a model
|
||||
alongside the web process, and offering a slot that OOMs the container the
|
||||
@@ -206,6 +256,7 @@ def test_the_ceiling_is_computed_not_stored(monkeypatch, tmp_path):
|
||||
ceiling when the container's limits change, with no row edit. A stored
|
||||
ceiling would keep authorising what the box no longer has."""
|
||||
lane = wl.LANES_BY_NAME["ml"]
|
||||
monkeypatch.setattr(wl, "container_cpu_count", lambda: 32)
|
||||
_point_memory_at(monkeypatch, tmp_path, str(34 * wl.GIB))
|
||||
big = wl.derived_ceiling(lane)
|
||||
(tmp_path / "memory.max").write_text(str(10 * wl.GIB))
|
||||
|
||||
Reference in New Issue
Block a user