fix: the ML dial offered slots the machine had no cores to feed (4295)
CI and images / lint (push) Successful in 2s
CI and images / extension-version (push) Successful in 3s
CI and images / frontend-build (push) Successful in 22s
CI and images / backend-lint-and-test (push) Successful in 32s
CI and images / integration (push) Successful in 2m21s
CI and images / sign-extension (push) Successful in 3s
CI and images / build-agent (push) Successful in 5s
CI and images / build-web (push) Successful in 1m42s
CI and images / smoke-web (push) Successful in 1m7s
CI and images / promote (push) Skipped

Operator's 2026-09-23 log: embed_image taking 107-246s each, ~49 slots in
flight by Little's law, and the daily CCIP sweep dying on its 1800s soft
limit in a numpy matmul. The billiard/pool.py frame in that traceback is
the soft-timeout signal handler, not a pool fault.

Two causes, both mine.

1. `derived_ceiling` computed the ML lane from MEMORY ALONE. Meanwhile
   `embedder.py` carried `_INTRA_OP_THREADS = 4` beside a comment reading
   "keep N_replicas x this within the cores allotted to ML" — a constraint
   stated where nothing could act on it. A large-memory host offered ~49
   slots, the operator took what the dial offered, and the lane asked the
   box for ~200 torch threads.

   The number moves onto the lane as `threads_per_slot`, the embedder
   reads it rather than restating it, and the ceiling is now the smaller
   of the two bounds. They fail differently on purpose: too little memory
   is honestly zero, because the first task would OOM the container; too
   few cores is merely slow, so it floors at one rather than making the
   lane unreachable on a small box.

2. `scheduled_ccip_auto_apply` scored one image per matmul, over every
   image in the library, on every daily run — ~119k products each too
   small to pay for its own BLAS setup. `char_maxima` does the same
   arithmetic in blocks bounded by elements, so its memory stays flat as
   either axis grows.

   Batching changes no arithmetic: a character's score for an image is a
   max over that image's figures AND that character's prototypes, and max
   does not care how it is grouped. Pinned against the old loop written
   out longhand, and against itself with the blocking forced to split
   every row.

The UI copy said the ML ceiling came from memory; it says cores or
memory, whichever runs out first.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
This commit is contained in:
2026-09-23 16:46:22 -04:00
co-authored by Claude Opus 5
parent 1353d346b3
commit 7f1693a40d
7 changed files with 277 additions and 33 deletions
+52
View File
@@ -35,6 +35,58 @@ DEFAULT_SIM_THRESHOLD = 0.85
_FIGURE_KINDS = ("face", "figure")
# How many cosine scores to hold in memory at once, per matmul block.
# 4M float32 is 16 MB — small enough to stay in cache-friendly territory on the
# shared ml lane, large enough that the per-call overhead stops mattering.
_MAX_SCORE_ELEMS = 4_000_000
def char_maxima(q_by_image, allref, seg, np, *, max_elems=_MAX_SCORE_ELEMS):
"""(n_images, n_chars) — each image's best cosine to each character.
`q_by_image` is one L2-normalised `(n_figures, dim)` array per image, in
the order the answer comes back in. `allref` is every character's
prototypes stacked, and `seg` their per-character start offsets into it.
## Why this is batched, and why that is safe
`scheduled_ccip_auto_apply` did this one image at a time — a `(nq, dim) @
(dim, total)` product per image, over every image in the library on every
run. At ~119k images that is 119k separate matmuls, each too small to pay
for its own BLAS setup, and on 2026-09-23 the daily sweep hit its 1800s
soft limit on the operator's instance.
Batching changes no arithmetic. The score a character gets for an image is
a max over that image's figures AND over that character's prototypes, and
max does not care in what order or grouping it is taken — so reducing the
prototype axis first (per row, inside a block) and the figure axis after
(per image, across blocks) gives exactly what the per-image loop gave.
That equivalence is what `test_char_maxima_matches_the_per_image_loop`
pins, against the naive form written out longhand.
Blocked by ROWS rather than done in one product, because the full score
matrix is (all figures in the chunk x every prototype) and that grows with
the library on both axes. The block bound is on elements, so the memory
this uses stays flat as either axis grows.
"""
counts = [len(q) for q in q_by_image]
rows = np.vstack(q_by_image)
total = max(int(allref.shape[0]), 1)
block = max(1, max_elems // total)
per_row = np.empty((rows.shape[0], len(seg)), dtype=np.float32)
for a in range(0, rows.shape[0], block):
scores = rows[a:a + block] @ allref.T
per_row[a:a + block] = np.maximum.reduceat(scores, seg, axis=1)
# Start offset of each image's rows. Every image has at least one figure —
# it is in `q_by_image` because a region produced it — so these strictly
# increase, which is what `reduceat` needs to reduce rather than pass a row
# through untouched.
starts = np.cumsum([0] + counts[:-1])
return np.maximum.reduceat(per_row, starts, axis=0)
async def _settings_threshold(session: AsyncSession) -> float:
val = (
await session.execute(