Operator's 2026-09-23 log: embed_image taking 107-246s each, ~49 slots in
flight by Little's law, and the daily CCIP sweep dying on its 1800s soft
limit in a numpy matmul. The billiard/pool.py frame in that traceback is
the soft-timeout signal handler, not a pool fault.
Two causes, both mine.
1. `derived_ceiling` computed the ML lane from MEMORY ALONE. Meanwhile
`embedder.py` carried `_INTRA_OP_THREADS = 4` beside a comment reading
"keep N_replicas x this within the cores allotted to ML" — a constraint
stated where nothing could act on it. A large-memory host offered ~49
slots, the operator took what the dial offered, and the lane asked the
box for ~200 torch threads.
The number moves onto the lane as `threads_per_slot`, the embedder
reads it rather than restating it, and the ceiling is now the smaller
of the two bounds. They fail differently on purpose: too little memory
is honestly zero, because the first task would OOM the container; too
few cores is merely slow, so it floors at one rather than making the
lane unreachable on a small box.
2. `scheduled_ccip_auto_apply` scored one image per matmul, over every
image in the library, on every daily run — ~119k products each too
small to pay for its own BLAS setup. `char_maxima` does the same
arithmetic in blocks bounded by elements, so its memory stays flat as
either axis grows.
Batching changes no arithmetic: a character's score for an image is a
max over that image's figures AND that character's prototypes, and max
does not care how it is grouped. Pinned against the old loop written
out longhand, and against itself with the blocking forced to split
every row.
The UI copy said the ML ceiling came from memory; it says cores or
memory, whichever runs out first.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
Consolidate duplication accrued across the ML tagging + settings backend,
behavior-preserving (over-DRY guard applied — the three auto-apply sweep
BODIES stay separate; only their shared inner helpers are extracted).
- _sigmoid / _conflict_scores / _insert_presentation_review (heads.py): the
score→prob transform (6 inlined sites), the presentation conflict signal
(2 sites), and the ring-loud PresentationReview insert (2 sites, single-
sourced so the mode column can't drift on the shared composite PK).
- _applied_or_rejected (training_data.py): the per-tag "applied ∪ rejected"
skip-set, byte-identical at 3 sweep sites (heads.py x2, tasks/ml.py ccip).
- ccip sweep divergence fixes: import ccip._FIGURE_KINDS + training_data._l2norm
instead of local copies that silently drift when the canonical changes.
- MLSettings.load / .load_sync classmethods (mirror ImportSettings); route all
8 scalar_one singleton reads through them (the session.get None-path stays).
- GET serializers for MLSettings + ImportSettings are now table-driven off the
same _EDITABLE tuples PATCH writes, so a new field can't be silently absent
from GET (the split that historically dropped fields).
- AUTO_APPLY_THRESHOLD_MIN/MAX constant single-sources the [0.5,0.999] operating
range across the service clamp + the 5 API validators.
- test_ml_dry_helpers.py pins _applied_or_rejected + _sigmoid.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NsmJSQxnNxGgtM5Yz4GAqi