Files
FabledCurator/backend/app/services/ml/ccip.py
T
bvandeusenandClaude Opus 5 7f1693a40d
CI and images / lint (push) Successful in 2s
CI and images / extension-version (push) Successful in 3s
CI and images / frontend-build (push) Successful in 22s
CI and images / backend-lint-and-test (push) Successful in 32s
CI and images / integration (push) Successful in 2m21s
CI and images / sign-extension (push) Successful in 3s
CI and images / build-agent (push) Successful in 5s
CI and images / build-web (push) Successful in 1m42s
CI and images / smoke-web (push) Successful in 1m7s
CI and images / promote (push) Skipped
fix: the ML dial offered slots the machine had no cores to feed (4295)
Operator's 2026-09-23 log: embed_image taking 107-246s each, ~49 slots in
flight by Little's law, and the daily CCIP sweep dying on its 1800s soft
limit in a numpy matmul. The billiard/pool.py frame in that traceback is
the soft-timeout signal handler, not a pool fault.

Two causes, both mine.

1. `derived_ceiling` computed the ML lane from MEMORY ALONE. Meanwhile
   `embedder.py` carried `_INTRA_OP_THREADS = 4` beside a comment reading
   "keep N_replicas x this within the cores allotted to ML" — a constraint
   stated where nothing could act on it. A large-memory host offered ~49
   slots, the operator took what the dial offered, and the lane asked the
   box for ~200 torch threads.

   The number moves onto the lane as `threads_per_slot`, the embedder
   reads it rather than restating it, and the ceiling is now the smaller
   of the two bounds. They fail differently on purpose: too little memory
   is honestly zero, because the first task would OOM the container; too
   few cores is merely slow, so it floors at one rather than making the
   lane unreachable on a small box.

2. `scheduled_ccip_auto_apply` scored one image per matmul, over every
   image in the library, on every daily run — ~119k products each too
   small to pay for its own BLAS setup. `char_maxima` does the same
   arithmetic in blocks bounded by elements, so its memory stays flat as
   either axis grows.

   Batching changes no arithmetic: a character's score for an image is a
   max over that image's figures AND that character's prototypes, and max
   does not care how it is grouped. Pinned against the old loop written
   out longhand, and against itself with the blocking forced to split
   every row.

The UI copy said the ML ceiling came from memory; it says cores or
memory, whichever runs out first.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-23 16:46:22 -04:00

359 lines
14 KiB
Python

"""CCIP few-shot character matcher (#114) — server-side, numpy on stored vectors.
CCIP is a FROZEN identity embedding; we don't train it. Instead the operator's
tagged characters become reference prototypes: a character tag's references are
the CCIP vectors of figure/face regions on images carrying that tag. To suggest
characters for a new image, we compare its figure-region CCIP vectors to every
character's references (multi-prototype: best match over a character's examples)
and surface the ones that clear a similarity threshold. No GPU here — the agent
already produced the vectors; this is cosine matching on what's stored.
v1 uses cosine similarity on the raw CCIP vectors with a tunable threshold; the
exact CCIP difference metric/threshold gets validated against the model during
the hands-on eval. numpy is imported lazily (API worker has it via pgvector).
"""
from sqlalchemy import exists, func, select
from sqlalchemy.ext.asyncio import AsyncSession
from ...models import (
CcipPrototypeState,
CharacterPrototype,
ImageRegion,
MLSettings,
Tag,
TagKind,
TagPositiveConfirmation,
)
from ...models.tag import image_tag
from .training_data import _AUTO_SOURCES
# Cosine-similarity floor to call a figure the same character. The live setting
# (ml_settings.ccip_match_threshold) drives it; this is only the fallback when no
# threshold is supplied AND no settings row exists.
DEFAULT_SIM_THRESHOLD = 0.85
_FIGURE_KINDS = ("face", "figure")
# How many cosine scores to hold in memory at once, per matmul block.
# 4M float32 is 16 MB — small enough to stay in cache-friendly territory on the
# shared ml lane, large enough that the per-call overhead stops mattering.
_MAX_SCORE_ELEMS = 4_000_000
def char_maxima(q_by_image, allref, seg, np, *, max_elems=_MAX_SCORE_ELEMS):
"""(n_images, n_chars) — each image's best cosine to each character.
`q_by_image` is one L2-normalised `(n_figures, dim)` array per image, in
the order the answer comes back in. `allref` is every character's
prototypes stacked, and `seg` their per-character start offsets into it.
## Why this is batched, and why that is safe
`scheduled_ccip_auto_apply` did this one image at a time — a `(nq, dim) @
(dim, total)` product per image, over every image in the library on every
run. At ~119k images that is 119k separate matmuls, each too small to pay
for its own BLAS setup, and on 2026-09-23 the daily sweep hit its 1800s
soft limit on the operator's instance.
Batching changes no arithmetic. The score a character gets for an image is
a max over that image's figures AND over that character's prototypes, and
max does not care in what order or grouping it is taken — so reducing the
prototype axis first (per row, inside a block) and the figure axis after
(per image, across blocks) gives exactly what the per-image loop gave.
That equivalence is what `test_char_maxima_matches_the_per_image_loop`
pins, against the naive form written out longhand.
Blocked by ROWS rather than done in one product, because the full score
matrix is (all figures in the chunk x every prototype) and that grows with
the library on both axes. The block bound is on elements, so the memory
this uses stays flat as either axis grows.
"""
counts = [len(q) for q in q_by_image]
rows = np.vstack(q_by_image)
total = max(int(allref.shape[0]), 1)
block = max(1, max_elems // total)
per_row = np.empty((rows.shape[0], len(seg)), dtype=np.float32)
for a in range(0, rows.shape[0], block):
scores = rows[a:a + block] @ allref.T
per_row[a:a + block] = np.maximum.reduceat(scores, seg, axis=1)
# Start offset of each image's rows. Every image has at least one figure —
# it is in `q_by_image` because a region produced it — so these strictly
# increase, which is what `reduceat` needs to reduce rather than pass a row
# through untouched.
starts = np.cumsum([0] + counts[:-1])
return np.maximum.reduceat(per_row, starts, axis=0)
async def _settings_threshold(session: AsyncSession) -> float:
val = (
await session.execute(
select(MLSettings.ccip_match_threshold).where(MLSettings.id == 1)
)
).scalar_one_or_none()
return float(val) if val is not None else DEFAULT_SIM_THRESHOLD
def _l2norm(mat, np):
n = np.linalg.norm(mat, axis=1, keepdims=True)
n[n == 0] = 1.0
return mat / n
# Single-shot cache of the (expensive) reference load, keyed on a cheap
# signature that changes exactly when references could: a character tag added/
# removed (n_char_tags) or a figure embedded (max/ n of ccip regions). Shared by
# the live matcher (every modal open) and the auto-apply sweep.
_REF_CACHE: dict = {"sig": None, "refs": None}
def _single_character_images():
"""Subquery of image ids carrying EXACTLY ONE character tag. References come
only from these — on a multi-character image the tag is image-level, so every
figure would otherwise pollute each character's prototype set (a 2-character
image tagged 'Velma' would make Daphne's figure a Velma reference)."""
return (
select(image_tag.c.image_record_id)
.join(Tag, Tag.id == image_tag.c.tag_id)
.where(Tag.kind == TagKind.character)
.group_by(image_tag.c.image_record_id)
.having(func.count() == 1)
)
def _hygiene_tagged_images():
"""Subquery of image ids carrying any SYSTEM tag (wip / banner / editor
screenshot). Training hygiene (#128): such images never contribute
reference prototypes — a faceless wip's figure region would otherwise
become an identity reference for the character it's tagged with."""
return (
select(image_tag.c.image_record_id)
.join(Tag, Tag.id == image_tag.c.tag_id)
.where(Tag.is_system.is_(True))
)
async def _ref_signature(session: AsyncSession) -> tuple:
n_tags = (
await session.execute(
select(func.count())
.select_from(image_tag)
.join(Tag, Tag.id == image_tag.c.tag_id)
.where(Tag.kind == TagKind.character)
)
).scalar_one()
n_regs, max_id = (
await session.execute(
select(func.count(), func.max(ImageRegion.id)).where(
ImageRegion.kind.in_(_FIGURE_KINDS),
ImageRegion.ccip_embedding.is_not(None),
)
)
).one()
# Hygiene applications must invalidate too: tagging an image `wip` changes
# the reference set without touching character-tag or region counts.
n_hygiene = (
await session.execute(
select(func.count())
.select_from(image_tag)
.join(Tag, Tag.id == image_tag.c.tag_id)
.where(Tag.is_system.is_(True))
)
).scalar_one()
return (n_tags, n_regs, max_id, n_hygiene)
def _positive_char_tag():
"""Condition on the joined character image_tag: HUMAN-applied or operator-
confirmed — NOT an unconfirmed auto-apply. Keeps an auto-tagged character from
self-seeding CCIP references, so a ccip_auto misfire can't reinforce itself
(milestone 139) — mirrors the head-training positive exclusion."""
return image_tag.c.source.not_in(_AUTO_SOURCES) | exists().where(
TagPositiveConfirmation.image_record_id == image_tag.c.image_record_id,
TagPositiveConfirmation.tag_id == image_tag.c.tag_id,
)
async def character_references(session: AsyncSession) -> dict[int, list]:
"""Per character-tag CCIP reference vectors: figure/face-region CCIP
embeddings on UNAMBIGUOUS (single-character) images carrying that tag.
Multi-prototype — several vectors per character. Cached on a cheap signature."""
sig = await _ref_signature(session)
if _REF_CACHE["sig"] == sig and _REF_CACHE["refs"] is not None:
return _REF_CACHE["refs"]
rows = (
await session.execute(
select(image_tag.c.tag_id, ImageRegion.ccip_embedding)
.select_from(ImageRegion)
.join(
image_tag,
image_tag.c.image_record_id == ImageRegion.image_record_id,
)
.join(Tag, Tag.id == image_tag.c.tag_id)
.where(Tag.kind == TagKind.character)
.where(_positive_char_tag())
.where(ImageRegion.kind.in_(_FIGURE_KINDS))
.where(ImageRegion.ccip_embedding.is_not(None))
.where(ImageRegion.image_record_id.in_(_single_character_images()))
.where(
ImageRegion.image_record_id.not_in(_hygiene_tagged_images())
)
)
).all()
refs: dict[int, list] = {}
for tag_id, vec in rows:
refs.setdefault(tag_id, []).append(vec)
_REF_CACHE.update(sig=sig, refs=refs)
return refs
async def _tag_names(session: AsyncSession, tag_ids: list[int]) -> dict[int, str]:
if not tag_ids:
return {}
return dict(
(
await session.execute(
select(Tag.id, Tag.name).where(Tag.id.in_(tag_ids))
)
).all()
)
# Per-character normalized prototype matrices, cached per process and refreshed
# INCREMENTALLY: only characters whose ccip_prototype_state.updated_at advanced
# are reloaded. This replaces the request-path rebuild of the ENTIRE reference
# blob (the ~4s stall, #1317) — the prototypes are precomputed off the request
# path by services.ml.character_prototypes (a beat + after each retrain).
_PROTO_CACHE: dict = {"mats": {}, "ver": {}}
async def _load_prototypes(session: AsyncSession) -> dict:
"""{tag_id: (P, D) L2-normalized prototype matrix} from character_prototype,
served from the in-process cache and reloading ONLY the characters whose
updated_at changed. Empty dict when the store isn't populated yet (cold start
→ match_image falls back to the legacy on-the-fly reference build)."""
import numpy as np
versions = dict(
(
await session.execute(
select(CcipPrototypeState.tag_id, CcipPrototypeState.updated_at)
)
).all()
)
mats = _PROTO_CACHE["mats"]
ver = _PROTO_CACHE["ver"]
# Forget characters that no longer have prototypes.
for tag_id in [t for t in mats if t not in versions]:
mats.pop(tag_id, None)
ver.pop(tag_id, None)
# Reload only the characters whose prototypes changed since we cached them.
stale = [t for t, u in versions.items() if ver.get(t) != u]
if stale:
rows = (
await session.execute(
select(
CharacterPrototype.tag_id, CharacterPrototype.ccip_embedding
).where(CharacterPrototype.tag_id.in_(stale))
)
).all()
by_tag: dict[int, list] = {}
for tag_id, vec in rows:
by_tag.setdefault(tag_id, []).append(
np.asarray(vec, dtype=np.float32)
)
for tag_id in stale:
vecs = by_tag.get(tag_id)
if vecs:
mats[tag_id] = _l2norm(np.vstack(vecs), np)
ver[tag_id] = versions[tag_id]
else:
mats.pop(tag_id, None)
ver.pop(tag_id, None)
return mats
async def match_image(
session: AsyncSession, image_id: int, threshold: float | None = None
) -> list[dict]:
"""Character suggestions for one image from its figure-region CCIP vectors:
[{tag_id, name, category:'character', score, source:'ccip'}], ranked.
Already-applied character tags are excluded. Empty if the image has no figure
CCIP vectors or no character references exist yet. threshold defaults to the
live ml_settings.ccip_match_threshold."""
import numpy as np
if threshold is None:
threshold = await _settings_threshold(session)
# Keep each figure region's bbox alongside its vector so a match can point at
# the figure that matched (#1206 grounding), not just the score.
fig_rows = (
await session.execute(
select(
ImageRegion.ccip_embedding,
ImageRegion.rx, ImageRegion.ry, ImageRegion.rw, ImageRegion.rh,
ImageRegion.kind, ImageRegion.detector_version,
).where(
ImageRegion.image_record_id == image_id,
ImageRegion.kind.in_(_FIGURE_KINDS),
ImageRegion.ccip_embedding.is_not(None),
)
)
).all()
if not fig_rows:
return []
# Prefer the precomputed prototype store (fast, incremental). On a cold start
# (store not yet populated post-deploy) fall back to the legacy on-the-fly
# reference build so character suggestions work immediately — the background
# refresh populates the store within ~15 min, after which this path is used
# and the per-accept ~4s rebuild is gone (#1317).
protos = await _load_prototypes(session)
refs = protos if protos else await character_references(session)
if not refs:
return []
applied = set(
(
await session.execute(
select(image_tag.c.tag_id).where(
image_tag.c.image_record_id == image_id
)
)
).scalars()
)
names = await _tag_names(session, [t for t in refs if t not in applied])
qvecs = [r[0] for r in fig_rows]
fig_meta = [
{"bbox": [rx, ry, rw, rh], "kind": kind, "detector": detector}
for _v, rx, ry, rw, rh, kind, detector in fig_rows
]
Q = _l2norm(np.vstack([np.asarray(v, dtype=np.float32) for v in qvecs]), np)
out = []
for tag_id, vecs in refs.items():
if tag_id in applied:
continue
# Prototype matrices are already L2-normalized; legacy refs are raw
# vector lists that still need stacking + normalizing.
R = vecs if protos else _l2norm(
np.vstack([np.asarray(v, dtype=np.float32) for v in vecs]), np
)
sims = Q @ R.T # (n_query_figures, n_references)
per_figure = sims.max(axis=1) # best reference cosine per figure
best_figure = int(per_figure.argmax())
best = float(per_figure[best_figure])
if best >= threshold:
out.append({
"tag_id": tag_id,
"name": names.get(tag_id, str(tag_id)),
"category": "character",
"score": round(best, 4),
"source": "ccip",
# the figure region that matched → grounds the character tag.
"grounding": fig_meta[best_figure],
})
out.sort(key=lambda d: d["score"], reverse=True)
return out