Build images / sign-extension (push) Successful in 4s
CI / lint (push) Failing after 2s
CI / extension-version (push) Successful in 2s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 27s
Build images / build-ml (push) Successful in 48s
CI / backend-lint-and-test (push) Successful in 1m7s
Build images / build-web (push) Successful in 40s
CI / integration (push) Successful in 4m1s
Milestone 328's acceptance test compared a database built by the real
0001..0087 chain against one built from the models, and found ~130
places where they disagree. This closes them.
Almost all were the MODEL being wrong, so almost all of this is model
edits with no DDL — the database already had these things, nothing in it
changes, and no deploy is needed for this part:
* 92 columns gained server_default. The models carried Python-side
`default=` only, so the ORM filled the value and the column had no
database default. Anything inserting outside the ORM behaved
differently from production.
* Eleven indexes that existed only in migrations are now declared:
the three backup_run reporting indexes, the two date-ordered
image_record browse indexes, import_task and presentation_review,
and the three task_run history indexes. All use text() for their DESC
ordering and postgresql_where for the partial one.
* Two UNIQUE indexes that autogenerate silently proposed DROPPING,
because neither is expressible as a UniqueConstraint:
uq_tag_name_kind_fandom — an EXPRESSION index over
(name, kind, COALESCE(fandom_id, 0))
uq_post_artist_external_id_null_source — PARTIAL, WHERE source_id
IS NULL
post.py already had a comment describing the second one. The comment
was right; nothing declared it.
* The two external_link enum CHECKs (host, status) — rule 36 territory,
and absent from the model entirely.
* Two indexes were named explicitly. A bare index=True generated
ix_tag_alias_canonical_tag_id where the database has
ix_tag_alias_canonical, so autogenerate proposed a drop+create of an
index that was already there under another name. Same for
tag_suggestion_rejection.
Only ONE thing needed DDL, as 0088: tag.fandom_id is declared
index=True but no migration ever created that index.
Deliberately NOT here: image_record.sha256. The model says unique=True;
0001 created a plain index. Duplicates are possible today and the ORM
believes otherwise. The fix depends on whether duplicates already exist
— if they do, that is a dedupe decision, not a constraint — so it waits
on an answer about live data.
The real severity of #3275 is not the squash. It is that --autogenerate
has been unsafe on this project: run against the old models it would
have proposed dropping eleven indexes and two uniqueness guarantees.
260 lines
12 KiB
Python
260 lines
12 KiB
Python
"""MLSettings — single-row table holding ML pipeline tunables."""
|
|
|
|
from datetime import datetime
|
|
|
|
from sqlalchemy import (
|
|
Boolean,
|
|
CheckConstraint,
|
|
DateTime,
|
|
Float,
|
|
Integer,
|
|
String,
|
|
func,
|
|
select,
|
|
)
|
|
from sqlalchemy.orm import Mapped, mapped_column
|
|
|
|
from .base import Base
|
|
|
|
|
|
class MLSettings(Base):
|
|
__tablename__ = "ml_settings"
|
|
# Bare name — Base.metadata's naming convention prepends ck_<table>_,
|
|
# producing the final ck_ml_settings_singleton (matches migration 0003).
|
|
__table_args__ = (CheckConstraint("id = 1", name="singleton"),)
|
|
|
|
id: Mapped[int] = mapped_column(Integer, primary_key=True)
|
|
# CPU whole-image embedding (B3, operator 2026-07-02). The ml-worker's ONLY
|
|
# processing role is the embed fallback for stacks WITHOUT a GPU agent — ON
|
|
# by default so a fresh install works with no agent. Stacks that run the
|
|
# agent and drop the ml-worker container turn this OFF so import hooks stop
|
|
# queueing embed work nothing will consume (the daily GPU 'embed' backfill
|
|
# covers those images instead).
|
|
cpu_embed_enabled: Mapped[bool] = mapped_column(
|
|
Boolean, nullable=False, default=True,
|
|
server_default="true",
|
|
)
|
|
# Video embedding (#747). Sample one frame every N seconds (fixed CADENCE, not
|
|
# a fixed count) so coverage reflects real screen time regardless of length;
|
|
# cap the total so a long video can't explode into hundreds of embeds. The
|
|
# per-frame SigLIP embeddings are mean-pooled. Operator-tunable.
|
|
video_frame_interval_seconds: Mapped[float] = mapped_column(
|
|
Float, nullable=False, default=4.0,
|
|
server_default="4",
|
|
)
|
|
video_max_frames: Mapped[int] = mapped_column(
|
|
Integer, nullable=False, default=64,
|
|
server_default="64",
|
|
)
|
|
# Tagging-v2 head training (#114). The head is the suggestion source that
|
|
# LEARNS from the operator's tags (replacing Camie + centroid). A concept
|
|
# needs >= head_min_positives labelled images before a head is trained;
|
|
# head_auto_apply_precision is the precision bar a head must clear (at some
|
|
# operating point) to "graduate" into earned auto-apply. Operator-tunable.
|
|
head_min_positives: Mapped[int] = mapped_column(
|
|
Integer, nullable=False, default=8,
|
|
server_default="8",
|
|
)
|
|
head_auto_apply_precision: Mapped[float] = mapped_column(
|
|
Float, nullable=False, default=0.97,
|
|
server_default="0.97",
|
|
)
|
|
# Earned auto-apply (#114). A graduated head fires (tags images without a
|
|
# human) when this master switch is on AND the head has at least
|
|
# head_auto_apply_min_positives clean labels — so a precise-looking but
|
|
# under-supported low-N head can't spray tags across the library. ON by
|
|
# default (operator-asked 2026-06-29: opt-OUT, not opt-in); the support +
|
|
# measured-precision gates keep it safe, and every auto-tag is reversible.
|
|
head_auto_apply_enabled: Mapped[bool] = mapped_column(
|
|
Boolean, nullable=False, default=True,
|
|
server_default="true",
|
|
)
|
|
head_auto_apply_min_positives: Mapped[int] = mapped_column(
|
|
# Support floor raised 30→50 (operator-asked 2026-07-06): a head needs
|
|
# more human labels before it may fire without a human.
|
|
Integer, nullable=False, default=50,
|
|
server_default="30",
|
|
)
|
|
# CCIP character-match cosine cut (#114). 0.85 default — the v1 flat 0.75
|
|
# over-fired (high-reference characters matched a scatter of images); 0.85
|
|
# keeps the confident single-character matches. Tunable from the agent card.
|
|
ccip_match_threshold: Mapped[float] = mapped_column(
|
|
Float, nullable=False, default=0.85,
|
|
server_default="0.85",
|
|
)
|
|
# CCIP auto-apply (#114). Confident matches (>= ccip_auto_apply_threshold,
|
|
# above the suggest cut) auto-tag on a daily sweep. ON by default (opt-out);
|
|
# single-character references + the high bar keep it safe, every tag reversible.
|
|
ccip_auto_apply_enabled: Mapped[bool] = mapped_column(
|
|
Boolean, nullable=False, default=True,
|
|
server_default="true",
|
|
)
|
|
ccip_auto_apply_threshold: Mapped[float] = mapped_column(
|
|
# Raised 0.92→0.95 (operator-asked 2026-07-06) so only very confident
|
|
# character matches auto-tag.
|
|
Float, nullable=False, default=0.95,
|
|
server_default="0.92",
|
|
)
|
|
# -- Presentation chrome auto-hide (#141) -------------------------------
|
|
# `banner` (chrome — clusters on UI, not content) auto-applies on the sweep
|
|
# with its OWN flat threshold (decoupled from content-head graduation) and is
|
|
# HIDDEN from the gallery. Hiding is consequential so it runs HIGH. When an
|
|
# image would be auto-hidden but ALSO scores >= presentation_conflict_threshold
|
|
# on a content head, it's still hidden but flagged for review
|
|
# (PresentationReview, mode='chrome') instead of buried silently. ON by default
|
|
# (opt-out); every auto-tag is reversible. NOTE (#1464): `wip` + `editor
|
|
# screenshot` are no longer chrome — they went to the PROCESS path below.
|
|
presentation_auto_apply_enabled: Mapped[bool] = mapped_column(
|
|
Boolean, nullable=False, default=True,
|
|
server_default="true",
|
|
)
|
|
presentation_auto_apply_threshold: Mapped[float] = mapped_column(
|
|
Float, nullable=False, default=0.90,
|
|
server_default="0.90",
|
|
)
|
|
presentation_conflict_threshold: Mapped[float] = mapped_column(
|
|
Float, nullable=False, default=0.50,
|
|
server_default="0.50",
|
|
)
|
|
# -- Process auto-apply (#1464) ----------------------------------------
|
|
# `wip` / `editor screenshot` are PROCESS art — unfinished pieces + program
|
|
# screenshots that must stay OUT of head/CCIP training but, unlike chrome,
|
|
# remain VISIBLE in the gallery (operator 2026-07-12). They auto-apply on the
|
|
# sweep with their OWN flat threshold and a PROVISIONAL source (`process_auto`,
|
|
# in training_data._AUTO_SOURCES) so the head NEVER trains on its own output —
|
|
# it learns only from title (`wip_title`) + manual labels, which breaks the
|
|
# runaway loop. When a process tag would be applied but the image ALSO scores
|
|
# >= process_conflict_threshold on a content head, it's flagged for review
|
|
# (PresentationReview, mode='process') rather than silently marked. OFF by
|
|
# default — a new whole-library auto-tagger is opt-in; every auto-tag reversible.
|
|
process_auto_apply_enabled: Mapped[bool] = mapped_column(
|
|
Boolean, nullable=False, default=False,
|
|
server_default="false",
|
|
)
|
|
process_auto_apply_threshold: Mapped[float] = mapped_column(
|
|
Float, nullable=False, default=0.90,
|
|
server_default="0.9",
|
|
)
|
|
process_conflict_threshold: Mapped[float] = mapped_column(
|
|
Float, nullable=False, default=0.50,
|
|
server_default="0.5",
|
|
)
|
|
# Default = SigLIP 2 (so400m, 512px) for new installs (migration 0069);
|
|
# existing libraries keep their stored value until the operator re-embeds.
|
|
embedder_model_version: Mapped[str] = mapped_column(
|
|
String(128), nullable=False, default="siglip2-so400m-patch16-512",
|
|
server_default="siglip2-so400m-patch16-512",
|
|
)
|
|
# The HF model NAME the embedder loads (server CPU embed + announced to the
|
|
# GPU agent in the lease). Operator-settable so the embedder is a choice, not
|
|
# a hardcode (#1190): set name + version together, then re-embed + retrain.
|
|
embedder_model_name: Mapped[str] = mapped_column(
|
|
String(128), nullable=False, default="google/siglip2-so400m-patch16-512",
|
|
server_default="google/siglip2-so400m-patch16-512",
|
|
)
|
|
# -- Crop proposers / detectors (#1202, #134) --------------------------
|
|
# WHERE-to-crop YOLO detectors feeding the crop→SigLIP bag + CCIP. Config
|
|
# lives HERE (DB) and is announced to the GPU agent in the lease — same as
|
|
# the embedder model — so it is UI-tunable with NO restart, and the agent's
|
|
# env is bootstrap-only. Each weights spec is an ultralytics builtin name,
|
|
# an http(s) URL, or "hf_repo::file" (agent's _resolve). enabled off (or an
|
|
# empty weights) skips that proposer. All ON by default (operator 2026-07-05)
|
|
# so a fresh install crops out-of-the-box.
|
|
# person: general COCO figure detector for Western/realistic art the anime
|
|
# person-detector misses → NMS-merged with imgutils → CCIP + concept.
|
|
detector_person_enabled: Mapped[bool] = mapped_column(
|
|
Boolean, nullable=False, default=True,
|
|
server_default="true",
|
|
)
|
|
detector_person_weights: Mapped[str] = mapped_column(
|
|
String(512), nullable=False, default="yolo11n.pt",
|
|
server_default="yolo11n.pt",
|
|
)
|
|
detector_person_conf: Mapped[float] = mapped_column(
|
|
Float, nullable=False, default=0.35,
|
|
server_default="0.35",
|
|
)
|
|
# anatomy: booru_yolo anime/furry/NSFW torso components → concept crops.
|
|
# Default = yolov11m_aa22 (26 classes, best mAP50-95 0.96), committed in the
|
|
# upstream repo so the URL resolves. License UNSTATED — fine for a private
|
|
# homelab (operator accepted #1202).
|
|
detector_anatomy_enabled: Mapped[bool] = mapped_column(
|
|
Boolean, nullable=False, default=True,
|
|
server_default="true",
|
|
)
|
|
detector_anatomy_weights: Mapped[str] = mapped_column(
|
|
String(512), nullable=False,
|
|
default=(
|
|
"https://github.com/aperveyev/booru_yolo/raw/main/models/"
|
|
"yolov11m_aa22.pt"
|
|
),
|
|
server_default="https://github.com/aperveyev/booru_yolo/raw/main/models/yolov11m_aa22.pt",
|
|
)
|
|
detector_anatomy_conf: Mapped[float] = mapped_column(
|
|
Float, nullable=False, default=0.30,
|
|
server_default="0.30",
|
|
)
|
|
# panel: comic page → panel regions → concept crops (Apache-2.0, YOLOv12x).
|
|
detector_panel_enabled: Mapped[bool] = mapped_column(
|
|
Boolean, nullable=False, default=True,
|
|
server_default="true",
|
|
)
|
|
detector_panel_weights: Mapped[str] = mapped_column(
|
|
String(512), nullable=False,
|
|
default="mosesb/best-comic-panel-detection::best.pt",
|
|
server_default="mosesb/best-comic-panel-detection::best.pt",
|
|
)
|
|
detector_panel_conf: Mapped[float] = mapped_column(
|
|
Float, nullable=False, default=0.30,
|
|
server_default="0.30",
|
|
)
|
|
# Per-frame caps bound the crop→embed explosion; max_regions is the hard
|
|
# per-job backstop; dedupe_iou drops near-duplicate crops before the embed.
|
|
detector_max_figures: Mapped[int] = mapped_column(
|
|
Integer, nullable=False, default=8,
|
|
server_default="8",
|
|
)
|
|
detector_max_components: Mapped[int] = mapped_column(
|
|
Integer, nullable=False, default=8,
|
|
server_default="8",
|
|
)
|
|
detector_max_panels: Mapped[int] = mapped_column(
|
|
Integer, nullable=False, default=8,
|
|
server_default="8",
|
|
)
|
|
detector_max_regions: Mapped[int] = mapped_column(
|
|
Integer, nullable=False, default=128,
|
|
server_default="128",
|
|
)
|
|
detector_dedupe_iou: Mapped[float] = mapped_column(
|
|
Float, nullable=False, default=0.85,
|
|
server_default="0.85",
|
|
)
|
|
# -- CCIP character prototypes (#1317) ---------------------------------
|
|
# The per-character reference set is precomputed + refreshed INCREMENTALLY
|
|
# (services.ml.character_prototypes) instead of rebuilt on the request path.
|
|
# ccip_ref_signature is the cheap GLOBAL gate — when it's unchanged the
|
|
# refresh no-ops; ccip_prototype_cap bounds the reference vectors kept per
|
|
# character so MATCH cost doesn't grow with a character's popularity.
|
|
ccip_ref_signature: Mapped[str | None] = mapped_column(
|
|
String(128), nullable=True
|
|
)
|
|
ccip_prototype_cap: Mapped[int] = mapped_column(
|
|
Integer, nullable=False, default=64,
|
|
server_default="64",
|
|
)
|
|
updated_at: Mapped[datetime] = mapped_column(
|
|
DateTime(timezone=True), nullable=False, server_default=func.now()
|
|
)
|
|
|
|
@classmethod
|
|
async def load(cls, session) -> MLSettings:
|
|
"""The singleton settings row (id=1), via an async session. Mirrors
|
|
ImportSettings.load — the shared singleton-loader pattern."""
|
|
return (await session.execute(select(cls).where(cls.id == 1))).scalar_one()
|
|
|
|
@classmethod
|
|
def load_sync(cls, session) -> MLSettings:
|
|
"""The singleton settings row (id=1), via a sync session."""
|
|
return session.execute(select(cls).where(cls.id == 1)).scalar_one()
|