Compare commits

..
4 Commits
Author SHA1 Message Date
bvandeusenandClaude Opus 4.8 099e1e664c refactor(patreon): DRY the campaigns-API request (#161)
CI / lint (push) Failing after 2s
CI / backend-lint-and-test (push) Successful in 28s
CI / frontend-build (push) Successful in 29s
CI / integration (push) Successful in 4m0s
_lookup_via_api and resolve_display_name shared ~90% of their body (same
endpoint, params, headers, error handling — differing only in which field
they pluck from data[0]). Extract _campaigns_api_first(vanity, cookies_path)
-> dict|None; callers pluck the campaign id vs the display name. Return-value
behavior preserved (the display-name path additionally gains the helper's more
granular warning logs). Covered by the existing test_patreon_resolver.py.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NsmJSQxnNxGgtM5Yz4GAqi
2026-07-13 21:41:34 -04:00
bvandeusenandClaude Opus 4.8 c87f8a1bb3 refactor(maintenance): DRY the recover_stalled head-run twins (#161)
recover_stalled_head_training_runs and recover_stalled_head_auto_apply_runs
were near-exact copies (coalesce-flip + keep-last-N prune, differing only in
model + two constants). Extract _recover_stalled_runs(model, stall_minutes,
keep_runs, label); the two tasks become thin wrappers. The other two recover
tasks are deliberately NOT folded in (library-audit has no prune tail; backup
uses a single started_at cutoff). test_recover_stalled_head_runs.py covers
both wrappers (stalled→error, fresh survives) — previously untested.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NsmJSQxnNxGgtM5Yz4GAqi
2026-07-13 21:41:34 -04:00
bvandeusenandClaude Opus 4.8 666b3a2ec8 refactor(ml): DRY pass — shared sweep helpers + table-driven settings (#161)
Consolidate duplication accrued across the ML tagging + settings backend,
behavior-preserving (over-DRY guard applied — the three auto-apply sweep
BODIES stay separate; only their shared inner helpers are extracted).

- _sigmoid / _conflict_scores / _insert_presentation_review (heads.py): the
  score→prob transform (6 inlined sites), the presentation conflict signal
  (2 sites), and the ring-loud PresentationReview insert (2 sites, single-
  sourced so the mode column can't drift on the shared composite PK).
- _applied_or_rejected (training_data.py): the per-tag "applied ∪ rejected"
  skip-set, byte-identical at 3 sweep sites (heads.py x2, tasks/ml.py ccip).
- ccip sweep divergence fixes: import ccip._FIGURE_KINDS + training_data._l2norm
  instead of local copies that silently drift when the canonical changes.
- MLSettings.load / .load_sync classmethods (mirror ImportSettings); route all
  8 scalar_one singleton reads through them (the session.get None-path stays).
- GET serializers for MLSettings + ImportSettings are now table-driven off the
  same _EDITABLE tuples PATCH writes, so a new field can't be silently absent
  from GET (the split that historically dropped fields).
- AUTO_APPLY_THRESHOLD_MIN/MAX constant single-sources the [0.5,0.999] operating
  range across the service clamp + the 5 API validators.
- test_ml_dry_helpers.py pins _applied_or_rejected + _sigmoid.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NsmJSQxnNxGgtM5Yz4GAqi
2026-07-13 21:41:24 -04:00
bvandeusenandClaude Opus 4.8 d80a5255ed feat(extension): in-app update prompt — popup banner + toolbar badge (#1489)
CI / lint (push) Successful in 3s
extension / lint (push) Successful in 10s
CI / frontend-build (push) Successful in 24s
CI / backend-lint-and-test (push) Successful in 39s
CI / integration (push) Successful in 3m51s
extension / lint (pull_request) Successful in 10s
The extension is installed per-instance from the operator's FC host, so Firefox's
static update_url can't apply (each instance has a different host) and updates
were fully manual. Add a self-hosted-friendly update surface that reuses the
existing public GET /api/extension/manifest ({version, latest_url, sha256}):

- lib/api.js: getExtensionManifest().
- background.js: checkForUpdateInfo() compares the instance's latest published
  version against runtime.getManifest().version (dotted-numeric compare so
  1.0.10 > 1.0.9); CHECK_UPDATE message handler; refreshUpdateBadge() sets a
  toolbar badge via browser.action; a daily browser.alarms check plus on
  startup/installed. New 'alarms' permission (non-prompting).
- popup: an 'Update available — vX' banner with an Update button that opens the
  signed XPI (web root, /api stripped like OPEN_ARTIST_PAGE) → Firefox's native
  install prompt. Never blocks the popup on a failed check.

No backend changes (endpoint already exists). Bump 1.0.8→1.0.9 so this ships;
from here on updates surface themselves instead of needing a manual reinstall.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-13 18:28:06 -04:00
19 changed files with 428 additions and 283 deletions
+1 -3
View File
@@ -256,9 +256,7 @@ async def lease():
if not await _agent_authed(session): if not await _agent_authed(session):
return jsonify({"error": "unauthorized"}), 401 return jsonify({"error": "unauthorized"}), 401
jobs = await GpuJobService(session).lease(agent_id, batch_size=batch) jobs = await GpuJobService(session).lease(agent_id, batch_size=batch)
ml = ( ml = await MLSettings.load(session)
await session.execute(select(MLSettings).where(MLSettings.id == 1))
).scalar_one()
# image rows for url/mime in one shot # image rows for url/mime in one shot
ids = [j.image_record_id for j in jobs] ids = [j.image_record_id for j in jobs]
imgs = { imgs = {
+17 -43
View File
@@ -4,6 +4,7 @@ from quart import Blueprint, jsonify, request
from ..extensions import get_session from ..extensions import get_session
from ..models import MLSettings from ..models import MLSettings
from ..services.ml.heads import AUTO_APPLY_THRESHOLD_MAX, AUTO_APPLY_THRESHOLD_MIN
ml_admin_bp = Blueprint("ml_admin", __name__, url_prefix="/api/ml") ml_admin_bp = Blueprint("ml_admin", __name__, url_prefix="/api/ml")
@@ -83,48 +84,21 @@ async def embedder_models():
@ml_admin_bp.route("/settings", methods=["GET"]) @ml_admin_bp.route("/settings", methods=["GET"])
async def get_settings(): async def get_settings():
from sqlalchemy import select
async with get_session() as session: async with get_session() as session:
s = ( s = await MLSettings.load(session)
await session.execute(select(MLSettings).where(MLSettings.id == 1)) # Table-driven off _EDITABLE (which PATCH also writes) so a new settings field
).scalar_one() # can never be silently absent from GET — the split that historically dropped
return jsonify( # fields. _EDITABLE already includes *_DETECTOR_FIELDS.
{ return jsonify({f: getattr(s, f) for f in _EDITABLE})
"cpu_embed_enabled": s.cpu_embed_enabled,
"video_frame_interval_seconds": s.video_frame_interval_seconds,
"video_max_frames": s.video_max_frames,
"embedder_model_version": s.embedder_model_version,
"head_min_positives": s.head_min_positives,
"head_auto_apply_precision": s.head_auto_apply_precision,
"head_auto_apply_enabled": s.head_auto_apply_enabled,
"head_auto_apply_min_positives": s.head_auto_apply_min_positives,
"ccip_match_threshold": s.ccip_match_threshold,
"ccip_auto_apply_enabled": s.ccip_auto_apply_enabled,
"ccip_auto_apply_threshold": s.ccip_auto_apply_threshold,
"presentation_auto_apply_enabled": s.presentation_auto_apply_enabled,
"presentation_auto_apply_threshold": s.presentation_auto_apply_threshold,
"presentation_conflict_threshold": s.presentation_conflict_threshold,
"process_auto_apply_enabled": s.process_auto_apply_enabled,
"process_auto_apply_threshold": s.process_auto_apply_threshold,
"process_conflict_threshold": s.process_conflict_threshold,
"embedder_model_name": s.embedder_model_name,
**{f: getattr(s, f) for f in _DETECTOR_FIELDS},
}
)
@ml_admin_bp.route("/settings", methods=["PATCH"]) @ml_admin_bp.route("/settings", methods=["PATCH"])
async def patch_settings(): async def patch_settings():
from sqlalchemy import select
body = await request.get_json() body = await request.get_json()
if not isinstance(body, dict): if not isinstance(body, dict):
return jsonify({"error": "body must be an object"}), 400 return jsonify({"error": "body must be an object"}), 400
async with get_session() as session: async with get_session() as session:
s = ( s = await MLSettings.load(session)
await session.execute(select(MLSettings).where(MLSettings.id == 1))
).scalar_one()
# Merge the patch over current values, then validate the result as a # Merge the patch over current values, then validate the result as a
# whole — the store-floor invariant couples three fields, so they # whole — the store-floor invariant couples three fields, so they
@@ -154,24 +128,24 @@ def _validate(p: dict) -> str | None:
# Head training (#114). # Head training (#114).
if int(p["head_min_positives"]) < 1: if int(p["head_min_positives"]) < 1:
return "head_min_positives must be >= 1" return "head_min_positives must be >= 1"
if not (0.5 <= float(p["head_auto_apply_precision"]) <= 0.999): if not (AUTO_APPLY_THRESHOLD_MIN <= float(p["head_auto_apply_precision"]) <= AUTO_APPLY_THRESHOLD_MAX):
return "head_auto_apply_precision must be between 0.5 and 0.999" return f"head_auto_apply_precision must be between {AUTO_APPLY_THRESHOLD_MIN} and {AUTO_APPLY_THRESHOLD_MAX}"
if int(p["head_auto_apply_min_positives"]) < 1: if int(p["head_auto_apply_min_positives"]) < 1:
return "head_auto_apply_min_positives must be >= 1" return "head_auto_apply_min_positives must be >= 1"
if not (0.5 <= float(p["ccip_match_threshold"]) <= 0.999): if not (AUTO_APPLY_THRESHOLD_MIN <= float(p["ccip_match_threshold"]) <= AUTO_APPLY_THRESHOLD_MAX):
return "ccip_match_threshold must be between 0.5 and 0.999" return f"ccip_match_threshold must be between {AUTO_APPLY_THRESHOLD_MIN} and {AUTO_APPLY_THRESHOLD_MAX}"
if not (0.5 <= float(p["ccip_auto_apply_threshold"]) <= 0.999): if not (AUTO_APPLY_THRESHOLD_MIN <= float(p["ccip_auto_apply_threshold"]) <= AUTO_APPLY_THRESHOLD_MAX):
return "ccip_auto_apply_threshold must be between 0.5 and 0.999" return f"ccip_auto_apply_threshold must be between {AUTO_APPLY_THRESHOLD_MIN} and {AUTO_APPLY_THRESHOLD_MAX}"
# Presentation chrome auto-hide (#141). Auto-apply runs high (hiding is # Presentation chrome auto-hide (#141). Auto-apply runs high (hiding is
# consequential); the conflict cut is a plain probability [0,1]. # consequential); the conflict cut is a plain probability [0,1].
if not (0.5 <= float(p["presentation_auto_apply_threshold"]) <= 0.999): if not (AUTO_APPLY_THRESHOLD_MIN <= float(p["presentation_auto_apply_threshold"]) <= AUTO_APPLY_THRESHOLD_MAX):
return "presentation_auto_apply_threshold must be between 0.5 and 0.999" return f"presentation_auto_apply_threshold must be between {AUTO_APPLY_THRESHOLD_MIN} and {AUTO_APPLY_THRESHOLD_MAX}"
if not (0.0 <= float(p["presentation_conflict_threshold"]) <= 1.0): if not (0.0 <= float(p["presentation_conflict_threshold"]) <= 1.0):
return "presentation_conflict_threshold must be between 0 and 1" return "presentation_conflict_threshold must be between 0 and 1"
# Process auto-apply (#1464). wip/editor stay VISIBLE so a false apply is # Process auto-apply (#1464). wip/editor stay VISIBLE so a false apply is
# low-harm (excludes-from-training + a review flag), but keep the same bar. # low-harm (excludes-from-training + a review flag), but keep the same bar.
if not (0.5 <= float(p["process_auto_apply_threshold"]) <= 0.999): if not (AUTO_APPLY_THRESHOLD_MIN <= float(p["process_auto_apply_threshold"]) <= AUTO_APPLY_THRESHOLD_MAX):
return "process_auto_apply_threshold must be between 0.5 and 0.999" return f"process_auto_apply_threshold must be between {AUTO_APPLY_THRESHOLD_MIN} and {AUTO_APPLY_THRESHOLD_MAX}"
if not (0.0 <= float(p["process_conflict_threshold"]) <= 1.0): if not (0.0 <= float(p["process_conflict_threshold"]) <= 1.0):
return "process_conflict_threshold must be between 0 and 1" return "process_conflict_threshold must be between 0 and 1"
# Embedder model swap (#1190): both must be non-empty. Changing them means a # Embedder model swap (#1190): both must be non-empty. Changing them means a
+3 -28
View File
@@ -66,34 +66,9 @@ _EXTDL_TOGGLE_FIELDS = (
async def get_import_settings(): async def get_import_settings():
async with get_session() as session: async with get_session() as session:
row = await ImportSettings.load(session) row = await ImportSettings.load(session)
return jsonify({ # Table-driven off _EDITABLE_FIELDS (which PATCH also writes) so a new field
"min_width": row.min_width, # can't be silently absent from GET.
"min_height": row.min_height, return jsonify({f: getattr(row, f) for f in _EDITABLE_FIELDS})
"skip_transparent": row.skip_transparent,
"transparency_threshold": row.transparency_threshold,
"skip_single_color": row.skip_single_color,
"single_color_threshold": row.single_color_threshold,
"single_color_tolerance": row.single_color_tolerance,
"phash_threshold": row.phash_threshold,
"download_rate_limit_seconds": row.download_rate_limit_seconds,
"download_validate_files": row.download_validate_files,
"download_schedule_default_seconds": row.download_schedule_default_seconds,
"download_event_retention_days": row.download_event_retention_days,
"download_failure_warning_threshold": row.download_failure_warning_threshold,
"series_suggest_enabled": row.series_suggest_enabled,
"series_suggest_threshold": row.series_suggest_threshold,
"extdl_mega_enabled": row.extdl_mega_enabled,
"extdl_gdrive_enabled": row.extdl_gdrive_enabled,
"extdl_mediafire_enabled": row.extdl_mediafire_enabled,
"extdl_dropbox_enabled": row.extdl_dropbox_enabled,
"extdl_pixeldrain_enabled": row.extdl_pixeldrain_enabled,
"translation_enabled": row.translation_enabled,
"interpreter_base_url": row.interpreter_base_url,
"translation_target_lang": row.translation_target_lang,
"translation_min_confidence": row.translation_min_confidence,
"wip_title_tagging_enabled": row.wip_title_tagging_enabled,
"wip_soft_title_tagging_enabled": row.wip_soft_title_tagging_enabled,
})
@settings_bp.route("/settings/import", methods=["PATCH"]) @settings_bp.route("/settings/import", methods=["PATCH"])
+12
View File
@@ -10,6 +10,7 @@ from sqlalchemy import (
Integer, Integer,
String, String,
func, func,
select,
) )
from sqlalchemy.orm import Mapped, mapped_column from sqlalchemy.orm import Mapped, mapped_column
@@ -212,3 +213,14 @@ class MLSettings(Base):
updated_at: Mapped[datetime] = mapped_column( updated_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now() DateTime(timezone=True), nullable=False, server_default=func.now()
) )
@classmethod
async def load(cls, session) -> "MLSettings":
"""The singleton settings row (id=1), via an async session. Mirrors
ImportSettings.load — the shared singleton-loader pattern."""
return (await session.execute(select(cls).where(cls.id == 1))).scalar_one()
@classmethod
def load_sync(cls, session) -> "MLSettings":
"""The singleton settings row (id=1), via a sync session."""
return session.execute(select(cls).where(cls.id == 1)).scalar_one()
@@ -150,9 +150,7 @@ def refresh_character_prototypes(
"""Incrementally refresh the prototype store. `full=True` rebuilds every """Incrementally refresh the prototype store. `full=True` rebuilds every
character regardless of the gate/fingerprints (nightly reconcile). Returns character regardless of the gate/fingerprints (nightly reconcile). Returns
{skipped, rebuilt, removed}; commits.""" {skipped, rebuilt, removed}; commits."""
settings = session.execute( settings = MLSettings.load_sync(session)
select(MLSettings).where(MLSettings.id == 1)
).scalar_one()
sig = _global_signature(session) sig = _global_signature(session)
if not full and settings.ccip_ref_signature == sig: if not full and settings.ccip_ref_signature == sig:
return {"skipped": True, "rebuilt": 0, "removed": 0} return {"skipped": True, "rebuilt": 0, "removed": 0}
@@ -204,9 +202,7 @@ def retract_auto_applied_ccip(session: Session) -> int:
n_retracted.""" n_retracted."""
import numpy as np import numpy as np
settings = session.execute( settings = MLSettings.load_sync(session)
select(MLSettings).where(MLSettings.id == 1)
).scalar_one()
if not settings.ccip_auto_apply_enabled: if not settings.ccip_auto_apply_enabled:
return 0 return 0
thr = float(settings.ccip_auto_apply_threshold) thr = float(settings.ccip_auto_apply_threshold)
+65 -62
View File
@@ -23,6 +23,7 @@ from datetime import UTC, datetime
from typing import Any from typing import Any
from sqlalchemy import delete, exists, func, select from sqlalchemy import delete, exists, func, select
from sqlalchemy.dialects.postgresql import insert as pg_insert
from sqlalchemy.ext.asyncio import AsyncSession from sqlalchemy.ext.asyncio import AsyncSession
from sqlalchemy.orm import Session from sqlalchemy.orm import Session
@@ -42,6 +43,7 @@ from ...models import (
from ...models.tag import CHROME_SYSTEM_TAGS, PROCESS_SYSTEM_TAGS, image_tag from ...models.tag import CHROME_SYSTEM_TAGS, PROCESS_SYSTEM_TAGS, image_tag
from .training_data import ( from .training_data import (
_AUTO_SOURCES, _AUTO_SOURCES,
_applied_or_rejected,
_auto_apply_point, _auto_apply_point,
_hygiene_excluded_ids, _hygiene_excluded_ids,
_ids_with_tag, _ids_with_tag,
@@ -61,6 +63,14 @@ MIN_POSITIVES_FLOOR = 8 # hard floor; settings.head_min_positives can raise
_UNLABELED_POOL = 4000 _UNLABELED_POOL = 4000
_EXAMPLES_MIN = 8 # need at least this many embedded +/- to fit a head _EXAMPLES_MIN = 8 # need at least this many embedded +/- to fit a head
# Auto-apply / match confidence operating range. Every graduated auto-apply or
# CCIP-match threshold the operator can set lives in this band, and the head
# precision target is clamped to it: below 0.5 "auto-apply" is meaningless, and
# 1.0 is unachievable so 0.999 is the ceiling. One source shared by the service
# clamp (_normalize_params) and the API validator (ml_admin._validate).
AUTO_APPLY_THRESHOLD_MIN = 0.5
AUTO_APPLY_THRESHOLD_MAX = 0.999
# Only these tag kinds get heads (the surfaced suggestion categories). # Only these tag kinds get heads (the surfaced suggestion categories).
_HEAD_KINDS = (TagKind.general, TagKind.character) _HEAD_KINDS = (TagKind.general, TagKind.character)
# tag.kind -> the suggestion category the rail groups under. # tag.kind -> the suggestion category the rail groups under.
@@ -78,6 +88,38 @@ _CATEGORY = {TagKind.general: "general", TagKind.character: "character"}
_SYSTEM_TAG_SUGGEST_FLOOR = 0.65 _SYSTEM_TAG_SUGGEST_FLOOR = 0.65
def _sigmoid(z, np):
"""Logistic sigmoid 1/(1+e^-z): the head score→probability transform. One home
for what was inlined at every scoring site (suggest, both sweeps, retract)."""
return 1.0 / (1.0 + np.exp(-z))
def _conflict_scores(Xn, Wc, bc, np):
"""The presentation conflict signal (#141): per row, the MAX content-head
probability and WHICH head produced it. Shared by the system-tag sweep's guard-2
and the soft-wip audit — both ask "does this ALSO look like real content?"."""
cprobs = _sigmoid(Xn @ Wc.T + bc, np)
return cprobs.max(axis=1), cprobs.argmax(axis=1)
def _insert_presentation_review(
session, *, image_record_id, tag_id, conflict_tag_id, conflict_score, mode,
):
"""Single-source the ring-loud PresentationReview row shape so the two writers
(system-tag sweep guard-2 + soft-wip audit) can't drift on columns or `mode` —
they share the (image_record_id, tag_id) composite PK, so a divergent `mode`
would be a silent first-writer-wins bug."""
session.execute(
pg_insert(PresentationReview)
.values(
image_record_id=image_record_id, tag_id=tag_id,
conflict_tag_id=conflict_tag_id, conflict_score=conflict_score,
mode=mode,
)
.on_conflict_do_nothing()
)
class HeadTrainingAlreadyRunning(Exception): class HeadTrainingAlreadyRunning(Exception):
"""Raised by start_head_training_run when a run is already in flight.""" """Raised by start_head_training_run when a run is already in flight."""
@@ -103,9 +145,7 @@ def start_head_training_run(session: Session, params: dict[str, Any]) -> int:
def _settings(session: Session) -> MLSettings: def _settings(session: Session) -> MLSettings:
return session.execute( return MLSettings.load_sync(session)
select(MLSettings).where(MLSettings.id == 1)
).scalar_one()
def _normalize_params(session: Session, params: dict[str, Any] | None) -> dict[str, Any]: def _normalize_params(session: Session, params: dict[str, Any] | None) -> dict[str, Any]:
@@ -124,7 +164,7 @@ def _normalize_params(session: Session, params: dict[str, Any] | None) -> dict[s
except (TypeError, ValueError): except (TypeError, ValueError):
cv_folds = DEFAULT_CV_FOLDS cv_folds = DEFAULT_CV_FOLDS
try: try:
precision_target = min(max(float(params.get("precision_target", s.head_auto_apply_precision)), 0.5), 0.999) precision_target = min(max(float(params.get("precision_target", s.head_auto_apply_precision)), AUTO_APPLY_THRESHOLD_MIN), AUTO_APPLY_THRESHOLD_MAX)
except (TypeError, ValueError): except (TypeError, ValueError):
precision_target = s.head_auto_apply_precision precision_target = s.head_auto_apply_precision
return { return {
@@ -536,7 +576,7 @@ async def score_image(
norms[norms == 0] = 1.0 norms[norms == 0] = 1.0
Xn = X / norms Xn = X / norms
Z = Xn @ heads["W"].T + heads["b"] # (B, H) Z = Xn @ heads["W"].T + heads["b"] # (B, H)
probs_bag = 1.0 / (1.0 + np.exp(-Z)) # (B, H) probs_bag = _sigmoid(Z, np) # (B, H)
probs = probs_bag.max(axis=0) # (H,) best over the bag probs = probs_bag.max(axis=0) # (H,) best over the bag
# ARGMAX beside the max: WHICH bag row won each head → the region that grounds # ARGMAX beside the max: WHICH bag row won each head → the region that grounds
# the tag (bag_meta[win]); None when the whole-image vector won (#1206). # the tag (bag_meta[win]); None when the whole-image vector won (#1206).
@@ -614,9 +654,7 @@ async def ground_applied_tag(
async def _settings_async(session: AsyncSession) -> MLSettings: async def _settings_async(session: AsyncSession) -> MLSettings:
return ( return await MLSettings.load(session)
await session.execute(select(MLSettings).where(MLSettings.id == 1))
).scalar_one()
# --- Earned auto-apply (sync, ml worker) --------------------------------- # --- Earned auto-apply (sync, ml worker) ---------------------------------
@@ -687,7 +725,6 @@ def auto_apply_sweep(
embeddings in chunks; commits per chunk on a real run. Returns embeddings in chunks; commits per chunk on a real run. Returns
{n_applied, concepts:[{tag_id,name,applied,scanned,threshold}]}.""" {n_applied, concepts:[{tag_id,name,applied,scanned,threshold}]}."""
import numpy as np import numpy as np
from sqlalchemy.dialects.postgresql import insert as pg_insert
settings = _settings(session) settings = _settings(session)
rows = _auto_apply_heads( rows = _auto_apply_heads(
@@ -704,18 +741,7 @@ def auto_apply_sweep(
names = [r.name for r in rows] names = [r.name for r in rows]
# Skip images that already carry, or have rejected, each tag. # Skip images that already carry, or have rejected, each tag.
skip = {tid: set() for tid in tag_ids} skip = _applied_or_rejected(session, tag_ids)
for tid in tag_ids:
for (iid,) in session.execute(
select(image_tag.c.image_record_id).where(image_tag.c.tag_id == tid)
):
skip[tid].add(iid)
for (iid,) in session.execute(
select(TagSuggestionRejection.image_record_id).where(
TagSuggestionRejection.tag_id == tid
)
):
skip[tid].add(iid)
applied = [0] * len(rows) applied = [0] * len(rows)
scanned = 0 scanned = 0
@@ -729,7 +755,7 @@ def auto_apply_sweep(
if not cids: if not cids:
continue continue
Xn = _l2norm(np.vstack([emb[i] for i in cids]).astype(np.float32), np) Xn = _l2norm(np.vstack([emb[i] for i in cids]).astype(np.float32), np)
probs = 1.0 / (1.0 + np.exp(-(Xn @ W.T + b))) # (N, H) probs = _sigmoid(Xn @ W.T + b, np) # (N, H)
scanned += len(cids) scanned += len(cids)
for h in range(len(rows)): for h in range(len(rows)):
tid = tag_ids[h] tid = tag_ids[h]
@@ -840,7 +866,6 @@ def system_tag_auto_apply_sweep(
enabled flag is set. numpy-only (no sklearn). Returns {n_applied, n_flagged, enabled flag is set. numpy-only (no sklearn). Returns {n_applied, n_flagged,
concepts}.""" concepts}."""
import numpy as np import numpy as np
from sqlalchemy.dialects.postgresql import insert as pg_insert
cfg = _SWEEP_MODES[mode] cfg = _SWEEP_MODES[mode]
settings = _settings(session) settings = _settings(session)
@@ -869,18 +894,7 @@ def system_tag_auto_apply_sweep(
valued = _valued_image_ids(session) valued = _valued_image_ids(session)
# Skip images that already carry, or have rejected, each presentation tag. # Skip images that already carry, or have rejected, each presentation tag.
skip = {tid: set() for tid in pres_tag_ids} skip = _applied_or_rejected(session, pres_tag_ids)
for tid in pres_tag_ids:
for (iid,) in session.execute(
select(image_tag.c.image_record_id).where(image_tag.c.tag_id == tid)
):
skip[tid].add(iid)
for (iid,) in session.execute(
select(TagSuggestionRejection.image_record_id).where(
TagSuggestionRejection.tag_id == tid
)
):
skip[tid].add(iid)
applied = [0] * len(pres) applied = [0] * len(pres)
n_flagged = 0 n_flagged = 0
@@ -895,11 +909,9 @@ def system_tag_auto_apply_sweep(
if not cids: if not cids:
continue continue
Xn = _l2norm(np.vstack([emb[i] for i in cids]).astype(np.float32), np) Xn = _l2norm(np.vstack([emb[i] for i in cids]).astype(np.float32), np)
probs = 1.0 / (1.0 + np.exp(-(Xn @ Wp.T + bp))) # (N, P) probs = _sigmoid(Xn @ Wp.T + bp, np) # (N, P)
if Wc is not None: if Wc is not None:
cprobs = 1.0 / (1.0 + np.exp(-(Xn @ Wc.T + bc))) # (N, C) max_c, arg_c = _conflict_scores(Xn, Wc, bc, np) # (N,), (N,)
max_c = cprobs.max(axis=1)
arg_c = cprobs.argmax(axis=1)
scanned += len(cids) scanned += len(cids)
for p in range(len(pres)): for p in range(len(pres)):
tid = pres_tag_ids[p] tid = pres_tag_ids[p]
@@ -924,15 +936,12 @@ def system_tag_auto_apply_sweep(
if Wc is not None and float(max_c[idx]) >= conflict_thr: if Wc is not None and float(max_c[idx]) >= conflict_thr:
n_flagged += 1 n_flagged += 1
if not dry_run: if not dry_run:
session.execute( _insert_presentation_review(
pg_insert(PresentationReview) session,
.values( image_record_id=iid, tag_id=tid,
image_record_id=iid, tag_id=tid, conflict_tag_id=conf_tag_ids[int(arg_c[idx])],
conflict_tag_id=conf_tag_ids[int(arg_c[idx])], conflict_score=float(max_c[idx]),
conflict_score=float(max_c[idx]), mode=mode,
mode=mode,
)
.on_conflict_do_nothing()
) )
if not dry_run: if not dry_run:
session.commit() session.commit()
@@ -956,7 +965,6 @@ def soft_wip_conflict_audit(session: Session, dry_run: bool = False) -> dict:
NOT remove the tag; the operator decides. No-op when there are no content heads. NOT remove the tag; the operator decides. No-op when there are no content heads.
numpy-only. Returns {n_scanned, n_flagged}.""" numpy-only. Returns {n_scanned, n_flagged}."""
import numpy as np import numpy as np
from sqlalchemy.dialects.postgresql import insert as pg_insert
from ..wip_title import WIP_TITLE_SOFT_SOURCE, resolve_wip_tag_id from ..wip_title import WIP_TITLE_SOFT_SOURCE, resolve_wip_tag_id
@@ -993,22 +1001,17 @@ def soft_wip_conflict_audit(session: Session, dry_run: bool = False) -> dict:
continue continue
scanned += len(cids) scanned += len(cids)
Xn = _l2norm(np.vstack([emb[i] for i in cids]).astype(np.float32), np) Xn = _l2norm(np.vstack([emb[i] for i in cids]).astype(np.float32), np)
cprobs = 1.0 / (1.0 + np.exp(-(Xn @ Wc.T + bc))) max_c, arg_c = _conflict_scores(Xn, Wc, bc, np)
max_c = cprobs.max(axis=1)
arg_c = cprobs.argmax(axis=1)
for k in range(len(cids)): for k in range(len(cids)):
if float(max_c[k]) >= conflict_thr: if float(max_c[k]) >= conflict_thr:
n_flagged += 1 n_flagged += 1
if not dry_run: if not dry_run:
session.execute( _insert_presentation_review(
pg_insert(PresentationReview) session,
.values( image_record_id=cids[k], tag_id=wip_id,
image_record_id=cids[k], tag_id=wip_id, conflict_tag_id=conf_tag_ids[int(arg_c[k])],
conflict_tag_id=conf_tag_ids[int(arg_c[k])], conflict_score=float(max_c[k]),
conflict_score=float(max_c[k]), mode="process",
mode="process",
)
.on_conflict_do_nothing()
) )
if not dry_run: if not dry_run:
session.commit() session.commit()
@@ -1062,7 +1065,7 @@ def retract_auto_applied_heads(session: Session) -> int:
continue continue
Xn = _l2norm(np.vstack([emb[i] for i in cids]).astype(np.float32), np) Xn = _l2norm(np.vstack([emb[i] for i in cids]).astype(np.float32), np)
w = np.asarray(weights, dtype=np.float32) w = np.asarray(weights, dtype=np.float32)
probs = 1.0 / (1.0 + np.exp(-(Xn @ w + float(bias)))) probs = _sigmoid(Xn @ w + float(bias), np)
below = [cids[k] for k in np.where(probs < float(thr))[0]] below = [cids[k] for k in np.where(probs < float(thr))[0]]
for iid in below: for iid in below:
session.execute( session.execute(
+18
View File
@@ -94,6 +94,24 @@ def _rejected_ids(session: Session, tag_id: int) -> list[int]:
] ]
def _applied_or_rejected(session: Session, tag_ids) -> dict[int, set[int]]:
"""Per-tag skip set for the auto-apply sweeps: every image that ALREADY carries
the tag (ANY source — not just training positives) OR has rejected it. A sweep
never re-applies to these. Shared by auto_apply_sweep + system_tag_auto_apply_sweep
(heads.py) and scheduled_ccip_auto_apply (tasks/ml.py). Callers mutate the returned
sets in-place to also dedupe within a single run."""
skip: dict[int, set[int]] = {}
for tid in tag_ids:
ids = {
r[0] for r in session.execute(
select(image_tag.c.image_record_id).where(image_tag.c.tag_id == tid)
).all()
}
ids.update(_rejected_ids(session, tid))
skip[tid] = ids
return skip
def _sample_unlabeled(session: Session, exclude: set[int], limit: int) -> list[int]: def _sample_unlabeled(session: Session, exclude: set[int], limit: int) -> list[int]:
"""Random image ids (with an embedding) NOT carrying the tag. Concepts are """Random image ids (with an embedding) NOT carrying the tag. Concepts are
sparse, so an untagged image is almost always a true negative.""" sparse, so an untagged image is almost always a true negative."""
+20 -36
View File
@@ -91,48 +91,46 @@ def _sync_lookup(vanity: str, cookies_path: str | None) -> str | None:
) )
def _lookup_via_api(vanity: str, cookies_path: str | None) -> str | None: def _campaigns_api_first(vanity: str, cookies_path: str | None) -> dict | None:
"""The first `data` object from Patreon's campaigns API filtered by vanity
(`?filter[vanity]=<vanity>&fields[campaign]=name`), or None on any failure
(network / non-200 / non-JSON / empty). The single request shape shared by
_lookup_via_api (plucks the campaign id) and resolve_display_name (plucks the
display name)."""
jar = _load_cookie_jar(cookies_path) jar = _load_cookie_jar(cookies_path)
headers = {
"User-Agent": _USER_AGENT,
"Accept": "application/vnd.api+json",
}
params = {
"filter[vanity]": vanity,
"fields[campaign]": "name",
}
try: try:
resp = requests.get( resp = requests.get(
_CAMPAIGNS_URL, _CAMPAIGNS_URL,
params=params, params={"filter[vanity]": vanity, "fields[campaign]": "name"},
headers=headers, headers={"User-Agent": _USER_AGENT, "Accept": "application/vnd.api+json"},
cookies=jar, cookies=jar,
timeout=_TIMEOUT_SECONDS, timeout=_TIMEOUT_SECONDS,
) )
except requests.RequestException as exc: except requests.RequestException as exc:
log.warning("Patreon campaigns API request failed for vanity=%s: %s", vanity, exc) log.warning("Patreon campaigns API request failed for vanity=%s: %s", vanity, exc)
return None return None
if resp.status_code != 200: if resp.status_code != 200:
log.warning( log.warning(
"Patreon campaigns API returned HTTP %d for vanity=%s", "Patreon campaigns API returned HTTP %d for vanity=%s",
resp.status_code, vanity, resp.status_code, vanity,
) )
return None return None
try: try:
payload = resp.json() payload = resp.json()
except ValueError as exc: except ValueError as exc:
log.warning("Patreon campaigns API returned non-JSON for vanity=%s: %s", vanity, exc) log.warning("Patreon campaigns API returned non-JSON for vanity=%s: %s", vanity, exc)
return None return None
data = payload.get("data") if isinstance(payload, dict) else None
if not isinstance(data, list) or not data or not isinstance(data[0], dict):
return None
return data[0]
if not isinstance(payload, dict):
def _lookup_via_api(vanity: str, cookies_path: str | None) -> str | None:
first = _campaigns_api_first(vanity, cookies_path)
if first is None:
return None return None
data = payload.get("data") campaign_id = first.get("id")
if not isinstance(data, list) or not data:
return None
first = data[0] if isinstance(data[0], dict) else None
campaign_id = first.get("id") if first else None
if not isinstance(campaign_id, str) or not campaign_id: if not isinstance(campaign_id, str) or not campaign_id:
return None return None
log.info("Resolved Patreon vanity=%s → campaign_id=%s", vanity, campaign_id) log.info("Resolved Patreon vanity=%s → campaign_id=%s", vanity, campaign_id)
@@ -144,24 +142,10 @@ def resolve_display_name(vanity: str, cookies_path: str | None) -> str | None:
(`fields[campaign]=name`), used to name the Artist at add-time (#130). None (`fields[campaign]=name`), used to name the Artist at add-time (#130). None
on any failure — the caller falls back to the vanity handle. Sync: call from on any failure — the caller falls back to the vanity handle. Sync: call from
an executor.""" an executor."""
jar = _load_cookie_jar(cookies_path) first = _campaigns_api_first(vanity, cookies_path)
try: if first is None:
resp = requests.get(
_CAMPAIGNS_URL,
params={"filter[vanity]": vanity, "fields[campaign]": "name"},
headers={"User-Agent": _USER_AGENT, "Accept": "application/vnd.api+json"},
cookies=jar,
timeout=_TIMEOUT_SECONDS,
)
if resp.status_code != 200:
return None
data = resp.json().get("data")
except (requests.RequestException, ValueError) as exc:
log.warning("Patreon name lookup failed for vanity=%s: %s", vanity, exc)
return None return None
if not isinstance(data, list) or not data or not isinstance(data[0], dict): name = (first.get("attributes") or {}).get("name")
return None
name = (data[0].get("attributes") or {}).get("name")
return name.strip() if isinstance(name, str) and name.strip() else None return name.strip() if isinstance(name, str) and name.strip() else None
+45 -72
View File
@@ -776,89 +776,62 @@ def recover_stalled_library_audit_runs() -> int:
return recovered return recovered
def _recover_stalled_runs(model, *, stall_minutes: int, keep_runs: int, label: str) -> int:
"""Shared recovery + retention sweep for the head run-tracking tables
(HeadTrainingRun / HeadAutoApplyRun, which share the
status/last_progress_at/started_at/finished_at/error/id columns): flip 'running'
rows with no progress past `stall_minutes` to 'error', then prune to the last
`keep_runs` (rule 89). Returns the number recovered. NOTE the two other recover
tasks are deliberately NOT folded in — library-audit has no prune tail and
backup uses a single started_at cutoff."""
SessionLocal = _sync_session_factory()
now = datetime.now(UTC)
cutoff = now - timedelta(minutes=stall_minutes)
with SessionLocal() as session:
result = session.execute(
update(model)
.where(model.status == "running")
.where(func.coalesce(model.last_progress_at, model.started_at) < cutoff)
.values(
status="error", finished_at=now,
error=f"stranded by recovery sweep (no progress for {stall_minutes} min)",
)
)
keep = session.execute(
select(model.id).order_by(model.id.desc()).limit(keep_runs)
).scalars().all()
if keep:
session.execute(delete(model).where(model.id.not_in(keep)))
session.commit()
recovered = result.rowcount or 0
if recovered:
log.info("%s: recovered %d rows", label, recovered)
return recovered
@celery.task(name="backend.app.tasks.maintenance.recover_stalled_head_training_runs") @celery.task(name="backend.app.tasks.maintenance.recover_stalled_head_training_runs")
def recover_stalled_head_training_runs() -> int: def recover_stalled_head_training_runs() -> int:
"""Flip HeadTrainingRun rows stuck in 'running' past the stall threshold to """Flip HeadTrainingRun rows stuck in 'running' past the stall threshold to
'error', and prune old runs to the last HEAD_TRAINING_KEEP_RUNS (retention, 'error', and prune old runs to the last HEAD_TRAINING_KEEP_RUNS (retention,
rule 89). Runs every 5 min on the maintenance lane; no-op when idle.""" rule 89). Runs every 5 min on the maintenance lane; no-op when idle."""
SessionLocal = _sync_session_factory() return _recover_stalled_runs(
now = datetime.now(UTC) HeadTrainingRun,
cutoff = now - timedelta(minutes=HEAD_TRAINING_STALL_THRESHOLD_MINUTES) stall_minutes=HEAD_TRAINING_STALL_THRESHOLD_MINUTES,
with SessionLocal() as session: keep_runs=HEAD_TRAINING_KEEP_RUNS,
result = session.execute( label="recover_stalled_head_training_runs",
update(HeadTrainingRun) )
.where(HeadTrainingRun.status == "running")
.where(
func.coalesce(
HeadTrainingRun.last_progress_at, HeadTrainingRun.started_at
)
< cutoff
)
.values(
status="error", finished_at=now,
error=(
f"stranded by recovery sweep (no progress for "
f"{HEAD_TRAINING_STALL_THRESHOLD_MINUTES} min)"
),
)
)
keep = session.execute(
select(HeadTrainingRun.id).order_by(HeadTrainingRun.id.desc())
.limit(HEAD_TRAINING_KEEP_RUNS)
).scalars().all()
if keep:
session.execute(
delete(HeadTrainingRun).where(HeadTrainingRun.id.not_in(keep))
)
session.commit()
recovered = result.rowcount or 0
if recovered:
log.info(
"recover_stalled_head_training_runs: recovered %d rows", recovered
)
return recovered
@celery.task(name="backend.app.tasks.maintenance.recover_stalled_head_auto_apply_runs") @celery.task(name="backend.app.tasks.maintenance.recover_stalled_head_auto_apply_runs")
def recover_stalled_head_auto_apply_runs() -> int: def recover_stalled_head_auto_apply_runs() -> int:
"""Flip stalled HeadAutoApplyRun 'running' rows to 'error' + prune to the """Flip stalled HeadAutoApplyRun 'running' rows to 'error' + prune to the
last HEAD_AUTO_APPLY_KEEP_RUNS (retention, rule 89). 5-min maintenance lane.""" last HEAD_AUTO_APPLY_KEEP_RUNS (retention, rule 89). 5-min maintenance lane."""
SessionLocal = _sync_session_factory() return _recover_stalled_runs(
now = datetime.now(UTC) HeadAutoApplyRun,
cutoff = now - timedelta(minutes=HEAD_AUTO_APPLY_STALL_THRESHOLD_MINUTES) stall_minutes=HEAD_AUTO_APPLY_STALL_THRESHOLD_MINUTES,
with SessionLocal() as session: keep_runs=HEAD_AUTO_APPLY_KEEP_RUNS,
result = session.execute( label="recover_stalled_head_auto_apply_runs",
update(HeadAutoApplyRun) )
.where(HeadAutoApplyRun.status == "running")
.where(
func.coalesce(
HeadAutoApplyRun.last_progress_at, HeadAutoApplyRun.started_at
)
< cutoff
)
.values(
status="error", finished_at=now,
error=(
f"stranded by recovery sweep (no progress for "
f"{HEAD_AUTO_APPLY_STALL_THRESHOLD_MINUTES} min)"
),
)
)
keep = session.execute(
select(HeadAutoApplyRun.id).order_by(HeadAutoApplyRun.id.desc())
.limit(HEAD_AUTO_APPLY_KEEP_RUNS)
).scalars().all()
if keep:
session.execute(
delete(HeadAutoApplyRun).where(HeadAutoApplyRun.id.not_in(keep))
)
session.commit()
recovered = result.rowcount or 0
if recovered:
log.info(
"recover_stalled_head_auto_apply_runs: recovered %d rows", recovered
)
return recovered
# Keep ~6 months of daily head-metric snapshots (enough to see tuning trends). # Keep ~6 months of daily head-metric snapshots (enough to see tuning trends).
+10 -30
View File
@@ -105,9 +105,7 @@ def embed_image(self, image_id: int) -> dict:
record = session.get(ImageRecord, image_id) record = session.get(ImageRecord, image_id)
if record is None: if record is None:
return {"status": "missing", "image_id": image_id} return {"status": "missing", "image_id": image_id}
settings = session.execute( settings = MLSettings.load_sync(session)
select(MLSettings).where(MLSettings.id == 1)
).scalar_one()
src = Path(record.path) src = Path(record.path)
is_vid = _is_video(src) is_vid = _is_video(src)
@@ -488,15 +486,10 @@ def scheduled_ccip_auto_apply() -> str:
from sqlalchemy import select as sa_select from sqlalchemy import select as sa_select
from sqlalchemy.dialects.postgresql import insert as pg_insert from sqlalchemy.dialects.postgresql import insert as pg_insert
from ..models import ImageRegion, MLSettings, Tag, TagKind, TagSuggestionRejection from ..models import ImageRegion, MLSettings, Tag, TagKind
from ..models.tag import image_tag from ..models.tag import image_tag
from ..services.ml.ccip import _FIGURE_KINDS
fig = ("face", "figure") from ..services.ml.training_data import _applied_or_rejected, _l2norm
def _l2(m):
n = np.linalg.norm(m, axis=1, keepdims=True)
n[n == 0] = 1.0
return m / n
SessionLocal = _sync_session_factory() SessionLocal = _sync_session_factory()
with SessionLocal() as session: with SessionLocal() as session:
@@ -521,7 +514,7 @@ def scheduled_ccip_auto_apply() -> str:
) )
.join(Tag, Tag.id == image_tag.c.tag_id) .join(Tag, Tag.id == image_tag.c.tag_id)
.where(Tag.kind == TagKind.character) .where(Tag.kind == TagKind.character)
.where(ImageRegion.kind.in_(fig)) .where(ImageRegion.kind.in_(_FIGURE_KINDS))
.where(ImageRegion.ccip_embedding.is_not(None)) .where(ImageRegion.ccip_embedding.is_not(None))
.where(ImageRegion.image_record_id.in_(single)) .where(ImageRegion.image_record_id.in_(single))
).all() ).all()
@@ -532,29 +525,16 @@ def scheduled_ccip_auto_apply() -> str:
for tid, vec in ref_rows: for tid, vec in ref_rows:
by_char.setdefault(tid, []).append(vec) by_char.setdefault(tid, []).append(vec)
ref_tags = list(by_char) ref_tags = list(by_char)
mats = [_l2(np.asarray(by_char[t], dtype=np.float32)) for t in ref_tags] mats = [_l2norm(np.asarray(by_char[t], dtype=np.float32), np) for t in ref_tags]
allref = np.vstack(mats) # (total, 768) allref = np.vstack(mats) # (total, 768)
seg = np.cumsum([0] + [len(m) for m in mats])[:-1] # per-char start seg = np.cumsum([0] + [len(m) for m in mats])[:-1] # per-char start
# Per character: images that already carry OR rejected the tag — skip. # Per character: images that already carry OR rejected the tag — skip.
skip = {t: set() for t in ref_tags} skip = _applied_or_rejected(session, ref_tags)
for t in ref_tags:
for (iid,) in session.execute(
sa_select(image_tag.c.image_record_id).where(
image_tag.c.tag_id == t
)
):
skip[t].add(iid)
for (iid,) in session.execute(
sa_select(TagSuggestionRejection.image_record_id).where(
TagSuggestionRejection.tag_id == t
)
):
skip[t].add(iid)
img_ids = list(session.execute( img_ids = list(session.execute(
sa_select(ImageRegion.image_record_id) sa_select(ImageRegion.image_record_id)
.where(ImageRegion.kind.in_(fig), ImageRegion.ccip_embedding.is_not(None)) .where(ImageRegion.kind.in_(_FIGURE_KINDS), ImageRegion.ccip_embedding.is_not(None))
.distinct() .distinct()
).scalars()) ).scalars())
@@ -566,7 +546,7 @@ def scheduled_ccip_auto_apply() -> str:
sa_select(ImageRegion.image_record_id, ImageRegion.ccip_embedding) sa_select(ImageRegion.image_record_id, ImageRegion.ccip_embedding)
.where( .where(
ImageRegion.image_record_id.in_(chunk), ImageRegion.image_record_id.in_(chunk),
ImageRegion.kind.in_(fig), ImageRegion.kind.in_(_FIGURE_KINDS),
ImageRegion.ccip_embedding.is_not(None), ImageRegion.ccip_embedding.is_not(None),
) )
).all() ).all()
@@ -574,7 +554,7 @@ def scheduled_ccip_auto_apply() -> str:
for iid, vec in rows: for iid, vec in rows:
by_img.setdefault(iid, []).append(vec) by_img.setdefault(iid, []).append(vec)
for iid, vecs in by_img.items(): for iid, vecs in by_img.items():
q = _l2(np.asarray(vecs, dtype=np.float32)) # (nq, 768) q = _l2norm(np.asarray(vecs, dtype=np.float32), np) # (nq, 768)
colmax = (q @ allref.T).max(axis=0) # (total,) colmax = (q @ allref.T).max(axis=0) # (total,)
charmax = np.maximum.reduceat(colmax, seg) # (n_chars,) charmax = np.maximum.reduceat(colmax, seg) # (n_chars,)
for ci in np.where(charmax >= thr)[0]: for ci in np.where(charmax >= thr)[0]:
+66
View File
@@ -31,6 +31,69 @@ browser.runtime.onInstalled.addListener(() => ensureInitialized());
browser.runtime.onStartup.addListener(() => ensureInitialized()); browser.runtime.onStartup.addListener(() => ensureInitialized());
ensureInitialized().catch(e => console.error('init failed:', e)); ensureInitialized().catch(e => console.error('init failed:', e));
// ---- Extension self-update check (#1489) ----
// Installed per-instance from the operator's FC host, so Firefox's static
// update_url can't apply (each instance has a different host). Instead ask the
// configured backend for the latest published version and nudge the operator to
// reinstall the freshly-signed XPI — surfaced as a popup banner (on demand) and
// a toolbar badge (daily). /api/extension/manifest is public and returns
// {version, latest_url, sha256}; the XPI is served from the web root (not /api).
function versionIsNewer(candidate, current) {
// Dotted numeric compare so 1.0.10 > 1.0.9 (a plain string compare wouldn't).
const a = String(candidate).split('.').map(n => parseInt(n, 10) || 0);
const b = String(current).split('.').map(n => parseInt(n, 10) || 0);
for (let i = 0; i < Math.max(a.length, b.length); i++) {
if ((a[i] || 0) !== (b[i] || 0)) return (a[i] || 0) > (b[i] || 0);
}
return false;
}
async function checkForUpdateInfo() {
await ensureInitialized();
if (!api.isConfigured()) return { updateAvailable: false, configured: false };
let info;
try {
info = await api.getExtensionManifest();
} catch (e) {
return { updateAvailable: false, error: e.message };
}
const currentVersion = browser.runtime.getManifest().version;
const latestVersion = info && info.version ? info.version : null;
// latest_url is served from the web root; strip the /api suffix off baseUrl
// (same transform as OPEN_ARTIST_PAGE).
const base = (api.baseUrl || '').replace(/\/+$/, '').replace(/\/api$/, '');
return {
updateAvailable: !!latestVersion && versionIsNewer(latestVersion, currentVersion),
currentVersion,
latestVersion,
xpiUrl: info && info.latest_url ? `${base}${info.latest_url}` : null,
};
}
async function refreshUpdateBadge() {
let r;
try { r = await checkForUpdateInfo(); } catch { return; }
try {
await browser.action.setBadgeText({ text: r.updateAvailable ? '↑' : '' });
if (r.updateAvailable) {
await browser.action.setBadgeBackgroundColor({ color: '#F4BA7A' });
await browser.action.setTitle({ title: `FabledCurator — update available (v${r.latestVersion})` });
} else {
await browser.action.setTitle({ title: 'FabledCurator' });
}
} catch { /* action API unavailable — non-fatal */ }
}
// Daily proactive check (needs the "alarms" permission). create() is idempotent
// by name, so re-running it on each event-page load is safe.
browser.alarms.create('fc-update-check', { periodInMinutes: 24 * 60, delayInMinutes: 1 });
browser.alarms.onAlarm.addListener((alarm) => {
if (alarm.name === 'fc-update-check') refreshUpdateBadge();
});
browser.runtime.onStartup.addListener(() => refreshUpdateBadge());
browser.runtime.onInstalled.addListener(() => refreshUpdateBadge());
// ---- Discord token capture via webRequest ---- // ---- Discord token capture via webRequest ----
browser.webRequest.onBeforeSendHeaders.addListener( browser.webRequest.onBeforeSendHeaders.addListener(
@@ -298,6 +361,9 @@ browser.runtime.onMessage.addListener(async (msg) => {
} }
} }
case 'CHECK_UPDATE':
return await checkForUpdateInfo();
default: default:
return { error: `Unknown message type: ${msg.type}` }; return { error: `Unknown message type: ${msg.type}` };
} }
+6
View File
@@ -89,6 +89,12 @@ class FabledCuratorAPI {
const qs = new URLSearchParams({ url }).toString(); const qs = new URLSearchParams({ url }).toString();
return this.request('GET', `/extension/probe?${qs}`); return this.request('GET', `/extension/probe?${qs}`);
} }
// Latest published extension version on this instance — drives the in-app
// update prompt. Public endpoint (no key needed, but request() sends it
// harmlessly). Returns {version, xpi_url, latest_url, sha256}.
getExtensionManifest() {
return this.request('GET', '/extension/manifest');
}
// Connection test = the cheapest read with auth. // Connection test = the cheapest read with auth.
testConnection() { testConnection() {
+3 -2
View File
@@ -1,7 +1,7 @@
{ {
"manifest_version": 3, "manifest_version": 3,
"name": "FabledCurator", "name": "FabledCurator",
"version": "1.0.8", "version": "1.0.9",
"description": "Export cookies from supported platforms to FabledCurator and add creators as sources in one click.", "description": "Export cookies from supported platforms to FabledCurator and add creators as sources in one click.",
"browser_specific_settings": { "browser_specific_settings": {
@@ -22,7 +22,8 @@
"tabs", "tabs",
"activeTab", "activeTab",
"webRequest", "webRequest",
"webRequestBlocking" "webRequestBlocking",
"alarms"
], ],
"host_permissions": [ "host_permissions": [
+1 -1
View File
@@ -1,6 +1,6 @@
{ {
"name": "fabledcurator-extension", "name": "fabledcurator-extension",
"version": "1.0.8", "version": "1.0.9",
"private": true, "private": true,
"description": "Firefox extension for FabledCurator", "description": "Firefox extension for FabledCurator",
"scripts": { "scripts": {
+11
View File
@@ -72,6 +72,17 @@ body {
.btn.block { display: block; width: 100%; margin-top: 8px; } .btn.block { display: block; width: 100%; margin-top: 8px; }
.btn.link { background: none; color: var(--on-surface-variant); padding: 4px; } .btn.link { background: none; color: var(--on-surface-variant); padding: 4px; }
.btn.link:hover { color: var(--accent); } .btn.link:hover { color: var(--accent); }
.btn.small { padding: 6px 12px; font-size: 13px; }
/* In-app update prompt (accent-tinted so it reads as an actionable notice). */
.update-banner {
display: flex; align-items: center; gap: 10px;
margin: 10px 10px 0; padding: 10px 12px;
background: rgba(244, 186, 122, 0.12);
border: 1px solid rgba(244, 186, 122, 0.4);
border-radius: 6px;
}
#update-text { flex: 1; font-size: 13px; }
.source-row .play { .source-row .play {
background: none; border: none; color: var(--on-surface-variant); background: none; border: none; color: var(--on-surface-variant);
+5
View File
@@ -20,6 +20,11 @@
</section> </section>
<section id="main-content" class="main hidden"> <section id="main-content" class="main hidden">
<div id="update-banner" class="update-banner hidden">
<span id="update-text"></span>
<button id="update-btn" class="btn primary small">Update</button>
</div>
<nav class="tabs"> <nav class="tabs">
<button class="tab active" data-tab="platforms">Platforms</button> <button class="tab active" data-tab="platforms">Platforms</button>
<button class="tab" data-tab="sources">Sources</button> <button class="tab" data-tab="sources">Sources</button>
+21
View File
@@ -14,6 +14,7 @@ async function init() {
setupEventListeners(); setupEventListeners();
showPlatformsLoading(); showPlatformsLoading();
testConnectionIfNeeded(); testConnectionIfNeeded();
checkForUpdate();
loadPlatformStatus().catch(e => showError(`Failed to load platforms: ${e.message}`)); loadPlatformStatus().catch(e => showError(`Failed to load platforms: ${e.message}`));
} catch (e) { } catch (e) {
showSetupRequired(); showSetupRequired();
@@ -63,6 +64,26 @@ function updateConnectionDot(connected) {
d.title = connected ? 'Connected to FabledCurator' : 'Disconnected'; d.title = connected ? 'Connected to FabledCurator' : 'Disconnected';
} }
// Nudge to reinstall when the configured instance publishes a newer signed XPI
// (the extension is self-hosted, so there's no Firefox auto-update). Never
// blocks the popup — a failed check just leaves the banner hidden.
async function checkForUpdate() {
try {
const r = await browser.runtime.sendMessage({ type: 'CHECK_UPDATE' });
if (r && r.updateAvailable && r.xpiUrl) showUpdateBanner(r);
} catch { /* non-fatal */ }
}
function showUpdateBanner(r) {
document.getElementById('update-text').textContent =
`Update available — v${r.latestVersion} (installed v${r.currentVersion})`;
// Opening the signed XPI triggers Firefox's native install prompt.
document.getElementById('update-btn').addEventListener('click', () => {
browser.tabs.create({ url: r.xpiUrl });
});
document.getElementById('update-banner').classList.remove('hidden');
}
async function loadPlatformStatus() { async function loadPlatformStatus() {
const status = await browser.runtime.sendMessage({ type: 'GET_PLATFORM_STATUS' }); const status = await browser.runtime.sendMessage({ type: 'GET_PLATFORM_STATUS' });
const c = document.getElementById('platforms-list'); const c = document.getElementById('platforms-list');
+59
View File
@@ -0,0 +1,59 @@
"""Shared ML helpers extracted in the DRY pass (milestone #161). These pin the
single sources the auto-apply sweeps now trust, so a future edit can't silently
drift them: `_applied_or_rejected` is the skip-set used by auto_apply_sweep,
system_tag_auto_apply_sweep (heads.py) and scheduled_ccip_auto_apply (tasks/ml.py);
`_sigmoid` is the head score→prob transform used at every scoring site."""
import pytest
from backend.app.models import ImageRecord, Tag, TagKind, TagSuggestionRejection
from backend.app.models.tag import image_tag
from backend.app.services.ml.training_data import _applied_or_rejected
def test_sigmoid_matches_naive_form():
import numpy as np
from backend.app.services.ml.heads import _sigmoid
z = np.array([-3.0, -0.5, 0.0, 1.5, 12.0], dtype=np.float32)
assert np.allclose(_sigmoid(z, np), 1.0 / (1.0 + np.exp(-z)))
assert float(_sigmoid(np.array([0.0]), np)[0]) == pytest.approx(0.5)
@pytest.mark.integration
def test_applied_or_rejected_unions_applied_any_source_and_rejected(db_sync):
a = Tag(name="dry-helper-a", kind=TagKind.general)
b = Tag(name="dry-helper-b", kind=TagKind.general)
db_sync.add_all([a, b])
db_sync.flush()
imgs = []
for i in range(5):
img = ImageRecord(
path=f"/images/dryhelp{i}.jpg", sha256=f"{i:064d}", size_bytes=1,
mime="image/jpeg", width=1, height=1, origin="imported_filesystem",
integrity_status="unknown", siglip_embedding=[0.0] * 1152,
)
db_sync.add(img)
imgs.append(img)
db_sync.flush()
# tag a: applied manually (img0), applied by an AUTO source (img1), rejected (img2).
db_sync.execute(image_tag.insert().values(
image_record_id=imgs[0].id, tag_id=a.id, source="manual"))
db_sync.execute(image_tag.insert().values(
image_record_id=imgs[1].id, tag_id=a.id, source="head_auto"))
db_sync.add(TagSuggestionRejection(image_record_id=imgs[2].id, tag_id=a.id))
# tag b: applied to img3 only.
db_sync.execute(image_tag.insert().values(
image_record_id=imgs[3].id, tag_id=b.id, source="manual"))
db_sync.flush()
skip = _applied_or_rejected(db_sync, [a.id, b.id])
# Applied-under-ANY-source (manual + head_auto) rejected, kept per-tag; the
# untouched image (img4) appears under neither tag.
assert skip[a.id] == {imgs[0].id, imgs[1].id, imgs[2].id}
assert skip[b.id] == {imgs[3].id}
assert imgs[4].id not in skip[a.id]
assert imgs[4].id not in skip[b.id]
+63
View File
@@ -0,0 +1,63 @@
"""recover_stalled_head_training_runs + recover_stalled_head_auto_apply_runs share
one helper (_recover_stalled_runs, DRY pass #161). These pin BOTH wrappers so the
shared source stays correct: a 'running' row with no progress past the stall
threshold flips to 'error'; a fresh 'running' row is left alone."""
from datetime import UTC, datetime, timedelta
import pytest
from sqlalchemy import select
from backend.app.models import HeadAutoApplyRun, HeadTrainingRun
pytestmark = pytest.mark.integration
def test_recover_stalled_head_training_runs_flips_stalled_keeps_fresh(db_sync):
from backend.app.tasks.maintenance import recover_stalled_head_training_runs
stale = HeadTrainingRun(
params={}, status="running",
last_progress_at=datetime.now(UTC) - timedelta(days=1),
)
fresh = HeadTrainingRun(
params={}, status="running", last_progress_at=datetime.now(UTC),
)
db_sync.add_all([stale, fresh])
db_sync.commit()
stale_id, fresh_id = stale.id, fresh.id
assert recover_stalled_head_training_runs.apply().get() == 1
db_sync.expire_all()
assert db_sync.execute(
select(HeadTrainingRun.status).where(HeadTrainingRun.id == stale_id)
).scalar_one() == "error"
assert db_sync.execute(
select(HeadTrainingRun.status).where(HeadTrainingRun.id == fresh_id)
).scalar_one() == "running"
def test_recover_stalled_head_auto_apply_runs_flips_stalled_keeps_fresh(db_sync):
from backend.app.tasks.maintenance import recover_stalled_head_auto_apply_runs
stale = HeadAutoApplyRun(
dry_run=False, params={}, status="running",
last_progress_at=datetime.now(UTC) - timedelta(days=1),
)
fresh = HeadAutoApplyRun(
dry_run=False, params={}, status="running",
last_progress_at=datetime.now(UTC),
)
db_sync.add_all([stale, fresh])
db_sync.commit()
stale_id, fresh_id = stale.id, fresh.id
assert recover_stalled_head_auto_apply_runs.apply().get() == 1
db_sync.expire_all()
assert db_sync.execute(
select(HeadAutoApplyRun.status).where(HeadAutoApplyRun.id == stale_id)
).scalar_one() == "error"
assert db_sync.execute(
select(HeadAutoApplyRun.status).where(HeadAutoApplyRun.id == fresh_id)
).scalar_one() == "running"