775941609dec48e0af03796721b406d423960647
100
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
775941609d |
feat(ui): confirm/keep auto-applied tags (milestone 139)
Auto-applied tags are provisional (they don't train the model + can be retracted until confirmed), so surface and confirm them: - Backend: list_for_image + get_image_with_tags now include `source` + a `confirmed` flag on each applied tag (via serialize_tag, image-scoped; defaulted for autocomplete/directory callers). - Frontend: TagChip badges an unconfirmed auto-tag with an "auto" pill + a one-click Keep/confirm (✓) → POST /images/<id>/tags/<id>/confirm, which promotes it to a training positive and shields it from the retraction sweep; TagPanel reloads so the badge + button drop once confirmed. Contract test for the source/confirmed payload. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
d3984ccb0d |
fix(ui): tag autocomplete no longer shows stale wrong-prefix results (race)
The keystroke debounce cleared the timer but not an already-fired fetch, so a
slower earlier-prefix response ("s") could land after "sex" and overwrite the
dropdown with wrong-prefix matches (operator-flagged with a "sex"→Stockings/
Super Mario screenshot). Gate each autocomplete response on a useInflightToken
(cancel on every keystroke, isCurrent() after the await) so only the latest
query's results are applied.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM
|
||
|
|
7d3a3b4a83 |
revert(ml): keep head auto-apply precision at 0.97 (operator: general tuning was fine)
Milestone 139 raised head_auto_apply_precision 0.97→0.98; operator confirmed the general-tag confidence was already well tuned, so revert that. The support floor (min_positives 30→50) and CCIP match confidence (0.92→0.95) stay. Migration 0081 (not yet deployed) edited to drop the precision bump. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
6684907577 |
feat(ui): "Reject rest" per suggestion category — confirm the good, reject the rest
Each category section header gets a subtle "Reject rest" action that dismisses every still-unhandled suggestion in it at once (store.dismissRemaining, parallel dispatch). Canonical tags persist a rejection and stay flagged (reversible, one-click un-reject); raw creates-new-tag rows drop client-side. Shows only when the section has unhandled items. No confirm dialog — it's fully reversible. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
2d44a26bdf |
feat(ml): auto-applied tags don't train a head unless confirmed (milestone 139)
Makes auto-apply truly "soft" for heads: _ids_with_tag (head positives) and _eligible_tag_ids (graduation count) now count human-applied + operator-confirmed tags only, via a shared _AUTO_SOURCES (head_auto/ccip_auto/ml_auto) exclusion. Unconfirmed auto-applied tags no longer train the head that judges them, so a misfire can't reinforce itself and the retraction sweep can actually drop it. Confirming a tag (TagPositiveConfirmation) promotes it to a positive AND protects it from retraction. sklearn-free tests. CCIP reference exclusion is the companion piece, next. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
0de726ed48 |
test(ml): bump auto-apply test head n_pos default 30→60 past the new floor (m139)
The stricter head_auto_apply_min_positives (30→50, migration 0081) dropped the _head helper's default n_pos=30 below the support floor, so the "supported head" sweep tests saw the head as ineligible (n_applied 0). Move the default to 60; the explicit n_pos=5 under-supported test stays correct. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
3006e84cc0 |
feat(ml): soft auto-apply — retract auto-tags now below threshold (milestone 139)
Daily scheduled_retract_auto_tags re-scores standing auto-applied tags and drops the ones the model no longer supports: - retract_auto_applied_heads: per graduated head, re-score its source='head_auto' images (bounded — only the images already carrying the auto-tag, not the whole library) and remove ones now < auto_apply_threshold. - retract_auto_applied_ccip: per source='ccip_auto' character tag, max-cosine the image's figure vectors vs that character's prototypes; remove ones now below the ccip auto-apply threshold. Both SKIP operator-confirmed tags (TagPositiveConfirmation) and are SILENT — a low score isn't proof the tag was wrong, so no hard negative is recorded (that's reserved for an operator removal). No-op unless the relevant auto-apply switch is on. New daily beat. sklearn-free tests for both paths + the disabled no-op. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
cbc3e11a53 |
feat(ml): stricter auto-apply defaults to cut misfires (milestone 139)
head_auto_apply_precision 0.97→0.98, head_auto_apply_min_positives 30→50, ccip_auto_apply_threshold 0.92→0.95 (operator-asked). Model defaults change for fresh installs; migration 0081 bumps the existing singleton row IFF still at the old default (won't clobber a deliberate operator change). ml_admin bounds already permit these. Fixed a stale comment in the auto-apply test. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
2cfbb284d5 |
feat(heads): incremental retraining — refit only changed tags (#1317 phase 2, m138)
train_all_heads is now incremental by default: a per-tag training-data fingerprint (positive + rejection count/latest-timestamp, stored on tag_head.train_fingerprint) means a manual Retrain refits ONLY the tags whose data changed — O(what you touched), not O(all heads). The nightly scheduled_train_heads passes full=True to reconcile sampled-negative + hygiene drift across every head. First incremental run after deploy still refits everyone (NULL fingerprints), stamping them, then it's incremental. The refit decision + fingerprint are split into sklearn-free helpers (_head_fingerprints, _heads_needing_retrain) so the incremental logic is unit-tested directly (train_head itself needs scikit-learn). Migration 0080. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
a94f6a2789 |
feat(ccip): matcher reads the incremental prototype store (#1317, m138 step 4)
match_image now sources character references from character_prototype via a per-character in-process cache (_load_prototypes) that reloads ONLY the characters whose ccip_prototype_state.updated_at advanced — no request-path rebuild, so the per-accept ~4s stall is gone once the store is populated. Cold start (store empty pre-first-refresh) falls back to the legacy on-the-fly reference build, so character suggestions work immediately post-deploy and the background refresh populates the store within ~15 min. Match math + grounding are unchanged; existing tests exercise the legacy fallback, and a new test covers matching from the populated prototype store. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
a1ed53136e |
feat(ccip): refresh task + beat + retrain hook for character prototypes (#1317, m138 step 3)
- refresh_character_prototypes celery task wraps the incremental builder (sync ml worker); returns skipped / rebuilt=N removed=N. - Beat: every ~15 min (cheap global-gate no-op when idle) + a nightly full=True reconcile as belt-and-suspenders. - train_heads enqueues it on success, so the Retrain button AND the nightly head retrain refresh CCIP on the SAME trigger — unified lifecycle, as asked. The initial (cold) full build loads the whole reference set once in the background, never on a /suggestions request. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
9504870c9a |
feat(ccip): incremental character-prototype builder (#1317, m138 step 2)
refresh_character_prototypes (sync, celery ml worker): - Cheap GLOBAL gate (a few COUNTs) → no-op when nothing that affects references changed since the last refresh (the operator's "only recompute if something was tagged" trigger). - Else a per-character fingerprint diff (one GROUP BY: ref count + max region id) rebuilds ONLY the characters whose references moved — each capped to MLSettings.ccip_prototype_cap — and drops characters that lost all refs. Cost scales with WHAT changed, not library size. Reuses ccip's reference predicate (single-character, non-hygiene, figure CCIP) so prototypes match the legacy matcher exactly. The async matcher (next step) will READ the table. Tests: gate no-op when idle, only-changed-character rebuild, capping, single-character exclusion, lost-reference cleanup. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
f24dc81764 |
feat(ccip): schema for precomputed incremental character prototypes (#1317, m138 step 1)
Foundation for making CCIP character references a precomputed, INCREMENTAL artifact instead of a request-path rebuild (kills the per-accept ~4s suggestions stall; cost will scale with change, not library size): - character_prototype: a character's reference CCIP vectors, capped to MLSettings.ccip_prototype_cap so match cost doesn't grow with popularity. - ccip_prototype_state: per-character fingerprint (ref count + max region id) + updated_at → drives per-character incremental rebuilds and the matcher cache's reload-only-what-advanced. - MLSettings.ccip_ref_signature (cheap global change gate) + ccip_prototype_cap. Migration 0079. Schema + models only — the builder service, refresh task/beat, and matcher rewrite land in the following steps. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
48be921b8d |
fix(ui): tag-rail no longer scroll-jumps on accept/reject; keep suggestions clear of the toast
- TagAutocomplete: focus the inner input with preventScroll so handing focus back after an accept/reject stops yanking the rail up to the field (which sits above the suggestions list — you had to scroll back down every time). Applies to both the image modal and the Explore rail. - ExploreView: pad the right rail's scroll bottom (88px) so the bottom-right snackbar floats over empty space instead of covering the last suggestions and their accept/reject controls. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
7939dba9ed |
feat(ui): grounding overlay leads with the hovered tag, crop origin as subtext (#1206)
The overlay label showed only the crop origin (booru:head / panel / person-m).
Put the hovered TAG on top as the headline and demote the crop origin to a small,
dimmed, monospace subline — the tag is what you're evaluating; the origin is just
provenance. Threads the tag name through the fcSuggestionHover payload ({g, tag})
from both setters (SuggestionItem for suggestions, TagChip for applied chips).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM
|
||
|
|
8f6547f8d7 |
refactor(subscriptions): remove the dead backfill dry-run preview (#1281)
PreviewDialog.vue was orphaned — nothing mounted it — so the dry-run preview was unreachable. Rather than wire it up, remove the whole chain: its unique value (a capped 3-page count of what a backfill would grab) is low, and its adjacent needs are already covered — auth validation by verify_source_credential, and actually fetching recent posts by "Check now". Operator decision 2026-07-06. Removed: - frontend: PreviewDialog.vue + sources store previewSource() - backend: POST /api/sources/<id>/preview route, download_backends.preview_source, IngestCore.preview() + its now-unused NativeIngestError import - tests: the 3 ingester preview tests Nothing else referenced the chain (verified). Shared campaign-resolution and ledger helpers stay — they're used by run()/verify_source_credential. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
2638cf1a35 |
style(tests): hoist test_api_tags grounding imports to module level (ruff I001)
CI ruff flagged the in-function import block; move them to the top mirroring test_ml_suggestions.py's import line, which ruff already accepts. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
9bb4211722 |
feat(ui): hover an applied tag chip → highlight its grounding crop (#133 step 4)
Applied tags aren't scored live, so compute the grounding on demand: run the
tag's head over the image's max-over-bag (whole-image + concept crops), argmax
→ the region that best explains the tag on this image, mirroring what
score_image records for live suggestions.
- heads.py: extract _image_bag (now shared by score_image) + ground_applied_tag.
Returns (grounding, has_head): has_head False = no head to localize with →
no overlay; grounding None = the whole-image vector won → whole-image frame.
- tags.py: GET /api/images/<id>/tags/<id>/grounding → {grounding, has_head}.
- TagChip/TagPanel: applied chips inject fcSuggestionHover and fetch grounding
on hover (cached per image+tag, race-guarded), reusing Step 3's overlay in
both the modal and Explore. No new frontend overlay code.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM
|
||
|
|
524a26c618 |
feat(ui): hover a suggestion → highlight the crop region it came from (#133 step 3)
The payoff: hover a suggestion in the rail and the exact crop that produced it
lights up on the image (a booru:head, a panel, a figure); a null-grounding tag
shows a subtle dashed whole-image frame ('global vector won, not a crop').
ImageCanvas gains a grounding overlay that tracks the <img>'s live bounding rect
(correct under object-fit letterboxing + pan/zoom) and draws the normalized bbox
+ a detector/kind label. SuggestionItem sets the hovered grounding via
provide/inject (no 4-level event relay through TagPanel/SuggestionsPanel/group);
ImageViewer AND ExploreView provide it + pass it to their canvas. Overlay is
pointer-events:none so it never blocks pan/zoom/click. Videos out of v1 scope.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM
|
||
|
|
dfe2fda564 |
feat(ml): CCIP character matches ground to the matched figure region (#133 step 2)
match_image now tracks WHICH query figure produced the winning cosine per
character (argmax over the per-figure best-reference sim) and attaches its bbox as
grounding {bbox,kind:'figure',detector}. SuggestionService carries it: a CCIP-only
character hit grounds to its figure; a 'both' hit keeps the head's localized crop
if it had one, else falls back to the CCIP figure — so corroborated characters
stay grounded. Test: a character match carries the matched figure's bbox+kind.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM
|
||
|
|
409724b981 |
feat(ml): argmax grounding in score_image → suggestions carry the winning crop (#133 step 1)
score_image now keeps the ARGMAX beside the max-over-bag: which bag row won each
head. The region query also selects bbox/kind/detector_version, a parallel
bag_meta maps each row → its region (None for the whole-image vector), and every
hit gains grounding {bbox,kind,detector} (null when the global vector won). Threaded
through SuggestionService (new Suggestion.grounding field) → /api/.../suggestions
payload. This is the data the #1206 hover-overlay draws. CCIP-only hits ground null
for now (figure grounding = step 2). Tests: winning crop grounds the tag with its
bbox+kind; whole-image win → grounding None.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM
|
||
|
|
ab362bc79c |
feat(ml): Settings → Tagging 'Crop proposers' card (#134 step 3)
Exposes the detector config (per-proposer enable + weights + confidence, caps, dedupe IoU) in Settings → Tagging, backed by MLSettings via /api/ml/settings. ml_admin adds the detector fields to _EDITABLE + GET payload + validation (conf 0..1, caps >=1, IoU 0..1). New CropProposersCard.vue (mirrors HeadsCard) with working defaults pre-filled, per-field live-save (no restart — the agent picks changes up on its next lease), weights-format help, switch-revert on error. Closes milestone #134: all three proposers are on out-of-the-box and tunable in the UI. Test: detector defaults GET + patch round-trip + range validation. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
a4df279343 |
feat(ml): lease announces detector config; agent builds proposers from it live (#134 step 2)
The GPU lease now carries the crop-proposer config from MLSettings in a per-job 'detectors' block (same pattern as embed_model_name). The agent's worker builds its Proposers from the announced config via _effective_cfg (lease block overlaid on env) + _proposers_for (rebuilds only when a config signature changes) — so an operator's UI edit takes effect on the next lease with NO restart, and env is now just the bootstrap fallback until the server announces. enabled-off maps to empty weights (proposer skipped); dedupe_iou + max_regions also come from the effective cfg. Test: lease announces the detectors block with the seeded default weights. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
62ec70b9e4 |
feat(ml): detector config in MLSettings with working defaults (#134 step 1)
Move the crop-proposer config (per-proposer enable + weights + conf, caps, dedupe IoU) into the DB so it's UI-tunable and can be announced to the GPU agent in the lease (like the embedder model) — no restart, agent env becomes bootstrap-only. Migration 0078 adds the columns with working server_defaults so existing rows + fresh installs crop out-of-the-box with all three proposers ON (operator: default-on): person=yolo11n.pt, anatomy=booru_yolo yolov11m_aa22 (URL, license unstated/private-homelab-OK), panel=mosesb best.pt. Plain columns, no CHECK enum. Steps 2 (lease announce + agent apply) and 3 (Settings UI) follow. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
0963bf0db3 |
feat(artist): resolve patreon + subscribestar display names at add-time (#130 step 5)
Parity with pixiv (operator ask): the extension add now resolves the real display name for our other native platforms too, not just the URL handle. patreon_resolver.resolve_display_name reads the campaigns API's attributes.name; SubscribeStarClient.resolve_display_name pulls the creator name off the profile page (og:title, else the <title> stripped of the SubscribeStar suffix). extension_service._resolve_artist_name dispatches per platform (pixiv=token, patreon/subscribestar=cookies via get_cookies_path), best-effort in an executor, falling back to the readable URL handle on any failure. Still all curator core — the extension is unchanged (sends only the URL). gallery-dl platforms keep the handle (readable, no native client). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
0c4b8aef8c |
feat(artist): pixiv display-name at add-time + identity-by-source (#130 steps 2+3)
Final piece of the artist decoupling. (1) Identity-by-source: quick_add_source resolves the artist by an existing (platform, url) Source first, so a re-add reuses the artist even after it was renamed (its frozen slug no longer matches the name) — a slug-based lookup would have duplicated it. (2) Pixiv naming: a new pixiv source resolves the real display name via the app API (PixivClient .resolve_display_name → /v1/user/detail) using the stored token, so the artist is 'Kurotsuchi Machi' not '12345678' — and its name-derived slug matches what a native download produces, unifying them. Falls back to the numeric id when no token/crypto. ExtensionService gains the crypto seam; the endpoint passes it. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
a69bd1baa8 |
feat(artist): move a source into another artist (#130 step 4)
Operator ask: a surface to merge new sources into existing artists (consolidate the singleton artist a fresh add spins up). Enabled by the #130 slug decoupling — the storage path is immutable, so re-attribution moves NO files. SourceService .reassign moves the source, re-points its posts (Post.source_id==S) and the images it contributed (ImageProvenance via S, scoped to the old artist so shared images aren't stolen), and deletes the old artist if it's left fully empty (else clears its subscription flag). POST /api/sources/<id>/reassign. Frontend: a 'Move…' action per source on the artist Management tab → artist-autocomplete picker → confirm → routes to the target (whose slug is stable). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
87d53db0cb |
feat(artist): editable display name + rename surface; drop name-uniqueness (#130 step 1)
First step of decoupling artist identity/storage/display. migration 0077 drops uq_artist_name so the display name is free text (two genuinely different creators can share a name); the slug stays the immutable, unique storage/identity key (the on-disk path component — untouched, so nothing moves). ArtistService.rename + PATCH /api/artists/<id> change the name ONLY. Frontend: inline pencil-edit on the artist header (mirrors TagCard), slug/route unaffected so no navigation. Fixes the operator's 'no surface to rename an artist' + the name-collision fragility. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
8838b325fb |
fix(recapture): link on-disk images to their post (#1288)
Recapture disk-skips already-downloaded media, and upsert_post_record only writes Post fields — so a pre-existing image (e.g. one pulled under the old gallery-dl path, imported bare with no post) stays orphaned even after its post record is (re)written. Confirmed on the operator's instance: 329 pixiv images with primary_post_id NULL, 694 pixiv posts with content but no linked images, 0 duplicate posts. Fix: the recapture relink channel now carries the media's post_id (2- → 3-tuple path/url/post_id), and phase 3 calls importer.link_existing_image_to_post — match the on-disk image by path, find its Post by (source, external_post_id), upsert image_provenance + primary_post_id. Factored the provenance-linking out of _apply_sidecar into a shared _attach_provenance so the fresh-import and recapture-backlink paths can't diverge. Idempotent; generic across native platforms (no-op for already-linked Patreon/SubscribeStar). Re-running recapture now repairs orphaned images; future walks never orphan. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
437bf4d37a |
feat(suggestions): group wip/banner/editor under a separate 'system' category
System tags are kind=general, so their suggestions previously landed in the General group. Give them their own 'system' suggestion category so the operator reviews them apart from content tags: _current_heads maps is_system heads to category 'system' (still trained as general heads, still gated by the 0.65 floor). Frontend: CATEGORY_ORDER/LABELS gain 'system'; SuggestionsPanel renders a 'System' group first (small, collapsible, open — false positives easy to spot and reject); the typed-dropdown shows the shield icon for system entries. Safe: system-tag suggestions always carry a canonical_tag_id, so the create-by-kind path (which would send 'system' as a TagKind) is never hit. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
f33808b977 |
fix(pixiv): capture ugoira frame timings in the post record (ordering bug)
The core writes the post record BEFORE extract_media, but the ugoira frame delays were only memoized DURING extract_media — so write_post_record never saw them and ugoira_frames was always empty in the record. Extract a memoized _ugoira_meta (frames + zip url share ONE /v1/ugoira/metadata call regardless of order) and inject client.fetch_ugoira_frames into the downloader (mirrors Patreon's content_fetcher) so write_post_record populates the frames itself. Zero extra API calls — the fetch is shared/memoized with extract_media. A recapture now backfills the timings onto existing ugoira posts. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
6c6e8bdb6d |
feat(heads): surface system-tag suggestions at a flat 0.65 confidence floor
System tags (wip/banner/editor) already get heads (kind=general) and aren't filtered from suggestions, but they surfaced only at each head's precision-tuned suggest_threshold — high enough to hide the borderline/false-positive guesses the operator wants to SEE and REJECT (hard-negative mining: 'negatively reinforce what isn't a system tag'). score_image now uses a flat _SYSTEM_TAG_SUGGEST_FLOOR (0.65, operator-set) for system-tag heads instead of their auto threshold; content-tag heads keep their own, and the typed-dropdown threshold_override still overrides everything. _current_heads carries Tag.is_system into the head meta to drive it. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
c9696a2faf |
fix(subscribestar): route .art creators to .adult; clear source failure on disable
Two pre-merge fixes: 1. SubscribeStar .art age wall: the 18+ cookie doesn't clear the age gate on the .art domain (keeps 302'ing to /age_confirmation_warning even with the cookie — Elasid #54116), but the same creator is reachable on .adult where the cookie works. _normalize_ss_host rewrites subscribestar.art → subscribestar.adult at request time (stored Source.url untouched), logged so it's visible in walk logs. .com/.adult pass through. 2. Disabling a source now clears its failure state (last_error, error_type, consecutive_failures) so subs you pause (not paying for) stop lingering as 'failing'. Only the explicit disable clears — an unrelated edit to an already-disabled source leaves state alone. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
544e30bfb9 |
fix(pixiv): match gallery-dl's exact on-disk filename to avoid re-download at cutover
The native downloader used the Windows-safe sanitize_segment, but gallery-dl on Linux (path-restrict auto→'/', path-remove default control chars, path-strip auto→'') replaces ONLY '/' and deletes control chars — the Windows-forbidden set (<>:"|?*) and trailing dots/spaces stay RAW in on-disk titles. Any pixiv title with those chars would therefore miss the tier-2 disk-skip and re-download the whole work at cutover (seen-ledger starts empty). Replace sanitize_segment with gdl_clean_filename, a byte-exact mirror of gallery-dl 1.32.5 build_filename (verified against path.py). Directory + template already matched; this closes the last parity gap. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
a7f715ec43 |
test: stub ctx gains the auth_token key run_download now reads
The real phase-1 ctx has always carried auth_token; the native branch now threads it into the adapter constructors, so the stub ctx must match the contract (kept the strict ctx[...] read — it catches exactly this drift). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
0bbcdee3bd |
feat(pixiv): flip dispatch to the native ingester (#129 step 4)
pixiv joins NATIVE_INGESTER_PLATFORMS: download/verify/preview and the recover/recapture UI actions now route through PixivIngester. Campaign id is parsed straight from the source URL (numeric user id — no network resolver), with a platform-aware resolution-failure message. auth_token now rides the uniform adapter construction (token platforms use it, cookie platforms accept-and-ignore), and the preview endpoint fetches/threads it. The legacy gallery-dl pixiv path is fully removed (PLATFORM_DEFAULTS entry + the refresh-token config branches in download/verify) per no-legacy policy; gallery-dl keeps hentaifoundry/discord/deviantart until they migrate/retire. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
4068a97392 |
test(pixiv): fix downloader tests — skip validation on stub bytes, no media extraction in post-record tests
The stub payload is PNG bytes regardless of target extension, so the real validator quarantined the .jpg cases; and extracting the ugoira work hit the API seam of a fake session with no .post. Validation/quarantine plumbing stays covered by the Patreon downloader tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
0563b2d750 |
feat(pixiv): ledger models + migration 0076 + PixivIngester adapter (#129 step 3)
pixiv_seen_media / pixiv_failed_media mirror the Patreon/SubscribeStar ledgers (keys are always synthesized <illust_id>:p<num> / <illust_id>:ugoira — pximg URLs carry no content hash). PixivIngester wires client/downloader/ ledgers into ingest_core with drift label 'Pixiv app API' and the new body_canary=False opt-out: caption-less pixiv artists are common, so the zero-bodies #862 alarm would false-positive here — the client's response-shape drift checks cover that failure class instead. auth_token joins the uniform adapter constructor (pixiv is the first token-auth native platform). verify_pixiv_credential = one OAuth refresh, no feed walk. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
7ef2ecd82f |
feat(pixiv): native downloader — gallery-dl layout parity + enriched post record (#129 step 2)
PixivDownloader writes originals to the exact pre-cutover gallery-dl layout
(<artist_slug>/pixiv/pixiv/{id}_{title[:50]}_{NN}.{ext} — flat, double
platform segment) so tier-2 disk-skip recognizes existing files. Post-first:
per-media sidecar is identity-only; the post record (_post_<id>.json — id
suffix because the flat layout would collide a bare _post.json) carries the
enrichment: tags + EN translations, rating from x_restrict, series,
view/bookmark/comment counts, AI flag, dimensions, author, and ugoira frame
delays (the zip has no timings). i.pximg.net media GETs ride the app-header
profile (403 without the app-api Referer).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM
|
||
|
|
86ae396914 |
feat(pixiv): native app-API client — gallery-dl-parity profile, post-first seams (#129 step 1)
PixivClient mirrors gallery-dl 1.32.5's PixivAppAPI request profile exactly (iOS app headers, OAuth refresh with X-Client-Time/X-Client-Hash, /v1/user/illusts pagination via next_url — whose query string doubles as the resumable page cursor). Post-first seams (post_record_key / post_is_gated / post_meta) + extract_media covering multi-page, single-page, ugoira zip (600x600→1920x1080 swap, frame delays memoized for the post record), and limit_* placeholder gating. No PHPSESSID web fallback: FC holds only the refresh token, same effective coverage as the gallery-dl path. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
65bd1c22c3 |
test: whole-table tag counts become non-system counts
The four remaining run-1895 failures were stale expectations, not predicate bugs — prune/reset returned the right counts, but these tests verified no-deletion by counting the ENTIRE tag table (or asserting the full kind set), which now includes the three seeded hygiene tags that survive prunes and resets by design. Filter is_system=false with a pointer to #128 so future system tags cannot re-break them. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
f77e75147d |
feat(tags): system-tag UI markers + full protection sweep (step 4 of #128)
UI: shield marker + tooltip on TagChip and TagCard; system tags hide rename/merge/delete affordances (chip kebab entirely — set-fandom never applies to their general kind; remove stays, un-tagging is normal use). Aliases stay available: mapping model outputs ONTO a system tag is useful. Directory cards carry is_system. Every destructive path that could take out a system row is now guarded, found by sweeping run 1891s off-by-three failures — each one was a surface that would have eaten the seeded tags: - prune-unused: predicate exempts is_system (they ship with zero applications and matched every unused condition) - reset-content: predicate exempts is_system AND keeps their applications — hygiene flags describe the file, not content tagging - admin tag DELETE: refused with system_tag error - normalize_existing_tags: scan excludes is_system — canonicalization would recase wip -> Wip behind TagService.rename's guard, breaking the name-keyed presentation lookup Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
723f023e6a |
feat(gallery): similar() hides presentation images (banner / editor screenshot)
Step 3 of milestone #128. Presentation-tagged images cluster on UI chrome rather than content, so near any one of them they fill the whole more-like-this grid. Excluded from candidates in the ONE whole-image similarity surface (gallery similar mode, explore walk, and RelatedStrip all ride GalleryService.similar) — the anchor itself may be a banner, and wip stays surfaced: only the training pipelines exclude it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
19744fa41d |
fix(tests): resync serial sequences after baseline restore
TRUNCATE ... RESTART IDENTITY resets every sequence to 1, and the baseline restore re-inserts seeded rows WITH their explicit ids — leaving each sequence pointing below MAX(id). Harmless while the only baseline rows lived in tables tests never sequence-insert into (ml_settings id=1); migration 0075 seeded tag rows and every Tag insert after the first truncate collided on pk_tag id=1 (205 failures, run 1888 — find_or_create then surfaced it as NoResultFound via its conflict-recovery re-select). setval every restored table with a serial id column past its restored rows. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
e6f128c894 |
feat(ml): training hygiene — system-tagged images are absent from other concepts training
Step 2 of milestone #128. _hygiene_excluded_ids (training_data.py) is the one shared predicate: images carrying any system tag are dropped from every OTHER concepts head training — not positives (a rough wip tagged as a character drags the head toward generic-sketch) and not rejection or sampled negatives (a wip OF character X is not evidence against X). A system tags own head trains on them unfiltered; that is what makes auto-flagging banners work. Selection is split out of train_head as the sklearn-free head_training_ids so CI (no sklearn) can pin the behavior. CCIP: reference prototypes skip hygiene-tagged images — a faceless wip figure region must never become an identity reference — and the ref cache signature now counts hygiene applications, since tagging an image wip changes the reference set without touching character/region counts. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
e9891ee9f3 |
feat(tags): system tags — is_system column, seeded hygiene tags, protection guards
Training hygiene step 1 (milestone #128). Migration 0075 adds tag.is_system and seeds wip / banner / editor screenshot (kind=general), ADOPTING an existing same-(name,kind) tag case-insensitively instead of duplicating. These rows drive the upcoming training exclusions, so they are protected: rename and merge-away refuse system tags (merge-INTO stays allowed — folding an operator's old hygiene tag into the system row is the intended move; merge is the only tag-delete path, so that guard covers deletion). is_system rides every tag serialization. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
581b778528 |
fix(modal): accept-by-known-id keeps the raw suggestion row identity for the drop
Spreading canonical_tag_id onto a raw suggestion changed its _keyOf identity, so _dropEverywhere missed the actual list row and the panel kept showing an already-accepted suggestion. Pass the resolved id as an option instead; pinned with a raw-suggestion spec. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
6d314d662f |
feat(modal): applied tags drop from search instantly; manual add == accepting a suggestion
Three tag-flow gaps in the view modal (and the Explore workspace, which shares TagPanel): - the type-to-add dropdown now filters both its sections against the imageledger applied tags reactively, so a just-added tag disappears from search the moment the chip rail updates instead of after a modal refresh - manually picking or creating a tag the model also suggested routes through the suggestion-accept flow: the acceptance is recorded for head training and the row leaves the panel, instead of the add silently bypassing the feedback loop - removing a tag reloads the suggestion lists, so a model-suggested tag returns to the suggestions area (flagged rejected, one-click reversible) rather than vanishing until the next modal open Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
b54243a1ff |
fix(subscribestar): inject the 18+ age cookie on every SubscribeStar domain
The cookie was pinned to .subscribestar.adult only; cookies are domain-scoped, so sources on subscribestar.art (Elasid, event #54116) never sent it and every poll 302d to /age_confirmation_warning. Emit one line per domain (.com/.adult/.art) with a per-domain presence check, and admit .art in the platform url_pattern. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
aa12a57f97 |
feat(recovery): surgical re-fetch for deep posts via ExternalLink reset
Operator-flagged: the recovered defective files live DEEP in their artists'
back-catalogues — the normal download cadence (by design, via the seen-gates)
will never re-walk them, so recovery's source re-check alone can't bring them
back. The durable per-post handle is the ExternalLink row, which survives the
image delete:
- services/external_links.refetch_links_for_post: reset settled links to
pending (fresh attempt budget, in-flight left alone) + dispatch their
fetches; sha-dedupe at import discards payload files that still exist, so
only the missing file lands.
- recover_defective_image now captures the image's post ids BEFORE the delete
cascades provenance away and resets those posts' links — future recoveries
are surgical automatically (response gains links_reset; source re-check
stays for gallery-dl-native files within walk reach).
- POST /api/admin/posts/refetch-external {external_post_id, source_id?} — the
manual tool for the three files recovered before this fix existed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM
|
||
|
|
b1cfbcc06a |
fix(agent): sleep mode — an empty queue sheds to one downloader and backs the lease poll off to 15 min
Operator-flagged on the deployed .5 build: the autoscaler grew the pool 1→8 against an EMPTY queue (an empty buffer read as 'GPU starving' regardless of WHY), and every downloader kept polling lease every 10s all night. - New idle signal straight from the lease results: an empty lease sets _idle, any jobs clear it. The occupancy-low branch now distinguishes three cases: queue empty → shed to ONE polling downloader; pinned at the bandwidth cap → shed toward 3; cap headroom + work flowing → grow. - Idle lease polls back off exponentially per downloader to IDLE_POLL_MAX_SECONDS (15 min) and reset the moment work appears — so an idle night costs one HTTP call per 15 min, and new work is noticed within at most ~15 min (operator-accepted trade-off). - UI hint: 'idle — queue empty, lease poll backed off'; /status gains idle. Agent build 2026-07-02.6. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
f08a64f7ae |
style(ia): wave 4 — one section-header language across every admin surface
The Subscriptions Settings tab's bare text-h6 headers adopt the same uppercase accent section-title + hint convention Maintenance/Cleanup use, with a one-line hint per section (extension / credentials / downloader / external file-hosts / schedule defaults). Every settings-ish surface now reads identically. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
dc7fa6eae2 |
feat(ia): wave 3 — Subscriptions landing answers 'what needs me, what came in?'
Daily-use reorder of the Subscriptions tab: needs-attention strip first (FailingSourcesCard moves up from below the Downloads fold — a broken subscription was invisible unless you went looking), then a new Recent arrivals card (real downloads only, no-change scans filtered out, artist links), then the source list. Both cards render nothing when there's nothing to say. Retry logic moves into the downloads store (retrySource / retryAllFailing) so the needs-attention card and the Downloads maintenance menu share one implementation — single-retry forces past cooldown, bulk keeps cooldown enforcement, same tally shape. The card's Logs button deep-links into the Downloads tab pre-filtered (?source_id now watched, not just read on mount). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
e039689eff |
feat(ia): wave 2 — Activity becomes the whole-app pulse; Overview gets the health strip
The Activity tab only knew Celery — the GPU agent (the majority of processing) and the download pipeline were invisible there. Two new self-polling panels: - GpuActivityPanel: queue depths + triage verdicts (defects / file-ok / unprobed, top reason buckets) with a jump to Maintenance -> Failed processing. The triage detail refetches only when the error count moves. - DownloadsActivityPanel: 24h stat chips + failing-source names with a jump into Subscriptions. Both panels join the Activity tab under Queues+workers AND double as the Overview health strip (side-by-side grid under the Celery summary) — one component set, so Overview answers 'is everything healthy?' across all systems. SystemStatsCards reviewed: content still accurate, left as-is. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
5b34c9221c |
feat(ia): wave 1 — Import tab dissolves, Maintenance regroups by system, one extension home
Settings IA per the approved A3 design (the old layout was the two-app merge fossilized): - Import tab retired: ImportTriggerPanel + ImportTaskList deleted (manual /import scans stay API-level; imports arrive via downloads/extension, heal via the Layer-2 auto-refetch sweep, and show in Activity). ImportFiltersForm moves to Maintenance → 'Ingestion & filters' and loads its own settings; the import store shrinks to settings-only (no remaining consumers of the scan/task-list machinery). Overview's pending banner now points at Activity. - Maintenance regrouped: Ingestion & filters / GPU agent & embeddings (GpuAgent, Failed processing, CPU embedding backfill) / Tagging (sliders, Heads, Aliases) / Library health (MissingFiles, Thumbnails, DB, Archive re-extract demoted last) / Storage. - One extension home: BrowserExtensionCard moves from Settings → Overview to Subscriptions → Settings, above the API key bar it authenticates. - Single-color import filter WIRED: skip_single_color/threshold existed since FC-2 but nothing read them (the audit module's docstring said as much) — now enforced on both import paths via the audit's canonical predicate (tolerance 30, matching the Cleanup card default; animated images exempt like the transparency check). Default stays off; test added. - Dead weight: PlaceholderView (zero refs) and the permanently-disabled 'Export failed logs (CSV — v2)' menu stub deleted; stale docs fixed (celery queue docstring, threshold comment citing retired tasks, ml package docstring, HeadsCard 'replaces Camie' blurb). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
19b962f1a7 |
feat(b3): ml-worker becomes optional — embed-only role, decoupled GPU coordination, cpu-embed switch
The ml-worker's ONLY processing role is now the CPU whole-image embed fallback (tag_and_embed renamed embed_image — Camie tagging was retired #1189 and the name kept implying otherwise; videos were already handled agent-style: frame sampling + mean-pool). Detection/cropping/CCIP stay GPU-agent-only, and their completion is judged per-pipeline: ccip by gpu_job rows, siglip by concept regions at the current model version — never by image_record.siglip_embedding. A CPU embed therefore can NEVER close crop work for the agent (regression test pins this; only the whole-image 'embed' job, the same artifact, is satisfied). Making removal actually safe (operator will drop the container): - GPU-queue coordination (enqueue_gpu_backfill, recover_orphaned_gpu_jobs, reprocess_gpu_jobs) moved verbatim to tasks/gpu_queue.py on the maintenance quick lane — it lived on the 'ml' queue only by module colocation, which made the ml-worker a hard dependency of the whole agent pipeline. - New ml_settings.cpu_embed_enabled (migration 0074, default ON so agent-less installs keep working): OFF stops the four import hooks queueing embed work nothing will consume and no-ops the manual backfill; switch lives on the renamed 'CPU embedding backfill' card. - NB heads training / auto-apply still run on the ml image (sklearn) — a stack that removes the container gives those up too. Deploy note: in-flight messages under the old task names are dropped by the new workers; the 60s orphan sweep + hourly backfill re-fire under the new names immediately. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
7c19ad91ed |
feat: cap-aware autoscaler + token-gated whole-instance tag reset (operator feedback)
Autoscaler (agent 2026-07-02.5): the buffer-occupancy signal alone would peg downloaders at DL_MAX while the bandwidth CAP — not concurrency — is the real constraint (8 streams sharing 8 MB/s move no more data than 4). Growth is now gated on the pipe having headroom (net < 85% of cap) and a pipe pinned at the cap (>= 95%) sheds streams down to 3; dead band prevents flapping. The UI hint says 'holding at the bandwidth cap' and /status reports bw_capped, so the behavior is legible without tests that need the ML stack. Reset content tagging: stays a FULL-instance reset (operator's call), but now lives in a fenced 'Danger zone' section on Cleanup and the apply is gated by a preview-derived confirm token (mirrors the Tier-C bulk-delete pattern — stale counts are rejected server-side). Copy no longer claims suggestions repopulate: it says plainly the heads' training examples are deleted and re-tagging starts fresh. Moved out of TagMaintenanceCard into DangerZoneCard. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
eaea4308fc |
chore: retire the tag-eval harness — it proved the heads system, job done (operator-approved)
The head-vs-centroid eval (#1130) existed to prove the 'frozen embedding + trained head' spine; the operator accepted the tagging system and dropped the harness. Removed per rule 22: TagEvalCard + store, /api/tag_eval blueprint, tag_eval_run ml task, recover-stalled-tag-eval-runs sweep + beat entry, TagEvalRun model + table (migration 0073), and its tests. The eval's data loaders + metric helpers were NOT eval-specific — the nightly heads trainer runs on them — so they moved verbatim to services/ml/training_data.py (heads.py import updated; behavior unchanged). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
a7abcc41ca |
feat(triage): failed-processing triage — probe errored files, flag defects, recover (#125 C1-C3)
An errored GPU job's stored reason is a suspicion; the file probe is the
verdict. A 15-min beat sweep (triage_gpu_errors) runs verify_integrity's own
probe (sha256 + decode) on each errored image ONCE and writes both verdicts:
ImageRecord.integrity_status and the new GpuJob.triage_status ('defect' |
'file_ok', migration 0072). Every classification logs at WARNING so it
surfaces in Logs/System Activity.
- 'defect' rows are excluded from /retry_errors (re-running a known-bad file
burns agent time re-minting the tombstone); response now reports
defects_kept and the GpuAgentCard toast says so.
- GET /api/gpu/errors: triage view — reason buckets (classify_reason),
probe verdicts, per-job detail. POST /errors/triage runs the sweep now.
- POST /api/gpu/errors/<id>/recover: reuses the Layer-2 refetch pattern —
delete the defective copy + record (full cascade takes the tombstones too)
and re-poll its subscription Source so a fresh copy re-imports and re-enters
the pipeline; 'no_source' when nothing pollable resolves.
- New 'Failed processing' card (GpuTriageCard) in Maintenance: verdict counts,
reason summary, probe-now, defect list with thumbnails + per-image Recover.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM
|
||
|
|
1f27189b8f |
chore: retire ml-backfill-daily beat + the spent purge-legacy action (operator-approved)
- ml-backfill-daily: the CPU tag_and_embed backfill raced the GPU agent's daily embed backfill for the same NULL-embedding images at ~100x the cost (B1 audit verdict, milestone #124). The backfill TASK stays — the manual /api/ml/backfill button remains the deliberate CPU fallback pending B3. - purge-legacy: one-time IR-migration cleanup, dry-run verified 0 targets on the live library before removal (A2 audit, milestone #123). Fully retired per rule 22: tile, store action, route, service fn, tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
95d2ae1d58 |
feat(agent): global bandwidth cap — the agent can't saturate the desktop's network
One shared TokenBucket (default 8 MB/s; BANDWIDTH_LIMIT_MB_S, 0 = unlimited; live MB/s dial + net readout in the control UI) is charged by every still download (streamed chunk reads) and every ffmpeg video stream (metered from outside via /proc/<pid>/io and SIGSTOP/SIGCONTed into budget). Why: D1 re-measurement 2026-07-02 — the idle link moves ~38 MB/s, but 8 unthrottled downloaders bufferbloated it to ~1-1.5 MB/s PER STREAM (operator's browser included). Capping the aggregate keeps the desktop usable and still beats the collapsed sweep throughput it replaces. Agent build 2026-07-02.4. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
31c416bc7b |
docs(beat): backfill comments no longer claim errored jobs are retried
Follow-through on the tombstone rule (
|
||
|
|
09e2772628 |
fix(gpu-jobs): end the error-tombstone loop — deliberate retry semantics + poison-job guards
The hourly ccip backfill's skip-list lacked 'error' (and the daily
siglip/embed variants re-gated failures on their missing results), so every
permanently-bad file got a fresh doomed job each run — ~24 duplicate error
rows/day per file, the perpetual 'unprocessable' flood. An errored job is now
a TOMBSTONE: no backfill re-enqueues it; retry is deliberate-only via
/retry_errors (an errored back-catalogue needs one button press after a
model swap).
One shared set of dedupe DELETEs (services/ml/gpu_jobs.error_dedupe_statements)
runs before every backfill and inside /retry_errors: error rows made moot by a
later pending/leased/done row go first, then older duplicates (newest reason
survives) — so the error count reads as distinct failing files and a retry
can't fan one file out into duplicate pending jobs. /retry_errors now returns
{requeued, pruned} and the toast shows both.
Poison-loop guards (release and lease-expiry burn no attempts, so a job that
stalls its transfer or crashes the agent every time cycled forever —
operator-observed jobs 99044/125288/131594/143131):
- agent: 3 in-session transient bounces (fetch or submit) → fail with the real
reason instead of another release; strikes never count while stopping, and
clear on submit success. Agent build 2026-07-02.3.
- server: the 60s orphan sweep (statements shared between the beat task and
GpuJobService so they can't drift) converts expired leases with >=5 lease
grants and pending jobs with >=10 to 'error', preserving the last stored
failure reason. Backstops old agent builds.
Tests: tombstone rule across all three backfill variants, moot-row pruning,
poison conversions, and the extended /retry_errors dedupe contract.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM
|
||
|
|
3d6201734c |
fix(settings): maintenance tiles start collapsed; remember manual open state
GpuAgentCard was hardcoded :open=true, HeadsCard opened whenever any head existed, TagEvalCard whenever a persisted run existed — so a fresh Settings load greeted the operator with several tiles already expanded. All three now force-open only while their task is actually running (the #877 resurface behavior on the busy-driven tiles is untouched). MaintenanceTile additionally persists MANUAL expand/collapse per tile in localStorage, so the section reloads the way the operator left it; a forced open while a task runs stays transient and is never saved as a preference. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM |
||
|
|
1b1d3732dc |
feat(agent): store ffmpeg's actual failure reason in the job's error field
sample_frames_from_url now returns (frames, reason) — reason carries the SPECIFIC cause on failure (ffmpeg's stderr tail, e.g. "moov atom not found", or the timeout) instead of only logging it agent-side. The worker folds it into the failure it reports, so curator's GpuJob.error reads e.g. no frames sampled from video — ffmpeg exit 183: moov atom not found ... instead of the bare "(unprocessable)". The errored-jobs list becomes self-describing: after a retry sweep, surviving errors name their real defect without needing the agent log. Return-value plumbing (not shared state) so concurrent downloaders stay isolated. Agent VERSION → 2026-07-02.2. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
686808d3f3 |
feat(gpu): "Retry errored jobs" — scoped requeue of errors only
After an agent-side fix (e.g. the short-video sampler), the errored jobs
(~2.8k) have exhausted their 3 attempts and stay parked: backfill skips
images that already have a job, and /reprocess is the nuclear option (it
resets the 179k DONE jobs too). There was no way to re-run just the errors.
POST /api/gpu/retry_errors resets every status='error' job (all task types)
to pending with attempts=0 and the stored error cleared — a small inline
UPDATE that returns {requeued: n} so the UI toast can show the count.
UI: a "Retry errored jobs" button on the GPU-agent card, right under the
queue tiles; disabled when errored==0. With the agent now logging ffmpeg's
stderr on failure, retrying also reveals which errors were real vs victims
of the fps-filter bug.
Test: retry_errors requeues the errored job (fresh attempts, error cleared)
and leaves done work untouched; asserts via column selects (Core-DML
gotcha), not ORM refresh.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
3a683d7feb |
fix(agent): short videos failed as "unprocessable" — fps filter emits 0 frames
Root cause of the "no frames sampled from video (unprocessable)" flood
(operator-flagged 2026-07-02, whole 62k-70k image block + others): the
sampler used `-vf fps=1/4`, and ffmpeg's fps filter emits round(duration/4)
frames — which is ZERO for any clip shorter than ~2s. Short animation loops
(0.5s, 1.75s — verified against two originals from different artists) are
complete, valid h264 videos; ffmpeg decoded them fine, emitted no frames,
exited 0, and the agent failed the job as unprocessable. Long videos worked,
so only the short-clip class flooded.
Fix: sample with select ("first frame always, then one per interval of
timestamp") + -fps_mode vfr, and scale=out_range=full so limited-range
yuv420p sources don't trip the mjpeg encoder's full-range strictness
(secondary failure observed on a 4440x2760 clip). Verified locally against
both failing originals (frames extracted, PIL-clean) and a synthetic 15s
video (4 frames at t=0/4/8/12 — long-video behavior unchanged).
Observability (why this hid for weeks): ffmpeg's stderr was discarded, so
every failure logged only "no frames sampled". stderr now goes to a temp
file and its tail is logged on any produced-no-frames/timeout failure — the
log names the actual ffmpeg reason from now on. Also: frames written before
a mid-stream ffmpeg error are now kept (partial > nothing).
VERSION → 2026-07-02.1.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
0a618db10c |
fix(agent): Status froze after one update — conchint.textContent destroyed #capn
THE root cause of "the Status section doesn't update" (chased across several
rounds; the backend was always healthy). `#capn` (the max-concurrency number)
was nested inside `#conchint`:
<div id=conchint>… · max <b id=capn>8</b></div>
and applyStatus() ran, every call: `capn.textContent=CAP` AND
`conchint.textContent = '…max '+CAP`. Setting conchint.textContent replaces
ALL of conchint's children — destroying the <b id=capn> node. So:
call 1: capn exists → tiles update → conchint.textContent DELETES capn
call 2+: `capn.textContent` → "capn is not defined" (ReferenceError) →
applyStatus throws on its FIRST line → aborts before any tile →
frozen.
This is exactly the observed "ticks a couple times then freezes", and why
/gpu + /logs (which never touch capn) kept updating fine.
The capn write was redundant anyway — conchint.textContent already renders
the max. Remove the nested <b id=capn> element and the capn.textContent line;
the hint still shows "· max N". VERSION → .10.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
713a11e394 |
fix(agent): server-side rate metrics + killable-on-stop ffmpeg
Two follow-ups from live debugging of "work/min never populates" and "stopped never reached". 1) jobs/min + downloads/min are now computed in the BACKEND on a fixed cadence (_rate_loop, EWMA) and reported ready-to-show. The rates were derived client-side from poll deltas with a dt<30s guard — but a backgrounded/unfocused browser tab throttles its timers to ~1/min, so every delta exceeded 30s and the guard blanked the rates forever. A server-side rate is independent of how often the tab polls. Frontend just displays s.jobs_per_min / s.downloads_per_min. VERSION → .9. 2) ffmpeg video sampling is now killable on Stop. A downloader stuck in a slow/reconnecting decode (observed: 47s, 230s for one video) couldn't see the stop signal until ffmpeg returned, so Stop detached still-running threads and work kept flowing long after — "stopped" that wasn't really stopped. sample_frames_from_url now runs ffmpeg via Popen and polls a `should_stop` callback every 0.5s, terminating (then killing) the process at once on Stop or the per-video timeout. A stop-killed job is handed back (transient), not failed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
6282e753a9 |
fix(agent): real start/stop state machine — kill the stuck "stopping" pill
The Status pill hung on "stopping" forever (operator-flagged 2026-07-01). Root cause: the backend had no lifecycle state — status() only returned running/stopped — so the UI FABRICATED "stopping" in JS as `!running && active>0`. That pill only cleared when the backend's `active` counter hit 0, but stop() (a) blocked the HTTP handler on lease-release calls to curator and (b) left `active>0` whenever a consumer wedged mid-submit/release to an overloaded curator → "stopping" that never resolved. Give the backend a real, truthful state it drives itself: stopped → starting → running → stopping → stopped - start(): → starting; a downloader flips it to running on its FIRST successful lease (so "running" means curator is actually answering, not just "Start was clicked"). If curator's down it honestly stays "starting". - stop(): → stopping; returns immediately (no handler block). A background monitor waits for the worker threads to actually exit, releases leases, then → stopped — bounded by STOPPING_TIMEOUT (20s) so a wedged submit can NEVER hold the UI in "stopping" again. In-flight work is handed back safely. - Buttons follow the real state (Start only from stopped; both disabled through the transition), so you can't fight a transition. - Log every Start/Stop button press (routes) and every transition (worker), so the Logs panel shows exactly what each button did. Frontend now trusts s.state (drops the active>0 hack); VERSION → .8. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
91ea06be79 |
feat(agent): Status shows smoothed jobs/min + downloads/min (replace jumpy gauges)
The buffer / on-GPU / downloader counts flip many times a second, so a 3s
status poll only ever samples noise — the tiles looked frozen (same value
twice) or random (wildly different), reading as "the Status section doesn't
update" when the backend was in fact live (operator-flagged 2026-07-01).
Replace the three instantaneous gauge tiles with two derived RATE tiles:
- jobs / min — GPU throughput, from the monotonic `processed` counter
- downloads / min — fetch throughput, from a new monotonic `downloaded`
counter (bumped when a job is decoded into the buffer)
Together they also show pipeline balance (dl/min > j/min ⇒ GPU-bound; the
reverse ⇒ GPU starved). Both are EWMA-smoothed over the poll deltas, clamped
at 0 (agent restart resets the counters), and skip a backgrounded-tab gap.
The still-useful instantaneous state is demoted, not lost: buffer stays as
the occupancy bar; downloaders/consumers/on-GPU move to the sub-line. `waited
out` (transient) gets promoted to a tile.
backend: worker.status() gains `downloaded`; `_bump(downloaded=)`.
frontend: retiled Status + rate math in applyStatus; VERSION → .7.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
98b2ac90dd |
refactor(agent): DRY pass on the GPU agent worker package
Consolidate genuine duplication in agent/fc_agent into single-source helpers (behavior-preserving; DRY Pass process #594): worker.py - _fail(jid, image_id, exc, verb) — 4 terminal "fail this job" blocks (downloader HTTP-fault + decode, consumer non-transient + generic). - _release(job_ids) (was _release_owned) — the one lease hand-back path; 6 inline release([jid])+unhold sites now route through it. - _stopped(stop_evt) + _abort_if_stopped(jid, stop_evt) — 4 stop-check -and-release blocks and every bare stop-check. - _timed(stage) contextmanager — ~8 monotonic()/_record() timing pairs; records only on clean exit, matching the old skip-on-raise behavior. - _ewma(prev, x, alpha) module fn — 3 EWMA updates in the autoscaler. client.py - _submit(path, payload) — submit / submit_embedding (retrying session). - _post_quiet(path, payload) — heartbeat / fail / release fire-and-forget. detectors.py - Proposers._top(detector, image, cap) — merges components() and panels(). config.py - _bool_env(name, default) — auto_start / auto_scale env parsing. Left alone (recorded): the xyxy→norm-xywh conversion duplicated across models.py/detectors.py (2 copies, independent wrapper modules — sharing would couple them), and the _ensure_embedder/_ensure_proposers pair (same lock shape, different concepts). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
8a0237eeea |
fix(agent): stream videos via ffmpeg-from-URL instead of downloading the whole file
The failing "poison" jobs were 800MB+ 4K VR videos: the agent pulled the ENTIRE file into memory (r.content) just to sample a few frames, which buffered ~1GB in RAM and — on any slow/contended media store — got cut off mid-download (ChunkedEncodingError), failed, and re-leased forever. Measured the media read at ~4–6 MB/s (raw off the share, curator out of the path), so no serving-layer tweak helps; the file simply shouldn't be fully downloaded. Environment-agnostic fix (works for any deployment, completes even when slow): - media.sample_frames_from_url(): point ffmpeg straight at curator's /images URL. It Range-reads only the video index + up to max_frames of content — never the whole file — and reconnect flags resume a dropped transfer instead of failing. Generous, env-tunable timeout (FFMPEG_TIMEOUT, default 1200s) = completion over speed. Removes the bytes-based sample_frames (dead once videos stream). - worker._download_decode: videos now stream (no fetch_image, no RAM blowup); stills still download+decode. On an ffmpeg miss, probe curator liveness (client.is_reachable) → fail the job if curator is up (unprocessable file, stops the infinite re-lease) vs release if curator is down (transient, survives a redeploy). Auth header passed so it works whether or not /images is gated. Build marker 2026-07-01.6. Refs issue #1225. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
aa0605585b |
fix(agent): log the REAL fetch/submit failure reason, not "curator unreachable"
"curator unreachable" was printed for every transient error, hiding whether a single file's transfer stalled (ReadTimeout — curator is up, that stream is slow) or curator itself is down (ConnectTimeout/ConnectionError) or errored (HTTP 5xx). Those need completely different fixes, and we've been diagnosing the download slowness blind. Add _transient_reason(exc) → a specific label (HTTP <code>, else the exception class: ReadTimeout / ConnectTimeout / ConnectionError / …) and use it in both transient paths: - downloader: "fetch failed job <id> (image <id>, ReadTimeout) — released, backing off" - consumer: "submit failed job <id> (<reason>) — released, re-lease later" Now the logs say which failure it actually is (and which image), so we can tell a slow/stalled transfer apart from an unreachable curator. Build marker 2026-07-01.5. Refs issue #1225. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
0fe1674753 |
perf(web): stream files in 4 MiB chunks + 4 hypercorn workers (fix 40s downloads)
The image library is on a CIFS/SMB share (mounted rsize=4 MiB, actimeo=1), and Quart's FileBody streams in 8 KiB chunks — so serving one large original was ~19k network round-trips to the storage server, i.e. 30–58s per download (operator-flagged). That's what starved the GPU agent (constant "curator unreachable" backoff) AND slowed the browser: every byte is read off CIFS and streamed through the Python app (no reverse-proxy sendfile), and only 2 hypercorn workers meant the agent + the browser's thumbnail grid queued behind each other. In-container fix, no new service: - Raise FileBody.buffer_size 8 KiB → 4 MiB in create_app, matching the mount's read size: one round-trip per read, ~500× fewer. buffer_size is the MAX read so small thumbnails still read in one gulp, and Range/mime/ETag/conditional handling lives on Response — all preserved. Guarded so a Quart-internal change can't break boot. - HYPERCORN_WORKERS default 2 → 4 so concurrent /images requests stop queuing. Expected: large-file transfers drop from ~40s toward link speed (a few seconds) for the agent and the browser. See issue #1223. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
c22f37d64d |
feat(gallery): sort by earliest post date across all posts (new default)
The gallery's newest/oldest sort keys off image_record.effective_date = COALESCE(primary post's post_date, created_at). The primary post is often the repost/download the file came from, so the grid led with download dates rather than when content was first posted (operator-flagged). Add a second materialized sort key, earliest_post_date = MIN(post_date) across ALL of an image's provenance posts (every post it appears in), else created_at — the original publish date. Mirrors the effective_date pattern so the sort stays a forward index scan. - alembic 0071: add earliest_post_date + index (DESC, id DESC); backfill created_at baseline then MIN over image_provenance ⋈ post. - importer: recompute earliest_post_date whenever a dated post is linked (MIN over the image's provenance, which now includes the just-added row). - gallery_service: new sorts posted_new / posted_old key off earliest_post_date; cursor + year/month grouping follow the active column transparently. - api: accept posted_new|posted_old; DEFAULT is now posted_new so the grid leads with original publish date. newest/oldest (effective_date) still available. - frontend: sort dropdown gains "Newest/Oldest post date" (default Newest post date); existing effective-date sorts relabelled "Newest/Oldest added". - tests: service test asserts posted_new/posted_old key off earliest_post_date; frontend default-sort omission test updated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
ccbb5cbc9e |
fix(agent): stop the downloader pool stampeding a slow curator (congestion collapse)
Operator hit an outage after the machine slept overnight: the agent showed "curator unreachable" in a loop while curator's API (lease) was actually fine and the browser could still load images — just slowly. Root cause is a feedback loop in the new pipeline: every download streams a full original through curator's single Python file-serving path, and the autoscaler grows DOWNLOADERS whenever the buffer is empty. When downloads are merely SLOW/failing, the buffer is empty for that reason — so the agent piled on more concurrent large-file GETs, saturating curator's web workers + NFS, which slowed curator (and its browser) further and produced more failures → more downloaders. Classic congestion collapse. - Failure-aware autoscaling: if transient download failures rose since the last decision, SHRINK the downloader pool toward the floor instead of growing — the empty buffer is caused by failures, not the GPU starving. It ramps back up only once downloads succeed again. - DL_MAX 24 → 8: 24 concurrent large-file downloads through one Python serving path is too many; 8 keeps a fast GPU fed without stampeding curator. - fetch_image timeout 180 → (10, 60): the read timeout is between-bytes, so a large-but-flowing download still completes, but a stuck/dead connection fails in 60s instead of hanging a downloader for 3 min and piling up stuck requests. Build marker 2026-07-01.4. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
ef3318aac1 |
feat(explore): more variance in the related rail (stronger MMR diversification)
Operator wants the Explore "related" rail to span more — the #1188 diversifier was tuned conservatively. Push all three knobs so it reaches further across clusters instead of clumping near the anchor: - MMR lam 0.55 → 0.40 — weight the diversity penalty harder (the main dial). - candidate pool min(200, max(limit*5, 60)) → min(400, max(limit*8, 100)) — a wider nearest-cosine pool so MMR has genuinely distinct neighbourhoods to pick from, not just the near-dupes. - pHash dup_threshold 6 → 8 — collapse more near-duplicate reposts/clones, freeing rail slots for distinct picks. Still deterministic (same set per image, just more spread) and relevance-anchored via the lam*sim-to-anchor term. Backend-only; no migration. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
7cdce0c474 |
feat(agent): temporal video dedup — drop near-duplicate frames before the GPU
Near-static videos are the dominant GPU load: sampled into up to 64 frames, each re-runs the whole detect→CCIP→SigLIP chain on ~identical content. Add a CPU perceptual-hash frame dedup upstream of the GPU so the redundant frames are never processed at all (not just their embeds). - media.dedupe_frames() + _dhash(): 8×8 difference-hash (64-bit) per frame; greedy keep — a frame survives only if its hash differs from every kept frame by >= min_distance bits (Hamming). A static run collapses to one frame; genuinely distinct scenes all survive. Order + frame_time preserved. - Called in worker._download_decode right after sample_frames, so it runs in the decode stage on the downloader thread (CPU) — the GPU consumers only ever see deduped frames, and buffered video items shrink (less RAM too). - Env-tunable FRAME_DEDUPE_DISTANCE (default 8; higher keeps more frames for brief localized changes an 8×8 hash can miss; 0 disables). Logs `video frames N→M` when it drops any, so video load reduction is visible. Complements the spatial per-frame crop dedup (2026-07-01.2); this is the temporal axis. Build marker 2026-07-01.3. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
eaae896858 |
feat(agent): dedupe near-duplicate crops before the SigLIP embed
Figure boxes are already NMS-merged (iou 0.6) and each YOLO detector self-NMSes, but the combined per-frame crop pile (figure→concept ∪ anatomy component→concept ∪ panel) was embedded with no cross-proposer dedup — so genuine near-duplicates slipped through (a figure box ≈ an anatomy component on a solo bust; overlapping booru head classes on one head), embedding the same region twice and burning a slot against max_regions. Add detectors.dedupe_crops(): a greedy, high-IoU (default 0.85), kind-aware pass over the pending (crop, template) list right before embed_batch — drop boxes that overlap ≥ iou within the same kind, keep the highest score. The high threshold is deliberate: it collapses only true near-identical boxes while preserving intentional nested crops across scopes (a whole figure vs a small head component sit well below it) and distinct kinds (concept vs panel). Env-tunable DEDUPE_IOU (≥1.0 disables). Runs on CPU before the GPU work, so it cuts both embed cost and region count. Temporal (cross-frame) dedup deferred. Build marker 2026-07-01.2. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
afef95a87d |
feat(agent): download/GPU producer-consumer pipeline + fix detector fuse crash
The agent workload is download-bound (download 400–5462ms vs GPU ~300–600ms), so the old N-slot serial chain (each slot: lease→download→decode→GPU→submit) left the fast GPU idle during every download. Rearchitect worker.py into a producer/consumer pipeline: downloader pool (autoscaled by BUFFER OCCUPANCY) → bounded queue → 1–2 GPU consumers (detect+embed→submit) - Downloaders are I/O-bound → many overlap; the autoscaler now tunes DOWNLOADER count by buffer fill (empty = GPU starving → add; full = outpacing GPU → add a 2nd consumer if it has util/VRAM headroom and lifts throughput, else trim). - Bounded buffer (12) = backpressure: a full buffer blocks downloaders, capping RAM + lease look-ahead. VRAM pressure sheds a consumer immediately. - Heartbeat thread keeps every held lease alive (buffered jobs wait on the GPU; curator's 180s TTL would otherwise reclaim them mid-buffer). - Preserves all resilience: lease exp-backoff, submit-path retry (#169), release-on-stop, region caps + video early-exit (#171). Stop drains BOTH pools and releases every held lease at once (single held-set as source of truth). - Consumers SHARE one embedder + proposers instance (a 2nd consumer adds concurrent inference, not N× VRAM — bounds the VRAM creep seen with N slots). - UI reworked for the pipeline: tiles show downloaders · buffer · on-GPU · processed · errors, a buffer-occupancy meter, and a consumers/waited-out line; the dial now tunes downloaders. Build marker 2026-07-01.1. Also fix the operator-flagged detector warning: yolo11n + the comic-panel model threw "'Conv' object has no attribute 'bn'" on every image (ultralytics' load- time Conv+BN fusion on a version-mismatched graph), silently disabling 2 of 3 crop proposers and spamming the log per image. Disable that fusion (unfused inference is correct, marginally slower) and permanently self-disable a proposer on the first inference failure instead of re-throwing forever. Refs milestone 122. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
83f1070a11 |
fix(agent): bound video GPU work — early-exit the frame loop at max_regions
Image 81602 turned out to be a 156 MB mp4, not a huge still: the agent samples up to 64 frames × ~32 regions/frame → ~2000 regions (the 413) and 64 frames of detect+CCIP+embed (the 38s). The MAX_REGIONS backstop (#171) only truncated the SUBMIT — the GPU work was already spent. Break out of the frame loop once accumulated regions reach max_regions, so a long video costs ~a few frames of GPU (~2-3s), not all 64 (~38s). The whole-image 'embed' task is unaffected (it mean-pools all frames and returns before this loop). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
c587ac667c |
fix(agent): cap figures + global region cap + reset active on stop
Three safety/robustness fixes from the operator's run logs: - Cap figures per frame (MAX_FIGURES, default 8) like components/panels already are. Uncapped, a huge/busy image yielded hundreds of figure boxes → hundreds of per-figure CCIP calls + crops → a 38s job AND a submit too big to accept (image 81602 looped on 413). This is the acute fix. - Global per-JOB backstop (MAX_REGIONS, default 128): if total regions still exceed the cap (long video), keep the highest-scoring and log the drop, so a submit body can never blow past curator's limit. - Stale "active" meter: stop() now resets _active to 0 (no slots remain, so the meter must read 0 at once), and _bump clamps at 0 so a slot finishing after the reset can't drive it negative. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
7e74fa767c |
fix(agent): load huge images — disable PIL decompression-bomb guard
Trusted local library, not an upload surface, so a legitimately large image (90–95M px, operator-flagged) must load. PIL only WARNS at the 89M-px default but RAISES DecompressionBombError at ~179M px, which would fail those jobs. Set Image.MAX_IMAGE_PIXELS = None. (The agent works off individual extracted files — curator's archive_extractor unpacks zip/cbz/rar/7z at import — so this is about big single images, not archives.) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
9f1148b110 |
chore(agent): modernise base → Ubuntu 24.04 / Python 3.12 / CUDA 12.9
Bump the GPU-agent base image from 12.4.1-cudnn-runtime-ubuntu22.04 (Python 3.10, CUDA 12.4, early-2024) to 12.9.2-cudnn-runtime-ubuntu24.04: - Ubuntu 24.04 LTS → Python 3.12 — one modern runtime, no more 3.10. - CUDA 12.9 + cuDNN 9 — current within the CUDA-12 / cuDNN-9 line that the default onnxruntime-gpu wheel AND torch cu124 are built against. NOT CUDA 13: ONNX Runtime's CUDA-13 support is still nascent (separate wheels + open "Unsupported CUDA version: 13" reports), and torch bundles cu124 anyway. The GPU (Ampere/Ada, 12 GB) is fine on either — this is a library-alignment call, not a hardware limit. - PIP_BREAK_SYSTEM_PACKAGES=1: 24.04 marks system Python externally-managed (PEP 668); a single-purpose container owns its environment, so global installs are fine and simplest. - agent/ruff.toml pinned to py312 (was py310) so CI lints against the real runtime; from __future__ import annotations stays (PEP 649 lazy annotations are 3.14, so self-refs still evaluate on 3.12). CI builds the image but has no GPU — validate on the desktop after pull that it starts and loads CUDAExecutionProvider (not CPU fallback). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
f01b59f390 |
fix(agent): py3.10 startup crash + submit-path retry; pin agent ruff to py310
The agent container (CUDA base, Python 3.10) crashed on startup with `NameError: name 'Config' is not defined` — an earlier `ruff --fix` unquoted the `from_env(cls) -> Config` self-reference, which is safe on CI's Python 3.14 (PEP 649 lazy annotations) but is evaluated at class-definition time on 3.10. CI lint/compile run on 3.14, so it slipped through. - config.py: `from __future__ import annotations` so the self-referential annotation is a string, never evaluated — works on 3.10 and every version. - agent/ruff.toml: pin the agent to `target-version = "py310"` (its real runtime) and inherit the root rules. Ruff now flags exactly this class as F821, so CI's lint lane catches it instead of shipping a broken image. (CI otherwise lints on 3.14, masking 3.10 issues.) - client.py: submit path now retries in-place. A dedicated session with a urllib3 Retry (connect/read/status, 0.5s backoff, 500/502/503/504, POST) so a momentary blip after the GPU work is done doesn't discard it and force a full re-download + recompute elsewhere. A duplicate submit after a lost response is a harmless 409 no-op. Lease/fetch keep the plain session + loop-level backoff. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
79269da802 |
fix(agent): prompt stop + lazy curator polling + build marker; add agent to CI
Addresses operator reports: Stop never finishes, the agent polls curator constantly, and stale-cached pages get mistaken for a failed deploy. - Stop is prompt: flip _running BEFORE any lock so /status + worker loops see "stopped" immediately, and add a stop/shrink checkpoint in _process (after decode, before the expensive detect+embed) that releases the job and bails — so a Stop doesn't wait out heavy GPU work. - Lazy curator polling: the queue snapshot is fetched only while a browser is actually watching (a /status hit within UI_IDLE_GRACE) and on a 5s cadence, not a constant background loop. The work loop's own lease/submit is curator's only visitor otherwise — nothing polls just to poll. - Build marker: VERSION is embedded in the page and reported on /status; the UI shows a "reload" banner when they differ, so a browser-cached page can't be mistaken for "the new image didn't deploy" (complements the no-store header). CI: the lint lane now also `ruff check`s agent/ and compileall-parses it, so the GPU agent is linted + syntax-checked before its image builds (build.yml only `docker build`s it). Fixed the agent's pre-existing UP037/B905 so it passes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
e6a7fe7d03 |
feat(agent): per-stage timing breakdown (lease/download/decode/gpu/submit)
Instrument the job pipeline so we can see where wall-clock actually goes and decide — on data, not theory — whether a download/compute split is worth building. Each stage is timed per job and a rolling breakdown is logged every 30s to the agent console, e.g.: timing/30s — lease 8ms · download 310ms · decode 40ms · gpu 165ms · submit 70ms | wall/job 585ms (214 jobs) - lease timed around client.lease() in the slot loop (per batch). - download = fetch_image; decode = image/frame decode; gpu = detect + CCIP + batched embed; submit = the results POST. One-time model load is excluded from the gpu figure. - Thread-safe accumulator (stage -> [sum, count]) summarised + reset by a small daemon reporter thread; logs only when there was work. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
181f1c6a27 |
perf(gpu-queue): partial indexes + two-phase lease so leasing stays O(batch)
The throughput bottleneck was curator-side, not the network. lease() claimed the lowest-id pending/expired jobs with `... ORDER BY id LIMIT n`, but with only a plain `status` index Postgres walked the primary key from id=1, skipping the entire prefix of already done/error rows before reaching pending ones. As `done` grew (69k+), every lease became an O(done) scan — leasing crawled, the DB saturated, and even /status (the queue GROUP BY count) stalled the agent. - Migration 0070 adds two partial indexes over just the live slice: pending rows indexed by id (hot path), and leased rows by lease_expires_at (crash-recovery + orphan sweep). They stay tiny no matter how large the done/error history. - lease() split into two phases so each uses a partial index: claim pending first (id-ordered, O(batch)); reclaim expired leases only when pending can't fill the batch. Same semantics (SKIP LOCKED, attempts++, expired reclaim). - Model __table_args__ declares the indexes so ORM and schema agree. - Test: a done-prefix at low ids must not stop the lease reaching pending. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
f0f031782d |
fix(agent): unfreeze status view + smoothed throughput-aware autoscaler + log pane
Operator: the status tiles (state/active/processed) and the Start/Stop buttons freeze while the GPU meters stay live. Root cause: /status made an INLINE blocking curator call (queue_status) on every poll, and with curator buried under a 112k-job backlog that call stalled — freezing the whole status refresh (the GPU bars survived because /gpu is a lock-free local read). Made worse by the old util-band autoscaler, which grew workers toward the 32 cap forever because util plateaus ~50% on this IO-bound load and never hit the 70 grow threshold — piling load onto curator and the agent process. - /status is now a pure in-memory read: worker.status() is lock-free, and the curator queue snapshot is refreshed by a background poller (never inline). - Autoscaler replaced with a smoothed, throughput-aware climb that SETTLES: samples util every 2s and EWMA-smooths it (raw util swings 0↔99), then every ~24s grows by one only while each grow keeps lifting smoothed jobs/s; when a grow stops helping it backs off one and holds, re-probing occasionally. No runaway, no flopping. - GPU util bar now shows a smoothed value: the agent's own EWMA (util_smooth, exposed on /gpu) when running, else smoothed client-side — so it glides instead of bouncing 0↔99. - act() aborts a slow Start/Stop POST after 8s so the buttons can't stick; the now-always-fast /status refresh recovers state regardless. - Log pane: bound the page to the viewport (height:100vh) so the Logs card scrolls INTERNALLY instead of overflowing off-screen; cap the ring buffer at 400 lines. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
82e1a4e127 |
fix(agent): send Cache-Control: no-store so a new image isn't masked by cache
The control page is a static string served with no cache headers, so after pulling a fresh agent image the browser kept showing the OLD UI until a hard refresh (operator-flagged). Add a no-store middleware covering the page and the status/gpu/logs polls. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
c2e9157822 |
feat(agent): graceful Start/Stop with starting/stopping states + instant status
Operator: the buttons fire but the status view doesn't reflect the change. Cause: act() ignored the POST's own status response and waited on the separate /status poll (which lags behind the curator queue call). Now: - act() applies the POST's returned status immediately for instant feedback, and shows an optimistic "starting"/"stopping" state (pulsing, buttons disabled) the moment it's clicked. - A stop that still has in-flight jobs draining shows "stopping" until active hits 0, then resolves to "stopped" on its own. - applyStatus() guards the /status-only fields (connection pill + queue) so the lean action response can't blank them — the Start/Stop path deliberately skips the slow curator call to stay snappy. Also de-duplicate GPU reads: read_gpu() now caches (1s TTL) with one probe at a time, and /status no longer spawns its own nvidia-smi — so the fast /gpu poll + autoscaler + /status share a single subprocess instead of piling up in the server thread pool (which was what made clicks feel dead under load). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
3b34230fbd |
fix(agent): stable util-band autoscaler + live GPU meters
Two operator-reported issues with the GPU agent: 1. Worker count flopped almost every cycle, spiking the GPU. The hill-climb probed +1, judged it over a too-short noisy throughput window, saw no clear gain and reverted -1 — every tick. Replace it with a GPU-utilization-band controller: HOLD while smoothed util sits in a healthy band, grow only on clear spare capacity (util below the low mark + VRAM headroom), shrink under saturation or memory pressure. Util is EWMA-smoothed and decisions are spaced (DECIDE_EVERY samples), so a noisy nvidia-smi reading can't move the pool. Load stays consistent instead of probe/reverting. 2. GPU util/VRAM bars only updated on manual refresh. They rode the /status poll, which blocks on the curator queue call (slow when curator is busy), so the meters froze between refreshes. Give them a dedicated /gpu endpoint (local nvidia-smi only, no curator round-trip) polled every 1.5s, and drop the curator queue-status timeout 15s -> 5s so /status itself stays snappy. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
c259d03618 |
fix(agent): revert full-width page, grow the Logs section to the bottom
Operator meant the LOG section should fill down the viewport (vertical), not the whole page going full-width horizontally. Restore the centered column (820px), make .wrap a full-height flex column, and let the Logs card flex to fill the remaining height to the bottom (drop the fixed 230px log-pane cap). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
2713c3f773 |
perf(agent): batch SigLIP crop embeds per image + load truncated images
Two issues surfaced by the live logs (GPU pegged at ~0% util, 0.5 jobs/s,
truncated-image failures):
- BATCH the SigLIP embeds: collect all of an image's crops (figure + booru_yolo
components + panels) and embed them in ONE forward pass instead of one
forward+lock per crop. The per-crop path serialised every crop through the
inference lock and starved the GPU (≈0% util, autoscaler stuck oscillating);
batching gives a real GPU-bound workload + far higher throughput. CCIP still
runs per figure inline.
- LOAD_TRUNCATED_IMAGES in the agent (matches the server embedder): slightly-
truncated scraped images now load instead of failing the job 3× then erroring
("image file is truncated (N bytes not processed)").
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa
|
||
|
|
9eaefac385 |
feat(agent): full-width control page, Copy-logs button, quiet HTTP log noise
- Page fills the viewport horizontally (drop the 780px cap). - Copy button on the Logs card → copies the console (clipboard API on localhost, textarea-execCommand fallback), with a brief "Copied" confirmation. - Silence httpx/httpcore/huggingface_hub/urllib3/filelock/uvicorn.access/ ultralytics to WARNING so the console shows agent activity (detector loads, job errors, autoscale moves) instead of per-request HF-download spam. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
c1b099e5a3 |
feat(agent): in-UI log console + a real styling pass on the control page
- logbuf.py: bounded in-memory log ring buffer + a logging.Handler on the root logger; GET /logs serves it; the control page polls it into a console pane — so runs are monitorable without `docker logs`. worker now logs autoscale moves (one line per change, with jobs/s + util + VRAM) and job failures (job + image + reason); detectors already log load/disable. - Restyled the whole control page: a proper dark layout with a header + live connection pill, cards (Control / Status / Logs), a styled Auto switch + worker stepper, status tiles, separate GPU-util and VRAM meters, and the log console. No longer feels like an afterthought; all the existing control hooks are preserved. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
6d7b17b0b5 |
feat(agent): autoscale the worker count (throughput hill-climb), Auto default-on
The new per-job workload (3 detectors + several SigLIP embeds) is far more GPU-bound than the old I/O-bound CCIP pass, so the right worker count shifted and is hard to guess. Add an Auto mode (default ON) that finds it: - _control_loop samples jobs/sec + GPU util/VRAM every ~6s and hill-climbs the target: grow while throughput keeps improving and VRAM stays under budget, revert a step that doesn't help, back off under memory pressure (VRAM >= 90%), then settle and periodically re-probe (the GPU/IO balance shifts over a run). - A manual concurrency set is an override → leaves Auto; an "Auto" toggle in the control UI re-enables it. status() reports `auto`; the dial reflects the auto-chosen count (read-only) while Auto is on. - AUTO_SCALE env (default on) + compose doc. Agent py-compiled (outside CI). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
359bc5a283 |
feat(ml): default to SigLIP 2 (new installs) + model dropdown, no free-text (#1203)
- Migration 0069: new installs default to SigLIP 2 (so400m, 512px, 1152-d drop-in) — UPDATE applies ONLY where no image is embedded yet (fresh install), so an existing library is NOT silently invalidated; it switches deliberately via the dropdown → Re-embed → Retrain. Column server_defaults moved to SigLIP 2. - GET /api/ml/embedder-models: server-authoritative supported list (SigLIP 2 512 recommended / 384 faster / SigLIP 1 384 original) so the UI never free-types. - GpuAgentCard: the two name/version text fields → a single model dropdown; Save sets name+version from the picked option (the current model is always selectable even if off-list). - embedder.py DEFAULT_MODEL_NAME unchanged (stays the baked local-dir SigLIP 1) to avoid a local-dir/weights mismatch; SigLIP 2 loads by HF name, cached on the ml-worker's persistent HF_HOME. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |
||
|
|
80f8eb4756 |
feat(gpu): re-process trigger to apply new crop detectors to the existing library (#1202)
The siglip/ccip backfills skip images that already have current-version regions, so adding crop detectors only affected NEW images — the back-catalogue would never be re-cropped. Add a reprocess trigger that resets every done/error job of a task back to pending, so the agent re-runs the FULL pipeline (figure detection + CCIP + concept/panel crops) over the whole library under the current detectors. - reprocess_gpu_jobs(task='ccip') task + POST /api/gpu/reprocess. - gpu store reprocess() + GpuAgentCard "Re-process library (re-detect + re-crop)" button with a confirm (it's heavy). - Test: a done job resets to pending (attempts cleared). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa |