Compare commits

..

114 Commits

Author SHA1 Message Date
bvandeusen 18d5c05639 Merge pull request 'fix(ml): per-task async engine for recompute_centroid (#881)' (#114) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
Build images / build-ml (push) Successful in 8s
Build images / build-web (push) Successful in 7s
CI / frontend-build (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 30s
CI / integration (push) Failing after 3m19s
2026-06-16 20:24:37 -04:00
bvandeusen 11ddfc3876 Merge pull request 'fix(maint): resurface dedup/gated-purge results after navigate-away (#877)' (#113) from dev into main
Build images / sign-extension (push) Successful in 2s
CI / lint (push) Successful in 2s
Build images / build-ml (push) Successful in 6s
CI / frontend-build (push) Successful in 18s
Build images / build-web (push) Successful in 15s
CI / backend-lint-and-test (push) Successful in 29s
CI / integration (push) Successful in 3m20s
2026-06-16 16:48:52 -04:00
bvandeusen 2b8ce86622 Merge pull request 'Gated Patreon posts: skip on ingest + cleanup tool (#874)' (#112) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 4s
Build images / build-ml (push) Successful in 8s
Build images / build-web (push) Successful in 11s
CI / frontend-build (push) Successful in 20s
CI / backend-lint-and-test (push) Successful in 34s
CI / integration (push) Successful in 3m27s
2026-06-16 15:25:34 -04:00
bvandeusen 49bee77cdc Merge pull request 'Video tag quality: cadence sampling + min-frame aggregation + ML thread cap (#747)' (#111) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 2s
Build images / build-ml (push) Successful in 8s
Build images / build-web (push) Successful in 10s
CI / frontend-build (push) Successful in 18s
CI / backend-lint-and-test (push) Successful in 30s
CI / integration (push) Successful in 3m15s
2026-06-16 14:08:40 -04:00
bvandeusen c209e3b37e Merge pull request 'Tier-1 video dedup: import-time + retroactive cleanup (#871)' (#110) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
Build images / build-ml (push) Successful in 7s
CI / frontend-build (push) Successful in 17s
CI / backend-lint-and-test (push) Successful in 27s
Build images / build-web (push) Successful in 26s
CI / integration (push) Successful in 3m16s
2026-06-16 08:55:38 -04:00
bvandeusen cffdd93418 Merge pull request 'Nested-archive extraction (#718) + post-first ingest (#67) + post-body canary (#862)' (#109) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
Build images / build-ml (push) Successful in 8s
Build images / build-web (push) Successful in 8s
CI / frontend-build (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 28s
CI / integration (push) Successful in 3m19s
2026-06-15 21:37:39 -04:00
bvandeusen fd84be40dd Merge pull request 'External-attach orphan fix (#859) + image-less post display' (#108) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
Build images / build-ml (push) Successful in 8s
Build images / build-web (push) Successful in 10s
CI / frontend-build (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 40s
CI / integration (push) Successful in 3m19s
2026-06-15 01:55:36 -04:00
bvandeusen 79f510d7f8 Merge pull request 'Merge dev → main: read post body from content_json_string (the empty-body fix) (#842)' (#107) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 2s
Build images / build-ml (push) Successful in 10s
Build images / build-web (push) Successful in 8s
CI / frontend-build (push) Successful in 19s
CI / backend-lint-and-test (push) Successful in 37s
CI / integration (push) Successful in 3m24s
2026-06-15 00:19:45 -04:00
bvandeusen 59181069da Merge pull request 'Merge dev → main: per-post stdout diagnostics + post-field write DRY (#842/#753)' (#106) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
CI / frontend-build (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 41s
Build images / build-web (push) Successful in 2m15s
Build images / build-ml (push) Successful in 2m55s
CI / integration (push) Successful in 3m11s
2026-06-14 23:45:39 -04:00
bvandeusen 428ecd8642 Merge pull request 'Merge dev → main: per-post body-capture diagnostics in the event UI (#842)' (#105) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
Build images / build-ml (push) Successful in 8s
Build images / build-web (push) Successful in 10s
CI / frontend-build (push) Successful in 22s
CI / backend-lint-and-test (push) Successful in 37s
CI / integration (push) Successful in 3m11s
2026-06-14 23:14:14 -04:00
bvandeusen ed1e04b831 Merge pull request 'Merge dev → main: recapture body-fetch fix + diagnostics (#842)' (#104) from dev into main
Build images / sign-extension (push) Successful in 2s
CI / lint (push) Successful in 3s
Build images / build-ml (push) Successful in 7s
Build images / build-web (push) Successful in 8s
CI / frontend-build (push) Successful in 19s
CI / backend-lint-and-test (push) Successful in 35s
CI / integration (push) Successful in 3m15s
2026-06-14 22:31:41 -04:00
bvandeusen f5156bd847 Merge pull request 'Merge dev → main: Recapture mode (#842) — re-grab post bodies/links + localize on-disk inline images' (#103) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
CI / frontend-build (push) Successful in 22s
CI / backend-lint-and-test (push) Successful in 32s
Build images / build-web (push) Successful in 2m30s
Build images / build-ml (push) Successful in 3m27s
CI / integration (push) Successful in 3m29s
2026-06-14 21:21:49 -04:00
bvandeusen dfc3922d24 Merge pull request 'Merge dev → main: #830 rich post capture + external-host downloads (+ #768/#789/#739)' (#102) from dev into main
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 3s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 48s
Build images / build-web (push) Successful in 2m21s
Build images / build-ml (push) Successful in 2m47s
CI / integration (push) Successful in 3m20s
2026-06-14 19:16:09 -04:00
bvandeusen 3eb08e926b Merge pull request 'fix(aliases): modal raw-key bug + alias visibility/management' (#101) from dev into main
CI / lint (push) Successful in 2s
Build images / sign-extension (push) Successful in 3s
Build images / build-ml (push) Successful in 7s
Build images / build-web (push) Successful in 10s
CI / frontend-build (push) Successful in 19s
CI / backend-lint-and-test (push) Successful in 26s
CI / integration (push) Successful in 3m13s
2026-06-12 14:02:00 -04:00
bvandeusen 9e81ced359 Merge pull request 'fix(images): percent-encode original-image URLs ('#' in paths 404'd)' (#100) from dev into main
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 3s
Build images / build-ml (push) Successful in 7s
Build images / build-web (push) Successful in 8s
CI / frontend-build (push) Successful in 20s
CI / backend-lint-and-test (push) Successful in 36s
CI / integration (push) Successful in 3m15s
2026-06-12 00:44:56 -04:00
bvandeusen 11e9f5af60 Merge pull request 'fix(browse): tabs and search on one row' (#99) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
Build images / build-ml (push) Successful in 7s
CI / frontend-build (push) Successful in 24s
CI / backend-lint-and-test (push) Successful in 29s
Build images / build-web (push) Successful in 11s
CI / integration (push) Successful in 3m13s
2026-06-12 00:28:38 -04:00
bvandeusen 909fa37b15 Merge pull request 'Browse search + series numbering rework + kebab fix' (#98) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
CI / frontend-build (push) Successful in 22s
CI / backend-lint-and-test (push) Successful in 28s
Build images / build-web (push) Successful in 2m17s
Build images / build-ml (push) Successful in 2m50s
CI / integration (push) Successful in 3m16s
2026-06-12 00:14:31 -04:00
bvandeusen dfab8f65ff Merge pull request 'feat(series): flat sequence + cosmetic dividers + pending staging (#789)' (#97) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
CI / frontend-build (push) Successful in 29s
CI / backend-lint-and-test (push) Successful in 38s
Build images / build-web (push) Successful in 2m13s
Build images / build-ml (push) Successful in 2m49s
CI / integration (push) Successful in 3m9s
2026-06-11 22:07:57 -04:00
bvandeusen 618f7cdc36 Merge pull request 'feat(ml): drop image_record.tagger_predictions — image_prediction is sole store (#768 step 3)' (#96) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
CI / frontend-build (push) Successful in 22s
CI / backend-lint-and-test (push) Successful in 25s
Build images / build-web (push) Successful in 2m29s
Build images / build-ml (push) Successful in 3m0s
CI / integration (push) Successful in 3m13s
2026-06-11 19:32:05 -04:00
bvandeusen 028ea33a7c Merge pull request 'fix(migration): make 0045 DDL-only; backfill image_prediction via batched task (#768)' (#95) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
CI / frontend-build (push) Successful in 19s
CI / backend-lint-and-test (push) Successful in 38s
Build images / build-web (push) Successful in 2m50s
Build images / build-ml (push) Successful in 2m58s
CI / integration (push) Successful in 3m15s
2026-06-11 09:22:22 -04:00
bvandeusen 444c1fb075 Merge pull request 'perf(migration): 0045 streams json_each (no materialize / no temp blowup)' (#94) from dev into main
Build images / sign-extension (push) Successful in 3s
Build images / build-web (push) Successful in 2m14s
Build images / build-ml (push) Successful in 2m36s
CI / lint (push) Successful in 2s
CI / frontend-build (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 32s
CI / integration (push) Successful in 3m8s
2026-06-10 22:07:16 -04:00
bvandeusen 26c68b0a75 Merge pull request 'fix(migration): 0045 guards json_each against scalar tagger_predictions' (#93) from dev into main
CI / lint (push) Successful in 2s
CI / frontend-build (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 27s
CI / integration (push) Successful in 3m13s
Build images / sign-extension (push) Successful in 2s
Build images / build-web (push) Successful in 6s
Build images / build-ml (push) Successful in 2m40s
2026-06-10 20:33:27 -04:00
bvandeusen e75427b19a Merge pull request '#768 steps 1+2: normalized image_prediction table (read cutover)' (#92) from dev into main
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 3s
CI / frontend-build (push) Successful in 22s
CI / backend-lint-and-test (push) Successful in 27s
Build images / build-web (push) Successful in 2m4s
Build images / build-ml (push) Successful in 2m42s
CI / integration (push) Successful in 3m8s
2026-06-10 20:15:26 -04:00
bvandeusen 5447fab987 Merge pull request 'Activity search + RecoverySweep fix + tagger_predictions shrink (#762, #764) + backup polish' (#91) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
CI / frontend-build (push) Successful in 19s
CI / backend-lint-and-test (push) Successful in 38s
Build images / build-web (push) Successful in 2m23s
Build images / build-ml (push) Successful in 2m51s
CI / integration (push) Successful in 3m10s
2026-06-10 14:33:35 -04:00
bvandeusen bad37e07b2 Merge pull request 'Browse hub, series rename, full-prediction dropdown + a DRY pass (7 sweeps)' (#90) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 2s
CI / backend-lint-and-test (push) Successful in 27s
CI / frontend-build (push) Successful in 44s
Build images / build-ml (push) Successful in 3m2s
CI / integration (push) Successful in 3m15s
Build images / build-web (push) Successful in 2m17s
2026-06-10 00:24:01 -04:00
bvandeusen 2bfc9936a1 Merge pull request 'UI batch: tagging flow, series browse, fandom chips, nav' (#89) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
CI / frontend-build (push) Successful in 26s
CI / backend-lint-and-test (push) Successful in 39s
Build images / build-web (push) Successful in 2m21s
Build images / build-ml (push) Successful in 2m52s
CI / integration (push) Successful in 3m24s
2026-06-09 20:48:19 -04:00
bvandeusen 4c6406ee18 Merge pull request 'fix(posts): link duplicate items to every post + prune bare shells' (#88) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
CI / frontend-build (push) Successful in 24s
CI / backend-lint-and-test (push) Successful in 28s
Build images / build-web (push) Successful in 2m19s
Build images / build-ml (push) Successful in 2m41s
CI / integration (push) Successful in 3m5s
2026-06-08 19:42:31 -04:00
bvandeusen bb47e80b3e Merge pull request 'fix(cleanup): unused-tags delete must use the same predicate as the preview' (#87) from dev into main
Build images / sign-extension (push) Successful in 2s
CI / lint (push) Successful in 2s
Build images / build-ml (push) Successful in 7s
Build images / build-web (push) Successful in 7s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 26s
CI / integration (push) Successful in 3m13s
2026-06-08 18:17:49 -04:00
bvandeusen dc1083b5e0 Merge pull request 'Unused-tag fandom fix + ML-worker logging/tuning + unified dropdown Enter' (#86) from dev into main
CI / lint (push) Successful in 3s
CI / frontend-build (push) Successful in 18s
CI / backend-lint-and-test (push) Successful in 27s
CI / integration (push) Successful in 3m16s
Build images / sign-extension (push) Successful in 3s
Build images / build-ml (push) Successful in 8s
Build images / build-web (push) Successful in 2m21s
2026-06-08 17:27:55 -04:00
bvandeusen e46893fefd Merge pull request 'Migration lock safety + remove the merge's full-table scan (the real 0040-hang fix)' (#85) from dev into main
Build images / sign-extension (push) Successful in 2s
CI / lint (push) Successful in 2s
CI / frontend-build (push) Successful in 25s
CI / backend-lint-and-test (push) Successful in 25s
Build images / build-web (push) Successful in 2m3s
Build images / build-ml (push) Successful in 2m33s
CI / integration (push) Successful in 3m7s
2026-06-08 01:03:06 -04:00
bvandeusen 0666e15211 Merge pull request 'Series manage redesign (FC-6.4) + migration/normalize hardening + UX fixes' (#84) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 25s
Build images / build-web (push) Successful in 1m58s
Build images / build-ml (push) Successful in 2m28s
CI / integration (push) Successful in 3m6s
2026-06-07 21:17:57 -04:00
bvandeusen 747390631d Merge pull request 'FC-6 series authoring + backup/NFS hardening + UX fixes' (#83) from dev into main
Build images / sign-extension (push) Successful in 2s
CI / lint (push) Successful in 2s
CI / frontend-build (push) Successful in 18s
CI / backend-lint-and-test (push) Successful in 27s
Build images / build-web (push) Successful in 2m5s
Build images / build-ml (push) Successful in 2m43s
CI / integration (push) Successful in 3m6s
2026-06-07 19:31:04 -04:00
bvandeusen e0d2a20588 Merge pull request 'Modal focus/keyboard polish, Camie-in-autocomplete, re-extract self-resume' (#82) from dev into main
Build images / sign-extension (push) Successful in 2s
CI / lint (push) Successful in 3s
CI / frontend-build (push) Successful in 25s
CI / backend-lint-and-test (push) Successful in 28s
Build images / build-web (push) Successful in 2m2s
Build images / build-ml (push) Successful in 2m32s
CI / integration (push) Successful in 3m9s
2026-06-07 12:17:43 -04:00
bvandeusen 1d84f67418 Merge pull request 'Modal: large centered spinner + kebab z-index fix' (#81) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
CI / frontend-build (push) Successful in 27s
CI / backend-lint-and-test (push) Successful in 29s
Build images / build-web (push) Successful in 2m3s
Build images / build-ml (push) Successful in 2m35s
CI / integration (push) Successful in 3m7s
2026-06-07 10:57:00 -04:00
bvandeusen 91265df3d6 Merge pull request 'Maintenance-queue health + modal/tagging keyboard pass' (#80) from dev into main
Build images / sign-extension (push) Successful in 2s
CI / lint (push) Successful in 3s
CI / frontend-build (push) Successful in 20s
CI / backend-lint-and-test (push) Successful in 25s
Build images / build-web (push) Successful in 2m5s
Build images / build-ml (push) Successful in 2m26s
CI / integration (push) Successful in 3m1s
2026-06-07 10:31:01 -04:00
bvandeusen 11acdb0322 Merge pull request 'Patreon: a missing media file_name is a fallback, not API drift' (#79) from dev into main
Build images / sign-extension (push) Successful in 2s
CI / lint (push) Successful in 2s
Build images / build-ml (push) Successful in 7s
Build images / build-web (push) Successful in 7s
CI / frontend-build (push) Successful in 20s
CI / backend-lint-and-test (push) Successful in 27s
CI / integration (push) Successful in 3m2s
2026-06-06 23:14:25 -04:00
bvandeusen 2eb9fd5dd0 Merge pull request 'Patreon: enforce the backfill time-box mid-post (stop soft-limit overruns)' (#78) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
Build images / build-ml (push) Successful in 8s
Build images / build-web (push) Successful in 8s
CI / frontend-build (push) Successful in 19s
CI / backend-lint-and-test (push) Successful in 27s
CI / integration (push) Successful in 3m3s
2026-06-06 23:00:29 -04:00
bvandeusen 01e5ce1410 Merge pull request 'Patreon: resolve creator campaign from a single-post URL' (#77) from dev into main
Build images / sign-extension (push) Successful in 2s
CI / lint (push) Successful in 2s
Build images / build-ml (push) Successful in 7s
Build images / build-web (push) Successful in 7s
CI / frontend-build (push) Successful in 24s
CI / backend-lint-and-test (push) Successful in 26s
CI / integration (push) Successful in 3m4s
2026-06-06 21:38:33 -04:00
bvandeusen 3bb94674cf Merge pull request 'Patreon download concurrency cap + immediate backfill kickoff on create' (#76) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 2s
Build images / build-ml (push) Successful in 7s
Build images / build-web (push) Successful in 7s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 26s
CI / integration (push) Successful in 3m1s
2026-06-06 21:27:22 -04:00
bvandeusen a75c602175 Merge pull request 'Subscriptions UX overhaul + Patreon /cw/ vanity fix' (#75) from dev into main
Build images / sign-extension (push) Successful in 2s
CI / lint (push) Successful in 2s
Build images / build-ml (push) Successful in 7s
Build images / build-web (push) Successful in 9s
CI / frontend-build (push) Successful in 24s
CI / backend-lint-and-test (push) Successful in 28s
CI / integration (push) Successful in 3m2s
2026-06-06 21:12:02 -04:00
bvandeusen ef8f4f7193 Merge pull request 'Tag-casing acronym fix, Patreon resolver hardening, archive diagnostics, post-card strip' (#74) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 2s
Build images / build-ml (push) Successful in 8s
Build images / build-web (push) Successful in 10s
CI / frontend-build (push) Successful in 22s
CI / backend-lint-and-test (push) Successful in 26s
CI / integration (push) Successful in 3m1s
2026-06-06 19:59:32 -04:00
bvandeusen ec3d27b219 Merge pull request 'Tag-maintenance sweep + bug-fix batch: #699 #700 #701 #709 #711 #712 #713 #714' (#73) from dev into main
Build images / sign-extension (push) Successful in 2s
CI / lint (push) Successful in 3s
Build images / build-ml (push) Successful in 8s
Build images / build-web (push) Successful in 9s
CI / frontend-build (push) Successful in 18s
CI / backend-lint-and-test (push) Successful in 27s
CI / integration (push) Successful in 3m2s
2026-06-06 16:49:00 -04:00
bvandeusen 03bd3b2eda Merge pull request 'Native Patreon ingester + download-engine ownership (plans #697, #703–#708)' (#72) from dev into main
Build images / sign-extension (push) Successful in 2s
CI / lint (push) Successful in 3s
Build images / build-ml (push) Successful in 7s
Build images / build-web (push) Successful in 10s
CI / backend-lint-and-test (push) Successful in 25s
CI / frontend-build (push) Successful in 32s
CI / integration (push) Successful in 3m1s
2026-06-06 12:57:31 -04:00
bvandeusen 7395e77d75 Merge pull request 'Smarter backfill: time-boxed chunks, run-until-done (plan #693)' (#71) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
Build images / build-ml (push) Successful in 8s
CI / backend-lint-and-test (push) Successful in 13s
Build images / build-web (push) Successful in 10s
CI / frontend-build (push) Successful in 18s
CI / integration (push) Successful in 2m58s
2026-06-05 16:33:06 -04:00
bvandeusen 575d817919 Merge pull request '#70 dev→main: cursor-paged backfill + mobile modal fixes' from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
Build images / build-ml (push) Successful in 8s
CI / backend-lint-and-test (push) Successful in 13s
Build images / build-web (push) Successful in 9s
CI / frontend-build (push) Successful in 18s
CI / integration (push) Successful in 2m59s
2026-06-05 11:10:02 -04:00
bvandeusen 2a8f7cd8b6 Merge pull request '#69 dev→main: release v26.06.04.0' from dev into main
CI / lint (push) Successful in 2s
CI / backend-lint-and-test (push) Successful in 12s
CI / frontend-build (push) Successful in 19s
CI / integration (push) Successful in 3m2s
Build images / sign-extension (push) Has been skipped
Build images / build-ml (push) Successful in 6s
Build images / build-web (push) Successful in 6s
2026-06-04 23:16:12 -04:00
bvandeusen 83f8af8090 Merge pull request 'dev→main: surface near-duplicate (pHash) control + reorder import tab' (#68) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 2s
Build images / build-ml (push) Successful in 6s
CI / backend-lint-and-test (push) Successful in 11s
Build images / build-web (push) Successful in 10s
CI / frontend-build (push) Successful in 20s
CI / integration (push) Successful in 2m56s
2026-06-04 21:39:44 -04:00
bvandeusen 9a2617c1a2 Merge pull request 'dev→main: post-card redesign (images→modal, in-place text expand)' (#67) from dev into main
CI / lint (push) Failing after 2s
Build images / sign-extension (push) Successful in 3s
Build images / build-ml (push) Successful in 6s
CI / backend-lint-and-test (push) Successful in 11s
Build images / build-web (push) Successful in 9s
CI / frontend-build (push) Successful in 24s
CI / integration (push) Successful in 2m57s
2026-06-04 17:32:50 -04:00
bvandeusen 81688815a0 Merge pull request 'dev→main: similar-search render fix + reset-content-tagging + scan persistence' (#66) from dev into main
Build images / sign-extension (push) Successful in 2s
CI / lint (push) Successful in 2s
Build images / build-ml (push) Successful in 9s
CI / backend-lint-and-test (push) Successful in 13s
Build images / build-web (push) Successful in 11s
CI / frontend-build (push) Successful in 22s
CI / integration (push) Successful in 2m58s
2026-06-04 16:59:52 -04:00
bvandeusen 773128c3bf Merge pull request 'dev→main: purpose-built mobile layout for subscriptions hub' (#65) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
Build images / build-ml (push) Successful in 6s
CI / backend-lint-and-test (push) Successful in 14s
Build images / build-web (push) Successful in 9s
CI / frontend-build (push) Successful in 38s
CI / integration (push) Successful in 2m56s
2026-06-04 13:43:56 -04:00
bvandeusen ce7b154ae9 Merge pull request 'dev→main: subscriptions table mobile card layout' (#64) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 4s
Build images / build-ml (push) Successful in 8s
CI / backend-lint-and-test (push) Successful in 14s
Build images / build-web (push) Successful in 10s
CI / frontend-build (push) Successful in 22s
CI / integration (push) Successful in 2m57s
2026-06-04 12:56:58 -04:00
bvandeusen 9430a9d9c3 Merge pull request 'dev→main: gallery similarity search (Phase 3) + UI/mobile polish' (#63) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
Build images / build-ml (push) Successful in 8s
CI / backend-lint-and-test (push) Successful in 12s
Build images / build-web (push) Successful in 9s
CI / frontend-build (push) Successful in 21s
CI / integration (push) Successful in 2m56s
2026-06-04 11:21:26 -04:00
bvandeusen 23aee56ce3 Merge pull request 'dev→main: showcase cascade + filter styling + DB maintenance + gallery filter Phase 2 + showcase decode-gate + CI perf' (#62) from dev into main
Build images / sign-extension (push) Successful in 2s
CI / lint (push) Successful in 3s
Build images / build-ml (push) Successful in 9s
CI / backend-lint-and-test (push) Successful in 12s
Build images / build-web (push) Successful in 11s
CI / frontend-build (push) Successful in 20s
CI / integration (push) Successful in 3m0s
2026-06-04 08:27:47 -04:00
bvandeusen 711abea567 Merge pull request 'Gallery speed + fandom editing + filters + pinned filter bar' (#61) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 2s
Build images / build-ml (push) Successful in 8s
CI / backend-lint-and-test (push) Successful in 18s
CI / frontend-build (push) Successful in 19s
Build images / build-web (push) Successful in 11s
CI / intimp (push) Successful in 3m48s
CI / intapi (push) Successful in 7m45s
CI / intcore (push) Successful in 8m29s
2026-06-04 00:07:19 -04:00
bvandeusen 844bb86802 Merge pull request 'fix(download): release DB connections across the gallery-dl subprocess' (#60) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 2s
Build images / build-ml (push) Successful in 10s
Build images / build-web (push) Successful in 7s
CI / frontend-build (push) Successful in 19s
CI / backend-lint-and-test (push) Successful in 25s
CI / intimp (push) Successful in 3m46s
CI / intapi (push) Successful in 7m41s
CI / intcore (push) Successful in 8m24s
2026-06-03 22:11:36 -04:00
bvandeusen a8f6a464aa Merge pull request 'fix(download): salvage soft-time-limit kills + fix timeout ladder' (#59) from dev into main
CI / lint (push) Successful in 2s
Build images / sign-extension (push) Successful in 3s
CI / backend-lint-and-test (push) Successful in 21s
CI / frontend-build (push) Successful in 24s
Build images / build-web (push) Successful in 2m37s
Build images / build-ml (push) Successful in 3m21s
CI / intimp (push) Successful in 3m40s
CI / intapi (push) Successful in 7m47s
CI / intcore (push) Successful in 8m15s
2026-06-03 19:35:45 -04:00
bvandeusen ab9922ad2e Merge pull request 'feat(artist): "new since last visit" badge + banner' (#58) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 4s
CI / frontend-build (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 23s
Build images / build-web (push) Successful in 2m9s
Build images / build-ml (push) Successful in 2m53s
CI / intimp (push) Successful in 3m36s
CI / intapi (push) Successful in 7m32s
CI / intcore (push) Successful in 8m9s
2026-06-03 16:20:54 -04:00
bvandeusen 0533807669 Merge pull request 'feat(ext): verify cookies in-browser before uploading (1.0.7)' (#57) from dev into main
CI / lint (push) Successful in 3s
CI / frontend-build (push) Successful in 20s
CI / backend-lint-and-test (push) Successful in 32s
extension / lint (push) Successful in 14s
Build images / sign-extension (push) Successful in 3m17s
Build images / build-ml (push) Successful in 3m28s
CI / intimp (push) Successful in 4m26s
Build images / build-web (push) Successful in 3m42s
CI / intapi (push) Successful in 9m32s
CI / intcore (push) Successful in 10m18s
2026-06-03 14:17:42 -04:00
bvandeusen 279dff3fb6 Merge pull request 'feat(ml): normalize Camie suggestion names to human-readable' (#56) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
CI / backend-lint-and-test (push) Successful in 15s
CI / frontend-build (push) Successful in 17s
Build images / build-web (push) Successful in 2m17s
Build images / build-ml (push) Successful in 2m55s
CI / intimp (push) Successful in 3m36s
CI / intapi (push) Successful in 7m38s
CI / intcore (push) Successful in 8m24s
2026-06-03 13:18:44 -04:00
bvandeusen 37e66cddc4 Merge pull request 'chore(modal): drop ?image=N soft-compat — pure overlay' (#55) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 4s
CI / frontend-build (push) Successful in 27s
CI / backend-lint-and-test (push) Successful in 38s
Build images / build-ml (push) Successful in 2m46s
Build images / build-web (push) Successful in 2m36s
CI / intimp (push) Successful in 3m42s
CI / intapi (push) Successful in 7m25s
CI / intcore (push) Successful in 8m12s
2026-06-02 19:35:04 -04:00
bvandeusen 9cf6b2d363 Merge pull request 'audit-g5 final + ML threshold default + kebab menu fix' (#54) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 4s
CI / backend-lint-and-test (push) Successful in 23s
CI / frontend-build (push) Successful in 27s
Build images / build-web (push) Successful in 3m0s
Build images / build-ml (push) Successful in 3m45s
CI / intimp (push) Successful in 3m46s
CI / intapi (push) Successful in 8m8s
CI / intcore (push) Successful in 8m34s
2026-06-02 19:09:49 -04:00
bvandeusen 6ef0fed41f Merge pull request 'audit-g5: architectural debt — 4 bundles (A/B/C/D)' (#53) from dev into main
CI / lint (push) Successful in 2s
Build images / sign-extension (push) Successful in 3s
Build images / build-ml (push) Successful in 9s
CI / backend-lint-and-test (push) Successful in 18s
Build images / build-web (push) Successful in 10s
CI / frontend-build (push) Successful in 26s
CI / intimp (push) Successful in 3m51s
CI / intapi (push) Successful in 7m38s
CI / intcore (push) Successful in 8m52s
2026-06-02 18:07:25 -04:00
bvandeusen 89b48f8f35 Merge pull request 'audit-g4: status-enum miss batch' (#52) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
Build images / build-ml (push) Successful in 11s
CI / backend-lint-and-test (push) Successful in 13s
Build images / build-web (push) Successful in 10s
CI / frontend-build (push) Successful in 35s
CI / intimp (push) Successful in 3m41s
CI / intapi (push) Successful in 7m25s
CI / intcore (push) Successful in 8m10s
2026-06-02 16:15:00 -04:00
bvandeusen d60e0b9494 Merge pull request 'audit-g3: lifecycle batch — recovery sweeps, retention, timeouts' (#51) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 2s
Build images / build-ml (push) Successful in 9s
CI / backend-lint-and-test (push) Successful in 12s
Build images / build-web (push) Successful in 7s
CI / frontend-build (push) Successful in 36s
CI / intimp (push) Successful in 3m39s
CI / intapi (push) Successful in 7m39s
CI / intcore (push) Successful in 8m10s
2026-06-02 14:49:28 -04:00
bvandeusen 9c27a2d3c7 Merge pull request 'audit-g2: async race / state-leak fixes across eight stores' (#50) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
Build images / build-ml (push) Successful in 8s
CI / backend-lint-and-test (push) Successful in 23s
CI / frontend-build (push) Successful in 37s
Build images / build-web (push) Successful in 13s
CI / intimp (push) Successful in 3m52s
CI / intapi (push) Successful in 8m15s
CI / intcore (push) Successful in 8m50s
2026-06-02 14:17:12 -04:00
bvandeusen 93e37681b7 Merge pull request 'audit-g1: six one-liner drift fixes from 2026-06-02 audit' (#49) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 2s
Build images / build-ml (push) Successful in 11s
CI / backend-lint-and-test (push) Successful in 11s
CI / frontend-build (push) Successful in 19s
Build images / build-web (push) Successful in 10s
CI / intimp (push) Successful in 3m42s
CI / intapi (push) Successful in 7m28s
CI / intcore (push) Successful in 8m3s
2026-06-02 13:29:17 -04:00
bvandeusen 64ca858574 Merge pull request 'UX fixes: suggestion-accept chip refresh, showcase endless feed, non-media downloads' (#48) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 2s
Build images / build-ml (push) Successful in 9s
CI / backend-lint-and-test (push) Successful in 12s
Build images / build-web (push) Successful in 9s
CI / frontend-build (push) Successful in 20s
CI / intimp (push) Successful in 3m52s
CI / intapi (push) Successful in 7m45s
CI / intcore (push) Successful in 8m20s
2026-06-02 08:26:58 -04:00
bvandeusen 9d0c0b7da8 Merge pull request 'fix(thumbnails): surface backfill counts + tighten validity check' (#47) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
Build images / build-ml (push) Successful in 8s
Build images / build-web (push) Successful in 9s
CI / frontend-build (push) Successful in 18s
CI / backend-lint-and-test (push) Successful in 18s
CI / intimp (push) Successful in 3m45s
CI / intapi (push) Successful in 7m32s
CI / intcore (push) Successful in 8m4s
2026-06-01 22:34:38 -04:00
bvandeusen 8e4d252ae4 Merge pull request 'fix(downloads): enqueue thumbnail + ML per attached image' (#46) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
Build images / build-ml (push) Successful in 8s
Build images / build-web (push) Successful in 8s
CI / backend-lint-and-test (push) Successful in 19s
CI / frontend-build (push) Successful in 26s
CI / intimp (push) Successful in 3m40s
CI / intapi (push) Successful in 7m51s
CI / intcore (push) Successful in 8m46s
2026-06-01 21:52:42 -04:00
bvandeusen fdd3e01f56 Merge pull request 'ux(failing-sources): visible row separators + clearer hover' (#45) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
Build images / build-ml (push) Successful in 10s
Build images / build-web (push) Successful in 10s
CI / backend-lint-and-test (push) Successful in 20s
CI / frontend-build (push) Successful in 33s
CI / intimp (push) Successful in 3m44s
CI / intapi (push) Successful in 7m48s
CI / intcore (push) Successful in 8m16s
2026-06-01 21:09:28 -04:00
bvandeusen c82fb308b6 Merge pull request 'Post.source_id refactor + tick/backfill modes + PARTIAL classifier + Mux fix + UX' (#44) from dev into main
Build images / build-ml (push) Failing after 2s
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
Build images / build-web (push) Successful in 12s
CI / backend-lint-and-test (push) Successful in 19s
CI / frontend-build (push) Successful in 31s
CI / intimp (push) Successful in 3m47s
CI / intapi (push) Successful in 7m52s
CI / intcore (push) Successful in 8m16s
2026-06-01 20:44:23 -04:00
bvandeusen 8cf8d2ca4d Merge pull request 'Modal Esc/overflow polish, artist-scoped post scroll, failing-sources Logs button' (#43) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 8s
CI / frontend-build (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 31s
Build images / build-web (push) Successful in 2m45s
Build images / build-ml (push) Successful in 3m31s
CI / intimp (push) Successful in 3m48s
CI / intapi (push) Successful in 7m11s
CI / intcore (push) Successful in 7m47s
2026-06-01 12:10:11 -04:00
bvandeusen b1d58bc3b8 Merge pull request 'fix(ci): POSIX-safe SHORT_SHA in build.yml' (#42) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
CI / backend-lint-and-test (push) Successful in 15s
CI / frontend-build (push) Successful in 20s
Build images / build-web (push) Successful in 2m17s
Build images / build-ml (push) Successful in 3m24s
CI / intimp (push) Successful in 3m22s
CI / intapi (push) Successful in 7m51s
CI / intcore (push) Successful in 8m5s
2026-06-01 08:03:17 -04:00
bvandeusen 65386f02a0 Merge pull request 'View modal batch: autofocus, suggestions UX, post-title click, retire copyright/artist' (#41) from dev into main
Build images / sign-extension (push) Successful in 2s
CI / lint (push) Successful in 2s
Build images / build-ml (push) Failing after 4s
Build images / build-web (push) Failing after 3s
CI / backend-lint-and-test (push) Successful in 15s
CI / frontend-build (push) Successful in 17s
CI / intimp (push) Successful in 3m46s
CI / intapi (push) Successful in 7m17s
CI / intcore (push) Successful in 8m2s
2026-06-01 07:01:42 -04:00
bvandeusen 667b05f14e Merge pull request 'Extension probe-and-add (v1.0.6) + per-commit image tags' (#40) from dev into main
CI / lint (push) Successful in 2s
Build images / build-ml (push) Failing after 2s
CI / frontend-build (push) Successful in 19s
CI / backend-lint-and-test (push) Successful in 22s
extension / lint (push) Successful in 8s
Build images / sign-extension (push) Successful in 1m51s
Build images / build-web (push) Failing after 4s
CI / intimp (push) Successful in 3m38s
CI / intapi (push) Successful in 7m20s
CI / intcore (push) Successful in 7m45s
2026-06-01 01:44:03 -04:00
bvandeusen 856e9104b4 Merge pull request 'Sidecar synthetic anchor cleanup + tier-gated classifier fix' (#39) from dev into main
CI / lint (push) Successful in 5s
CI / frontend-build (push) Successful in 22s
CI / backend-lint-and-test (push) Successful in 28s
CI / intimp (push) Successful in 3m54s
Build images / sign-extension (push) Has been skipped
Build images / build-web (push) Successful in 1m1s
Build images / build-ml (push) Successful in 1m21s
CI / intapi (push) Successful in 7m31s
CI / intcore (push) Successful in 8m0s
2026-06-01 00:16:58 -04:00
bvandeusen 0397642b21 Merge pull request 'Showcase cadence tuning + cooldown-aware bulk retry' (#38) from dev into main 2026-05-30 23:50:36 -04:00
bvandeusen 237575447d Merge pull request 'Thumbnail URL fix + archive daemon fix + batched initial loads' (#37) from dev into main 2026-05-30 22:01:43 -04:00
bvandeusen ed358757dc Merge pull request 'Most-overdue-first scheduling + rich timeout diagnostics' (#36) from dev into main 2026-05-30 14:30:41 -04:00
bvandeusen d181f4afb8 Merge pull request 'Downloads burst-prevention + maintenance-menu fix + gdl timeout' (#35) from dev into main 2026-05-30 11:43:18 -04:00
bvandeusen 2886fa4997 Merge pull request 'Tooltip !important fix — 104cac5 follow-up after Vite CSS reorder' (#34) from dev into main 2026-05-30 00:02:24 -04:00
bvandeusen f256f587ee Merge pull request 'UI batch + I1–I6 service passes + download-event recovery sweep' (#33) from dev into main 2026-05-29 22:46:16 -04:00
bvandeusen 384d8d5e50 Merge pull request 'Dashboard insights + project-wide DRY pass' (#32) from dev into main 2026-05-28 15:38:26 -04:00
bvandeusen 319e8c1d18 Merge pull request 'v26.05.28.0: downloads dashboard + task-resilience overhaul (timeouts, archive split, 3-layer poison-pill defense)' (#31) from dev into main 2026-05-28 00:45:00 -04:00
bvandeusen 9075d8eadd Merge pull request 'v26.05.27.2: subscribestar + HF cookie quirks, platforms package refactor, showcase IR-parity, secure-context audit' (#30) from dev into main 2026-05-27 21:34:02 -04:00
bvandeusen 88e53e5b86 Merge pull request 'v26.05.27.1: subscriptions hub + post-card merge + sidecar audit' (#29) from dev into main 2026-05-27 17:12:48 -04:00
bvandeusen 37e8b796a1 Merge pull request 'v26.05.27.0: PostCard redesign + IR-style tag suffix + drop meta/rating + extension v1.0.4 CSP fix' (#28) from dev into main 2026-05-27 11:31:18 -04:00
bvandeusen 4e82208926 Merge pull request 'v26.05.26.5 — extension CORS unblock + UI gap closes + CI workflow cleanup' (#27) from dev into main 2026-05-26 20:15:07 -04:00
bvandeusen 52fff00353 Merge pull request 'v26.05.26.4 — hotfix: migration 0022 pre-DELETE colliding ImageProvenance before UPDATE' (#26) from dev into main 2026-05-26 18:06:20 -04:00
bvandeusen c14338cbce Merge pull request 'v26.05.26.3 — hotfix: migration 0022 pre-merge across ENTIRE (canonical+others) group' (#25) from dev into main 2026-05-26 17:52:59 -04:00
bvandeusen 8c36dd28b0 Merge pull request 'v26.05.26.2 — hotfix: alembic 0022 Post-collision pre-merge + ci.yml cache continue-on-error' (#24) from dev into main 2026-05-26 16:50:43 -04:00
bvandeusen 88cfb3dd02 Merge pull request 'v26.05.26.1 — thumb backfill, modal redesign, recovery sweep race-safety, artist view redesign, extension fixes' (#23) from dev into main 2026-05-26 16:32:00 -04:00
bvandeusen 5d4f223b71 Merge pull request 'Release v26.05.25.7 — FC-Cleanup tab + UniqueViolation fix + error modal + extension install fix' (#22) from dev into main 2026-05-26 08:26:46 -04:00
bvandeusen 05090c6e85 Merge pull request 'Release v26.05.25.7 — animated-WebP worker fix + FC-Cleanup backend' (#21) from dev into main 2026-05-26 01:48:13 -04:00
bvandeusen 3a577d5ade Merge pull request 'fix(ext-ci): use browser_download_url + curl -f + ZIP magic check (XPI silently corrupt)' (#20) from dev into main 2026-05-26 00:43:02 -04:00
bvandeusen f4fe02e346 Merge pull request 'fix(ext-ci): drop actions/upload-artifact (Forgejo doesn't support v4+ GHES)' (#19) from dev into main 2026-05-25 23:33:40 -04:00
bvandeusen e766197d99 Merge pull request 'fix(ext-ci): jq→python + bump ext to 1.0.3 + rollback-on-upload-failure' (#18) from dev into main 2026-05-25 23:14:51 -04:00
bvandeusen 3872e1dda9 Merge pull request 'fix(ext-ci): web-ext v8 .cjs config workaround' (#17) from dev into main 2026-05-25 22:49:14 -04:00
bvandeusen 9814f3dbaf Merge pull request 'Release v26.05.25.5 — Extension publish refactor, deep-scan IR-parity, archive-import perf, artist Settings tab' (#16) from dev into main 2026-05-25 22:44:59 -04:00
bvandeusen b214460fdb Merge pull request 'Release v26.05.25.4 — importer ext sanitize fix, CI shard split, BrowserExtensionCard on Overview' (#15) from dev into main 2026-05-25 21:11:50 -04:00
bvandeusen ac55d0e8d8 Merge pull request 'fix(ext-ci): match AMO-renamed signed XPI' (#14) from dev into main 2026-05-25 18:22:50 -04:00
bvandeusen 89a89e0ded Merge pull request 'Release v26.05.25.3 — ML embedder SigLIP fix, import-UX, extension publish' (#13) from dev into main 2026-05-25 17:56:50 -04:00
bvandeusen 4e9aac2c05 Merge pull request 'v26.05.25.2: supersede + sidecar enrichment, scan toast feedback, CI uv + pip cache + durations' (#12) from dev into main 2026-05-25 14:30:25 -04:00
bvandeusen 2879ac6f2b Merge pull request 'v26.05.25.1: maintenance sweep + Camie v2 + corrupt-file handling + post-date gallery + clear-stuck escape hatch' (#11) from dev into main 2026-05-25 12:57:46 -04:00
bvandeusen b8dce6c483 Merge pull request 'FC-3h + FC-3k: backup first-class + admin destructive actions' (#10) from dev into main 2026-05-25 01:41:53 -04:00
bvandeusen d1c0b82a22 Merge pull request 'v26.05.24.3: FC-3i System Activity dashboard + migration backup-gate retired + modal Escape' (#9) from dev into main 2026-05-24 21:47:53 -04:00
bvandeusen 5526b8dc78 Merge pull request 'v26.05.24.2: IR Post/Provenance restore + modal artist fallback' (#8) from dev into main 2026-05-24 14:30:06 -04:00
bvandeusen 16eb7075c4 Merge pull request 'v26.05.24.1: FC-3g Firefox extension + worker resilience + UI/migration fixes' (#7) from dev into main 2026-05-24 12:52:31 -04:00
bvandeusen 885dcf64f3 Merge pull request 'v26.05.24.0: TopNav re-fix (flex 1 1 0 side cells)' (#6) from dev into main 2026-05-23 22:49:29 -04:00
bvandeusen f2f6b6d25e Merge pull request 'v26.05.23.3: dogfood UX polish + accurate active-batch stats' (#5) from dev into main 2026-05-23 22:05:59 -04:00
bvandeusen 0822240fde Merge pull request 'v26.05.23.2: serve /images + artist cleanup migrator' (#4) from dev into main 2026-05-23 12:19:16 -04:00
bvandeusen 27f7f3fd01 Merge pull request 'v26.05.23.1: migration durability + dogfood UX' (#3) from dev into main 2026-05-23 11:21:33 -04:00
bvandeusen c5bf564f53 Merge dev: v26.05.23.0 migration follow-ups (#2)
pg_dump + zstd in runtime image, lift Quart body cap to 1 GiB. See PR #2.
2026-05-22 22:37:06 -04:00
bvandeusen 602c7d275d Merge dev: FC-1 → FC-5 v1 build (#1)
First merge of `dev` into `main` for FabledCurator. Brings FC-1 (Foundation) through FC-5 (Migration tooling) onto `main`. See PR #1 body for the full stage rollup.
2026-05-22 14:15:45 -04:00
180 changed files with 2014 additions and 14602 deletions
-38
View File
@@ -329,41 +329,3 @@ jobs:
file: Dockerfile.ml
push: true
tags: ${{ steps.tag.outputs.tags }}
# The desktop GPU agent (#114) — published so the operator pulls + runs it on
# the GPU machine instead of building locally. Independent of web/ml (its own
# CUDA + onnxruntime-gpu image, context = agent/). Same tag cadence.
build-agent:
runs-on: python-ci
container:
image: git.fabledsword.com/bvandeusen/ci-python:3.14
steps:
- uses: actions/checkout@v4
- name: Determine tag
id: tag
run: |
SHORT_SHA=$(printf '%s' "$GITHUB_SHA" | cut -c1-7)
if [ "${GITHUB_REF#refs/tags/}" != "${GITHUB_REF}" ]; then
TAG_NAME="${GITHUB_REF#refs/tags/}"
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator-agent:${TAG_NAME}" >> "$GITHUB_OUTPUT"
elif [ "${GITHUB_REF##*/}" = "main" ]; then
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator-agent:main,git.fabledsword.com/bvandeusen/fabledcurator-agent:latest,git.fabledsword.com/bvandeusen/fabledcurator-agent:c-${SHORT_SHA}" >> "$GITHUB_OUTPUT"
else
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator-agent:dev" >> "$GITHUB_OUTPUT"
fi
- name: Login to Forgejo registry
uses: docker/login-action@v3
with:
registry: git.fabledsword.com
username: ${{ github.actor }}
password: ${{ secrets.RELEASE_TOKEN }}
- name: Build and push agent image
uses: docker/build-push-action@v5
with:
context: agent
file: agent/Dockerfile
push: true
tags: ${{ steps.tag.outputs.tags }}
-27
View File
@@ -1,27 +0,0 @@
# FabledCurator GPU agent — runs on the desktop with the GPU.
# CUDA + cuDNN runtime so onnxruntime-gpu can use the card (it needs cuDNN 9 —
# the plain -runtime image lacks it: "libcudnn.so.9: cannot open shared object
# file"); ffmpeg for video frames.
FROM nvidia/cuda:12.4.1-cudnn-runtime-ubuntu22.04
ENV DEBIAN_FRONTEND=noninteractive PYTHONUNBUFFERED=1
RUN apt-get update \
&& apt-get install -y --no-install-recommends python3 python3-pip ffmpeg \
&& rm -rf /var/lib/apt/lists/*
WORKDIR /app
# torch from the CUDA-12.4 wheel index (matches the base image); its wheels
# bundle their own CUDA + cuDNN and coexist with onnxruntime-gpu. Installed
# first + separately so the GPU build of torch is deterministic and layer-cached.
RUN pip3 install --no-cache-dir torch==2.6.0 --index-url https://download.pytorch.org/whl/cu124
COPY requirements.txt .
RUN pip3 install --no-cache-dir -r requirements.txt
COPY fc_agent ./fc_agent
# imgutils ONNX models + the transformers SigLIP weights both cache here; mount
# a volume to persist them across restarts (the SigLIP download is ~3.5 GB once).
ENV HF_HOME=/models
EXPOSE 8770
# The control UI; the worker is started from it (or POST /start).
CMD ["uvicorn", "fc_agent.app:app", "--host", "0.0.0.0", "--port", "8770"]
-71
View File
@@ -1,71 +0,0 @@
# FabledCurator GPU agent
A desktop-GPU worker that embeds characters (CCIP) + figure crops for
FabledCurator. It talks to FC **only over HTTP** — it leases jobs, fetches image
pixels, runs the models on your GPU, and posts results back. Your FC database and
Redis stay private; the agent never touches them.
You run it when you want a burst and stop it to reclaim the card.
## 0. Host prerequisite — NVIDIA Container Toolkit
Docker needs the toolkit to hand the GPU to a container (else: *"could not select
device driver nvidia with capabilities [[gpu]]"*). On Arch/CachyOS:
```sh
sudo pacman -S nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
# verify:
docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi
```
## 1. Get a token
In FC: **Settings → Tagging → GPU agent → Generate token** (or Rotate). Copy it.
## 2. Pull (CI publishes it alongside the web/ml images)
```sh
docker pull git.fabledsword.com/bvandeusen/fabledcurator-agent:latest
```
> Local build for development instead: `docker build -t fc-gpu-agent agent/`
## 3. Run (on the machine with the GPU)
```sh
docker run --rm --gpus all -p 8770:8770 \
-e FC_URL=http://curator.traefik.internal \
-e FC_TOKEN=<paste-the-token> \
-v fc-agent-models:/models \
git.fabledsword.com/bvandeusen/fabledcurator-agent:latest
```
Then open <http://localhost:8770> — the control page. Click **Start** to begin
draining the queue; **Pause**/**Stop** to yield the GPU. The `-v fc-agent-models`
volume caches the downloaded ONNX models so restarts are fast.
Kick off a backfill from FC (**GPU agent card → Queue character embedding**), then
watch the queue counts on the control page (or FC's card) drain.
## Config (env)
| var | default | meaning |
|---|---|---|
| `FC_URL` | `http://localhost:8000` | FC base URL |
| `FC_TOKEN` | — | the bearer token (required) |
| `AGENT_ID` | `desktop-agent` | identifies this agent's leases |
| `BATCH_SIZE` | `4` | jobs leased per round (still processed one at a time) |
| `CCIP_MODEL` | imgutils default | CCIP model name |
| `DETECTOR_LEVEL` | `m` | person-detector size: `n` < `s` < `m` < `x` |
| `POLL_IDLE_SECONDS` | `10` | wait between empty leases |
## ⚠️ Verify on first run
This part can't be CI-tested (no GPU/models in CI), so confirm against your
installed `dghs-imgutils` (`pip show dghs-imgutils`) — see `fc_agent/models.py`:
- `imgutils.detect.detect_person(image, level=...)` returns
`[((x0,y0,x1,y1), label, score), ...]`.
- `imgutils.metrics.ccip_extract_feature(image, model=...)` returns a vector
(768-d for caformer). If you want the F1-0.94 variant, set
`CCIP_MODEL=ccip-caformer_b36-24` (verify the exact string in imgutils).
If FC's matcher under/over-fires, tune the cosine threshold in
`backend/app/services/ml/ccip.py` (`DEFAULT_SIM_THRESHOLD`) and use
`GET /api/ccip/overview` + `/api/ccip/images/<id>` to spot-check.
## CPU fallback
Swap `onnxruntime-gpu``onnxruntime` in `requirements.txt` and drop `--gpus all`
to grind it slowly on the server instead. Same agent, no card.
-53
View File
@@ -1,53 +0,0 @@
# FabledCurator GPU agent — desktop run via docker compose.
#
# Usage:
# 1. Generate a token: FC → Settings → Tagging → GPU agent → Generate token.
# 2. Create a .env next to this file:
# FC_URL=http://curator.traefik.internal
# FC_TOKEN=<paste-the-token>
# # optional: CCIP_MODEL=ccip-caformer_b36-24 (the F1-0.94 variant)
# 3. docker compose up -d (pulls the published image)
# 4. Open http://localhost:8770 → Start. Pause/Stop hands the GPU back.
# docker compose down to stop the container entirely.
#
# Surviving a curator redeploy (you're away, can't touch the agent):
# - A running agent rides out curator being unreachable on its own — it retries
# leasing with capped backoff and resumes when the server is back. In-flight
# work is handed back (not failed), so a redeploy never poisons good jobs.
# - AUTO_START=1 (below) also resumes the worker if the AGENT container itself
# restarts (host reboot / crash via `restart: unless-stopped`) — no click.
#
# Needs the NVIDIA Container Toolkit installed on the host for --gpus.
services:
fc-gpu-agent:
image: git.fabledsword.com/bvandeusen/fabledcurator-agent:latest
pull_policy: always
ports:
- "8770:8770"
environment:
FC_URL: ${FC_URL:-http://curator.traefik.internal}
FC_TOKEN: ${FC_TOKEN:?set FC_TOKEN in .env (FC → GPU agent → Generate token)}
CCIP_MODEL: ${CCIP_MODEL:-}
DETECTOR_LEVEL: ${DETECTOR_LEVEL:-m}
BATCH_SIZE: ${BATCH_SIZE:-4}
# Resume the worker automatically on container start (survive a reboot /
# crash-restart while you're away). Set to 0 to require a manual Start.
AUTO_START: ${AUTO_START:-1}
# Crop embedder (SigLIP concept bag): float16 keeps VRAM low on a shared
# desktop GPU; the model itself is announced by the server.
SIGLIP_DTYPE: ${SIGLIP_DTYPE:-float16}
volumes:
# Persist the downloaded ONNX models so restarts are fast.
- fc-agent-models:/models
restart: unless-stopped
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
volumes:
fc-agent-models:
View File
-134
View File
@@ -1,134 +0,0 @@
"""FastAPI control surface for the agent (served on localhost).
Start / stop the worker pool, tune the worker count live (trades desktop
responsiveness for throughput), and watch GPU load + progress + the server-side
queue. Config is env-seeded; the worker count is adjustable here on the fly.
"""
from fastapi import FastAPI, Request
from fastapi.responses import HTMLResponse, JSONResponse
from .config import Config
from .gpu import read_gpu
from .worker import Worker
cfg = Config.from_env()
worker = Worker(cfg)
app = FastAPI(title="FabledCurator GPU agent")
@app.on_event("startup")
def _maybe_autostart() -> None:
# With AUTO_START set, a container restart (host reboot, or `restart:
# unless-stopped` after a crash) resumes the worker on its own — the slots
# then ride out a still-down curator via lease backoff. Lets the agent
# survive a redeploy with nobody at the desktop to click Start.
if cfg.auto_start and cfg.token:
worker.start()
@app.get("/", response_class=HTMLResponse)
def index() -> str:
return _PAGE
@app.post("/start")
def start():
worker.start()
return JSONResponse(worker.status())
@app.post("/stop")
def stop():
worker.stop()
return JSONResponse(worker.status())
@app.post("/concurrency")
async def concurrency(request: Request):
body = await request.json()
worker.set_concurrency(int(body.get("value", 1)))
return JSONResponse(worker.status())
@app.get("/status")
def status():
s = worker.status()
s["fc_url"] = cfg.fc_url
s["configured"] = bool(cfg.token)
s["gpu"] = read_gpu()
try:
s["queue"] = worker.client.queue_status()
except Exception:
s["queue"] = None
return JSONResponse(s)
_PAGE = """<!doctype html><html><head><meta charset=utf-8>
<title>FabledCurator GPU agent</title>
<style>
body{font:14px system-ui;margin:2rem;max-width:680px;background:#14171a;color:#e8e8e8}
h1{font-size:18px} button{font:14px system-ui;padding:.5rem 1rem;border:0;border-radius:6px;
margin-right:.5rem;cursor:pointer;color:#fff} .start{background:#2e7d32}.stop{background:#b3261e}
.step{background:#33373b;padding:.4rem .7rem;font-weight:700}
.stat{display:inline-block;margin-right:1.5rem;vertical-align:top}
.n{font-size:22px;font-weight:700} code{background:#222;padding:2px 6px;border-radius:4px}
.q,.gpu{margin-top:1rem;color:#9aa} .bar{height:8px;border-radius:4px;background:#222;overflow:hidden;
max-width:320px;margin-top:4px} .bar>i{display:block;height:100%;background:#3f7d3f}
.row{margin:.8rem 0}
</style></head><body>
<h1>FabledCurator GPU agent</h1>
<p>FC: <code id=fc>—</code> · token <code id=cfg>—</code></p>
<div class=row>
<button class=start onclick=act('start')>Start</button>
<button class=stop onclick=act('stop')>Stop</button>
</div>
<div class=row>
workers
<button class=step onclick=setc(-1)></button>
<input id=conc type=number min=1 value=1
style="width:3.5rem;font:700 16px system-ui;text-align:center;background:#222;color:#e8e8e8;border:1px solid #444;border-radius:6px;padding:.3rem"
onchange="setv(this.value)">
<button class=step onclick=setc(1)>+</button>
<span class=cap style=color:#9aa>(more = overlap I/O, fill the GPU) max <b id=capn>8</b></span>
</div>
<div class=row>
<span class=stat><span class=n id=state>stopped</span><br>state</span>
<span class=stat><span class=n id=active>0</span><br>active now</span>
<span class=stat><span class=n id=done>0</span><br>processed</span>
<span class=stat><span class=n id=err>0</span><br>errors</span>
<span class=stat><span class=n id=wait>0</span><br>waited out</span>
</div>
<div id=banner style="display:none;margin:.6rem 0;padding:.5rem .8rem;border-radius:6px;background:#5a4a17;color:#ffe28a">
curator unreachable — holding work + retrying, will resume on its own (no restart needed)
</div>
<div class=gpu id=gpu>GPU — …</div>
<div class=bar><i id=gpubar style=width:0%></i></div>
<div class=q id=queue></div>
<script>
let CAP=8
async function act(p){await fetch('/'+p,{method:'POST'});refresh()}
function setc(d){ setv((parseInt(conc.value||'1'))+d) }
async function setv(v){
v=Math.max(1,Math.min(CAP,parseInt(v)||1)); conc.value=v
await fetch('/concurrency',{method:'POST',headers:{'Content-Type':'application/json'},
body:JSON.stringify({value:v})});refresh()
}
async function refresh(){
const s=await (await fetch('/status')).json()
CAP=s.max_concurrency||8; capn.textContent=CAP
state.textContent=s.state; active.textContent=s.active; done.textContent=s.processed
err.textContent=s.errors; fc.textContent=s.fc_url; wait.textContent=s.transient||0
// Running but the queue read failed → curator is unreachable; show we're
// riding it out rather than erroring.
banner.style.display=(s.state==='running' && !s.queue)?'block':'none'
if(document.activeElement!==conc) conc.value=s.concurrency
conc.max=CAP
cfg.textContent=s.configured?'set':'MISSING'
if(s.gpu){
gpu.textContent=`GPU — ${s.gpu.util_pct}% util · VRAM ${s.gpu.mem_used_mb}/${s.gpu.mem_total_mb} MB · ${s.gpu.temp_c}°C`
gpubar.style.width=Math.round(100*s.gpu.mem_used_mb/s.gpu.mem_total_mb)+'%'
} else { gpu.textContent='GPU — n/a (CPU fallback?)'; gpubar.style.width='0%' }
queue.textContent=s.queue?`queue — pending ${s.queue.pending} · in flight ${s.queue.leased} · done ${s.queue.done} · errored ${s.queue.error}`:'queue — unreachable'
}
refresh(); setInterval(refresh,3000)
</script></body></html>"""
-85
View File
@@ -1,85 +0,0 @@
"""HTTP client for the FabledCurator GPU-job API.
The agent's ONLY contact with FC — lease/submit/heartbeat/fail + fetch image
bytes, all over HTTP with the bearer token. No DB/Redis.
"""
import requests
from requests.adapters import HTTPAdapter
class FcClient:
def __init__(self, base_url: str, token: str, agent_id: str):
self.base = base_url.rstrip("/")
self.agent_id = agent_id
self.s = requests.Session()
self.s.headers["Authorization"] = f"Bearer {token}"
# Many worker threads share this Session; the default pool (10) would
# throttle them + spam "connection pool is full". Size it for the cap.
adapter = HTTPAdapter(pool_connections=64, pool_maxsize=64)
self.s.mount("http://", adapter)
self.s.mount("https://", adapter)
def lease(self, batch_size: int) -> list[dict]:
r = self.s.post(
f"{self.base}/api/gpu/jobs/lease",
json={"agent_id": self.agent_id, "batch_size": batch_size},
timeout=30,
)
r.raise_for_status()
return r.json().get("jobs", [])
def submit(self, job_id: int, regions: list[dict], replace_kinds: list[str]) -> dict:
r = self.s.post(
f"{self.base}/api/gpu/jobs/submit",
json={
"agent_id": self.agent_id, "job_id": job_id,
"regions": regions, "replace_kinds": replace_kinds,
},
timeout=120,
)
r.raise_for_status()
return r.json()
def heartbeat(self, job_ids: list[int]) -> None:
try:
self.s.post(
f"{self.base}/api/gpu/jobs/heartbeat",
json={"agent_id": self.agent_id, "job_ids": job_ids},
timeout=30,
)
except requests.RequestException:
pass
def fail(self, job_id: int, error: str) -> None:
try:
self.s.post(
f"{self.base}/api/gpu/jobs/fail",
json={"agent_id": self.agent_id, "job_id": job_id, "error": error},
timeout=30,
)
except requests.RequestException:
pass
def release(self, job_ids: list[int]) -> None:
# Graceful hand-back on stop so orphaned work is re-leased at once.
if not job_ids:
return
try:
self.s.post(
f"{self.base}/api/gpu/jobs/release",
json={"agent_id": self.agent_id, "job_ids": job_ids},
timeout=30,
)
except requests.RequestException:
pass
def fetch_image(self, image_url: str) -> bytes:
# image_url is a server-relative path ("/images/...").
r = self.s.get(f"{self.base}{image_url}", timeout=180)
r.raise_for_status()
return r.content
def queue_status(self) -> dict:
r = self.s.get(f"{self.base}/api/gpu/status", timeout=15)
r.raise_for_status()
return r.json()
-36
View File
@@ -1,36 +0,0 @@
"""Agent config, all from env (the control container is configured at run)."""
import os
from dataclasses import dataclass
@dataclass
class Config:
fc_url: str # base URL of the FabledCurator web service
token: str # the bearer token from Settings → Tagging → GPU agent
agent_id: str # identifies this agent's leases
batch_size: int # jobs a worker leases per round
concurrency: int # INITIAL parallel workers (tunable live from the UI)
ccip_model: str # imgutils CCIP model name ("" → imgutils default)
detector_level: str # imgutils person-detector level: n|s|m|x
poll_idle_seconds: float # wait between empty leases
embed_dtype: str # torch dtype for the crop embedder: float16|float32
embed_model_override: str # force a SigLIP-family model ("" → use the one
# the server announces in the lease)
auto_start: bool # start the worker pool on boot (so a container restart
# resumes processing without anyone clicking Start)
@classmethod
def from_env(cls) -> "Config":
return cls(
fc_url=os.environ.get("FC_URL", "http://localhost:8000").rstrip("/"),
token=os.environ.get("FC_TOKEN", ""),
agent_id=os.environ.get("AGENT_ID", "desktop-agent"),
batch_size=int(os.environ.get("BATCH_SIZE", "4")),
concurrency=int(os.environ.get("CONCURRENCY", "1")),
ccip_model=os.environ.get("CCIP_MODEL", ""),
detector_level=os.environ.get("DETECTOR_LEVEL", "m"),
poll_idle_seconds=float(os.environ.get("POLL_IDLE_SECONDS", "10")),
embed_dtype=os.environ.get("SIGLIP_DTYPE", "float16"),
embed_model_override=os.environ.get("EMBED_MODEL_NAME", ""),
auto_start=os.environ.get("AUTO_START", "").lower() in ("1", "true", "yes"),
)
-36
View File
@@ -1,36 +0,0 @@
"""Crop primitive — vendored from backend/app/services/ml/crops.py so the agent
is self-contained. Keep in sync if the floor logic changes."""
from PIL import Image
MIN_CROP_FRACTION = 0.10
MIN_CROP_PX = 64
def crop_region(
img: Image.Image,
bbox: tuple[float, float, float, float],
*,
pad: float = 0.0,
min_fraction: float = MIN_CROP_FRACTION,
min_px: int = MIN_CROP_PX,
) -> Image.Image | None:
"""Crop a NORMALIZED bbox (x, y, w, h in [0,1]); None if below the size
floor (max of a fraction-of-short-side and an absolute pixel floor)."""
iw, ih = img.size
x, y, w, h = bbox
px, py, pw, ph = x * iw, y * ih, w * iw, h * ih
if pad:
px -= pw * pad / 2.0
py -= ph * pad / 2.0
pw *= (1.0 + pad)
ph *= (1.0 + pad)
left = max(0, int(round(px)))
top = max(0, int(round(py)))
right = min(iw, int(round(px + pw)))
bottom = min(ih, int(round(py + ph)))
if right <= left or bottom <= top:
return None
floor = max(min_px, int(min_fraction * min(iw, ih)))
if min(right - left, bottom - top) < floor:
return None
return img.crop((left, top, right, bottom)).convert("RGB")
-69
View File
@@ -1,69 +0,0 @@
"""Crop EMBEDDER for the concept bag — model-agnostic (CLIP/SigLIP-family).
The server trains its per-concept heads in the embedding space of whatever model
its `embedder_model_version` names; a crop must be embedded with the SAME model
or its vector lands in a different coordinate system and every head misfires. So
the model identity (HF name + version) is ANNOUNCED BY THE SERVER in the lease —
nothing here is hardcoded to SigLIP. Whatever name the server sends is loaded via
transformers `get_image_features` (the CLIP/SigLIP-family image-tower call); a
non-CLIP backbone (e.g. a DINO encoder) would need its own pooling adapter.
torch on CUDA, fp16 by default to keep VRAM low on a shared desktop GPU — the
tiny fp16-vs-fp32 difference is negligible for the linear heads (cosine ~0.999).
A single inference lock serializes the forward pass: the pipeline is I/O-bound,
so the GPU isn't the bottleneck, and one model shared across worker threads is
safest behind a lock.
"""
import threading
import numpy as np
from PIL import Image
class CropEmbedder:
def __init__(self, model_name: str, dtype: str = "float16"):
self._name = model_name
self._dtype_name = dtype
self._model = None
self._processor = None
self._torch = None
self._device = None
self._dt = None
self._load_lock = threading.Lock()
self._infer_lock = threading.Lock()
@property
def model_name(self) -> str:
return self._name
def load(self) -> None:
if self._model is not None:
return
with self._load_lock:
if self._model is not None:
return
import torch
from transformers import AutoImageProcessor, AutoModel
self._torch = torch
self._device = "cuda" if torch.cuda.is_available() else "cpu"
dt = getattr(torch, self._dtype_name, torch.float16)
if self._device == "cpu":
dt = torch.float32 # fp16 matmul is unsupported/slow on CPU
self._dt = dt
self._processor = AutoImageProcessor.from_pretrained(self._name)
model = AutoModel.from_pretrained(self._name, torch_dtype=dt)
model.eval().to(self._device)
self._model = model
def embed(self, image: Image.Image) -> list[float]:
"""A crop → its embedding as a plain float list, ready to POST."""
self.load()
torch = self._torch
enc = self._processor(images=image, return_tensors="pt")
pixel_values = enc["pixel_values"].to(self._device, self._dt)
with self._infer_lock, torch.no_grad():
out = self._model.get_image_features(pixel_values=pixel_values)
pooled = out.pooler_output if hasattr(out, "pooler_output") else out
vec = pooled[0].float().cpu().numpy().astype(np.float32).reshape(-1)
return vec.tolist()
-30
View File
@@ -1,30 +0,0 @@
"""GPU load readout via nvidia-smi (present in the container thanks to the
NVIDIA Container Toolkit's `utility` capability). Returns None if unavailable —
the UI just shows n/a (e.g. CPU-fallback run)."""
import subprocess
def read_gpu() -> dict | None:
try:
out = subprocess.run(
[
"nvidia-smi",
"--query-gpu=utilization.gpu,memory.used,memory.total,temperature.gpu",
"--format=csv,noheader,nounits",
],
capture_output=True, text=True, timeout=5, check=True,
).stdout.strip().splitlines()
except (OSError, subprocess.SubprocessError):
return None
if not out:
return None
parts = [p.strip() for p in out[0].split(",")]
try:
return {
"util_pct": int(float(parts[0])),
"mem_used_mb": int(float(parts[1])),
"mem_total_mb": int(float(parts[2])),
"temp_c": int(float(parts[3])),
}
except (ValueError, IndexError):
return None
-63
View File
@@ -1,63 +0,0 @@
"""Image + video handling. Stills load directly; videos are sampled into frames
(ffmpeg) at the cadence FC sends — so a video becomes a bag of per-frame
instances, each with a timestamp."""
import io
import os
import subprocess
import tempfile
from PIL import Image
def is_video(mime: str) -> bool:
return bool(mime) and (mime.startswith("video/") or mime in {"image/gif"})
def to_rgb(img: Image.Image) -> Image.Image:
"""RGB, flattening any transparency onto white first. A naive convert('RGB')
on a palette-with-transparency image (common for character PNGs on a clear
background) lets PIL guess the transparent pixels — usually black artifacts
that bleed into the crop + the embedding (and the "should be converted to
RGBA" warning). Compositing over white gives a clean, consistent background."""
if img.mode in ("RGBA", "LA", "PA") or (
img.mode == "P" and "transparency" in img.info
):
img = img.convert("RGBA")
bg = Image.new("RGBA", img.size, (255, 255, 255, 255))
return Image.alpha_composite(bg, img).convert("RGB")
return img.convert("RGB")
def load_image(data: bytes) -> Image.Image:
return to_rgb(Image.open(io.BytesIO(data)))
def sample_frames(
data: bytes, interval_seconds: float, max_frames: int
) -> list[tuple[float, Image.Image]]:
"""Extract up to max_frames frames at one-every-interval_seconds via ffmpeg.
Returns [(timestamp_seconds, frame)]. Empty on failure (caller falls back)."""
interval = max(0.5, float(interval_seconds or 4.0))
cap = max(1, int(max_frames or 64))
with tempfile.TemporaryDirectory() as tmp:
src = os.path.join(tmp, "in")
with open(src, "wb") as fh:
fh.write(data)
pattern = os.path.join(tmp, "f_%05d.jpg")
try:
subprocess.run(
[
"ffmpeg", "-nostdin", "-loglevel", "error", "-i", src,
"-vf", f"fps=1/{interval}", "-frames:v", str(cap),
"-q:v", "3", pattern,
],
check=True, timeout=600,
)
except (subprocess.SubprocessError, FileNotFoundError):
return []
out: list[tuple[float, Image.Image]] = []
names = sorted(n for n in os.listdir(tmp) if n.startswith("f_"))
for i, name in enumerate(names[:cap]):
with Image.open(os.path.join(tmp, name)) as im:
out.append((round(i * interval, 2), to_rgb(im)))
return out
-39
View File
@@ -1,39 +0,0 @@
"""imgutils model wrappers — the figure DETECTOR + the CCIP EMBEDDER.
⚠️ VERIFY ON FIRST RUN: the exact imgutils function names/signatures + the CCIP
model string can drift between dghs-imgutils releases. These are the two seams to
check against your installed version (`pip show dghs-imgutils`):
- detect_person(image, level=...) -> [((x0,y0,x1,y1), label, score), ...]
- ccip_extract_feature(image, model=...) -> a vector (768-d for caformer)
imgutils auto-downloads the ONNX models from HuggingFace on first use; GPU is
used when onnxruntime-gpu is installed.
"""
import numpy as np
from PIL import Image
def detect_figures(image: Image.Image, level: str = "m") -> list[tuple[tuple, float | None]]:
"""Person/figure bounding boxes, NORMALIZED (x, y, w, h in [0,1]) + score.
Returns [] if detection finds nothing (caller falls back to whole-image)."""
from imgutils.detect import detect_person
iw, ih = image.size
out = []
for (x0, y0, x1, y1), _label, score in detect_person(image, level=level):
out.append((
(x0 / iw, y0 / ih, (x1 - x0) / iw, (y1 - y0) / ih),
float(score),
))
return out
def ccip_vector(image: Image.Image, model: str | None = None) -> list[float]:
"""The CCIP identity embedding of a (cropped) character image, as a plain
float list ready to POST."""
from imgutils.metrics import ccip_extract_feature
feat = (
ccip_extract_feature(image, model=model)
if model else ccip_extract_feature(image)
)
return np.asarray(feat, dtype=np.float32).reshape(-1).tolist()
-274
View File
@@ -1,274 +0,0 @@
"""The lease → fetch → detect+embed → submit loop, run by a pool of worker
slots whose count is tunable live from the UI.
Each slot is an independent loop (its own leases; the server's SKIP-LOCKED lease
keeps them from colliding). More slots = more GPU load + throughput; the model is
loaded once and shared, so slots add concurrent inference, not N× model VRAM.
That's the dial the operator turns to trade desktop responsiveness for speed.
Stop (or shrinking the pool) RELEASES a slot's still-leased jobs immediately so
orphaned work is re-picked at once rather than waiting out the lease.
"""
import threading
import requests
from . import media, models
from .client import FcClient
from .config import Config
from .crops import crop_region
# Cap on the lease-retry backoff: when curator is unreachable (e.g. you redeploy
# it while away), each slot retries leasing with exponential backoff up to this
# many seconds, then resumes within this window once the server is back — no
# restart needed.
MAX_BACKOFF_SECONDS = 60.0
def _is_transient(exc: "requests.RequestException") -> bool:
"""A server/transport problem (wait it out) vs a job-specific fault (fail it).
No response → connection refused/timeout → curator is down → transient. With
a response: 5xx, auth (401/403, e.g. a token blip on redeploy), 408/409/429
(timeout / our lease reclaimed / rate-limited) are all 'not this job's fault'.
A specific 4xx like 404 (image gone) / 400 IS the job's fault → fail it."""
resp = getattr(exc, "response", None)
if resp is None:
return True
return resp.status_code >= 500 or resp.status_code in (401, 403, 408, 409, 429)
# Generous cap: the pipeline is usually I/O-bound (downloading + decoding images
# over HTTP), so the GPU stays underused until many workers overlap that I/O.
# Push it up while watching the GPU util + VRAM in the UI.
MAX_CONCURRENCY = 32
# Fallbacks only — the server ANNOUNCES the embedding model (name + version) in
# the lease so the agent stays model-agnostic and in lock-step with the space
# the heads were trained in. These cover an older server that doesn't send them.
DEFAULT_EMBED_MODEL = "google/siglip-so400m-patch14-384"
DEFAULT_EMBED_VERSION = "siglip-so400m-patch14-384"
class _Slot:
"""One worker loop. `inflight` = jobs leased but not yet processed, so a
graceful stop can hand them back."""
__slots__ = ("stop", "inflight")
def __init__(self):
self.stop = threading.Event()
self.inflight: list[int] = []
class Worker:
def __init__(self, cfg: Config):
self.cfg = cfg
self.client = FcClient(cfg.fc_url, cfg.token, cfg.agent_id)
self._lock = threading.Lock()
self._running = False
self._target = max(1, min(MAX_CONCURRENCY, cfg.concurrency))
self._slots: list[_Slot] = []
self.processed = 0
self.errors = 0
self.transient = 0 # jobs handed back due to a server outage (NOT
# failed) — the "waiting out curator" counter
self._active = 0 # slots currently mid-image
# The crop embedder (SigLIP-family) is built lazily on the first job that
# needs it, from the model the server announces — one shared instance.
self._embedder = None
self._embedder_lock = threading.Lock()
# --- control -----------------------------------------------------------
def start(self):
with self._lock:
self._running = True
self._reconcile_locked()
def stop(self):
with self._lock:
self._running = False
slots, self._slots = self._slots, []
for s in slots:
s.stop.set() # each slot releases its inflight on exit
def set_concurrency(self, n: int):
with self._lock:
self._target = max(1, min(MAX_CONCURRENCY, int(n)))
if self._running:
self._reconcile_locked()
def _reconcile_locked(self):
while len(self._slots) < self._target:
slot = _Slot()
self._slots.append(slot)
threading.Thread(target=self._loop, args=(slot,), daemon=True).start()
while len(self._slots) > self._target:
self._slots.pop().stop.set()
def status(self) -> dict:
with self._lock:
return {
"state": "running" if self._running else "stopped",
"concurrency": self._target,
"max_concurrency": MAX_CONCURRENCY,
"workers": len(self._slots),
"active": self._active,
"processed": self.processed,
"errors": self.errors,
"transient": self.transient,
}
def _bump(self, *, processed=0, errors=0, active=0, transient=0):
with self._lock:
self.processed += processed
self.errors += errors
self.transient += transient
self._active += active
# --- per-slot loop -----------------------------------------------------
def _loop(self, slot: _Slot):
backoff = self.cfg.poll_idle_seconds
while not slot.stop.is_set() and self._running:
try:
jobs = self.client.lease(self.cfg.batch_size)
backoff = self.cfg.poll_idle_seconds # server answered → reset
except Exception:
# curator unreachable (redeploy, network drop): wait it out with
# exponential backoff, capped — resume on our own when it returns.
self._interruptible_sleep(slot, backoff)
backoff = min(backoff * 2, MAX_BACKOFF_SECONDS)
continue
if not jobs:
self._interruptible_sleep(slot, self.cfg.poll_idle_seconds)
continue
slot.inflight = [j["job_id"] for j in jobs]
for job in jobs:
if slot.stop.is_set() or not self._running:
break
ok = self._process(job)
slot.inflight = [i for i in slot.inflight if i != job["job_id"]]
if not ok:
# Server went away mid-batch: hand the rest back (best effort)
# and back off instead of hammering a recovering server or
# burning the jobs' attempt budgets on fail().
if slot.inflight:
self.client.release(slot.inflight)
slot.inflight = []
self._interruptible_sleep(slot, backoff)
backoff = min(backoff * 2, MAX_BACKOFF_SECONDS)
break
if slot.inflight:
self.client.heartbeat(slot.inflight)
# Graceful hand-back of anything leased but not processed.
if slot.inflight:
self.client.release(slot.inflight)
slot.inflight = []
def _interruptible_sleep(self, slot: _Slot, seconds: float):
"""Sleep, but wake immediately if the slot is told to stop — so a Stop or
a pool-shrink doesn't hang for a full backoff window."""
slot.stop.wait(timeout=seconds)
def _ensure_embedder(self, model_name: str):
if self._embedder is not None:
return self._embedder
with self._embedder_lock:
if self._embedder is None:
from .embedder import CropEmbedder
self._embedder = CropEmbedder(model_name, self.cfg.embed_dtype)
return self._embedder
def _process(self, job: dict) -> bool:
"""Process one job. Returns True when handled (completed, or hard-failed
because the job itself is bad) and False on a TRANSPORT error (curator
unreachable / 5xx / our lease was reclaimed mid-flight) — which is not
the job's fault, so the caller backs off and the job is left to be
re-leased rather than fail()ed into its attempt budget."""
self._bump(active=1)
try:
data = self.client.fetch_image(job["image_url"])
if media.is_video(job.get("mime", "")):
frames = media.sample_frames(
data, job.get("frame_interval_seconds", 4.0),
job.get("max_frames", 64),
) or [(None, media.load_image(data))]
else:
frames = [(None, media.load_image(data))]
# task picks what to produce per crop:
# 'siglip' (backfill existing images) → concept (SigLIP) regions
# ONLY, so it never churns their figure/CCIP regions or the
# character-reference cache.
# 'ccip' / 'both' (a new image's first pass) → figure (CCIP) AND
# concept (SigLIP) in one go, off the same crop.
task = job.get("task") or "ccip"
want_ccip = task in ("ccip", "both")
want_siglip = task in ("ccip", "siglip", "both")
replace_kinds = (
["concept"] if task == "siglip" else ["figure", "face", "concept"]
)
embed_version = job.get("embed_version") or DEFAULT_EMBED_VERSION
embedder = None
if want_siglip:
model_name = (
self.cfg.embed_model_override
or job.get("embed_model_name")
or DEFAULT_EMBED_MODEL
)
embedder = self._ensure_embedder(model_name)
regions = []
ccip_ev = self.cfg.ccip_model or "ccip-default"
dv = f"person-{self.cfg.detector_level}"
for t, frame in frames:
figs = models.detect_figures(frame, self.cfg.detector_level)
if not figs:
figs = [((0.0, 0.0, 1.0, 1.0), None)] # whole-frame fallback
for bbox, score in figs:
crop = crop_region(frame, bbox)
if crop is None:
continue
if want_ccip:
regions.append({
"kind": "figure",
"bbox": list(bbox),
"frame_time": t,
"score": score,
"ccip_embedding": models.ccip_vector(
crop, self.cfg.ccip_model or None
),
"embedding_version": ccip_ev,
"detector_version": dv,
})
if want_siglip:
regions.append({
"kind": "concept",
"bbox": list(bbox),
"frame_time": t,
"score": score,
"siglip_embedding": embedder.embed(crop),
"embedding_version": embed_version,
"detector_version": dv,
})
self.client.submit(job["job_id"], regions, replace_kinds)
self._bump(processed=1)
return True
except requests.RequestException as exc:
if _is_transient(exc):
# curator down/redeploying, a 5xx, or our lease was reclaimed
# while we worked. NOT the job's fault — hand it back (best
# effort; no-ops if the server is still down, then the server's
# orphan-recovery reclaims it) and signal the loop to wait.
self._bump(transient=1)
self.client.release([job["job_id"]])
return False
# A job-specific HTTP fault (404 image gone, 400) → fail it so it
# doesn't re-lease forever.
self._bump(errors=1)
self.client.fail(job["job_id"], str(exc)[:500])
return True
except Exception as exc: # noqa: BLE001 — a genuine job fault: report it
self._bump(errors=1)
self.client.fail(job["job_id"], str(exc)[:500])
return True
finally:
self._bump(active=-1)
-15
View File
@@ -1,15 +0,0 @@
# CCIP + figure detection (ONNX models, auto-downloaded from HuggingFace).
dghs-imgutils>=0.4
# GPU inference for the ONNX models. Swap to onnxruntime (CPU) for a slow
# server-side fallback run.
onnxruntime-gpu
# The crop EMBEDDER (concept bag). torch is installed separately in the
# Dockerfile from the CUDA-12.4 wheel index so the GPU build is deterministic;
# transformers loads whatever SigLIP-family model the server announces.
transformers>=4.45
# Control surface + HTTP.
fastapi
uvicorn[standard]
requests
pillow
numpy
@@ -1,82 +0,0 @@
"""subscribestar_seen_media + subscribestar_failed_media: per-source ledgers
Revision ID: 0054
Revises: 0053
Create Date: 2026-06-17
SubscribeStar native ingester (phase 1 of the gallery-dl → native-core
migration). Mirrors the Patreon ledger tables (0037/0038): a seen-ledger so
routine walks skip already-ingested media (recovery bypasses it) and a
dead-letter ledger so persistently-failing media stops re-burning backfill
chunks. `filehash` is a CDN content hash when present, else a synthesized
``<post_id>:<filename>`` key — hence String(128). UNIQUE (source_id, filehash)
is the upsert key on each.
"""
from typing import Sequence, Union
import sqlalchemy as sa
from alembic import op
revision: str = "0054"
down_revision: Union[str, None] = "0053"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
op.create_table(
"subscribestar_seen_media",
sa.Column("id", sa.Integer, primary_key=True),
sa.Column(
"source_id",
sa.Integer,
sa.ForeignKey("source.id", ondelete="CASCADE"),
nullable=False,
index=True,
),
sa.Column("filehash", sa.String(128), nullable=False),
sa.Column("post_id", sa.String(64), nullable=True),
sa.Column(
"seen_at",
sa.DateTime(timezone=True),
nullable=False,
server_default=sa.text("NOW()"),
),
sa.UniqueConstraint(
"source_id", "filehash", name="uq_subscribestar_seen_media_source_id"
),
)
op.create_table(
"subscribestar_failed_media",
sa.Column("id", sa.Integer, primary_key=True),
sa.Column(
"source_id",
sa.Integer,
sa.ForeignKey("source.id", ondelete="CASCADE"),
nullable=False,
index=True,
),
sa.Column("filehash", sa.String(128), nullable=False),
sa.Column("attempts", sa.Integer, nullable=False, server_default="1"),
sa.Column("last_error", sa.Text, nullable=True),
sa.Column(
"first_failed_at",
sa.DateTime(timezone=True),
nullable=False,
server_default=sa.text("NOW()"),
),
sa.Column(
"last_failed_at",
sa.DateTime(timezone=True),
nullable=False,
server_default=sa.text("NOW()"),
),
sa.UniqueConstraint(
"source_id", "filehash", name="uq_subscribestar_failed_media_source_id"
),
)
def downgrade() -> None:
op.drop_table("subscribestar_failed_media")
op.drop_table("subscribestar_seen_media")
@@ -1,55 +0,0 @@
"""image_provenance: from_attachment_id (which archive an image was extracted from)
Milestone #87. When an image is pulled out of a .zip/.rar, record WHICH archive
PostAttachment it came from, so the provenance UI can show the single archive a
file lives inside instead of every attachment on the post. Nullable FK with
ON DELETE SET NULL — a loose (non-archive) download leaves it NULL, and deleting
the archive attachment forgets the linkage without destroying the (image, post)
provenance edge. Existing rows are NULL until the reextract backfill stamps them.
Revision ID: 0055
Revises: 0054
Create Date: 2026-06-22
"""
from typing import Sequence, Union
import sqlalchemy as sa
from alembic import op
revision: str = "0055"
down_revision: Union[str, None] = "0054"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
op.add_column(
"image_provenance",
sa.Column("from_attachment_id", sa.Integer(), nullable=True),
)
op.create_index(
"ix_image_provenance_from_attachment_id",
"image_provenance",
["from_attachment_id"],
)
op.create_foreign_key(
"fk_image_provenance_from_attachment",
"image_provenance",
"post_attachment",
["from_attachment_id"],
["id"],
ondelete="SET NULL",
)
def downgrade() -> None:
op.drop_constraint(
"fk_image_provenance_from_attachment",
"image_provenance",
type_="foreignkey",
)
op.drop_index(
"ix_image_provenance_from_attachment_id",
table_name="image_provenance",
)
op.drop_column("image_provenance", "from_attachment_id")
-43
View File
@@ -1,43 +0,0 @@
"""tag_eval_run: persisted head-vs-centroid tagging eval runs (#1130)
Milestone #114 slice 1. A long ml-queue eval whose full report must SURVIVE
navigation, so the run + report live in a row the admin card rehydrates from
(mirrors library_audit_run). running -> ready / error.
Revision ID: 0056
Revises: 0055
Create Date: 2026-06-28
"""
from typing import Sequence, Union
import sqlalchemy as sa
from alembic import op
from sqlalchemy.dialects.postgresql import JSONB
revision: str = "0056"
down_revision: Union[str, None] = "0055"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
op.create_table(
"tag_eval_run",
sa.Column("id", sa.Integer(), primary_key=True),
sa.Column("params", JSONB(), nullable=False),
sa.Column("status", sa.String(length=16), nullable=False, server_default="running"),
sa.Column(
"started_at", sa.DateTime(timezone=True), nullable=False,
server_default=sa.func.now(),
),
sa.Column("finished_at", sa.DateTime(timezone=True), nullable=True),
sa.Column("report", JSONB(), nullable=True),
sa.Column("error", sa.Text(), nullable=True),
sa.Column("last_progress_at", sa.DateTime(timezone=True), nullable=True),
)
op.create_index("ix_tag_eval_run_status", "tag_eval_run", ["status"])
def downgrade() -> None:
op.drop_index("ix_tag_eval_run_status", table_name="tag_eval_run")
op.drop_table("tag_eval_run")
@@ -1,40 +0,0 @@
"""tag_positive_confirmation: operator-affirmed correct positives (#1130)
Mirror of tag_suggestion_rejection. "Keep" on a doubted positive records here so
the eval's doubts list stops resurfacing confirmed-correct images every run.
Revision ID: 0057
Revises: 0056
Create Date: 2026-06-28
"""
from typing import Sequence, Union
import sqlalchemy as sa
from alembic import op
revision: str = "0057"
down_revision: Union[str, None] = "0056"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
op.create_table(
"tag_positive_confirmation",
sa.Column(
"image_record_id", sa.Integer(),
sa.ForeignKey("image_record.id", ondelete="CASCADE"), primary_key=True,
),
sa.Column(
"tag_id", sa.Integer(),
sa.ForeignKey("tag.id", ondelete="CASCADE"), primary_key=True, index=True,
),
sa.Column(
"confirmed_at", sa.DateTime(timezone=True), nullable=False,
server_default=sa.func.now(),
),
)
def downgrade() -> None:
op.drop_table("tag_positive_confirmation")
-95
View File
@@ -1,95 +0,0 @@
"""tag_head + head_training_run: production heads that learn from tags (#114)
The eval (#1130) proved the frozen-embedding + trained-head spine; this lands its
production form. tag_head stores one logistic-regression head per concept (the
new suggestion source, replacing Camie + centroid); head_training_run tracks the
batch that (re)trains them. Adds two head-training tunables to ml_settings.
Revision ID: 0058
Revises: 0057
Create Date: 2026-06-28
"""
from typing import Sequence, Union
import sqlalchemy as sa
from alembic import op
from pgvector.sqlalchemy import Vector
from sqlalchemy.dialects.postgresql import JSONB
revision: str = "0058"
down_revision: Union[str, None] = "0057"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
_HEAD_DIM = 1152
def upgrade() -> None:
op.create_table(
"tag_head",
sa.Column(
"tag_id", sa.Integer(),
sa.ForeignKey("tag.id", ondelete="CASCADE"), primary_key=True,
),
sa.Column("embedding_version", sa.String(length=128), nullable=False),
sa.Column("weights", Vector(_HEAD_DIM), nullable=False),
sa.Column("bias", sa.Float(), nullable=False),
sa.Column("suggest_threshold", sa.Float(), nullable=False),
sa.Column("auto_apply_threshold", sa.Float(), nullable=True),
sa.Column("n_pos", sa.Integer(), nullable=False),
sa.Column("n_neg", sa.Integer(), nullable=False),
sa.Column("ap", sa.Float(), nullable=False),
sa.Column("precision_cv", sa.Float(), nullable=False),
sa.Column("recall", sa.Float(), nullable=False),
sa.Column(
"trained_at", sa.DateTime(timezone=True), nullable=False,
server_default=sa.func.now(),
),
sa.Column("metrics", JSONB(), nullable=True),
)
op.create_table(
"head_training_run",
sa.Column("id", sa.Integer(), primary_key=True),
sa.Column("params", JSONB(), nullable=False),
sa.Column(
"status", sa.String(length=16), nullable=False,
server_default="running",
),
sa.Column(
"started_at", sa.DateTime(timezone=True), nullable=False,
server_default=sa.func.now(),
),
sa.Column("finished_at", sa.DateTime(timezone=True), nullable=True),
sa.Column("n_trained", sa.Integer(), nullable=True),
sa.Column("n_skipped", sa.Integer(), nullable=True),
sa.Column("error", sa.Text(), nullable=True),
sa.Column("last_progress_at", sa.DateTime(timezone=True), nullable=True),
)
op.create_index(
"ix_head_training_run_status", "head_training_run", ["status"],
)
# Head-training tunables on the ml_settings singleton.
op.add_column(
"ml_settings",
sa.Column(
"head_min_positives", sa.Integer(), nullable=False,
server_default="8",
),
)
op.add_column(
"ml_settings",
sa.Column(
"head_auto_apply_precision", sa.Float(), nullable=False,
server_default="0.97",
),
)
def downgrade() -> None:
op.drop_column("ml_settings", "head_auto_apply_precision")
op.drop_column("ml_settings", "head_min_positives")
op.drop_index("ix_head_training_run_status", table_name="head_training_run")
op.drop_table("head_training_run")
op.drop_table("tag_head")
-70
View File
@@ -1,70 +0,0 @@
"""head_auto_apply_run + earned-auto-apply settings (#114)
A graduated head can apply its tag without a human, gated by a master switch +
a support floor. head_auto_apply_run tracks each sweep / dry-run preview.
Revision ID: 0059
Revises: 0058
Create Date: 2026-06-29
"""
from typing import Sequence, Union
import sqlalchemy as sa
from alembic import op
from sqlalchemy.dialects.postgresql import JSONB
revision: str = "0059"
down_revision: Union[str, None] = "0058"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
op.create_table(
"head_auto_apply_run",
sa.Column("id", sa.Integer(), primary_key=True),
sa.Column(
"dry_run", sa.Boolean(), nullable=False, server_default=sa.false()
),
sa.Column("params", JSONB(), nullable=False),
sa.Column(
"status", sa.String(length=16), nullable=False,
server_default="running",
),
sa.Column(
"started_at", sa.DateTime(timezone=True), nullable=False,
server_default=sa.func.now(),
),
sa.Column("finished_at", sa.DateTime(timezone=True), nullable=True),
sa.Column("n_applied", sa.Integer(), nullable=True),
sa.Column("report", JSONB(), nullable=True),
sa.Column("error", sa.Text(), nullable=True),
sa.Column("last_progress_at", sa.DateTime(timezone=True), nullable=True),
)
op.create_index(
"ix_head_auto_apply_run_status", "head_auto_apply_run", ["status"],
)
op.add_column(
"ml_settings",
sa.Column(
"head_auto_apply_enabled", sa.Boolean(), nullable=False,
server_default=sa.true(), # opt-out: on by default (operator-asked)
),
)
op.add_column(
"ml_settings",
sa.Column(
"head_auto_apply_min_positives", sa.Integer(), nullable=False,
server_default="30",
),
)
def downgrade() -> None:
op.drop_column("ml_settings", "head_auto_apply_min_positives")
op.drop_column("ml_settings", "head_auto_apply_enabled")
op.drop_index(
"ix_head_auto_apply_run_status", table_name="head_auto_apply_run"
)
op.drop_table("head_auto_apply_run")
-74
View File
@@ -1,74 +0,0 @@
"""head_metric + head_metrics_snapshot: auto-apply observability (#114)
Running misfire/under-fire counters per concept (captured at correction time,
since image_tag.source is lost on delete) + a daily per-concept time-series so
the operator can tune the precision target + support floor from real data.
Revision ID: 0060
Revises: 0059
Create Date: 2026-06-29
"""
from typing import Sequence, Union
import sqlalchemy as sa
from alembic import op
revision: str = "0060"
down_revision: Union[str, None] = "0059"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
op.create_table(
"head_metric",
sa.Column(
"tag_id", sa.Integer(),
sa.ForeignKey("tag.id", ondelete="CASCADE"), primary_key=True,
),
sa.Column("n_misfires", sa.Integer(), nullable=False, server_default="0"),
sa.Column("n_underfires", sa.Integer(), nullable=False, server_default="0"),
sa.Column(
"updated_at", sa.DateTime(timezone=True), nullable=False,
server_default=sa.func.now(),
),
)
op.create_table(
"head_metrics_snapshot",
sa.Column("id", sa.Integer(), primary_key=True),
sa.Column(
"tag_id", sa.Integer(),
sa.ForeignKey("tag.id", ondelete="CASCADE"),
),
sa.Column("name", sa.String(length=255), nullable=False),
sa.Column(
"snapshot_at", sa.DateTime(timezone=True), nullable=False,
server_default=sa.func.now(),
),
sa.Column("n_auto_applied", sa.Integer(), nullable=False, server_default="0"),
sa.Column("n_misfires", sa.Integer(), nullable=False, server_default="0"),
sa.Column("n_underfires", sa.Integer(), nullable=False, server_default="0"),
sa.Column("ap", sa.Float(), nullable=True),
sa.Column("precision_cv", sa.Float(), nullable=True),
sa.Column("recall", sa.Float(), nullable=True),
sa.Column("n_pos", sa.Integer(), nullable=True),
)
op.create_index(
"ix_head_metrics_snapshot_tag_id", "head_metrics_snapshot", ["tag_id"],
)
op.create_index(
"ix_head_metrics_snapshot_snapshot_at", "head_metrics_snapshot",
["snapshot_at"],
)
def downgrade() -> None:
op.drop_index(
"ix_head_metrics_snapshot_snapshot_at", table_name="head_metrics_snapshot"
)
op.drop_index(
"ix_head_metrics_snapshot_tag_id", table_name="head_metrics_snapshot"
)
op.drop_table("head_metrics_snapshot")
op.drop_table("head_metric")
-59
View File
@@ -1,59 +0,0 @@
"""image_region: detected/proposed regions + their crop embeddings (#114)
Storage backbone of the crop pipeline. A region = normalized bbox + the crop's
embedding (CCIP for face/figure → character id; SigLIP for concept regions →
head bag-of-embeddings). Also serves as grounded-tag bbox provenance.
Revision ID: 0061
Revises: 0060
Create Date: 2026-06-29
"""
from typing import Sequence, Union
import sqlalchemy as sa
from alembic import op
from pgvector.sqlalchemy import Vector
revision: str = "0061"
down_revision: Union[str, None] = "0060"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
_CCIP_DIM = 768
_SIGLIP_DIM = 1152
def upgrade() -> None:
op.create_table(
"image_region",
sa.Column("id", sa.Integer(), primary_key=True),
sa.Column(
"image_record_id", sa.Integer(),
sa.ForeignKey("image_record.id", ondelete="CASCADE"), nullable=False,
),
sa.Column("kind", sa.String(length=16), nullable=False),
# Video/animated: source frame timestamp (seconds); NULL for stills.
sa.Column("frame_time", sa.Float(), nullable=True),
sa.Column("rx", sa.Float(), nullable=False),
sa.Column("ry", sa.Float(), nullable=False),
sa.Column("rw", sa.Float(), nullable=False),
sa.Column("rh", sa.Float(), nullable=False),
sa.Column("score", sa.Float(), nullable=True),
sa.Column("detector_version", sa.String(length=64), nullable=True),
sa.Column("crop_version", sa.String(length=64), nullable=True),
sa.Column("embedding_version", sa.String(length=128), nullable=True),
sa.Column("ccip_embedding", Vector(_CCIP_DIM), nullable=True),
sa.Column("siglip_embedding", Vector(_SIGLIP_DIM), nullable=True),
sa.Column(
"created_at", sa.DateTime(timezone=True), nullable=False,
server_default=sa.func.now(),
),
)
op.create_index(
"ix_image_region_image_record_id", "image_region", ["image_record_id"],
)
def downgrade() -> None:
op.drop_index("ix_image_region_image_record_id", table_name="image_region")
op.drop_table("image_region")
-55
View File
@@ -1,55 +0,0 @@
"""gpu_job: the HTTP-leased GPU work queue for the desktop agent (#114)
The agent stays HTTP-only — the server enqueues per-(image, task) jobs here and
the agent leases/submits over the web API; Redis/Postgres stay private.
Revision ID: 0062
Revises: 0061
Create Date: 2026-06-29
"""
from typing import Sequence, Union
import sqlalchemy as sa
from alembic import op
revision: str = "0062"
down_revision: Union[str, None] = "0061"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
op.create_table(
"gpu_job",
sa.Column("id", sa.Integer(), primary_key=True),
sa.Column(
"image_record_id", sa.Integer(),
sa.ForeignKey("image_record.id", ondelete="CASCADE"), nullable=False,
),
sa.Column("task", sa.String(length=32), nullable=False),
sa.Column(
"status", sa.String(length=16), nullable=False,
server_default="pending",
),
sa.Column("lease_token", sa.String(length=64), nullable=True),
sa.Column("leased_at", sa.DateTime(timezone=True), nullable=True),
sa.Column("lease_expires_at", sa.DateTime(timezone=True), nullable=True),
sa.Column("attempts", sa.Integer(), nullable=False, server_default="0"),
sa.Column("error", sa.Text(), nullable=True),
sa.Column(
"created_at", sa.DateTime(timezone=True), nullable=False,
server_default=sa.func.now(),
),
sa.Column(
"updated_at", sa.DateTime(timezone=True), nullable=False,
server_default=sa.func.now(),
),
)
op.create_index("ix_gpu_job_image_record_id", "gpu_job", ["image_record_id"])
op.create_index("ix_gpu_job_status", "gpu_job", ["status"])
def downgrade() -> None:
op.drop_index("ix_gpu_job_status", table_name="gpu_job")
op.drop_index("ix_gpu_job_image_record_id", table_name="gpu_job")
op.drop_table("gpu_job")
@@ -1,33 +0,0 @@
"""ml_settings.ccip_match_threshold — tunable CCIP character-match cut (#114)
The v1 matcher used a flat 0.75 cosine; live data showed that over-fires (a
high-reference character matched a scatter of images). 0.85 keeps the confident
single-character matches and drops the noise. Tunable from the GPU agent card.
Revision ID: 0063
Revises: 0062
Create Date: 2026-06-29
"""
from typing import Sequence, Union
import sqlalchemy as sa
from alembic import op
revision: str = "0063"
down_revision: Union[str, None] = "0062"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
op.add_column(
"ml_settings",
sa.Column(
"ccip_match_threshold", sa.Float(), nullable=False,
server_default="0.85",
),
)
def downgrade() -> None:
op.drop_column("ml_settings", "ccip_match_threshold")
-42
View File
@@ -1,42 +0,0 @@
"""ml_settings: CCIP auto-apply switch + threshold (#114)
Confident CCIP character matches auto-tag (source='ccip_auto') on a daily sweep,
so identity tags keep flowing without pressing a button. ON by default (opt-out,
like head auto-apply); the high threshold (0.92, above the 0.85 suggest cut) +
single-character references keep it safe, and every auto-tag is reversible.
Revision ID: 0064
Revises: 0063
Create Date: 2026-06-30
"""
from typing import Sequence, Union
import sqlalchemy as sa
from alembic import op
revision: str = "0064"
down_revision: Union[str, None] = "0063"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
op.add_column(
"ml_settings",
sa.Column(
"ccip_auto_apply_enabled", sa.Boolean(), nullable=False,
server_default=sa.true(),
),
)
op.add_column(
"ml_settings",
sa.Column(
"ccip_auto_apply_threshold", sa.Float(), nullable=False,
server_default="0.92",
),
)
def downgrade() -> None:
op.drop_column("ml_settings", "ccip_auto_apply_threshold")
op.drop_column("ml_settings", "ccip_auto_apply_enabled")
-8
View File
@@ -20,14 +20,11 @@ def all_blueprints() -> list[Blueprint]:
from .artist import artist_bp
from .artists import artists_bp
from .attachments import attachments_bp
from .ccip import ccip_bp
from .cleanup import cleanup_bp
from .credentials import credentials_bp
from .downloads import downloads_bp
from .extension import extension_bp
from .gallery import gallery_bp
from .gpu import gpu_bp
from .heads import heads_bp
from .import_admin import import_admin_bp
from .ml_admin import ml_admin_bp
from .platforms import platforms_bp
@@ -39,7 +36,6 @@ def all_blueprints() -> list[Blueprint]:
from .suggestions import suggestions_bp
from .system_activity import system_activity_bp
from .system_backup import system_backup_bp
from .tag_eval import tag_eval_bp
from .tags import tags_bp
from .thumbnails import thumbnails_bp
return [
@@ -60,10 +56,6 @@ def all_blueprints() -> list[Blueprint]:
suggestions_bp,
allowlist_bp,
aliases_bp,
tag_eval_bp,
heads_bp,
gpu_bp,
ccip_bp,
ml_admin_bp,
thumbnails_bp,
sources_bp,
+39 -76
View File
@@ -39,31 +39,6 @@ def _bulk_image_confirm_token(image_ids: list[int]) -> str:
return digest[:8]
async def _run_dry_run_op(service_fn, **service_kwargs):
"""Shared body for the Tier-A dry-run/apply endpoints: read the `dry_run`
flag, run the cleanup_service predicate under `run_sync`, and return its
result dict. The SAME `service_fn` drives both preview and apply (the flag
just toggles), so a handler physically can't let its preview diverge from
its delete (rule 93). Default False preserves the existing contract — the UI
always passes `dry_run` explicitly (true to preview, false to apply). Extra
service kwargs (e.g. `source_id`) pass straight through."""
body = await request.get_json(silent=True) or {}
dry_run = bool(body.get("dry_run", False))
async with get_session() as session:
result = await session.run_sync(
lambda sync_sess: service_fn(sync_sess, dry_run=dry_run, **service_kwargs)
)
return jsonify(result)
def _queued(async_result):
"""Standard 202 for an operator-triggered maintenance task: hand the UI the
Celery task id so it can tail /maintenance/task-result (or the activity
dashboard) for the summary. (trigger_vacuum stays bespoke — the UI doesn't
poll it, so it returns no task id.)"""
return jsonify({"task_id": async_result.id, "status": "queued"}), 202
@admin_bp.route("/artists/<slug>/cascade-delete", methods=["POST"])
async def artist_cascade_delete(slug: str):
body = await request.get_json(silent=True) or {}
@@ -171,30 +146,6 @@ async def tag_merge(dest_id: int):
if not isinstance(source_id, int) or source_id == dest_id:
return _bad("invalid_source_id", detail="source_id must be int and differ from dest")
# dry_run: non-mutating preview (counts + sample) so the operator can
# confirm the target before the irreversible merge (#8, rule 93 parity).
if body.get("dry_run"):
async with get_session() as session:
try:
p = await TagService(session).merge_preview(
source_id=source_id, target_id=dest_id,
)
except TagValidationError as exc:
return _bad("tag_not_found", status=404, detail=str(exc))
return jsonify({
"preview": {
"source_id": p.source_id, "source_name": p.source_name,
"target_id": p.target_id, "target_name": p.target_name,
"compatible": p.compatible,
"images_moving": p.images_moving,
"images_already_on_target": p.images_already_on_target,
"source_total": p.source_total,
"series_pages": p.series_pages,
"will_alias": p.will_alias,
"sample_thumbnails": p.sample_thumbnails,
},
})
async with get_session() as session:
try:
result = await TagService(session).merge(
@@ -242,7 +193,16 @@ async def tags_prune_unused():
re-call with dry_run=false."""
from ..services.cleanup_service import prune_unused_tags
return await _run_dry_run_op(prune_unused_tags)
body = await request.get_json(silent=True) or {}
dry_run = bool(body.get("dry_run", False))
async with get_session() as session:
result = await session.run_sync(
lambda sync_sess: prune_unused_tags(
sync_sess, dry_run=dry_run,
)
)
return jsonify(result)
@admin_bp.route("/posts/prune-bare", methods=["POST"])
@@ -254,27 +214,16 @@ async def posts_prune_bare():
prune itself, so the preview can't diverge from the delete."""
from ..services.cleanup_service import prune_bare_posts
return await _run_dry_run_op(prune_bare_posts)
@admin_bp.route("/posts/reconcile-duplicates", methods=["POST"])
async def posts_reconcile_duplicates():
"""Tier-A: unify duplicate post rows for the same real post — the gallery-dl
(attachment-id) + native (post-id) duplicates — onto ONE post-id-keyed keeper,
moving image/provenance/attachment/link rows over. Images are untouched.
dry_run=true returns {groups, posts_to_merge, sample}; dry_run=false applies
and returns {groups, merged, sample}. Optional source_id scopes to one source.
Same find_duplicate_post_groups predicate drives preview + apply (rule 93)."""
from ..services.cleanup_service import reconcile_duplicate_posts
body = await request.get_json(silent=True) or {}
raw_source = body.get("source_id")
try:
source_id = int(raw_source) if raw_source is not None else None
except (TypeError, ValueError):
return _bad("invalid_source_id", detail="source_id must be an integer")
dry_run = bool(body.get("dry_run", False))
return await _run_dry_run_op(reconcile_duplicate_posts, source_id=source_id)
async with get_session() as session:
result = await session.run_sync(
lambda sync_sess: prune_bare_posts(
sync_sess, dry_run=dry_run,
)
)
return jsonify(result)
@admin_bp.route("/tags/purge-legacy", methods=["POST"])
@@ -287,7 +236,14 @@ async def tags_purge_legacy():
operator confirms with dry_run=false."""
from ..services.cleanup_service import purge_legacy_tags
return await _run_dry_run_op(purge_legacy_tags)
body = await request.get_json(silent=True) or {}
dry_run = bool(body.get("dry_run", False))
async with get_session() as session:
result = await session.run_sync(
lambda sync_sess: purge_legacy_tags(sync_sess, dry_run=dry_run)
)
return jsonify(result)
@admin_bp.route("/tags/reset-content", methods=["POST"])
@@ -301,7 +257,14 @@ async def tags_reset_content():
Irreversible except via DB backup restore."""
from ..services.cleanup_service import reset_content_tagging
return await _run_dry_run_op(reset_content_tagging)
body = await request.get_json(silent=True) or {}
dry_run = bool(body.get("dry_run", False))
async with get_session() as session:
result = await session.run_sync(
lambda sync_sess: reset_content_tagging(sync_sess, dry_run=dry_run)
)
return jsonify(result)
@admin_bp.route("/tags/normalize", methods=["POST"])
@@ -327,7 +290,7 @@ async def tags_normalize():
from ..tasks.admin import normalize_tags_task
async_result = normalize_tags_task.delay()
return _queued(async_result)
return jsonify({"task_id": async_result.id, "status": "queued"}), 202
@admin_bp.route("/maintenance/db-stats", methods=["GET"])
@@ -384,7 +347,7 @@ async def trigger_reextract_archives():
from ..tasks.admin import reextract_archive_attachments_task
async_result = reextract_archive_attachments_task.delay()
return _queued(async_result)
return jsonify({"task_id": async_result.id, "status": "queued"}), 202
@admin_bp.route("/maintenance/prune-missing-files", methods=["POST"])
@@ -397,7 +360,7 @@ async def trigger_prune_missing_files():
from ..tasks.admin import prune_missing_file_records_task
async_result = prune_missing_file_records_task.delay()
return _queued(async_result)
return jsonify({"task_id": async_result.id, "status": "queued"}), 202
@admin_bp.route("/maintenance/dedup-videos", methods=["POST"])
@@ -413,7 +376,7 @@ async def trigger_dedup_videos():
body = await request.get_json(silent=True) or {}
dry_run = bool(body.get("dry_run", True)) # default to the SAFE preview
async_result = dedup_videos_task.delay(dry_run=dry_run)
return _queued(async_result)
return jsonify({"task_id": async_result.id, "status": "queued"}), 202
@admin_bp.route("/maintenance/purge-gated-previews", methods=["POST"])
@@ -429,7 +392,7 @@ async def trigger_purge_gated_previews():
body = await request.get_json(silent=True) or {}
dry_run = bool(body.get("dry_run", True)) # default to the SAFE preview
async_result = purge_gated_previews_task.delay(dry_run=dry_run)
return _queued(async_result)
return jsonify({"task_id": async_result.id, "status": "queued"}), 202
@admin_bp.route("/maintenance/task-result/<task_id>", methods=["GET"])
-25
View File
@@ -20,37 +20,12 @@ async def list_allowlist():
"tag_name": r.tag_name,
"tag_kind": r.tag_kind,
"min_confidence": r.min_confidence,
"applied_count": r.applied_count,
"coverage_count": r.coverage_count,
}
for r in rows
]
)
@allowlist_bp.route("/tags/<int:tag_id>/allowlist/coverage", methods=["GET"])
async def coverage(tag_id: int):
"""Live "at threshold T, a sweep would cover ~N images" projection for the
allowlist tuning dashboard. Defaults to the tag's stored threshold."""
raw = request.args.get("threshold")
async with get_session() as session:
svc = AllowlistService(session)
if raw is not None:
try:
threshold = float(raw)
except ValueError:
return jsonify({"error": "threshold must be a float"}), 400
if not (0 < threshold <= 1):
return jsonify({"error": "threshold must be in (0, 1]"}), 400
else:
row = await session.get(TagAllowlist, tag_id)
if row is None:
return jsonify({"error": "not on allowlist"}), 404
threshold = row.min_confidence
count = await svc.coverage(tag_id, threshold)
return jsonify({"count": count, "threshold": threshold})
@allowlist_bp.route("/tags/<int:tag_id>/allowlist", methods=["GET"])
async def get_one(tag_id: int):
async with get_session() as session:
-124
View File
@@ -1,124 +0,0 @@
"""CCIP / region observability API (#114) — read-only, analysis-shaped.
So the work can be checked through an API as the agent fills in vectors: overall
coverage (regions by kind, how many images have figure CCIP vectors, which
characters have enough reference examples to match on) + a per-image drill-down
(its regions + the CCIP character matches it would get). Mirrors the heads
metrics endpoint; no GPU, just reads what's stored.
"""
from quart import Blueprint, jsonify
from sqlalchemy import distinct, func, select
from ..extensions import get_session
from ..models import ImageRegion, Tag, TagKind
from ..models.tag import image_tag
from ..services.ml.ccip import match_image
ccip_bp = Blueprint("ccip", __name__, url_prefix="/api/ccip")
_FIGURE_KINDS = ("face", "figure")
@ccip_bp.route("/overview", methods=["GET"])
async def overview():
async with get_session() as session:
by_kind = dict(
(
await session.execute(
select(ImageRegion.kind, func.count()).group_by(ImageRegion.kind)
)
).all()
)
images_with_figure_ccip = (
await session.execute(
select(func.count(distinct(ImageRegion.image_record_id)))
.where(ImageRegion.kind.in_(_FIGURE_KINDS))
.where(ImageRegion.ccip_embedding.is_not(None))
)
).scalar_one()
# Concept-crop (SigLIP bag) coverage — how far the back-catalogue embed
# has progressed, so the max-over-bag scorer's reach is checkable.
images_with_concept_siglip = (
await session.execute(
select(func.count(distinct(ImageRegion.image_record_id)))
.where(ImageRegion.kind == "concept")
.where(ImageRegion.siglip_embedding.is_not(None))
)
).scalar_one()
# Per-character reference counts (no vectors loaded) — which characters
# have enough examples to match on.
ref_rows = (
await session.execute(
select(image_tag.c.tag_id, Tag.name, func.count())
.select_from(ImageRegion)
.join(
image_tag,
image_tag.c.image_record_id == ImageRegion.image_record_id,
)
.join(Tag, Tag.id == image_tag.c.tag_id)
.where(Tag.kind == TagKind.character)
.where(ImageRegion.kind.in_(_FIGURE_KINDS))
.where(ImageRegion.ccip_embedding.is_not(None))
.group_by(image_tag.c.tag_id, Tag.name)
.order_by(func.count().desc())
)
).all()
versions = [
v for (v,) in (
await session.execute(
select(distinct(ImageRegion.embedding_version))
)
).all() if v
]
auto_applied = (
await session.execute(
select(func.count()).select_from(image_tag).where(
image_tag.c.source == "ccip_auto"
)
)
).scalar_one()
return jsonify({
"regions_by_kind": by_kind,
"images_with_figure_ccip": images_with_figure_ccip,
"images_with_concept_siglip": images_with_concept_siglip,
"characters_with_references": len(ref_rows),
"character_references": [
{"tag_id": t, "name": n, "n_refs": c} for (t, n, c) in ref_rows
],
"embedding_versions": versions,
"auto_applied": auto_applied,
})
@ccip_bp.route("/images/<int:image_id>", methods=["GET"])
async def image_detail(image_id: int):
"""An image's stored regions + the CCIP character matches it would get —
for spot-checking the agent's output + the matcher."""
async with get_session() as session:
regions = (
await session.execute(
select(ImageRegion)
.where(ImageRegion.image_record_id == image_id)
.order_by(ImageRegion.id)
)
).scalars().all()
matches = await match_image(session, image_id)
return jsonify({
"image_id": image_id,
"regions": [
{
"id": r.id,
"kind": r.kind,
"bbox": [r.rx, r.ry, r.rw, r.rh],
"frame_time": r.frame_time,
"score": r.score,
"detector_version": r.detector_version,
"embedding_version": r.embedding_version,
"has_ccip": r.ccip_embedding is not None,
"has_siglip": r.siglip_embedding is not None,
}
for r in regions
],
"ccip_matches": matches,
})
+7 -23
View File
@@ -37,30 +37,16 @@ def _parse_filters():
"""Parse the composable gallery filters from query args, returning
``(filters_dict, sort)``. Raises ValueError (→ 400) on malformed ids/dates.
The structured tag filter (#6) is AND-of-OR plus exclusions:
- `tag_id` accepts a single id or a comma-separated list — all ANDed
(the include common case; back-compat).
- `tag_or` is REPEATABLE; each instance is a comma-separated OR-group, and
the image must match at least one tag from EACH group (groups ANDed).
- `tag_not` is a comma-separated exclude list (image must carry none).
`media` is image|video; `sort` is newest|oldest; `platform` selects one
platform (or the UNSOURCED_PLATFORM sentinel); `untagged`/`no_artist` are
boolean flags; `date_from`/`date_to` are inclusive calendar-day bounds
(date_to is widened by a day so the whole day is covered by the service's
half-open `< date_to`)."""
`tag_id` accepts a single id or a comma-separated list (AND); `media` is
image|video; `sort` is newest|oldest; `platform` selects one platform
(or the UNSOURCED_PLATFORM sentinel); `untagged`/`no_artist` are boolean
flags; `date_from`/`date_to` are inclusive calendar-day bounds (date_to is
widened by a day so the whole day is covered by the service's half-open
`< date_to`)."""
tag_raw = request.args.get("tag_id")
tag_ids = (
[int(x) for x in tag_raw.split(",") if x.strip()] if tag_raw else None
) or None
tag_or_groups = [
grp for raw in request.args.getlist("tag_or")
if (grp := [int(x) for x in raw.split(",") if x.strip()])
] or None
not_raw = request.args.get("tag_not")
tag_exclude = (
[int(x) for x in not_raw.split(",") if x.strip()] if not_raw else None
) or None
post_id_raw = request.args.get("post_id")
post_id = int(post_id_raw) if post_id_raw else None
artist_id_raw = request.args.get("artist_id")
@@ -78,9 +64,7 @@ def _parse_filters():
date_to += timedelta(days=1) # inclusive of the date_to calendar day
filters = {
"tag_ids": tag_ids, "post_id": post_id, "artist_id": artist_id,
"media_type": media_type,
"tag_or_groups": tag_or_groups, "tag_exclude": tag_exclude,
"platform": platform,
"media_type": media_type, "platform": platform,
"untagged": untagged, "no_artist": no_artist,
"date_from": date_from, "date_to": date_to,
}
-220
View File
@@ -1,220 +0,0 @@
"""GPU-job API (#114): the HTTP surface the desktop agent pulls work from.
The agent stays HTTP-only — it leases jobs, fetches image pixels via the normal
FC image URLs, and submits embeddings/regions back, all over this API. Redis and
Postgres are never exposed. The agent endpoints are gated by a bearer token
(Authorization: Bearer <token>) stored in AppSetting; the admin endpoints
(token / backfill / status) ride the browser session like the rest of FC's
homelab admin.
"""
import secrets
from quart import Blueprint, jsonify, request
from sqlalchemy import func, select
from sqlalchemy.dialects.postgresql import insert as pg_insert
from ..extensions import get_session
from ..models import AppSetting, GpuJob, ImageRecord, MLSettings
from ..services.gallery_service import image_url
from ..services.ml.embedder import MODEL_NAME as EMBED_MODEL_NAME
from ..services.ml.gpu_jobs import GpuJobService
from ..services.ml.regions import RegionService
gpu_bp = Blueprint("gpu", __name__, url_prefix="/api/gpu")
_TOKEN_KEY = "gpu_agent_token"
def _bearer() -> str | None:
h = request.headers.get("Authorization", "")
return h[7:].strip() if h.startswith("Bearer ") else None
async def _agent_authed(session) -> bool:
supplied = _bearer()
if not supplied:
return False
stored = (
await session.execute(
select(AppSetting.value).where(AppSetting.key == _TOKEN_KEY)
)
).scalar_one_or_none()
return stored is not None and secrets.compare_digest(supplied, stored)
# --- Admin (browser): token + backfill + status -------------------------
@gpu_bp.route("/token", methods=["GET"])
async def get_token():
async with get_session() as session:
tok = (
await session.execute(
select(AppSetting.value).where(AppSetting.key == _TOKEN_KEY)
)
).scalar_one_or_none()
return jsonify({"token": tok, "configured": tok is not None})
@gpu_bp.route("/token/rotate", methods=["POST"])
async def rotate_token():
token = secrets.token_urlsafe(32)
async with get_session() as session:
await session.execute(
pg_insert(AppSetting)
.values(key=_TOKEN_KEY, value=token)
.on_conflict_do_update(index_elements=["key"], set_={"value": token})
)
await session.commit()
return jsonify({"token": token})
@gpu_bp.route("/status", methods=["GET"])
async def status():
async with get_session() as session:
rows = (
await session.execute(
select(GpuJob.status, func.count()).group_by(GpuJob.status)
)
).all()
counts = dict(rows)
return jsonify({
"pending": counts.get("pending", 0),
"leased": counts.get("leased", 0),
"done": counts.get("done", 0),
"error": counts.get("error", 0),
})
@gpu_bp.route("/backfill", methods=["POST"])
async def backfill():
"""Enqueue a job for every image that doesn't already have one for `task`."""
body = await request.get_json(silent=True) or {}
task = str(body.get("task") or "ccip")
from ..tasks.ml import enqueue_gpu_backfill
r = enqueue_gpu_backfill.delay(task)
return jsonify({"celery_task_id": r.id, "task": task}), 202
# --- Agent (bearer token): lease / submit / heartbeat / fail ------------
@gpu_bp.route("/jobs/lease", methods=["POST"])
async def lease():
body = await request.get_json(silent=True) or {}
agent_id = str(body.get("agent_id") or "agent")
try:
batch = min(max(int(body.get("batch_size", 8)), 1), 64)
except (TypeError, ValueError):
batch = 8
async with get_session() as session:
if not await _agent_authed(session):
return jsonify({"error": "unauthorized"}), 401
jobs = await GpuJobService(session).lease(agent_id, batch_size=batch)
ml = (
await session.execute(select(MLSettings).where(MLSettings.id == 1))
).scalar_one()
# image rows for url/mime in one shot
ids = [j.image_record_id for j in jobs]
imgs = {
i.id: i for i in (
await session.execute(
select(ImageRecord).where(ImageRecord.id.in_(ids))
)
).scalars()
} if ids else {}
await session.commit()
out = []
for j in jobs:
img = imgs.get(j.image_record_id)
if img is None:
continue
out.append({
"job_id": j.id,
"image_id": j.image_record_id,
"task": j.task,
"mime": img.mime,
"image_url": image_url(img.path),
# For video/animated: the agent samples at this cadence.
"frame_interval_seconds": ml.video_frame_interval_seconds,
"max_frames": ml.video_max_frames,
# The embedding model the agent must use for concept crops, so
# its region vectors land in the SAME space the heads trained in.
# Server-announced → the agent stays model-agnostic; a swap is a
# server setting + a re-embed migration, never an agent change.
"embed_model_name": EMBED_MODEL_NAME,
"embed_version": ml.embedder_model_version,
})
return jsonify({"jobs": out})
@gpu_bp.route("/jobs/heartbeat", methods=["POST"])
async def heartbeat():
body = await request.get_json(silent=True) or {}
agent_id = str(body.get("agent_id") or "agent")
job_ids = [int(x) for x in (body.get("job_ids") or [])]
async with get_session() as session:
if not await _agent_authed(session):
return jsonify({"error": "unauthorized"}), 401
n = await GpuJobService(session).heartbeat(agent_id, job_ids)
await session.commit()
return jsonify({"extended": n})
@gpu_bp.route("/jobs/submit", methods=["POST"])
async def submit():
"""Store a job's regions + close it. regions: [{kind, bbox:[x,y,w,h],
frame_time?, score?, *_version?, ccip_embedding?, siglip_embedding?}].
replace_kinds defaults to the kinds present in the submitted regions."""
body = await request.get_json(silent=True) or {}
agent_id = str(body.get("agent_id") or "agent")
job_id = body.get("job_id")
regions = body.get("regions") or []
if job_id is None:
return jsonify({"error": "job_id required"}), 400
kinds = body.get("replace_kinds") or sorted({r["kind"] for r in regions})
async with get_session() as session:
if not await _agent_authed(session):
return jsonify({"error": "unauthorized"}), 401
job = await session.get(GpuJob, int(job_id))
if job is None or job.status != "leased" or job.lease_token != agent_id:
return jsonify({"error": "lease_invalid"}), 409
if kinds:
await RegionService(session).replace_regions(
job.image_record_id, kinds, regions
)
await GpuJobService(session).complete(agent_id, int(job_id))
await session.commit()
return jsonify({"ok": True, "stored": len(regions)})
@gpu_bp.route("/jobs/fail", methods=["POST"])
async def fail():
body = await request.get_json(silent=True) or {}
agent_id = str(body.get("agent_id") or "agent")
job_id = body.get("job_id")
if job_id is None:
return jsonify({"error": "job_id required"}), 400
async with get_session() as session:
if not await _agent_authed(session):
return jsonify({"error": "unauthorized"}), 401
ok = await GpuJobService(session).fail(
agent_id, int(job_id), str(body.get("error") or "")
)
await session.commit()
return jsonify({"ok": ok})
@gpu_bp.route("/jobs/release", methods=["POST"])
async def release():
"""Graceful stop: the agent hands its still-leased jobs back to pending so
they're picked up immediately instead of waiting out the lease."""
body = await request.get_json(silent=True) or {}
agent_id = str(body.get("agent_id") or "agent")
job_ids = [int(x) for x in (body.get("job_ids") or [])]
async with get_session() as session:
if not await _agent_authed(session):
return jsonify({"error": "unauthorized"}), 401
n = await GpuJobService(session).release(agent_id, job_ids)
await session.commit()
return jsonify({"released": n})
-285
View File
@@ -1,285 +0,0 @@
"""Heads API (#114): train + inspect the per-concept heads that power
suggestions (replacing Camie + centroid).
POST /api/heads/train — (re)train all eligible heads (one run at a time).
GET /api/heads — status: head count, last-trained, running run, the
per-concept head table (strength + auto-apply ready),
and recent training runs. The card rehydrates from
here so status survives navigation.
"""
from quart import Blueprint, jsonify, request
from sqlalchemy import desc, func, select
from ..extensions import get_session
from ..models import (
HeadAutoApplyRun,
HeadMetric,
HeadMetricsSnapshot,
HeadTrainingRun,
Tag,
TagHead,
)
from ..models.tag import image_tag
from ..services.ml.heads import (
HeadAutoApplyAlreadyRunning,
HeadAutoApplyDisabled,
HeadTrainingAlreadyRunning,
start_head_auto_apply_run,
start_head_training_run,
)
heads_bp = Blueprint("heads", __name__, url_prefix="/api/heads")
def _serialize_run(run: HeadTrainingRun) -> dict:
return {
"id": run.id,
"params": run.params,
"status": run.status,
"started_at": run.started_at.isoformat() if run.started_at else None,
"finished_at": run.finished_at.isoformat() if run.finished_at else None,
"n_trained": run.n_trained,
"n_skipped": run.n_skipped,
"error": run.error,
}
@heads_bp.route("/train", methods=["POST"])
async def train():
body = await request.get_json(silent=True) or {}
params = body.get("params") or body or {}
async with get_session() as session:
try:
run_id = await session.run_sync(
lambda s: start_head_training_run(s, params)
)
except HeadTrainingAlreadyRunning as running:
return jsonify({
"error": "training_already_running",
"running_id": int(running.args[0]),
}), 409
await session.commit()
return jsonify({"run_id": run_id, "status": "running"}), 202
@heads_bp.route("", methods=["GET"])
async def status():
async with get_session() as session:
count, last_trained = (
await session.execute(
select(func.count(), func.max(TagHead.trained_at))
)
).one()
graduated = (
await session.execute(
select(func.count()).where(
TagHead.auto_apply_threshold.is_not(None)
)
)
).scalar_one()
running = (
await session.execute(
select(HeadTrainingRun.id)
.where(HeadTrainingRun.status == "running")
.order_by(HeadTrainingRun.id.desc())
.limit(1)
)
).scalar_one_or_none()
runs = (
await session.execute(
select(HeadTrainingRun)
.order_by(HeadTrainingRun.id.desc())
.limit(10)
)
).scalars().all()
# The per-concept table: strongest first, capped for the admin card.
head_rows = (
await session.execute(
select(
TagHead.tag_id, Tag.name, Tag.kind,
TagHead.n_pos, TagHead.n_neg, TagHead.ap,
TagHead.precision_cv, TagHead.recall,
TagHead.auto_apply_threshold, TagHead.trained_at,
)
.join(Tag, Tag.id == TagHead.tag_id)
.order_by(desc(TagHead.ap))
.limit(500)
)
).all()
heads = [
{
"tag_id": r.tag_id,
"name": r.name,
"category": r.kind.value if hasattr(r.kind, "value") else str(r.kind),
"n_pos": r.n_pos,
"n_neg": r.n_neg,
"ap": r.ap,
"precision": r.precision_cv,
"recall": r.recall,
"auto_apply": r.auto_apply_threshold is not None,
"trained_at": r.trained_at.isoformat() if r.trained_at else None,
}
for r in head_rows
]
return jsonify({
"head_count": count,
"graduated_count": graduated,
"last_trained_at": last_trained.isoformat() if last_trained else None,
"running_id": running,
"runs": [_serialize_run(r) for r in runs],
"heads": heads,
})
def _serialize_apply_run(run: HeadAutoApplyRun) -> dict:
return {
"id": run.id,
"dry_run": run.dry_run,
"status": run.status,
"started_at": run.started_at.isoformat() if run.started_at else None,
"finished_at": run.finished_at.isoformat() if run.finished_at else None,
"n_applied": run.n_applied,
"report": run.report,
"error": run.error,
}
@heads_bp.route("/auto-apply", methods=["POST"])
async def auto_apply():
"""Trigger an earned-auto-apply sweep. {dry_run:true} previews (writes
nothing); a real sweep needs head_auto_apply_enabled on."""
body = await request.get_json(silent=True) or {}
params = {"dry_run": bool(body.get("dry_run", False))}
async with get_session() as session:
try:
run_id = await session.run_sync(
lambda s: start_head_auto_apply_run(s, params)
)
except HeadAutoApplyAlreadyRunning as running:
return jsonify({
"error": "auto_apply_already_running",
"running_id": int(running.args[0]),
}), 409
except HeadAutoApplyDisabled:
return jsonify({"error": "auto_apply_disabled"}), 400
await session.commit()
return jsonify({"run_id": run_id, "status": "running"}), 202
@heads_bp.route("/auto-apply", methods=["GET"])
async def auto_apply_status():
async with get_session() as session:
running = (
await session.execute(
select(HeadAutoApplyRun.id)
.where(HeadAutoApplyRun.status == "running")
.order_by(HeadAutoApplyRun.id.desc())
.limit(1)
)
).scalar_one_or_none()
runs = (
await session.execute(
select(HeadAutoApplyRun)
.order_by(HeadAutoApplyRun.id.desc())
.limit(10)
)
).scalars().all()
return jsonify({
"running_id": running,
"runs": [_serialize_apply_run(r) for r in runs],
})
@heads_bp.route("/metrics", methods=["GET"])
async def metrics():
"""Auto-apply observability: per-concept current counts (volume, misfires,
under-fires, realized misfire rate, head quality) + the daily time-series so
the operator can tune the precision target + support floor from real data."""
async with get_session() as session:
head_rows = (
await session.execute(
select(
TagHead.tag_id, Tag.name, TagHead.ap, TagHead.precision_cv,
TagHead.recall, TagHead.auto_apply_threshold, TagHead.n_pos,
).join(Tag, Tag.id == TagHead.tag_id)
)
).all()
heads = {r.tag_id: r for r in head_rows}
metric_rows = (
await session.execute(
select(
HeadMetric.tag_id, HeadMetric.n_misfires, HeadMetric.n_underfires
)
)
).all()
mets = {r.tag_id: r for r in metric_rows}
applied = dict(
(
await session.execute(
select(image_tag.c.tag_id, func.count())
.where(image_tag.c.source == "head_auto")
.group_by(image_tag.c.tag_id)
)
).all()
)
names = {r.tag_id: r.name for r in head_rows}
# Names for metric-only tags (head pruned but corrections recorded).
missing = [t for t in mets if t not in names]
if missing:
for tid, nm in (
await session.execute(
select(Tag.id, Tag.name).where(Tag.id.in_(missing))
)
).all():
names[tid] = nm
concepts = []
for tid in set(heads) | set(mets):
h = heads.get(tid)
m = mets.get(tid)
n_applied = applied.get(tid, 0)
n_mis = m.n_misfires if m else 0
denom = n_applied + n_mis
concepts.append({
"tag_id": tid,
"name": names.get(tid, str(tid)),
"n_auto_applied": n_applied,
"n_misfires": n_mis,
"n_underfires": m.n_underfires if m else 0,
# Of everything this head ever auto-applied, the fraction you
# removed — the misfire rate (null until something fired).
"misfire_rate": round(n_mis / denom, 4) if denom else None,
"ap": h.ap if h else None,
"precision_cv": h.precision_cv if h else None,
"recall": h.recall if h else None,
"auto_apply": bool(h and h.auto_apply_threshold is not None),
"n_pos": h.n_pos if h else None,
})
concepts.sort(key=lambda c: (c["n_misfires"], c["n_auto_applied"]), reverse=True)
snaps = (
await session.execute(
select(HeadMetricsSnapshot)
.order_by(HeadMetricsSnapshot.snapshot_at.desc())
.limit(1000)
)
).scalars().all()
return jsonify({
"concepts": concepts,
"snapshots": [
{
"tag_id": s.tag_id,
"name": s.name,
"snapshot_at": s.snapshot_at.isoformat() if s.snapshot_at else None,
"n_auto_applied": s.n_auto_applied,
"n_misfires": s.n_misfires,
"n_underfires": s.n_underfires,
"ap": s.ap,
"precision_cv": s.precision_cv,
"recall": s.recall,
"n_pos": s.n_pos,
}
for s in snaps
],
})
-25
View File
@@ -17,13 +17,6 @@ _EDITABLE = (
"video_frame_interval_seconds",
"video_max_frames",
"video_min_tag_frames",
"head_min_positives",
"head_auto_apply_precision",
"head_auto_apply_enabled",
"head_auto_apply_min_positives",
"ccip_match_threshold",
"ccip_auto_apply_enabled",
"ccip_auto_apply_threshold",
)
@@ -47,13 +40,6 @@ async def get_settings():
"video_min_tag_frames": s.video_min_tag_frames,
"tagger_model_version": s.tagger_model_version,
"embedder_model_version": s.embedder_model_version,
"head_min_positives": s.head_min_positives,
"head_auto_apply_precision": s.head_auto_apply_precision,
"head_auto_apply_enabled": s.head_auto_apply_enabled,
"head_auto_apply_min_positives": s.head_auto_apply_min_positives,
"ccip_match_threshold": s.ccip_match_threshold,
"ccip_auto_apply_enabled": s.ccip_auto_apply_enabled,
"ccip_auto_apply_threshold": s.ccip_auto_apply_threshold,
}
)
@@ -114,17 +100,6 @@ def _validate(p: dict) -> str | None:
return "video_min_tag_frames must be >= 1"
if p["video_min_tag_frames"] > p["video_max_frames"]:
return "video_min_tag_frames cannot exceed video_max_frames"
# Head training (#114).
if int(p["head_min_positives"]) < 1:
return "head_min_positives must be >= 1"
if not (0.5 <= float(p["head_auto_apply_precision"]) <= 0.999):
return "head_auto_apply_precision must be between 0.5 and 0.999"
if int(p["head_auto_apply_min_positives"]) < 1:
return "head_auto_apply_min_positives must be >= 1"
if not (0.5 <= float(p["ccip_match_threshold"]) <= 0.999):
return "ccip_match_threshold must be between 0.5 and 0.999"
if not (0.5 <= float(p["ccip_auto_apply_threshold"]) <= 0.999):
return "ccip_auto_apply_threshold must be between 0.5 and 0.999"
return None
+6 -51
View File
@@ -3,31 +3,12 @@
from quart import Blueprint, jsonify, request
from ..extensions import get_session
from ..models import Tag, TagAllowlist
from ..services.ml.allowlist import AllowlistService
from ..services.ml.suggestions import SuggestionService
suggestions_bp = Blueprint("suggestions", __name__, url_prefix="/api")
async def _accept_payload(session, svc, newly_added: bool, tag_id: int) -> dict:
"""Shape the accept/alias response. When accepting newly allowlists a tag,
include the coverage PROJECTION (at the tag's threshold) so the UI can show
a non-blocking "auto-applying to ~N images" toast — the actual apply runs
async via apply_allowlist_tags, so this is an estimate, not a post-hoc
count (#7)."""
payload = {"allowlisted": newly_added}
if newly_added:
tag = await session.get(Tag, tag_id)
row = await session.get(TagAllowlist, tag_id)
payload["tag_id"] = tag_id
payload["tag_name"] = tag.name if tag is not None else None
payload["projected_count"] = await svc.coverage(
tag_id, row.min_confidence if row is not None else 0.90,
)
return payload
@suggestions_bp.route("/images/<int:image_id>/suggestions", methods=["GET"])
async def get_suggestions(image_id: int):
# ?min=<float> overrides the configured per-category thresholds so the typed
@@ -61,10 +42,6 @@ async def get_suggestions(image_id: int):
# modal's "Treat as alias"/"Remove alias" affordances.
"raw_name": s.raw_name,
"via_alias": s.via_alias,
# operator dismissed this tag for this image — surfaced
# (not dropped) so the rail can show it rejected + offer
# one-click un-reject.
"rejected": s.rejected,
}
for s in items
]
@@ -83,15 +60,13 @@ async def accept_suggestion(image_id: int):
return jsonify({"error": "tag_id required"}), 400
tag_id = body["tag_id"]
async with get_session() as session:
svc = AllowlistService(session)
newly_added = await svc.accept(image_id, tag_id)
payload = await _accept_payload(session, svc, newly_added, tag_id)
newly_added = await AllowlistService(session).accept(image_id, tag_id)
await session.commit()
if newly_added:
from ..tasks.ml import apply_allowlist_tags
apply_allowlist_tags.delay(tag_id=tag_id)
return jsonify(payload)
return "", 204
@suggestions_bp.route(
@@ -102,24 +77,19 @@ async def alias_suggestion(image_id: int):
required = {"alias_string", "alias_category", "canonical_tag_id"}
if not body or not required.issubset(body):
return jsonify({"error": f"required: {sorted(required)}"}), 400
canonical_tag_id = body["canonical_tag_id"]
async with get_session() as session:
svc = AllowlistService(session)
newly_added = await svc.add_alias_and_accept(
newly_added = await AllowlistService(session).add_alias_and_accept(
image_id,
body["alias_string"],
body["alias_category"],
canonical_tag_id,
)
payload = await _accept_payload(
session, svc, newly_added, canonical_tag_id,
body["canonical_tag_id"],
)
await session.commit()
if newly_added:
from ..tasks.ml import apply_allowlist_tags
apply_allowlist_tags.delay(tag_id=canonical_tag_id)
return jsonify(payload)
apply_allowlist_tags.delay(tag_id=body["canonical_tag_id"])
return "", 204
@suggestions_bp.route(
@@ -135,21 +105,6 @@ async def dismiss_suggestion(image_id: int):
return "", 204
@suggestions_bp.route(
"/images/<int:image_id>/suggestions/undismiss", methods=["POST"]
)
async def undismiss_suggestion(image_id: int):
"""Reverse a per-image dismissal (reject-recovery). Idempotent — undoing a
tag that isn't rejected is a no-op delete."""
body = await request.get_json()
if not body or "tag_id" not in body:
return jsonify({"error": "tag_id required"}), 400
async with get_session() as session:
await AllowlistService(session).undismiss(image_id, body["tag_id"])
await session.commit()
return "", 204
@suggestions_bp.route("/suggestions/bulk", methods=["POST"])
async def bulk_suggestions():
body = await request.get_json()
-70
View File
@@ -1,70 +0,0 @@
"""Tag-eval API (#1130): trigger + revisit the head-vs-centroid eval.
The run + full report live in the tag_eval_run row, so the admin card rehydrates
from GET (history / detail) on mount — the report survives navigation rather than
living in transient frontend state.
"""
from quart import Blueprint, jsonify, request
from sqlalchemy import select
from ..extensions import get_session
from ..models import TagEvalRun
from ..services.ml.tag_eval import EvalAlreadyRunning, start_tag_eval_run
tag_eval_bp = Blueprint("tag_eval", __name__, url_prefix="/api/tag-eval")
def _serialize(run: TagEvalRun, *, include_report: bool) -> dict:
out = {
"id": run.id,
"params": run.params,
"status": run.status,
"started_at": run.started_at.isoformat() if run.started_at else None,
"finished_at": run.finished_at.isoformat() if run.finished_at else None,
"error": run.error,
}
if include_report:
out["report"] = run.report
return out
@tag_eval_bp.route("", methods=["POST"])
async def create():
body = await request.get_json(silent=True) or {}
params = body.get("params") or body or {}
async with get_session() as session:
try:
run_id = await session.run_sync(
lambda s: start_tag_eval_run(s, params)
)
except EvalAlreadyRunning as running:
return jsonify({
"error": "eval_already_running",
"running_id": int(running.args[0]),
}), 409
await session.commit()
return jsonify({"run_id": run_id, "status": "running"}), 202
@tag_eval_bp.route("", methods=["GET"])
async def history():
try:
limit = min(int(request.args.get("limit", "20")), 100)
except ValueError:
return jsonify({"error": "invalid_limit"}), 400
async with get_session() as session:
rows = (await session.execute(
select(TagEvalRun).order_by(TagEvalRun.id.desc()).limit(limit)
)).scalars().all()
# List is light — no full report (the detail endpoint carries it).
return jsonify({"runs": [_serialize(r, include_report=False) for r in rows]})
@tag_eval_bp.route("/<int:run_id>", methods=["GET"])
async def detail(run_id: int):
async with get_session() as session:
run = await session.get(TagEvalRun, run_id)
if run is None:
return jsonify({"error": "not_found"}), 404
return jsonify(_serialize(run, include_report=True))
+1 -16
View File
@@ -2,11 +2,10 @@
from quart import Blueprint, jsonify, request
from sqlalchemy import exists, select
from sqlalchemy.dialects.postgresql import insert as pg_insert
from sqlalchemy.exc import IntegrityError
from ..extensions import get_session
from ..models import Tag, TagKind, TagPositiveConfirmation
from ..models import Tag, TagKind
from ..models.tag_allowlist import TagAllowlist
from ..services.bulk_tag_service import BulkTagService
from ..services.ml.aliases import AliasService
@@ -184,20 +183,6 @@ async def remove_tag_from_image(image_id: int, tag_id: int):
return "", 204
@tags_bp.route("/images/<int:image_id>/tags/<int:tag_id>/confirm", methods=["POST"])
async def confirm_tag_on_image(image_id: int, tag_id: int):
"""Operator affirmed an applied tag is correct ("keep" on a doubted positive).
Idempotent; recorded so the eval's doubts list stops resurfacing it (#1130)."""
async with get_session() as session:
await session.execute(
pg_insert(TagPositiveConfirmation)
.values(image_record_id=image_id, tag_id=tag_id)
.on_conflict_do_nothing(index_elements=["image_record_id", "tag_id"])
)
await session.commit()
return "", 204
@tags_bp.route("/tags/<int:tag_id>", methods=["GET"])
async def get_tag(tag_id: int):
"""Resolve a single tag (used by the gallery to label its active
-68
View File
@@ -61,33 +61,7 @@ def make_celery() -> Celery:
# Heavy ML tasks need fair dispatch — see ImageRepo's precedent.
task_acks_late=True,
worker_prefetch_multiplier=1,
# Broker resilience (2026-06-24): a swarm overlay-network blip after a
# redeploy left Redis healthy but transiently unreachable, and a worker
# starting in that window crash-looped on the initial broker connect
# (kombu OperationalError) instead of waiting it out — needing a manual
# Redis reset to recover. Retry the broker FOREVER (None) on startup and
# at runtime so a transient outage self-heals when routing returns,
# rather than the worker exiting.
broker_connection_retry_on_startup=True,
broker_connection_retry=True,
broker_connection_max_retries=None,
# Redis-transport socket options (apply to the BROKER connection): a
# short connect timeout + TCP keepalive so a dead/blocked socket is
# noticed and retried, and a periodic health check that proactively
# reconnects a live worker through a network hiccup.
broker_transport_options={
"socket_connect_timeout": 5,
"socket_timeout": 30,
"socket_keepalive": True,
"retry_on_timeout": True,
"health_check_interval": 30,
},
# Same hardening for the Redis RESULT backend (separate connection pool).
redis_socket_connect_timeout=5,
redis_socket_timeout=30,
redis_socket_keepalive=True,
redis_retry_on_timeout=True,
redis_backend_health_check_interval=30,
beat_schedule={
"recover-interrupted-tasks": {
"task": "backend.app.tasks.maintenance.recover_interrupted_tasks",
@@ -109,36 +83,6 @@ def make_celery() -> Celery:
"task": "backend.app.tasks.ml.apply_allowlist_tags",
"schedule": 86400.0,
},
"train-heads-nightly": {
"task": "backend.app.tasks.ml.scheduled_train_heads",
"schedule": 86400.0, # passive cadence; manual retrain stays available
},
"apply-head-tags-daily": {
"task": "backend.app.tasks.ml.scheduled_apply_head_tags",
"schedule": 86400.0, # no-op unless head_auto_apply_enabled
},
"recover-orphaned-gpu-jobs": {
"task": "backend.app.tasks.ml.recover_orphaned_gpu_jobs",
"schedule": 60.0, # quick pickup of work a dead agent orphaned
},
"enqueue-ccip-backfill-hourly": {
"task": "backend.app.tasks.ml.enqueue_gpu_backfill",
"schedule": 3600.0, # auto-feed new images (+ retry errored) so
"args": ("ccip",), # the queue keeps moving without the button
},
"enqueue-siglip-backfill-daily": {
"task": "backend.app.tasks.ml.enqueue_gpu_backfill",
"schedule": 86400.0, # drain the concept-crop back-catalogue +
"args": ("siglip",), # retry failed embeds, no button needed
},
"ccip-auto-apply-daily": {
"task": "backend.app.tasks.ml.scheduled_ccip_auto_apply",
"schedule": 86400.0, # no-op unless ccip_auto_apply_enabled
},
"snapshot-head-metrics-daily": {
"task": "backend.app.tasks.maintenance.snapshot_head_metrics",
"schedule": 86400.0,
},
"integrity-verify-weekly": {
"task": "backend.app.tasks.maintenance.verify_integrity",
"schedule": 604800.0, # weekly
@@ -186,18 +130,6 @@ def make_celery() -> Celery:
"task": "backend.app.tasks.maintenance.recover_stalled_library_audit_runs",
"schedule": 300.0,
},
"recover-stalled-tag-eval-runs": {
"task": "backend.app.tasks.maintenance.recover_stalled_tag_eval_runs",
"schedule": 300.0,
},
"recover-stalled-head-training-runs": {
"task": "backend.app.tasks.maintenance.recover_stalled_head_training_runs",
"schedule": 300.0,
},
"recover-stalled-head-auto-apply-runs": {
"task": "backend.app.tasks.maintenance.recover_stalled_head_auto_apply_runs",
"schedule": 300.0,
},
"recover-stalled-import-batches": {
"task": "backend.app.tasks.maintenance.recover_stalled_import_batches",
"schedule": 300.0,
+1 -9
View File
@@ -69,15 +69,7 @@ def _queue_for(task) -> str:
return "ml"
if name.startswith("backend.app.tasks.thumbnail."):
return "thumbnail"
if name.startswith((
"backend.app.tasks.download.",
# External file-host fetches share the download lane (celery_app
# routes external.* → download). Mirror it here or TaskRun.queue
# lies 'default' for them, so per-queue dashboard filters and the
# per-queue threshold override miss them — the same gap the
# 2026-06-02 audit fixed for backup/admin/library_audit.
"backend.app.tasks.external.",
)):
if name.startswith("backend.app.tasks.download."):
return "download"
if name.startswith("backend.app.tasks.scan."):
return "scan"
-22
View File
@@ -8,15 +8,9 @@ from .base import Base
from .credential import Credential
from .download_event import DownloadEvent
from .external_link import ExternalLink
from .gpu_job import GpuJob
from .head_auto_apply_run import HeadAutoApplyRun
from .head_metric import HeadMetric
from .head_metrics_snapshot import HeadMetricsSnapshot
from .head_training_run import HeadTrainingRun
from .image_prediction import ImagePrediction
from .image_provenance import ImageProvenance
from .image_record import ImageRecord
from .image_region import ImageRegion
from .import_batch import ImportBatch
from .import_settings import ImportSettings
from .import_task import ImportTask
@@ -30,14 +24,9 @@ from .series_chapter import SeriesChapter
from .series_page import SeriesPage
from .series_suggestion import SeriesSuggestion
from .source import Source
from .subscribestar_failed_media import SubscribeStarFailedMedia
from .subscribestar_seen_media import SubscribeStarSeenMedia
from .tag import Tag, TagKind, image_tag
from .tag_alias import TagAlias
from .tag_allowlist import TagAllowlist
from .tag_eval_run import TagEvalRun
from .tag_head import TagHead
from .tag_positive_confirmation import TagPositiveConfirmation
from .tag_reference_embedding import TagReferenceEmbedding
from .tag_suggestion_rejection import TagSuggestionRejection
from .task_run import TaskRun
@@ -52,8 +41,6 @@ __all__ = [
"Credential",
"PatreonFailedMedia",
"PatreonSeenMedia",
"SubscribeStarFailedMedia",
"SubscribeStarSeenMedia",
"Post",
"PostAttachment",
"SeriesChapter",
@@ -62,27 +49,18 @@ __all__ = [
"ImageRecord",
"ImagePrediction",
"ImageProvenance",
"ImageRegion",
"Tag",
"TagKind",
"image_tag",
"DownloadEvent",
"ExternalLink",
"GpuJob",
"ImportBatch",
"ImportTask",
"ImportSettings",
"LibraryAuditRun",
"MLSettings",
"HeadAutoApplyRun",
"HeadMetric",
"HeadMetricsSnapshot",
"HeadTrainingRun",
"TagAlias",
"TagAllowlist",
"TagEvalRun",
"TagHead",
"TagPositiveConfirmation",
"TagReferenceEmbedding",
"TagSuggestionRejection",
"TaskRun",
-50
View File
@@ -1,50 +0,0 @@
"""GpuJob — a unit of GPU work the desktop agent pulls over HTTP (#114).
The durable work list that lets the agent stay HTTP-only: the server enqueues a
job per (image, task) — e.g. detect figures + CCIP-embed — and the agent LEASES a
batch, computes on its GPU, then SUBMITS results, all over the already-exposed web
API. Redis/Postgres stay private. A lease has an expiry; the lease query itself
re-claims expired leases (agent died / stopped mid-batch), so the queue is
self-healing without a separate sweep. One job is per ITEM; the agent fans a
VIDEO out into per-frame instances internally (see image_region.frame_time).
State: pending → leased → done | error (a failure under the attempt cap returns to
pending for another agent).
"""
from datetime import datetime
from sqlalchemy import DateTime, ForeignKey, Integer, String, Text, func
from sqlalchemy.orm import Mapped, mapped_column
from .base import Base
class GpuJob(Base):
__tablename__ = "gpu_job"
id: Mapped[int] = mapped_column(Integer, primary_key=True)
image_record_id: Mapped[int] = mapped_column(
ForeignKey("image_record.id", ondelete="CASCADE"), index=True
)
# What to compute, e.g. 'ccip' (detect figures + CCIP-embed) or 'siglip_region'.
task: Mapped[str] = mapped_column(String(32), nullable=False)
status: Mapped[str] = mapped_column(
String(16), nullable=False, default="pending", index=True
)
# pending | leased | done | error
lease_token: Mapped[str | None] = mapped_column(String(64), nullable=True)
leased_at: Mapped[datetime | None] = mapped_column(
DateTime(timezone=True), nullable=True
)
lease_expires_at: Mapped[datetime | None] = mapped_column(
DateTime(timezone=True), nullable=True
)
attempts: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
error: Mapped[str | None] = mapped_column(Text, nullable=True)
created_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now()
)
updated_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now()
)
-46
View File
@@ -1,46 +0,0 @@
"""HeadAutoApplyRun — persisted lifecycle of an earned-auto-apply sweep (#114).
A graduated head can apply its tag to images it scores above the head's
auto-apply threshold, without a human. This row tracks one such sweep (or a
dry-run PREVIEW of it) so the result survives navigation and the admin card can
show what fired / what would fire. Mirrors HeadTrainingRun. State machine:
running → ready / error. The `report` JSONB holds per-concept counts
(applied / projected / scanned).
"""
from datetime import datetime
from typing import Any
from sqlalchemy import Boolean, DateTime, Integer, String, Text, func
from sqlalchemy.dialects.postgresql import JSONB
from sqlalchemy.orm import Mapped, mapped_column
from .base import Base
class HeadAutoApplyRun(Base):
__tablename__ = "head_auto_apply_run"
id: Mapped[int] = mapped_column(Integer, primary_key=True)
# dry_run=True is a PREVIEW: scores + counts what WOULD apply, writes nothing
# (preview/apply parity, rule 93).
dry_run: Mapped[bool] = mapped_column(Boolean, nullable=False, default=False)
params: Mapped[dict[str, Any]] = mapped_column(JSONB, nullable=False)
status: Mapped[str] = mapped_column(
String(16), nullable=False, default="running", index=True
)
# running | ready | error
started_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now()
)
finished_at: Mapped[datetime | None] = mapped_column(
DateTime(timezone=True), nullable=True
)
# Total tags applied across all heads this sweep (0 for a clean dry-run).
n_applied: Mapped[int | None] = mapped_column(Integer, nullable=True)
# Per-concept breakdown: [{tag_id, name, applied, scanned, threshold}, ...].
report: Mapped[dict[str, Any] | None] = mapped_column(JSONB, nullable=True)
error: Mapped[str | None] = mapped_column(Text, nullable=True)
last_progress_at: Mapped[datetime | None] = mapped_column(
DateTime(timezone=True), nullable=True
)
-32
View File
@@ -1,32 +0,0 @@
"""HeadMetric — running correction counters per concept (#114 observability).
Earned auto-apply fires graduated heads; to TUNE it we need to know how often a
head's auto-applied tag was wrong (the operator removed it = a MISFIRE) and how
often the operator had to add a tag a head exists for by hand (an UNDER-FIRE,
the head missed it). image_tag.source is lost when a row is deleted, so these
are captured as durable cumulative counters at correction time — they survive
head retrain/prune (keyed by tag, not by the head row). The daily snapshot reads
them into the time-series.
"""
from datetime import datetime
from sqlalchemy import DateTime, ForeignKey, Integer, func
from sqlalchemy.orm import Mapped, mapped_column
from .base import Base
class HeadMetric(Base):
__tablename__ = "head_metric"
tag_id: Mapped[int] = mapped_column(
ForeignKey("tag.id", ondelete="CASCADE"), primary_key=True
)
# An auto-applied (source='head_auto') tag the operator later REMOVED.
n_misfires: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
# A tag with a head that the operator added by HAND (the head missed it).
n_underfires: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
updated_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now()
)
@@ -1,38 +0,0 @@
"""HeadMetricsSnapshot — a daily per-concept time-series point (#114).
The "amount of change over time" reporting the operator asked for: once a day,
record each concept's auto-applied VOLUME (current head_auto tags), cumulative
misfires/under-fires, and the head's measured quality. Plotting these rows over
time shows whether auto-apply is landing better/worse and whether tagging more is
sharpening a concept — the signal for tuning the precision target + support floor.
"""
from datetime import datetime
from sqlalchemy import DateTime, Float, ForeignKey, Integer, String, func
from sqlalchemy.orm import Mapped, mapped_column
from .base import Base
class HeadMetricsSnapshot(Base):
__tablename__ = "head_metrics_snapshot"
id: Mapped[int] = mapped_column(Integer, primary_key=True)
tag_id: Mapped[int] = mapped_column(
ForeignKey("tag.id", ondelete="CASCADE"), index=True
)
# Denormalized so a snapshot stays readable even if the tag is later renamed.
name: Mapped[str] = mapped_column(String(255), nullable=False)
snapshot_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now(), index=True
)
# Current count of source='head_auto' applications still standing.
n_auto_applied: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
n_misfires: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
n_underfires: Mapped[int] = mapped_column(Integer, nullable=False, default=0)
# The head's measured quality at snapshot time (null if no head exists).
ap: Mapped[float | None] = mapped_column(Float, nullable=True)
precision_cv: Mapped[float | None] = mapped_column(Float, nullable=True)
recall: Mapped[float | None] = mapped_column(Float, nullable=True)
n_pos: Mapped[int | None] = mapped_column(Integer, nullable=True)
-44
View File
@@ -1,44 +0,0 @@
"""HeadTrainingRun — persisted lifecycle of a head-training batch (#114).
Mirrors TagEvalRun so the run SURVIVES navigation and the admin card can show
live + historical status instead of holding it in transient frontend state.
Training is idempotent (it upserts tag_head rows), so a SIGKILL'd run is harmless
— a maintenance recovery sweep flips a stalled `running` row to `error`, and the
next run re-trains. State machine: running → ready / error.
"""
from datetime import datetime
from typing import Any
from sqlalchemy import DateTime, Integer, String, Text, func
from sqlalchemy.dialects.postgresql import JSONB
from sqlalchemy.orm import Mapped, mapped_column
from .base import Base
class HeadTrainingRun(Base):
__tablename__ = "head_training_run"
id: Mapped[int] = mapped_column(Integer, primary_key=True)
# Training parameters: {min_positives, neg_ratio, precision_target, ...}.
params: Mapped[dict[str, Any]] = mapped_column(JSONB, nullable=False)
status: Mapped[str] = mapped_column(
String(16), nullable=False, default="running", index=True
)
# running | ready | error
started_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now()
)
finished_at: Mapped[datetime | None] = mapped_column(
DateTime(timezone=True), nullable=True
)
# How many concepts got a (re)trained head vs were skipped (too few labels).
n_trained: Mapped[int | None] = mapped_column(Integer, nullable=True)
n_skipped: Mapped[int | None] = mapped_column(Integer, nullable=True)
error: Mapped[str | None] = mapped_column(Text, nullable=True)
# Last time the task made progress — the recovery sweep tells a live run from
# a SIGKILL'd one by this (mirrors TagEvalRun).
last_progress_at: Mapped[datetime | None] = mapped_column(
DateTime(timezone=True), nullable=True
)
-10
View File
@@ -41,16 +41,6 @@ class ImageProvenance(Base):
source_id: Mapped[int | None] = mapped_column(
ForeignKey("source.id", ondelete="SET NULL"), nullable=True, index=True
)
# The archive PostAttachment this image was extracted FROM, when it came
# out of a .zip/.rar rather than as a loose file (milestone #87). Lets the
# provenance UI show the exact archive a file lives inside instead of every
# attachment on the post. NULL for loose downloads and pre-backfill rows.
# SET NULL so deleting the archive attachment never destroys the (image,
# post) edge — it just forgets which archive it came from.
from_attachment_id: Mapped[int | None] = mapped_column(
ForeignKey("post_attachment.id", ondelete="SET NULL"),
nullable=True, index=True,
)
captured_metadata: Mapped[dict | None] = mapped_column(JSON, nullable=True)
captured_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now()
-62
View File
@@ -1,62 +0,0 @@
"""ImageRegion — a detected/proposed sub-region of an image + its crop embedding.
The storage backbone of the crop pipeline (#114). A region is a normalized bbox
plus the embedding of its crop:
- kind='face' / 'figure' → embedded by CCIP for cross-artist character identity.
- kind='concept' → embedded by SigLIP, a localized instance for a concept head's
bag-of-embeddings (a concept is "present if ANY instance matches").
One row carries the embedding appropriate to its kind (the other is null). The
bbox doubles as grounded-tag provenance (hover a tag → highlight its region; a
wrong box is a precise negative). The GPU agent writes these via the job API;
the few-shot character matcher + bag scorer read them — both server-side, no GPU.
"""
from datetime import datetime
from pgvector.sqlalchemy import Vector
from sqlalchemy import DateTime, Float, ForeignKey, Integer, String, func
from sqlalchemy.orm import Mapped, mapped_column
from .base import Base
CCIP_DIM = 768 # deepghs/imgutils CCIP character embedding
SIGLIP_DIM = 1152 # matches image_record.siglip_embedding
class ImageRegion(Base):
__tablename__ = "image_region"
id: Mapped[int] = mapped_column(Integer, primary_key=True)
image_record_id: Mapped[int] = mapped_column(
ForeignKey("image_record.id", ondelete="CASCADE"), index=True
)
# 'frame' (a whole video frame → SigLIP bag) | 'face' | 'figure' (→ CCIP
# character id) | 'concept' (→ SigLIP head bag).
kind: Mapped[str] = mapped_column(String(16), nullable=False)
# For video/animated media: the source frame's timestamp in SECONDS. NULL for
# static images. Lets a video be a BAG of per-frame instances (fixes the
# mean-embedding muddle) + grounds a tag to "appears at 0:42".
frame_time: Mapped[float | None] = mapped_column(Float, nullable=True)
# Normalized bbox in [0,1]: top-left (rx, ry) + size (rw, rh). Named rx/ry/…
# rather than x/y/by to dodge SQL keyword ambiguity ('by').
rx: Mapped[float] = mapped_column(Float, nullable=False)
ry: Mapped[float] = mapped_column(Float, nullable=False)
rw: Mapped[float] = mapped_column(Float, nullable=False)
rh: Mapped[float] = mapped_column(Float, nullable=False)
# Proposer/detector confidence (null for deterministic proposers).
score: Mapped[float | None] = mapped_column(Float, nullable=True)
# Version stamps so a re-detect / re-crop / re-embed can be gated (compute
# once; only redo when the producing model version changes).
detector_version: Mapped[str | None] = mapped_column(String(64), nullable=True)
crop_version: Mapped[str | None] = mapped_column(String(64), nullable=True)
embedding_version: Mapped[str | None] = mapped_column(String(128), nullable=True)
# Exactly one is set, per kind.
ccip_embedding: Mapped[list[float] | None] = mapped_column(
Vector(CCIP_DIM), nullable=True
)
siglip_embedding: Mapped[list[float] | None] = mapped_column(
Vector(SIGLIP_DIM), nullable=True
)
created_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now()
)
+1 -47
View File
@@ -2,15 +2,7 @@
from datetime import datetime
from sqlalchemy import (
Boolean,
CheckConstraint,
DateTime,
Float,
Integer,
String,
func,
)
from sqlalchemy import CheckConstraint, DateTime, Float, Integer, String, func
from sqlalchemy.orm import Mapped, mapped_column
from .base import Base
@@ -63,44 +55,6 @@ class MLSettings(Base):
video_min_tag_frames: Mapped[int] = mapped_column(
Integer, nullable=False, default=3
)
# Tagging-v2 head training (#114). The head is the suggestion source that
# LEARNS from the operator's tags (replacing Camie + centroid). A concept
# needs >= head_min_positives labelled images before a head is trained;
# head_auto_apply_precision is the precision bar a head must clear (at some
# operating point) to "graduate" into earned auto-apply. Operator-tunable.
head_min_positives: Mapped[int] = mapped_column(
Integer, nullable=False, default=8
)
head_auto_apply_precision: Mapped[float] = mapped_column(
Float, nullable=False, default=0.97
)
# Earned auto-apply (#114). A graduated head fires (tags images without a
# human) when this master switch is on AND the head has at least
# head_auto_apply_min_positives clean labels — so a precise-looking but
# under-supported low-N head can't spray tags across the library. ON by
# default (operator-asked 2026-06-29: opt-OUT, not opt-in); the support +
# measured-precision gates keep it safe, and every auto-tag is reversible.
head_auto_apply_enabled: Mapped[bool] = mapped_column(
Boolean, nullable=False, default=True
)
head_auto_apply_min_positives: Mapped[int] = mapped_column(
Integer, nullable=False, default=30
)
# CCIP character-match cosine cut (#114). 0.85 default — the v1 flat 0.75
# over-fired (high-reference characters matched a scatter of images); 0.85
# keeps the confident single-character matches. Tunable from the agent card.
ccip_match_threshold: Mapped[float] = mapped_column(
Float, nullable=False, default=0.85
)
# CCIP auto-apply (#114). Confident matches (>= ccip_auto_apply_threshold,
# above the suggest cut) auto-tag on a daily sweep. ON by default (opt-out);
# single-character references + the high bar keep it safe, every tag reversible.
ccip_auto_apply_enabled: Mapped[bool] = mapped_column(
Boolean, nullable=False, default=True
)
ccip_auto_apply_threshold: Mapped[float] = mapped_column(
Float, nullable=False, default=0.92
)
tagger_model_version: Mapped[str] = mapped_column(
String(128), nullable=False, default="camie-tagger-v2"
)
@@ -1,44 +0,0 @@
"""SubscribeStarFailedMedia — per-source dead-letter ledger of SubscribeStar
media that keeps failing to download/validate.
Mirror of PatreonFailedMedia. Media that fails every walk (404'd CDN URL,
deleted post, persistently-corrupt bytes) would otherwise re-error forever and
re-burn backfill chunks. After ``attempts`` reaches the dead-letter threshold
the ingester skips it on routine tick/backfill walks (recovery still
re-attempts). A later clean download clears the row.
`filehash` is the same per-media key the seen-ledger uses (CDN content hash or a
synthesized ``<post_id>:<filename>`` key) — hence String(128). UNIQUE
(source_id, filehash) is the upsert key.
"""
from datetime import datetime
from sqlalchemy import ForeignKey, Integer, String, Text, UniqueConstraint, func
from sqlalchemy.orm import Mapped, mapped_column
from sqlalchemy.types import DateTime
from .base import Base
class SubscribeStarFailedMedia(Base):
__tablename__ = "subscribestar_failed_media"
__table_args__ = (
UniqueConstraint(
"source_id", "filehash", name="uq_subscribestar_failed_media_source_id"
),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True)
source_id: Mapped[int] = mapped_column(
ForeignKey("source.id", ondelete="CASCADE"), nullable=False, index=True
)
filehash: Mapped[str] = mapped_column(String(128), nullable=False)
attempts: Mapped[int] = mapped_column(Integer, nullable=False, default=1)
last_error: Mapped[str | None] = mapped_column(Text, nullable=True)
first_failed_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now()
)
last_failed_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now()
)
@@ -1,40 +0,0 @@
"""SubscribeStarSeenMedia — per-source ledger of SubscribeStar media already
downloaded+processed.
Mirror of PatreonSeenMedia for the SubscribeStar native ingester (replacing
gallery-dl). One queryable row per (source, media) so routine walks skip media
we've already ingested; recovery mode bypasses the ledger to re-walk.
`filehash` is a CDN content hash when the media URL carries one, else a
synthesized ``<post_id>:<filename>`` key (SubscribeStar URLs aren't always
content-addressed) — hence String(128) rather than 32.
"""
from datetime import datetime
from sqlalchemy import ForeignKey, Integer, String, UniqueConstraint, func
from sqlalchemy.orm import Mapped, mapped_column
from sqlalchemy.types import DateTime
from .base import Base
class SubscribeStarSeenMedia(Base):
__tablename__ = "subscribestar_seen_media"
__table_args__ = (
# Dedup key the downloader upserts against: one ledger row per
# (source, media). A second sighting of the same media is a no-op.
UniqueConstraint(
"source_id", "filehash", name="uq_subscribestar_seen_media_source_id"
),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True)
source_id: Mapped[int] = mapped_column(
ForeignKey("source.id", ondelete="CASCADE"), nullable=False, index=True
)
filehash: Mapped[str] = mapped_column(String(128), nullable=False)
post_id: Mapped[str | None] = mapped_column(String(64), nullable=True)
seen_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now()
)
-45
View File
@@ -1,45 +0,0 @@
"""TagEvalRun — persisted lifecycle of a head-vs-centroid tagging eval (#1130).
Mirrors LibraryAuditRun so the result SURVIVES navigation: the run + its full
report live in this row, and the admin card rehydrates from it on mount instead
of holding the report in transient frontend state. State machine:
running → ready / error. The async ml-queue task writes `report` (JSONB) when
done; a maintenance recovery sweep flips a stalled `running` row to `error`.
"""
from datetime import datetime
from typing import Any
from sqlalchemy import DateTime, Integer, String, Text, func
from sqlalchemy.dialects.postgresql import JSONB
from sqlalchemy.orm import Mapped, mapped_column
from .base import Base
class TagEvalRun(Base):
__tablename__ = "tag_eval_run"
id: Mapped[int] = mapped_column(Integer, primary_key=True)
# The eval parameters: {concepts: [...], curve_points: [...], neg_ratio,
# cv_folds, ...} — echoed back so the report is self-describing.
params: Mapped[dict[str, Any]] = mapped_column(JSONB, nullable=False)
status: Mapped[str] = mapped_column(
String(16), nullable=False, default="running", index=True,
)
# running | ready | error
started_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now(),
)
finished_at: Mapped[datetime | None] = mapped_column(
DateTime(timezone=True), nullable=True,
)
# The full result: per-concept metrics (head vs centroid), learning-curve
# points, and example image ids. Null until the task finishes.
report: Mapped[dict[str, Any] | None] = mapped_column(JSONB, nullable=True)
error: Mapped[str | None] = mapped_column(Text, nullable=True)
# Last time the task made progress — the recovery sweep tells a live run
# from a SIGKILL'd one by this (mirrors LibraryAuditRun).
last_progress_at: Mapped[datetime | None] = mapped_column(
DateTime(timezone=True), nullable=True,
)
-77
View File
@@ -1,77 +0,0 @@
"""TagHead — a small per-concept classifier trained on the operator's tags.
Milestone #114, tagging-v2: the production form of the head the eval (#1130)
proved. One row per concept (general or character) that has enough labelled
positives. The head is a logistic-regression boundary over the FROZEN SigLIP
embedding (L2-normalized), trained on the operator's positives + negatives
(rejections + sampled unlabeled). It REPLACES the Camie prediction + per-tag
centroid as the suggestion source — and unlike them it LEARNS: every accept /
reject re-trains it sharper.
Scoring (suggestion path, API worker, NO numpy): p = sigmoid(weights · x̂ + bias)
where x̂ is the L2-normalized image embedding. Surface as a suggestion when
p >= suggest_threshold; auto-apply only once auto_apply_threshold is set (the
head "graduated" — a precision-targeted operating point was achievable). The
thresholds come from CROSS-VALIDATED out-of-fold scores so they're honest, not
in-sample-optimistic; the deployable weights are fit on all data.
"""
from datetime import datetime
from typing import Any
from pgvector.sqlalchemy import Vector
from sqlalchemy import (
DateTime,
Float,
ForeignKey,
Integer,
String,
func,
)
from sqlalchemy.dialects.postgresql import JSONB
from sqlalchemy.orm import Mapped, mapped_column
from .base import Base
# Matches image_record.siglip_embedding's dimensionality — the head operates in
# the same space. A model-version change re-embeds AND retrains (embedding_version
# guards staleness).
HEAD_DIM = 1152
class TagHead(Base):
__tablename__ = "tag_head"
# One head per concept tag; cascade so deleting a tag retires its head.
tag_id: Mapped[int] = mapped_column(
ForeignKey("tag.id", ondelete="CASCADE"), primary_key=True
)
# The embedding the head was trained against (image_record's
# embedder_model_version). A mismatch with the current embedder means the
# head is stale and must be retrained, not scored.
embedding_version: Mapped[str] = mapped_column(String(128), nullable=False)
# Logistic-regression coefficients over the L2-normalized embedding, stored
# as a pgvector for compactness + a future in-DB dot-product path. NOT a
# similarity target, just a serialized weight vector.
weights: Mapped[list[float]] = mapped_column(Vector(HEAD_DIM), nullable=False)
bias: Mapped[float] = mapped_column(Float, nullable=False)
# Probability cutoff for SURFACING as a suggestion (F1-best on CV scores).
suggest_threshold: Mapped[float] = mapped_column(Float, nullable=False)
# Probability cutoff for EARNED auto-apply: the operating point that holds
# precision >= the configured target while maximizing recall. NULL = the head
# hasn't graduated (can't auto-apply without a human yet).
auto_apply_threshold: Mapped[float | None] = mapped_column(Float, nullable=True)
# Training-set sizes + cross-validated quality, surfaced in the admin card so
# the operator can see which concepts are strong / need more tags.
n_pos: Mapped[int] = mapped_column(Integer, nullable=False)
n_neg: Mapped[int] = mapped_column(Integer, nullable=False)
ap: Mapped[float] = mapped_column(Float, nullable=False)
# 'precision' is a SQL reserved word → store as precision_cv (the
# cross-validated precision at the suggest operating point).
precision_cv: Mapped[float] = mapped_column(Float, nullable=False)
recall: Mapped[float] = mapped_column(Float, nullable=False)
trained_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now()
)
# Extra detail (auto-apply operating point, F1, etc.) — non-load-bearing.
metrics: Mapped[dict[str, Any] | None] = mapped_column(JSONB, nullable=True)
@@ -1,28 +0,0 @@
"""TagPositiveConfirmation — operator affirmed an applied tag is correct.
The mirror of TagSuggestionRejection (#1130). When the operator "keeps" a
positive the head doubts (low-scoring), record it so the eval's doubts list
stops resurfacing the same confirmed-correct images every run. Does not change
training (it's already a positive) — purely a "I've reviewed this" marker.
"""
from datetime import datetime
from sqlalchemy import DateTime, ForeignKey, func
from sqlalchemy.orm import Mapped, mapped_column
from .base import Base
class TagPositiveConfirmation(Base):
__tablename__ = "tag_positive_confirmation"
image_record_id: Mapped[int] = mapped_column(
ForeignKey("image_record.id", ondelete="CASCADE"), primary_key=True
)
tag_id: Mapped[int] = mapped_column(
ForeignKey("tag.id", ondelete="CASCADE"), primary_key=True, index=True
)
confirmed_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now()
)
+5 -217
View File
@@ -18,11 +18,10 @@ from pathlib import Path
from typing import Any
from sqlalchemy import delete, func, or_, select, update
from sqlalchemy.orm import Session, aliased
from sqlalchemy.orm import Session
from ..models import (
Artist,
ExternalLink,
ImageProvenance,
ImageRecord,
LibraryAuditRun,
@@ -37,7 +36,6 @@ from ..models.series_page import SeriesPage
from ..models.tag import image_tag
from ..utils import safe_probe
from .importer import _VIDEO_DUP_ASPECT_TOL, _VIDEO_DUP_DURATION_TOL_SECONDS
from .platforms import PLATFORMS
log = logging.getLogger(__name__)
@@ -151,22 +149,11 @@ def project_bulk_image_delete(
def count_tag_associations(session: Session, *, tag_id: int) -> int:
"""Images affected by deleting this tag — the Tier-B blast-radius prompt.
Mirrors the gallery/directory membership predicate: images carrying the tag
DIRECTLY, plus — when it's a fandom — images carrying one of its characters
(member.fandom_id == tag_id). DISTINCT so each image counts once. Without
the character leg a fandom would report 0 here yet its delete still strips
the fandom off every character, badly understating the prompt."""
member = aliased(Tag)
"""COUNT(*) FROM image_tag WHERE tag_id=?. For Tier-B prompt."""
return session.execute(
select(func.count(image_tag.c.image_record_id.distinct())).where(
or_(
image_tag.c.tag_id == tag_id,
image_tag.c.tag_id.in_(
select(member.id).where(member.fandom_id == tag_id)
),
)
)
select(func.count())
.select_from(image_tag)
.where(image_tag.c.tag_id == tag_id)
).scalar_one()
@@ -520,205 +507,6 @@ def prune_bare_posts(session: Session, *, dry_run: bool = False) -> dict:
return {"deleted": result.rowcount or 0, "sample_names": sample}
# -- duplicate-post reconciliation (gallery-dl → native migration) ----------
# An artist first downloaded by gallery-dl gets Post rows keyed by the per-
# ATTACHMENT id (gallery-dl's `id`); a later native walk keys the SAME real post
# by the post id. The two never dedup (uq_post_source_external_id is on
# external_post_id) → duplicate post rows. The real post id is recoverable in-DB
# from raw_metadata["post_id"] (both eras store the sidecar there). We unify each
# group onto ONE post row keyed the way the CURRENT native downloader keys it
# (post id), so future native walks match and the dup can't recur. Images are
# untouched (content-addressed/deduped already); only post rows + their link
# rows move. Milestone #73 / note #917.
def _canonical_post_id(post: Post) -> str | None:
"""The real platform post id used to group duplicate rows: raw_metadata
['post_id'] when present (the true id, stored by both gallery-dl and native
imports), else external_post_id. None when neither is usable."""
rm = post.raw_metadata or {}
pid = rm.get("post_id")
if pid is not None and str(pid).strip():
return str(pid)
return post.external_post_id or None
def find_duplicate_post_groups(
session: Session, *, source_id: int | None = None,
) -> list[list[Post]]:
"""Groups of >1 Post that are the SAME real post (same source_id + canonical
post id) — the gallery-dl(attachment-id) / native(post-id) duplicates. Shared
by the dry-run preview and the live reconcile (preview/apply parity)."""
stmt = select(Post)
if source_id is not None:
stmt = stmt.where(Post.source_id == source_id)
groups: dict[tuple, list[Post]] = {}
for post in session.execute(stmt.order_by(Post.id)).scalars().all():
cpid = _canonical_post_id(post)
if not cpid:
continue
groups.setdefault((post.source_id, cpid), []).append(post)
return [posts for posts in groups.values() if len(posts) > 1]
def _choose_keeper(posts: list[Post], cpid: str) -> Post:
"""The surviving row for a dup group: prefer one already keyed by the
canonical post id (the native format we keep), then the most complete
(has description, has date), then the lowest id for stability."""
native = [p for p in posts if p.external_post_id == cpid]
pool = native or posts
return min(pool, key=lambda p: (not p.description, not p.post_date, p.id))
def _repoint_post_links(session: Session, loser_id: int, keeper_id: int) -> None:
"""Move every link row from a loser post to the keeper, conflict-safe against
each table's uniqueness (drop the loser's row when the keeper already has the
equivalent, else re-point). Images themselves are never touched."""
# ImageRecord.primary_post_id — no uniqueness; straight re-point.
session.execute(
update(ImageRecord)
.where(ImageRecord.primary_post_id == loser_id)
.values(primary_post_id=keeper_id)
)
# ImageProvenance — unique (image_record_id, post_id).
dup_imgs = select(ImageProvenance.image_record_id).where(
ImageProvenance.post_id == keeper_id
)
# Before dropping the colliding loser rows, carry their from_attachment_id
# (which archive the file came out of, milestone #87) onto the keeper's
# surviving row when the keeper didn't record one. For the gallery-dl→native
# case this very milestone targets, the keeper is the native stub (no
# archive) and the loser is the gallery-dl row that extracted the member, so
# a blind delete would silently lose the containing-archive linkage.
for img_id, att_id in session.execute(
select(ImageProvenance.image_record_id, ImageProvenance.from_attachment_id)
.where(
ImageProvenance.post_id == loser_id,
ImageProvenance.image_record_id.in_(dup_imgs),
ImageProvenance.from_attachment_id.is_not(None),
)
).all():
session.execute(
update(ImageProvenance)
.where(
ImageProvenance.post_id == keeper_id,
ImageProvenance.image_record_id == img_id,
ImageProvenance.from_attachment_id.is_(None),
)
.values(from_attachment_id=att_id)
)
session.execute(
delete(ImageProvenance).where(
ImageProvenance.post_id == loser_id,
ImageProvenance.image_record_id.in_(dup_imgs),
)
)
session.execute(
update(ImageProvenance)
.where(ImageProvenance.post_id == loser_id)
.values(post_id=keeper_id)
)
# PostAttachment — partial unique (post_id, sha256) where post_id NOT NULL.
dup_shas = select(PostAttachment.sha256).where(PostAttachment.post_id == keeper_id)
session.execute(
delete(PostAttachment).where(
PostAttachment.post_id == loser_id,
PostAttachment.sha256.in_(dup_shas),
)
)
session.execute(
update(PostAttachment)
.where(PostAttachment.post_id == loser_id)
.values(post_id=keeper_id)
)
# ExternalLink — unique (post_id, url).
dup_urls = select(ExternalLink.url).where(ExternalLink.post_id == keeper_id)
session.execute(
delete(ExternalLink).where(
ExternalLink.post_id == loser_id,
ExternalLink.url.in_(dup_urls),
)
)
session.execute(
update(ExternalLink)
.where(ExternalLink.post_id == loser_id)
.values(post_id=keeper_id)
)
def _fill_missing_post_fields(keeper: Post, loser: Post) -> None:
"""Backfill the keeper's empty metadata from a loser that has it (the native
stub is often bare; the gallery-dl row carries date/title/body/raw_metadata)."""
if not keeper.post_date and loser.post_date:
keeper.post_date = loser.post_date
if not keeper.post_title and loser.post_title:
keeper.post_title = loser.post_title
if not keeper.description and loser.description:
keeper.description = loser.description
if not keeper.raw_metadata and loser.raw_metadata:
keeper.raw_metadata = loser.raw_metadata
if not keeper.attachment_count and loser.attachment_count:
keeper.attachment_count = loser.attachment_count
def _canonical_post_url(post: Post, cpid: str) -> str | None:
"""Permalink for the unified post, via the platform's derive_post_url hook
(subscribestar/patreon synthesize `…/posts/<id>`). Falls back to the keeper's
existing url when no hook applies."""
platform = (post.raw_metadata or {}).get("category")
info = PLATFORMS.get(platform) if platform else None
if info is not None and info.derive_post_url is not None:
derived = info.derive_post_url({"post_id": cpid})
if derived:
return derived
return post.post_url
def reconcile_duplicate_posts(
session: Session, *, source_id: int | None = None, dry_run: bool = False,
) -> dict:
"""Unify duplicate post rows (gallery-dl attachment-id + native post-id) onto
one keeper per real post, re-keyed to the post id. Images untouched.
Returns:
dry_run=True: {"groups": G, "posts_to_merge": L, "sample": [...]}
dry_run=False: {"groups": G, "merged": L, "sample": [...]}
where L = rows that would be (were) deleted after merging into keepers. The
SAME find_duplicate_post_groups predicate drives preview and apply (rule 93).
"""
groups = find_duplicate_post_groups(session, source_id=source_id)
sample: list[dict] = []
losers_total = 0
for posts in groups:
cpid = _canonical_post_id(posts[0])
keeper = _choose_keeper(posts, cpid)
losers = [p for p in posts if p.id != keeper.id]
losers_total += len(losers)
if len(sample) < 50:
sample.append({
"post_id": cpid,
"rows": len(posts),
"keeper_id": keeper.id,
"title": keeper.post_title or f"Post {cpid}",
})
if dry_run:
continue
for loser in losers:
_repoint_post_links(session, loser.id, keeper.id)
_fill_missing_post_fields(keeper, loser)
keeper.external_post_id = cpid
new_url = _canonical_post_url(keeper, cpid)
if new_url:
keeper.post_url = new_url
session.flush()
for loser in losers:
session.delete(loser)
if dry_run:
return {"groups": len(groups), "posts_to_merge": losers_total, "sample": sample}
session.commit()
return {"groups": len(groups), "merged": losers_total, "sample": sample}
# Legacy tags FC no longer uses, in two shapes:
# (1) kinds the tag input never produces — archive/post/artist.
# provenance (post grouping) + archive membership are their own
+20 -50
View File
@@ -24,32 +24,19 @@ import asyncio
from pathlib import Path
from .gallery_dl import DownloadResult, ErrorType
from .native_ingest_common import NativeIngestError
from .patreon_ingester import PatreonIngester
from .patreon_resolver import extract_vanity, resolve_campaign_id_for_source
from .subscribestar_ingester import SubscribeStarIngester
# Platforms whose download + verify go through the native ingester rather than
# gallery-dl. gallery-dl still serves the rest (hentaifoundry, discord, pixiv,
# deviantart) until they migrate too.
NATIVE_INGESTER_PLATFORMS = frozenset({"patreon", "subscribestar"})
# gallery-dl. gallery-dl still serves every other platform (subscribestar,
# hentaifoundry, discord, pixiv, deviantart) unchanged.
NATIVE_INGESTER_PLATFORMS = frozenset({"patreon"})
# Mirrors patreon_resolver._CAMPAIGNS_URL — surfaced in resolution-failure
# messages so the operator sees the exact lookup endpoint that was hit.
_CAMPAIGNS_API = "https://www.patreon.com/api/campaigns"
def _native_ingester_cls(platform: str):
"""The native ingester class for `platform` (uniform constructor signature).
A call-time lookup (not a module-level dict captured at import) so tests can
monkeypatch db_mod.PatreonIngester / SubscribeStarIngester and have the
dispatch pick up the replacement."""
return {
"patreon": PatreonIngester,
"subscribestar": SubscribeStarIngester,
}[platform]
def uses_native_ingester(platform: str) -> bool:
"""True when `platform` is served by the native ingester (not gallery-dl).
The single predicate the download path and verify both route on."""
@@ -94,36 +81,22 @@ async def run_download(
return result, None
async def _resolve_native_campaign_id(
platform: str, url: str, cookies_path: str | None, overrides: dict,
) -> tuple[str | None, str | None]:
"""`(campaign_id, resolved_campaign_id)` for a native source. SubscribeStar's
feed id IS the creator URL (no lookup → resolved None). Patreon resolves the
campaign id from the vanity URL (resolved non-None when a lookup actually ran,
so phase 3 caches it)."""
if platform == "subscribestar":
return url, None
return await resolve_campaign_id_for_source(url, cookies_path, overrides)
async def _run_native_ingester(
ctx: dict, source_config, mode: str | None, gdl, sync_session_factory,
) -> tuple[DownloadResult, str | None]:
"""Run the native ingester for a native platform in a worker thread (sync
requests/subprocess). Patreon resolves a campaign id from the vanity URL;
SubscribeStar's feed id is the creator URL itself. A campaign id we cannot
resolve is a loud NOT_FOUND — never a silent empty success.
"""Patreon (today the only native platform): resolve the campaign id, then run
the native ingester in a worker thread (it is sync requests/subprocess).
`resolved_campaign_id` is non-None only when a lookup ran this call, so phase
3 caches it the way the old gallery-dl retry did.
`resolved_campaign_id` is non-None only when we had to look it up from the
vanity URL this run, so phase 3 caches it the way the old gallery-dl retry
did. A campaign id we cannot resolve is a loud NOT_FOUND — never a silent
empty success.
"""
platform = ctx["platform"]
overrides = ctx["config_overrides"] or {}
campaign_id, resolved_campaign_id = await _resolve_native_campaign_id(
platform, ctx["url"], ctx["cookies_path"], overrides
campaign_id, resolved_campaign_id = await resolve_campaign_id_for_source(
ctx["url"], ctx["cookies_path"], overrides
)
if not campaign_id:
# Only reachable for Patreon (SubscribeStar's campaign id is the URL).
url = ctx["url"]
vanity = extract_vanity(url)
return (
@@ -131,7 +104,7 @@ async def _run_native_ingester(
success=False,
url=url,
artist_slug=ctx["artist_slug"],
platform=platform,
platform="patreon",
error_type=ErrorType.NOT_FOUND,
error_message=(
f"Could not resolve Patreon campaign id. source_url={url!r}; "
@@ -154,7 +127,7 @@ async def _run_native_ingester(
if source_config.sleep_request is not None
else max(0.5, rate_limit / 4)
)
ingester = _native_ingester_cls(platform)(
ingester = PatreonIngester(
images_root=gdl.images_root,
cookies_path=ctx["cookies_path"],
session_factory=sync_session_factory,
@@ -201,8 +174,10 @@ async def preview_source(
"""
import asyncio
campaign_id, _ = await _resolve_native_campaign_id(
platform, url, cookies_path, config_overrides or {}
from .patreon_client import PatreonAPIError
campaign_id, _ = await resolve_campaign_id_for_source(
url, cookies_path, config_overrides or {}
)
if not campaign_id:
vanity = extract_vanity(url)
@@ -213,7 +188,7 @@ async def preview_source(
"(cookies expired, or the creator moved/renamed?)."
)
}
ingester = _native_ingester_cls(platform)(
ingester = PatreonIngester(
images_root=images_root,
cookies_path=cookies_path,
session_factory=sync_session_factory,
@@ -224,7 +199,7 @@ async def preview_source(
None,
lambda: ingester.preview(source_id, campaign_id, page_limit=page_limit),
)
except NativeIngestError as exc:
except PatreonAPIError as exc:
return {"error": f"Couldn't preview: {exc}"}
return result
@@ -246,12 +221,7 @@ async def verify_source_credential(
"""
if uses_native_ingester(platform):
# Native ingester platforms verify via their own lightweight auth probe
# (one authenticated feed fetch). SubscribeStar's probe takes the creator
# URL directly; Patreon's resolves the campaign id first.
if platform == "subscribestar":
from .subscribestar_ingester import verify_subscribestar_credential
return await verify_subscribestar_credential(url, cookies_path, config_overrides)
# (resolve campaign id + one authenticated API page). Patreon today.
from .patreon_ingester import verify_patreon_credential
return await verify_patreon_credential(url, cookies_path, config_overrides)
+27 -71
View File
@@ -9,8 +9,7 @@ call `fetch_external()` directly.
No single tool covers all five hosts, so a small registry maps host → fetch
function behind one signature:
fetch_external(host, url, dest_dir, *, read_timeout, total_timeout, should_stop)
-> FetchResult
fetch_external(host, url, dest_dir, *, timeout, should_stop) -> FetchResult
Backends:
- dropbox : force the direct-download variant (dl=1) + stream GET.
@@ -31,7 +30,6 @@ import logging
import os
import re
import subprocess
import time
from collections.abc import Callable
from dataclasses import dataclass, field
from pathlib import Path
@@ -42,19 +40,7 @@ import requests
log = logging.getLogger(__name__)
_CHUNK = 1 << 16
# Two distinct limits, because conflating them (the old single 3000s value) meant
# a stalled HTTP connection tied up a download-worker slot + the per-host lock for
# ~50 min before failing (operator-flagged 2026-06-17):
# * READ timeout — max idle gap between bytes on an HTTP socket. A stalled host
# (socket open, nothing flowing) is the common failure mode; a short read
# timeout fails it fast. This is what requests' `timeout` actually enforces —
# per-read, never a total.
# * TOTAL budget — generous wall-clock cap for a file that IS actively
# transferring (big films/packs). Enforced as a deadline across chunks, since
# no HTTP client timeout bounds the total. Also the subprocess total for mega.
_CONNECT_TIMEOUT = 30.0
_READ_TIMEOUT = 60.0
_TOTAL_TIMEOUT = 1800.0 # 30 min per fetch
_DEFAULT_TIMEOUT = 600.0
_USER_AGENT = (
"Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/122.0.0.0 Safari/537.36"
@@ -136,11 +122,9 @@ def _filename_from(resp: requests.Response, url: str, fallback: str) -> str:
def _stream_to_file(resp: requests.Response, dest: Path,
should_stop: Callable[[], bool],
*, deadline: float | None = None) -> int:
should_stop: Callable[[], bool]) -> int:
"""Stream a response body to `dest` (atomic via .part). Returns byte count.
Honors should_stop and the total-budget `deadline` (a time.monotonic() value)
between chunks; the partial file is removed on either abort."""
Honors should_stop between chunks (partial file removed)."""
part = dest.with_name(dest.name + ".part")
total = 0
try:
@@ -148,8 +132,6 @@ def _stream_to_file(resp: requests.Response, dest: Path,
for chunk in resp.iter_content(chunk_size=_CHUNK):
if should_stop():
raise ExternalFetchError("stopped")
if deadline is not None and time.monotonic() > deadline:
raise ExternalFetchError("exceeded total fetch budget")
if chunk:
fh.write(chunk)
total += len(chunk)
@@ -160,29 +142,21 @@ def _stream_to_file(resp: requests.Response, dest: Path,
return total
def _get_to_dir(url: str, dest_dir: Path, *, read_timeout: float,
total_timeout: float, should_stop: Callable[[], bool],
fallback: str, headers: dict | None = None) -> FetchResult:
# (connect, read): a short read timeout fails a stalled socket fast; the total
# budget is enforced separately as a deadline across chunks (requests has no
# total-download timeout).
resp = _http_get(
url, timeout=(_CONNECT_TIMEOUT, read_timeout), headers=headers, stream=True
)
def _get_to_dir(url: str, dest_dir: Path, *, timeout: float,
should_stop: Callable[[], bool], fallback: str,
headers: dict | None = None) -> FetchResult:
resp = _http_get(url, timeout=timeout, headers=headers, stream=True)
if resp.status_code != 200:
return FetchResult(error=f"HTTP {resp.status_code} for {url}")
name = _filename_from(resp, url, fallback)
dest = dest_dir / name
written = _stream_to_file(
resp, dest, should_stop, deadline=time.monotonic() + total_timeout
)
written = _stream_to_file(resp, dest, should_stop)
return FetchResult(files=[dest], bytes=written)
# -- per-host fetchers -----------------------------------------------------
def _fetch_dropbox(url: str, dest_dir: Path, *, read_timeout: float,
total_timeout: float,
def _fetch_dropbox(url: str, dest_dir: Path, *, timeout: float,
should_stop: Callable[[], bool]) -> FetchResult:
# Force the direct-download variant: dl=1 (Dropbox serves an HTML preview
# for dl=0). Rewrite/insert the param rather than string-replace so ?dl=0,
@@ -191,44 +165,35 @@ def _fetch_dropbox(url: str, dest_dir: Path, *, read_timeout: float,
q = dict(parse_qsl(parts.query))
q["dl"] = "1"
direct = urlunsplit(parts._replace(query=urlencode(q)))
return _get_to_dir(direct, dest_dir, read_timeout=read_timeout,
total_timeout=total_timeout, should_stop=should_stop,
fallback="dropbox-file")
return _get_to_dir(direct, dest_dir, timeout=timeout,
should_stop=should_stop, fallback="dropbox-file")
def _fetch_pixeldrain(url: str, dest_dir: Path, *, read_timeout: float,
total_timeout: float,
def _fetch_pixeldrain(url: str, dest_dir: Path, *, timeout: float,
should_stop: Callable[[], bool]) -> FetchResult:
# /u/{id} (and /l/{id}) → the API file endpoint.
file_id = urlsplit(url).path.rstrip("/").split("/")[-1]
if not file_id:
return FetchResult(error=f"no pixeldrain id in {url}")
api = f"https://pixeldrain.com/api/file/{file_id}"
return _get_to_dir(api, dest_dir, read_timeout=read_timeout,
total_timeout=total_timeout, should_stop=should_stop,
fallback=f"{file_id}.bin")
return _get_to_dir(api, dest_dir, timeout=timeout,
should_stop=should_stop, fallback=f"{file_id}.bin")
def _fetch_mediafire(url: str, dest_dir: Path, *, read_timeout: float,
total_timeout: float,
def _fetch_mediafire(url: str, dest_dir: Path, *, timeout: float,
should_stop: Callable[[], bool]) -> FetchResult:
page = _http_get(url, timeout=(_CONNECT_TIMEOUT, read_timeout), stream=False)
page = _http_get(url, timeout=timeout, stream=False)
if page.status_code != 200:
return FetchResult(error=f"HTTP {page.status_code} for mediafire page")
m = _MEDIAFIRE_RE.search(page.text or "")
if not m:
return FetchResult(error="mediafire direct link not found on page")
return _get_to_dir(m.group(1), dest_dir, read_timeout=read_timeout,
total_timeout=total_timeout, should_stop=should_stop,
fallback="mediafire-file")
return _get_to_dir(m.group(1), dest_dir, timeout=timeout,
should_stop=should_stop, fallback="mediafire-file")
def _fetch_gdrive(url: str, dest_dir: Path, *, read_timeout: float,
total_timeout: float,
def _fetch_gdrive(url: str, dest_dir: Path, *, timeout: float,
should_stop: Callable[[], bool]) -> FetchResult:
# gdown manages its own HTTP session/timeouts; the task's celery hard limit is
# the outer backstop. read_timeout/total_timeout are accepted for a uniform
# registry signature but not separately enforceable here.
out = _gdown_download(url, str(dest_dir))
if not out:
return FetchResult(error="gdown returned no file (quota / private?)")
@@ -238,13 +203,10 @@ def _fetch_gdrive(url: str, dest_dir: Path, *, read_timeout: float,
return FetchResult(files=[p], bytes=p.stat().st_size)
def _fetch_mega(url: str, dest_dir: Path, *, read_timeout: float,
total_timeout: float,
def _fetch_mega(url: str, dest_dir: Path, *, timeout: float,
should_stop: Callable[[], bool]) -> FetchResult:
before = set(dest_dir.iterdir()) if dest_dir.exists() else set()
# megatools is a subprocess: its timeout IS a total wall-clock cap (the read
# timeout has no analogue here), so the total budget applies directly.
_run_mega_get(url, str(dest_dir), timeout=total_timeout)
_run_mega_get(url, str(dest_dir), timeout=timeout)
new = [p for p in dest_dir.iterdir() if p not in before and p.is_file()]
if not new:
return FetchResult(error="mega-get wrote no new file")
@@ -263,24 +225,18 @@ SUPPORTED_HOSTS = tuple(_REGISTRY)
def fetch_external(host: str, url: str, dest_dir: Path, *,
read_timeout: float = _READ_TIMEOUT,
total_timeout: float = _TOTAL_TIMEOUT,
timeout: float = _DEFAULT_TIMEOUT,
should_stop: Callable[[], bool] = lambda: False) -> FetchResult:
"""Fetch `url` (a `host` link) into `dest_dir`. Returns a FetchResult; never
raises — any backend error (transport, read/total timeout, non-200, scrape
miss, subprocess failure, stop) is captured on `.error` so the worker can
record it and move on.
`read_timeout` fails a stalled HTTP socket fast (idle gap between bytes);
`total_timeout` is the generous wall-clock cap for a large file that is
actively transferring (and the subprocess total for mega)."""
raises — any backend error (transport, non-200, scrape miss, subprocess
failure, stop) is captured on `.error` so the worker can record it and move
on."""
fetcher = _REGISTRY.get(host)
if fetcher is None:
return FetchResult(error=f"unsupported host {host!r}")
dest_dir.mkdir(parents=True, exist_ok=True)
try:
return fetcher(url, dest_dir, read_timeout=read_timeout,
total_timeout=total_timeout, should_stop=should_stop)
return fetcher(url, dest_dir, timeout=timeout, should_stop=should_stop)
except requests.RequestException as exc:
return FetchResult(error=f"transport error: {exc}")
except subprocess.TimeoutExpired:
+5 -2
View File
@@ -298,8 +298,11 @@ class GalleryDLService:
# removed at the plan-#697 cutover — it now uses the native ingester
# (services/patreon_ingester.py), not gallery-dl.
PLATFORM_DEFAULTS = {
# subscribestar removed — it's a native-ingester platform now (#71); the
# remaining entries are the gallery-dl platforms not yet migrated.
"subscribestar": {
"content_types": ["all"],
"directory": ["{date:%Y-%m-%d}_{id}_{title[:40]}"],
"filename": "{num:>02}_{filename}.{extension}",
},
"hentaifoundry": {
"content_types": ["all"],
"directory": [],
+18 -65
View File
@@ -25,13 +25,7 @@ from sqlalchemy.orm import aliased
from ..models import Artist, ImageProvenance, ImageRecord, Post, Source, Tag
from ..models.tag import image_tag
from .pagination import decode_cursor, encode_cursor
from .tag_query import (
fandom_join_alias,
image_in_any_tag_scope,
image_in_tag_scope,
serialize_tag,
tag_columns,
)
from .tag_query import fandom_join_alias, serialize_tag, tag_columns
# Reserved `platform` filter value selecting images with NO platformed
# provenance (filesystem imports). Returned by facets() as a null-valued
@@ -142,15 +136,11 @@ def image_url(path: str) -> str:
return f"/images/{quote(rel, safe='/')}"
def _require_single_filter(
tag_ids, post_id, artist_id, tag_or_groups=None, tag_exclude=None,
) -> None:
def _require_single_filter(tag_ids, post_id, artist_id) -> None:
"""post_id is the post-detail view — it can't combine with the
composable filters. tag_ids / tag_or_groups / tag_exclude + artist_id
(+ media_type) compose freely (AND)."""
if post_id is not None and (
tag_ids or artist_id is not None or tag_or_groups or tag_exclude
):
composable filters. tag_ids + artist_id (+ media_type) compose freely
(AND)."""
if post_id is not None and (tag_ids or artist_id is not None):
raise ValueError(
"post_id cannot be combined with tag or artist filters"
)
@@ -158,7 +148,6 @@ def _require_single_filter(
def _apply_scope(
stmt, *, tag_ids, post_id, artist_id, media_type,
tag_or_groups=None, tag_exclude=None,
platform=None, untagged=False, no_artist=False,
date_from=None, date_to=None,
):
@@ -169,14 +158,7 @@ def _apply_scope(
be present on `stmt` (the artist/platform paths alias Post/Source inside
their own EXISTS).
Tag filtering is one structured model (#6): AND-of-OR plus exclusions.
- tag_ids: image must carry ALL of them — one correlated EXISTS per tag
(the AND-of-singletons "include" common case; light editor + back-compat).
- tag_or_groups: list of OR-groups; the image must carry AT LEAST ONE tag
from EACH group — one EXISTS(tag_id IN group) per group, AND'd across
groups. (advanced editor)
- tag_exclude: image must carry NONE of these — a single NOT EXISTS(tag_id
IN exclude). (light "exclude" chips + advanced NOT)
- tag_ids: image must carry ALL of them — one correlated EXISTS per tag.
- post_id / artist_id: provenance EXISTS (post_id is exclusive, guarded
by _require_single_filter).
- media_type: 'image' | 'video' narrows by mime prefix.
@@ -186,17 +168,13 @@ def _apply_scope(
- no_artist: ImageRecord.artist_id IS NULL.
- date_from / date_to: half-open [from, to) bounds on effective_date.
"""
# Every tag clause goes through image_in_tag_scope/_any: a fandom tag also
# matches images carrying any of its characters (Tag.fandom_id). Include,
# OR-group, and exclude are all symmetric on that membership.
for tid in tag_ids or []:
stmt = stmt.where(image_in_tag_scope(tid))
for group in tag_or_groups or []:
if not group:
continue # an empty OR-group would match nothing; treat as absent
stmt = stmt.where(image_in_any_tag_scope(group))
if tag_exclude:
stmt = stmt.where(~image_in_any_tag_scope(tag_exclude))
stmt = stmt.where(
exists().where(
image_tag.c.image_record_id == ImageRecord.id,
image_tag.c.tag_id == tid,
)
)
prov = _provenance_clause(post_id, artist_id)
if prov is not None:
stmt = stmt.where(prov)
@@ -318,8 +296,6 @@ class GalleryService:
artist_id: int | None = None,
media_type: str | None = None,
sort: str = "newest",
tag_or_groups: list[list[int]] | None = None,
tag_exclude: list[int] | None = None,
platform: str | None = None,
untagged: bool = False,
no_artist: bool = False,
@@ -328,9 +304,7 @@ class GalleryService:
) -> GalleryPage:
if limit < 1 or limit > 200:
raise ValueError("limit must be between 1 and 200")
_require_single_filter(
tag_ids, post_id, artist_id, tag_or_groups, tag_exclude,
)
_require_single_filter(tag_ids, post_id, artist_id)
eff = _effective_date_col()
stmt = select(ImageRecord, Post.post_date, eff.label("eff"))
@@ -338,7 +312,6 @@ class GalleryService:
stmt = _apply_scope(
stmt, tag_ids=tag_ids, post_id=post_id,
artist_id=artist_id, media_type=media_type,
tag_or_groups=tag_or_groups, tag_exclude=tag_exclude,
platform=platform, untagged=untagged, no_artist=no_artist,
date_from=date_from, date_to=date_to,
)
@@ -386,8 +359,6 @@ class GalleryService:
post_id: int | None = None,
artist_id: int | None = None,
media_type: str | None = None,
tag_or_groups: list[list[int]] | None = None,
tag_exclude: list[int] | None = None,
platform: str | None = None,
untagged: bool = False,
no_artist: bool = False,
@@ -401,13 +372,10 @@ class GalleryService:
year_col, month_col, func.count(ImageRecord.id).label("cnt")
)
stmt = _outer_join_primary_post(stmt)
_require_single_filter(
tag_ids, post_id, artist_id, tag_or_groups, tag_exclude,
)
_require_single_filter(tag_ids, post_id, artist_id)
stmt = _apply_scope(
stmt, tag_ids=tag_ids, post_id=post_id,
artist_id=artist_id, media_type=media_type,
tag_or_groups=tag_or_groups, tag_exclude=tag_exclude,
platform=platform, untagged=untagged, no_artist=no_artist,
date_from=date_from, date_to=date_to,
)
@@ -419,8 +387,6 @@ class GalleryService:
self, year: int, month: int, tag_ids: list[int] | None = None,
post_id: int | None = None, artist_id: int | None = None,
media_type: str | None = None, sort: str = "newest",
tag_or_groups: list[list[int]] | None = None,
tag_exclude: list[int] | None = None,
platform: str | None = None, untagged: bool = False,
no_artist: bool = False, date_from: datetime | None = None,
date_to: datetime | None = None,
@@ -437,13 +403,10 @@ class GalleryService:
extract("month", eff) == month,
)
stmt = _outer_join_primary_post(stmt)
_require_single_filter(
tag_ids, post_id, artist_id, tag_or_groups, tag_exclude,
)
_require_single_filter(tag_ids, post_id, artist_id)
stmt = _apply_scope(
stmt, tag_ids=tag_ids, post_id=post_id,
artist_id=artist_id, media_type=media_type,
tag_or_groups=tag_or_groups, tag_exclude=tag_exclude,
platform=platform, untagged=untagged, no_artist=no_artist,
date_from=date_from, date_to=date_to,
)
@@ -464,10 +427,7 @@ class GalleryService:
async def facets(
self, *, tag_ids: list[int] | None = None,
post_id: int | None = None, artist_id: int | None = None,
media_type: str | None = None,
tag_or_groups: list[list[int]] | None = None,
tag_exclude: list[int] | None = None,
platform: str | None = None,
media_type: str | None = None, platform: str | None = None,
untagged: bool = False, no_artist: bool = False,
date_from: datetime | None = None, date_to: datetime | None = None,
) -> GalleryFacets:
@@ -477,13 +437,10 @@ class GalleryService:
No outer join is needed — every clause is a correlated EXISTS or a
column predicate on ImageRecord.
"""
_require_single_filter(
tag_ids, post_id, artist_id, tag_or_groups, tag_exclude,
)
_require_single_filter(tag_ids, post_id, artist_id)
common = {
"tag_ids": tag_ids, "post_id": post_id,
"artist_id": artist_id, "media_type": media_type,
"tag_or_groups": tag_or_groups, "tag_exclude": tag_exclude,
}
# total — the full active filter (the headline result count).
@@ -558,10 +515,7 @@ class GalleryService:
async def similar(
self, image_id: int, limit: int = 100, *,
tag_ids: list[int] | None = None, artist_id: int | None = None,
media_type: str | None = None,
tag_or_groups: list[list[int]] | None = None,
tag_exclude: list[int] | None = None,
platform: str | None = None,
media_type: str | None = None, platform: str | None = None,
untagged: bool = False, no_artist: bool = False,
date_from: datetime | None = None, date_to: datetime | None = None,
) -> list[GalleryImage] | None:
@@ -593,7 +547,6 @@ class GalleryService:
stmt = _apply_scope(
stmt, tag_ids=tag_ids, post_id=None,
artist_id=artist_id, media_type=media_type,
tag_or_groups=tag_or_groups, tag_exclude=tag_exclude,
platform=platform, untagged=untagged, no_artist=no_artist,
date_from=date_from, date_to=date_to,
)
+1 -58
View File
@@ -17,7 +17,7 @@ from enum import StrEnum
from pathlib import Path
from PIL import Image
from sqlalchemy import select, update
from sqlalchemy import select
from sqlalchemy.exc import IntegrityError
from sqlalchemy.orm import Session
@@ -506,12 +506,6 @@ class Importer:
artist_use = artist if artist is not None else self._resolve_artist(source)
post = self._post_for_sidecar(source, artist_use)
member_ids: list[int] = []
# Every member image touched (new + superseded + deduped), so the
# from_attachment_id stamp below covers files that already existed in the
# library and were merely re-linked to this post — those matter most
# (the HR copy a bundle re-ships). Separate from member_ids, which is
# the NEWLY-imported subset feeding the ImportResult contract.
member_record_ids: set[int] = set()
# Per-outcome tally so the "no images" reason names the ACTUAL cause
# (#718): nested-archive packs, all-deduped (benign), unsupported formats,
# or failed/corrupt members — instead of one catch-all string.
@@ -522,21 +516,11 @@ class Importer:
self._collect_archive_members(
source, attribution=source, source_row=source_row,
depth=0, member_ids=member_ids, counts=counts,
member_record_ids=member_record_ids,
)
# Preserve the archive itself (links to the same Post/Artist).
self._capture_attachment(
source, post=post, artist=artist_use, resolved=True
)
# Stamp each member's provenance row for THIS post with the archive it
# came out of (milestone #87). Done as a post-pass rather than threaded
# through _import_media/_apply_sidecar so the many dedup/supersede
# branches stay untouched. NULL-only so a re-extract never re-stamps and
# the backfill (reextract task → this same path) is idempotent. Nested
# members link to this OUTER archive — the only one stored as a blob.
self._stamp_member_archive(
post.id if post is not None else None, source, member_record_ids,
)
if member_ids:
return ImportResult(
status="imported", image_id=member_ids[0],
@@ -571,7 +555,6 @@ class Importer:
self, archive_path: Path, *, attribution: Path,
source_row: Source | None, depth: int,
member_ids: list[int], counts: dict,
member_record_ids: set[int],
) -> None:
"""Extract `archive_path` and import its image/video members, RECURSING
into nested archives (#718). Members attribute to `attribution` — the
@@ -607,7 +590,6 @@ class Importer:
member_path, attribution=attribution,
source_row=source_row, depth=depth + 1,
member_ids=member_ids, counts=counts,
member_record_ids=member_record_ids,
)
continue
counts["media"] += 1
@@ -619,16 +601,10 @@ class Importer:
)
if res.status in ("imported", "superseded") and res.image_id:
member_ids.append(res.image_id)
member_record_ids.add(res.image_id)
elif res.status == "skipped" and res.skip_reason in (
SkipReason.duplicate_hash, SkipReason.duplicate_phash
):
counts["deduped"] += 1
# A deduped member still links provenance to this post
# (enrich-on-duplicate); record it so its archive origin
# gets stamped too.
if res.image_id:
member_record_ids.add(res.image_id)
else:
counts["failed"] += 1
except Exception as exc: # noqa: BLE001 — defensive per level; keep going
@@ -637,39 +613,6 @@ class Importer:
archive_path.name, depth, exc,
)
def _stamp_member_archive(
self, post_id: int | None, archive_source: Path, member_record_ids: set[int],
) -> None:
"""Record which archive each extracted member came from (milestone #87).
Resolves the archive's own PostAttachment (by post + sha — it was just
captured) and stamps from_attachment_id on every member's provenance row
FOR THIS POST. NULL-only, so re-extracting the same archive (the backfill
path) never overwrites and stays idempotent. No-op when the archive isn't
post-attached (filesystem import with no post) or yielded no members.
"""
if post_id is None or not member_record_ids:
return
sha = _sha256_of(archive_source)
att_id = self.session.execute(
select(PostAttachment.id).where(
PostAttachment.post_id == post_id,
PostAttachment.sha256 == sha,
)
).scalar_one_or_none()
if att_id is None:
return
self.session.execute(
update(ImageProvenance)
.where(
ImageProvenance.image_record_id.in_(member_record_ids),
ImageProvenance.post_id == post_id,
ImageProvenance.from_attachment_id.is_(None),
)
.values(from_attachment_id=att_id)
)
self.session.commit()
@staticmethod
def _video_aspect_matches(w, h, cw, ch) -> bool:
"""True when two (w,h) pairs share an aspect ratio within tolerance.
+6 -65
View File
@@ -36,7 +36,6 @@ from sqlalchemy import delete, func, select, text
from sqlalchemy.dialects.postgresql import insert as pg_insert
from .gallery_dl import DownloadResult, ErrorType, make_run_stats
from .native_ingest_common import NativeAuthError, NativeDriftError
log = logging.getLogger(__name__)
@@ -89,7 +88,6 @@ class Ingester:
ledger_key: Callable[[object], str],
platform: str,
error_base: type[Exception],
drift_label: str | None = None,
):
self.client = client
self.downloader = downloader
@@ -101,10 +99,6 @@ class Ingester:
self._ledger_key = ledger_key
self._platform = platform
self._error_base = error_base
# Human label for the API_DRIFT message ("<label> changed — ingester needs
# update"). Defaults to the platform name; adapters pass a richer phrase
# (e.g. "Patreon API", "SubscribeStar markup").
self._drift_label = drift_label or platform
# -- public ------------------------------------------------------------
@@ -240,15 +234,6 @@ class Ingester:
),
)
# #899 L1: emit run milestones through the real logger (not only the
# in-memory log_lines → DownloadResult.stdout, which is persisted to the
# DownloadEvent ONLY at phase 3). A worker SIGKILL/OOM/hard-time-limit
# mid-walk would otherwise leave NO trace; these land in the container log
# in real time regardless of whether the event gets finalized.
log.info(
"%s ingest START (%s): source=%s campaign=%s resume_cursor=%s",
self._platform, mode, source_id, campaign_id, resume_cursor,
)
try:
for post, included, page_cursor in self.client.iter_posts(
campaign_id, cursor=resume_cursor
@@ -258,15 +243,6 @@ class Ingester:
# after it. Carried as DownloadResult.cursor (plan #704).
if page_cursor and page_cursor != emitted_cursor:
emitted_cursor = page_cursor
# #899 L1: a per-page breadcrumb in the container log (pages can
# be minutes apart on image-dense backfills) — survives a worker
# kill so the operator sees how far a since-died walk got.
log.info(
"%s ingest progress (%s, source=%s): posts=%d downloaded=%d "
"skipped=%d errors=%d quarantined=%d gated=%d cursor=%s",
self._platform, mode, source_id, posts_processed, downloaded,
skipped_count, errors, quarantined, gated_skipped, emitted_cursor,
)
# plan #705 #6: persist the cursor at each page boundary so a
# worker SIGKILL mid-chunk resumes near the crash, not the
# chunk start. (phase 3 still writes the final cursor — same
@@ -468,10 +444,6 @@ class Ingester:
# posts), so don't re-write them — return PARTIAL (reads as "ok/progress",
# the lifecycle no-ops since state is gone) instead of a false "complete".
if stopped:
log.info(
"%s ingest STOPPED by operator (%s, source=%s): %d file(s) this chunk",
self._platform, mode, source_id, downloaded,
)
return _result(
success=False, return_code=-1,
error_type=ErrorType.PARTIAL,
@@ -489,7 +461,7 @@ class Ingester:
log_lines.append(f"{quarantined} media item(s) quarantined (invalid)")
if dead_lettered:
log_lines.append(f"{dead_lettered} media item(s) skipped (dead-lettered)")
summary = (
log_lines.append(
f"{self._platform} ingest ({mode}): {downloaded} downloaded, "
f"{skipped_count} skipped, {quarantined} quarantined, "
f"{dead_lettered} dead-lettered, {errors} error(s), "
@@ -503,9 +475,6 @@ class Ingester:
+ (", reached end" if reached_bottom else "")
+ (", time-boxed" if budget_hit else "")
)
log_lines.append(summary)
# #899 L1: also to the container log (survives event-finalization failure).
log.info("%s (source=%s)", summary, source_id)
if budget_hit:
# A chunk that hit its time-box but made forward progress is a
@@ -632,41 +601,13 @@ class Ingester:
# -- failure mapping (adapter overrides) -------------------------------
def _failure_result(self, exc: Exception, _result) -> DownloadResult:
"""Map a platform client-error to a loud, typed failed DownloadResult
NEVER a silent zero-download "success". The mapping is shared across
platforms via the NativeAuthError/NativeDriftError taxonomy (the platform
client raises subclasses), so a new platform gets it for free:
- NativeAuthError → AUTH_ERROR (rotate the credential)
- NativeDriftError → API_DRIFT (the ingester/scraper needs updating)
- HTTP 429 / 404 → RATE_LIMITED / NOT_FOUND
- other HTTP status→ HTTP_ERROR; transport failure → NETWORK_ERROR
Auth/Drift are matched first (they also carry a status_code in some paths).
"""
message = str(exc)
if isinstance(exc, NativeAuthError):
error_type = ErrorType.AUTH_ERROR
elif isinstance(exc, NativeDriftError):
error_type = ErrorType.API_DRIFT
message = f"{self._drift_label} changed — ingester needs update: {message}"
else:
status = getattr(exc, "status_code", None)
if status == 429:
error_type = ErrorType.RATE_LIMITED
elif status == 404:
error_type = ErrorType.NOT_FOUND
elif status is not None:
error_type = ErrorType.HTTP_ERROR
else:
error_type = ErrorType.NETWORK_ERROR
log.warning("%s ingest failed (%s): %s", self._platform, error_type.value, message)
result = _result(
"""Map a platform client-error to a typed failed DownloadResult. The base
gives a safe default; adapters override with their exception taxonomy."""
log.warning("%s ingest failed: %s", self._platform, exc)
return _result(
success=False, return_code=1,
error_type=error_type, error_message=message,
error_type=ErrorType.UNKNOWN_ERROR, error_message=str(exc),
)
# plan #708 B1: carry the server's Retry-After up to the cooldown.
if error_type == ErrorType.RATE_LIMITED:
result.retry_after_seconds = getattr(exc, "retry_after", None)
return result
# -- seen-ledger (short-lived sessions) --------------------------------
+10 -84
View File
@@ -5,18 +5,11 @@ image_tag AND to tag_allowlist; per-image removal/dismiss writes a rejection.
from collections.abc import Sequence
from dataclasses import dataclass
from sqlalchemy import and_, delete, distinct, func, or_, select
from sqlalchemy import delete, select
from sqlalchemy.dialects.postgresql import insert
from sqlalchemy.ext.asyncio import AsyncSession
from ...models import (
ImagePrediction,
MLSettings,
Tag,
TagAlias,
TagAllowlist,
TagSuggestionRejection,
)
from ...models import MLSettings, Tag, TagAllowlist, TagSuggestionRejection
from ...models.tag import image_tag
from .aliases import AliasService
@@ -27,8 +20,6 @@ class AllowlistRow:
tag_name: str
tag_kind: str
min_confidence: float
applied_count: int # image_tag rows currently carrying this tag
coverage_count: int # images a sweep WOULD cover at min_confidence
class AllowlistService:
@@ -89,12 +80,6 @@ class AllowlistService:
)
await self.session.execute(stmt)
async def undismiss(self, image_id: int, tag_id: int) -> None:
"""Undo a per-image dismissal — drop the TagSuggestionRejection so the
suggestion reverts to a live (un-rejected) state. Backs the rail's
one-click reject-recovery (operator-asked 2026-06-27)."""
await self._clear_rejection(image_id, tag_id)
async def reject_applied_tag(self, image_id: int, tag_id: int) -> None:
"""Operator removed an applied tag from an image. Remove the
image_tag row AND record a rejection so the allowlist won't
@@ -131,44 +116,6 @@ class AllowlistService:
delete(TagAllowlist).where(TagAllowlist.tag_id == tag_id)
)
async def _coverage_match(self, tag: Tag):
"""The predicate over image_prediction rows that resolve to `tag`,
mirroring tasks.ml._confidence_for_tag's resolution: a prediction whose
raw_name equals the tag name (any category), OR an alias maps
(raw_name, category) -> this tag. Returns a SQLAlchemy boolean clause.
"""
alias_rows = (
await self.session.execute(
select(TagAlias.alias_string, TagAlias.alias_category).where(
TagAlias.canonical_tag_id == tag.id
)
)
).all()
name_clause = ImagePrediction.raw_name == tag.name
alias_clauses = [
and_(
ImagePrediction.raw_name == a,
ImagePrediction.category == c,
)
for a, c in alias_rows
]
return or_(name_clause, *alias_clauses) if alias_clauses else name_clause
async def coverage(self, tag_id: int, threshold: float) -> int:
"""How many distinct images a sweep WOULD cover for this tag at
`threshold`: images with a resolving prediction scoring >= threshold.
The gross candidate pool (NOT minus already-applied/rejected) — it's
the tuning signal for "lower the threshold and ~N more images qualify".
"""
tag = await self.session.get(Tag, tag_id)
if tag is None:
return 0
match = await self._coverage_match(tag)
stmt = select(
func.count(distinct(ImagePrediction.image_record_id))
).where(ImagePrediction.score >= threshold, match)
return (await self.session.execute(stmt)).scalar_one()
async def list_all(self) -> Sequence[AllowlistRow]:
stmt = (
select(
@@ -181,33 +128,12 @@ class AllowlistService:
.order_by(Tag.name.asc())
)
rows = (await self.session.execute(stmt)).all()
tag_ids = [r[0] for r in rows]
# Applied counts in ONE grouped query (vs N per-row counts).
applied: dict[int, int] = {}
if tag_ids:
applied = dict(
(
await self.session.execute(
select(image_tag.c.tag_id, func.count())
.where(image_tag.c.tag_id.in_(tag_ids))
.group_by(image_tag.c.tag_id)
)
).all()
return [
AllowlistRow(
tag_id=r[0],
tag_name=r[1],
tag_kind=r[2].value if hasattr(r[2], "value") else str(r[2]),
min_confidence=r[3],
)
result = []
for r in rows:
# Coverage is per-tag (alias set differs); allowlist is small.
cov = await self.coverage(r[0], r[3])
result.append(
AllowlistRow(
tag_id=r[0],
tag_name=r[1],
tag_kind=r[2].value if hasattr(r[2], "value") else str(r[2]),
min_confidence=r[3],
applied_count=applied.get(r[0], 0),
coverage_count=cov,
)
)
return result
for r in rows
]
-180
View File
@@ -1,180 +0,0 @@
"""CCIP few-shot character matcher (#114) — server-side, numpy on stored vectors.
CCIP is a FROZEN identity embedding; we don't train it. Instead the operator's
tagged characters become reference prototypes: a character tag's references are
the CCIP vectors of figure/face regions on images carrying that tag. To suggest
characters for a new image, we compare its figure-region CCIP vectors to every
character's references (multi-prototype: best match over a character's examples)
and surface the ones that clear a similarity threshold. No GPU here — the agent
already produced the vectors; this is cosine matching on what's stored.
v1 uses cosine similarity on the raw CCIP vectors with a tunable threshold; the
exact CCIP difference metric/threshold gets validated against the model during
the hands-on eval. numpy is imported lazily (API worker has it via pgvector).
"""
from sqlalchemy import func, select
from sqlalchemy.ext.asyncio import AsyncSession
from ...models import ImageRegion, MLSettings, Tag, TagKind
from ...models.tag import image_tag
# Cosine-similarity floor to call a figure the same character. The live setting
# (ml_settings.ccip_match_threshold) drives it; this is only the fallback when no
# threshold is supplied AND no settings row exists.
DEFAULT_SIM_THRESHOLD = 0.85
_FIGURE_KINDS = ("face", "figure")
async def _settings_threshold(session: AsyncSession) -> float:
val = (
await session.execute(
select(MLSettings.ccip_match_threshold).where(MLSettings.id == 1)
)
).scalar_one_or_none()
return float(val) if val is not None else DEFAULT_SIM_THRESHOLD
def _l2norm(mat, np):
n = np.linalg.norm(mat, axis=1, keepdims=True)
n[n == 0] = 1.0
return mat / n
# Single-shot cache of the (expensive) reference load, keyed on a cheap
# signature that changes exactly when references could: a character tag added/
# removed (n_char_tags) or a figure embedded (max/ n of ccip regions). Shared by
# the live matcher (every modal open) and the auto-apply sweep.
_REF_CACHE: dict = {"sig": None, "refs": None}
def _single_character_images():
"""Subquery of image ids carrying EXACTLY ONE character tag. References come
only from these — on a multi-character image the tag is image-level, so every
figure would otherwise pollute each character's prototype set (a 2-character
image tagged 'Velma' would make Daphne's figure a Velma reference)."""
return (
select(image_tag.c.image_record_id)
.join(Tag, Tag.id == image_tag.c.tag_id)
.where(Tag.kind == TagKind.character)
.group_by(image_tag.c.image_record_id)
.having(func.count() == 1)
)
async def _ref_signature(session: AsyncSession) -> tuple:
n_tags = (
await session.execute(
select(func.count())
.select_from(image_tag)
.join(Tag, Tag.id == image_tag.c.tag_id)
.where(Tag.kind == TagKind.character)
)
).scalar_one()
n_regs, max_id = (
await session.execute(
select(func.count(), func.max(ImageRegion.id)).where(
ImageRegion.kind.in_(_FIGURE_KINDS),
ImageRegion.ccip_embedding.is_not(None),
)
)
).one()
return (n_tags, n_regs, max_id)
async def character_references(session: AsyncSession) -> dict[int, list]:
"""Per character-tag CCIP reference vectors: figure/face-region CCIP
embeddings on UNAMBIGUOUS (single-character) images carrying that tag.
Multi-prototype — several vectors per character. Cached on a cheap signature."""
sig = await _ref_signature(session)
if _REF_CACHE["sig"] == sig and _REF_CACHE["refs"] is not None:
return _REF_CACHE["refs"]
rows = (
await session.execute(
select(image_tag.c.tag_id, ImageRegion.ccip_embedding)
.select_from(ImageRegion)
.join(
image_tag,
image_tag.c.image_record_id == ImageRegion.image_record_id,
)
.join(Tag, Tag.id == image_tag.c.tag_id)
.where(Tag.kind == TagKind.character)
.where(ImageRegion.kind.in_(_FIGURE_KINDS))
.where(ImageRegion.ccip_embedding.is_not(None))
.where(ImageRegion.image_record_id.in_(_single_character_images()))
)
).all()
refs: dict[int, list] = {}
for tag_id, vec in rows:
refs.setdefault(tag_id, []).append(vec)
_REF_CACHE.update(sig=sig, refs=refs)
return refs
async def _tag_names(session: AsyncSession, tag_ids: list[int]) -> dict[int, str]:
if not tag_ids:
return {}
return dict(
(
await session.execute(
select(Tag.id, Tag.name).where(Tag.id.in_(tag_ids))
)
).all()
)
async def match_image(
session: AsyncSession, image_id: int, threshold: float | None = None
) -> list[dict]:
"""Character suggestions for one image from its figure-region CCIP vectors:
[{tag_id, name, category:'character', score, source:'ccip'}], ranked.
Already-applied character tags are excluded. Empty if the image has no figure
CCIP vectors or no character references exist yet. threshold defaults to the
live ml_settings.ccip_match_threshold."""
import numpy as np
if threshold is None:
threshold = await _settings_threshold(session)
qvecs = (
await session.execute(
select(ImageRegion.ccip_embedding).where(
ImageRegion.image_record_id == image_id,
ImageRegion.kind.in_(_FIGURE_KINDS),
ImageRegion.ccip_embedding.is_not(None),
)
)
).scalars().all()
if not qvecs:
return []
refs = await character_references(session)
if not refs:
return []
applied = set(
(
await session.execute(
select(image_tag.c.tag_id).where(
image_tag.c.image_record_id == image_id
)
)
).scalars()
)
names = await _tag_names(session, [t for t in refs if t not in applied])
Q = _l2norm(np.vstack([np.asarray(v, dtype=np.float32) for v in qvecs]), np)
out = []
for tag_id, vecs in refs.items():
if tag_id in applied:
continue
R = _l2norm(np.vstack([np.asarray(v, dtype=np.float32) for v in vecs]), np)
best = float((Q @ R.T).max()) # best (query figure, reference) cosine
if best >= threshold:
out.append({
"tag_id": tag_id,
"name": names.get(tag_id, str(tag_id)),
"category": "character",
"score": round(best, 4),
"source": "ccip",
})
out.sort(key=lambda d: d["score"], reverse=True)
return out
-73
View File
@@ -1,73 +0,0 @@
"""Shared crop primitive for the region/crop pipeline (#114).
One model- and transport-agnostic function sits at the trunk of both crop jobs:
- CCIP characters: a face/figure detector proposes regions → crop → CCIP-embed.
- SigLIP concepts: head-guided / saliency proposes regions → crop → SigLIP-embed.
Only the PROPOSER (where to crop) and the EMBEDDER (what to run) differ; the crop
itself — including the lower-bound size floor below which a region is too small to
embed reliably — is identical, so it lives here and both jobs call it.
The actual detector + embedders run in the GPU agent; this is pure Pillow so it's
importable + testable anywhere (and the agent imports it for the crop step).
"""
from __future__ import annotations
from PIL import Image
# Size floor: a region must be at least this big on its SHORTER edge to be worth
# embedding — a smaller crop is a blurry upscale carrying little real signal, and
# unbounded tiny crops would explode the bag. Expressed as BOTH a fraction of the
# image's short side and an absolute pixel floor; the larger of the two wins.
MIN_CROP_FRACTION = 0.10
MIN_CROP_PX = 64
def _to_pixels(bbox: tuple[float, float, float, float], w: int, h: int):
"""Normalized (x, y, w, h) in [0,1] → pixel (x, y, w, h)."""
x, y, bw, bh = bbox
return x * w, y * h, bw * w, bh * h
def crop_region(
img: Image.Image,
bbox: tuple[float, float, float, float],
*,
pad: float = 0.0,
min_fraction: float = MIN_CROP_FRACTION,
min_px: int = MIN_CROP_PX,
out_size: int | None = None,
) -> Image.Image | None:
"""Crop a NORMALIZED bbox (x, y, w, h in [0,1]) from img.
- pad: grow the box by this fraction on each side (e.g. 0.15 = +15% context),
clamped to the image bounds.
- Returns None when the resulting region is below the size floor (too small to
embed reliably) — the caller skips embedding it.
- out_size: if given, resize the crop to out_size×out_size; otherwise return
the raw crop and let the embedder do its own preprocessing.
"""
iw, ih = img.size
px, py, pw, ph = _to_pixels(bbox, iw, ih)
if pad:
px -= pw * pad / 2.0
py -= ph * pad / 2.0
pw *= (1.0 + pad)
ph *= (1.0 + pad)
left = max(0, int(round(px)))
top = max(0, int(round(py)))
right = min(iw, int(round(px + pw)))
bottom = min(ih, int(round(py + ph)))
if right <= left or bottom <= top:
return None
floor = max(min_px, int(min_fraction * min(iw, ih)))
if min(right - left, bottom - top) < floor:
return None
crop = img.crop((left, top, right, bottom)).convert("RGB")
if out_size:
crop = crop.resize((out_size, out_size))
return crop
-177
View File
@@ -1,177 +0,0 @@
"""GPU-job queue engine (#114): enqueue / lease / heartbeat / complete / fail
/ release / recover_orphaned.
Backs the HTTP API the desktop agent pulls work from. The lease claims pending
OR expired-leased jobs with FOR UPDATE SKIP LOCKED, so concurrent agents/workers
never grab the same job. Orphan recovery is three-layered: a graceful agent stop
calls release() to hand its in-flight jobs back instantly; a hard crash is caught
by recover_orphaned() (a 60s beat sweep) which resets expired leases to pending;
and the lease itself reclaims expired leases as a final backstop. Result-writing
(regions) is done by the API handler via RegionService; complete() just closes.
"""
from datetime import UTC, datetime, timedelta
from sqlalchemy import and_, or_, select, update
from sqlalchemy.ext.asyncio import AsyncSession
from ...models import GpuJob
# Lease window. Kept comfortably above any single job (a capped-frame video embed
# is tens of seconds) so a live, heartbeating worker is never falsely expired,
# but short enough that a hard crash recovers fast once the sweep fires.
DEFAULT_LEASE_TTL = 180 # seconds an agent holds a job before it can be re-leased
DEFAULT_BATCH = 8
MAX_ATTEMPTS = 3
class GpuJobService:
def __init__(self, session: AsyncSession):
self.session = session
async def enqueue(self, image_id: int, task: str) -> GpuJob | None:
"""Queue a (image, task) job. Idempotent: returns None if one is already
pending/leased for the same pair (no duplicate work)."""
dup = (
await self.session.execute(
select(GpuJob.id).where(
GpuJob.image_record_id == image_id,
GpuJob.task == task,
GpuJob.status.in_(["pending", "leased"]),
)
)
).first()
if dup:
return None
job = GpuJob(image_record_id=image_id, task=task, status="pending")
self.session.add(job)
await self.session.flush()
return job
async def lease(
self, token: str, batch_size: int = DEFAULT_BATCH, ttl: int = DEFAULT_LEASE_TTL
) -> list[GpuJob]:
"""Claim up to batch_size pending (or expired-leased) jobs for `token`."""
now = datetime.now(UTC)
picked = (
await self.session.execute(
select(GpuJob.id)
.where(
or_(
GpuJob.status == "pending",
and_(
GpuJob.status == "leased",
GpuJob.lease_expires_at < now,
),
)
)
.order_by(GpuJob.id)
.limit(batch_size)
.with_for_update(skip_locked=True)
)
).scalars().all()
if not picked:
return []
await self.session.execute(
update(GpuJob)
.where(GpuJob.id.in_(picked))
.values(
status="leased", lease_token=token, leased_at=now,
lease_expires_at=now + timedelta(seconds=ttl),
attempts=GpuJob.attempts + 1, updated_at=now,
)
)
# populate_existing: overwrite identity-map copies with the post-UPDATE
# values so the returned jobs reflect the new lease/attempts, not stale
# pre-lease state.
return list(
(
await self.session.execute(
select(GpuJob)
.where(GpuJob.id.in_(picked))
.order_by(GpuJob.id)
.execution_options(populate_existing=True)
)
).scalars()
)
async def heartbeat(
self, token: str, job_ids: list[int], ttl: int = DEFAULT_LEASE_TTL
) -> int:
"""Extend the lease on the agent's in-flight jobs. Returns rows touched."""
now = datetime.now(UTC)
res = await self.session.execute(
update(GpuJob)
.where(
GpuJob.id.in_(job_ids),
GpuJob.lease_token == token,
GpuJob.status == "leased",
)
.values(lease_expires_at=now + timedelta(seconds=ttl), updated_at=now)
)
return res.rowcount or 0
async def complete(self, token: str, job_id: int) -> bool:
"""Close a leased job (after its results were stored). False if the job
isn't leased by this token (a stale/expired submit)."""
job = await self.session.get(GpuJob, job_id)
if job is None or job.status != "leased" or job.lease_token != token:
return False
job.status = "done"
job.lease_token = None
job.lease_expires_at = None
job.error = None
job.updated_at = datetime.now(UTC)
return True
async def fail(self, token: str, job_id: int, error: str) -> bool:
"""Report a failure: re-queue (pending) until MAX_ATTEMPTS, then 'error'."""
job = await self.session.get(GpuJob, job_id)
if job is None or job.lease_token != token:
return False
if job.attempts >= MAX_ATTEMPTS:
job.status = "error"
else:
job.status = "pending"
job.lease_token = None
job.lease_expires_at = None
job.error = (error or "")[:1000]
job.updated_at = datetime.now(UTC)
return True
async def release(self, token: str, job_ids: list[int]) -> int:
"""Hand the agent's still-leased jobs back to pending NOW (graceful stop),
so another worker picks them up immediately instead of waiting out the
lease. Scoped to the token's own leases. Returns rows released."""
if not job_ids:
return 0
now = datetime.now(UTC)
res = await self.session.execute(
update(GpuJob)
.where(
GpuJob.id.in_(job_ids),
GpuJob.lease_token == token,
GpuJob.status == "leased",
)
.values(
status="pending", lease_token=None, leased_at=None,
lease_expires_at=None, updated_at=now,
)
)
return res.rowcount or 0
async def recover_orphaned(self) -> int:
"""Reset every expired lease back to pending — catches agents that died
mid-job (no graceful release). Run on a short beat so the queue recovers
+ reads honestly even when no worker is actively leasing. Returns rows
recovered."""
now = datetime.now(UTC)
res = await self.session.execute(
update(GpuJob)
.where(GpuJob.status == "leased", GpuJob.lease_expires_at < now)
.values(
status="pending", lease_token=None, leased_at=None,
lease_expires_at=None, updated_at=now,
)
)
return res.rowcount or 0
-490
View File
@@ -1,490 +0,0 @@
"""Production heads: train + score the per-concept classifiers (#114).
The eval (#1130, tag_eval.py) proved the spine; this is its production form.
- TRAIN (sync, ml worker — needs scikit-learn): for every general/character tag
with enough labelled positives, fit a logistic-regression head on the FROZEN
SigLIP embeddings (positives + negatives = rejections + sampled unlabeled),
derive an honest suggest threshold + earned-auto-apply point from CROSS-
VALIDATED scores, and upsert a TagHead row. Reuses tag_eval's proven data
loaders + metric helpers so production heads match the eval's measured numbers.
- SCORE (async, API worker — numpy via pgvector, NO scikit-learn): score one
image's embedding against all current heads → the suggestions the rail shows,
REPLACING Camie predictions + per-tag centroids.
scikit-learn is imported lazily inside the train path so the API worker can still
import this module to enqueue training + to score (scoring needs only numpy).
"""
from __future__ import annotations
import logging
from datetime import UTC, datetime
from typing import Any
from sqlalchemy import delete, func, select
from sqlalchemy.ext.asyncio import AsyncSession
from sqlalchemy.orm import Session
from ...models import (
HeadAutoApplyRun,
HeadTrainingRun,
ImageRecord,
ImageRegion,
MLSettings,
Tag,
TagHead,
TagKind,
TagSuggestionRejection,
)
from ...models.tag import image_tag
from .tag_eval import (
_auto_apply_point,
_ids_with_tag,
_l2norm,
_load_embeddings,
_metrics_from_scores,
_rejected_ids,
_safe_folds,
_sample_unlabeled,
)
log = logging.getLogger(__name__)
DEFAULT_NEG_RATIO = 3
DEFAULT_CV_FOLDS = 5
MIN_POSITIVES_FLOOR = 8 # hard floor; settings.head_min_positives can raise it
_UNLABELED_POOL = 4000
_EXAMPLES_MIN = 8 # need at least this many embedded +/- to fit a head
# Only these tag kinds get heads (the surfaced suggestion categories).
_HEAD_KINDS = (TagKind.general, TagKind.character)
# tag.kind -> the suggestion category the rail groups under.
_CATEGORY = {TagKind.general: "general", TagKind.character: "character"}
class HeadTrainingAlreadyRunning(Exception):
"""Raised by start_head_training_run when a run is already in flight."""
def start_head_training_run(session: Session, params: dict[str, Any]) -> int:
"""Create a HeadTrainingRun (status='running') + dispatch the ml-queue task.
Returns the run id. One training run at a time (light guard)."""
existing = session.execute(
select(HeadTrainingRun.id).where(HeadTrainingRun.status == "running")
).scalar_one_or_none()
if existing is not None:
raise HeadTrainingAlreadyRunning(existing)
norm = _normalize_params(session, params)
run = HeadTrainingRun(
params=norm, status="running", last_progress_at=datetime.now(UTC)
)
session.add(run)
session.flush()
run_id = run.id
from ...tasks.ml import train_heads as _task
_task.delay(run_id)
return run_id
def _settings(session: Session) -> MLSettings:
return session.execute(
select(MLSettings).where(MLSettings.id == 1)
).scalar_one()
def _normalize_params(session: Session, params: dict[str, Any] | None) -> dict[str, Any]:
params = params or {}
s = _settings(session)
try:
min_pos = max(MIN_POSITIVES_FLOOR, int(params.get("min_positives", s.head_min_positives)))
except (TypeError, ValueError):
min_pos = max(MIN_POSITIVES_FLOOR, s.head_min_positives)
try:
neg_ratio = max(1, int(params.get("neg_ratio", DEFAULT_NEG_RATIO)))
except (TypeError, ValueError):
neg_ratio = DEFAULT_NEG_RATIO
try:
cv_folds = max(2, int(params.get("cv_folds", DEFAULT_CV_FOLDS)))
except (TypeError, ValueError):
cv_folds = DEFAULT_CV_FOLDS
try:
precision_target = min(max(float(params.get("precision_target", s.head_auto_apply_precision)), 0.5), 0.999)
except (TypeError, ValueError):
precision_target = s.head_auto_apply_precision
return {
"min_positives": min_pos,
"neg_ratio": neg_ratio,
"cv_folds": cv_folds,
"precision_target": round(precision_target, 4),
}
def _embedder_version(session: Session) -> str:
return _settings(session).embedder_model_version
def _eligible_tag_ids(session: Session, min_pos: int) -> list[int]:
"""Concept tags (general/character) with >= min_pos labelled images — the
set that gets a head. Counts all sources; source-aware filtering (#1133) is
a separate, optional refinement."""
rows = session.execute(
select(Tag.id)
.join(image_tag, image_tag.c.tag_id == Tag.id)
.where(Tag.kind.in_(_HEAD_KINDS))
.group_by(Tag.id)
.having(func.count(image_tag.c.image_record_id) >= min_pos)
).all()
return [r[0] for r in rows]
def train_all_heads(
session: Session, params: dict[str, Any], run: HeadTrainingRun | None = None
) -> dict[str, int]:
"""(Re)train a head for every eligible concept; prune heads whose tag is no
longer eligible. Commits per head so a SIGKILL leaves trained heads durable
(training is idempotent). Returns {n_trained, n_skipped}."""
import numpy as np
cfg = _normalize_params(session, params)
embedding_version = _embedder_version(session)
eligible = _eligible_tag_ids(session, cfg["min_positives"])
eligible_set = set(eligible)
trained = 0
skipped = 0
for i, tag_id in enumerate(eligible):
try:
ok = train_head(session, tag_id, embedding_version, cfg, np)
except Exception:
log.exception("train_head failed for tag %d", tag_id)
ok = False
session.commit()
trained += int(ok)
skipped += int(not ok)
if run is not None and i % 10 == 0:
run.last_progress_at = datetime.now(UTC)
session.commit()
# Retire heads whose concept dropped out of the eligible set (lost its
# positives, or the tag was re-kinded) so stale heads can't keep suggesting.
if eligible_set:
session.execute(delete(TagHead).where(TagHead.tag_id.not_in(eligible_set)))
else:
session.execute(delete(TagHead))
session.commit()
return {"n_trained": trained, "n_skipped": skipped}
def train_head(
session: Session, tag_id: int, embedding_version: str, cfg: dict, np
) -> bool:
"""Fit + upsert one head. Returns True if a head was written, False if the
concept had too few usable examples to train (the row is then removed)."""
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import StratifiedKFold, cross_val_predict
pos_ids = _ids_with_tag(session, tag_id)
if len(pos_ids) < cfg["min_positives"]:
session.execute(delete(TagHead).where(TagHead.tag_id == tag_id))
return False
pos_set = set(pos_ids)
rejected = [i for i in _rejected_ids(session, tag_id) if i not in pos_set]
want_neg = max(len(pos_ids) * cfg["neg_ratio"], _EXAMPLES_MIN * 4)
sampled = _sample_unlabeled(
session, pos_set | set(rejected), min(_UNLABELED_POOL, want_neg)
)
neg_ids = rejected + [i for i in sampled if i not in pos_set]
emb = _load_embeddings(session, pos_ids + neg_ids)
pos = [emb[i] for i in pos_ids if i in emb]
neg = [emb[i] for i in neg_ids if i in emb]
if len(pos) < _EXAMPLES_MIN or len(neg) < _EXAMPLES_MIN:
session.execute(delete(TagHead).where(TagHead.tag_id == tag_id))
return False
X = np.vstack(pos + neg).astype(np.float32)
y = np.array([1] * len(pos) + [0] * len(neg))
Xn = _l2norm(X, np)
clf = LogisticRegression(max_iter=1000, class_weight="balanced")
cv = StratifiedKFold(
n_splits=_safe_folds(y, cfg["cv_folds"], np), shuffle=True, random_state=0
)
# Honest thresholds from out-of-fold scores; deployable weights from a final
# fit on ALL the data.
cv_probs = cross_val_predict(clf, Xn, y, cv=cv, method="predict_proba")[:, 1]
metrics = _metrics_from_scores(y, cv_probs, np)
auto = _auto_apply_point(y, cv_probs, cfg["precision_target"], np)
clf.fit(Xn, y)
head = session.get(TagHead, tag_id)
if head is None:
head = TagHead(tag_id=tag_id)
session.add(head)
head.embedding_version = embedding_version
head.weights = clf.coef_[0].astype(np.float32).tolist()
head.bias = float(clf.intercept_[0])
head.suggest_threshold = float(metrics["threshold"])
head.auto_apply_threshold = float(auto["threshold"]) if auto else None
head.n_pos = len(pos)
head.n_neg = len(neg)
head.ap = float(metrics["ap"])
head.precision_cv = float(metrics["precision"])
head.recall = float(metrics["recall"])
head.trained_at = datetime.now(UTC)
head.metrics = {"f1": metrics["f1"], "auto_apply": auto}
return True
# --- Scoring (async, API worker) -----------------------------------------
# Score one image against every current head to produce the rail's suggestions.
# A tiny in-process cache holds the stacked weight matrix keyed on (count,
# max(trained_at)) so a retrain invalidates it without per-request weight loads.
_HEADS_CACHE: dict[str, Any] = {"key": None, "heads": None}
async def _current_heads(session: AsyncSession, embedding_version: str):
"""Stacked (W, b, thresholds, tag_id/name/category) for heads matching the
current embedding, cached until the next retrain."""
import numpy as np
sig = (
await session.execute(
select(func.count(), func.max(TagHead.trained_at)).where(
TagHead.embedding_version == embedding_version
)
)
).one()
key = f"{embedding_version}:{sig[0]}:{sig[1].isoformat() if sig[1] else '-'}"
cached = _HEADS_CACHE.get("heads")
if cached is not None and _HEADS_CACHE.get("key") == key:
return cached
rows = (
await session.execute(
select(
TagHead.tag_id, Tag.name, Tag.kind,
TagHead.weights, TagHead.bias,
TagHead.suggest_threshold, TagHead.auto_apply_threshold,
)
.join(Tag, Tag.id == TagHead.tag_id)
.where(TagHead.embedding_version == embedding_version)
)
).all()
if not rows:
loaded = {"W": None, "rows": []}
else:
W = np.vstack([np.asarray(r.weights, dtype=np.float32) for r in rows])
b = np.asarray([r.bias for r in rows], dtype=np.float32)
thr = np.asarray([r.suggest_threshold for r in rows], dtype=np.float32)
meta = [
{
"tag_id": r.tag_id,
"name": r.name,
"category": _CATEGORY.get(r.kind, "general"),
"auto_apply_threshold": r.auto_apply_threshold,
}
for r in rows
]
loaded = {"W": W, "b": b, "thr": thr, "meta": meta}
_HEADS_CACHE["key"] = key
_HEADS_CACHE["heads"] = loaded
return loaded
async def score_image(
session: AsyncSession, image_id: int, threshold_override: float | None = None,
) -> list[dict]:
"""Suggestions for one image from the trained heads: [{tag_id, name,
category, score}], ranked. A concept surfaces when its score clears the
head's own suggest_threshold — or, when threshold_override is given (the
typed-dropdown "show everything" mode), that flat floor instead (0 → every
head). Empty if the image has no embedding or no heads exist yet.
MAX-OVER-BAG: the image is scored as a BAG of embeddings — the whole-image
vector PLUS every concept-region crop the agent embedded (same model
version) — and each head takes its MAX score across the bag. A small/local
concept (glasses, a stomach bulge) that the whole-image vector washes out
can still surface from the crop where it dominates. The whole-image vector is
always in the bag, so this can never score lower than whole-image alone."""
import numpy as np
img = await session.get(ImageRecord, image_id)
if img is None or img.siglip_embedding is None:
return []
settings = await _settings_async(session)
heads = await _current_heads(session, settings.embedder_model_version)
if heads["W"] is None:
return []
bag = [np.asarray(img.siglip_embedding, dtype=np.float32)]
region_vecs = (
await session.execute(
select(ImageRegion.siglip_embedding)
.where(ImageRegion.image_record_id == image_id)
.where(ImageRegion.siglip_embedding.is_not(None))
.where(ImageRegion.embedding_version == settings.embedder_model_version)
)
).all()
for (vec,) in region_vecs:
if vec is not None:
bag.append(np.asarray(vec, dtype=np.float32))
X = np.vstack(bag) # (B, D)
norms = np.linalg.norm(X, axis=1, keepdims=True)
norms[norms == 0] = 1.0
Xn = X / norms
Z = Xn @ heads["W"].T + heads["b"] # (B, H)
probs = (1.0 / (1.0 + np.exp(-Z))).max(axis=0) # (H,) best over the bag
out = []
for i, p in enumerate(probs):
cut = threshold_override if threshold_override is not None else heads["thr"][i]
if p >= cut:
m = heads["meta"][i]
out.append({
"tag_id": m["tag_id"],
"name": m["name"],
"category": m["category"],
"score": float(p),
})
out.sort(key=lambda d: d["score"], reverse=True)
return out
async def _settings_async(session: AsyncSession) -> MLSettings:
return (
await session.execute(select(MLSettings).where(MLSettings.id == 1))
).scalar_one()
# --- Earned auto-apply (sync, ml worker) ---------------------------------
# A graduated head can apply its tag to images it scores above the head's
# auto_apply_threshold, without a human. Gated by a master switch + a support
# floor so a precise-looking but under-supported head can't spray tags.
_AUTO_APPLY_CHUNK = 5000
class HeadAutoApplyAlreadyRunning(Exception):
"""Raised when an auto-apply sweep is already in flight."""
class HeadAutoApplyDisabled(Exception):
"""Raised when a real (non-dry-run) sweep is requested but the master
switch (head_auto_apply_enabled) is off."""
def start_head_auto_apply_run(session: Session, params: dict[str, Any]) -> int:
"""Create a HeadAutoApplyRun + dispatch the ml-queue sweep. dry_run previews
(writes nothing); a real sweep needs the master switch on. One run at a time."""
dry_run = bool((params or {}).get("dry_run", False))
existing = session.execute(
select(HeadAutoApplyRun.id).where(HeadAutoApplyRun.status == "running")
).scalar_one_or_none()
if existing is not None:
raise HeadAutoApplyAlreadyRunning(existing)
if not dry_run and not _settings(session).head_auto_apply_enabled:
raise HeadAutoApplyDisabled()
run = HeadAutoApplyRun(
dry_run=dry_run, params={"dry_run": dry_run}, status="running",
last_progress_at=datetime.now(UTC),
)
session.add(run)
session.flush()
run_id = run.id
from ...tasks.ml import apply_head_tags as _task
_task.delay(run_id)
return run_id
def _auto_apply_heads(session: Session, embedding_version: str, min_pos: int):
"""Eligible heads to fire: graduated (auto_apply_threshold set), enough
support, current embedding. Returns the row list (tag_id/name/weights/...)."""
return session.execute(
select(
TagHead.tag_id, Tag.name, TagHead.weights, TagHead.bias,
TagHead.auto_apply_threshold,
)
.join(Tag, Tag.id == TagHead.tag_id)
.where(TagHead.embedding_version == embedding_version)
.where(TagHead.auto_apply_threshold.is_not(None))
.where(TagHead.n_pos >= min_pos)
).all()
def auto_apply_sweep(
session: Session, run: HeadAutoApplyRun, dry_run: bool
) -> dict[str, Any]:
"""Score every embedded image against the eligible heads and apply (or, for
dry_run, just count) each head's tag where score >= its auto_apply_threshold
and the tag isn't already applied or rejected on that image. Streams
embeddings in chunks; commits per chunk on a real run. Returns
{n_applied, concepts:[{tag_id,name,applied,scanned,threshold}]}."""
import numpy as np
from sqlalchemy.dialects.postgresql import insert as pg_insert
settings = _settings(session)
rows = _auto_apply_heads(
session, settings.embedder_model_version,
settings.head_auto_apply_min_positives,
)
if not rows:
return {"n_applied": 0, "concepts": []}
W = np.vstack([np.asarray(r.weights, dtype=np.float32) for r in rows])
b = np.asarray([r.bias for r in rows], dtype=np.float32)
thr = np.asarray([r.auto_apply_threshold for r in rows], dtype=np.float32)
tag_ids = [r.tag_id for r in rows]
names = [r.name for r in rows]
# Skip images that already carry, or have rejected, each tag.
skip = {tid: set() for tid in tag_ids}
for tid in tag_ids:
for (iid,) in session.execute(
select(image_tag.c.image_record_id).where(image_tag.c.tag_id == tid)
):
skip[tid].add(iid)
for (iid,) in session.execute(
select(TagSuggestionRejection.image_record_id).where(
TagSuggestionRejection.tag_id == tid
)
):
skip[tid].add(iid)
applied = [0] * len(rows)
scanned = 0
all_ids = list(session.execute(
select(ImageRecord.id).where(ImageRecord.siglip_embedding.is_not(None))
).scalars())
for start in range(0, len(all_ids), _AUTO_APPLY_CHUNK):
chunk = all_ids[start:start + _AUTO_APPLY_CHUNK]
emb = _load_embeddings(session, chunk)
cids = [i for i in chunk if i in emb]
if not cids:
continue
Xn = _l2norm(np.vstack([emb[i] for i in cids]).astype(np.float32), np)
probs = 1.0 / (1.0 + np.exp(-(Xn @ W.T + b))) # (N, H)
scanned += len(cids)
for h in range(len(rows)):
tid = tag_ids[h]
for idx in np.where(probs[:, h] >= thr[h])[0]:
iid = cids[int(idx)]
if iid in skip[tid]:
continue
skip[tid].add(iid)
applied[h] += 1
if not dry_run:
session.execute(
pg_insert(image_tag)
.values(image_record_id=iid, tag_id=tid, source="head_auto")
.on_conflict_do_nothing()
)
if not dry_run:
session.commit()
run.last_progress_at = datetime.now(UTC)
session.commit()
concepts = [
{"tag_id": tag_ids[h], "name": names[h], "applied": applied[h],
"scanned": scanned, "threshold": float(thr[h])}
for h in range(len(rows))
]
return {"n_applied": sum(applied), "concepts": concepts}
-59
View File
@@ -1,59 +0,0 @@
"""Region read/write for the crop pipeline (#114).
The GPU agent's results endpoint calls replace_regions() to store a freshly
detected/embedded set; the character matcher + concept-bag scorer read via
get_regions(). Replacement is scoped BY KIND so the figure pipeline and the
concept pipeline don't clobber each other.
"""
from typing import Any
from sqlalchemy import delete, select
from sqlalchemy.ext.asyncio import AsyncSession
from ...models import ImageRegion
class RegionService:
def __init__(self, session: AsyncSession):
self.session = session
async def get_regions(
self, image_id: int, kinds: list[str] | None = None
) -> list[ImageRegion]:
stmt = select(ImageRegion).where(ImageRegion.image_record_id == image_id)
if kinds:
stmt = stmt.where(ImageRegion.kind.in_(kinds))
return list(
(await self.session.execute(stmt.order_by(ImageRegion.id))).scalars()
)
async def replace_regions(
self, image_id: int, kinds: list[str], regions: list[dict[str, Any]]
) -> int:
"""Replace this image's regions OF THE GIVEN KINDS with `regions` (a
re-detect/re-propose supersedes the prior set without touching other
kinds). Each region dict: {kind, bbox:(x,y,w,h), score?, detector_version?,
crop_version?, embedding_version?, ccip_embedding?, siglip_embedding?}.
Returns the number inserted."""
await self.session.execute(
delete(ImageRegion)
.where(ImageRegion.image_record_id == image_id)
.where(ImageRegion.kind.in_(kinds))
)
n = 0
for r in regions:
rx, ry, rw, rh = r["bbox"]
self.session.add(ImageRegion(
image_record_id=image_id, kind=r["kind"],
frame_time=r.get("frame_time"),
rx=rx, ry=ry, rw=rw, rh=rh,
score=r.get("score"),
detector_version=r.get("detector_version"),
crop_version=r.get("crop_version"),
embedding_version=r.get("embedding_version"),
ccip_embedding=r.get("ccip_embedding"),
siglip_embedding=r.get("siglip_embedding"),
))
n += 1
return n
+205 -73
View File
@@ -1,23 +1,24 @@
"""The suggestion read-path: trained HEADS score one image's frozen embedding
into alias-resolved, category-grouped, ranked suggestions.
Tagging-v2 (#114): suggestions now come from the per-concept heads that LEARN
from the operator's tags (services/ml/heads.py) — the Camie prediction source
and the per-tag SigLIP centroid have been REMOVED. A head exists only for an
existing concept tag, so every suggestion is a canonical tag (no raw model key,
no alias remap, no creates-new). Rejected tags stay in the list FLAGGED (not
dropped) so the rail can show + reverse a dismissal.
"""The suggestion read-path: raw predictions + centroids -> alias-resolved,
threshold-filtered, category-grouped, ranked suggestions for one image.
"""
from dataclasses import dataclass, field
from sqlalchemy import select
from sqlalchemy import func, select
from sqlalchemy.ext.asyncio import AsyncSession
from ...models import ImageRecord, TagSuggestionRejection
from ...models import (
ImagePrediction,
ImageRecord,
MLSettings,
Tag,
TagSuggestionRejection,
)
from ...models.tag import image_tag
from .ccip import match_image as ccip_match_image
from .heads import score_image
from .aliases import AliasService
from .centroids import CentroidService
from .tag_name import normalize as normalize_tag_name
from .tagger import SURFACED_CATEGORIES
@dataclass(frozen=True)
@@ -28,7 +29,7 @@ class Suggestion:
display_name: str
category: str
score: float
source: str # 'head' | 'ccip' | 'both' (Camie tagger/centroid removed in v2)
source: str # 'tagger' | 'centroid' | 'both'
creates_new_tag: bool
# raw_name = the booru model vocab key behind this suggestion. It's the key
# an alias MUST be stored under (resolution looks up the raw key), so the
@@ -38,11 +39,6 @@ class Suggestion:
# via_alias = this suggestion was surfaced because an operator alias remapped
# the raw prediction to this canonical tag. Lets the UI mark it + offer undo.
via_alias: bool = False
# rejected = the operator dismissed this tag for this image (a stored
# TagSuggestionRejection). It stays in the list — flagged, not dropped — so
# the rejection is VISIBLE and REVERSIBLE in the rail (misclick recovery,
# operator-asked 2026-06-27) instead of silently vanishing or re-suggesting.
rejected: bool = False
@dataclass
@@ -53,24 +49,67 @@ class SuggestionList:
class SuggestionService:
def __init__(self, session: AsyncSession):
self.session = session
self.aliases = AliasService(session)
self.centroids = CentroidService(session)
async def _settings(self) -> MLSettings:
return (
await self.session.execute(select(MLSettings).where(MLSettings.id == 1))
).scalar_one()
async def _load_predictions(self, image_id: int) -> dict:
"""Predictions for one image from the normalized image_prediction
table (#768), in the {raw_name: {category, confidence}} shape the rest
of this service consumed from the old JSON column — so all downstream
threshold/alias/merge logic is unchanged."""
rows = (
await self.session.execute(
select(
ImagePrediction.raw_name,
ImagePrediction.category,
ImagePrediction.score,
).where(ImagePrediction.image_record_id == image_id)
)
).all()
return {
r.raw_name: {"category": r.category, "confidence": r.score}
for r in rows
}
def _threshold_for(
self, s: MLSettings, category: str, override: float | None = None,
) -> float:
# 'artist' (FC-2d-vii-c) and 'copyright' (2026-06-01) retired;
# both fall through to the 1.01 "never surfaces" default like any
# unsurfaced category.
# override (the typed-dropdown "show everything the model saw" mode)
# applies to the surfaced categories only — unsurfaced ones are already
# skipped before the threshold check, so they can't leak in.
if override is not None:
return override
return {
"character": s.suggestion_threshold_character,
"general": s.suggestion_threshold_general,
}.get(category, 1.01)
async def for_image(
self, image_id: int, threshold_override: float | None = None,
self, image_id: int, *, threshold_override: float | None = None,
) -> SuggestionList:
"""Head-scored suggestions for one image, grouped by category and ranked.
"""Ranked suggestions for one image.
Each trained head scores the image's frozen embedding; a concept surfaces
when its score clears the head's own suggest threshold. threshold_override
(used by the typed tag-input dropdown's "show everything" mode) replaces
that per-head cut with a flat floor (0 → every head), so a low-scoring
concept can still be typed + picked in canonical formatting.
Already-applied tags are dropped; rejected tags stay FLAGGED and sink to
the bottom of their category so a dismissal is visible + reversible."""
threshold_override surfaces EVERY stored tagger prediction (down to the
ingest STORE_FLOOR) regardless of the configured per-category suggestion
thresholds — backs the tag-input dropdown's "search all of the model's
predictions, including low-confidence ones, in the canonical formatting"
mode (operator-asked 2026-06-09). The Suggestions panel still calls with
no override so it stays the curated above-threshold list."""
img = await self.session.get(ImageRecord, image_id)
if img is None:
return SuggestionList()
settings = await self._settings()
predictions: dict = await self._load_predictions(image_id)
applied = set(
(
await self.session.execute(
@@ -90,50 +129,148 @@ class SuggestionService:
).scalars().all()
)
hits = await score_image(
self.session, image_id, threshold_override=threshold_override
)
# CCIP character matches OVERLAY the SigLIP character heads — a
# complementary, identity-specialized signal with different failure modes
# (CCIP needs a detected figure; heads work whole-image). Merged by tag:
# 'both' when they corroborate, taking the higher score.
ccip_hits = await ccip_match_image(self.session, image_id)
# --- Camie predictions ---
# candidates carry (raw_name, display_name, category, confidence).
# raw_name = the booru-formatted vocab key, kept for alias_map
# lookup since alias rows are hand-curated against raw keys.
# display_name = normalize_tag_name(raw_name) — what the operator
# sees AND what gets written to tag.name on Accept.
candidates: list[tuple[str, str, str, float]] = []
for name, p in predictions.items():
category = p.get("category", "general")
if category not in SURFACED_CATEGORIES:
continue
conf = float(p.get("confidence", 0.0))
if conf < self._threshold_for(settings, category, threshold_override):
continue
display = normalize_tag_name(name)
if display is None:
# emoticon / pure-punctuation vocab entry — drop entirely
continue
candidates.append((name, display, category, conf))
merged: dict[tuple[str, int], dict] = {}
for h in hits:
merged[(h["category"], h["tag_id"])] = {
"name": h["name"], "score": h["score"], "source": "head",
}
for c in ccip_hits:
key = ("character", c["tag_id"])
ex = merged.get(key)
if ex is not None:
ex["source"] = "both"
ex["score"] = max(ex["score"], c["score"])
alias_map = await self.aliases.resolve_many(
[(raw, c) for raw, _disp, c, _conf in candidates]
)
merged: dict[object, Suggestion] = {}
def _merge(key, sug: Suggestion):
existing = merged.get(key)
if existing is None:
merged[key] = sug
elif sug.score > existing.score:
merged[key] = Suggestion(
canonical_tag_id=existing.canonical_tag_id,
display_name=existing.display_name,
category=existing.category,
score=sug.score,
source="both"
if existing.source != sug.source
else existing.source,
creates_new_tag=existing.creates_new_tag,
# Keep the alias identity from `existing`: the tagger pass
# (which carries raw_name / via_alias) runs before centroid
# augmentation, so it's always the first writer for a key.
raw_name=existing.raw_name,
via_alias=existing.via_alias,
)
for raw, display, category, conf in candidates:
canonical = alias_map.get((raw, category))
if canonical is not None:
if canonical.id in applied or canonical.id in rejected:
continue
_merge(
canonical.id,
Suggestion(
canonical_tag_id=canonical.id,
display_name=canonical.name,
category=category,
score=conf,
source="tagger",
creates_new_tag=False,
raw_name=raw,
via_alias=True,
),
)
else:
merged[key] = {
"name": c["name"], "score": c["score"], "source": "ccip",
}
# Case-insensitive match on BOTH the raw camie key AND
# the normalized form — covers legacy underscore-named
# Tag rows accepted before normalization shipped, AND
# any tag the operator created with the human form.
existing_tag = (
await self.session.execute(
select(Tag).where(
func.lower(Tag.name).in_(
[raw.lower(), display.lower()]
)
)
)
).scalars().first()
if existing_tag is not None:
if (
existing_tag.id in applied
or existing_tag.id in rejected
):
continue
_merge(
existing_tag.id,
Suggestion(
canonical_tag_id=existing_tag.id,
display_name=existing_tag.name,
category=category,
score=conf,
source="tagger",
creates_new_tag=False,
raw_name=raw,
via_alias=False,
),
)
else:
_merge(
f"raw:{display}:{category}",
Suggestion(
canonical_tag_id=None,
display_name=display,
category=category,
score=conf,
source="tagger",
creates_new_tag=True,
raw_name=raw,
via_alias=False,
),
)
# --- Centroid augmentation ---
hits = await self.centroids.find_similar_tags(image_id, limit=30)
for hit in hits:
if hit.similarity < settings.centroid_similarity_threshold:
continue
if hit.tag_id in applied or hit.tag_id in rejected:
continue
tag = await self.session.get(Tag, hit.tag_id)
if tag is None:
continue
cat = tag.kind.value if hasattr(tag.kind, "value") else str(tag.kind)
display_cat = cat if cat in SURFACED_CATEGORIES else "general"
_merge(
tag.id,
Suggestion(
canonical_tag_id=tag.id,
display_name=tag.name,
category=display_cat,
score=hit.similarity,
source="centroid",
creates_new_tag=False,
),
)
result = SuggestionList()
for (cat, tag_id), m in merged.items():
if tag_id in applied:
continue
result.by_category.setdefault(cat, []).append(
Suggestion(
canonical_tag_id=tag_id,
display_name=m["name"],
category=cat,
score=m["score"],
source=m["source"],
creates_new_tag=False,
rejected=tag_id in rejected,
)
)
for sug in merged.values():
result.by_category.setdefault(sug.category, []).append(sug)
for cat in result.by_category:
# Live suggestions first (by score), rejected ones sink to the
# bottom of the category — visible for recovery, out of the way.
result.by_category[cat].sort(key=lambda s: (s.rejected, -s.score))
result.by_category[cat].sort(key=lambda s: s.score, reverse=True)
return result
async def for_selection(
@@ -160,11 +297,6 @@ class SuggestionService:
for s in items:
if s.canonical_tag_id is None or s.creates_new_tag:
continue
# for_image keeps rejected tags (flagged) for the rail;
# bulk consensus must still ignore them — a tag dismissed on
# an image isn't a suggestion for that image.
if s.rejected:
continue
st = stats.get(s.canonical_tag_id)
if st is None:
st = {
-430
View File
@@ -1,430 +0,0 @@
"""Head-vs-centroid tagging eval (#1130, milestone #114 slice 1).
Proves the "frozen embedding + small trained head (with negatives)" spine on the
operator's OWN data, reusing the SigLIP embeddings already stored on
image_record. For each concept tag it compares:
- CENTROID baseline (the old approach): cosine to the mean of positive vectors.
- HEAD (the new approach): logistic regression trained on positives + negatives.
and reports cross-validated precision/recall/AP for both, a LEARNING CURVE
(accuracy as the number of tagged positives grows), and example image ids to
eyeball.
numpy + scikit-learn are imported LAZILY inside run_eval so the API worker (base
image, no ML stack) can still import start_tag_eval_run to enqueue the ml-queue
task — the heavy compute only runs on the ml worker.
"""
from __future__ import annotations
import logging
from datetime import UTC, datetime
from typing import Any
from sqlalchemy import func, select
from sqlalchemy.orm import Session
from ...models import (
ImageRecord,
Tag,
TagEvalRun,
TagKind,
TagPositiveConfirmation,
TagSuggestionRejection,
)
from ...models.tag import image_tag
log = logging.getLogger(__name__)
# The operator's real concept list (mix of whole-ish + small/local cues). The
# admin trigger can override; this is the default eval set.
DEFAULT_CONCEPTS = [
"glasses", "cat", "dog", "horse", "goblin",
"cum", "lactation", "fellatio", "xray", "stomach bulge",
]
DEFAULT_CURVE_POINTS = [10, 30, 100, 300]
DEFAULT_NEG_RATIO = 3 # negatives per positive (rejections + sampled unlabeled)
DEFAULT_CV_FOLDS = 5
MIN_POSITIVES = 8 # below this, a concept can't be evaluated meaningfully
_UNLABELED_POOL = 4000 # cap on sampled unlabeled rows pulled per concept
_EXAMPLES_K = 12
def start_tag_eval_run(session: Session, params: dict[str, Any]) -> int:
"""Create a TagEvalRun (status='running') and dispatch the ml-queue task.
Returns the new run id. Light guard: one running eval at a time."""
existing = session.execute(
select(TagEvalRun.id).where(TagEvalRun.status == "running")
).scalar_one_or_none()
if existing is not None:
raise EvalAlreadyRunning(existing)
norm = _normalize_params(params)
run = TagEvalRun(params=norm, status="running", last_progress_at=datetime.now(UTC))
session.add(run)
session.flush()
run_id = run.id
# Same enqueue-by-import pattern api/suggestions.py uses for ml tasks; the
# commit happens in the API handler so row + dispatch are visible together.
from ...tasks.ml import tag_eval_run as _task
_task.delay(run_id)
return run_id
class EvalAlreadyRunning(Exception):
"""Raised by start_tag_eval_run when an eval is already in flight."""
def _normalize_params(params: dict[str, Any] | None) -> dict[str, Any]:
params = params or {}
concepts = [str(c).strip() for c in (params.get("concepts") or []) if str(c).strip()]
try:
neg_ratio = max(1, int(params.get("neg_ratio", DEFAULT_NEG_RATIO)))
except (TypeError, ValueError):
neg_ratio = DEFAULT_NEG_RATIO
try:
cv_folds = max(2, int(params.get("cv_folds", DEFAULT_CV_FOLDS)))
except (TypeError, ValueError):
cv_folds = DEFAULT_CV_FOLDS
try:
auto_top_n = min(max(int(params.get("auto_top_n", 0) or 0), 0), 200)
except (TypeError, ValueError):
auto_top_n = 0
try:
precision_target = min(max(float(params.get("precision_target", 0.97)), 0.5), 0.999)
except (TypeError, ValueError):
precision_target = 0.97
# No explicit concepts and auto-discovery off → fall back to the hand list.
if not concepts and not auto_top_n:
concepts = list(DEFAULT_CONCEPTS)
curve = params.get("curve_points") or DEFAULT_CURVE_POINTS
curve = sorted({int(n) for n in curve if int(n) > 0})
return {
"concepts": concepts,
"neg_ratio": neg_ratio,
"cv_folds": cv_folds,
"auto_top_n": auto_top_n,
"precision_target": round(precision_target, 4),
"curve_points": curve,
}
def _top_general_concepts(session: Session, n: int, min_count: int) -> list[str]:
"""The n most-tagged general (concept) tags with >= min_count images — a fast
server-side way to broaden the eval beyond the hand-picked list (counts all
sources; source-aware filtering is a separate concern)."""
rows = session.execute(
select(Tag.name)
.join(image_tag, image_tag.c.tag_id == Tag.id)
.where(Tag.kind == TagKind.general)
.group_by(Tag.id)
.having(func.count(image_tag.c.image_record_id) >= min_count)
.order_by(func.count(image_tag.c.image_record_id).desc())
.limit(n)
).all()
return [r[0] for r in rows]
def _resolve_tag_id(session: Session, name: str) -> int | None:
"""Case-insensitive tag-name match; if several share a name, take the one
applied to the most images (the one the operator actually uses)."""
rows = session.execute(
select(Tag.id, func.count(image_tag.c.image_record_id))
.outerjoin(image_tag, image_tag.c.tag_id == Tag.id)
.where(func.lower(Tag.name) == name.lower())
.group_by(Tag.id)
.order_by(func.count(image_tag.c.image_record_id).desc())
).all()
return rows[0][0] if rows else None
def _ids_with_tag(session: Session, tag_id: int) -> list[int]:
return [
r[0] for r in session.execute(
select(image_tag.c.image_record_id).where(image_tag.c.tag_id == tag_id)
).all()
]
def _rejected_ids(session: Session, tag_id: int) -> list[int]:
return [
r[0] for r in session.execute(
select(TagSuggestionRejection.image_record_id)
.where(TagSuggestionRejection.tag_id == tag_id)
).all()
]
def _confirmed_ids(session: Session, tag_id: int) -> set[int]:
"""Positives the operator explicitly affirmed ('keep') — excluded from the
doubts list so confirmed-correct images don't resurface every run."""
return {
r[0] for r in session.execute(
select(TagPositiveConfirmation.image_record_id)
.where(TagPositiveConfirmation.tag_id == tag_id)
).all()
}
def _sample_unlabeled(session: Session, exclude: set[int], limit: int) -> list[int]:
"""Random image ids (with an embedding) NOT carrying the tag. Concepts are
sparse, so an untagged image is almost always a true negative."""
stmt = (
select(ImageRecord.id)
.where(ImageRecord.siglip_embedding.is_not(None))
.order_by(func.random())
.limit(limit)
)
if exclude:
stmt = stmt.where(ImageRecord.id.not_in(exclude))
return [r[0] for r in session.execute(stmt).all()]
def _load_embeddings(session: Session, ids: list[int]) -> dict[int, Any]:
import numpy as np
out: dict[int, Any] = {}
if not ids:
return out
# Chunk the IN list to stay well under psycopg's parameter ceiling.
for i in range(0, len(ids), 2000):
chunk = ids[i:i + 2000]
for rid, emb in session.execute(
select(ImageRecord.id, ImageRecord.siglip_embedding)
.where(ImageRecord.id.in_(chunk))
.where(ImageRecord.siglip_embedding.is_not(None))
).all():
out[rid] = np.asarray(emb, dtype=np.float32)
return out
def run_eval(session: Session, params: dict[str, Any]) -> dict[str, Any]:
"""Compute the full report. Per-concept failures are captured, not fatal."""
import numpy as np
cfg = _normalize_params(params)
# Auto-discovery: union the explicit concepts with the top-N most-tagged
# general tags (server-side, fast) so the eval can broaden itself.
concepts = list(cfg["concepts"])
if cfg["auto_top_n"]:
seen = {c.lower() for c in concepts}
for name in _top_general_concepts(session, cfg["auto_top_n"], MIN_POSITIVES):
if name.lower() not in seen:
concepts.append(name)
seen.add(name.lower())
cfg["concepts"] = concepts
concepts_out = []
for name in cfg["concepts"]:
try:
concepts_out.append(_eval_concept(session, name, cfg, np))
except Exception as exc: # one bad concept shouldn't kill the run
log.exception("tag-eval concept %r failed", name)
concepts_out.append({"name": name, "skipped": f"error: {exc}"})
return {
"generated_at": datetime.now(UTC).isoformat(),
"params": cfg,
"concepts": concepts_out,
}
def _eval_concept(session: Session, name: str, cfg: dict, np) -> dict[str, Any]:
tag_id = _resolve_tag_id(session, name)
if tag_id is None:
return {"name": name, "skipped": "no such tag"}
pos_ids = _ids_with_tag(session, tag_id)
if len(pos_ids) < MIN_POSITIVES:
return {"name": name, "tag_id": tag_id, "n_pos": len(pos_ids),
"skipped": f"too few positives (<{MIN_POSITIVES})"}
neg_ratio = cfg["neg_ratio"]
pos_set = set(pos_ids)
rejected = [i for i in _rejected_ids(session, tag_id) if i not in pos_set]
want_neg = max(len(pos_ids) * neg_ratio, _EXAMPLES_K * 4)
sampled = _sample_unlabeled(session, pos_set | set(rejected),
min(_UNLABELED_POOL, want_neg))
neg_ids = rejected + [i for i in sampled if i not in pos_set]
emb = _load_embeddings(session, pos_ids + neg_ids)
pos = [(i, emb[i]) for i in pos_ids if i in emb]
neg = [(i, emb[i]) for i in neg_ids if i in emb]
if len(pos) < MIN_POSITIVES or len(neg) < MIN_POSITIVES:
return {"name": name, "tag_id": tag_id, "n_pos": len(pos),
"n_neg": len(neg), "skipped": "too few embedded examples"}
ids = np.array([i for i, _ in pos] + [i for i, _ in neg])
X = np.vstack([v for _, v in pos] + [v for _, v in neg]).astype(np.float32)
y = np.array([1] * len(pos) + [0] * len(neg))
Xn = _l2norm(X, np)
head = _eval_head(Xn, y, cfg["cv_folds"], cfg["precision_target"], np)
centroid = _eval_centroid(Xn, y, cfg["cv_folds"], np)
curve = _learning_curve(Xn, y, cfg["curve_points"], neg_ratio, np)
confirmed = _confirmed_ids(session, tag_id)
examples = _examples(session, Xn, y, ids, np, set(rejected), confirmed)
return {
"name": name, "tag_id": tag_id,
"n_pos": len(pos), "n_neg": len(neg),
"n_rejected": len(rejected),
"head": head, "centroid": centroid,
"curve": curve, "examples": examples,
}
def _l2norm(X, np):
n = np.linalg.norm(X, axis=1, keepdims=True)
n[n == 0] = 1.0
return X / n
def _metrics_from_scores(y, scores, np) -> dict[str, float]:
from sklearn.metrics import average_precision_score, precision_recall_curve
ap = float(average_precision_score(y, scores))
prec, rec, thr = precision_recall_curve(y, scores)
f1 = (2 * prec * rec) / np.clip(prec + rec, 1e-9, None)
best = int(np.argmax(f1))
# thr has len = len(prec)-1; map best index safely.
t = float(thr[min(best, len(thr) - 1)]) if len(thr) else 0.5
return {
"ap": round(ap, 4),
"precision": round(float(prec[best]), 4),
"recall": round(float(rec[best]), 4),
"f1": round(float(f1[best]), 4),
"threshold": round(t, 4),
}
def _safe_folds(y, folds, np) -> int:
minority = int(min(np.bincount(y)))
return max(2, min(folds, minority))
def _eval_head(Xn, y, folds, target, np) -> dict[str, float]:
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import StratifiedKFold, cross_val_predict
clf = LogisticRegression(max_iter=1000, class_weight="balanced")
cv = StratifiedKFold(n_splits=_safe_folds(y, folds, np), shuffle=True,
random_state=0)
probs = cross_val_predict(clf, Xn, y, cv=cv, method="predict_proba")[:, 1]
m = _metrics_from_scores(y, probs, np)
m["auto_apply"] = _auto_apply_point(y, probs, target, np)
return m
def _auto_apply_point(y, scores, target, np) -> dict | None:
"""The auto-apply operating point: the threshold that yields the MOST recall
while holding precision >= target. This answers 'could this concept fire
without a human, and how much would it catch?' Returns None if no threshold
reaches the precision target (concept not auto-apply-ready)."""
from sklearn.metrics import precision_recall_curve
prec, rec, thr = precision_recall_curve(y, scores)
best = None # (threshold, precision, recall) maximizing recall s.t. prec>=target
for i in range(len(thr)): # thr[i] corresponds to prec[i], rec[i]
if prec[i] >= target and (best is None or rec[i] > best[2]):
best = (float(thr[i]), float(prec[i]), float(rec[i]))
if best is None:
return None
return {
"target": round(float(target), 4),
"threshold": round(best[0], 4),
"precision": round(best[1], 4),
"recall": round(best[2], 4),
}
def _eval_centroid(Xn, y, folds, np) -> dict[str, float]:
"""Cross-validated cosine-to-positive-mean — the OLD method's quality."""
from sklearn.model_selection import StratifiedKFold
cv = StratifiedKFold(n_splits=_safe_folds(y, folds, np), shuffle=True,
random_state=0)
scores = np.zeros(len(y), dtype=np.float32)
for train, test in cv.split(Xn, y):
c = Xn[train][y[train] == 1].mean(axis=0)
cn = c / (np.linalg.norm(c) or 1.0)
scores[test] = Xn[test] @ cn
return _metrics_from_scores(y, scores, np)
def _learning_curve(Xn, y, points, neg_ratio, np) -> list[dict[str, float]]:
"""Hold out a fixed test split; train the head on a growing number of
positives and watch AP/F1 climb — answers 'does tagging more sharpen it?'"""
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
rng = np.random.default_rng(0)
idx = np.arange(len(y))
try:
tr, te = train_test_split(idx, test_size=0.3, stratify=y, random_state=0)
except ValueError:
return []
tr_pos = tr[y[tr] == 1]
tr_neg = tr[y[tr] == 0]
out = []
for n in points:
if n > len(tr_pos):
break
sp = rng.choice(tr_pos, size=n, replace=False)
nn = min(len(tr_neg), n * neg_ratio)
sn = rng.choice(tr_neg, size=nn, replace=False)
sub = np.concatenate([sp, sn])
clf = LogisticRegression(max_iter=1000, class_weight="balanced")
clf.fit(Xn[sub], y[sub])
prob = clf.predict_proba(Xn[te])[:, 1]
m = _metrics_from_scores(y[te], prob, np)
out.append({"n_pos": int(n), "ap": m["ap"], "f1": m["f1"]})
return out
def _examples(session, Xn, y, ids, np, rejected_set, confirmed_set) -> dict[str, list[dict]]:
"""Train on all data, then surface: top-scoring negatives the operator has
NOT already rejected (= fresh suggestions) and lowest-scoring POSITIVES the
operator has NOT already confirmed (= unreviewed doubts). Excluding rejected
ids stops an adjudicated near-miss from resurfacing in 'would suggest';
excluding confirmed ids stops a 'kept' correct positive from resurfacing in
'head doubts' every run. Resolves thumbnail urls for a self-contained report."""
from sklearn.linear_model import LogisticRegression
clf = LogisticRegression(max_iter=1000, class_weight="balanced")
clf.fit(Xn, y)
s = clf.predict_proba(Xn)[:, 1]
neg_idx = np.where(y == 0)[0]
pos_idx = np.where(y == 1)[0]
top_neg = []
for i in neg_idx[np.argsort(s[neg_idx])[::-1]]: # high score → low
rid = int(ids[i])
if rid in rejected_set:
continue # already told the head 'no' — don't re-suggest it
top_neg.append(rid)
if len(top_neg) >= _EXAMPLES_K:
break
low_pos = []
for i in pos_idx[np.argsort(s[pos_idx])]: # low score → high
rid = int(ids[i])
if rid in confirmed_set:
continue # already kept/confirmed — don't re-doubt it
low_pos.append(rid)
if len(low_pos) >= _EXAMPLES_K:
break
thumbs = _resolve_thumbs(session, top_neg + low_pos)
return {
"head_would_suggest": [thumbs[i] for i in top_neg if i in thumbs],
"head_doubts_positive": [thumbs[i] for i in low_pos if i in thumbs],
}
def _resolve_thumbs(session, ids: list[int]) -> dict[int, dict]:
from ..gallery_service import thumbnail_url
out: dict[int, dict] = {}
if not ids:
return out
for rid, tp, sha, mime in session.execute(
select(
ImageRecord.id, ImageRecord.thumbnail_path,
ImageRecord.sha256, ImageRecord.mime,
).where(ImageRecord.id.in_(ids))
).all():
out[rid] = {"id": rid, "thumbnail_url": thumbnail_url(tp, sha, mime)}
return out
@@ -1,345 +0,0 @@
"""Shared primitives for the native-ingest platform adapters (Patreon,
SubscribeStar, …) — the single home for logic the per-platform client/downloader
modules would otherwise each copy.
DRY pass 2026-06-17 (#899): these used to live in `patreon_*` with the
SubscribeStar modules importing patreon privates (wrong owner + sibling-coupling).
They're platform-agnostic, so they live here and both adapters import them. The
per-platform modules keep only what genuinely differs (feed parsing, the media
shape, Patreon's Mux/yt-dlp video branch + detail-fetch enrichment).
FC runs on a plain-HTTP homelab; nothing here uses a secure-context Web API.
"""
from __future__ import annotations
import contextlib
import http.cookiejar
import json
import logging
import os
import time
from dataclasses import dataclass
from datetime import datetime
from pathlib import Path
from urllib.parse import urlsplit
import requests
from ..utils.paths import filehash_from_url, safe_ext
from .file_validator import is_validatable, quarantine_file, validate_file
log = logging.getLogger(__name__)
_USER_AGENT = (
"Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/122.0.0.0 Safari/537.36"
)
# 429 backoff (plan #703): ride out a transient API rate-limit instead of failing
# the whole walk. Honor the server's Retry-After; else exponential, capped.
_MAX_429_RETRIES = 3
_BACKOFF_BASE_SECONDS = 2.0
_BACKOFF_CAP_SECONDS = 30.0
# Media-download tuning (shared by every platform downloader).
_TIMEOUT_SECONDS = 120.0
_CHUNK = 1 << 16
_MAX_MEDIA_RETRIES = 3
_TRANSIENT_TRANSPORT_EXC = (
requests.ConnectionError,
requests.Timeout,
requests.exceptions.ChunkedEncodingError,
)
_TITLE_MAX = 40
# Windows/gallery-dl path-restrict forbidden set + path separators.
_FORBIDDEN = set('<>:"/\\|?*')
# -- shared exception taxonomy --------------------------------------------
# Every native client raises one of these (platform subclasses keep an
# isinstance-distinct platform name AND the semantic Auth/Drift class), so the
# base Ingester._failure_result can map them platform-agnostically.
class NativeIngestError(Exception):
"""Base for a native-ingest client failure. `status_code` carries the HTTP
status when the failure was an HTTP response (None for transport/parse);
`retry_after` carries the server's 429 Retry-After hint so the cooldown can
match it (plan #708 B1)."""
def __init__(
self,
message: str,
*,
status_code: int | None = None,
retry_after: float | None = None,
):
super().__init__(message)
self.status_code = status_code
self.retry_after = retry_after
class NativeAuthError(NativeIngestError):
"""Authentication/authorization failure — expired/missing credential or an
insufficient tier. The fix is rotating the credential, NOT updating the
ingester. Maps to error_type 'auth_error'."""
class NativeDriftError(NativeIngestError):
"""A response did not match the shape the ingester depends on (JSON:API field
set, or scraped HTML structure). Fail loud so the import step flags 'the
platform changed' instead of silently importing nothing. Maps to API_DRIFT."""
# -- HTTP session ----------------------------------------------------------
def make_session(
cookies_path: str | Path | None,
*,
accept: str = "*/*",
extra_headers: dict | None = None,
) -> requests.Session:
"""Build a requests.Session loaded with the Netscape cookies.txt
CredentialService materializes. `accept` sets the Accept header (the JSON:API
vs HTML feed differ); `extra_headers` adds platform headers (e.g.
X-Requested-With). Missing/unparseable cookies log a warning, never fail."""
session = requests.Session()
headers = {"User-Agent": _USER_AGENT, "Accept": accept}
if extra_headers:
headers.update(extra_headers)
session.headers.update(headers)
if cookies_path and os.path.isfile(str(cookies_path)):
try:
jar = http.cookiejar.MozillaCookieJar(str(cookies_path))
jar.load(ignore_discard=True, ignore_expires=True)
session.cookies = jar # type: ignore[assignment]
except (OSError, http.cookiejar.LoadError) as exc:
log.warning("Could not load cookies from %s: %s", cookies_path, exc)
return session
def retry_after_seconds(
resp: requests.Response,
attempt: int,
*,
base: float = _BACKOFF_BASE_SECONDS,
cap: float = _BACKOFF_CAP_SECONDS,
) -> float:
"""Backoff for a 429: the numeric Retry-After header if present, else
exponential base·2^(attempt-1), both capped."""
header = resp.headers.get("Retry-After")
if header:
try:
return min(float(header), cap)
except (TypeError, ValueError):
pass
return min(base * (2 ** max(0, attempt - 1)), cap)
# -- filename / path helpers -----------------------------------------------
def sanitize_segment(name: str) -> str:
"""Make `name` safe for one filesystem path segment: replace separators, the
Windows-forbidden set, and control chars with `_`; strip trailing dots/spaces
(gallery-dl path-restrict). Never empty (falls back to `_`)."""
out = ["_" if (ch in _FORBIDDEN or ord(ch) < 32) else ch for ch in name]
cleaned = "".join(out).rstrip(". ")
return cleaned or "_"
def basename_from_url(url: str) -> str:
"""Derive a sane filename from a URL when the media has no name: path basename
with a junk-extension guard (safe_ext), bounded stem; falls back to the URL's
content hash, then "file"."""
path = urlsplit(url).path
base = os.path.basename(path)
if base:
ext = safe_ext(base)
stem = base[: -len(Path(base).suffix)] if Path(base).suffix else base
stem = stem[:120] or "file"
return f"{stem}{ext}"
return filehash_from_url(url) or "file"
def post_dir_name(post: dict) -> str:
"""`<YYYY-MM-DD>_<post_id>_<title40>` matching gallery-dl's layout (date prefix
omitted when published_at is missing/unparseable; title is empty for platforms
with no title field). Accepts both ISO and trailing-`Z` published_at."""
post_id = str(post.get("id") or "")
attrs = post.get("attributes") or {}
title = attrs.get("title")
title40 = (title if isinstance(title, str) else "")[:_TITLE_MAX]
published = attrs.get("published_at")
date_prefix = None
if isinstance(published, str) and published:
s = published.strip()
if s.endswith("Z"):
s = s[:-1] + "+00:00"
try:
date_prefix = f"{datetime.fromisoformat(s):%Y-%m-%d}"
except ValueError:
date_prefix = None
raw = f"{date_prefix}_{post_id}_{title40}" if date_prefix else f"{post_id}_{title40}"
return sanitize_segment(raw)
# -- per-item download outcomes (shared dataclasses) -----------------------
@dataclass
class MediaOutcome:
"""Per-media result of a download_post pass. status ∈ downloaded /
skipped_seen / skipped_disk / quarantined / error. `path` is the on-disk file
(downloaded / skipped_disk), the quarantine dest (quarantined), or None;
`error` is the failure/validation reason (error/quarantined) else None."""
media: object
status: str
path: Path | None
error: str | None
@dataclass
class PostRecordOutcome:
"""Result of write_post_record — mirrors the MediaOutcome contract so the core
reports per-post handling. `path` is the _post.json sidecar (None when the post
had no id); the rest is the captured body's shape for the run log."""
path: Path | None
post_type: str | None
title: str | None
body_chars: int
# -- base downloader (shared fetch/validate plumbing) ----------------------
class BaseNativeDownloader:
"""Shared download plumbing for native-platform downloaders: the streaming
GET (transient-retry + Range-resume) and file validation/quarantine. Platform
downloaders subclass this and implement `download_post` / `write_post_record`
/ the per-media sidecar (and any platform-specific fetch, e.g. Patreon's
Mux/yt-dlp video branch). PURE: no DB; the seen-skip is an injected predicate.
"""
def __init__(
self,
images_root: Path,
cookies_path: str | None = None,
*,
platform: str,
validate: bool = True,
rate_limit: float = 0.0,
session: requests.Session | None = None,
):
self.images_root = Path(images_root)
self.cookies_path = str(cookies_path) if cookies_path else None
self.platform = platform
self._validate = validate
self._rate_limit = rate_limit or 0.0
self.session = session if session is not None else make_session(cookies_path)
# -- download seams ----------------------------------------------------
def _fetch_get(self, url: str, dest: Path) -> Path:
"""Stream `url` to a .part then atomic-rename to `dest`."""
part = dest.with_name(dest.name + ".part")
try:
self._fetch_to_file(url, part)
except Exception:
with contextlib.suppress(OSError):
part.unlink()
raise
os.replace(part, dest)
return dest
def _fetch_to_file(self, url: str, dest: Path) -> None:
"""Stream a URL to `dest`, retrying TRANSIENT failures (transport blips,
429 honoring Retry-After, 5xx) with backoff + resume-from-disk (Range);
failing fast on permanent 4xx (404/403). Resume: a retry with bytes on
disk asks `Range: bytes=<have>-`; 206 → append, 200 → restart clean, 416 →
already complete. The caller stages into a `.part` so a non-range server
never corrupts the output."""
attempt = 0
while True:
have = dest.stat().st_size if dest.exists() else 0
headers = {"Range": f"bytes={have}-"} if have > 0 else None
try:
resp = self.session.get(
url, stream=True, timeout=_TIMEOUT_SECONDS, headers=headers,
)
if (resp.status_code == 429 or resp.status_code >= 500) \
and attempt < _MAX_MEDIA_RETRIES:
attempt += 1
delay = retry_after_seconds(resp, attempt)
log.warning(
"%s media transient HTTP %d (%s) — backing off %.1fs "
"(retry %d/%d)",
self.platform, resp.status_code, url, delay, attempt,
_MAX_MEDIA_RETRIES,
)
time.sleep(delay)
continue
if have > 0 and resp.status_code == 416:
return
resp.raise_for_status()
mode = "ab" if (have > 0 and resp.status_code == 206) else "wb"
with open(dest, mode) as fh:
for chunk in resp.iter_content(chunk_size=_CHUNK):
if chunk:
fh.write(chunk)
return
except _TRANSIENT_TRANSPORT_EXC as exc:
if attempt >= _MAX_MEDIA_RETRIES:
raise
attempt += 1
delay = min(2.0 * (2 ** (attempt - 1)), _BACKOFF_CAP_SECONDS)
log.warning(
"%s media transport error (%s) — backing off %.1fs "
"(retry %d/%d): %s",
self.platform, url, delay, attempt, _MAX_MEDIA_RETRIES, exc,
)
time.sleep(delay)
# -- validation --------------------------------------------------------
def _validate_path(
self, path: Path, artist_slug: str, source_url: str | None = None
) -> tuple[str | None, Path | None]:
"""Validate a freshly-written file; quarantine if bad (shared
file_validator move + provenance sidecar). Returns (reason,
quarantine_dest) when quarantined, else (None, None). Logs the quarantine
so a corrupt file is visible in the worker logs, not just counted (#899
L2)."""
if not self._validate or not is_validatable(path):
return None, None
try:
result = validate_file(path)
except Exception as exc:
log.warning("Validator raised on %s: %s", path, exc)
return None, None
if result.ok:
return None, None
dest = quarantine_file(
self.images_root, path, artist_slug, self.platform,
url=source_url, result=result,
)
reason = result.reason or "validation failed"
log.warning(
"%s quarantined %s (%s) — %s",
self.platform, dest or path, artist_slug, reason,
)
return reason, (dest or path)
# -- sidecar (per-media, minimal) --------------------------------------
def _write_minimal_sidecar(
self, post: dict, media_path: Path, *, source_url: str | None = None
) -> Path:
"""Post-first per-media sidecar (#856): image identity ONLY
(category/id/source_url). The post body/links live solely in _post.json."""
data: dict = {"category": self.platform, "id": str(post.get("id") or "")}
if source_url:
data["source_url"] = source_url
sidecar_path = media_path.with_suffix(".json")
sidecar_path.write_text(json.dumps(data, indent=2))
return sidecar_path
+101 -21
View File
@@ -26,7 +26,9 @@ FC runs on a plain-HTTP homelab; nothing here uses a secure-context Web API.
from __future__ import annotations
import http.cookiejar
import logging
import os
import re
import time
from collections.abc import Iterator
@@ -37,23 +39,46 @@ from urllib.parse import parse_qs, urlsplit
import requests
from ..utils.paths import filehash_from_url
from ..utils.paths import filehash_from_url, safe_ext
from ..utils.prosemirror import post_body_html
from .native_ingest_common import (
_MAX_429_RETRIES,
NativeAuthError,
NativeDriftError,
NativeIngestError,
basename_from_url,
make_session,
retry_after_seconds,
)
log = logging.getLogger(__name__)
_POSTS_URL = "https://www.patreon.com/api/posts"
_USER_AGENT = (
"Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/122.0.0.0 Safari/537.36"
)
_TIMEOUT_SECONDS = 30.0
# 429 backoff (plan #703): ride out a transient API rate-limit instead of
# failing the whole walk (which would stamp RATE_LIMITED → platform-wide
# cooldown → every Patreon source dark). Honor the server's `Retry-After`;
# otherwise exponential base·2^(n-1), capped. Only after the retries are
# exhausted does the 429 propagate as terminal RATE_LIMITED.
_MAX_429_RETRIES = 3
_BACKOFF_BASE_SECONDS = 2.0
_BACKOFF_CAP_SECONDS = 30.0
def _retry_after_seconds(
resp: requests.Response,
attempt: int,
*,
base: float = _BACKOFF_BASE_SECONDS,
cap: float = _BACKOFF_CAP_SECONDS,
) -> float:
"""Backoff delay for a 429: the `Retry-After` seconds header if present and
numeric, else exponential `base·2^(attempt-1)`, both capped. (HTTP-date form
of Retry-After is rare here and falls through to exponential.)"""
header = resp.headers.get("Retry-After")
if header:
try:
return min(float(header), cap)
except (TypeError, ValueError):
pass
return min(base * (2 ** max(0, attempt - 1)), cap)
# JSON:API request contract (observed from real traffic — see module plan).
_INCLUDE = (
"campaign,access_rules,attachments,attachments_media,audio,images,media,"
@@ -76,12 +101,30 @@ _FIELDS_CAMPAIGN = "name,url"
_CONTENT_IMG_RE = re.compile(r"<img\b[^>]*?\bsrc=[\"']([^\"']+)[\"']", re.IGNORECASE)
class PatreonAPIError(NativeIngestError):
"""Base for native Patreon client failures. status_code / retry_after are
inherited from NativeIngestError (HTTP status; 429 Retry-After hint)."""
class PatreonAPIError(Exception):
"""Base for native Patreon client failures.
`status_code` carries the HTTP status when the failure was an HTTP response
(None for transport-level / parse failures), so the ingester can map it to a
DownloadResult.error_type (429 → rate_limited, 404 → not_found, …).
`retry_after` carries the server's `Retry-After` seconds on a terminal 429, so
the platform cooldown can match the server's hint instead of a flat default
(plan #708 B1).
"""
def __init__(
self,
message: str,
*,
status_code: int | None = None,
retry_after: float | None = None,
):
super().__init__(message)
self.status_code = status_code
self.retry_after = retry_after
class PatreonAuthError(PatreonAPIError, NativeAuthError):
class PatreonAuthError(PatreonAPIError):
"""Authentication / authorization failure — missing or expired session
cookies, an insufficient pledge tier, or an HTML login/challenge page served
where JSON was expected. DISTINCT from drift: the fix is rotating the
@@ -89,7 +132,7 @@ class PatreonAuthError(PatreonAPIError, NativeAuthError):
"""
class PatreonDriftError(PatreonAPIError, NativeDriftError):
class PatreonDriftError(PatreonAPIError):
"""A JSON response did not match the JSON:API shape we depend on.
Raised for: a missing top-level `data` list, `data` not a list, or a media
@@ -125,12 +168,49 @@ class MediaItem:
post_id: str
def _load_session(cookies_path: str | Path | None) -> requests.Session:
session = requests.Session()
session.headers.update(
{
"User-Agent": _USER_AGENT,
"Accept": "application/vnd.api+json",
}
)
if cookies_path and os.path.isfile(str(cookies_path)):
try:
jar = http.cookiejar.MozillaCookieJar(str(cookies_path))
jar.load(ignore_discard=True, ignore_expires=True)
session.cookies = jar # type: ignore[assignment]
except (OSError, http.cookiejar.LoadError) as exc:
log.warning("Could not load Patreon cookies from %s: %s", cookies_path, exc)
return session
def _filehash(url: str) -> str | None:
# Delegate to the shared extractor (utils.paths) so capture-time persistence
# and render-time inline-image matching use the EXACT same identity.
return filehash_from_url(url)
def _basename_from_url(url: str) -> str:
"""Derive a sane filename from a URL when the media has no file_name.
Strips query/fragment, takes the path basename, and drops a junk
extension (the importer._safe_ext gotcha) so we never write base64 noise
as a name. Falls back to the filehash, then to "file".
"""
path = urlsplit(url).path
base = os.path.basename(path)
if base:
ext = safe_ext(base)
stem = base[: -len(Path(base).suffix)] if Path(base).suffix else base
# Keep the stem bounded; URL-encoded stems can be enormous.
stem = stem[:120] or "file"
return f"{stem}{ext}"
fh = _filehash(url)
return fh or "file"
def parse_cursor_from_url(url: str | None) -> str | None:
"""Extract the `page[cursor]` query param from a links.next URL."""
if not url:
@@ -158,7 +238,7 @@ class PatreonClient:
max_retries: int = _MAX_429_RETRIES,
):
self.cookies_path = str(cookies_path) if cookies_path else None
self._session = make_session(cookies_path, accept="application/vnd.api+json")
self._session = _load_session(cookies_path)
# Politeness: seconds to sleep before each /api/posts page fetch (paces
# the rate-limited API endpoint). 0 = no pacing. plan #703.
self._request_sleep = request_sleep or 0.0
@@ -203,7 +283,7 @@ class PatreonClient:
# through to the terminal RATE_LIMITED raise below.
if resp.status_code == 429 and attempt < self._max_retries:
attempt += 1
delay = retry_after_seconds(resp, attempt)
delay = _retry_after_seconds(resp, attempt)
log.warning(
"Patreon 429 (campaign_id=%s) — backing off %.1fs (retry %d/%d)",
campaign_id, delay, attempt, self._max_retries,
@@ -336,7 +416,7 @@ class PatreonClient:
# uses. A genuine schema change shows up as no URL (above) or a media id
# absent from `included` (caller), not a missing name.
file_name = attrs.get("file_name")
filename = file_name if isinstance(file_name, str) and file_name else basename_from_url(url)
filename = file_name if isinstance(file_name, str) and file_name else _basename_from_url(url)
return MediaItem(
url=url,
filename=filename,
@@ -382,7 +462,7 @@ class PatreonClient:
items.append(
MediaItem(
url=large_url,
filename=basename_from_url(large_url),
filename=_basename_from_url(large_url),
kind="image_large",
filehash=_filehash(large_url),
post_id=post_id,
@@ -401,7 +481,7 @@ class PatreonClient:
filename = (
pf_name
if isinstance(pf_name, str) and pf_name
else basename_from_url(pf_url)
else _basename_from_url(pf_url)
)
items.append(
MediaItem(
@@ -423,7 +503,7 @@ class PatreonClient:
items.append(
MediaItem(
url=src,
filename=basename_from_url(src),
filename=_basename_from_url(src),
kind="content",
filehash=_filehash(src),
post_id=post_id,
+233 -25
View File
@@ -27,33 +27,46 @@ FC runs on a plain-HTTP homelab; nothing here uses a secure-context Web API.
from __future__ import annotations
import contextlib
import json
import logging
import os
import subprocess
import time
from collections.abc import Callable
from dataclasses import dataclass
from datetime import datetime
from pathlib import Path
from urllib.parse import urlsplit
import requests
from ..utils.prosemirror import post_body_html
from .native_ingest_common import (
from .file_validator import is_validatable, quarantine_file, validate_file
from .patreon_client import (
_BACKOFF_CAP_SECONDS,
_MAX_MEDIA_RETRIES,
BaseNativeDownloader,
MediaOutcome,
PostRecordOutcome,
post_dir_name,
sanitize_segment,
_load_session,
_retry_after_seconds,
)
log = logging.getLogger(__name__)
# yt-dlp subprocess wall-clock per attempt (video only; the shared HTTP fetch
# budgets live in BaseNativeDownloader).
_TITLE_MAX = 40
_TIMEOUT_SECONDS = 120.0
_CHUNK = 1 << 16
# Retry a media GET that hits a TRANSIENT failure within the same pass (plan
# #705 #8): a transport blip (connection reset / timeout / truncated stream), a
# 429, or a 5xx. PERMANENT failures (404 gone, 403 forbidden) fail fast straight
# to the error/dead-letter path — no point re-fetching them. Keeps a momentary
# network hiccup from becoming a per-item error that waits for the next walk.
_MAX_MEDIA_RETRIES = 3
# requests transport errors worth retrying (vs. an HTTPError, which is a real
# server response and is classified by status code).
_TRANSIENT_TRANSPORT_EXC = (
requests.ConnectionError,
requests.Timeout,
requests.exceptions.ChunkedEncodingError,
)
# Referer/Origin yt-dlp must send for Mux-hosted Patreon video. Mux's JWT
# playback policy checks Referer/Origin on every request, so yt-dlp must send
@@ -65,6 +78,26 @@ _VIDEO_HEADERS = {
"Origin": "https://www.patreon.com",
}
# Characters Windows/gallery-dl path-restrict forbids, plus path separators.
_FORBIDDEN = set('<>:"/\\|?*')
def _sanitize(name: str) -> str:
"""Make `name` safe for a single filesystem path segment.
Replaces path separators, the Windows-forbidden set <>:"/\\|?* and control
characters with `_`, then strips trailing dots/spaces (gallery-dl
path-restrict behavior). Never returns empty (falls back to "_").
"""
out = []
for ch in name:
if ch in _FORBIDDEN or ord(ch) < 32:
out.append("_")
else:
out.append(ch)
cleaned = "".join(out).rstrip(". ")
return cleaned or "_"
def _is_video_url(url: str) -> bool:
parts = urlsplit(url)
@@ -73,12 +106,73 @@ def _is_video_url(url: str) -> bool:
return parts.path.lower().endswith(".m3u8")
class PatreonDownloader(BaseNativeDownloader):
"""Download resolved Patreon media to gallery-dl's on-disk layout. Subclasses
BaseNativeDownloader for the shared streaming GET (transient-retry +
Range-resume) and validation/quarantine; adds the Mux/HLS yt-dlp video branch
and the detail-fetch body enrichment. PURE: no DB. `_run_ytdlp` is
monkeypatchable and the HTTP session is the injectable `session=` seam.
def _post_dir_name(post: dict) -> str:
"""Build the post directory name matching gallery-dl's layout."""
post_id = str(post.get("id") or "")
attrs = post.get("attributes") or {}
title = attrs.get("title")
title = title if isinstance(title, str) else ""
title40 = title[:_TITLE_MAX]
published = attrs.get("published_at")
date_prefix = None
if isinstance(published, str) and published:
s = published.strip()
if s.endswith("Z"):
s = s[:-1] + "+00:00"
try:
dt = datetime.fromisoformat(s)
except ValueError:
dt = None
if dt is not None:
date_prefix = f"{dt:%Y-%m-%d}"
if date_prefix:
raw = f"{date_prefix}_{post_id}_{title40}"
else:
raw = f"{post_id}_{title40}"
return _sanitize(raw)
@dataclass
class MediaOutcome:
"""Per-media result of a download_post pass.
status is one of: "downloaded", "skipped_seen", "skipped_disk",
"quarantined", "error". `path` is the final on-disk path for "downloaded"
(the actual yt-dlp output for video), the path that already existed for
"skipped_disk", or the _quarantine destination for "quarantined"; None for
"skipped_seen" and (usually) "error". `error` carries the failure/validation
reason for "error"/"quarantined", else None.
"""
media: object # MediaItem (avoid importing the name for a bare annotation)
status: str
path: Path | None
error: str | None
@dataclass
class PostRecordOutcome:
"""Result of write_post_record — mirrors the download_post → MediaOutcome
contract so the engine reports per-post handling without re-reading the post.
`path` is the _post.json sidecar (None when the post had no id); the rest is
the captured body's shape (post_type + final char count) for the run log.
"""
path: Path | None
post_type: str | None
title: str | None
body_chars: int
class PatreonDownloader:
"""Download resolved Patreon media to gallery-dl's on-disk layout.
PURE: no DB. The HTTP session and the yt-dlp invocation are injectable seams
so tests run without network or a real subprocess:
- pass `session=` to stub `session.get`, or monkeypatch `_fetch_to_file`.
- monkeypatch `_run_ytdlp` to avoid spawning yt-dlp.
"""
def __init__(
@@ -91,16 +185,23 @@ class PatreonDownloader(BaseNativeDownloader):
session: requests.Session | None = None,
content_fetcher: Callable[[str], str | None] | None = None,
):
super().__init__(
images_root, cookies_path, platform="patreon",
validate=validate, rate_limit=rate_limit, session=session,
)
self.images_root = Path(images_root)
self.cookies_path = str(cookies_path) if cookies_path else None
self._validate = validate
# Best-effort enrichment seam: (post_id) -> full HTML body, or None. The
# feed endpoint often omits `content`; the adapter wires this to
# PatreonClient.fetch_post_detail_content so the sidecar captures the
# real body (formatting + inline <img> + external <a href> links).
# None in unit tests / when enrichment isn't wanted.
self._content_fetcher = content_fetcher
# Politeness: seconds to sleep before each actual media download (paces
# the CDN; honors ImportSettings.download_rate_limit_seconds, the same
# value gallery-dl used as its between-downloads `sleep`). 0 = no pacing.
# Applied only to real downloads, not to seen/disk skips. plan #703.
self._rate_limit = rate_limit or 0.0
# Build a cookie-loaded session the same way patreon_client does, so the
# CDN GETs carry the creator's auth.
self.session = session if session is not None else _load_session(cookies_path)
# -- public ------------------------------------------------------------
@@ -133,7 +234,7 @@ class PatreonDownloader(BaseNativeDownloader):
would tier-1 skip it — so the engine can backfill source_filehash for
inline-image localization. Genuinely-missing seen media is NOT refetched.
"""
post_dir = self.images_root / artist_slug / "patreon" / post_dir_name(post)
post_dir = self.images_root / artist_slug / "patreon" / _post_dir_name(post)
outcomes: list[MediaOutcome] = []
for i, media in enumerate(media_items, start=1):
@@ -178,7 +279,7 @@ class PatreonDownloader(BaseNativeDownloader):
return MediaOutcome(media=media, status="skipped_seen", path=None, error=None)
nn = f"{index:02d}"
final_name = sanitize_segment(f"{nn}_{media.filename}")
final_name = _sanitize(f"{nn}_{media.filename}")
media_path = post_dir / final_name
# tier-2: already on disk.
@@ -232,9 +333,87 @@ class PatreonDownloader(BaseNativeDownloader):
self._write_sidecar(post, out_path, source_url=media.url)
return MediaOutcome(media=media, status="downloaded", path=out_path, error=None)
# -- video (Mux/HLS via yt-dlp) ----------------------------------------
# The plain-GET streaming path (_fetch_get / _fetch_to_file) and
# _validate_path are inherited from BaseNativeDownloader.
# -- download seams ----------------------------------------------------
def _fetch_get(self, url: str, dest: Path) -> Path:
"""Stream `url` to a .part file then atomic-rename to `dest`.
Thin wrapper over `_fetch_to_file` so tests can stub either the whole
GET path (`_fetch_to_file`) or just `session.get`.
"""
part = dest.with_name(dest.name + ".part")
try:
self._fetch_to_file(url, part)
except Exception:
with contextlib.suppress(OSError):
part.unlink()
raise
os.replace(part, dest)
return dest
def _fetch_to_file(self, url: str, dest: Path) -> None:
"""Stream a non-video URL to `dest` via the (stubbable) session, retrying
TRANSIENT failures within the same pass (plan #705 #8) and RESUMING from
the bytes already on disk via a Range request when a retry follows a
mid-download cut (plan #708 B5).
Retried (backoff): transport blips (connection reset / timeout /
truncated stream — incl. mid-download), HTTP 429 (honoring Retry-After),
and 5xx. Failed fast (no retry → HTTPError → per-item error → dead-letter
path): 4xx other than 429 (404 gone, 403 forbidden) — re-fetching a
permanent failure is pointless.
Resume: on a retry, if bytes already landed in `dest`, ask for the rest
with `Range: bytes=<have>-`. A 206 means the server honored it → append; a
200 means it ignored it (served the whole file) → start clean. The caller
(_fetch_get) stages into a `.part`, so a non-range server never corrupts
the output — the worst case is re-downloading from zero, as before.
"""
attempt = 0
while True:
have = dest.stat().st_size if dest.exists() else 0
headers = {"Range": f"bytes={have}-"} if have > 0 else None
try:
resp = self.session.get(
url, stream=True, timeout=_TIMEOUT_SECONDS, headers=headers,
)
if (resp.status_code == 429 or resp.status_code >= 500) \
and attempt < _MAX_MEDIA_RETRIES:
attempt += 1
delay = _retry_after_seconds(resp, attempt)
log.warning(
"Patreon media transient HTTP %d (%s) — backing off "
"%.1fs (retry %d/%d)",
resp.status_code, url, delay, attempt, _MAX_MEDIA_RETRIES,
)
time.sleep(delay)
continue
# A Range that starts at/past EOF (we already have the whole file)
# comes back 416 — the bytes we kept ARE the file.
if have > 0 and resp.status_code == 416:
return
# 2xx → ok; 4xx-non-429 (or an exhausted 429/5xx) → HTTPError
# (permanent for this pass) → not caught below → per-item error.
resp.raise_for_status()
# 206 → server honored the Range; append after the kept bytes.
# Anything else (200) → it served the whole file → start clean.
mode = "ab" if (have > 0 and resp.status_code == 206) else "wb"
with open(dest, mode) as fh:
for chunk in resp.iter_content(chunk_size=_CHUNK):
if chunk:
fh.write(chunk)
return
except _TRANSIENT_TRANSPORT_EXC as exc:
if attempt >= _MAX_MEDIA_RETRIES:
raise # exhausted → terminal error outcome
attempt += 1
delay = min(2.0 * (2 ** (attempt - 1)), _BACKOFF_CAP_SECONDS)
log.warning(
"Patreon media transport error (%s) — backing off %.1fs "
"(retry %d/%d): %s",
url, delay, attempt, _MAX_MEDIA_RETRIES, exc,
)
time.sleep(delay)
def _run_ytdlp(self, url: str, dest: Path, headers: dict) -> Path | None:
"""Invoke yt-dlp to fetch a Mux/HLS stream to (around) `dest`.
@@ -316,6 +495,35 @@ class PatreonDownloader(BaseNativeDownloader):
return cand
return None
# -- validation --------------------------------------------------------
def _validate_path(
self, path: Path, artist_slug: str, source_url: str | None = None
) -> tuple[str | None, Path | None]:
"""Validate a freshly-written file; quarantine if bad.
Uses the shared `file_validator.quarantine_file` — same move + provenance
sidecar gallery-dl writes (the native path used to skip the sidecar; that
parity gap is closed here). Returns `(reason, quarantine_dest)` when
quarantined (dest is the original path if the move itself failed), else
`(None, None)` (ok / not validatable / disabled). plan #704: the dest is
surfaced so the run reports a real quarantined-paths list.
"""
if not self._validate or not is_validatable(path):
return None, None
try:
result = validate_file(path)
except Exception as exc:
log.warning("Validator raised on %s: %s", path, exc)
return None, None
if result.ok:
return None, None
dest = quarantine_file(
self.images_root, path, artist_slug, "patreon",
url=source_url, result=result,
)
return (result.reason or "validation failed"), (dest or path)
# -- sidecar -----------------------------------------------------------
def _write_sidecar(
@@ -406,7 +614,7 @@ class PatreonDownloader(BaseNativeDownloader):
return PostRecordOutcome(
path=None, post_type=post_type, title=title, body_chars=0,
)
post_dir = self.images_root / artist_slug / "patreon" / post_dir_name(post)
post_dir = self.images_root / artist_slug / "patreon" / _post_dir_name(post)
post_dir.mkdir(parents=True, exist_ok=True)
path = self._write_sidecar_data(post, post_dir / "_post.json")
# _write_sidecar_data has by now memoized any detail-fetched body onto
+51 -4
View File
@@ -35,8 +35,15 @@ from collections.abc import Callable
from pathlib import Path
from ..models import PatreonFailedMedia, PatreonSeenMedia
from .gallery_dl import DownloadResult, ErrorType
from .ingest_core import DEAD_LETTER_THRESHOLD, Ingester
from .patreon_client import MediaItem, PatreonAPIError, PatreonClient
from .patreon_client import (
MediaItem,
PatreonAPIError,
PatreonAuthError,
PatreonClient,
PatreonDriftError,
)
from .patreon_downloader import PatreonDownloader
from .patreon_resolver import extract_vanity, resolve_campaign_id_for_source
@@ -122,11 +129,51 @@ class PatreonIngester(Ingester):
ledger_key=_ledger_key,
platform="patreon",
error_base=PatreonAPIError,
# API_DRIFT message phrasing; the base Ingester._failure_result owns
# the auth/drift/HTTP→error_type mapping now (shared across platforms).
drift_label="Patreon API",
)
# -- failure mapping (Patreon exception taxonomy) ----------------------
def _failure_result(self, exc: Exception, _result) -> DownloadResult:
"""Map a client-level exception to a loud, typed failed DownloadResult.
We NEVER return a silent zero-download "success" — the whole point of the
native ingester is to fail RED when Patreon's API shape or our auth
changes. The typed mapping lets FailingSourcesCard render the right chip
and tells the operator what to do:
- PatreonAuthError → AUTH_ERROR (rotate cookies)
- PatreonDriftError → API_DRIFT (ingester field-set/parser needs update)
- HTTP 429 / 404 → RATE_LIMITED / NOT_FOUND
- other HTTP status → HTTP_ERROR; transport failure → NETWORK_ERROR
PatreonAuthError and PatreonDriftError both subclass PatreonAPIError, so
they must be matched before the generic HTTP/transport fallthrough.
"""
message = str(exc)
if isinstance(exc, PatreonAuthError):
error_type = ErrorType.AUTH_ERROR
elif isinstance(exc, PatreonDriftError):
error_type = ErrorType.API_DRIFT
message = f"Patreon API changed — ingester needs update: {message}"
else: # generic PatreonAPIError: HTTP non-2xx (status_code set) or transport
status = getattr(exc, "status_code", None)
if status == 429:
error_type = ErrorType.RATE_LIMITED
elif status == 404:
error_type = ErrorType.NOT_FOUND
elif status is not None:
error_type = ErrorType.HTTP_ERROR
else:
error_type = ErrorType.NETWORK_ERROR
log.warning("Patreon ingest failed (%s): %s", error_type.value, message)
result = _result(
success=False, return_code=1,
error_type=error_type, error_message=message,
)
# plan #708 B1: carry the server's Retry-After up to the cooldown.
if error_type == ErrorType.RATE_LIMITED:
result.retry_after_seconds = getattr(exc, "retry_after", None)
return result
async def verify_patreon_credential(
url: str,
+3 -4
View File
@@ -21,10 +21,9 @@ from ..config import get_config
log = logging.getLogger(__name__)
# Platforms walked one-at-a-time. gallery-dl platforms are intentionally NOT
# here: each runs as a self-pacing subprocess and they're lower-volume. The
# native-ingester platforms are serialized (one paced scrape/API walk at a time).
# Add a platform here to cap it to a single concurrent walk.
SERIALIZED_PLATFORMS = frozenset({"patreon", "subscribestar"})
# here: each runs as a self-pacing subprocess and they're lower-volume. Add a
# platform here to cap it to a single concurrent walk.
SERIALIZED_PLATFORMS = frozenset({"patreon"})
_LOCK_PREFIX = "fc:download_lock:"
+1 -31
View File
@@ -66,10 +66,6 @@ class ProvenanceService:
).scalars().all()
return [_attachment_dict(a) for a in rows]
async def _attachment_by_id(self, attachment_id: int) -> list[dict]:
att = await self.session.get(PostAttachment, attachment_id)
return [_attachment_dict(att)] if att is not None else []
async def for_image(self, image_id: int) -> dict | None:
rec = await self.session.get(ImageRecord, image_id)
if rec is None:
@@ -89,33 +85,7 @@ class ProvenanceService:
)
rows = (await self.session.execute(stmt)).all()
post_ids = [ip.post_id for ip, _p, _s, _a in rows]
# Prefer the EXACT archive this file came out of (milestone #87): if the
# originating post's provenance row records from_attachment_id, the image
# was extracted from that one .zip/.rar, so show only it — not the dozens
# of unrelated archives a "High Resolution Files" bundle post carries.
from_att_id = next(
(
ip.from_attachment_id
for ip, _p, _s, _a in rows
if ip.post_id == rec.primary_post_id
and ip.from_attachment_id is not None
),
None,
)
if from_att_id is not None:
attachments = await self._attachment_by_id(from_att_id)
else:
# No recorded containing archive (loose download, or pre-backfill):
# scope to the originating post only, not every pHash-linked post.
# primary_post_id is the post this file was actually captured from;
# fall back to all linked posts when it's unset (older rows /
# filesystem imports).
attach_post_ids = (
[rec.primary_post_id]
if rec.primary_post_id is not None
else post_ids
)
attachments = await self._attachments_for_posts(attach_post_ids)
attachments = await self._attachments_for_posts(post_ids)
return {
"image_id": image_id,
"provenance": [
@@ -1,602 +0,0 @@
"""Native SubscribeStar client — the SubscribeStar adapter's read path.
Unlike Patreon (a clean JSON:API), SubscribeStar has NO public API: gallery-dl
scrapes its HTML, and so do we. The platform-agnostic core (`ingest_core`) only
calls `client.iter_posts` / `extract_media` / the optional seams, so the
HTML-scrape divergence is contained entirely to this module.
Feed shape (characterized from a live sample 2026-06-17 — Scribe note
"SubscribeStar HTML characterization"):
- Page 1 is the creator page HTML: GET <base>/<slug>.
- Each page embeds `data-role="infinite_scroll-next_page" href="/posts?...&page=N
&slug=<slug>&sort_by=newest"`. That path returns JSON {"html": "<...posts...>"};
the fragment carries the NEXT page link, until it runs out.
- A post is `<div class="post is-shown ..." data-id="<post_id>" ...>`; body lives
in `.post-content .trix-content`; date in `.post-date`; media in a `data-gallery`
JSON manifest on `.uploads-images`.
`campaign_id` for SubscribeStar is the full creator URL (e.g.
https://subscribestar.adult/sabu) — the host (.com vs .adult) and slug both come
from it, so no separate resolver is needed.
Drift is loud on purpose (mirrors patreon_client): an HTML login/age-gate where
the feed was expected is AUTH (rotate cookies); a feed whose post structure we
can't parse at all is DRIFT (the scraper needs updating). FC runs on a
plain-HTTP homelab; nothing here uses a secure-context Web API.
"""
from __future__ import annotations
import json
import logging
import re
import time
from collections.abc import Iterator
from dataclasses import dataclass
from datetime import datetime
from html import unescape
from pathlib import Path
from urllib.parse import urljoin, urlsplit
import requests
from ..utils.paths import filehash_from_url
from .native_ingest_common import (
_MAX_429_RETRIES,
NativeAuthError,
NativeDriftError,
NativeIngestError,
basename_from_url,
make_session,
retry_after_seconds,
)
log = logging.getLogger(__name__)
_TIMEOUT_SECONDS = 30.0
# Match gallery-dl's default (cookies-only) request profile EXACTLY — the proven
# way to read SubscribeStar without tripping its /verify_subscriber gate. In that
# mode gallery-dl's base Extractor._init_session sends a Firefox UA, Accept: */*,
# Accept-Language, and a same-site Referer (root/) on EVERY request — including
# the first creator-page GET — and uses NO X-Requested-With on either the creator
# page or the "load more" JSON endpoint (it GETs and parses the body as JSON).
# Our prior Chrome UA + missing Referer + XHR toggling looked enough unlike a
# browser that SubscribeStar 302'd the adult-creator page to /<slug>/verify_
# subscriber even with valid cookies (cheunart, 2026-06-17).
_FIREFOX_UA = (
"Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:140.0) "
"Gecko/20100101 Firefox/140.0"
)
_GDL_HEADERS = {
"User-Agent": _FIREFOX_UA,
"Accept-Language": "en-US,en;q=0.5",
}
# A post block opens with this wrapper; we slice the page between consecutive
# occurrences (regex can't match the balanced close of nested divs, so chunk-
# per-post is the robust approach gallery-dl also uses).
#
# Delimiter is the GENERIC `<div class="post ` (trailing space — matches the post
# container regardless of the classes that follow), exactly what gallery-dl's
# `_pagination` splits on. We previously keyed on `<div class="post is-shown`,
# but `is-shown` is added by SubscribeStar's infinite-scroll JS when a post
# scrolls into view — it's present in a browser-SAVED page but ABSENT from the
# raw server HTML we (and gallery-dl) actually fetch, so the raw feed parsed to
# zero posts → false drift (cheunart, 2026-06-17). The space rules out the
# hyphenated siblings (`post-content`/`post-date`/`post-body`/`post-uploads`).
_POST_OPEN = '<div class="post '
_POST_ID_RE = re.compile(r'data-id="(\d+)"')
# The date may be plain text OR wrapped in an <a> permalink — image posts wrap it
# (`<div class="post-date"><a href="/posts/ID">DATE</a></div>`), text-only posts
# don't. gallery-dl's `_data_from_post` handles both: take text up to the first
# `</`, then whatever follows the last `>`. A naive
# `<div class="post-date">([^<]+)</div>` regex matched ONLY the unwrapped case, so
# every image post got a null date and sorted to the top (cheunart 2026-06-17).
_DATE_OPEN = 'class="post-date">'
# The body lives between these two LITERAL markers — exactly gallery-dl's
# `_data_from_post` extraction: from the post_content-text wrapper open to the
# youtube-uploads div that always follows post-content. A balanced-</div> regex
# either returned empty or over-captured into sibling upload divs AND the
# "View next posts (N / M)" pagination counter (the "264 / 265" body bug,
# cheunart 2026-06-17). The trix editor wraps rich bodies in a full
# `<html><body>…</body></html>` document, so strip to the body inner when present.
_CONTENT_OPEN = '<div class="post-content" data-role="post_content-text">'
_CONTENT_CLOSE = '</div><div class="post-uploads for-youtube"'
_GALLERY_RE = re.compile(r'data-gallery="([^"]*)"')
# Document + audio attachments are NOT in data-gallery — gallery-dl's
# _media_from_post scrapes them from their own `uploads-docs` / `uploads-audios`
# sections, splitting on each preview block. Some posts deliver content ONLY
# through these (PDFs/zips, audio), so we must walk them too.
_DOC_SPLIT_RE = re.compile(r'class="doc_preview[" ]')
_AUDIO_SPLIT_RE = re.compile(r'class="audio_preview-data[" ]')
_NEXT_PAGE_RE = re.compile(
r'data-role="infinite_scroll-next_page"\s+href="([^"]+)"'
)
# Signals an auth/age wall served in place of the feed (cookies expired or the
# age cookie missing) rather than a real — possibly empty — creator feed.
_LOGIN_MARKERS = ("/session/new", 'data-role="sign_in"', "age_confirmation_warning")
# An interstitial served INSTEAD of the feed — characterized so a drift error
# reports the actual cause (bot challenge / age gate / login) rather than a bare
# "markup changed". (name, substrings to look for, case-insensitive.)
_INTERSTITIAL_MARKERS = (
("cloudflare/bot-challenge", (
"just a moment", "cf-challenge", "challenge-platform", "cf_chl",
"attention required", "enable javascript and cookies",
)),
("age-gate", (
"18 or older", "adult content", "age_confirmation", "i am over",
"confirm your age", "must be 18",
)),
("login", ("/session/new", 'data-role="sign_in"', "sign in", "log in")),
("captcha", ("g-recaptcha", "hcaptcha", "captcha")),
)
def _describe_page(html: str) -> str:
"""A short, log-safe description of an unexpected page: its <title> + which
known interstitial it resembles (bot challenge / age gate / login / captcha)."""
m = re.search(r"<title[^>]*>(.*?)</title>", html, re.IGNORECASE | re.DOTALL)
title = unescape(m.group(1).strip())[:120] if m else "(no <title>)"
low = html.lower()
hits = [name for name, needles in _INTERSTITIAL_MARKERS
if any(n.lower() in low for n in needles)]
return f"title={title!r}; resembles: {'/'.join(hits) if hits else 'unrecognized'}"
class SubscribeStarAPIError(NativeIngestError):
"""Base for native SubscribeStar client failures. status_code / retry_after
are inherited from NativeIngestError."""
class SubscribeStarAuthError(SubscribeStarAPIError, NativeAuthError):
"""Auth/authorization failure — expired cookies, missing age cookie, or an
HTML login/age wall served where the feed was expected. Fix = rotate the
credential, not update the scraper. Maps to error_type 'auth_error'."""
class SubscribeStarDriftError(SubscribeStarAPIError, NativeDriftError):
"""The feed HTML did not match the structure we scrape (no recognizable post
blocks AND no known empty-feed state). The scrape analog of API drift — fail
loud so the import step flags 'SubscribeStar changed its markup' instead of
silently importing nothing."""
@dataclass
class MediaItem:
"""One resolved downloadable item belonging to a SubscribeStar post.
Fields mirror patreon_client.MediaItem (so the downloader is structurally the
same) plus `media_id` — the stable per-upload gallery id. SubscribeStar's
full-res URL is an opaque `/post_uploads?payload=...` (not content-addressed),
so `filehash` is usually None and the ledger keys on `<post_id>:<media_id>`.
"""
url: str
filename: str
kind: str
filehash: str | None
post_id: str
media_id: str
def _extr(text: str, start: str, end: str) -> str:
"""Substring between the first `start` and the next `end` after it (gallery-
dl's `text.extr`); '' when either marker is absent."""
i = text.find(start)
if i < 0:
return ""
i += len(start)
j = text.find(end, i)
if j < 0:
return ""
return text[i:j]
def _extract_content(chunk: str) -> str:
"""The post body HTML — gallery-dl's `_data_from_post` content rule: between
the post_content-text wrapper and the youtube-uploads div, with the trix
editor's `<html><body>…</body></html>` wrapper stripped to its inner."""
content = _extr(chunk, _CONTENT_OPEN, _CONTENT_CLOSE)
if "<html><body>" in content:
content = _extr(content, "<body>", "</body>")
return content.strip()
def _attachment_item(
frag: str, base: str, post_id: str, *, kind: str, title_marker: str, url_attr: str
) -> MediaItem | None:
"""One doc/audio attachment from its preview block — gallery-dl's
`_media_from_post` attachment/audio fields. `url_attr` is the URL-bearing
attribute (`href="` for docs, `src="` for audio); `title_marker` precedes the
display name. Returns None when the block carries no URL."""
rel = unescape(_extr(frag, url_attr, '"'))
if not rel:
return None
url = urljoin(base + "/", rel)
upload_id = _extr(frag, 'data-upload-id="', '"')
name = unescape(_extr(frag, title_marker, "<")).strip()
return MediaItem(
url=url,
filename=name or basename_from_url(url),
kind=kind,
filehash=filehash_from_url(url),
post_id=post_id,
media_id=str(upload_id or ""),
)
def _split_creator_url(campaign_id: str) -> tuple[str, str]:
"""`campaign_id` is the creator URL → (base, slug).
base = scheme://host (preserving .com vs .adult); slug = first path segment.
"""
parts = urlsplit(campaign_id)
base = f"{parts.scheme or 'https'}://{parts.netloc}"
slug = parts.path.strip("/").split("/")[0] if parts.path else ""
return base, slug
def _parse_ss_datetime(text: str) -> str | None:
"""SubscribeStar renders human dates: 'Jun 17, 2026 03:19 am' (and an
'Updated on <date>' variant). Return ISO-8601 (UTC-naive) or None."""
s = unescape(text or "").strip()
if s.lower().startswith("updated on "):
s = s[len("updated on "):].strip()
for fmt in ("%b %d, %Y %I:%M %p", "%b %d, %Y"):
try:
return datetime.strptime(s, fmt).isoformat()
except ValueError:
continue
return None
class SubscribeStarClient:
"""Synchronous SubscribeStar HTML-scrape read client. Construct with a path
to a Netscape cookies.txt (the same file CredentialService.get_cookies_path
materializes, already carrying the age cookie via augment_cookies)."""
def __init__(
self,
cookies_path: str | Path | None,
*,
request_sleep: float = 0.0,
max_retries: int = _MAX_429_RETRIES,
):
self.cookies_path = str(cookies_path) if cookies_path else None
# gallery-dl-parity request profile (Firefox UA, Accept: */*, Accept-
# Language; Referer is stamped per-walk once the creator base is known).
self._session = make_session(
cookies_path, accept="*/*", extra_headers=_GDL_HEADERS
)
self._request_sleep = request_sleep or 0.0
self._max_retries = max_retries
# -- request -----------------------------------------------------------
def _get(self, url: str, *, headers: dict | None = None) -> requests.Response:
if self._request_sleep > 0:
time.sleep(self._request_sleep)
attempt = 0
while True:
try:
resp = self._session.get(url, timeout=_TIMEOUT_SECONDS, headers=headers)
except requests.RequestException as exc:
raise SubscribeStarAPIError(
f"SubscribeStar request failed ({url}): {exc}"
) from exc
if resp.status_code == 429 and attempt < self._max_retries:
attempt += 1
delay = retry_after_seconds(resp, attempt)
log.warning(
"SubscribeStar 429 (%s) — backing off %.1fs (retry %d/%d)",
url, delay, attempt, self._max_retries,
)
time.sleep(delay)
continue
break
# gallery-dl's own gating signal: SubscribeStar 302-redirects an
# unauthenticated / age-unconfirmed request to /verify_subscriber or
# /age_confirmation_warning (lands as a 200 on that URL). Treat it as auth,
# not drift — the fix is to rotate cookies / set the age cookie.
if resp.history and (
"/verify_subscriber" in resp.url
or "/age_confirmation_warning" in resp.url
):
raise SubscribeStarAuthError(
f"SubscribeStar redirected to {resp.url} — auth/age wall "
f"(rotate cookies or set the 18+ age cookie; requested {url})",
status_code=resp.status_code,
)
if resp.status_code in (401, 403):
raise SubscribeStarAuthError(
f"SubscribeStar returned HTTP {resp.status_code} — auth rejected "
f"(cookies expired or tier insufficient; {url})",
status_code=resp.status_code,
)
if resp.status_code != 200:
retry_after = None
if resp.status_code == 429:
hdr = resp.headers.get("Retry-After")
if hdr:
try:
retry_after = float(hdr)
except (TypeError, ValueError):
retry_after = None
raise SubscribeStarAPIError(
f"SubscribeStar returned HTTP {resp.status_code} ({url})",
status_code=resp.status_code,
retry_after=retry_after,
)
return resp
def _feed_html(self, url: str) -> str:
"""Page 1: the creator page (full HTML document)."""
resp = self._get(url)
text = resp.text or ""
if any(m in text for m in _LOGIN_MARKERS) and _POST_OPEN not in text:
raise SubscribeStarAuthError(
f"SubscribeStar served a login/age wall instead of the feed "
f"(cookies expired or age cookie missing; {url})"
)
return text
def _loadmore_html(self, url: str) -> str:
"""Subsequent pages: a plain GET (gallery-dl uses no XHR header) to the
`/posts?...` endpoint, which returns JSON {"html": "..."}."""
resp = self._get(url)
try:
payload = resp.json()
except ValueError as exc:
raise SubscribeStarAuthError(
f"SubscribeStar 'load more' returned non-JSON (session expired?; "
f"{url}): {exc}"
) from exc
html = payload.get("html") if isinstance(payload, dict) else None
return html if isinstance(html, str) else ""
# -- parsing -----------------------------------------------------------
@staticmethod
def _post_chunks(html: str) -> list[str]:
"""Slice the page into one chunk per post (between consecutive wrapper
opens). Regex can't match a post's balanced close, so each chunk runs to
the next post's open (or end of fragment) — enough to scope per-post
field extraction."""
starts = [m.start() for m in re.finditer(re.escape(_POST_OPEN), html)]
chunks = []
for i, start in enumerate(starts):
end = starts[i + 1] if i + 1 < len(starts) else len(html)
chunks.append(html[start:end])
return chunks
def _parse_post(self, chunk: str) -> dict | None:
m = _POST_ID_RE.search(chunk)
if not m:
return None
post_id = m.group(1)
# gallery-dl date method: text up to first '</', then after the last '>'
# — handles both plain and <a>-wrapped dates (see _DATE_OPEN comment).
raw_date = _extr(chunk, _DATE_OPEN, "</").rpartition(">")[2]
published = _parse_ss_datetime(raw_date) if raw_date else None
content = _extract_content(chunk)
return {
"id": post_id,
"attributes": {
# SubscribeStar has no title field; the importer synthesizes a
# display title from the body's first line (sidecar util).
"title": "",
"content": content,
"published_at": published,
"post_type": "subscribestar",
},
# Raw chunk retained so extract_media parses the data-gallery manifest
# and post_is_gated can scan for the locked-teaser marker.
"_html": chunk,
}
def _parse_posts(self, html: str) -> list[dict]:
posts = []
for chunk in self._post_chunks(html):
post = self._parse_post(chunk)
if post is not None:
posts.append(post)
if posts:
dated = sum(1 for p in posts if p["attributes"].get("published_at"))
bodied = sum(1 for p in posts if p["attributes"].get("content"))
log.info(
"SubscribeStar parsed %d posts (%d dated, %d with body)",
len(posts), dated, bodied,
)
# Canary for this exact failure class: posts parsed but NONE got a
# date or a body, while the raw markers ARE present → our extraction
# diverged from the live markup (e.g. the <a>-wrapped date bug). Log
# the marker counts so the cause is diagnosable from the worker log
# alone, without re-fetching the authed page.
if dated == 0 or bodied == 0:
log.warning(
"SubscribeStar parse canary: %d posts but dated=%d bodied=%d; "
"raw markers post-date=%d post_content-text=%d data-gallery=%d "
"— extraction likely diverged from the live markup",
len(posts), dated, bodied,
html.count('class="post-date"'),
html.count("post_content-text"),
html.count("data-gallery"),
)
return posts
@staticmethod
def _next_page_href(html: str) -> str | None:
m = _NEXT_PAGE_RE.search(html)
return unescape(m.group(1)) if m else None
def extract_media(self, post: dict, included_index: dict) -> list[MediaItem]:
"""Resolve downloadable media (gallery-dl's `_media_from_post`): the
per-post `data-gallery` JSON manifest (images/videos), PLUS document
attachments (`uploads-docs`) and audio (`uploads-audios`) — some posts
deliver content only through the latter two. `/previews` (locked teaser)
gallery items are skipped. `included_index` is unused (media is inline)."""
chunk = post.get("_html") or ""
base = post.get("_base") or "https://www.subscribestar.com"
post_id = str(post.get("id") or "")
items: list[MediaItem] = []
for gm in _GALLERY_RE.finditer(chunk):
try:
gallery = json.loads(unescape(gm.group(1)))
except (ValueError, TypeError):
continue
if not isinstance(gallery, list):
continue
for it in gallery:
if not isinstance(it, dict):
continue
rel = it.get("url")
if not isinstance(rel, str) or not rel:
continue
# gallery-dl's _media_from_post: a gallery item whose URL is under
# /previews is a locked/blurred TEASER, not the real file — skip it
# (the SubscribeStar analog of the Patreon gated-preview bug #874).
# This is why a locked post yields no downloadable media.
if "/previews" in rel:
continue
url = urljoin(base + "/", rel)
media_id = str(it.get("id") or "")
name = it.get("original_filename")
filename = name if isinstance(name, str) and name else basename_from_url(url)
items.append(
MediaItem(
url=url,
filename=filename,
kind=str(it.get("type") or "image"),
filehash=filehash_from_url(url),
post_id=post_id,
media_id=media_id,
)
)
# Document attachments (uploads-docs → doc_preview blocks): href URL.
docs = _extr(chunk, 'class="uploads-docs"', 'class="post-edit_form"')
for frag in _DOC_SPLIT_RE.split(docs)[1:]:
item = _attachment_item(
frag, base, post_id, kind="attachment",
title_marker='doc_preview-title">', url_attr='href="',
)
if item is not None:
items.append(item)
# Audio attachments (uploads-audios → audio_preview-data blocks): src URL.
audios = _extr(chunk, 'class="uploads-audios"', 'class="post-edit_form"')
for frag in _AUDIO_SPLIT_RE.split(audios)[1:]:
item = _attachment_item(
frag, base, post_id, kind="audio",
title_marker='audio_preview-title">', url_attr='src="',
)
if item is not None:
items.append(item)
return items
@staticmethod
def post_meta(post: dict) -> dict:
"""Title + date for the preview sample. Title is synthesized from the body
(SubscribeStar has no title field)."""
attrs = post.get("attributes") or {}
return {"title": None, "date": attrs.get("published_at")}
@staticmethod
def post_is_gated(post: dict) -> bool:
"""True when the subscriber cannot view this post (locked teaser). #874
"no stub for gated content": the core skips it ENTIRELY (no media, no
post-record), so we don't replicate a hollow teaser in Curator.
SubscribeStar does NOT expose downloadable preview media for locked posts
(no `data-gallery`), so the Patreon "downloaded blurred junk" failure
can't happen here — this gate only prevents capturing an empty teaser
stub. Detect the locked marker conservatively (default to NOT gated when
absent, so we never over-filter an accessible text post). The exact marker
is being confirmed against a live locked sample; until then we gate only on
the explicit lock-overlay class SubscribeStar renders on a paywalled post.
"""
chunk = post.get("_html") or ""
return 'class="post-content-locked"' in chunk or "for-locked_content" in chunk
@staticmethod
def post_record_key(post: dict) -> tuple[str, str] | None:
"""`(ledger_key, post_id)` for a post's seen-ledger entry (the `post:<id>`
synthetic key gates post-record capture through the same ledger as media),
or None when the post has no id."""
pid = post.get("id")
pid = str(pid) if pid is not None else ""
if not pid:
return None
return (f"post:{pid}", pid)
# -- iteration ---------------------------------------------------------
def iter_posts(
self, campaign_id: str, cursor: str | None = None
) -> Iterator[tuple[dict, dict, str | None]]:
"""Yield (post, {}, page_cursor) for every post in the feed.
`campaign_id` is the creator URL. `cursor` is the relative "load more"
href that fetches a page (None → page 1, the creator page HTML). The
yielded `page_cursor` is the href that FETCHED this post's page, so the
core checkpoints a value that re-fetches the same page on resume (matching
the Patreon cursor contract).
"""
base, slug = _split_creator_url(campaign_id)
if not slug:
raise SubscribeStarDriftError(
f"Could not extract a creator slug from {campaign_id!r}"
)
# Same-site Referer (gallery-dl sends root/ on every request).
self._session.headers["Referer"] = f"{base}/"
current = cursor
first_page = True
while True:
page_cursor = current
if current is None:
html = self._feed_html(f"{base}/{slug}")
else:
html = self._loadmore_html(urljoin(base + "/", current))
posts = self._parse_posts(html)
if first_page and not posts and _POST_OPEN not in html:
# Page 1 with no recognizable post wrappers at all: either a brand
# new creator with zero posts, or our scraper is stale. The core's
# body canary catches systematic emptiness across a populated feed;
# here we only raise if the feed container itself is missing.
if 'data-role="posts_container-list"' not in html:
# Report what we actually got — length + the page's identity —
# so a served interstitial (bot challenge / age gate / login)
# is named instead of mislabeled "markup changed".
looks_json = html.lstrip()[:1] in ("{", "[")
raise SubscribeStarDriftError(
f"SubscribeStar feed for {slug!r} had no posts and no "
f"recognizable feed container ({len(html)} bytes"
f"{', looks like JSON' if looks_json else ''}; "
f"{_describe_page(html)})"
)
for post in posts:
post["_base"] = base
yield post, {}, page_cursor
next_href = self._next_page_href(html)
if not next_href:
return
current = next_href
first_page = False
# -- verify ------------------------------------------------------------
def verify_auth(self, campaign_id: str) -> tuple[bool | None, str]:
"""Cheap auth probe: fetch the first feed page and report whether the
credential authenticated, without downloading anything."""
base, slug = _split_creator_url(campaign_id)
if not slug:
return None, f"Couldn't parse a SubscribeStar creator from {campaign_id!r}"
self._session.headers["Referer"] = f"{base}/"
try:
html = self._feed_html(f"{base}/{slug}")
except SubscribeStarAuthError as exc:
return False, f"SubscribeStar rejected the credential — {exc}"
except SubscribeStarAPIError as exc:
return None, f"Couldn't verify (network/HTTP issue): {exc}"
if _POST_OPEN in html or 'data-role="posts_container-list"' in html:
return True, "Credentials valid — the SubscribeStar feed loaded."
return None, "Couldn't verify — SubscribeStar feed shape unrecognized."
@@ -1,206 +0,0 @@
"""Native SubscribeStar media downloader — the SubscribeStar counterpart to
patreon_downloader.
Given a SubscribeStar post and its resolved `MediaItem`s
(subscribestar_client.extract_media), download the media to gallery-dl's on-disk
layout (so existing gallery-dl downloads are recognized on disk and not
re-fetched at cutover), write the post-first sidecars the importer consumes, and
report per-media outcomes.
Simpler than the Patreon downloader: SubscribeStar serves every upload as a
direct file via `/post_uploads?payload=...` (plain GET — no Mux/HLS, no yt-dlp),
and the full post body is already present in the feed HTML (no detail-endpoint
enrichment). PURE: no DB; the seen-skip is an injected predicate.
On-disk layout (matches gallery-dl's subscribestar config
`{date:%Y-%m-%d}_{id}_{title[:40]}` / `{num:>02}_{filename}`): SubscribeStar posts
have no title, so the directory is `<YYYY-MM-DD>_<post_id>_` — the display title is
synthesized by the importer from the body, so the bare dir name is on-disk only.
FC runs on a plain-HTTP homelab; nothing here uses a secure-context Web API.
"""
from __future__ import annotations
import json
import logging
import time
from collections.abc import Callable
from pathlib import Path
import requests
from .native_ingest_common import (
BaseNativeDownloader,
MediaOutcome,
PostRecordOutcome,
post_dir_name,
sanitize_segment,
)
log = logging.getLogger(__name__)
class SubscribeStarDownloader(BaseNativeDownloader):
"""Download resolved SubscribeStar media to gallery-dl's on-disk layout.
Subclasses BaseNativeDownloader for the shared streaming GET (transient-retry +
Range-resume) and validation/quarantine. No video branch (SubscribeStar serves
files directly via /post_uploads) and no detail-fetch (the body is already in
the feed HTML). PURE: no DB."""
def __init__(
self,
images_root: Path,
cookies_path: str | None = None,
*,
validate: bool = True,
rate_limit: float = 0.0,
session: requests.Session | None = None,
):
super().__init__(
images_root, cookies_path, platform="subscribestar",
validate=validate, rate_limit=rate_limit, session=session,
)
# -- public ------------------------------------------------------------
def download_post(
self,
post: dict,
media_items: list,
artist_slug: str,
*,
is_seen: Callable[[object], bool] = lambda m: False,
should_stop: Callable[[], bool] = lambda: False,
recapture: bool = False,
) -> list[MediaOutcome]:
"""Download every media item of one post; return per-item outcomes.
Mirrors PatreonDownloader.download_post (two-tier skip, mid-post time-box,
recapture surfacing) minus the video branch."""
post_dir = self.images_root / artist_slug / "subscribestar" / post_dir_name(post)
outcomes: list[MediaOutcome] = []
for i, media in enumerate(media_items, start=1):
if should_stop():
break
try:
outcomes.append(
self._download_one(
post, media, post_dir, artist_slug, i, is_seen,
recapture=recapture,
)
)
except Exception as exc: # resilient: isolate one item's failure
log.warning(
"SubscribeStar media failed (post %s, item %d): %s",
post.get("id"), i, exc,
)
outcomes.append(
MediaOutcome(media=media, status="error", path=None, error=str(exc))
)
return outcomes
# -- per-item ----------------------------------------------------------
def _download_one(
self,
post: dict,
media,
post_dir: Path,
artist_slug: str,
index: int,
is_seen: Callable[[object], bool],
*,
recapture: bool = False,
) -> MediaOutcome:
seen = is_seen(media)
if seen and not recapture:
return MediaOutcome(media=media, status="skipped_seen", path=None, error=None)
nn = f"{index:02d}"
media_path = post_dir / sanitize_segment(f"{nn}_{media.filename}")
if media_path.exists(): # tier-2: already on disk
return MediaOutcome(
media=media, status="skipped_disk", path=media_path, error=None
)
# recapture: a seen item not on disk is NOT re-downloaded (recovery's job).
if seen:
return MediaOutcome(media=media, status="skipped_seen", path=None, error=None)
post_dir.mkdir(parents=True, exist_ok=True)
if self._rate_limit > 0:
time.sleep(self._rate_limit)
out_path = self._fetch_get(media.url, media_path)
reason, quarantine_dest = self._validate_path(out_path, artist_slug, media.url)
if reason is not None:
return MediaOutcome(
media=media, status="quarantined", path=quarantine_dest, error=reason,
)
self._write_sidecar(post, out_path, source_url=media.url)
return MediaOutcome(media=media, status="downloaded", path=out_path, error=None)
# The plain-GET streaming path (_fetch_get / _fetch_to_file) and
# _validate_path are inherited from BaseNativeDownloader.
# -- sidecar -----------------------------------------------------------
def _write_sidecar(
self, post: dict, media_path: Path, *, source_url: str | None = None
) -> Path:
"""Per-media sidecar — post-first (#856): image identity ONLY
(category/id/source_url). The post body/links live solely in _post.json."""
return self._write_sidecar_data(
post, media_path.with_suffix(".json"), source_url=source_url, minimal=True,
)
def _write_sidecar_data(
self, post: dict, sidecar_path: Path, *, source_url: str | None = None,
minimal: bool = False,
) -> Path:
"""Serialize the post's metadata. minimal=True → per-media sidecar (image
identity only); else the full post record (body/title/date/url)."""
if minimal:
data = {"category": "subscribestar", "id": str(post.get("id") or "")}
if source_url:
data["source_url"] = source_url
sidecar_path.write_text(json.dumps(data, indent=2))
return sidecar_path
attrs = post.get("attributes") or {}
content = attrs.get("content")
# SubscribeStar synthesizes the post permalink from the id (matches the
# platforms/subscribestar.py derive_post_url helper).
pid = str(post.get("id") or "")
data = {
"category": "subscribestar",
"id": pid,
"post_id": pid, # ensures derive_post_url has its key on the native path
"title": attrs.get("title") if isinstance(attrs.get("title"), str) else "",
"content": content if isinstance(content, str) else "",
"published_at": attrs.get("published_at"),
}
if source_url:
data["source_url"] = source_url
sidecar_path.write_text(json.dumps(data, indent=2))
return sidecar_path
def write_post_record(self, post: dict, artist_slug: str) -> PostRecordOutcome:
"""Write the post-first `_post.json` (body/links/metadata) — the sole
writer of the post record on the native path. SubscribeStar's body is
already in the feed HTML, so no detail-fetch is needed."""
attrs = post.get("attributes") or {}
title = attrs.get("title") if isinstance(attrs.get("title"), str) else None
post_type = attrs.get("post_type") if isinstance(attrs.get("post_type"), str) else None
pid = str(post.get("id") or "")
if not pid:
return PostRecordOutcome(
path=None, post_type=post_type, title=title, body_chars=0,
)
post_dir = self.images_root / artist_slug / "subscribestar" / post_dir_name(post)
post_dir.mkdir(parents=True, exist_ok=True)
path = self._write_sidecar_data(post, post_dir / "_post.json")
body = attrs.get("content")
body_chars = len(body) if isinstance(body, str) else 0
return PostRecordOutcome(
path=path, post_type=post_type, title=title, body_chars=body_chars,
)
@@ -1,109 +0,0 @@
"""Native SubscribeStar ingester — the SubscribeStar ADAPTER over the
platform-agnostic core (`ingest_core.Ingester`).
Thin counterpart to patreon_ingester: wires the SubscribeStar client/downloader/
ledger models/constraints/key into the core and supplies the SubscribeStar
failure mapping. The three modes (tick / backfill / recovery / recapture), the
seen + dead-letter ledgers, cursor checkpointing, and the post-first capture all
live in the core — identical to Patreon. `download_service.download_source`
drives `SubscribeStarIngester.run` exactly as it drives the Patreon one.
`campaign_id` is the creator URL (the client derives host + slug from it), so no
campaign-id resolver is needed. FC runs on a plain-HTTP homelab; nothing here
uses a secure-context Web API.
"""
from __future__ import annotations
import asyncio
import logging
from collections.abc import Callable
from pathlib import Path
from ..models import SubscribeStarFailedMedia, SubscribeStarSeenMedia
from .ingest_core import DEAD_LETTER_THRESHOLD, Ingester
from .subscribestar_client import MediaItem, SubscribeStarAPIError, SubscribeStarClient
from .subscribestar_downloader import SubscribeStarDownloader
__all__ = [
"DEAD_LETTER_THRESHOLD",
"SubscribeStarIngester",
"_ledger_key",
"verify_subscribestar_credential",
]
log = logging.getLogger(__name__)
_LEDGER_KEY_MAX = 128
def _ledger_key(media: MediaItem) -> str:
"""Stable per-media identity for the cross-run seen-ledger. SubscribeStar's
full-res URL is an opaque `/post_uploads?payload=...` (no content hash), so
`media.filehash` is normally None and the stable proxy is the gallery item id
scoped to its post: `<post_id>:<media_id>`. Bounded to the column width."""
if media.filehash:
return media.filehash
return f"{media.post_id}:{media.media_id}"[:_LEDGER_KEY_MAX]
class SubscribeStarIngester(Ingester):
"""Walk a SubscribeStar creator's posts, download unseen media, return a
`DownloadResult`. A thin adapter over `ingest_core.Ingester`; `client` /
`downloader` are injectable seams so unit tests run without network."""
def __init__(
self,
images_root: Path,
cookies_path: str | None,
session_factory: Callable[[], object],
*,
validate: bool = True,
rate_limit: float = 0.0,
request_sleep: float = 0.0,
client: SubscribeStarClient | None = None,
downloader: SubscribeStarDownloader | None = None,
):
self.images_root = Path(images_root)
self.cookies_path = str(cookies_path) if cookies_path else None
resolved_client = (
client
if client is not None
else SubscribeStarClient(cookies_path, request_sleep=request_sleep)
)
resolved_downloader = (
downloader
if downloader is not None
else SubscribeStarDownloader(
self.images_root, cookies_path, validate=validate, rate_limit=rate_limit,
)
)
super().__init__(
client=resolved_client,
downloader=resolved_downloader,
session_factory=session_factory,
seen_model=SubscribeStarSeenMedia,
failed_model=SubscribeStarFailedMedia,
seen_constraint="uq_subscribestar_seen_media_source_id",
failed_constraint="uq_subscribestar_failed_media_source_id",
ledger_key=_ledger_key,
platform="subscribestar",
error_base=SubscribeStarAPIError,
# API_DRIFT message phrasing; the base Ingester._failure_result owns
# the auth/drift/HTTP→error_type mapping (shared across platforms).
drift_label="SubscribeStar markup",
)
async def verify_subscribestar_credential(
url: str,
cookies_path: str | None,
overrides: dict | None,
) -> tuple[bool | None, str]:
"""Native SubscribeStar credential probe — fetches ONE feed page via
SubscribeStarClient.verify_auth. `campaign_id` is just the creator URL (no
resolver). Returns the uniform `(ok, message)` contract so
download_backends.verify_credential treats it like the gallery-dl probe."""
client = SubscribeStarClient(cookies_path)
loop = asyncio.get_running_loop()
return await loop.run_in_executor(None, client.verify_auth, url)
+17 -48
View File
@@ -53,28 +53,12 @@ class TagDirectoryService:
raise ValueError("limit must be between 1 and 200")
fandom = aliased(Tag)
member = aliased(Tag)
# image_count aggregates the tag's images INCLUDING, for a fandom, every
# image carrying one of its characters. `member` is the tag actually on
# the image; an image counts for the outer Tag if that applied tag IS the
# Tag (direct) OR is a character whose fandom_id is the Tag (the fandom
# leg). Both legs correlate the OUTER Tag.id at a SINGLE level — a nested
# `tag_id IN (SELECT ... WHERE member.fandom_id == Tag.id)` does NOT
# correlate and silently counts every fandom-character globally (~every
# tag collapses to the same inflated number). DISTINCT so an image with
# the fandom AND ≥1 of its characters counts once; a correlated scalar
# subquery keeps it 1 row per tag with no group_by.
count_col = (
select(func.count(image_tag.c.image_record_id.distinct()))
.select_from(image_tag)
.join(member, member.id == image_tag.c.tag_id)
.where(or_(image_tag.c.tag_id == Tag.id, member.fandom_id == Tag.id))
.correlate(Tag)
.scalar_subquery()
.label("image_count")
)
stmt = select(Tag, fandom.name.label("fandom_name"), count_col).outerjoin(
fandom, Tag.fandom_id == fandom.id
count_col = func.count(image_tag.c.image_record_id).label("image_count")
stmt = (
select(Tag, fandom.name.label("fandom_name"), count_col)
.outerjoin(fandom, Tag.fandom_id == fandom.id)
.outerjoin(image_tag, image_tag.c.tag_id == Tag.id)
.group_by(Tag.id, fandom.name)
)
if kind is not None:
stmt = stmt.where(Tag.kind == kind)
@@ -117,34 +101,19 @@ class TagDirectoryService:
async def _previews(self, tag_ids: list[int]) -> dict[int, list[str]]:
if not tag_ids:
return {}
# Preview pool per card = images tagged with the card's tag DIRECTLY,
# UNION images carrying a character whose fandom_id is the card's tag (so
# a fandom card previews its characters' images). UNION dedups the
# overlap. For non-fandom cards the via-fandom leg is empty.
member = aliased(Tag)
direct = select(
image_tag.c.tag_id.label("card_id"),
image_tag.c.image_record_id.label("image_record_id"),
).where(image_tag.c.tag_id.in_(tag_ids))
via_fandom = (
select(
member.fandom_id.label("card_id"),
image_tag.c.image_record_id.label("image_record_id"),
)
.select_from(image_tag)
.join(member, member.id == image_tag.c.tag_id)
.where(member.fandom_id.in_(tag_ids))
)
scoped = direct.union(via_fandom).subquery()
rn = func.row_number().over(
partition_by=scoped.c.card_id,
order_by=scoped.c.image_record_id.desc(),
partition_by=image_tag.c.tag_id,
order_by=image_tag.c.image_record_id.desc(),
).label("rn")
sub = select(
scoped.c.card_id.label("tag_id"),
scoped.c.image_record_id.label("image_record_id"),
rn,
).subquery()
sub = (
select(
image_tag.c.tag_id.label("tag_id"),
image_tag.c.image_record_id.label("image_record_id"),
rn,
)
.where(image_tag.c.tag_id.in_(tag_ids))
.subquery()
)
stmt = (
select(
sub.c.tag_id,
+2 -48
View File
@@ -1,10 +1,4 @@
"""Shared tag-with-fandom query columns, serialization, and membership.
`image_in_tag_scope`/`image_in_any_tag_scope` are the SINGLE source of truth for
"which images belong to a tag" — a fandom tag also owns the images of its
characters (Tag.fandom_id), derived at query time rather than materialized. The
gallery scope, directory count/previews, and the cleanup impact count all route
through them so a fandom can never under-count its characters' images.
"""Shared tag-with-fandom query columns + serialization.
Resolving a character tag's fandom NAME via a Tag self-join (Tag.fandom_id -> the
fandom Tag) and serializing the canonical
@@ -17,47 +11,7 @@ TagDirectoryService selects the FULL Tag ORM plus an image-count aggregate (a
different select shape), so it keeps its own variant — not folded in here.
"""
from sqlalchemy import exists, or_, select
from sqlalchemy.orm import aliased
from ..models import ImageRecord, Tag
from ..models.tag import image_tag
def _fandom_member_char_ids(tids):
"""Subquery of character tag ids owned by any fandom in `tids` (via
Tag.fandom_id). Empty for non-fandom tids, so callers can OR it in
unconditionally — a fandom aggregates its characters' images, every other
kind degrades to direct-only."""
char = aliased(Tag)
return select(char.id).where(char.fandom_id.in_(tids))
def image_in_tag_scope(tid):
"""Correlated EXISTS: the current ImageRecord belongs to tag `tid` — either
tagged with it DIRECTLY, or (when `tid` is a fandom) carrying a character
whose fandom_id == `tid`. This is the SINGLE membership predicate shared by
the gallery scope, the directory count/previews, and the cleanup impact
count, so a fandom never under-counts the images of its characters."""
return exists().where(
image_tag.c.image_record_id == ImageRecord.id,
or_(
image_tag.c.tag_id == tid,
image_tag.c.tag_id.in_(_fandom_member_char_ids([tid])),
),
)
def image_in_any_tag_scope(tids):
"""Like `image_in_tag_scope` but for an OR-group: the image carries AT LEAST
ONE of `tids` directly, or a character of any fandom in `tids`."""
return exists().where(
image_tag.c.image_record_id == ImageRecord.id,
or_(
image_tag.c.tag_id.in_(tids),
image_tag.c.tag_id.in_(_fandom_member_char_ids(tids)),
),
)
from ..models import Tag
def fandom_join_alias():
+1 -159
View File
@@ -9,7 +9,7 @@ from sqlalchemy import and_, case, exists, func, select, text, update
from sqlalchemy.dialects.postgresql import insert as pg_insert
from sqlalchemy.ext.asyncio import AsyncSession
from ..models import HeadMetric, Tag, TagHead, TagKind, image_tag
from ..models import Tag, TagKind, image_tag
from ..models.tag_allowlist import TagAllowlist
from ..models.tag_reference_embedding import TagReferenceEmbedding
from .db_helpers import get_or_create
@@ -79,23 +79,6 @@ class MergeResult:
source_deleted: bool
@dataclass(frozen=True)
class MergePreview:
"""Non-mutating projection of what merge(source→target) would do (#8).
Shares the apply's predicates (rule 93) so the preview can't drift."""
source_id: int
source_name: str
target_id: int
target_name: str
compatible: bool # same kind + fandom — would apply succeed?
images_moving: int # source links that move (image lacks target)
images_already_on_target: int # source links dropped (image has target)
source_total: int # images carrying source
series_pages: int # series pages repointed
will_alias: bool # source name kept as a protective alias
sample_thumbnails: list[str] # a few of the moving images
class TagService:
def __init__(self, session: AsyncSession):
self.session = session
@@ -215,18 +198,6 @@ class TagService:
async def add_to_image(self, image_id: int, tag_id: int, source: str = "manual") -> None:
"""Idempotent: re-adding an existing tag does nothing."""
# A genuinely-new MANUAL add of a tag that already has a head is an
# UNDER-FIRE signal — the auto-system should have caught it (#114 obs).
is_new = source == "manual" and (
await self.session.execute(
select(image_tag.c.tag_id).where(
and_(
image_tag.c.image_record_id == image_id,
image_tag.c.tag_id == tag_id,
)
)
)
).first() is None
stmt = pg_insert(image_tag).values(
image_record_id=image_id, tag_id=tag_id, source=source
)
@@ -234,22 +205,8 @@ class TagService:
index_elements=["image_record_id", "tag_id"]
)
await self.session.execute(stmt)
if is_new:
await self._note_under_fire(tag_id)
async def remove_from_image(self, image_id: int, tag_id: int) -> None:
# Removing an auto-applied (source='head_auto') tag is a MISFIRE — read
# the source BEFORE deleting, since it's lost with the row (#114 obs).
src = (
await self.session.execute(
select(image_tag.c.source).where(
and_(
image_tag.c.image_record_id == image_id,
image_tag.c.tag_id == tag_id,
)
)
)
).scalar_one_or_none()
await self.session.execute(
image_tag.delete().where(
and_(
@@ -258,31 +215,6 @@ class TagService:
)
)
)
if src == "head_auto":
await self._bump_metric(tag_id, "n_misfires")
async def _note_under_fire(self, tag_id: int) -> None:
"""Count an under-fire only when the tag actually has a head."""
has_head = (
await self.session.execute(
select(TagHead.tag_id).where(TagHead.tag_id == tag_id)
)
).first() is not None
if has_head:
await self._bump_metric(tag_id, "n_underfires")
async def _bump_metric(self, tag_id: int, column: str) -> None:
"""Increment a HeadMetric counter (upsert), keyed by tag so it survives
head retrain/prune."""
col = HeadMetric.__table__.c[column]
await self.session.execute(
pg_insert(HeadMetric)
.values(tag_id=tag_id, **{column: 1})
.on_conflict_do_update(
index_elements=["tag_id"],
set_={column: col + 1, "updated_at": func.now()},
)
)
async def list_for_image(self, image_id: int) -> Sequence:
"""Tags on an image, ordered (kind, name). Each row carries the fandom's
@@ -444,96 +376,6 @@ class TagService:
await self.session.flush()
return tag
async def merge_preview(
self, source_id: int, target_id: int
) -> MergePreview:
"""Read-only dry-run of merge(source→target): the same counts the apply
would produce, plus a thumbnail sample of the images that move. Raises
TagValidationError only for missing/self-merge (so the UI can render a
kind/fandom-mismatch warning rather than a hard error)."""
if source_id == target_id:
raise TagValidationError("Cannot merge a tag into itself")
source = await self.session.get(Tag, source_id)
target = await self.session.get(Tag, target_id)
if source is None or target is None:
raise TagValidationError("Tag not found")
compatible = source.kind == target.kind and (
(source.fandom_id or 0) == (target.fandom_id or 0)
)
# Mirrors _repoint_image_tags: links whose image already has target are
# dropped (already_on_target); the rest move.
target_images = select(image_tag.c.image_record_id).where(
image_tag.c.tag_id == target_id
)
moving = await self.session.scalar(
select(func.count())
.select_from(image_tag)
.where(
image_tag.c.tag_id == source_id,
image_tag.c.image_record_id.notin_(target_images),
)
)
already = await self.session.scalar(
select(func.count())
.select_from(image_tag)
.where(
image_tag.c.tag_id == source_id,
image_tag.c.image_record_id.in_(target_images),
)
)
from ..models.series_page import SeriesPage
series_pages = await self.session.scalar(
select(func.count())
.select_from(SeriesPage)
.where(SeriesPage.series_tag_id == source_id)
)
will_alias = await self._keep_as_alias(source_id)
sample = await self._merge_sample_thumbnails(source_id, target_images)
return MergePreview(
source_id=source_id,
source_name=source.name,
target_id=target_id,
target_name=target.name,
compatible=compatible,
images_moving=moving or 0,
images_already_on_target=already or 0,
source_total=(moving or 0) + (already or 0),
series_pages=series_pages or 0,
will_alias=will_alias,
sample_thumbnails=sample,
)
async def _merge_sample_thumbnails(
self, source_id: int, target_images
) -> list[str]:
"""Up to 6 thumbnails of the images that WOULD move (carry source, not
target) — a sanity check that the merge targets the right content."""
from ..models import ImageRecord
from .gallery_service import thumbnail_url
rows = (
await self.session.execute(
select(
ImageRecord.thumbnail_path,
ImageRecord.sha256,
ImageRecord.mime,
)
.join(image_tag, image_tag.c.image_record_id == ImageRecord.id)
.where(
image_tag.c.tag_id == source_id,
image_tag.c.image_record_id.notin_(target_images),
)
.limit(6)
)
).all()
return [thumbnail_url(tp, sha, mime) for tp, sha, mime in rows]
async def merge(self, source_id: int, target_id: int) -> MergeResult:
"""Transactionally repoint every FK from source→target, optionally
keep source's name as a tagger alias, delete source. Atomic: any
+6 -9
View File
@@ -41,19 +41,16 @@ IMAGES_ROOT = Path("/images")
# After this many failed attempts a link is dead-lettered (skipped by routine
# sweeps; an operator recovery still re-attempts). Mirrors the ingester ledger.
DEAD_LETTER_THRESHOLD = 3
# Per-fetch read + total budgets now live in external_fetch (a short read
# timeout fails a stalled host fast; a generous total caps a big-but-flowing
# download). The celery soft/hard time_limit is the outer backstop above those.
# Per-fetch wall-clock budget — films/packs are large; the celery time_limit is
# the backstop above this.
_FETCH_TIMEOUT = 3000.0
# Links enqueued per sweep — bounds the burst when a big backfill records many.
_SWEEP_BATCH = 50
# Dead rows older than this are pruned (retention).
_RETENTION_DAYS = 30
# Per-host serialize lock. TTL is a safety net for a worker that dies holding it
# (normal completion/error releases it in `finally`); sized just past the fetch
# total budget (30 min) so a dead worker can't wedge a host's links much longer
# than one fetch would have taken.
# Per-host serialize lock.
_LOCK_PREFIX = "fc:extdl_lock:"
_LOCK_TTL = 2400
_LOCK_TTL = 3600
_SERIALIZE_COUNTDOWN = 120
_MAX_SERIALIZE_WAITS = 30
@@ -180,7 +177,7 @@ def fetch_external_link(self, link_id: int, _serialize_waits: int = 0) -> dict:
/ str(post.external_post_id) / str(link_id)
)
try:
result = fetch_external(host, url, post_dir)
result = fetch_external(host, url, post_dir, timeout=_FETCH_TIMEOUT)
with SessionLocal() as session:
link = session.get(ExternalLink, link_id)
if not result.ok:
-223
View File
@@ -13,15 +13,12 @@ from ..celery_app import celery
from ..models import (
BackupRun,
DownloadEvent,
HeadAutoApplyRun,
HeadTrainingRun,
ImageRecord,
ImportBatch,
ImportSettings,
ImportTask,
LibraryAuditRun,
Source,
TagEvalRun,
TaskRun,
)
from ..utils.phash import compute_phash
@@ -96,15 +93,6 @@ BACKUP_DB_STALL_THRESHOLD_MINUTES = 40
# Library audit: scan_library_for_rule has time_limit=7500s (2h5m).
# 2h15m gives a 10-min buffer.
LIBRARY_AUDIT_STALL_THRESHOLD_MINUTES = 135
# tag-eval (#1130) has a 30-min soft limit; flag a run with no progress past 40.
TAG_EVAL_STALL_THRESHOLD_MINUTES = 40
TAG_EVAL_KEEP_RUNS = 20
# head training (#114) has a 60-min soft limit; flag no-progress past 75.
HEAD_TRAINING_STALL_THRESHOLD_MINUTES = 75
HEAD_TRAINING_KEEP_RUNS = 20
# head auto-apply (#114) shares the 60-min soft limit; flag past 75.
HEAD_AUTO_APPLY_STALL_THRESHOLD_MINUTES = 75
HEAD_AUTO_APPLY_KEEP_RUNS = 20
# Import batches finalize only after every child ImportTask hits a
# terminal state. The recovery sweep targets the case where every
# task is done but the batch never got its closing UPDATE
@@ -159,15 +147,6 @@ TASK_STUCK_THRESHOLD_MINUTES: dict[str, int] = {
"backend.app.tasks.backup.restore_images_task": 420,
# Library audit scans the full library — 2h hard limit.
"backend.app.tasks.library_audit.scan_library_for_rule": 130,
# External file-host fetches (mega/gdrive/film packs) can run to the task's
# 60-min hard limit (time_limit=3600) — the fetcher's own read/total budgets
# (external_fetch) cap a single fetch below that, but this stays the outer
# backstop. Without an override these healthy in-flight fetches were
# phantom-flagged 'RecoverySweep' before their own timeout/error could
# surface (operator-flagged 2026-06-17, target 414 swept at 6.6min). A
# task-name override beats the queue threshold whatever queue the row records
# (it recorded 'default' before the celery_signals fix → download). 65 = 60+5.
"backend.app.tasks.external.fetch_external_link": 65,
}
@@ -721,208 +700,6 @@ def recover_stalled_library_audit_runs() -> int:
return recovered
@celery.task(name="backend.app.tasks.maintenance.recover_stalled_tag_eval_runs")
def recover_stalled_tag_eval_runs() -> int:
"""Flip TagEvalRun rows stuck in 'running' past the stall threshold to
'error', and prune old runs to the last TAG_EVAL_KEEP_RUNS (retention,
rule 89). Runs every 5 min on the maintenance lane; no-op when idle."""
SessionLocal = _sync_session_factory()
now = datetime.now(UTC)
cutoff = now - timedelta(minutes=TAG_EVAL_STALL_THRESHOLD_MINUTES)
with SessionLocal() as session:
result = session.execute(
update(TagEvalRun)
.where(TagEvalRun.status == "running")
.where(
func.coalesce(TagEvalRun.last_progress_at, TagEvalRun.started_at)
< cutoff
)
.values(
status="error", finished_at=now,
error=(
f"stranded by recovery sweep (no progress for "
f"{TAG_EVAL_STALL_THRESHOLD_MINUTES} min)"
),
)
)
# Retention: keep only the most recent N runs.
keep = session.execute(
select(TagEvalRun.id).order_by(TagEvalRun.id.desc())
.limit(TAG_EVAL_KEEP_RUNS)
).scalars().all()
if keep:
session.execute(
delete(TagEvalRun).where(TagEvalRun.id.not_in(keep))
)
session.commit()
recovered = result.rowcount or 0
if recovered:
log.info("recover_stalled_tag_eval_runs: recovered %d rows", recovered)
return recovered
@celery.task(name="backend.app.tasks.maintenance.recover_stalled_head_training_runs")
def recover_stalled_head_training_runs() -> int:
"""Flip HeadTrainingRun rows stuck in 'running' past the stall threshold to
'error', and prune old runs to the last HEAD_TRAINING_KEEP_RUNS (retention,
rule 89). Runs every 5 min on the maintenance lane; no-op when idle."""
SessionLocal = _sync_session_factory()
now = datetime.now(UTC)
cutoff = now - timedelta(minutes=HEAD_TRAINING_STALL_THRESHOLD_MINUTES)
with SessionLocal() as session:
result = session.execute(
update(HeadTrainingRun)
.where(HeadTrainingRun.status == "running")
.where(
func.coalesce(
HeadTrainingRun.last_progress_at, HeadTrainingRun.started_at
)
< cutoff
)
.values(
status="error", finished_at=now,
error=(
f"stranded by recovery sweep (no progress for "
f"{HEAD_TRAINING_STALL_THRESHOLD_MINUTES} min)"
),
)
)
keep = session.execute(
select(HeadTrainingRun.id).order_by(HeadTrainingRun.id.desc())
.limit(HEAD_TRAINING_KEEP_RUNS)
).scalars().all()
if keep:
session.execute(
delete(HeadTrainingRun).where(HeadTrainingRun.id.not_in(keep))
)
session.commit()
recovered = result.rowcount or 0
if recovered:
log.info(
"recover_stalled_head_training_runs: recovered %d rows", recovered
)
return recovered
@celery.task(name="backend.app.tasks.maintenance.recover_stalled_head_auto_apply_runs")
def recover_stalled_head_auto_apply_runs() -> int:
"""Flip stalled HeadAutoApplyRun 'running' rows to 'error' + prune to the
last HEAD_AUTO_APPLY_KEEP_RUNS (retention, rule 89). 5-min maintenance lane."""
SessionLocal = _sync_session_factory()
now = datetime.now(UTC)
cutoff = now - timedelta(minutes=HEAD_AUTO_APPLY_STALL_THRESHOLD_MINUTES)
with SessionLocal() as session:
result = session.execute(
update(HeadAutoApplyRun)
.where(HeadAutoApplyRun.status == "running")
.where(
func.coalesce(
HeadAutoApplyRun.last_progress_at, HeadAutoApplyRun.started_at
)
< cutoff
)
.values(
status="error", finished_at=now,
error=(
f"stranded by recovery sweep (no progress for "
f"{HEAD_AUTO_APPLY_STALL_THRESHOLD_MINUTES} min)"
),
)
)
keep = session.execute(
select(HeadAutoApplyRun.id).order_by(HeadAutoApplyRun.id.desc())
.limit(HEAD_AUTO_APPLY_KEEP_RUNS)
).scalars().all()
if keep:
session.execute(
delete(HeadAutoApplyRun).where(HeadAutoApplyRun.id.not_in(keep))
)
session.commit()
recovered = result.rowcount or 0
if recovered:
log.info(
"recover_stalled_head_auto_apply_runs: recovered %d rows", recovered
)
return recovered
# Keep ~6 months of daily head-metric snapshots (enough to see tuning trends).
HEAD_METRICS_SNAPSHOT_RETENTION_DAYS = 180
@celery.task(name="backend.app.tasks.maintenance.snapshot_head_metrics")
def snapshot_head_metrics() -> int:
"""Daily per-concept observability point (#114): record each head-bearing
concept's auto-applied volume, cumulative misfires/under-fires, and the
head's measured quality — the time-series the operator tunes from. Prunes
points older than the retention window."""
from ..models import (
HeadMetric,
HeadMetricsSnapshot,
Tag,
TagHead,
)
from ..models.tag import image_tag
SessionLocal = _sync_session_factory()
now = datetime.now(UTC)
with SessionLocal() as session:
heads = {
r.tag_id: r for r in session.execute(
select(
TagHead.tag_id, TagHead.ap, TagHead.precision_cv,
TagHead.recall, TagHead.n_pos,
)
)
}
metrics = {
r.tag_id: r for r in session.execute(
select(
HeadMetric.tag_id, HeadMetric.n_misfires, HeadMetric.n_underfires
)
)
}
# .all() first: dict() of a bare Result tries the mapping protocol (a
# Result exposes .keys()) and subscripts it, which fails.
applied = dict(
session.execute(
select(image_tag.c.tag_id, func.count())
.where(image_tag.c.source == "head_auto")
.group_by(image_tag.c.tag_id)
).all()
)
tag_ids = set(heads) | set(metrics)
if not tag_ids:
return 0
names = dict(
session.execute(
select(Tag.id, Tag.name).where(Tag.id.in_(tag_ids))
).all()
)
for tid in tag_ids:
h = heads.get(tid)
m = metrics.get(tid)
session.add(HeadMetricsSnapshot(
tag_id=tid, name=names.get(tid, str(tid)),
snapshot_at=now,
n_auto_applied=applied.get(tid, 0),
n_misfires=m.n_misfires if m else 0,
n_underfires=m.n_underfires if m else 0,
ap=h.ap if h else None,
precision_cv=h.precision_cv if h else None,
recall=h.recall if h else None,
n_pos=h.n_pos if h else None,
))
session.execute(
delete(HeadMetricsSnapshot).where(
HeadMetricsSnapshot.snapshot_at
< now - timedelta(days=HEAD_METRICS_SNAPSHOT_RETENTION_DAYS)
)
)
session.commit()
return len(tag_ids)
@celery.task(name="backend.app.tasks.maintenance.recover_stalled_import_batches")
def recover_stalled_import_batches() -> int:
"""Finalize ImportBatch rows stuck in running past the hard limit
-397
View File
@@ -538,400 +538,3 @@ def recompute_centroids(self) -> int:
for tid in drifted:
recompute_centroid.delay(tid)
return len(drifted)
@celery.task(
name="backend.app.tasks.ml.tag_eval_run",
bind=True,
# The head-vs-centroid eval (#1130) loads embeddings + fits sklearn heads
# for several concepts — minutes, not seconds. Runs on the ml queue because
# only that worker has numpy/scikit-learn.
soft_time_limit=1800, time_limit=2100,
)
def tag_eval_run(self, run_id: int) -> str:
"""Compute the eval report into the persisted TagEvalRun row so it survives
navigation (the admin card rehydrates from the row, not transient state)."""
from datetime import UTC, datetime
from ..models import TagEvalRun
from ..services.ml.tag_eval import run_eval
SessionLocal = _sync_session_factory()
with SessionLocal() as session:
run = session.get(TagEvalRun, run_id)
if run is None:
return "missing"
run.last_progress_at = datetime.now(UTC)
session.commit()
try:
report = run_eval(session, run.params)
except SoftTimeLimitExceeded:
run.status = "error"
run.error = "timed out"
run.finished_at = datetime.now(UTC)
session.commit()
raise
except Exception as exc:
log.exception("tag_eval_run %d failed", run_id)
run.status = "error"
run.error = str(exc)
run.finished_at = datetime.now(UTC)
session.commit()
return "error"
run.report = report
run.status = "ready"
run.finished_at = datetime.now(UTC)
session.commit()
return "ready"
@celery.task(
name="backend.app.tasks.ml.train_heads",
bind=True,
# Trains a logistic-regression head per eligible concept over stored SigLIP
# embeddings — minutes for a full library. Runs on the ml queue (only that
# worker has scikit-learn). Commits per head so a kill leaves progress.
soft_time_limit=3600, time_limit=3900,
)
def train_heads(self, run_id: int) -> str:
"""(Re)train all eligible concept heads into tag_head, tracked by the
HeadTrainingRun row so the admin card shows live + historical status."""
from datetime import UTC, datetime
from ..models import HeadTrainingRun
from ..services.ml.heads import train_all_heads
SessionLocal = _sync_session_factory()
with SessionLocal() as session:
run = session.get(HeadTrainingRun, run_id)
if run is None:
return "missing"
run.last_progress_at = datetime.now(UTC)
session.commit()
try:
result = train_all_heads(session, run.params, run)
except SoftTimeLimitExceeded:
run.status = "error"
run.error = "timed out"
run.finished_at = datetime.now(UTC)
session.commit()
raise
except Exception as exc:
log.exception("train_heads %d failed", run_id)
run.status = "error"
run.error = str(exc)
run.finished_at = datetime.now(UTC)
session.commit()
return "error"
run.n_trained = result["n_trained"]
run.n_skipped = result["n_skipped"]
run.status = "ready"
run.finished_at = datetime.now(UTC)
session.commit()
return "ready"
@celery.task(name="backend.app.tasks.ml.scheduled_train_heads")
def scheduled_train_heads() -> str:
"""Nightly passive retrain (#114): fold the day's accepts/rejects + any
newly-eligible concepts into the heads without the operator clicking. Skips
if a run is already in flight (one at a time). Creates + COMMITS the run row
before dispatching so the ml-queue worker can always find it."""
from datetime import UTC, datetime
from sqlalchemy import select as sa_select
from ..models import HeadTrainingRun
SessionLocal = _sync_session_factory()
with SessionLocal() as session:
running = session.execute(
sa_select(HeadTrainingRun.id).where(HeadTrainingRun.status == "running")
).scalar_one_or_none()
if running is not None:
return "already running"
run = HeadTrainingRun(
params={"source": "scheduled"}, status="running",
last_progress_at=datetime.now(UTC),
)
session.add(run)
session.commit()
run_id = run.id
train_heads.delay(run_id)
return "dispatched"
@celery.task(
name="backend.app.tasks.ml.apply_head_tags",
bind=True,
# Scores the whole library against the graduated heads and applies their
# tags (or, dry_run, just counts). Streams embeddings in chunks; numpy only,
# but ml queue keeps it off the API workers. Commits per chunk.
soft_time_limit=3600, time_limit=3900,
)
def apply_head_tags(self, run_id: int) -> str:
"""Run an earned-auto-apply sweep into the persisted HeadAutoApplyRun row."""
from datetime import UTC, datetime
from ..models import HeadAutoApplyRun
from ..services.ml.heads import auto_apply_sweep
SessionLocal = _sync_session_factory()
with SessionLocal() as session:
run = session.get(HeadAutoApplyRun, run_id)
if run is None:
return "missing"
run.last_progress_at = datetime.now(UTC)
session.commit()
try:
result = auto_apply_sweep(session, run, run.dry_run)
except SoftTimeLimitExceeded:
run.status = "error"
run.error = "timed out"
run.finished_at = datetime.now(UTC)
session.commit()
raise
except Exception as exc:
log.exception("apply_head_tags %d failed", run_id)
run.status = "error"
run.error = str(exc)
run.finished_at = datetime.now(UTC)
session.commit()
return "error"
run.n_applied = result["n_applied"]
run.report = {"concepts": result["concepts"]}
run.status = "ready"
run.finished_at = datetime.now(UTC)
session.commit()
return "ready"
@celery.task(name="backend.app.tasks.ml.scheduled_apply_head_tags")
def scheduled_apply_head_tags() -> str:
"""Daily passive auto-apply sweep (#114) — only when the master switch is on.
Skips if a sweep is already in flight. Creates + COMMITS the run before
dispatching so the worker always finds it."""
from datetime import UTC, datetime
from sqlalchemy import select as sa_select
from ..models import HeadAutoApplyRun, MLSettings
SessionLocal = _sync_session_factory()
with SessionLocal() as session:
enabled = session.execute(
sa_select(MLSettings.head_auto_apply_enabled).where(MLSettings.id == 1)
).scalar_one_or_none()
if not enabled:
return "disabled"
running = session.execute(
sa_select(HeadAutoApplyRun.id).where(HeadAutoApplyRun.status == "running")
).scalar_one_or_none()
if running is not None:
return "already running"
run = HeadAutoApplyRun(
dry_run=False, params={"dry_run": False, "source": "scheduled"},
status="running", last_progress_at=datetime.now(UTC),
)
session.add(run)
session.commit()
run_id = run.id
apply_head_tags.delay(run_id)
return "dispatched"
@celery.task(name="backend.app.tasks.ml.enqueue_gpu_backfill")
def enqueue_gpu_backfill(task_name: str) -> int:
"""Enqueue a gpu_job for every image that still needs `task_name` (one
INSERT…SELECT, so it scales to a full library). The desktop agent drains the
queue over HTTP. Returns the number enqueued.
'siglip' gates on the RESULT (no concept region yet) rather than on a prior
job, so it picks up the back-catalogue of images that were CCIP-embedded
before concept crops existed, and retries images whose concept embed failed —
without re-touching their figure/CCIP regions."""
from sqlalchemy import exists, insert, literal
from sqlalchemy import select as sa_select
from ..models import GpuJob, ImageRecord, ImageRegion
SessionLocal = _sync_session_factory()
with SessionLocal() as session:
if task_name == "siglip":
has_concept = exists().where(
ImageRegion.image_record_id == ImageRecord.id,
ImageRegion.kind == "concept",
)
queued = exists().where(
GpuJob.image_record_id == ImageRecord.id,
GpuJob.task == "siglip",
GpuJob.status.in_(["pending", "leased"]),
)
sel = sa_select(
ImageRecord.id, literal("siglip"), literal("pending")
).where(~has_concept).where(~queued)
else:
already = exists().where(
GpuJob.image_record_id == ImageRecord.id,
GpuJob.task == task_name,
GpuJob.status.in_(["pending", "leased", "done"]),
)
sel = sa_select(
ImageRecord.id, literal(task_name), literal("pending")
).where(~already)
# RETURNING + count: result.rowcount is unreliable for INSERT…SELECT.
rows = session.execute(
insert(GpuJob)
.from_select(["image_record_id", "task", "status"], sel)
.returning(GpuJob.id)
).fetchall()
session.commit()
return len(rows)
@celery.task(name="backend.app.tasks.ml.recover_orphaned_gpu_jobs")
def recover_orphaned_gpu_jobs() -> int:
"""Reset expired GPU-job leases back to pending — recovers work orphaned by an
agent that died mid-job (no graceful release). Short beat cadence so orphans
get picked back up quickly + the queue counts read honestly. Returns the
number recovered."""
from datetime import UTC, datetime
from sqlalchemy import update
from ..models import GpuJob
SessionLocal = _sync_session_factory()
with SessionLocal() as session:
now = datetime.now(UTC)
res = session.execute(
update(GpuJob)
.where(GpuJob.status == "leased", GpuJob.lease_expires_at < now)
.values(
status="pending", lease_token=None, leased_at=None,
lease_expires_at=None, updated_at=now,
)
)
session.commit()
return res.rowcount or 0
@celery.task(
name="backend.app.tasks.ml.scheduled_ccip_auto_apply",
soft_time_limit=1800, time_limit=2100,
)
def scheduled_ccip_auto_apply() -> str:
"""Auto-tag confident CCIP character matches (source='ccip_auto') so identity
tags keep flowing without a button. No-op unless ccip_auto_apply_enabled.
References come only from single-character images (unambiguous); a tag is
applied where any figure's best cosine to a character's prototypes clears
ccip_auto_apply_threshold and it isn't already applied/rejected. Reversible."""
import numpy as np
from sqlalchemy import func
from sqlalchemy import select as sa_select
from sqlalchemy.dialects.postgresql import insert as pg_insert
from ..models import ImageRegion, MLSettings, Tag, TagKind, TagSuggestionRejection
from ..models.tag import image_tag
fig = ("face", "figure")
def _l2(m):
n = np.linalg.norm(m, axis=1, keepdims=True)
n[n == 0] = 1.0
return m / n
SessionLocal = _sync_session_factory()
with SessionLocal() as session:
s = session.get(MLSettings, 1)
if s is None or not s.ccip_auto_apply_enabled:
return "disabled"
thr = float(s.ccip_auto_apply_threshold)
single = (
sa_select(image_tag.c.image_record_id)
.join(Tag, Tag.id == image_tag.c.tag_id)
.where(Tag.kind == TagKind.character)
.group_by(image_tag.c.image_record_id)
.having(func.count() == 1)
)
ref_rows = session.execute(
sa_select(image_tag.c.tag_id, ImageRegion.ccip_embedding)
.select_from(ImageRegion)
.join(
image_tag,
image_tag.c.image_record_id == ImageRegion.image_record_id,
)
.join(Tag, Tag.id == image_tag.c.tag_id)
.where(Tag.kind == TagKind.character)
.where(ImageRegion.kind.in_(fig))
.where(ImageRegion.ccip_embedding.is_not(None))
.where(ImageRegion.image_record_id.in_(single))
).all()
if not ref_rows:
return "no-references"
by_char: dict[int, list] = {}
for tid, vec in ref_rows:
by_char.setdefault(tid, []).append(vec)
ref_tags = list(by_char)
mats = [_l2(np.asarray(by_char[t], dtype=np.float32)) for t in ref_tags]
allref = np.vstack(mats) # (total, 768)
seg = np.cumsum([0] + [len(m) for m in mats])[:-1] # per-char start
# Per character: images that already carry OR rejected the tag — skip.
skip = {t: set() for t in ref_tags}
for t in ref_tags:
for (iid,) in session.execute(
sa_select(image_tag.c.image_record_id).where(
image_tag.c.tag_id == t
)
):
skip[t].add(iid)
for (iid,) in session.execute(
sa_select(TagSuggestionRejection.image_record_id).where(
TagSuggestionRejection.tag_id == t
)
):
skip[t].add(iid)
img_ids = list(session.execute(
sa_select(ImageRegion.image_record_id)
.where(ImageRegion.kind.in_(fig), ImageRegion.ccip_embedding.is_not(None))
.distinct()
).scalars())
applied = 0
chunk_n = 500
for start in range(0, len(img_ids), chunk_n):
chunk = img_ids[start:start + chunk_n]
rows = session.execute(
sa_select(ImageRegion.image_record_id, ImageRegion.ccip_embedding)
.where(
ImageRegion.image_record_id.in_(chunk),
ImageRegion.kind.in_(fig),
ImageRegion.ccip_embedding.is_not(None),
)
).all()
by_img: dict[int, list] = {}
for iid, vec in rows:
by_img.setdefault(iid, []).append(vec)
for iid, vecs in by_img.items():
q = _l2(np.asarray(vecs, dtype=np.float32)) # (nq, 768)
colmax = (q @ allref.T).max(axis=0) # (total,)
charmax = np.maximum.reduceat(colmax, seg) # (n_chars,)
for ci in np.where(charmax >= thr)[0]:
t = ref_tags[int(ci)]
if iid in skip[t]:
continue
skip[t].add(iid)
session.execute(
pg_insert(image_tag)
.values(
image_record_id=iid, tag_id=t, source="ccip_auto",
)
.on_conflict_do_nothing()
)
applied += 1
session.commit()
return f"applied={applied}"
+3 -7
View File
@@ -72,14 +72,10 @@ import PipelineStatusChip from './PipelineStatusChip.vue'
const system = useSystemStore()
onMounted(() => system.refreshHealth())
// Every route with a meta.title is a nav entry. Order by meta.navOrder —
// router.getRoutes() does NOT guarantee declaration order, so explicit numbers
// pin the sequence (e.g. Explore after Gallery). Routes without one fall to the
// end. Auto-tracks future routes.
// Same mechanism the old sidebar used: every route with a meta.title is a
// nav entry, in router declaration order. Auto-tracks future routes.
const navRoutes = computed(() =>
router.getRoutes()
.filter(r => r.meta?.title)
.sort((a, b) => (a.meta.navOrder ?? 999) - (b.meta.navOrder ?? 999))
router.getRoutes().filter(r => r.meta?.title)
)
// Content links for the centered desktop row — everything EXCEPT Settings,
// which is config and gets pinned to the right edge instead.
@@ -1,50 +1,49 @@
<template>
<MaintenanceTile
icon="mdi-image-size-select-small"
title="Minimum dimensions"
blurb="Delete images below the import min-width / min-height threshold."
destructive
>
<p class="fc-muted text-body-2 mb-3">
Mirrors the import-time <code>min_width</code> / <code>min_height</code>
filter, applied retroactively to the existing library.
</p>
<v-card class="fc-clean-card">
<CardHeading icon="mdi-image-size-select-small" title="Minimum dimensions" />
<v-card-text>
<p class="fc-muted text-body-2 mb-3">
Find and delete images smaller than the threshold. Mirrors the
import-time <code>min_width</code> / <code>min_height</code>
filter, applied retroactively to the existing library.
</p>
<v-row dense>
<v-col cols="6">
<v-text-field
v-model.number="minW" label="Min width (px)" type="number"
min="0" density="compact" hide-details
/>
</v-col>
<v-col cols="6">
<v-text-field
v-model.number="minH" label="Min height (px)" type="number"
min="0" density="compact" hide-details
/>
</v-col>
</v-row>
<v-row dense>
<v-col cols="6">
<v-text-field
v-model.number="minW" label="Min width (px)" type="number"
min="0" density="compact" hide-details
/>
</v-col>
<v-col cols="6">
<v-text-field
v-model.number="minH" label="Min height (px)" type="number"
min="0" density="compact" hide-details
/>
</v-col>
</v-row>
<div class="d-flex align-center mt-3" style="gap: 10px;">
<v-btn
color="accent" variant="flat" rounded="pill"
prepend-icon="mdi-magnify"
:loading="busy"
@click="onPreview"
>Preview</v-btn>
<span v-if="preview" class="text-body-2">
<strong>{{ preview.count }}</strong> image(s) would be deleted.
</span>
</div>
<div class="d-flex align-center mt-3" style="gap: 10px;">
<v-btn
color="accent" variant="flat" rounded="pill"
prepend-icon="mdi-magnify"
:loading="busy"
@click="onPreview"
>Preview</v-btn>
<span v-if="preview" class="text-body-2">
<strong>{{ preview.count }}</strong> image(s) would be deleted.
</span>
</div>
<v-btn
v-if="preview && preview.count > 0"
class="mt-3"
color="error" variant="flat" rounded="pill"
prepend-icon="mdi-delete"
@click="onDeleteClick"
>Delete {{ preview.count }} matching...</v-btn>
v-if="preview && preview.count > 0"
class="mt-3"
color="error" variant="flat" rounded="pill"
prepend-icon="mdi-delete"
@click="onDeleteClick"
>Delete {{ preview.count }} matching...</v-btn>
</v-card-text>
<DestructiveConfirmModal
v-model="showModal"
@@ -56,7 +55,7 @@
:description="`Width < ${minW} OR height < ${minH}`"
@confirm="onConfirmedDelete"
/>
</MaintenanceTile>
</v-card>
</template>
<script setup>
@@ -65,7 +64,7 @@ import { onMounted, ref } from 'vue'
import DestructiveConfirmModal from '../modal/DestructiveConfirmModal.vue'
import { useCleanupStore } from '../../stores/cleanup.js'
import MaintenanceTile from '../common/MaintenanceTile.vue'
import CardHeading from '../common/CardHeading.vue'
// Backend's preview response hands the full Tier-C confirm token back
// as `confirm_token` (e.g. `delete-min-dim-1a2b3c4d`); passed straight
@@ -117,3 +116,7 @@ async function onConfirmedDelete(token) {
}
}
</script>
<style scoped>
.fc-clean-card { border-radius: 8px; }
</style>
@@ -1,82 +1,79 @@
<template>
<MaintenanceTile
icon="mdi-palette-swatch"
title="Single-color audit"
blurb="Find & delete near-solid / placeholder images."
destructive
:open="audit?.status === 'running'"
>
<p class="fc-muted text-body-2 mb-3">
Scan library for images dominated by one color within the
tolerance. Catches placeholder / solid-fill / error-page images
that slipped through the import filter. Same background-scan
cadence as the transparency audit.
</p>
<v-row dense>
<v-col cols="6">
<v-text-field
v-model.number="threshold" label="Threshold (01)"
type="number" min="0" max="1" step="0.01"
density="compact" hide-details
:disabled="audit && audit.status === 'running'"
/>
</v-col>
<v-col cols="6">
<v-text-field
v-model.number="tolerance" label="Color tolerance (0441)"
type="number" min="0" max="441"
density="compact" hide-details
:disabled="audit && audit.status === 'running'"
/>
</v-col>
</v-row>
<v-btn
v-if="!audit || audit.status !== 'running'"
class="mt-3"
color="accent" variant="flat" rounded="pill"
prepend-icon="mdi-magnify-scan"
:loading="busy"
@click="onStart"
>Scan library</v-btn>
<div v-if="audit && audit.status === 'running'" class="mt-3">
<v-progress-linear indeterminate color="accent" />
<div class="text-body-2 mt-2 d-flex align-center" style="gap: 10px;">
<span>
Scanning {{ audit.scanned_count }} checked,
{{ audit.matched_count }} matched
</span>
<v-btn
variant="text" size="small" color="warning" rounded="pill"
@click="onCancel"
>Cancel</v-btn>
</div>
</div>
<div v-if="audit && audit.status === 'ready'" class="mt-3">
<p class="text-body-2 mb-2">
Scan complete. <strong>{{ audit.matched_count }}</strong>
image(s) match.
<v-card class="fc-clean-card">
<CardHeading icon="mdi-palette-swatch" title="Single-color audit" />
<v-card-text>
<p class="fc-muted text-body-2 mb-3">
Scan library for images dominated by one color within the
tolerance. Catches placeholder / solid-fill / error-page images
that slipped through the import filter. Same background-scan
cadence as the transparency audit.
</p>
<v-row dense>
<v-col cols="6">
<v-text-field
v-model.number="threshold" label="Threshold (01)"
type="number" min="0" max="1" step="0.01"
density="compact" hide-details
:disabled="audit && audit.status === 'running'"
/>
</v-col>
<v-col cols="6">
<v-text-field
v-model.number="tolerance" label="Color tolerance (0441)"
type="number" min="0" max="441"
density="compact" hide-details
:disabled="audit && audit.status === 'running'"
/>
</v-col>
</v-row>
<v-btn
v-if="audit.matched_count > 0"
color="error" variant="flat" rounded="pill"
prepend-icon="mdi-delete"
@click="onApplyClick"
>Delete {{ audit.matched_count }} matching...</v-btn>
</div>
v-if="!audit || audit.status !== 'running'"
class="mt-3"
color="accent" variant="flat" rounded="pill"
prepend-icon="mdi-magnify-scan"
:loading="busy"
@click="onStart"
>Scan library</v-btn>
<v-alert
v-if="audit && audit.status === 'error'"
type="error" variant="tonal" density="compact" class="mt-3"
>Scan failed: {{ audit.error }}</v-alert>
<div v-if="audit && audit.status === 'running'" class="mt-3">
<v-progress-linear indeterminate color="accent" />
<div class="text-body-2 mt-2 d-flex align-center" style="gap: 10px;">
<span>
Scanning {{ audit.scanned_count }} checked,
{{ audit.matched_count }} matched
</span>
<v-btn
variant="text" size="small" color="warning" rounded="pill"
@click="onCancel"
>Cancel</v-btn>
</div>
</div>
<v-alert
v-if="audit && audit.status === 'applied'"
type="success" variant="tonal" density="compact" class="mt-3"
>Applied matched images deleted.</v-alert>
<div v-if="audit && audit.status === 'ready'" class="mt-3">
<p class="text-body-2 mb-2">
Scan complete. <strong>{{ audit.matched_count }}</strong>
image(s) match.
</p>
<v-btn
v-if="audit.matched_count > 0"
color="error" variant="flat" rounded="pill"
prepend-icon="mdi-delete"
@click="onApplyClick"
>Delete {{ audit.matched_count }} matching...</v-btn>
</div>
<v-alert
v-if="audit && audit.status === 'error'"
type="error" variant="tonal" density="compact" class="mt-3"
>Scan failed: {{ audit.error }}</v-alert>
<v-alert
v-if="audit && audit.status === 'applied'"
type="success" variant="tonal" density="compact" class="mt-3"
>Applied matched images deleted.</v-alert>
</v-card-text>
<DestructiveConfirmModal
v-if="audit"
@@ -89,7 +86,7 @@
description="Permanently deletes images matched by the single-color scan."
@confirm="onConfirmedApply"
/>
</MaintenanceTile>
</v-card>
</template>
<script setup>
@@ -97,7 +94,7 @@ import { toast } from '../../utils/toast.js'
import { onMounted, onUnmounted, ref } from 'vue'
import DestructiveConfirmModal from '../modal/DestructiveConfirmModal.vue'
import MaintenanceTile from '../common/MaintenanceTile.vue'
import CardHeading from '../common/CardHeading.vue'
import { useCleanupStore } from '../../stores/cleanup.js'
const store = useCleanupStore()
@@ -188,3 +185,7 @@ async function onConfirmedApply(token) {
}
}
</script>
<style scoped>
.fc-clean-card { border-radius: 8px; }
</style>
@@ -1,69 +1,66 @@
<template>
<MaintenanceTile
icon="mdi-checkerboard"
title="Transparency audit"
blurb="Find & delete images that are mostly transparent."
destructive
:open="audit?.status === 'running'"
>
<p class="fc-muted text-body-2 mb-3">
Scan library for images whose transparent-pixel fraction exceeds
the threshold. Animated WebPs / GIFs are skipped (the import-side
rule does the same). Runs as a background task ~50ms per image,
so a 57k library takes ~50 minutes.
</p>
<v-text-field
v-model.number="threshold" label="Transparency threshold (01)"
type="number" min="0" max="1" step="0.01" density="compact" hide-details
:disabled="audit && audit.status === 'running'"
class="mb-3"
/>
<v-btn
v-if="!audit || audit.status !== 'running'"
color="accent" variant="flat" rounded="pill"
prepend-icon="mdi-magnify-scan"
:loading="busy"
@click="onStart"
>Scan library</v-btn>
<div v-if="audit && audit.status === 'running'" class="mt-3">
<v-progress-linear indeterminate color="accent" />
<div class="text-body-2 mt-2 d-flex align-center" style="gap: 10px;">
<span>
Scanning {{ audit.scanned_count }} checked,
{{ audit.matched_count }} matched
</span>
<v-btn
variant="text" size="small" color="warning" rounded="pill"
@click="onCancel"
>Cancel</v-btn>
</div>
</div>
<div v-if="audit && audit.status === 'ready'" class="mt-3">
<p class="text-body-2 mb-2">
Scan complete. <strong>{{ audit.matched_count }}</strong>
image(s) match.
<v-card class="fc-clean-card">
<CardHeading icon="mdi-checkerboard" title="Transparency audit" />
<v-card-text>
<p class="fc-muted text-body-2 mb-3">
Scan library for images whose transparent-pixel fraction exceeds
the threshold. Animated WebPs / GIFs are skipped (the import-side
rule does the same). Runs as a background task ~50ms per image,
so a 57k library takes ~50 minutes.
</p>
<v-text-field
v-model.number="threshold" label="Transparency threshold (01)"
type="number" min="0" max="1" step="0.01" density="compact" hide-details
:disabled="audit && audit.status === 'running'"
class="mb-3"
/>
<v-btn
v-if="audit.matched_count > 0"
color="error" variant="flat" rounded="pill"
prepend-icon="mdi-delete"
@click="onApplyClick"
>Delete {{ audit.matched_count }} matching...</v-btn>
</div>
v-if="!audit || audit.status !== 'running'"
color="accent" variant="flat" rounded="pill"
prepend-icon="mdi-magnify-scan"
:loading="busy"
@click="onStart"
>Scan library</v-btn>
<v-alert
v-if="audit && audit.status === 'error'"
type="error" variant="tonal" density="compact" class="mt-3"
>Scan failed: {{ audit.error }}</v-alert>
<div v-if="audit && audit.status === 'running'" class="mt-3">
<v-progress-linear indeterminate color="accent" />
<div class="text-body-2 mt-2 d-flex align-center" style="gap: 10px;">
<span>
Scanning {{ audit.scanned_count }} checked,
{{ audit.matched_count }} matched
</span>
<v-btn
variant="text" size="small" color="warning" rounded="pill"
@click="onCancel"
>Cancel</v-btn>
</div>
</div>
<v-alert
v-if="audit && audit.status === 'applied'"
type="success" variant="tonal" density="compact" class="mt-3"
>Applied matched images deleted.</v-alert>
<div v-if="audit && audit.status === 'ready'" class="mt-3">
<p class="text-body-2 mb-2">
Scan complete. <strong>{{ audit.matched_count }}</strong>
image(s) match.
</p>
<v-btn
v-if="audit.matched_count > 0"
color="error" variant="flat" rounded="pill"
prepend-icon="mdi-delete"
@click="onApplyClick"
>Delete {{ audit.matched_count }} matching...</v-btn>
</div>
<v-alert
v-if="audit && audit.status === 'error'"
type="error" variant="tonal" density="compact" class="mt-3"
>Scan failed: {{ audit.error }}</v-alert>
<v-alert
v-if="audit && audit.status === 'applied'"
type="success" variant="tonal" density="compact" class="mt-3"
>Applied matched images deleted.</v-alert>
</v-card-text>
<DestructiveConfirmModal
v-if="audit"
@@ -76,7 +73,7 @@
description="Permanently deletes images matched by the transparency scan."
@confirm="onConfirmedApply"
/>
</MaintenanceTile>
</v-card>
</template>
<script setup>
@@ -84,7 +81,7 @@ import { toast } from '../../utils/toast.js'
import { onMounted, onUnmounted, ref } from 'vue'
import DestructiveConfirmModal from '../modal/DestructiveConfirmModal.vue'
import MaintenanceTile from '../common/MaintenanceTile.vue'
import CardHeading from '../common/CardHeading.vue'
import { useCleanupStore } from '../../stores/cleanup.js'
const store = useCleanupStore()
@@ -171,3 +168,7 @@ async function onConfirmedApply(token) {
}
}
</script>
<style scoped>
.fc-clean-card { border-radius: 8px; }
</style>
@@ -1,121 +0,0 @@
<!--
Compact, expandable maintenance/cleanup action tile. Collapsed it shows just an
icon + short title + one-line purpose, so a section of these tiles in a grid is
scannable at a glance; clicking the header expands the full controls / preview /
result UI inline (the operator opted for "compact tiles in a grid, detail on
expand", 2026-06-18, replacing the old long stack of full-width cards).
Usage: wrap a card's action body in the default slot; pass icon/title/blurb.
`destructive` tints the icon error-red for delete actions. `open` can be forced
(e.g. keep a running task's tile expanded). Keyboard accessible: the header is a
real <button> with aria-expanded + focus ring.
-->
<template>
<v-card class="fc-tile" :class="{ 'fc-tile--open': isOpen }">
<button
type="button"
class="fc-tile__head"
:aria-expanded="isOpen"
@click="toggle"
>
<v-icon
:icon="icon"
:color="destructive ? 'error' : 'accent'"
class="fc-tile__icon"
/>
<span class="fc-tile__text">
<span class="fc-tile__title">{{ title }}</span>
<span class="fc-tile__blurb">{{ blurb }}</span>
</span>
<v-icon
:icon="isOpen ? 'mdi-chevron-up' : 'mdi-chevron-down'"
class="fc-tile__chev"
/>
</button>
<v-expand-transition>
<div v-show="isOpen" class="fc-tile__body">
<slot />
</div>
</v-expand-transition>
</v-card>
</template>
<script setup>
import { computed, ref, watch } from 'vue'
const props = defineProps({
icon: { type: String, required: true },
title: { type: String, required: true },
blurb: { type: String, default: '' },
destructive: { type: Boolean, default: false },
// Force-open (e.g. a tile whose task is mid-run). When set, the operator can
// still collapse it locally, but a change to `open` re-applies.
open: { type: Boolean, default: false },
})
const local = ref(props.open)
watch(() => props.open, (v) => { local.value = v })
const isOpen = computed(() => local.value)
function toggle() {
local.value = !local.value
}
</script>
<style scoped>
.fc-tile {
border-radius: 8px;
overflow: hidden;
}
.fc-tile__head {
display: flex;
align-items: center;
gap: 12px;
width: 100%;
padding: 12px 14px;
background: transparent;
border: none;
text-align: left;
cursor: pointer;
color: inherit;
}
.fc-tile__head:hover {
background: rgba(var(--v-theme-on-surface), 0.04);
}
.fc-tile__head:focus-visible {
outline: 2px solid rgb(var(--v-theme-accent));
outline-offset: -2px;
}
.fc-tile__icon {
flex: 0 0 auto;
}
.fc-tile__text {
display: flex;
flex-direction: column;
min-width: 0;
flex: 1 1 auto;
}
.fc-tile__title {
font-weight: 600;
font-size: 0.95rem;
line-height: 1.2;
}
.fc-tile__blurb {
font-size: 0.8rem;
line-height: 1.25;
color: rgb(var(--v-theme-on-surface-variant));
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
}
.fc-tile--open .fc-tile__blurb {
white-space: normal;
}
.fc-tile__chev {
flex: 0 0 auto;
color: rgb(var(--v-theme-on-surface-variant));
}
.fc-tile__body {
padding: 0 14px 14px;
}
</style>
@@ -1,68 +0,0 @@
<template>
<!-- Thin debounced tag-search autocomplete: emits `pick` with the chosen tag
and self-clears. Shared by the gallery's advanced tag-query builder; the
filter bar's own inline search predates this and also folds in artists. -->
<v-autocomplete
v-model="selected"
:items="items"
:loading="loading"
item-title="name" item-value="id"
no-filter hide-details density="compact" variant="outlined"
:placeholder="placeholder"
prepend-inner-icon="mdi-tag-plus-outline"
:aria-label="placeholder"
@update:search="onSearch"
@update:model-value="onPick"
>
<template #item="{ props: itemProps, item }">
<v-list-item v-bind="itemProps" :title="item.raw.name">
<template #subtitle>
{{ item.raw.fandom_name ? `character · ${item.raw.fandom_name}` : item.raw.kind }}
</template>
</v-list-item>
</template>
<template #no-data>
<v-list-item :title="searchedOnce ? 'No tags match' : 'Type to search tags'" />
</template>
</v-autocomplete>
</template>
<script setup>
import { ref, onBeforeUnmount } from 'vue'
import { useApi } from '../../composables/useApi.js'
defineProps({ placeholder: { type: String, default: 'Add tag…' } })
const emit = defineEmits(['pick'])
const api = useApi()
const selected = ref(null)
const items = ref([])
const loading = ref(false)
const searchedOnce = ref(false)
let debounce = null
function onSearch (q) {
if (debounce) clearTimeout(debounce)
if (!q || !q.trim()) { items.value = []; return }
debounce = setTimeout(async () => {
loading.value = true
try {
items.value = (await api.get('/api/tags/autocomplete', { params: { q, limit: 10 } })) || []
} catch {
items.value = []
} finally {
loading.value = false
searchedOnce.value = true
}
}, 250)
}
function onPick (id) {
const tag = items.value.find((i) => i.id === id)
selected.value = null
items.value = []
if (tag) emit('pick', tag)
}
onBeforeUnmount(() => { if (debounce) clearTimeout(debounce) })
</script>
@@ -27,34 +27,11 @@
</v-autocomplete>
<div class="fc-filterbar__chips">
<!-- Include tag chips: click the body to flip to exclude, to remove
(#6 light editor one model, two editors). -->
<v-chip
v-for="id in store.filter.tag_ids" :key="`t${id}`"
size="small" closable :color="chipColor(id)" variant="tonal"
class="fc-filterbar__tagchip"
title="Click to exclude this tag instead"
@click="toggleTagPolarity(id)"
@click:close="removeTag(id)"
>{{ store.tagLabels[id] || `#${id}` }}</v-chip>
<!-- Exclude tag chips (red, minus). Click body to flip back to include. -->
<v-chip
v-for="id in store.filter.tag_exclude" :key="`x${id}`"
size="small" closable color="error" variant="tonal"
class="fc-filterbar__tagchip"
title="Excluded — click to include instead"
@click="toggleTagPolarity(id)"
@click:close="removeExclude(id)"
><v-icon start size="x-small">mdi-minus</v-icon>{{ store.tagLabels[id] || `#${id}` }}</v-chip>
<!-- Advanced OR-groups can't render as flat chips; one affordance opens
the builder where they live. -->
<v-chip
v-if="orGroupCount"
size="small" color="accent" variant="tonal"
prepend-icon="mdi-filter-cog-outline"
title="Open the advanced tag filter"
@click="advancedOpen = true"
>{{ orGroupCount }} OR-group{{ orGroupCount > 1 ? 's' : '' }}</v-chip>
<v-chip
v-if="store.filter.artist_id"
size="small" closable color="accent" variant="tonal"
@@ -91,13 +68,6 @@
@update:model-value="setSort"
/>
<v-btn
:color="advancedActive ? 'accent' : undefined"
:variant="advancedActive ? 'tonal' : 'text'"
size="small" prepend-icon="mdi-filter-cog-outline"
@click="advancedOpen = true"
>Advanced{{ advancedCount ? ` (${advancedCount})` : '' }}</v-btn>
<v-btn
:color="refineOpen ? 'accent' : undefined"
:variant="refineOpen || hasRefineFilters ? 'tonal' : 'text'"
@@ -113,18 +83,6 @@
</div>
<GalleryFacetPanel v-if="refineOpen" />
<v-dialog v-model="advancedOpen" max-width="640" scrollable>
<TagQueryBuilder
v-if="advancedOpen"
:include="store.filter.tag_ids"
:or-groups="store.filter.tag_or"
:exclude="store.filter.tag_exclude"
:labels="store.tagLabels"
@apply="onAdvancedApply"
@close="advancedOpen = false"
/>
</v-dialog>
</div>
</template>
@@ -135,7 +93,6 @@ import { useApi } from '../../composables/useApi.js'
import { cloneFilter, filterToQuery, useGalleryStore } from '../../stores/gallery.js'
import { useTagStore } from '../../stores/tags.js'
import GalleryFacetPanel from './GalleryFacetPanel.vue'
import TagQueryBuilder from './TagQueryBuilder.vue'
const store = useGalleryStore()
const tagStore = useTagStore()
@@ -160,19 +117,8 @@ const refineCount = computed(() => {
})
const hasRefineFilters = computed(() => refineCount.value > 0)
// The structured tag filter beyond plain includes (OR-groups + excludes) —
// drives the Advanced button's active state + its count badge.
const advancedOpen = ref(false)
const orGroupCount = computed(() => store.filter.tag_or.length)
const advancedCount = computed(
() => store.filter.tag_or.length + store.filter.tag_exclude.length,
)
const advancedActive = computed(() => advancedOpen.value || advancedCount.value > 0)
const hasActiveFilters = computed(() =>
store.filter.tag_ids.length > 0 ||
store.filter.tag_exclude.length > 0 ||
store.filter.tag_or.length > 0 ||
store.filter.artist_id != null ||
store.filter.media_type != null ||
store.filter.sort !== 'newest' ||
@@ -249,34 +195,6 @@ function onPick(value) {
function removeTag(id) {
pushFilter((n) => { n.tag_ids = n.tag_ids.filter((t) => t !== id) })
}
function removeExclude(id) {
pushFilter((n) => { n.tag_exclude = n.tag_exclude.filter((t) => t !== id) })
}
// Flip a tag between include and exclude in place (light-editor toggle). A tag
// is only ever in one of the two lists.
function toggleTagPolarity(id) {
pushFilter((n) => {
if (n.tag_ids.includes(id)) {
n.tag_ids = n.tag_ids.filter((t) => t !== id)
if (!n.tag_exclude.includes(id)) n.tag_exclude.push(id)
} else {
n.tag_exclude = n.tag_exclude.filter((t) => t !== id)
if (!n.tag_ids.includes(id)) n.tag_ids.push(id)
}
})
}
// The advanced builder hands back the whole tag model at once.
function onAdvancedApply({ tag_ids, tag_or, tag_exclude, labels }) {
for (const [id, name] of Object.entries(labels || {})) {
store.noteTagLabel(Number(id), name)
}
pushFilter((n) => {
n.tag_ids = tag_ids
n.tag_or = tag_or
n.tag_exclude = tag_exclude
})
advancedOpen.value = false
}
function clearArtist() {
store.noteArtistLabel(null)
pushFilter((n) => { n.artist_id = null })
@@ -339,12 +257,6 @@ function pushFilter(mutate) {
}
.fc-filterbar__search { max-width: 320px; min-width: 200px; }
.fc-filterbar__chips { display: flex; align-items: center; gap: 6px; flex-wrap: wrap; }
/* The tag chips' bodies toggle include/exclude — signal they're clickable. */
.fc-filterbar__tagchip { cursor: pointer; }
.fc-filterbar__tagchip:focus-visible {
outline: 2px solid rgb(var(--v-theme-accent));
outline-offset: 1px;
}
.fc-filterbar__sort { max-width: 150px; }
/* Phones: the search's 200px min-width jams the wrapping bar. Give search its
@@ -1,236 +0,0 @@
<template>
<v-card class="fc-tqb">
<v-card-title class="fc-tqb__head">
<v-icon icon="mdi-filter-cog-outline" size="small" class="mr-2" />
Advanced tag filter
<v-spacer />
<v-btn
icon="mdi-close" variant="text" size="small"
aria-label="Close advanced tag filter" @click="$emit('close')"
/>
</v-card-title>
<v-card-text>
<p class="text-body-2 fc-tqb__hint">
An image must match <strong>every</strong> group below, and a group
matches when the image carries <strong>any</strong> tag in it
(AND of ORs). Excluded tags are never shown.
</p>
<!-- AND-of-OR groups -->
<div class="fc-tqb__section">
<div class="fc-tqb__section-title">Must match all of</div>
<p v-if="!groups.length" class="fc-tqb__empty">
No groups yet add a tag below to start.
</p>
<template v-for="(g, gi) in groups" :key="gi">
<div v-if="gi > 0" class="fc-tqb__and">AND</div>
<div class="fc-tqb__group">
<div class="fc-tqb__group-chips">
<template v-for="(id, ci) in g" :key="id">
<span v-if="ci > 0" class="fc-tqb__or">or</span>
<v-chip
size="small" closable variant="tonal"
:color="kindColor(id)"
@click:close="removeFromGroup(gi, id)"
>{{ label(id) }}</v-chip>
</template>
<span v-if="!g.length" class="fc-tqb__empty">empty</span>
</div>
<div class="fc-tqb__group-actions">
<TagPicker
placeholder="or…" class="fc-tqb__picker"
@pick="(t) => addToGroup(gi, t)"
/>
<v-btn
icon="mdi-delete-outline" variant="text" size="small"
aria-label="Remove this group" @click="removeGroup(gi)"
/>
</div>
</div>
</template>
<div class="fc-tqb__add">
<TagPicker
placeholder="Add a tag (new group)…" class="fc-tqb__picker"
@pick="addNewGroup"
/>
</div>
</div>
<v-divider class="my-4" />
<!-- exclude -->
<div class="fc-tqb__section">
<div class="fc-tqb__section-title">Exclude (must NOT have)</div>
<div class="fc-tqb__group-chips">
<v-chip
v-for="id in excludeIds" :key="id"
size="small" closable color="error" variant="tonal"
@click:close="removeExclude(id)"
>
<v-icon start size="x-small">mdi-minus</v-icon>{{ label(id) }}
</v-chip>
<span v-if="!excludeIds.length" class="fc-tqb__empty">none</span>
</div>
<div class="fc-tqb__add">
<TagPicker
placeholder="Exclude a tag…" class="fc-tqb__picker"
@pick="addExclude"
/>
</div>
</div>
</v-card-text>
<v-card-actions class="fc-tqb__actions">
<v-btn
variant="text" size="small" :disabled="!isDirty && isEmpty"
@click="clearAll"
>Clear</v-btn>
<v-spacer />
<v-btn variant="text" @click="$emit('close')">Cancel</v-btn>
<v-btn color="accent" variant="flat" rounded="pill" @click="apply">
Apply filter
</v-btn>
</v-card-actions>
</v-card>
</template>
<script setup>
import { ref, computed } from 'vue'
import { useTagStore } from '../../stores/tags.js'
import TagPicker from '../common/TagPicker.vue'
// One editor for the whole structured tag model. It works purely on "AND-of-OR
// groups + exclude": singleton includes (tag_ids) and OR-groups (tag_or) are
// unified into one `groups` list here, then split back apart on Apply so the
// URL stays compact (singletons serialize to tag_id, multi-tag groups to
// tag_or). Writes the SAME model the light chips edit — two editors, one model.
const props = defineProps({
include: { type: Array, default: () => [] }, // tag_ids
orGroups: { type: Array, default: () => [] }, // tag_or
exclude: { type: Array, default: () => [] }, // tag_exclude
labels: { type: Object, default: () => ({}) }, // id -> name (best effort)
})
const emit = defineEmits(['apply', 'close'])
const tagStore = useTagStore()
// Seed local draft from props (the dialog is recreated on each open, so a
// setup-time seed is the current filter every time).
const groups = ref([
...props.include.map((id) => [id]),
...props.orGroups.map((g) => [...g]),
])
const excludeIds = ref([...props.exclude])
// Names/kinds learned as tags are picked, layered over the labels passed in.
const nameMap = ref({ ...props.labels })
const kindMap = ref({})
function label (id) { return nameMap.value[id] || `#${id}` }
function kindColor (id) { return tagStore.colorFor(kindMap.value[id] || 'general') }
function _note (tag) {
nameMap.value = { ...nameMap.value, [tag.id]: tag.name }
kindMap.value = { ...kindMap.value, [tag.id]: tag.kind }
}
function _dropFromGroups (id) {
groups.value = groups.value
.map((g) => g.filter((t) => t !== id))
.filter((g) => g.length)
}
function addToGroup (gi, tag) {
_note(tag)
excludeIds.value = excludeIds.value.filter((t) => t !== tag.id) // can't be both
if (!groups.value[gi].includes(tag.id)) groups.value[gi].push(tag.id)
}
function addNewGroup (tag) {
_note(tag)
excludeIds.value = excludeIds.value.filter((t) => t !== tag.id)
groups.value.push([tag.id])
}
function removeFromGroup (gi, id) {
groups.value[gi] = groups.value[gi].filter((t) => t !== id)
if (!groups.value[gi].length) groups.value.splice(gi, 1)
}
function removeGroup (gi) { groups.value.splice(gi, 1) }
function addExclude (tag) {
_note(tag)
_dropFromGroups(tag.id) // a tag can't be both required and excluded
if (!excludeIds.value.includes(tag.id)) excludeIds.value.push(tag.id)
}
function removeExclude (id) {
excludeIds.value = excludeIds.value.filter((t) => t !== id)
}
function clearAll () { groups.value = []; excludeIds.value = [] }
const isEmpty = computed(() => !groups.value.length && !excludeIds.value.length)
const isDirty = computed(() =>
JSON.stringify(groups.value) !== JSON.stringify(props.orGroups.length || props.include.length
? [...props.include.map((id) => [id]), ...props.orGroups]
: []) ||
JSON.stringify(excludeIds.value) !== JSON.stringify(props.exclude),
)
function apply () {
const clean = groups.value.filter((g) => g.length)
emit('apply', {
tag_ids: clean.filter((g) => g.length === 1).map((g) => g[0]),
tag_or: clean.filter((g) => g.length > 1),
tag_exclude: [...excludeIds.value],
labels: nameMap.value, // so freshly-picked chips show their name on return
})
}
</script>
<style scoped>
.fc-tqb__head { display: flex; align-items: center; }
.fc-tqb__hint {
color: rgb(var(--v-theme-on-surface-variant));
margin-bottom: 16px;
}
.fc-tqb__section { margin-bottom: 4px; }
.fc-tqb__section-title {
font-size: 12px; font-weight: 600; letter-spacing: 0.04em;
text-transform: uppercase;
color: rgb(var(--v-theme-on-surface-variant));
margin-bottom: 8px;
}
.fc-tqb__and {
font-size: 11px; font-weight: 700; letter-spacing: 0.08em;
color: rgb(var(--v-theme-accent));
margin: 6px 0 6px 2px;
}
.fc-tqb__group {
display: flex; align-items: center; gap: 12px;
flex-wrap: wrap;
padding: 8px 10px;
border: 1px solid rgba(var(--v-theme-on-surface), 0.12);
border-radius: 10px;
background: rgba(var(--v-theme-on-surface), 0.03);
}
.fc-tqb__group-chips {
display: flex; align-items: center; gap: 6px; flex-wrap: wrap;
flex: 1 1 220px; min-width: 0;
}
.fc-tqb__or {
font-size: 11px; font-style: italic;
color: rgb(var(--v-theme-on-surface-variant));
}
.fc-tqb__group-actions {
display: flex; align-items: center; gap: 4px;
}
.fc-tqb__picker { min-width: 180px; max-width: 240px; }
.fc-tqb__add { margin-top: 10px; max-width: 280px; }
.fc-tqb__empty {
font-size: 12px; font-style: italic;
color: rgb(var(--v-theme-on-surface-variant));
}
.fc-tqb__actions { padding: 8px 16px 16px; }
</style>
@@ -1,112 +0,0 @@
<template>
<section v-if="img" class="fc-meta" aria-label="Image details">
<!-- #4a: dimensions / size / type sit as a compact top block alongside a
small save action (operator-asked 2026-06-26) meta on the left, the
download control shrunk to a floppy-disk icon on the right. -->
<dl class="fc-meta__grid">
<div v-if="img.width && img.height" class="fc-meta__item">
<dt>Dimensions</dt><dd>{{ img.width }} × {{ img.height }}</dd>
</div>
<div v-if="img.size_bytes" class="fc-meta__item">
<dt>Size</dt><dd>{{ humanSize(img.size_bytes) }}</dd>
</div>
<div v-if="img.mime" class="fc-meta__item">
<dt>Type</dt><dd>{{ shortType(img.mime) }}</dd>
</div>
</dl>
<!-- #4b: floppy-disk save icon = download; the kebab menu keeps Copy link.
Image-to-clipboard is OFF the table on the plain-HTTP origin (rule 95),
so the secondary action copies the link TEXT, which works everywhere. -->
<div class="fc-meta__actions">
<v-btn
icon="mdi-content-save" size="small" variant="tonal" color="accent"
aria-label="Download image" title="Download" @click="download"
/>
<v-menu location="bottom end">
<template #activator="{ props }">
<v-btn
v-bind="props" icon="mdi-dots-vertical" size="small" variant="text"
aria-label="More download options"
/>
</template>
<v-list density="compact">
<v-list-item prepend-icon="mdi-link-variant" title="Copy link" @click="copyLink" />
<v-list-item prepend-icon="mdi-content-save" title="Download" @click="download" />
</v-list>
</v-menu>
</div>
</section>
</template>
<script setup>
import { computed } from 'vue'
import { useModalStore } from '../../stores/modal.js'
import { copyText } from '../../utils/clipboard.js'
import { toast } from '../../utils/toast.js'
// `image` lets a non-modal surface (the Explore workspace) render the same
// meta + download block for its anchor. Defaults to the modal store's current
// image so the image modal is unchanged.
const props = defineProps({ image: { type: Object, default: null } })
const modal = useModalStore()
const img = computed(() => props.image ?? modal.current)
function humanSize (bytes) {
const units = ['B', 'KB', 'MB', 'GB']
let n = bytes
let i = 0
while (n >= 1024 && i < units.length - 1) { n /= 1024; i++ }
return `${n < 10 && i > 0 ? n.toFixed(1) : Math.round(n)} ${units[i]}`
}
function shortType (mime) { return mime.split('/')[1]?.toUpperCase() || mime }
function absoluteUrl () {
const origin = typeof window !== 'undefined' ? window.location.origin : ''
return origin + img.value.image_url
}
function download () {
// Same-origin /images/* link — the download attribute lets the browser save
// the original (filename derived from the path) instead of navigating to it.
const a = document.createElement('a')
a.href = img.value.image_url
a.setAttribute('download', '')
document.body.appendChild(a)
a.click()
a.remove()
}
async function copyLink () {
try {
await copyText(absoluteUrl())
toast({ text: 'Link copied to clipboard', type: 'success' })
} catch {
toast({ text: 'Could not copy the link', type: 'error' })
}
}
</script>
<style scoped>
.fc-meta {
padding: 14px 16px 0;
display: flex; align-items: flex-start; gap: 12px;
}
.fc-meta__grid {
flex: 1 1 auto; min-width: 0;
display: flex; flex-wrap: wrap; gap: 4px 18px; margin: 0;
}
.fc-meta__actions {
flex: 0 0 auto;
display: flex; align-items: center; gap: 2px;
}
.fc-meta__item { display: flex; flex-direction: column; }
.fc-meta__item dt {
font-size: 10px; text-transform: uppercase; letter-spacing: 0.06em;
color: rgb(var(--v-theme-on-surface-variant));
}
.fc-meta__item dd {
margin: 0; font-size: 13px; font-variant-numeric: tabular-nums;
color: rgb(var(--v-theme-on-surface));
}
</style>
+5 -31
View File
@@ -73,20 +73,11 @@
</div>
<aside v-if="modal.current" class="fc-viewer__side">
<!-- Provenance, then the meta + save block just beneath it, then tags
+ suggestions all scroll together here (operator-asked
2026-06-26: meta/download sits directly above Tags, under
Provenance). -->
<div class="fc-viewer__side-main">
<ProvenancePanel />
<ImageMetaBar />
<TagPanel />
</div>
<!-- while Related is PINNED to the bottom of the rail so it stays
reachable no matter how long Tags/Suggestions run (operator-asked
2026-06-26). Non-blocking: fetches its own similar set; collapses
silently (and takes no footer space) if empty/slow/failed. -->
<RelatedStrip class="fc-viewer__related" />
<ProvenancePanel />
<TagPanel />
<!-- Non-blocking: fetches its own similar set after the modal is up;
collapses silently if empty/slow/failed (see RelatedStrip). -->
<RelatedStrip />
</aside>
</div>
</div>
@@ -99,7 +90,6 @@ import { useModalStore } from '../../stores/modal.js'
import { arrowNavAllowed, isTextEntry } from '../../utils/textEntry.js'
import ImageCanvas from './ImageCanvas.vue'
import VideoCanvas from './VideoCanvas.vue'
import ImageMetaBar from './ImageMetaBar.vue'
import TagPanel from './TagPanel.vue'
import ProvenancePanel from './ProvenancePanel.vue'
import RelatedStrip from './RelatedStrip.vue'
@@ -317,18 +307,6 @@ function nextFrame() {
width: var(--fc-side-w); flex-shrink: 0;
background: rgb(var(--v-theme-surface));
border-left: 1px solid rgb(var(--v-theme-surface-light));
/* Flex column: a scrolling main area + a pinned Related footer. */
display: flex; flex-direction: column; min-height: 0;
}
.fc-viewer__side-main {
flex: 1 1 auto; min-height: 0; overflow-y: auto;
}
.fc-viewer__related {
/* Pinned to the bottom; capped so a tall strip can't swallow the rail —
it scrolls internally past the cap. RelatedStrip's own border-top draws
the divider; it self-collapses (no space) when there's nothing to show. */
flex: 0 0 auto;
max-height: 45%;
overflow-y: auto;
}
@@ -359,10 +337,6 @@ function nextFrame() {
border-left: none;
border-top: 1px solid rgb(var(--v-theme-surface-light));
}
/* The whole body scrolls on mobile — don't nest scrolls or pin Related; let
the rail flow naturally (Related lands at the end). */
.fc-viewer__side-main { flex: none; overflow: visible; }
.fc-viewer__related { max-height: none; overflow: visible; }
/* Re-center the prev/next arrows over the 55vh image band (their base
top:50% would land on the scrolling panel); next uses the full width
now that the panel is below, not beside. Close + integrity badge keep
@@ -70,21 +70,14 @@
</div>
<div v-if="attachments.length" class="fc-prov__attach">
<h4 class="fc-prov__attach-title">
{{ attachments.length === 1 ? 'Attachment' : `Attachments (${attachments.length})` }}
</h4>
<!-- Scroll-capped: a single post can carry dozens of archives (HR
bundle posts), which previously ballooned the panel past the
viewport. Mirror the cards' independent-scroll treatment. -->
<div class="fc-prov__attach-list">
<a
v-for="at in attachments" :key="at.id"
class="fc-prov__attach-row"
:href="at.download_url" :download="at.original_filename"
> {{ at.original_filename }}
<span class="fc-prov__attach-size">({{ at.size_bytes }} B)</span>
</a>
</div>
<h4 class="fc-prov__attach-title">Attachments</h4>
<a
v-for="at in attachments" :key="at.id"
class="fc-prov__attach-row"
:href="at.download_url" :download="at.original_filename"
>⬇ {{ at.original_filename }}
<span class="fc-prov__attach-size">({{ at.size_bytes }} B)</span>
</a>
</div>
</section>
</template>
@@ -97,19 +90,9 @@ import { useProvenanceStore } from '../../stores/provenance.js'
import { formatPostDate } from '../../utils/date.js'
import { toPlainText } from '../../utils/htmlSanitize.js'
// `imageId`/`image` let a non-modal surface (the Explore workspace) render
// provenance for its anchor. Default to the modal store's current image so the
// image modal is unchanged. Provenance is its own system (loaded by id via the
// provenance store), so it only needs the right id + the artist fallback.
const props = defineProps({
imageId: { type: Number, default: null },
image: { type: Object, default: null },
})
const modal = useModalStore()
const prov = useProvenanceStore()
const router = useRouter()
const effectiveId = computed(() => props.imageId ?? modal.currentImageId)
const effectiveImage = computed(() => props.image ?? modal.current)
// Per-post description collapse state (keyed by provenance_id). Default
// collapsed so multiple posts don't each eat ~180px of the panel the
@@ -117,21 +100,21 @@ const effectiveImage = computed(() => props.image ?? modal.current)
// 2026-05-28. Reset when the viewed image changes.
const expanded = reactive({})
function toggleDesc(id) { expanded[id] = !expanded[id] }
watch(() => effectiveId.value, () => {
watch(() => modal.currentImageId, () => {
for (const k of Object.keys(expanded)) delete expanded[k]
})
watch(
() => effectiveId.value,
() => modal.currentImageId,
(id) => { if (id != null) prov.loadForImage(id) },
{ immediate: true }
)
const state = computed(() =>
effectiveId.value == null ? null : prov.imageProv(effectiveId.value)
modal.currentImageId == null ? null : prov.imageProv(modal.currentImageId)
)
const fallbackArtist = computed(() => effectiveImage.value?.artist || null)
const fallbackArtist = computed(() => modal.current?.artist || null)
const showArtistFallback = computed(() => {
const st = state.value
@@ -246,15 +229,6 @@ function openPost(postId, artistId) {
font-size: 13px; color: rgb(var(--v-theme-on-surface-variant));
margin: 8px 0 4px;
}
.fc-prov__attach-list {
/* ~7 rows visible before scrolling; the clipped row hints at more.
Hairline scrollbar matching .fc-prov__cards. */
max-height: 180px;
overflow-y: auto;
scrollbar-width: thin;
scrollbar-color: rgb(var(--v-theme-surface-light)) transparent;
padding-right: 4px;
}
.fc-prov__attach-row {
display: block; font-size: 13px; text-decoration: none;
color: rgb(var(--v-theme-accent)); padding: 2px 0;
+5 -21
View File
@@ -5,18 +5,11 @@
<div v-if="show" class="fc-related">
<div class="fc-related__head">
<span class="fc-related__title">Related</span>
<div class="fc-related__actions">
<v-btn
size="x-small" variant="text" color="accent"
prepend-icon="mdi-compass-outline"
@click="explore"
>Explore</v-btn>
<v-btn
size="x-small" variant="text" color="accent"
:disabled="loading || !results.length"
@click="seeAll"
>See all similar</v-btn>
</div>
<v-btn
size="x-small" variant="text" color="accent"
:disabled="loading || !results.length"
@click="seeAll"
>See all similar</v-btn>
</div>
<div class="fc-related__row">
<template v-if="loading">
@@ -101,14 +94,6 @@ function seeAll() {
modal.close()
router.push({ name: 'gallery', query: { similar_to: String(id) } })
}
// #94: leave the modal and open the dedicated Explore walk anchored here.
function explore() {
const id = modal.current?.id
if (!id) return
modal.close()
router.push({ name: 'explore', params: { imageId: String(id) } })
}
</script>
<style scoped>
@@ -120,7 +105,6 @@ function explore() {
display: flex; align-items: center; justify-content: space-between;
margin-bottom: 8px;
}
.fc-related__actions { display: flex; align-items: center; gap: 2px; }
.fc-related__title {
font-size: 0.7rem; text-transform: uppercase; letter-spacing: 0.06em;
color: rgb(var(--v-theme-on-surface-variant));
@@ -1,54 +1,29 @@
<template>
<!-- Chip-card row: visible border + hover/focus state unifies the
name, score, and action buttons as one "object" (operator-asked
2026-06-01). The row itself is informational; the green / red
verdict pair + 3-dot alias menu are the action affordances. -->
<div class="fc-suggestion" :class="{ 'fc-suggestion--rejected': suggestion.rejected }">
2026-06-01). The row itself is informational; the explicit
Accept button + 3-dot menu are the action affordances. -->
<div class="fc-suggestion">
<span class="fc-suggestion__name">
{{ suggestion.display_name }}
<span v-if="suggestion.rejected" class="fc-suggestion__rejected-tag"
title="You rejected this for this image — un-reject to recover">rejected</span>
<span v-else-if="suggestion.creates_new_tag" class="fc-suggestion__new"
<span v-if="suggestion.creates_new_tag" class="fc-suggestion__new"
title="No matching tag yet — accepting creates it">+ new</span>
<span v-else-if="suggestion.via_alias" class="fc-suggestion__alias"
:title="`Mapped from the tagger's “${suggestion.raw_name}” via an alias`">alias</span>
</span>
<span class="fc-suggestion__score">{{ scorePct }}</span>
<!-- Green / red pair (operator-asked 2026-06-28) mirrors the eval
card's verdict buttons: ✓ accepts the tag (positive), ✗ dismisses it
for this image (records a TagSuggestionRejection — a hard negative the
heads train on). Together they occupy ~the footprint of the old single
Accept pill, so rejecting is now a one-click peer of accepting rather
than buried in the kebab. When the row is already rejected the ✗ swaps
to an undo (↶) so the rejection is reversible in place. -->
<div class="fc-suggestion__acts">
<button
class="fc-act fc-act--yes" type="button"
:aria-label="`Accept ${suggestion.display_name}`"
:title="`Yes — tag ${suggestion.display_name}`"
@click="$emit('accept', suggestion)"
><v-icon size="16">mdi-check</v-icon></button>
<button
v-if="suggestion.rejected"
class="fc-act fc-act--undo" type="button"
:aria-label="`Un-reject ${suggestion.display_name}`"
:title="`Undo — restore ${suggestion.display_name} as a suggestion`"
@click="$emit('undismiss', suggestion)"
><v-icon size="16">mdi-undo-variant</v-icon></button>
<button
v-else
class="fc-act fc-act--no" type="button"
:aria-label="`Reject ${suggestion.display_name}`"
:title="`No — not ${suggestion.display_name}`"
@click="$emit('dismiss', suggestion)"
><v-icon size="16">mdi-close</v-icon></button>
</div>
<v-btn
class="fc-suggestion__accept"
size="small" variant="tonal" color="accent"
density="compact" rounded="pill"
:aria-label="`Accept ${suggestion.display_name}`"
@click="$emit('accept', suggestion)"
>
Accept
</v-btn>
<!-- Modal-safe kebab is baked into KebabMenu (this row lives in the
teleported image modal — #711). Only rendered when an alias action
applies — dismiss now lives on the red ✗, so a centroid hit with no
alias option has no menu. -->
teleported image modal #711). -->
<KebabMenu
v-if="hasMenu"
class="fc-suggestion__menu" size="small" variant="outlined"
:label="`More actions for ${suggestion.display_name}`"
>
@@ -67,6 +42,9 @@
>
<v-list-item-title>Remove alias</v-list-item-title>
</v-list-item>
<v-list-item @click="$emit('dismiss', suggestion)">
<v-list-item-title>Dismiss for this image</v-list-item-title>
</v-list-item>
</KebabMenu>
</div>
</template>
@@ -76,15 +54,9 @@ import { computed } from 'vue'
import KebabMenu from '../common/KebabMenu.vue'
const props = defineProps({ suggestion: { type: Object, required: true } })
defineEmits(['accept', 'alias', 'remove-alias', 'dismiss', 'undismiss'])
defineEmits(['accept', 'alias', 'remove-alias', 'dismiss'])
const scorePct = computed(() => `${Math.round(props.suggestion.score * 100)}%`)
// Kebab now only carries alias actions: show it when this suggestion can be
// aliased (raw model key, not yet aliased) or is already aliased (so it can be
// un-aliased). Centroid hits (no raw_name, no alias) have an empty menu → hide.
const hasMenu = computed(() =>
Boolean(props.suggestion.raw_name) || Boolean(props.suggestion.via_alias)
)
</script>
<style scoped>
@@ -132,51 +104,12 @@ const hasMenu = computed(() =>
color: rgb(var(--v-theme-on-surface-variant, var(--v-theme-on-surface)));
font-family: 'JetBrains Mono', monospace;
}
/* Green ✓ / red ✗ verdict pair — same circular language as the eval card
(TagEvalCard .fc-act) so accept/reject read identically across surfaces. */
.fc-suggestion__acts {
flex: 0 0 auto; display: flex; gap: 4px;
}
.fc-act {
width: 26px; height: 26px; border-radius: 50%; border: none; cursor: pointer;
display: flex; align-items: center; justify-content: center; color: #fff;
opacity: 0.9; transition: transform 0.1s, opacity 0.1s;
}
.fc-act:hover { opacity: 1; transform: scale(1.1); }
.fc-act:focus-visible {
outline: 2px solid rgb(var(--v-theme-accent)); outline-offset: 1px;
}
.fc-act--yes { background: rgb(var(--v-theme-success)); }
.fc-act--no { background: rgb(var(--v-theme-error)); }
/* Undo reads as neutral-secondary, not a verdict: outlined, not filled. */
.fc-act--undo {
background: transparent; color: rgb(var(--v-theme-on-surface-variant));
border: 1px solid rgb(var(--v-theme-on-surface-variant), 0.5);
/* Vuetify's compact density doesn't shrink the tonal button enough
for a tight row; clamp the min-width so Accept stays compact. */
.fc-suggestion__accept :deep(.v-btn__content) {
font-size: 12px; letter-spacing: 0.02em;
}
.fc-suggestion__menu {
flex: 0 0 auto;
}
/* Rejected state: the row stays put (recovery), dimmed + red-edged so it
reads as "handled, negative" without shouting over live suggestions. */
.fc-suggestion--rejected {
border-color: rgb(var(--v-theme-error), 0.4);
background: rgb(var(--v-theme-error), 0.06);
}
.fc-suggestion--rejected .fc-suggestion__name {
color: rgb(var(--v-theme-on-surface-variant));
text-decoration: line-through;
text-decoration-color: rgb(var(--v-theme-error), 0.6);
}
.fc-suggestion__rejected-tag {
display: inline-block;
font-size: 10px; font-weight: 600;
color: rgb(var(--v-theme-error));
background: rgb(var(--v-theme-error), 0.12);
border: 1px solid rgb(var(--v-theme-error), 0.4);
padding: 1px 6px; border-radius: 999px;
margin-left: 6px;
text-transform: uppercase; letter-spacing: 0.04em;
text-decoration: none;
}
</style>
@@ -18,7 +18,6 @@
@alias="$emit('alias', $event)"
@remove-alias="$emit('remove-alias', $event)"
@dismiss="$emit('dismiss', $event)"
@undismiss="$emit('undismiss', $event)"
/>
</div>
</div>
@@ -34,7 +33,7 @@ const props = defineProps({
collapsible: { type: Boolean, default: false },
defaultOpen: { type: Boolean, default: true }
})
defineEmits(['accept', 'alias', 'remove-alias', 'dismiss', 'undismiss'])
defineEmits(['accept', 'alias', 'remove-alias', 'dismiss'])
const open = ref(props.collapsible ? props.defaultOpen : true)
</script>
@@ -12,25 +12,22 @@
No suggestions above threshold.
</div>
<!-- Flows in the rail's main scroll area; Related is pinned to the bottom
of the rail (ImageViewer side layout), so a long suggestion set no
longer needs an internal scroll cap to keep Related reachable. -->
<div v-else class="fc-suggestions__list">
<template v-else>
<SuggestionsCategoryGroup
v-for="cat in peopleCats" :key="cat"
v-show="store.byCategory[cat] && store.byCategory[cat].length"
:label="labelFor(cat)" :items="store.byCategory[cat] || []"
@accept="onAccept" @alias="onAlias" @remove-alias="onRemoveAlias"
@dismiss="onDismiss" @undismiss="onUndismiss"
@dismiss="store.dismiss"
/>
<SuggestionsCategoryGroup
v-if="store.byCategory.general && store.byCategory.general.length"
label="General" :items="store.byCategory.general"
collapsible :default-open="true"
@accept="onAccept" @alias="onAlias" @remove-alias="onRemoveAlias"
@dismiss="onDismiss" @undismiss="onUndismiss"
@dismiss="store.dismiss"
/>
</div>
</template>
<v-dialog v-model="aliasDialog" max-width="480">
<AliasPickerDialog
@@ -50,25 +47,12 @@ import { useModalStore } from '../../stores/modal.js'
import SuggestionsCategoryGroup from './SuggestionsCategoryGroup.vue'
import AliasPickerDialog from './AliasPickerDialog.vue'
const props = defineProps({
imageId: { type: Number, required: true },
// The tagging host whose chip rail to refresh after an accept. Defaults to
// the modal store (image modal); the Explore workspace passes its anchor host
// so the same panel refreshes the right surface. See TagPanel.
host: { type: Object, default: null },
})
// 'accepted'/'dismissed' let the parent return focus to the tag input after a
// suggestion is accepted OR rejected, so the operator keeps the keyboard flow on
// the input without re-clicking (operator-asked 2026-06-08, 2026-06-30).
const emit = defineEmits(['accepted', 'dismissed'])
// Reject (✗) / un-reject (↶): apply the store change, then signal the parent to
// re-focus the tag input — same return-to-input behaviour as accept.
function onDismiss (s) { store.dismiss(s); emit('dismissed') }
function onUndismiss (s) { store.undismiss(s); emit('dismissed') }
const props = defineProps({ imageId: { type: Number, required: true } })
// 'accepted' lets the parent return focus to the tag input after a suggestion is
// applied (operator-asked 2026-06-08).
const emit = defineEmits(['accepted'])
const store = useSuggestionsStore()
const modalStore = useModalStore()
const host = props.host || modalStore
const modal = useModalStore()
// 'artist' (FC-2d-vii-c) and 'copyright' (2026-06-01) retired as
// suggestion categories. Only 'character' remains as a people-style
@@ -96,7 +80,7 @@ watch(() => props.imageId, (id) => {
async function onAccept(s) {
try {
await store.accept(s)
await host.reloadTags()
await modal.reloadTags()
emit('accepted')
} catch (e) {
toast({ text: `Accept failed: ${e.message}`, type: 'error' })
@@ -110,7 +94,7 @@ async function onAliasConfirm(canonicalTagId) {
try {
await store.aliasAccept(aliasTarget.value, canonicalTagId)
aliasDialog.value = false
await host.reloadTags()
await modal.reloadTags()
emit('accepted')
} catch (e) {
toast({ text: `Alias failed: ${e.message}`, type: 'error' })
@@ -141,12 +125,6 @@ async function onRemoveAlias(s) {
color: rgb(var(--v-theme-on-surface));
margin-bottom: 8px;
}
.fc-suggestions__list {
/* No internal scroll cap: the rail now scrolls suggestions in its single
main scroll area while Related is pinned to the bottom (ImageViewer side
layout), so suggestions flow naturally without a nested scrollbar. */
padding-right: 4px;
}
.fc-suggestions__skeleton { display: flex; flex-direction: column; gap: 8px; }
.fc-suggestions__skel-row {
height: 18px; border-radius: 4px;
@@ -72,7 +72,7 @@
</v-icon>
</template>
<v-list-item-title>
{{ createLabel }}
Create "{{ parsedName }}" as {{ parsedKind }}
</v-list-item-title>
</v-list-item>
</template>
@@ -178,34 +178,14 @@ watch(query, () => {
}, 200)
})
// A same-name character ALREADY exists. Characters are unique by
// (name, kind, fandom), so this is still a valid distinct tag in another fandom.
const sameNameCharExists = computed(() =>
parsedKind.value === 'character' &&
hits.value.some(h =>
h.kind === 'character' && h.name.toLowerCase() === parsedName.value.toLowerCase(),
),
)
const allowCreate = computed(() => {
const q = parsedName.value
if (!q) return false
// Characters disambiguate by fandom, so a same-named character in a DIFFERENT
// fandom is a valid new tag — always offer Create (the fandom picker resolves
// it; find_or_create is idempotent if you re-pick the same fandom). Other
// kinds are unique by (name, kind): an exact match means it already exists.
if (parsedKind.value === 'character') return true
return !hits.value.some(h =>
h.name.toLowerCase() === q.toLowerCase() && h.kind === parsedKind.value,
)
})
const createLabel = computed(() =>
sameNameCharExists.value
? `Create another "${parsedName.value}" character (different fandom)`
: `Create "${parsedName.value}" as ${parsedKind.value}`,
)
function scorePct (s) { return `${Math.round(s.score * 100)}%` }
// This image's suggestions that match the typed query, minus any the server
@@ -222,10 +202,6 @@ const suggestionHits = computed(() => {
const out = []
for (const list of Object.values(suggestions.allByCategory)) {
for (const s of list || []) {
// Rejected suggestions now stay in allByCategory (flagged) so the panel
// can show + un-reject them; keep them OUT of the type-to-add dropdown,
// whose job is finding a tag to ADD (un-reject lives in the panel).
if (s.rejected) continue
const key = `${s.category}:${s.display_name.toLowerCase()}`
if (!s.display_name.toLowerCase().includes(q)) continue
if (seen.has(key)) continue
+1 -12
View File
@@ -3,10 +3,6 @@
<v-chip
size="small" closable
:color="store.colorFor(tag.kind)" variant="tonal"
class="fc-tag-chip__nav"
role="link"
:title="`Browse images tagged “${tag.name}”`"
@click="$emit('navigate', tag)"
@click:close="$emit('remove', tag.id)"
>
<v-icon start size="x-small">{{ iconFor(tag.kind) }}</v-icon>
@@ -34,7 +30,7 @@ import { useTagStore } from '../../stores/tags.js'
import KebabMenu from '../common/KebabMenu.vue'
const props = defineProps({ tag: { type: Object, required: true } })
defineEmits(['remove', 'rename', 'set-fandom', 'navigate'])
defineEmits(['remove', 'rename', 'set-fandom'])
const store = useTagStore()
@@ -57,13 +53,6 @@ function iconFor (k) { return KIND_ICONS[k] || 'mdi-tag' }
<style scoped>
.fc-tag-chip { display: inline-flex; align-items: center; gap: 1px; }
/* The chip body navigates to the filtered gallery (#5); signal it's clickable.
The close ✕ (remove) and the sibling kebab stay as the explicit controls. */
.fc-tag-chip__nav { cursor: pointer; }
.fc-tag-chip__nav:focus-visible {
outline: 2px solid rgb(var(--v-theme-accent));
outline-offset: 1px;
}
.fc-tag-chip__kebab { opacity: 0.7; }
.fc-tag-chip:hover .fc-tag-chip__kebab { opacity: 1; }
.fc-tag-chip__fandom { opacity: 0.7; font-size: 0.85em; }

Some files were not shown because too many files have changed in this diff Show More