Compare commits

...
Author SHA1 Message Date
bvandeusen 9eb946b21b ci(extension): sign on dev too, and bundle the XPI into :dev (step 6)
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 4s
CI / frontend-build (push) Successful in 24s
CI / backend-lint-and-test (push) Successful in 31s
Build images / build-ml (push) Successful in 2m40s
CI / integration (push) Successful in 3m55s
Build images / sign-extension (push) Successful in 4m43s
Build images / build-web (push) Successful in 2m11s
Build images / build-agent (push) Successful in 10m13s
The step the milestone exists for. sign-extension ungates from main-only
to main-or-dev, and build-web downloads the XPI on dev as well, so a dev
push produces an image carrying the extension that is being developed
rather than requiring a merge to try one.

Not two signatures. The version is the commit TIME of the newest packaged
extension change, so dev and main derive the SAME number for the same
source. A dev push that changes the extension signs it; the merge to main
finds the ext-<version> release already there, hits the cache, and bundles
the byte-identical XPI into :latest with no second AMO call. One signature
per extension CHANGE, shared by both channels. That property is what makes
two channels affordable at all, and it is why step 4 had to land first:
ungating this while the version was still the hand-set 1.0.11 would have
found the existing ext-1.0.11 release, skipped AMO, and bundled main's
stale XPI into :dev — a dev channel confidently serving old code.

Tags stay excluded. The tag path deliberately skips signing and polls for
the release instead (the 2026-05-27 race).

The ext-<version> release's target_commitish moves from the literal "main"
to $GITHUB_SHA. Either branch can create that release now, and tagging a
dev-signed XPI against a main commit that need not even contain the source
it was built from is a lie that costs nothing to avoid.

Known, not addressed here: two concurrent builds that both derive the same
unsigned version will both call AMO and the loser gets a 409. The window
already existed between main and tag pushes; dev signing widens it. It
fails loudly rather than shipping anything wrong, and the rollback trap
cleans up the empty release. Filed separately.

Also unchanged here: ci.yml's manual-bump guard is still in place and
still false. It does not fire on this commit — nothing packaged changed —
but it will fail the lane on the next extension change, demanding a bump
that no longer decides anything. Step 5 next.
2026-08-27 10:56:58 -04:00
bvandeusen 5447a40e97 ci(extension): the derived version drives signing (milestone 271 step 4)
CI / extension-version (push) Successful in 3s
Build images / build-agent (push) Successful in 7s
CI / backend-lint-and-test (push) Successful in 30s
extension / lint (push) Successful in 27s
Build images / sign-extension (push) Skipped
CI / lint (push) Successful in 3s
CI / frontend-build (push) Successful in 23s
Build images / build-web (push) Successful in 2m8s
Build images / build-ml (push) Successful in 2m48s
CI / integration (push) Successful in 3m52s
Cutover. sign-extension no longer reads the version out of the repo — it
runs packaging.sh version and stamps the result into manifest.json and
package.json in the working tree before web-ext sees them. Never
committed back: the commit carrying the bump would itself be a change to
the extension and would move the version again.

Shadow mode ends here, in both build.yml and ci.yml. It had one job —
validate the formula at zero cost before a real AMO version was burned —
and CI confirmed it on 239b1ed: shadow: manual=1.0.11 derived=1.0.3499884.

build-web re-derives rather than being handed the value, so it gains
fetch-depth: 0. It was the outstanding landmine: a depth-1 clone derives a
WRONG, too-low version rather than failing, and would then 404 fetching a
release that exists under its real name. sign-extension and
extension-version already had full history.

New guard, and it stays permanently: refuse to sign when the derived
version is strictly OLDER than the highest ext-* release already signed.
Firefox rejects a downgrade and AMO never releases a burned version, so
backwards is unrecoverable — it strands every install that took the higher
one. Strictly older, not older-or-equal: equality is the ordinary case,
an unchanged extension deriving the same version it did last build, which
is exactly what makes the ext-<version> cache hit and holds AMO to one
call per extension CHANGE rather than per push. The release list is
paginated because ext-* shares it with the v* tags, and the bound fails
rather than calling the highest it happened to see the highest there is.

First derived value is 1.0.3499884 against a highest-signed ext-1.0.10, so
the backfill direction is right by six orders of magnitude. 1.0.11 sits in
the repo and was never signed; nothing is stranded by skipping past it.

Still main-only. Step 6 ungates sign-extension to dev, which is what
actually puts an XPI on :dev.

Note for step 5: ci.yml's manual-bump guard is now false. It still demands
a hand bump when a packaged file changes, and that bump no longer decides
anything — the derived value overwrites it at build time. Harmless but
pointless, and it should be retired before the next extension change.
2026-08-27 10:45:17 -04:00
bvandeusen 239b1ed8d9 ci: build :dev images again so the dev channel can carry a build
Build images / sign-extension (push) Skipped
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 29s
extension / lint (push) Successful in 30s
Build images / build-web (push) Successful in 2m23s
Build images / build-ml (push) Successful in 3m20s
CI / integration (push) Successful in 3m52s
Build images / build-agent (push) Successful in 9m26s
build.yml triggered on main and tags only. The 2026-05-26 comment gave
the reason: "operator tests from :latest after merge-to-main, not from
the dev branch image. Saves one full docker build per dev push."

That trade has since been named as a fault. Family rule 147 — main IS
production, test on :dev, never by shipping — and rule 146 — a rolling
channel refreshes itself, and a channel that can only be refreshed by
shipping is not a channel. 146's note on 147 describes this exact shape:
the pressure to test by shipping does not come from carelessness, it
comes from :dev being unable to carry the build.

Two live consequences, not hypotheticals:
  - docker-compose.yml pins fabledcurator:dev, an image nothing has
    published since May. The registry-image path of the documented
    quick-start could not have worked.
  - trying an extension change required merging to main, because
    sign-extension is gated to main and :dev did not exist to carry an
    XPI. Shipping was the only way to test.

All three images build on dev. Deliberate: a :dev web image paired with
a stale :dev ml or agent is a worse trap than no dev channel, because
the mismatch surfaces as a runtime failure rather than a missing tag.
The cost the 2026-05-26 note was avoiding is real and is now paid on
every dev push — layer reuse should keep ml's cost to the COPY layers,
but if it bites, narrowing is a `paths:` filter away.

:dev only. The dev path never writes :c-<sha>: that is the rollback unit
(rule 145), and a rolling tag may legitimately carry newer contents than
the :c-<sha> of the same commit.

This does NOT yet put an XPI on :dev — sign-extension is still gated to
main, and ungating it has to wait for the derived version to control
publishing, or dev would sign the hand-set 1.0.11, hit the existing
cache and ship main's stale XPI. That is the next step.
2026-08-27 09:26:58 -04:00
bvandeusen cd5444e3ae ci(extension): derive the version from commit TIME, not commit count (#3092)
Rule 149: an artifact's ordering key must be time-derived, never a commit
count. packaging.sh's cmd_patch was a count.

Why that matters here rather than in the abstract. A count is per-branch:
dev and main count different histories of the SAME code. Today only main
signs, so nothing has ordered the two against each other and the fault is
invisible. The moment dev also publishes an extension, the two versions
order by which branch accumulated more commits rather than by which is
newer — and a squash-merge makes it permanent, because main gains one
commit where dev gained five. dev then climbs away from main and a dev
install can never cross back.

That is Roundtable's 2026-08-24 incident (Scribe #2993) in a different
repo: their versionCode was the branch's commit count, and it produced a
channel you could enter and not leave. Measured on this repo today the
old formula gives main=23, dev=24 — one apart, which is exactly how the
inversion stays invisible until it strands somebody.

New formula: minutes since 2020-01-01 of the LATEST commit touching a
packaged extension file. Same anchor and unit Roundtable settled on.

Commit time, not build time, and the difference is load-bearing:

  - stable while the extension is unchanged, so the ext-<version>
    signature cache still hits and AMO is called once per extension
    CHANGE rather than once per push. Build-time minutes would re-sign
    on every push and never let two channels share a signature.
  - after a merge, main sees the same commit and derives the same
    number, so :latest reuses the signature :dev already produced for
    byte-identical code. Same code, same version, one signing.
  - monotonic: max() over a set that only gains members. Verified
    across all 24 extension-touching commits, zero non-monotonic steps.
  - reproducible from any checkout.

Derives 1.0.3499884 on dev, 1.0.3465860 on main — both far above the
last hand-set 1.0.11, so milestone 271's backfill guard is satisfied by
construction rather than by an offset.

Still shadow-only: nothing reads the derived value yet. Both shadow
steps log it, and ci.yml's runs on dev too, so both channels' numbers
are visible — that is the pair that has to stay ordered. Prior shadow
observations describe the OLD formula and prove nothing about this one,
so the window restarts; ci.yml says so at the step.

New requirement recorded in ci-requirements.md: a depth-1 clone derives
a wrong, too-low value rather than failing, so fetch-depth: 0 is
load-bearing wherever packaging.sh version is called.

Refs #3092, milestone 271
2026-08-27 09:26:58 -04:00
bvandeusen 5a0e1bbd03 perf(ml): batch the auto-apply sweeps' image_tag inserts (#3072)
CI / extension-version (push) Successful in 3s
CI / lint (push) Successful in 3s
CI / frontend-build (push) Successful in 27s
CI / backend-lint-and-test (push) Successful in 32s
CI / integration (push) Successful in 4m2s
Item 1 of #3072. Both sweeps issued a single-row pg_insert(image_tag)
from inside their per-image loop. Steady state that is nothing; a first
sweep over a back-catalogue is one round-trip per applied tag, tens of
thousands of them. Each chunk now collects its rows and writes them in
one statement.

The ticket suggested one insert per chunk PER TAG. A single multi-row
VALUES carries every tag at once, so it is one statement per chunk full
stop — and the sweeps already accumulate across all heads before they
commit, so nothing had to be restructured to allow it.

Not a new helper: wip_title.apply_wip_image_tags was already doing the
chunked ON CONFLICT DO NOTHING insert, so that shape is extracted to
services/image_tag_apply.insert_image_tags and all three writers share
it. The extraction deliberately leaves wip_title's pre-SELECT behind
rather than pulling it into the shared function — the sweeps don't need
it (their `skip` sets already exclude applied and rejected images) and
it exists only to produce an accurate count, which the sweeps also
compute themselves. So the shared primitive returns nothing: psycopg
reports rowcount -1 for a multi-row ON CONFLICT DO NOTHING insert, and a
count taken from the statement would be a lie rather than an
approximation.

Ordering note for the system-tag sweep: tag rows are now written after
that chunk's PresentationReview rows rather than interleaved before
them. Safe — PresentationReview FKs to image_record and tag, not to
image_tag.

Chunk size stays 5000: 5000 rows x 3 bound params = 15000, inside
Postgres' 65535-parameter ceiling with room to spare.

tests/test_image_tag_apply.py covers the primitive directly, since it is
now the single place three writers can be wrong at once — most
importantly that a re-run never restamps a hand-applied tag's source,
which would silently poison head training (it excludes the auto
sources).

Left alone: _insert_presentation_review is still per-row, and the
retract path still deletes per-row. Both operate on sets that are small
by construction, unlike the apply path.

Refs #3072
2026-08-27 07:51:23 -04:00
bvandeusen 1ac448d881 refactor: four small cleanups from the review pass (#3072)
CI / extension-version (push) Successful in 2s
CI / backend-lint-and-test (push) Successful in 27s
CI / integration (push) Successful in 4m57s
CI / lint (push) Successful in 3s
CI / frontend-build (push) Successful in 23s
Items 2-5 of #3072. Item 1 (the per-row sweep inserts) is separate.

2. .fc-bad was not merely duplicated — it is .fc-weak under a second
   name. Both local definitions were `color: rgb(var(--v-theme-error))`,
   identical to the global .fc-weak, and GpuAgentCard was already using
   .fc-weak to colour exactly what GpuActivityPanel coloured .fc-bad (an
   errored count, red when non-zero). So rather than promoting a synonym
   to app.css, both call sites now use .fc-weak and the local defs are
   gone. app.css's status-colour comment records why there is no .fc-bad,
   next to the existing note on why .fc-ok is deliberately NOT global.

3. GalleryItem.vue's obsidian literals now use --v-theme-background,
   which IS obsidian (vuetify-theme.js maps background -> surfaces.
   obsidian). Preferred over --fc-chrome-rgb: same value, but that
   variable is named for the nav fade, not for the palette entry.

   The ticket said these were the only three real uses in the tree. They
   are not — GalleryItem itself had two more in the artist-label
   gradient (fixed here, so the file is now consistent), and ~13 more
   live in SeriesView, SeriesReaderView, ImageViewer, ArtistHeader,
   ExploreView and GalleryFilterBar. Those are a separate sweep, filed
   rather than folded in here.

4. The attachment download path had two hand-formatted copies. One
   definition now, `attachment_download_url`, next to the model both
   serializers already import. The test pins it by MATCHING the built
   path against the app's real URL map rather than comparing to a
   literal — a string-equality test would still pass after someone
   renamed the route, which is the drift the helper exists to prevent.

5. Extension API key now compares with hmac.compare_digest. Compared as
   BYTES, not str: compare_digest's str form raises TypeError on
   non-ASCII, and this value comes straight from an attacker-controlled
   header, so the str form would turn a junk key into a 500 instead of a
   403. Low stakes either way — the API is unauthenticated-by-design on
   a LAN — but it costs nothing.

Refs #3072
2026-08-27 07:48:18 -04:00
bvandeusen bfc5135f19 docs: true up README — status, the five pieces, CI (#3070)
CI / extension-version (push) Successful in 2s
CI / lint (push) Successful in 3s
CI / frontend-build (push) Successful in 25s
CI / backend-lint-and-test (push) Successful in 28s
CI / integration (push) Successful in 5m0s
README.md was last touched in 4aff9c5 (2026-05-14) and several of its
most visible lines had gone false:

- "Pre-v1. Not yet functional." — FC has been continuously deployed for
  months. Replaced with what main/dev actually mean for what's running.
- "Node 22 pre-installed" — ci-requirements.md and extension.yml both say
  node 24, and frontend/package.json requires >=24.
- "Runner label python-ci — a runner with Python 3.14, ruff and Node 22
  pre-installed ... The runner image (runner-base:python-ci) is built
  from CI-Runner/CI-python/" — describes runs-on as selecting the
  toolchain. It doesn't: runs-on is a scheduling label, and every job
  names its own container.image (ci-python:3.14, or node:24-bookworm-slim
  for the extension lane).
- "Both ci.yml and build.yml use this label" — there are three workflows.

The CI section now points at ci-requirements.md rather than restating
it, so the two can't drift apart again; that file is current and is the
one the CI-runner process expects.

Added a "What's in here" table for the five deployable pieces — the
extension, the GPU agent and the ML image were unmentioned, three of the
five. Also corrected the RELEASE_TOKEN write:release scope, which is no
longer "for future release-cutting workflows": it backs the ext-<version>
releases that cache the signed XPI.

Refs #3070
2026-08-27 07:39:49 -04:00
bvandeusenandClaude Opus 5 89155478a8 test(refetch): cover the Layer-2 auto-refetch remediation (#3071)
CI / extension-version (push) Successful in 2s
CI / lint (push) Successful in 2s
CI / frontend-build (push) Successful in 24s
CI / backend-lint-and-test (push) Successful in 28s
CI / integration (push) Successful in 3m59s
refetch_service was the only module under backend/app/services/ with no
test file — and not an inert one: it runs unattended off the recovery
sweep and deletes a file from disk before asking a downloader to replace
it. The frontend cites it by name as the reason the Import tab could be
retired ("imports heal themselves").

The ticket described it as having zero direct coverage. That is true of
the module, but not of the code: test_api_import_admin.py already drives
the happy path end-to-end through the refetch route — file deleted, task
flagged, one dispatch, second attempt a no-op. These 17 tests therefore
target what the route tests cannot reach rather than restating them:

  * every branch of resolve_refetch_source — disabled Source, a
    `sidecar:<platform>:<slug>` synthetic anchor, a platform mismatch,
    the gallery-dl `NN_` numbering-prefix sidecar, the lowest-id pick
    among several candidates, and each of the five ways it declines
    (no sidecar, unreadable JSON, non-object JSON, no platform, no
    artist folder / no matching Artist row).

  * that the file SURVIVES when nothing re-pollable resolves. This is
    the assertion the module exists for: `no_source` is the common case
    on a filesystem-only library, where the file on disk is the
    operator's only copy. The route-level no_source test cannot catch a
    regression here — its path never existed, so an unconditional unlink
    would pass it.

  * that the `refetched` bound is checked BEFORE the unlink, so a second
    sweep leaves the re-downloaded file alone rather than deleting it
    again.

  * that an unlink failure is logged and stepped over, not raised —
    a raise would abort the whole sweep for every other poison-pill row
    in the batch. Exercised with a real IsADirectoryError rather than a
    patched pathlib.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 07:33:14 -04:00
bvandeusenandClaude Opus 5 516521e7b0 refactor(platforms): drop migration 0088 — no deviantart rows exist (#3069)
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
CI / integration (push) Successful in 3m43s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 28s
Operator confirms the instance has never used DeviantArt, so there is
nothing for 0088 to quiesce. The migration only ever had two jobs —
disable leftover `source` rows and delete a stale `credential` row — and
both were guards against data that does not exist here.

Removing it rather than keeping a no-op: a migration that runs on every
deploy to touch zero rows is a permanent cost paid for a hypothetical,
and it would read to a future reader as evidence that DeviantArt sources
once existed. `platform` has no CHECK constraint, so retiring the key
needs no schema change of its own.

alembic head returns to 0087.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 07:26:02 -04:00
bvandeusenandClaude Opus 5 ddf896078c refactor(platforms): retire deviantart end-to-end (#3069)
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 4s
CI / frontend-build (push) Successful in 23s
extension / lint (push) Successful in 26s
CI / backend-lint-and-test (push) Successful in 28s
CI / integration (push) Successful in 3m43s
Executes the 2026-07-05 product decision (FC downloaders = art-dedicated
services only), which removed Twitter/X and Bluesky but left deviantart
fully wired for seven weeks — the half-retired state rule 22 exists to
prevent.

Removed: the PlatformInfo module and its registry entry, the gallery-dl
extractor block, extension_service's artist-page pattern, the extension's
PLATFORMS + PLATFORM_ARTIST_PATTERNS entries, its manifest host permission
and content-script match, the frontend icon/colour/label, and the operator-
facing "supported platforms" list that still advertised it.

Two judgment calls, both recorded in migration 0088:

  * existing `source` rows are DISABLED, not deleted. The row is the only
    record of the artist's DeviantArt URL. Disabling is also required for
    correctness rather than tidiness: with the platform unregistered the
    download path falls through to gallery-dl, which carries its OWN
    deviantart extractor, so an enabled row would have kept downloading
    from a dropped platform.
  * the `credential` row IS deleted — a live session cookie for a site FC
    will never call again.

Adds the invariant whose absence is why manifest.json drifted in the first
place: nothing tied its domain lists back to the platform table. The
extension suite now asserts both directions, plus that no host permission
belongs to an unclaimed domain (`*://*/*` exempted — FC is self-hosted at
an operator-chosen URL the extension cannot enumerate).

Extension version 1.0.10 -> 1.0.11: ci.yml's guard hard-fails a packaged
extension change without a bump. No release is cut — build.yml's
sign-extension job only runs on main.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 07:24:08 -04:00
bvandeusenandClaude Opus 5 2e0f8f8c61 feat(cleanup): reclaim orphaned attachments — rows and store blobs (#3068)
CI / extension-version (push) Successful in 3s
CI / backend-lint-and-test (push) Successful in 28s
CI / lint (push) Successful in 2s
CI / frontend-build (push) Successful in 24s
CI / integration (push) Successful in 3m47s
PostAttachment's two FKs are both ON DELETE SET NULL, so a deleted post or
artist left the row behind rather than taking it. Nothing ever pruned those
rows, and nothing in the repo had ever unlinked a file under the attachment
store — so both rows and bytes accumulated permanently, invisible to every
existing diagnostic.

Why a disk->DB reconciliation rather than a row sweep: the store is
sha-addressed and idempotent, so ONE blob backs MANY rows. Deleting a row
does not free its blob, and since the artist cascade (#3066) now deletes
its attachment rows outright, a freed blob has no DB pointer left to find
it by. Walking the store and asking "does any row still reference this
sha?" catches orphans from every cause, including ones no future delete
path will think to report.

Preview and apply share `_orphan_attachment_conditions` (rule 93). The
dry-run derives its surviving-sha set by NEGATING that same predicate, so
it is honest about blobs the delete would free rather than counting them as
still-referenced — the one place this was easy to get backwards, so it has
its own parity test.

Guards, each with a reason:
- A blob is written before its row commits, so a just-stored file legitimately
  has no referencing row. Files under 6h are never judged — same guard and
  reasoning as ORPHAN_TEMP_MIN_AGE_HOURS.
- `.partial` staging files belong to cleanup_orphaned_temp_files; skipped
  rather than raced.
- The sha is parsed as the first 64 chars, not via Path.stem: store() takes
  the extension from the source filename, and a URL-encoded basename yields a
  multi-dot suffix that would make stem eat part of the sha.
- A 900s walk budget reports partial=True instead of running to the task's
  hard limit (rule 89).
- TASK_STUCK_THRESHOLD_MINUTES override at 30 (= time_limit 25 + 5). Without
  it a healthy 20-minute walk is phantom-flagged 'RecoverySweep' at the bare
  5-min default — the #883 failure class; its invariant test is mirrored here.

Defaults to the safe preview at both the task and the route, unlike the other
maintenance triggers: this apply unlinks files. Operator-triggered only,
never on a beat.

Ships with its UI (rule 27): AttachmentReclaimCard in Cleanup → Duplicates &
leftovers, built on the existing useMaintenanceTask/MaintenanceTile shapes, so
a run survives navigating away. Surfaces files_failed and partial explicitly,
since both change what the numbers mean.

Also promotes humanBytes to utils/bytes.js — it was byte-identical in
VideoDedupCard and GatedPurgeCard and this card would have been the third
copy. The three divergent `formatBytes` helpers are deliberately left alone.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 22:37:24 -04:00
bvandeusenandClaude Opus 5 2ce467e347 fix(cleanup): artist cascade preview counts posts and attachments (#3067)
CI / lint (push) Successful in 5s
CI / backend-lint-and-test (push) Successful in 34s
CI / extension-version (push) Successful in 5s
CI / frontend-build (push) Successful in 23s
CI / integration (push) Successful in 4m32s
`project_artist_cascade` is documented as "a read-only projection of what
delete_artist_cascade would touch" and drives the Tier-C confirm dialog,
but it counted only images, sources, thumbs, import_tasks and bytes. It
never counted posts or attachments, both of which the apply destroys.

That is silent in the worst case. `Post.artist_id` is ondelete=CASCADE, so
every post goes whether or not it carried an image — and FC has a large
body-only post population (#1288 measured 694 pixiv posts with text and no
images). Such an artist previewed as `images: 0`, reading as "empty, safe
to remove", while the apply destroyed every captured body, description,
external-link set and raw_metadata snapshot. The danger-zone card already
promised "every image, source, post, and attachment" — the copy was honest
and the numbers were not.

Root of the drift: the preview re-derived its own predicates instead of
sharing the apply's, the same shape as the 2026-06-08 fandom-tag deletion
that rule 93 exists for. So rather than bolt on two counts, both halves now
build from shared `_artist_{images,posts,attachments}_conditions` helpers,
following the `_unused_tag_conditions` / `_bare_post_conditions` style
already in the file. Sources keep no helper — the apply doesn't query them
either, it gets them from the Artist.sources ORM cascade.

The apply also now reports `posts_deleted` (counted before the delete,
since the CASCADE leaves nothing to count after). Rule 93's second half
asks for the apply to be tested, and parity is only assertable if both
halves state the number.

Adds a preview/apply parity test that runs both against one artist and
asserts the three pairs agree AND that the rows actually went, plus a
body-only-artist test covering the case that motivated this.

The confirm dialog's counts grid renders every key, so posts and
attachments surface there automatically; the prose summary line names
posts explicitly, since that is the number that changes how an artist
reads.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 22:23:29 -04:00
bvandeusenandClaude Opus 5 39cf81aea6 fix(cleanup): clear an artist's attachments before the cascade delete (#3066)
CI / lint (push) Successful in 2s
CI / extension-version (push) Successful in 3s
CI / frontend-build (push) Successful in 24s
CI / backend-lint-and-test (push) Successful in 28s
CI / integration (push) Successful in 3m49s
`delete_artist_cascade` could abort partway through, and it aborted after
the irreversible half. Deleting an artist CASCADEs to Post (post.artist_id
is ondelete=CASCADE), which SET NULLs post_attachment.post_id — and
`uq_post_attachment_null_post_sha` is a partial UNIQUE on sha256 ALONE
WHERE post_id IS NULL. So any two of that artist's attachments sharing a
sha collapse onto one another and raise.

That shape is ordinary, not corrupt: `_capture_attachment` deliberately
writes one row per post over a single sha-addressed blob, so a creator who
attaches the same pdf to two posts already has two such rows. A
pre-existing filesystem-import row (post_id NULL) with the same sha
collides on its own.

The images and their on-disk files are deleted and committed in 500-row
batches BEFORE the artist row is touched, so the failure landed after
them: images gone, artist and posts alive, files unrecoverable.

Fix: delete the artist's post_attachment rows explicitly first, matched by
artist_id OR by the owning post's artist (artist_id is nullable, so neither
arm alone covers every row). `_repoint_post_links` already guards the
identical collision class in the reconcile path; this is its artist-cascade
counterpart. Migration 0043 reasoned only about upgrade-time safety and
never about this later SET NULL.

The sha-addressed blobs are deliberately left on disk: one blob backs many
rows, so unlinking needs a refcount pass, and this Tier-C op must not
delete bytes its own preview never disclosed.

Adds `attachments_deleted` to the summary, and two regression tests — the
same sha on two posts, and an unrelated NULL-post row that must survive.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 22:02:51 -04:00
Claude 11dd324f89 fix(extension): exclude the test/ and scripts/ directory entries from the XPI
CI / lint (push) Successful in 2s
CI / extension-version (push) Successful in 3s
extension / lint (push) Successful in 19s
CI / frontend-build (push) Successful in 20s
CI / backend-lint-and-test (push) Successful in 38s
CI / integration (push) Successful in 3m57s
extension / lint (pull_request) Successful in 34s
The XPI-content check added in 1c6452e did its job on its first run: the
archive carried empty `test/` and `scripts/` entries.

`test/**` matches the files inside a directory but not the directory entry
itself, and web-ext writes an entry for the directory separately -- so the
contents were correctly excluded while the empty directories shipped anyway.
Nothing harmful reached users (no dev code, just two empty entries), but our
single declaration claimed these do not ship and something was shipping.

Both forms are now listed per directory. The bare name alone would not do:
minimatch's `test` does not match `test/url.spec.js`, so dropping the glob
would ship the contents instead.

Verified locally: the derived version is unchanged at 1.0.19, confirming the
added entries are redundant for the git pathspec and only affect web-ext.

Refs #2400

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 16:20:54 -04:00
Claude 1c6452e10e ci(extension): shadow the derived version + verify real XPI contents
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
CI / frontend-build (push) Successful in 21s
extension / lint (push) Failing after 28s
CI / backend-lint-and-test (push) Successful in 47s
CI / integration (push) Successful in 4m1s
Milestone #271 steps 2 and 3. Neither changes what gets published.

STEP 2 -- shadow mode.

build.yml's sign-extension and ci.yml's extension-version guard now log the
version that WOULD be derived from git history alongside the hand-maintained
one. Nothing reads the derived value, and neither site can fail because of it.

This exists because `web-ext sign` is one-shot per version: AMO 409s on a
repeat, so a wrong formula burns a real version number that cannot be
reclaimed. Comparing the two across real builds is the only way to validate it
at zero cost. sign-extension runs on main only, so main pushes are the sole
source of truth for whether the derived number moves exactly when the shipped
extension changes -- the dev-side log is a convenience, not the evidence.

sign-extension now checks out with fetch-depth: 0. The derived version is a
commit count and a depth-1 clone cannot produce one.

STEP 3 -- XPI content verification.

Every other packaging assertion checks our declaration against itself. This is
the first that asks web-ext what it ACTUALLY wrote into the archive.

That assumption was both unverified and fragile: `test/**` only survives to
web-ext because callers `set -f` before substituting it, so losing that
quoting would silently start shipping dev files with no other signal. The step
builds the XPI and asserts test/, scripts/, vitest.config.js, package.json,
package-lock.json, README.md and node_modules are absent -- and, because an
over-matching exclusion would break the extension at runtime rather than at
build time, that manifest.json, all four lib/*.js and every UI directory are
present.

unzip is installed only when missing; node:24-bookworm-slim may not carry it.

Refs #2399, #2400

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 16:15:12 -04:00
Claude 597b91d29b refactor(extension): one definition of what ships in the XPI
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 4s
extension / lint (push) Successful in 20s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 44s
CI / integration (push) Successful in 3m59s
Milestone #271 step 1. Groundwork for deriving the extension version from git;
no behavior change yet -- nothing consumes `version` so far.

"Which files end up in the XPI" was stated in two places and about to become
three. Three hand-kept copies of one fact is what allowed #2397, where the
publish path could republish a stale XPI because its cache key had no link to
the content it stood for.

New extension/scripts/packaging.sh holds the single declaration and exposes:
  ignore       web-ext --ignore-files values
  pathspec     :(exclude)extension/... for git
  version      <MAJOR.MINOR from manifest>.<commit count over packaged files>
  major-minor  /  patch

Consumers now delegate instead of restating it:
- extension/package.json -- all four web-ext scripts
- .forgejo/workflows/ci.yml -- the extension-version guard's exclusions
- (step 4) the rev-list that derives the version

scripts/** joins the non-packaged set; the script must not ship to users.

Two shell hazards, both load-bearing:

The script runs `set -euf`. Its lists are iterated with deliberate word
splitting, and without -f the shell ALSO globs them -- invoking `pathspec`
from a directory where test/ exists (exactly how ci.yml calls it) would expand
`test/**` into the individual spec files and silently stop covering anything
added later. A caller's own `set -f` cannot prevent this: the script is a
separate sh process and does not inherit it.

Callers additionally need their own `set -f` for the substituted RESULT, which
is a different expansion. version.spec.js asserts every --ignore-files caller
sets it, that the pathspec comes through with `test/**` literal and no
.spec.js paths, and that neither consumer has reinstated a hardcoded list --
the easy future regression is "simplifying" by inlining one again.

Verified: all five subcommands plus the usage/exit-2 path. Derived version on
main (8300029) is 1.0.19, matching dev. Last published is 1.0.10, so the
eventual cutover moves strictly upward and needs no offset -- Firefox refuses
downgrades. (An earlier note recorded 18; that was measured against a stale
origin/main from before the PR #234 merge.)

Refs #2398

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 15:33:11 -04:00
Claude f9111c06a7 test(extension): unit suite for lib/ + version-consistency specs
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 4s
CI / frontend-build (push) Successful in 21s
extension / lint (push) Successful in 19s
CI / backend-lint-and-test (push) Successful in 43s
CI / integration (push) Successful in 3m53s
extension / lint (pull_request) Successful in 19s
extension/ had no test harness at all -- web-ext lint was the only signal, so
the URL-normalization fix in 8214afe shipped with nothing exercising it.

Adds vitest (mirroring frontend/vitest.config.js) and three specs:

- url.spec.js       normalizeApiUrl / webRootFromApiUrl, including the #2393
                    regression: instance-root input must reach /api/credentials,
                    idempotence, trailing-slash and whitespace handling, and
                    that empty input never yields a bare "/api" (which
                    isConfigured() would read as configured).
- platforms.spec.js getPlatformFromUrl / isArtistPage, pinning the #1485
                    regression -- all three Patreon creator URL shapes
                    (bare, /c/, /cw/) plus inner pages, with nav pages
                    excluded -- and table-integrity checks.
- version.spec.js   manifest.json and package.json versions in lockstep,
                    AMO-safe version format, and url.js ordered before api.js
                    in background.scripts (classic scripts share one scope, so
                    a reorder is a runtime ReferenceError with no build signal).

Specs load lib/*.js by evaluating the real file as a classic script
(test/helpers/loadLib.js) instead of adding module.exports shims to production
code that would never run in the browser. The suite therefore exercises exactly
the bytes packaged into the XPI.

Two packaging consequences, both handled:

- web-ext would otherwise bundle test/ and vitest.config.js INTO the XPI;
  both are now in --ignore-files across all four web-ext scripts.
- ci.yml's extension-version guard must ignore the same paths, or editing a
  spec would demand a pointless version bump. The dangerous drift direction is
  the opposite one -- a guard exclusion for a file that DOES ship would let a
  real change pass unnoticed -- so version.spec.js asserts every :(exclude) in
  ci.yml appears in --ignore-files.

extension.yml also triggers on ci.yml now, since version.spec.js reads it.

Refs #2393, #2397

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 23:47:18 -04:00
Claude c37a180c3c ci: guard the extension publish path against a missed version bump
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 2s
CI / frontend-build (push) Successful in 26s
CI / backend-lint-and-test (push) Successful in 45s
CI / integration (push) Successful in 3m59s
build.yml's sign-extension keys its AMO-signing cache purely on the version
string in extension/package.json. If an ext-<version> release already has an
XPI, signing is skipped and build-web bakes that OLD signed XPI into :latest.
Nothing in that path inspects whether extension/ actually changed, so a
forgotten bump ships a stale extension on a fully green build -- silently, and
as the default outcome of forgetting. AMO can't backstop it either: it 409s on
re-signing a version, which is precisely why the cache exists.

New extension-version job, pure git + text, no deps or services:

1. Unconditional consistency check. manifest.json and package.json versions
   must match. web-ext sign reads manifest.json (package.json is in
   --ignore-files and isn't even inside the XPI), so AMO signs the manifest
   version; build.yml keys its cache, release tag, XPI filename -- and so the
   version /api/extension/manifest reports to the update prompt -- on
   package.json. Divergence either 409s at AMO or ships an XPI whose update
   prompt lies about what's installed.

2. Changed-without-bump check. If any PACKAGED file under extension/ differs,
   the version must have moved. Exclusions mirror --ignore-files so a Renovate
   web-ext devDep bump in package.json doesn't falsely demand one.

Compared against main rather than the previous push: the publish decision is
made at merge-to-main against whatever ext-<version> exists, so "differs from
main" is the question that matters. Diffing against the previous dev push
would demand a fresh bump on every iteration, inflating the version to buy
nothing.

Bumping stays manual -- making it automatic requires rewriting the version in
CI and committing back to a protected branch, which this workflow deliberately
avoided. This only ensures a missed bump can no longer be silent.

Refs #2393

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 20:20:44 -04:00
Claude 8214afee1e fix(extension): normalize FC URL so credential push doesn't 405
CI / lint (push) Successful in 3s
extension / lint (push) Successful in 38s
CI / frontend-build (push) Successful in 40s
CI / backend-lint-and-test (push) Successful in 2m58s
CI / integration (push) Successful in 4m56s
The stored apiUrl was required to already carry the `/api` suffix, since
api.js builds requests as `${baseUrl}/credentials`. The options label read
"FC base URL", so entering the instance root -- the natural reading --
sent every request one path segment short: POST /credentials hit the Vue
SPA catch-all and came back 405, and GET /extension/manifest 404'd.

Worse, Test Connection reported success on it: the catch-all answers GET
/credentials with 200 HTML, so `r.ok` was true and the only affordance
meant to catch this misconfiguration actively masked it.

Normalize instead of validate (rules 92, 26):

- New lib/url.js: normalizeApiUrl / webRootFromApiUrl, one source shared
  by the background client and the options page. Accepts either the
  instance root or the API root.
- api.js normalizes on read, so configs already stored in the broken form
  heal themselves without the operator reopening Settings.
- options.js stores the canonical form, echoes back what it saved, and
  the test now asserts a JSON content-type -- killing the false green.
- 404/405 in request() now names the URL and points at the setting.
- Options label/placeholder state that both forms work.

Version 1.0.9 -> 1.0.10 in BOTH manifest.json and package.json; build.yml
resolves the release version from package.json, and a stale value there
would hit the cached ext-1.0.9 asset and republish the old XPI unsigned
against the new code.

Refs #2393

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 19:37:38 -04:00
bvandeusen 306de50f61 docs: Fabled-Git, not Forgejo, in ci-requirements
CI / lint (push) Successful in 3s
CI / backend-lint-and-test (push) Successful in 52s
CI / frontend-build (push) Successful in 25s
CI / integration (push) Successful in 4m14s
The instance has run Gitea since the migration. Also fixes a dead rulebook
pointer: the topic was renamed forgejo.md -> fabled-git.md, so the
"CI philosophy" reference pointed at a file that no longer exists.

Prose only — no workflow or path change. Scribe issue #2272.
2026-07-31 23:42:47 -04:00
bvandeusenandClaude Opus 4.8 57e52433d0 feat(agent): idle-unload GPU models to free VRAM when the queue is idle
CI / lint (push) Successful in 3s
CI / frontend-build (push) Successful in 22s
CI / backend-lint-and-test (push) Successful in 38s
CI / integration (push) Successful in 4m3s
The SigLIP embedder + YOLO proposers load lazily then stay resident for the
container's whole lifetime — a 24/7 agent with an empty queue squats on ~5GB of
VRAM doing nothing (operator-observed: 4900MiB held at GPU-util 8% / P8). Sleep
mode only sheds downloaders + poll cadence; even a UI Stop left the models loaded.

Add a monitor thread that unloads the torch-owned models after
cfg.idle_unload_seconds (env IDLE_UNLOAD_SECONDS, default 300; 0 disables) with
the GPU genuinely idle (active==0, buffer drained, no job completed in the
window), then torch.cuda.empty_cache() to hand the blocks back to the driver.
They reload lazily on the next job via the existing _ensure_embedder /
_proposers_for. Covers both sleep-mode idle and a full Stop. Surfaced in
/status (models_loaded) and the agent UI pipe line; the VRAM meter drops too.

Residual: imgutils CCIP/person ONNX sessions + the CUDA context stay resident
(no clean unload API) — idle VRAM drops substantially, not to zero.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TbrA36zNczjVhrM6cWThQa
2026-07-17 12:57:31 -04:00
bvandeusenandClaude Opus 4.8 ec66ea5f83 refactor(ui): settings-card primitives + fix threshold clamp / card misgroup (#161)
CI / lint (push) Successful in 2s
CI / frontend-build (push) Successful in 22s
CI / backend-lint-and-test (push) Successful in 33s
CI / integration (push) Successful in 3m52s
extension / lint (pull_request) Successful in 10s
Tier-3 frontend DRY for the ML settings cards, plus the F-D2 clamp bug and the
F-D3 card misgrouping.

New primitives (components/common + composables):
- <SettingToggleRow> — the accent-icon + .fc-section-h label + right-aligned
  switch row (HeadsCard x3, CropProposersCard). iconColor prop absorbs the
  on/off dim.
- <SettingNumberField> — compact numeric field that CLAMPS to [min,max] on
  commit. This fixes F-D2: HeadsCard/CropProposersCard previously sent
  Number(raw) straight to the API, so an out-of-range threshold bounced off the
  400 validator (only TranslationCard clamped). density prop for the grid cards.
- useSettingSave(patchFn) — the busy + patch + toast + revert-on-failure flow
  each card hand-rolled (HeadsCard x6 handlers, CropProposersCard, MLBackfillCard,
  VideoEmbeddingCard). Returns ok/false for the optimistic-switch revert.

Adopted in HeadsCard, CropProposersCard, MLBackfillCard (handler only — its
plain labelled switch is a different affordance), VideoEmbeddingCard.

F-D3: MLThresholdSliders.vue actually rendered a "Video embedding" (frame-
sampling) card but sat under "Tagging → Suggestion thresholds". Renamed it
VideoEmbeddingCard.vue and moved it to the "GPU agent & embeddings" section.

Left deliberately (over-DRY guard): TranslationCard uses an inline error ALERT
(not a toast), already clamps its confidence with a NaN fallback, and lives on
the ImportStore — a genuinely different save pattern, so forcing it onto
useSettingSave would change its UX.

Behaviour-preserving refactor; CI has no Vue type-check so this needs a live
UI pass (toggles persist + revert on failure, thresholds clamp on blur, video
card now under Embeddings).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NsmJSQxnNxGgtM5Yz4GAqi
2026-07-13 22:28:47 -04:00
bvandeusenandClaude Opus 4.8 e92570a31e refactor(extension): DRY the web-root transform + cookie-export flow (#161)
CI / lint (push) Successful in 2s
extension / lint (push) Successful in 11s
CI / frontend-build (push) Successful in 22s
CI / backend-lint-and-test (push) Successful in 30s
CI / integration (push) Successful in 3m53s
Behavior-preserving (extension JS has lint-only CI + your manual test):
- api.webRoot() single-sources the baseUrl→web-root transform (strip trailing
  slash + /api) that was copy-pasted in background.js's self-update check and
  OPEN_ARTIST_PAGE, whose comments even cross-referenced each other.
- exportPlatformCookies(key) shares the extract→verify→upload spine between
  EXPORT_COOKIES (single) and EXPORT_ALL_COOKIES; it returns a structured
  outcome so each caller keeps its own response/skip messages verbatim.
- popup.js mutedNote(text) replaces the "centered muted note" div hand-rolled
  in the platform-loading, sources-loading, and empty-sources renderers.

No version bump — no behavior change, so it rides the next real ext release
rather than forcing an AMO re-sign + reinstall.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NsmJSQxnNxGgtM5Yz4GAqi
2026-07-13 21:57:14 -04:00
bvandeusenandClaude Opus 4.8 a2d1ed935d refactor(ui): DRY the settings-card CSS tokens + fix unstyled headers (#161)
CI / lint (push) Successful in 3s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 35s
CI / integration (push) Successful in 3m57s
- Promote .fc-section-h to a global token (app.css). It was copied identically
  into 4 cards, and TranslationCard used the class with NO local def — so its
  section headers rendered unstyled. Now fixed everywhere.
- Promote .fc-good / .fc-weak status colours to globals; delete the local copies
  in the GPU/heads cards. (.fc-ok stays local — divergent: on-surface in
  HeadsCard vs success in QueuesTable. .fc-bad stays — different name.)
- Delete 10 identical local .fc-muted redefinitions that crept back after the
  2026-06-09 sweep; the global utility already covers them.
- DbMaintenanceCard: opacity:0.6 muted text → the .fc-muted token (the exact
  anti-pattern that token's comment forbids).
- HeadsCard: collapse byte-identical ratePct() into pct().

CSS-only + one template class swap; no logic change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NsmJSQxnNxGgtM5Yz4GAqi
2026-07-13 21:51:23 -04:00
bvandeusenandClaude Opus 4.8 05df51b749 fix(ml): drop unnecessary quotes on MLSettings.load annotations (UP037)
CI / lint (push) Successful in 3s
CI / frontend-build (push) Successful in 28s
CI / backend-lint-and-test (push) Successful in 38s
CI / integration (push) Successful in 3m51s
Python 3.14 evaluates annotations lazily, so the self-referential return
type needs no forward-ref quotes — matches ImportSettings.load. Fixes the
ruff lint lane on the DRY-pass push.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NsmJSQxnNxGgtM5Yz4GAqi
2026-07-13 21:49:30 -04:00
bvandeusenandClaude Opus 4.8 099e1e664c refactor(patreon): DRY the campaigns-API request (#161)
CI / lint (push) Failing after 2s
CI / backend-lint-and-test (push) Successful in 28s
CI / frontend-build (push) Successful in 29s
CI / integration (push) Successful in 4m0s
_lookup_via_api and resolve_display_name shared ~90% of their body (same
endpoint, params, headers, error handling — differing only in which field
they pluck from data[0]). Extract _campaigns_api_first(vanity, cookies_path)
-> dict|None; callers pluck the campaign id vs the display name. Return-value
behavior preserved (the display-name path additionally gains the helper's more
granular warning logs). Covered by the existing test_patreon_resolver.py.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NsmJSQxnNxGgtM5Yz4GAqi
2026-07-13 21:41:34 -04:00
bvandeusenandClaude Opus 4.8 c87f8a1bb3 refactor(maintenance): DRY the recover_stalled head-run twins (#161)
recover_stalled_head_training_runs and recover_stalled_head_auto_apply_runs
were near-exact copies (coalesce-flip + keep-last-N prune, differing only in
model + two constants). Extract _recover_stalled_runs(model, stall_minutes,
keep_runs, label); the two tasks become thin wrappers. The other two recover
tasks are deliberately NOT folded in (library-audit has no prune tail; backup
uses a single started_at cutoff). test_recover_stalled_head_runs.py covers
both wrappers (stalled→error, fresh survives) — previously untested.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NsmJSQxnNxGgtM5Yz4GAqi
2026-07-13 21:41:34 -04:00
bvandeusenandClaude Opus 4.8 666b3a2ec8 refactor(ml): DRY pass — shared sweep helpers + table-driven settings (#161)
Consolidate duplication accrued across the ML tagging + settings backend,
behavior-preserving (over-DRY guard applied — the three auto-apply sweep
BODIES stay separate; only their shared inner helpers are extracted).

- _sigmoid / _conflict_scores / _insert_presentation_review (heads.py): the
  score→prob transform (6 inlined sites), the presentation conflict signal
  (2 sites), and the ring-loud PresentationReview insert (2 sites, single-
  sourced so the mode column can't drift on the shared composite PK).
- _applied_or_rejected (training_data.py): the per-tag "applied ∪ rejected"
  skip-set, byte-identical at 3 sweep sites (heads.py x2, tasks/ml.py ccip).
- ccip sweep divergence fixes: import ccip._FIGURE_KINDS + training_data._l2norm
  instead of local copies that silently drift when the canonical changes.
- MLSettings.load / .load_sync classmethods (mirror ImportSettings); route all
  8 scalar_one singleton reads through them (the session.get None-path stays).
- GET serializers for MLSettings + ImportSettings are now table-driven off the
  same _EDITABLE tuples PATCH writes, so a new field can't be silently absent
  from GET (the split that historically dropped fields).
- AUTO_APPLY_THRESHOLD_MIN/MAX constant single-sources the [0.5,0.999] operating
  range across the service clamp + the 5 API validators.
- test_ml_dry_helpers.py pins _applied_or_rejected + _sigmoid.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NsmJSQxnNxGgtM5Yz4GAqi
2026-07-13 21:41:24 -04:00
bvandeusenandClaude Opus 4.8 d80a5255ed feat(extension): in-app update prompt — popup banner + toolbar badge (#1489)
CI / lint (push) Successful in 3s
extension / lint (push) Successful in 10s
CI / frontend-build (push) Successful in 24s
CI / backend-lint-and-test (push) Successful in 39s
CI / integration (push) Successful in 3m51s
extension / lint (pull_request) Successful in 10s
The extension is installed per-instance from the operator's FC host, so Firefox's
static update_url can't apply (each instance has a different host) and updates
were fully manual. Add a self-hosted-friendly update surface that reuses the
existing public GET /api/extension/manifest ({version, latest_url, sha256}):

- lib/api.js: getExtensionManifest().
- background.js: checkForUpdateInfo() compares the instance's latest published
  version against runtime.getManifest().version (dotted-numeric compare so
  1.0.10 > 1.0.9); CHECK_UPDATE message handler; refreshUpdateBadge() sets a
  toolbar badge via browser.action; a daily browser.alarms check plus on
  startup/installed. New 'alarms' permission (non-prompting).
- popup: an 'Update available — vX' banner with an Update button that opens the
  signed XPI (web root, /api stripped like OPEN_ARTIST_PAGE) → Firefox's native
  install prompt. Never blocks the popup on a failed check.

No backend changes (endpoint already exists). Bump 1.0.8→1.0.9 so this ships;
from here on updates surface themselves instead of needing a manual reinstall.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-13 18:28:06 -04:00
97 changed files with 3415 additions and 749 deletions
+188 -25
View File
@@ -2,10 +2,18 @@ name: Build images
on:
push:
# `:dev` builds dropped 2026-05-26 — operator tests from `:latest` after
# merge-to-main, not from the dev branch image. Saves one full docker
# build per dev push.
branches: [main]
# `:dev` builds were dropped 2026-05-26 to save a docker build per dev
# push, on the reasoning that "operator tests from `:latest` after
# merge-to-main". Restored 2026-08-27: that is testing by shipping, and
# family rules 146/147 now name it directly — `main` IS production, and a
# channel that can only be refreshed by shipping is not a channel. The
# pressure to merge in order to try something does not come from
# carelessness; it comes from `:dev` being unable to carry the build.
#
# All three images build on dev, deliberately: a `:dev` web image paired
# with a stale `:dev` ml or agent is a worse trap than no dev channel at
# all, since the mismatch only shows up as a runtime failure.
branches: [main, dev]
# Tag-push triggers an immutable per-version image build (e.g.
# `:v26.05.26.5`) — gives a real rollback story alongside the floating
# `:main` / `:latest`. Layer reuse keeps the registry-storage cost
@@ -25,28 +33,135 @@ jobs:
# Forgejo release exists yet, otherwise downloads the cached signed XPI.
# Result is uploaded as an Actions artifact for build-web to consume.
#
# Why this lives in build.yml (not a separate workflow): the merge-commit's
# docker image tagged `:latest` MUST carry the XPI. A separate sign workflow
# racing build.yml leaves `:latest` without the XPI for ~5min (until the
# commit-back triggers another build). Inline ordering eliminates the race.
# Why this lives in build.yml (not a separate workflow): the image a push
# publishes MUST carry the XPI. A separate sign workflow racing build.yml
# leaves that image without one for ~5min (until the commit-back triggers
# another build). Inline ordering eliminates the race.
# Cache strategy: Forgejo Release Assets — picked 2026-05-25 over Generic
# Packages (cleaner API surface) and commit-back-to-side-branch (no extra
# branch to manage). AMO blocks re-signing the same version (returns 409),
# so signing is intentionally one-shot per version bump.
# so signing is intentionally one-shot per version.
#
# BOTH branches sign (milestone 271 step 6, 2026-08-27). Not two signatures:
# the version is the commit TIME of the newest packaged-extension change, so
# dev and main derive the SAME number for the same extension source. A dev
# push that changes the extension signs it; the merge to main then finds the
# ext-<version> release already there, hits the cache, and bundles the
# byte-identical XPI into `:latest` with no second AMO call. One signature
# per extension CHANGE, shared by both channels — that is what makes two
# channels affordable, and it is why step 4 (derived version) had to land
# first. Ungating this while the version was still the hand-set 1.0.11 would
# have hit the existing ext-1.0.11 cache and bundled MAIN's stale XPI into
# `:dev` — a dev channel confidently serving old code.
#
# Tags stay excluded: the tag path deliberately skips signing and polls for
# the release instead (see build-web's race note, 2026-05-27).
sign-extension:
if: github.ref == 'refs/heads/main'
if: github.ref == 'refs/heads/main' || github.ref == 'refs/heads/dev'
runs-on: python-ci
container:
image: git.fabledsword.com/bvandeusen/ci-python:3.14
steps:
- uses: actions/checkout@v4
with:
# Full history is load-bearing, not a convenience: the version this
# job signs is derived from the commit TIME of the newest packaged
# extension change. A depth-1 clone sees one commit and derives a
# wrong, too-low value rather than failing (ci-requirements.md).
fetch-depth: 0
- name: Resolve extension version
# The version is DERIVED, not read from the repo (milestone 271 step 4,
# cut over 2026-08-27). `packaging.sh version` returns MAJOR.MINOR from
# manifest.json plus a patch component that is the commit TIME of the
# newest change to a PACKAGED extension file, in minutes since
# 2020-01-01 — family rule 149, never a commit count, which orders by
# branch rather than by recency.
#
# The committed "version" in manifest.json / package.json no longer
# decides anything: the stamp step below overwrites it in the working
# tree before web-ext ever reads it. It is deliberately NOT committed
# back — the commit carrying the bump would itself be a change to the
# extension and would move the version again. The repo holds the source;
# the build derives the label.
- name: Derive extension version
id: extver
run: |
VERSION=$(grep -E '"version"' extension/package.json | head -1 | sed -E 's/.*"version"[[:space:]]*:[[:space:]]*"([^"]+)".*/\1/')
set -eu
VERSION=$(sh extension/scripts/packaging.sh version)
echo "version=$VERSION" >> "$GITHUB_OUTPUT"
echo "Resolved extension version: $VERSION"
echo "Derived extension version: $VERSION"
# Firefox refuses a downgrade and AMO never releases a burned version,
# so a version that moves BACKWARDS is unrecoverable: it strands every
# install that already took the higher one. Two ways it could happen —
# a checkout without full history (derives too low), or a rewritten
# history that drops the newest packaged commit.
#
# The test is `derived < highest already signed`, strictly. Equality is
# the ORDINARY case, not a fault: an unchanged extension derives the same
# version it did last build, which is exactly what lets the ext-<version>
# cache hit and holds AMO to one call per extension CHANGE. Only moving
# backwards is a failure, so this runs on every path — cache hit
# included — rather than only before a sign.
- name: Guard — the derived version must never go backwards
env:
TOKEN: ${{ secrets.RELEASE_TOKEN }}
DERIVED: ${{ steps.extver.outputs.version }}
run: |
python3 - <<'PY'
import json, os, sys, urllib.request
API = ("https://git.fabledsword.com/api/v1/repos/"
"bvandeusen/FabledCurator/releases")
headers = {"Authorization": "token " + os.environ["TOKEN"]}
# Paginated rather than first-page-only: ext-* releases share this
# list with the v* release tags, so one page would start missing them
# as those accumulate. The bound FAILS rather than silently scanning
# part of the list and calling the highest it saw the highest there is.
tags = []
for page in range(1, 21):
req = urllib.request.Request(
f"{API}?limit=50&page={page}", headers=headers)
with urllib.request.urlopen(req, timeout=30) as resp:
batch = json.load(resp)
if not batch:
break
tags += [r.get("tag_name", "") for r in batch]
else:
sys.exit("guard: >1000 releases — pagination bound reached")
def parse(v):
try:
return tuple(int(part) for part in v.split("."))
except ValueError:
return None
derived_s = os.environ["DERIVED"]
derived = parse(derived_s)
if derived is None:
sys.exit(f"guard: derived version {derived_s!r} is not numeric")
signed = sorted(
(v, t) for t in tags if t.startswith("ext-")
for v in [parse(t[4:])] if v
)
if not signed:
print("guard: no ext-* release yet — nothing to go backwards from")
raise SystemExit(0)
hi, hi_tag = signed[-1]
print(f"guard: derived={derived_s} highest already signed={hi_tag}")
if derived < hi:
sys.exit(
f"REFUSING TO SIGN: derived {derived_s} is OLDER than the "
f"already-signed {hi_tag}. Firefox would reject it as a "
f"downgrade, and AMO will not release the burned version. "
f"First thing to check: did this job check out with "
f"fetch-depth: 0?"
)
print("guard: ok")
PY
- name: Check Forgejo release-asset cache
id: cache
@@ -84,6 +199,29 @@ jobs:
# removal — sign-extension's job is just to ensure the cache
# exists on Forgejo; the build-web side reads it independently).
# web-ext signs whatever manifest.json says, so the derived value has to
# reach the tree before signing. package.json is written too: the two are
# required to agree (ci.yml's guard), and a local `npm run build` reads
# it. Working tree only — never committed, per the note on the derive
# step.
- name: Stamp the derived version into manifest.json + package.json
env:
DERIVED: ${{ steps.extver.outputs.version }}
run: |
python3 - <<'PY'
import json, os
version = os.environ["DERIVED"]
for path in ("extension/manifest.json", "extension/package.json"):
with open(path) as fh:
doc = json.load(fh)
doc["version"] = version
with open(path, "w") as fh:
json.dump(doc, fh, indent=2)
fh.write("\n")
print(f"{path}: version -> {version}")
PY
- name: Sign via AMO (cache miss)
if: steps.cache.outputs.cached != 'true'
run: |
@@ -110,6 +248,11 @@ jobs:
# created it so an upload failure below can roll back (don't
# leave an empty release tombstone that the next run's
# cache-check mistakes for a partial-failure state).
#
# target_commitish is the signing commit, not a branch name: since
# step 6 either branch can create this release, and hard-coding
# `main` would tag a dev-signed XPI against a main commit that may
# not even contain the extension source it was built from.
STATUS=$(curl -s -o release.json -w "%{http_code}" \
-H "Authorization: token $TOKEN" \
"https://git.fabledsword.com/api/v1/repos/bvandeusen/FabledCurator/releases/tags/ext-$VERSION" || echo 000)
@@ -117,7 +260,7 @@ jobs:
CREATED_BY_US=false
else
curl -s -X POST -H "Authorization: token $TOKEN" -H "Content-Type: application/json" \
-d "{\"tag_name\":\"ext-$VERSION\",\"name\":\"Extension $VERSION (signed XPI cache)\",\"body\":\"Internal cache for the signed XPI consumed by build.yml's build-web job. Not a user-facing FC release.\",\"target_commitish\":\"main\"}" \
-d "{\"tag_name\":\"ext-$VERSION\",\"name\":\"Extension $VERSION (signed XPI cache)\",\"body\":\"Internal cache for the signed XPI consumed by build.yml's build-web job. Not a user-facing FC release.\",\"target_commitish\":\"$GITHUB_SHA\"}" \
-o release.json \
"https://git.fabledsword.com/api/v1/repos/bvandeusen/FabledCurator/releases"
CREATED_BY_US=true
@@ -160,19 +303,31 @@ jobs:
build-web:
needs: [sign-extension]
# sign-extension is main-only; on dev it's skipped, build-web still runs.
# sign-extension runs on main and dev, and is skipped on a tag push (which
# polls for the release instead). Either is fine to build on; a FAILED sign
# is not — this condition lets success and skipped through, so a failure
# skips build-web rather than shipping an image without the XPI.
if: always() && (needs.sign-extension.result == 'success' || needs.sign-extension.result == 'skipped')
runs-on: python-ci
container:
image: git.fabledsword.com/bvandeusen/ci-python:3.14
steps:
- uses: actions/checkout@v4
with:
# Full history: this job RE-DERIVES the extension version rather than
# being handed it, and a depth-1 clone derives a wrong, too-low value
# rather than failing — which would 404 the download of a release
# that exists perfectly well under its real name.
fetch-depth: 0
- name: Download signed XPI from Forgejo release asset (main + tags)
# Fires on main-push AND on tag-push. Tag-push builds re-package the
# same source code as the preceding main-push build but with an
# immutable version tag — they need the XPI too, otherwise the
# versioned image ships without the signed extension.
- name: Download signed XPI from Forgejo release asset
# Fires on every trigger shape. dev and main each bundle the XPI their
# own sign-extension just published — that is the whole point of the
# channel work (milestone 271 step 6): the dev image carries the
# extension being developed, rather than requiring a merge to try it.
# Tag-push builds re-package the same source as the preceding main-push
# build but with an immutable version tag — they need the XPI too,
# otherwise the versioned image ships without the signed extension.
#
# Tag-push vs main-push race (operator-flagged 2026-05-27 after
# v26.05.27.0 hit it): a release cut fires BOTH workflows almost
@@ -184,12 +339,18 @@ jobs:
# for up to 10min total) before giving up. Main-push's signing
# eventually wins and tag-push picks the release up on a later
# iteration.
if: github.ref == 'refs/heads/main' || startsWith(github.ref, 'refs/tags/')
if: github.ref == 'refs/heads/main' || github.ref == 'refs/heads/dev' || startsWith(github.ref, 'refs/tags/')
env:
TOKEN: ${{ secrets.RELEASE_TOKEN }}
run: |
set -eux
VERSION=$(grep -E '"version"' extension/package.json | head -1 | sed -E 's/.*"version"[[:space:]]*:[[:space:]]*"([^"]+)".*/\1/')
# Re-derived, not read from the repo: sign-extension published
# ext-<derived>, and the committed version has been inert since
# milestone 271 step 4. Both jobs run `packaging.sh version` over the
# same commit, so they agree by construction — and if they ever
# didn't, this download 404s and the build fails loudly instead of
# shipping a stale XPI.
VERSION=$(sh extension/scripts/packaging.sh version)
# Poll for the ext-<version> release. main-push's sign-extension
# step (AMO round-trip, 1-5min) needs to finish + upload before
# tag-push can fetch. 30s * 20 = up to 10min wait, then hard-fail.
@@ -255,9 +416,11 @@ jobs:
# rollback unit"). Rollback to any commit
# becomes `docker pull …:c-<sha>` without a
# release ceremony.
# anything else → safety net; shouldn't fire given the `on:`
# config above. Tag :dev to surface the
# unexpected run in the registry.
# refs/heads/dev → push to dev: publish :dev, the rolling test
# channel (family rule 146). Rolling means it may
# carry newer contents than the :c-<sha> of the
# same commit; it never writes :c-<sha> itself,
# because that is the rollback unit (rule 145).
# POSIX-safe substring (the runner shell is dash/BusyBox sh, not
# bash — `${var:0:7}` errors with "Bad substitution"; cut works
# everywhere). Operator-flagged 2026-06-01 after first :c-<sha>
+112
View File
@@ -2,6 +2,7 @@ name: CI
# CI lanes per FabledRulebook/forgejo.md "CI philosophy":
# - lint: ruff only, no dep install — fast-fail for the common lint bounce.
# - extension-version: guards the extension publish path (see the job).
# - backend-lint-and-test: `pytest -m "not integration"`, no service containers.
# - frontend-build: vitest unit + vite build.
# - integration: pgvector + redis service containers; alembic + `pytest -m integration`.
@@ -41,6 +42,117 @@ jobs:
# catching syntax errors before the image build.
run: python -m compileall -q agent/fc_agent
# Guards the extension publish path, which has no self-correcting behavior.
#
# build.yml's sign-extension job keys its AMO-signing cache purely on the
# version string in extension/package.json: if an `ext-<version>` Forgejo
# release already carries an XPI, signing is SKIPPED and that old signed XPI
# is what build-web bakes into `:latest`. Nothing in that path inspects
# whether extension/ actually changed — so a forgotten version bump ships a
# stale extension on a fully green build, silently. (AMO can't help: it 409s
# on re-signing a version, which is exactly why the cache exists.)
#
# This job makes that case loud, on the dev push, instead of invisible at
# merge-to-main. It is pure git + text work — no deps, no services.
extension-version:
runs-on: python-ci
container:
image: git.fabledsword.com/bvandeusen/ci-python:3.14
steps:
- uses: actions/checkout@v4
with:
# Full history: the check diffs against the push's `before` SHA (or
# the PR base), which a depth-1 clone wouldn't contain.
fetch-depth: 0
- name: Extension version guard
env:
BEFORE: ${{ github.event.before }}
PR_BASE: ${{ github.event.pull_request.base.sha }}
run: |
set -eu
# busybox sh on the act_runner — no bashisms (family rule).
ver() { grep -E '"version"' "$1" | head -1 | sed -E 's/.*"version"[[:space:]]*:[[:space:]]*"([^"]+)".*/\1/'; }
PKG=$(ver extension/package.json)
MAN=$(ver extension/manifest.json)
test -n "$PKG" || { echo "ERROR: no version found in extension/package.json"; exit 1; }
test -n "$MAN" || { echo "ERROR: no version found in extension/manifest.json"; exit 1; }
# (1) Unconditional: the two version strings must agree. `web-ext sign`
# reads manifest.json (package.json sits in --ignore-files and isn't
# even inside the XPI), so AMO signs MAN and Firefox installs MAN.
# build.yml keys its cache, release tag, XPI filename — and therefore
# the version /api/extension/manifest reports to the update prompt —
# on PKG. Divergence either hard-fails at AMO or ships a mislabelled
# XPI whose update prompt lies about what's installed.
if [ "$MAN" != "$PKG" ]; then
echo "ERROR: extension version mismatch."
echo " extension/manifest.json = $MAN <- what AMO signs / Firefox installs"
echo " extension/package.json = $PKG <- what CI caches, names, and reports"
echo "Set both to the same value."
exit 1
fi
# (2) If the SHIPPED extension changed, the version must have moved.
#
# Compare against MAIN, not against the previous push. The publish
# decision is made at merge-to-main against whatever ext-<version>
# already exists, so "differs from main" is the question that matters.
# Diffing against the previous dev push instead would demand a fresh
# bump on every iteration — push, tweak the extension again, and CI
# would insist on a second bump that buys nothing, inflating the
# version for no reason. On a main push there is no "main to compare
# to" yet, so fall back to that push's own before-SHA.
if [ "${GITHUB_REF##*/}" = "main" ]; then
BASE="${BEFORE:-}"
else
BASE=$(git rev-parse --verify -q origin/main 2>/dev/null || git rev-parse --verify -q main 2>/dev/null || echo "")
# PR base is the fallback when main isn't in the clone at all.
[ -n "$BASE" ] || BASE="${PR_BASE:-}"
fi
case "$BASE" in
''|0000000000000000000000000000000000000000)
echo "No usable base ref (no main in clone / first push) — skipping the bump check."
echo "OK: extension version $PKG"
exit 0
;;
esac
if ! git cat-file -e "$BASE^{commit}" 2>/dev/null; then
echo "Base commit $BASE not in this clone — skipping the bump check."
echo "OK: extension version $PKG"
exit 0
fi
# The exclusion list is NOT written out here — it comes from
# extension/scripts/packaging.sh, the one definition of what ships,
# shared with web-ext's --ignore-files and the derived-version patch
# count. Three hand-kept copies of that fact is how #2397 happened.
#
# `set -f` is required around the substitution: without it the shell
# globs `test/**` against the working tree and silently narrows it.
set -f
CHANGED=$(git diff --name-only "$BASE" HEAD -- extension/ $(sh extension/scripts/packaging.sh pathspec))
set +f
if [ -z "$CHANGED" ]; then
echo "No packaged extension files changed since $BASE — nothing to guard."
echo "OK: extension version $PKG"
exit 0
fi
echo "Packaged extension files changed since $BASE:"
echo "$CHANGED" | sed 's/^/ /'
PKG_OLD=$(git show "$BASE:extension/package.json" 2>/dev/null | grep -E '"version"' | head -1 | sed -E 's/.*"version"[[:space:]]*:[[:space:]]*"([^"]+)".*/\1/')
if [ -z "$PKG_OLD" ]; then
echo "Could not read the base version — skipping the bump check."
echo "OK: extension version $PKG"
exit 0
fi
if [ "$PKG_OLD" = "$PKG" ]; then
echo "ERROR: packaged extension files changed but the version is still $PKG."
echo "build.yml would find the existing ext-$PKG release, skip AMO signing,"
echo "and bake the OLD signed XPI into :latest — a green build shipping stale code."
echo "Bump the version in BOTH extension/package.json and extension/manifest.json."
exit 1
fi
echo "OK: extension version $PKG_OLD -> $PKG"
backend-lint-and-test:
runs-on: python-ci
container:
+56 -3
View File
@@ -1,5 +1,5 @@
name: extension
# Lint-only workflow. The sign-and-publish dance moved into build.yml's
# Lint + unit tests. The sign-and-publish dance moved into build.yml's
# `sign-extension` job (2026-05-25) — `:latest` now always bundles the XPI
# because sign-extension runs as a build-web dependency in the SAME workflow,
# eliminating the prior race between build.yml and a separate extension.yml.
@@ -10,10 +10,15 @@ on:
paths:
- 'extension/**'
- '.forgejo/workflows/extension.yml'
# test/version.spec.js asserts ci.yml's extension-version guard never
# ignores a file web-ext actually packages, so a ci.yml-only edit can
# break this suite and must trigger it.
- '.forgejo/workflows/ci.yml'
pull_request:
branches: [main]
paths:
- 'extension/**'
- '.forgejo/workflows/ci.yml'
workflow_dispatch:
jobs:
@@ -23,7 +28,55 @@ jobs:
image: node:24-bookworm-slim
steps:
- uses: actions/checkout@v4
- name: Install web-ext
run: cd extension && npm install --no-save --no-audit --no-fund
# Not --no-save: vitest and web-ext are both real devDependencies now,
# and the suite needs vitest resolvable from node_modules.
- name: Install dev dependencies
run: cd extension && npm install --no-audit --no-fund
- name: Lint
run: cd extension && npm run lint
# Pure-logic specs over lib/url.js and lib/platforms.js plus manifest /
# package version-consistency checks. No browser, no network.
- name: Unit tests
run: cd extension && npm run test:unit
# Everything else about packaging is asserted against our own declaration
# of what ships. This is the only check that asks web-ext what it ACTUALLY
# put in the archive. Until now that was an unverified assumption about
# glob semantics — and a fragile one: `test/**` reaches web-ext intact
# only because callers `set -f` first, so losing that quoting would
# silently start shipping dev files with no other signal.
- name: Verify XPI contents
run: |
set -eu
command -v unzip >/dev/null 2>&1 || { apt-get update -qq && apt-get install -y -qq unzip; }
cd extension
npm run build
ZIP=$(ls web-ext-artifacts/*.zip | head -1)
echo "=== packaged entries in $ZIP ==="
unzip -Z1 "$ZIP" | sort
echo "=== end ==="
ENTRIES=$(unzip -Z1 "$ZIP")
fail=0
# Must NOT ship: repo infrastructure with no business in a user's browser.
for pat in 'test/' 'scripts/' 'vitest.config.js' 'package.json' 'package-lock.json' 'README.md' 'node_modules/' 'web-ext-artifacts/'; do
if echo "$ENTRIES" | grep -q "^$pat"; then
echo "ERROR: '$pat' was packaged into the XPI but must not be"
fail=1
fi
done
# Must ship: if an exclusion pattern ever over-matches, the extension
# breaks at runtime rather than at build time, so assert presence too.
for req in 'manifest.json' 'lib/url.js' 'lib/api.js' 'lib/platforms.js' 'lib/cookies.js'; do
if ! echo "$ENTRIES" | grep -q "^$req$"; then
echo "ERROR: '$req' is missing from the XPI"
fail=1
fi
done
for dir in 'background/' 'popup/' 'options/' 'content/' 'icons/'; do
if ! echo "$ENTRIES" | grep -q "^$dir"; then
echo "ERROR: nothing from '$dir' was packaged"
fail=1
fi
done
[ "$fail" -eq 0 ] || exit 1
echo "XPI contents verified."
+35 -6
View File
@@ -6,7 +6,21 @@ Combines what was [ImageRepo](https://git.fabledsword.com/bvandeusen/ImageRepo)
## Status
Pre-v1. Not yet functional.
In production. `main` is continuously deployed — every merge to `main` builds
and publishes `:latest` images, so whatever is on `main` is what is running.
Day-to-day work happens on `dev`, which publishes `:dev` images.
## What's in here
Five deployable pieces, built by `.forgejo/workflows/build.yml`:
| Piece | Built from | Image | Role |
| --- | --- | --- | --- |
| **Web / workers** | `Dockerfile` | `fabledcurator` | Quart API + the built Vue SPA in one image. `entrypoint.sh` picks the role: `web`, `worker`, `scheduler`. The `maintenance-long` service is a second `worker` pinned to the long-running maintenance queue. |
| **ML worker** | `Dockerfile.ml` | `fabledcurator-ml` | Same app, plus `requirements-ml.txt` — tagging and embedding models that run in-container. |
| **GPU agent** | `agent/Dockerfile` | `fabledcurator-agent` | Optional desktop-GPU worker (`agent/`). Leases jobs over **HTTP only** — never touches the database or Redis. Run it for a burst, stop it to reclaim the card. See `agent/README.md`. |
| **Firefox extension** | `extension/` | signed XPI | MV3 extension: pushes platform session cookies into FC and adds a creator as a Source in one click. AMO-signed on `main` only, then bundled into the web image and served from Settings → Maintenance. See `extension/README.md`. |
| **Data** | — | `pgvector/pgvector:pg16`, `redis:7-alpine` | Postgres with pgvector for embeddings; Redis as the Celery broker. |
## Quick start
@@ -29,22 +43,37 @@ docker compose -f docker-compose.yml up -d
# (skips the override so containers pull registry images)
```
The GPU agent is deployed separately, on the machine with the card —
`agent/docker-compose.yml`, not this stack.
## Deployment posture
FabledCurator is designed to run inside a self-hosted homelab environment over plain HTTP. If you want TLS, terminate it at your reverse proxy. The app does not generate certificates, redirect to HTTPS, or set HSTS.
## CI / Forgejo setup
The repo's workflows expect:
Three workflows: `ci.yml` (lint, extension-version guard, backend unit tests,
frontend build, integration), `extension.yml` (extension lint, vitest, XPI
content verification), and `build.yml` (sign + publish).
- **Runner label `python-ci`** — a Forgejo runner with Python 3.14, ruff, and Node 22 pre-installed. Both `ci.yml` and `build.yml` use this label. The runner image (`runner-base:python-ci`) is built from `CI-Runner/CI-python/` in the operator's workspace; `make push` from that directory builds and pushes a new image when toolchain pins change.
- **Repo secret `RELEASE_TOKEN`** — a Forgejo PAT with the following scopes:
**The toolchain each job runs in is its `container.image`, not its `runs-on`
label.** `runs-on: python-ci` only schedules the job onto a runner; every job
then names the image it actually wants. `ci-requirements.md` is the current,
authoritative list of images and per-job installs — read that rather than a
copy here, so the two can't drift.
The repo expects one secret:
- **`RELEASE_TOKEN`** — a Forgejo PAT with:
- `write:package` + `read:package` — for `docker push` to `git.fabledsword.com`
- `write:release` — for future release-cutting workflows
- `write:issue` — for future issue-management automation
- `write:release` — for the `ext-<version>` releases that cache the signed XPI
- `write:issue` — for issue-management automation
Generate at https://git.fabledsword.com/user/settings/applications. The injected `GITHUB_TOKEN` cannot be used because it lacks `write:package`.
AMO signing additionally needs `MOZILLA_AMO_JWT_KEY` / `MOZILLA_AMO_JWT_SECRET`; it runs on
`main` only and is cached per version, since AMO rejects a re-signed version.
## License
Personal project; use at your own discretion.
+4 -1
View File
@@ -21,7 +21,7 @@ log = logging.getLogger("fc_agent.app")
# Bump on every agent change. The page embeds this and /status reports it; the UI
# warns to reload when they differ — so a stale browser-cached page can't be
# mistaken for "the new image didn't deploy". (Belt-and-braces with no-store.)
VERSION = "2026-07-02.6 · sleep mode: an empty queue sheds to one downloader and backs the lease poll off to 15 min"
VERSION = "2026-07-17.1 · idle model-unload: after ~5 min idle the GPU models release their VRAM and reload on the next job (env IDLE_UNLOAD_SECONDS, 0=off) · sleep mode sheds to one downloader"
logbuf.install()
cfg = Config.from_env()
@@ -334,9 +334,12 @@ _PAGE = """<!doctype html><html><head><meta charset=utf-8>
waited.textContent=s.transient||0
// Instantaneous pool state → demoted to the sub-line, where its jumpiness reads
// as live churn rather than a "broken" headline metric.
// '=== false' (not falsy) so a stale page that doesn't send models_loaded shows
// nothing; when the idle monitor unloads, the VRAM meter drops alongside this.
pipe.textContent='downloaders '+(s.downloaders!=null?s.downloaders:'')+' · consumers '+(s.consumers!=null?s.consumers:'')+' · on GPU '+(s.active||0)
+' · net '+(s.net_mb_s!=null?s.net_mb_s.toFixed(1):'')+' MB/s'
+(s.bandwidth_limit_mb_s>0?(' / cap '+s.bandwidth_limit_mb_s):'')
+(s.models_loaded===false?' · GPU models unloaded (idle — reload on next job)':'')
if(document.activeElement!==bw && s.bandwidth_limit_mb_s!=null) bw.value=s.bandwidth_limit_mb_s
// Buffer occupancy bar (also driven here so it tracks the /status cadence).
if(s.buffer!=null && s.buffer_max){ const p=Math.round(100*s.buffer/s.buffer_max)
+10
View File
@@ -51,6 +51,12 @@ class Config:
bandwidth_limit_mb_s: float # aggregate download cap in MEGABYTES/s across
# all downloaders + video streams (0 = unlimited);
# tunable live from the agent UI
idle_unload_seconds: float # after this long with the GPU idle (nothing in
# flight, queue empty or Stopped), unload the
# SigLIP embedder + YOLO proposers to free their
# VRAM; they reload lazily on the next job. A
# 24/7 agent otherwise squats on ~5GB doing
# nothing. 0 disables (keep models warm forever).
@classmethod
def from_env(cls) -> Config:
@@ -87,4 +93,8 @@ class Config:
# link to ~1-1.5 MB/s per stream, browser included). Raise it (or 0)
# from the agent UI on wired/faster networks.
bandwidth_limit_mb_s=float(os.environ.get("BANDWIDTH_LIMIT_MB_S", "8")),
# 5 min: long enough that a lull between job bursts doesn't thrash the
# (few-second) reload, short enough that an agent left running with an
# empty queue hands its VRAM back promptly.
idle_unload_seconds=float(os.environ.get("IDLE_UNLOAD_SECONDS", "300")),
)
+15
View File
@@ -170,6 +170,13 @@ class YoloProposer:
))
return out
def unload(self) -> None:
"""Drop the loaded YOLO so its VRAM can be reclaimed; detect() reloads it
lazily on the next job. Leaves _ok untouched — a healthy proposer comes
back, but one that self-disabled on a fault stays off."""
with self._lock:
self._model = None
class Proposers:
"""The agent's proposer set, built from config. Each detector is optional —
@@ -216,3 +223,11 @@ class Proposers:
def panels(self, image):
return self._top(self._panel, image, self.cfg.max_panels)
def unload(self) -> None:
"""Release every loaded proposer's YOLO (idle VRAM reclaim). The worker
also drops its reference to this Proposers and rebuilds a fresh one via
_proposers_for on the next job, so this is belt-and-braces."""
for p in (self._person, self._anatomy, self._panel):
if p is not None:
p.unload()
+15
View File
@@ -75,3 +75,18 @@ class CropEmbedder:
pooled = out.pooler_output if hasattr(out, "pooler_output") else out
arr = pooled.float().cpu().numpy().astype(np.float32)
return [row.reshape(-1).tolist() for row in arr]
def unload(self) -> bool:
"""Drop the loaded model so its VRAM can be reclaimed — the idle monitor
calls this after a spell with no work so an idle agent doesn't squat on
the card; the next embed() reloads it lazily (a few seconds). Held under
BOTH the load and inference locks so it can never race a concurrent load
or an in-flight forward pass. Returns True if a model was actually
released (the caller then runs one empty_cache() to hand the freed blocks
back to the driver)."""
with self._load_lock, self._infer_lock:
if self._model is None:
return False
self._model = None
self._processor = None
return True
+75
View File
@@ -57,6 +57,15 @@ MAX_BACKOFF_SECONDS = 60.0
# up on their own.
IDLE_POLL_MAX_SECONDS = 900.0
# Idle VRAM reclaim (operator 2026-07-17): the SigLIP embedder + YOLO proposers
# load lazily and then stay warm for fast job bursts — but a 24/7 agent with an
# empty queue would otherwise squat on that VRAM (~5GB on the operator's card)
# indefinitely while doing nothing. So a monitor unloads them after
# cfg.idle_unload_seconds with the GPU genuinely idle (nothing in flight, buffer
# drained); they reload lazily on the next job. This is just how often the
# monitor wakes to check — it bounds how soon past the threshold the unload fires.
IDLE_UNLOAD_CHECK_INTERVAL = 30.0
# A job whose fetch dies transiently this many times IN ONE SESSION stops being
# handed back and is failed instead. Transient handbacks (release) burn no
# attempts on the server, so a poisoned transfer — an original that stalls the
@@ -268,6 +277,11 @@ class Worker:
self._proposers_sig = None # detector-config signature the current
# proposers were built for (#134)
self._proposers_lock = threading.Lock()
# Monotonic time of the last GPU activity (a consumer finishing a job).
# The idle monitor unloads the warm models once this goes stale by
# cfg.idle_unload_seconds — see _idle_unload_loop.
self._last_gpu_activity = time.monotonic()
threading.Thread(target=self._idle_unload_loop, daemon=True).start()
# --- held-lease bookkeeping --------------------------------------------
def _hold(self, job_ids) -> None:
@@ -608,6 +622,9 @@ class Worker:
"net_mb_s": round(self._net_mb_s, 1), # observed aggregate rate
"bw_capped": self._bw_capped, # autoscaler holding at the cap (UI hint)
"idle": self._idle, # queue empty → poll backed off (UI hint)
# Whether the GPU models are currently resident (False after an idle
# unload freed their VRAM) — a plain bool read, UI hint only.
"models_loaded": self._embedder is not None or self._proposers is not None,
}
def _bump(self, *, processed=0, downloaded=0, errors=0, active=0, transient=0):
@@ -788,6 +805,9 @@ class Worker:
self._bump(processed=1)
finally:
self._bump(active=-1)
# Mark the GPU busy-until-now so the idle monitor starts its
# unload countdown from when work actually stopped, not before.
self._last_gpu_activity = time.monotonic()
def _ensure_embedder(self, model_name: str):
if self._embedder is not None:
@@ -845,6 +865,61 @@ class Worker:
self._proposers_sig = sig
return self._proposers
def _unload_models(self) -> bool:
"""Release the GPU-resident models (SigLIP embedder + YOLO proposers) so an
idle agent hands their VRAM back instead of squatting on the card. They
reload lazily on the next job (_ensure_embedder / _proposers_for) — a
few seconds' cost paid only when work actually resumes. Dropping the
shared instances under their build locks means a concurrent job either
sees the old instance (before) or rebuilds a fresh one (after); the idle
monitor only calls this with nothing in flight, so no inference is using
them. Returns True if anything was released."""
released = False
with self._embedder_lock:
if self._embedder is not None:
self._embedder.unload()
self._embedder = None
released = True
with self._proposers_lock:
if self._proposers is not None:
self._proposers.unload()
self._proposers = None
self._proposers_sig = None
released = True
if released:
try:
import torch
if torch.cuda.is_available():
# torch's caching allocator holds freed blocks; hand them back
# to the driver so nvidia-smi actually reflects the drop.
torch.cuda.empty_cache()
except Exception: # noqa: BLE001 — torch absent / CPU-only → nothing to free
pass
return released
def _idle_unload_loop(self) -> None:
"""Unload the warm GPU models after a stretch of inactivity so a 24/7
agent with an empty queue doesn't hold ~5GB of VRAM doing nothing. Fires
only when nothing is in flight (active == 0 AND the buffer is drained) and
no job has completed for cfg.idle_unload_seconds — a window long enough
that a brief lull between bursts doesn't thrash reload/unload. Covers BOTH
sleep mode (queue empty, pipeline still running) and a full Stop; the
models reload lazily on the next job. idle_unload_seconds <= 0 disables it."""
idle_after = self.cfg.idle_unload_seconds
if idle_after <= 0:
return
while True:
time.sleep(IDLE_UNLOAD_CHECK_INTERVAL)
if self._embedder is None and self._proposers is None:
continue # nothing loaded → nothing to free
if self._active != 0 or not self._buffer.empty():
continue # work in flight → keep them warm
if time.monotonic() - self._last_gpu_activity < idle_after:
continue # not idle long enough yet
if self._unload_models():
log.info("idle %.0fs — unloaded GPU models, freed VRAM "
"(reload on next job)", idle_after)
def _consume(self, job: dict, frames: list, stop_evt: threading.Event) -> bool:
"""Detect + embed the decoded frames and submit the result. Returns True
when the job was completed (→ count it processed), False otherwise: a
+16
View File
@@ -459,6 +459,22 @@ async def trigger_prune_missing_files():
return _queued(async_result)
@admin_bp.route("/maintenance/reclaim-attachments", methods=["POST"])
async def trigger_reclaim_attachments():
"""Reclaim orphaned attachments (#3068). Body {"dry_run": bool}: dry_run
(the DEFAULT here) projects the orphan rows and unreferenced store blobs
without touching either; dry_run=false deletes the rows then unlinks every
blob no surviving row references. Maintenance queue; operator-triggered
only — never an unattended sweep, since the apply unlinks files. Returns the
Celery task id — poll /maintenance/task-result/<id> for the summary."""
from ..tasks.admin import reclaim_orphaned_attachments_task
body = await request.get_json(silent=True) or {}
dry_run = bool(body.get("dry_run", True)) # default to the SAFE preview
async_result = reclaim_orphaned_attachments_task.delay(dry_run=dry_run)
return _queued(async_result)
@admin_bp.route("/maintenance/dedup-videos", methods=["POST"])
async def trigger_dedup_videos():
"""Tier-1 video dedup (#871). Body {"dry_run": bool}: dry_run=true previews
+10 -1
View File
@@ -6,6 +6,7 @@ from __future__ import annotations
import asyncio
import hashlib
import hmac
import re
from pathlib import Path
@@ -41,7 +42,15 @@ async def _ext_key_required(session) -> bool:
stored = (await session.execute(
select(AppSetting.value).where(AppSetting.key == "extension_api_key")
)).scalar_one_or_none()
return stored is not None and supplied == stored
if stored is None:
return False
# compare_digest, not `==`: the stored key is a shared secret, and a
# short-circuiting compare leaks its prefix through timing. Costs nothing
# here — it is not that this route is exposed (#3072). Compared as BYTES:
# compare_digest's str form rejects non-ASCII with TypeError, and this
# header is attacker-supplied, so a str compare would turn a junk key into
# a 500 instead of a 403.
return hmac.compare_digest(supplied.encode("utf-8"), stored.encode("utf-8"))
def _extract_version(xpi_name: str) -> str:
+1 -3
View File
@@ -256,9 +256,7 @@ async def lease():
if not await _agent_authed(session):
return jsonify({"error": "unauthorized"}), 401
jobs = await GpuJobService(session).lease(agent_id, batch_size=batch)
ml = (
await session.execute(select(MLSettings).where(MLSettings.id == 1))
).scalar_one()
ml = await MLSettings.load(session)
# image rows for url/mime in one shot
ids = [j.image_record_id for j in jobs]
imgs = {
+17 -43
View File
@@ -4,6 +4,7 @@ from quart import Blueprint, jsonify, request
from ..extensions import get_session
from ..models import MLSettings
from ..services.ml.heads import AUTO_APPLY_THRESHOLD_MAX, AUTO_APPLY_THRESHOLD_MIN
ml_admin_bp = Blueprint("ml_admin", __name__, url_prefix="/api/ml")
@@ -83,48 +84,21 @@ async def embedder_models():
@ml_admin_bp.route("/settings", methods=["GET"])
async def get_settings():
from sqlalchemy import select
async with get_session() as session:
s = (
await session.execute(select(MLSettings).where(MLSettings.id == 1))
).scalar_one()
return jsonify(
{
"cpu_embed_enabled": s.cpu_embed_enabled,
"video_frame_interval_seconds": s.video_frame_interval_seconds,
"video_max_frames": s.video_max_frames,
"embedder_model_version": s.embedder_model_version,
"head_min_positives": s.head_min_positives,
"head_auto_apply_precision": s.head_auto_apply_precision,
"head_auto_apply_enabled": s.head_auto_apply_enabled,
"head_auto_apply_min_positives": s.head_auto_apply_min_positives,
"ccip_match_threshold": s.ccip_match_threshold,
"ccip_auto_apply_enabled": s.ccip_auto_apply_enabled,
"ccip_auto_apply_threshold": s.ccip_auto_apply_threshold,
"presentation_auto_apply_enabled": s.presentation_auto_apply_enabled,
"presentation_auto_apply_threshold": s.presentation_auto_apply_threshold,
"presentation_conflict_threshold": s.presentation_conflict_threshold,
"process_auto_apply_enabled": s.process_auto_apply_enabled,
"process_auto_apply_threshold": s.process_auto_apply_threshold,
"process_conflict_threshold": s.process_conflict_threshold,
"embedder_model_name": s.embedder_model_name,
**{f: getattr(s, f) for f in _DETECTOR_FIELDS},
}
)
s = await MLSettings.load(session)
# Table-driven off _EDITABLE (which PATCH also writes) so a new settings field
# can never be silently absent from GET — the split that historically dropped
# fields. _EDITABLE already includes *_DETECTOR_FIELDS.
return jsonify({f: getattr(s, f) for f in _EDITABLE})
@ml_admin_bp.route("/settings", methods=["PATCH"])
async def patch_settings():
from sqlalchemy import select
body = await request.get_json()
if not isinstance(body, dict):
return jsonify({"error": "body must be an object"}), 400
async with get_session() as session:
s = (
await session.execute(select(MLSettings).where(MLSettings.id == 1))
).scalar_one()
s = await MLSettings.load(session)
# Merge the patch over current values, then validate the result as a
# whole — the store-floor invariant couples three fields, so they
@@ -154,24 +128,24 @@ def _validate(p: dict) -> str | None:
# Head training (#114).
if int(p["head_min_positives"]) < 1:
return "head_min_positives must be >= 1"
if not (0.5 <= float(p["head_auto_apply_precision"]) <= 0.999):
return "head_auto_apply_precision must be between 0.5 and 0.999"
if not (AUTO_APPLY_THRESHOLD_MIN <= float(p["head_auto_apply_precision"]) <= AUTO_APPLY_THRESHOLD_MAX):
return f"head_auto_apply_precision must be between {AUTO_APPLY_THRESHOLD_MIN} and {AUTO_APPLY_THRESHOLD_MAX}"
if int(p["head_auto_apply_min_positives"]) < 1:
return "head_auto_apply_min_positives must be >= 1"
if not (0.5 <= float(p["ccip_match_threshold"]) <= 0.999):
return "ccip_match_threshold must be between 0.5 and 0.999"
if not (0.5 <= float(p["ccip_auto_apply_threshold"]) <= 0.999):
return "ccip_auto_apply_threshold must be between 0.5 and 0.999"
if not (AUTO_APPLY_THRESHOLD_MIN <= float(p["ccip_match_threshold"]) <= AUTO_APPLY_THRESHOLD_MAX):
return f"ccip_match_threshold must be between {AUTO_APPLY_THRESHOLD_MIN} and {AUTO_APPLY_THRESHOLD_MAX}"
if not (AUTO_APPLY_THRESHOLD_MIN <= float(p["ccip_auto_apply_threshold"]) <= AUTO_APPLY_THRESHOLD_MAX):
return f"ccip_auto_apply_threshold must be between {AUTO_APPLY_THRESHOLD_MIN} and {AUTO_APPLY_THRESHOLD_MAX}"
# Presentation chrome auto-hide (#141). Auto-apply runs high (hiding is
# consequential); the conflict cut is a plain probability [0,1].
if not (0.5 <= float(p["presentation_auto_apply_threshold"]) <= 0.999):
return "presentation_auto_apply_threshold must be between 0.5 and 0.999"
if not (AUTO_APPLY_THRESHOLD_MIN <= float(p["presentation_auto_apply_threshold"]) <= AUTO_APPLY_THRESHOLD_MAX):
return f"presentation_auto_apply_threshold must be between {AUTO_APPLY_THRESHOLD_MIN} and {AUTO_APPLY_THRESHOLD_MAX}"
if not (0.0 <= float(p["presentation_conflict_threshold"]) <= 1.0):
return "presentation_conflict_threshold must be between 0 and 1"
# Process auto-apply (#1464). wip/editor stay VISIBLE so a false apply is
# low-harm (excludes-from-training + a review flag), but keep the same bar.
if not (0.5 <= float(p["process_auto_apply_threshold"]) <= 0.999):
return "process_auto_apply_threshold must be between 0.5 and 0.999"
if not (AUTO_APPLY_THRESHOLD_MIN <= float(p["process_auto_apply_threshold"]) <= AUTO_APPLY_THRESHOLD_MAX):
return f"process_auto_apply_threshold must be between {AUTO_APPLY_THRESHOLD_MIN} and {AUTO_APPLY_THRESHOLD_MAX}"
if not (0.0 <= float(p["process_conflict_threshold"]) <= 1.0):
return "process_conflict_threshold must be between 0 and 1"
# Embedder model swap (#1190): both must be non-empty. Changing them means a
+3 -28
View File
@@ -66,34 +66,9 @@ _EXTDL_TOGGLE_FIELDS = (
async def get_import_settings():
async with get_session() as session:
row = await ImportSettings.load(session)
return jsonify({
"min_width": row.min_width,
"min_height": row.min_height,
"skip_transparent": row.skip_transparent,
"transparency_threshold": row.transparency_threshold,
"skip_single_color": row.skip_single_color,
"single_color_threshold": row.single_color_threshold,
"single_color_tolerance": row.single_color_tolerance,
"phash_threshold": row.phash_threshold,
"download_rate_limit_seconds": row.download_rate_limit_seconds,
"download_validate_files": row.download_validate_files,
"download_schedule_default_seconds": row.download_schedule_default_seconds,
"download_event_retention_days": row.download_event_retention_days,
"download_failure_warning_threshold": row.download_failure_warning_threshold,
"series_suggest_enabled": row.series_suggest_enabled,
"series_suggest_threshold": row.series_suggest_threshold,
"extdl_mega_enabled": row.extdl_mega_enabled,
"extdl_gdrive_enabled": row.extdl_gdrive_enabled,
"extdl_mediafire_enabled": row.extdl_mediafire_enabled,
"extdl_dropbox_enabled": row.extdl_dropbox_enabled,
"extdl_pixeldrain_enabled": row.extdl_pixeldrain_enabled,
"translation_enabled": row.translation_enabled,
"interpreter_base_url": row.interpreter_base_url,
"translation_target_lang": row.translation_target_lang,
"translation_min_confidence": row.translation_min_confidence,
"wip_title_tagging_enabled": row.wip_title_tagging_enabled,
"wip_soft_title_tagging_enabled": row.wip_soft_title_tagging_enabled,
})
# Table-driven off _EDITABLE_FIELDS (which PATCH also writes) so a new field
# can't be silently absent from GET.
return jsonify({f: getattr(row, f) for f in _EDITABLE_FIELDS})
@settings_bp.route("/settings/import", methods=["PATCH"])
+2 -1
View File
@@ -27,7 +27,7 @@ from .patreon_seen_media import PatreonSeenMedia
from .pixiv_failed_media import PixivFailedMedia
from .pixiv_seen_media import PixivSeenMedia
from .post import Post
from .post_attachment import PostAttachment
from .post_attachment import PostAttachment, attachment_download_url
from .presentation_review import PresentationReview
from .series_chapter import SeriesChapter
from .series_page import SeriesPage
@@ -58,6 +58,7 @@ __all__ = [
"SubscribeStarSeenMedia",
"Post",
"PostAttachment",
"attachment_download_url",
"PresentationReview",
"SeriesChapter",
"SeriesPage",
+12
View File
@@ -10,6 +10,7 @@ from sqlalchemy import (
Integer,
String,
func,
select,
)
from sqlalchemy.orm import Mapped, mapped_column
@@ -212,3 +213,14 @@ class MLSettings(Base):
updated_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now()
)
@classmethod
async def load(cls, session) -> MLSettings:
"""The singleton settings row (id=1), via an async session. Mirrors
ImportSettings.load — the shared singleton-loader pattern."""
return (await session.execute(select(cls).where(cls.id == 1))).scalar_one()
@classmethod
def load_sync(cls, session) -> MLSettings:
"""The singleton settings row (id=1), via a sync session."""
return session.execute(select(cls).where(cls.id == 1)).scalar_one()
+12
View File
@@ -65,3 +65,15 @@ class PostAttachment(Base):
captured_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now()
)
def attachment_download_url(attachment_id: int) -> str:
"""The path that streams this attachment's bytes.
Both serializers that expose an attachment to the frontend
(`provenance_service`, `post_feed_service`) built this literal themselves,
so changing the route in `api/attachments.py` meant two edits and only one
would be remembered (#3072). `test_attachment_download_url` pins it against
the app's registered rule, so the drift is caught rather than trusted to.
"""
return f"/api/attachments/{attachment_id}/download"
+256 -6
View File
@@ -48,6 +48,47 @@ log = logging.getLogger(__name__)
_VIDEO_DURATION_UNKNOWN = -1.0
# -- artist-cascade predicates (rule 93: ONE definition, preview + apply) ---
# project_artist_cascade (preview) and delete_artist_cascade (apply) both build
# their queries from these. The preview used to re-derive its own — which is how
# it came to count images and stay silent about posts and attachments while the
# apply destroyed both. Same failure shape as the 2026-06-08 fandom-tag
# deletion, where a re-implemented delete predicate diverged from the preview's.
# Returned as condition LISTS spread into `.where(*conds)`, matching
# _unused_tag_conditions / _bare_post_conditions below.
def _artist_images_conditions(artist_id: int) -> list:
"""Images the cascade deletes (rows AND their on-disk files)."""
return [ImageRecord.artist_id == artist_id]
def _artist_posts_conditions(artist_id: int) -> list:
"""Posts the cascade destroys. The apply never names these — post.artist_id
is ondelete=CASCADE, so Postgres takes them when the artist row goes — which
is exactly why the preview has to name them: an artist whose posts are
body-only (no images) otherwise previews as `images: 0` and reads as an
empty artist, while every captured body/description/external-link set is
destroyed."""
return [Post.artist_id == artist_id]
def _artist_attachments_conditions(artist_id: int) -> list:
"""Attachments the cascade deletes. Matched by artist_id OR by the owning
post's artist: artist_id is nullable (_capture_attachment leaves it NULL
when no artist resolved), so neither arm alone covers every row. The
sha-addressed blobs are NOT unlinked (one blob backs many rows) — these are
row counts, and the bytes are not part of this operation's footprint."""
return [
or_(
PostAttachment.artist_id == artist_id,
PostAttachment.post_id.in_(
select(Post.id).where(*_artist_posts_conditions(artist_id))
),
)
]
def project_artist_cascade(session: Session, *, slug: str) -> dict:
"""Read-only projection of what delete_artist_cascade would touch.
@@ -56,12 +97,17 @@ def project_artist_cascade(session: Session, *, slug: str) -> dict:
"artist": {"id": int, "name": str, "slug": str},
"projected": {
"images": int,
"posts": int, # hard-deleted by the post.artist_id CASCADE
"attachments": int, # rows deleted; the sha-addressed blobs stay
"sources": int,
"thumbs": int, # images with a thumbnail_path set
"import_tasks": int, # ImportTask rows referencing the artist's images
"bytes_on_disk": int, # SUM(image_record.size_bytes) — column is NOT NULL
},
}
Every count is built from the shared `_artist_*_conditions` predicates the
apply uses, so the two halves cannot drift (rule 93).
Raises LookupError if slug not found. No mutations.
"""
from ..models.import_task import ImportTask
@@ -73,36 +119,49 @@ def project_artist_cascade(session: Session, *, slug: str) -> dict:
if artist is None:
raise LookupError(f"artist slug not found: {slug!r}")
images_conds = _artist_images_conditions(artist.id)
images_count = session.execute(
select(func.count(ImageRecord.id))
.where(ImageRecord.artist_id == artist.id)
select(func.count(ImageRecord.id)).where(*images_conds)
).scalar_one()
posts_count = session.execute(
select(func.count(Post.id))
.where(*_artist_posts_conditions(artist.id))
).scalar_one()
attachments_count = session.execute(
select(func.count(PostAttachment.id))
.where(*_artist_attachments_conditions(artist.id))
).scalar_one()
# Sources have no shared predicate: the apply never queries them either, it
# gets them from the Artist.sources ORM cascade. Counted directly here.
sources_count = session.execute(
select(func.count(Source.id))
.where(Source.artist_id == artist.id)
).scalar_one()
thumbs_count = session.execute(
select(func.count(ImageRecord.id))
.where(ImageRecord.artist_id == artist.id)
.where(*images_conds)
.where(ImageRecord.thumbnail_path.is_not(None))
).scalar_one()
import_tasks_count = session.execute(
select(func.count(ImportTask.id))
.where(
ImportTask.result_image_id.in_(
select(ImageRecord.id).where(ImageRecord.artist_id == artist.id)
select(ImageRecord.id).where(*images_conds)
)
)
).scalar_one()
bytes_on_disk = session.execute(
select(func.coalesce(func.sum(ImageRecord.size_bytes), 0))
.where(ImageRecord.artist_id == artist.id)
.where(*images_conds)
).scalar_one()
return {
"artist": {"id": artist.id, "name": artist.name, "slug": artist.slug},
"projected": {
"images": images_count,
"posts": posts_count,
"attachments": attachments_count,
"sources": sources_count,
"thumbs": thumbs_count,
"import_tasks": import_tasks_count,
@@ -277,6 +336,10 @@ def delete_artist_cascade(
series_page / tag_suggestion_rejection from ImageRecord delete,
and source / post / download_event / etc. from Artist delete
(via Artist.sources cascade="all, delete-orphan").
The artist's post_attachment rows are cleared EXPLICITLY before the
artist row goes — see the comment at that step; leaving them to the
cascade aborts the whole delete on a unique violation.
"""
artist = session.get(Artist, artist_id)
if artist is None:
@@ -287,11 +350,22 @@ def delete_artist_cascade(
"files_deleted": 0,
"thumbs_deleted": 0,
"import_tasks_nulled": 0,
"posts_deleted": 0,
"attachments_deleted": 0,
"files_failed": 0,
},
}
artist_info = {"id": artist.id, "name": artist.name, "slug": artist.slug}
# Counted BEFORE the delete: Postgres takes these via the post.artist_id
# CASCADE when the artist row goes, so afterwards there is nothing left to
# count. Reported so the summary can be checked against the preview's
# `posts` — the parity rule 93 asks for is only testable if both halves
# actually state the number.
posts_deleted = session.execute(
select(func.count(Post.id)).where(*_artist_posts_conditions(artist.id))
).scalar_one()
images_deleted = 0
files_deleted = 0
thumbs_deleted = 0
@@ -300,7 +374,7 @@ def delete_artist_cascade(
while True:
rows = session.execute(
select(ImageRecord)
.where(ImageRecord.artist_id == artist.id)
.where(*_artist_images_conditions(artist.id))
.limit(500)
).scalars().all()
if not rows:
@@ -323,6 +397,28 @@ def delete_artist_cascade(
# source_path_prefix matching that's out of scope here.
import_tasks_nulled = 0
# Clear the artist's attachments BEFORE the artist row, or the delete below
# aborts. Deleting an artist CASCADEs to Post (post.artist_id is
# ondelete=CASCADE), which SET NULLs post_attachment.post_id — and
# `uq_post_attachment_null_post_sha` is a partial UNIQUE on sha256 ALONE
# WHERE post_id IS NULL, so any two of this artist's attachments sharing a
# sha collapse onto one another and raise. That is an ORDINARY shape, not a
# corrupt one: _capture_attachment deliberately writes one row per post over
# a single sha-addressed blob (a creator who attaches the same pdf to two
# posts has two rows), and a pre-existing filesystem-import row with the same
# sha and a NULL post_id collides on its own. Migration 0043 reasoned only
# about upgrade-time safety and never about this later SET NULL.
# _repoint_post_links guards the identical collision class in the reconcile
# path; this is its artist-cascade counterpart.
#
# Which rows count as the artist's — and why the blobs are left on disk —
# is _artist_attachments_conditions, shared with the preview.
attachments_deleted = session.execute(
delete(PostAttachment)
.where(*_artist_attachments_conditions(artist.id))
).rowcount or 0
session.commit()
session.delete(artist)
session.commit()
@@ -333,6 +429,8 @@ def delete_artist_cascade(
"files_deleted": files_deleted,
"thumbs_deleted": thumbs_deleted,
"import_tasks_nulled": import_tasks_nulled,
"posts_deleted": posts_deleted,
"attachments_deleted": attachments_deleted,
"files_failed": files_failed,
},
}
@@ -1494,3 +1592,155 @@ def purge_gated_previews(
"ledger_cleared": ledger_cleared,
"posts_deleted": posts_deleted,
}
# -- orphaned attachment reclamation ---------------------------------------
# PostAttachment's two FKs are both ON DELETE SET NULL, so a deleted post or
# artist leaves the row behind rather than taking it. Nothing ever pruned those
# rows, and nothing has ever unlinked a file under the attachment store — so
# both rows and bytes accumulated permanently and were invisible to every
# existing diagnostic.
#
# Why this is a DISK->DB reconciliation rather than a row sweep: the store is
# sha-addressed and idempotent (attachment_store.store), so ONE blob backs MANY
# rows. Deleting a row therefore does not free its blob, and — since the artist
# cascade now deletes its attachment rows outright — a freed blob has no DB
# pointer left to find it by. Walking the store and asking "does any row still
# reference this sha?" catches orphans from every cause, including ones no
# future delete path will think to report.
# A blob is written by attachment_store.store BEFORE its row is inserted and
# committed, so a just-stored file legitimately has no referencing row for a
# moment. Same guard, same reasoning as ORPHAN_TEMP_MIN_AGE_HOURS in
# tasks/maintenance.py: never judge a file younger than this.
_ATTACHMENT_ORPHAN_MIN_AGE_HOURS = 6
# Wall-clock budget for the store walk (rule 89). A library with a large
# attachment store shouldn't be able to run this past its soft time limit; on
# exhaustion it reports partial=True and the operator re-runs to finish.
_ATTACHMENT_RECLAIM_BUDGET_SECONDS = 900
# The store names files `<sha256><ext>`. Parse the sha as the first 64 chars
# rather than via Path.stem: store() takes the extension straight from the
# source filename, and a URL-encoded basename yields a multi-dot "suffix"
# (see [[reference_url_encoded_basename_suffix]]) that would make stem eat part
# of the sha. Validating the 64 chars as hex also skips anything else in the
# tree that isn't a stored blob.
_SHA256_HEX_LEN = 64
def _orphan_attachment_conditions() -> list:
"""PostAttachment rows belonging to nothing: both FKs nulled by a deleted
post AND a deleted artist. A row with post_id NULL but an artist_id is the
deliberate filesystem-import case (importer._capture_attachment writes it
that way) and is NOT an orphan — it is still attributed."""
return [
PostAttachment.post_id.is_(None),
PostAttachment.artist_id.is_(None),
]
def _is_sha_named(name: str) -> bool:
"""True when `name` starts with a 64-char lowercase-hex sha256."""
if len(name) < _SHA256_HEX_LEN:
return False
head = name[:_SHA256_HEX_LEN]
return all(c in "0123456789abcdef" for c in head)
def reclaim_orphaned_attachments(
session: Session, *, images_root: Path, dry_run: bool = False,
) -> dict:
"""Prune unattributed PostAttachment rows, then unlink store blobs that no
surviving row references.
Returns (same discovery keys either way, so the UI renders one shape):
{"rows": int, # orphan rows found / deleted
"files": int, # unreferenced blobs found / unlinked
"bytes": int, # their total size
"scanned": int, # blobs examined
"skipped_recent": int, # blobs under the min-age guard
"files_failed": int, # unlink raised (apply only)
"partial": bool} # walk hit the time budget
dry_run computes exactly what the apply would do and mutates nothing — the
surviving-sha set is derived by NEGATING the same orphan predicate the
delete uses, so the preview cannot disagree with the apply (rule 93).
"""
started = time.monotonic()
orphan_conds = _orphan_attachment_conditions()
if dry_run:
rows = session.execute(
select(func.count(PostAttachment.id)).where(*orphan_conds)
).scalar_one()
else:
rows = session.execute(
delete(PostAttachment).where(*orphan_conds)
).rowcount or 0
session.commit()
# Shas that still have a home. In the apply path the orphan rows are already
# gone, so `NOT orphan` is redundant but harmless; in the dry-run path it is
# what makes the projection honest about blobs the delete would free. One
# predicate, one query, both modes.
surviving_shas = set(session.execute(
select(PostAttachment.sha256).where(~and_(*orphan_conds)).distinct()
).scalars())
root = Path(images_root) / "attachments"
cutoff = (
datetime.now(UTC).timestamp()
- _ATTACHMENT_ORPHAN_MIN_AGE_HOURS * 3600
)
files = 0
freed_bytes = 0
scanned = 0
skipped_recent = 0
files_failed = 0
partial = False
if root.is_dir():
for path in root.rglob("*"):
if time.monotonic() - started >= _ATTACHMENT_RECLAIM_BUDGET_SECONDS:
partial = True
break
# .partial staging files belong to cleanup_orphaned_temp_files —
# leave them alone rather than racing an in-flight store().
if path.suffix in (".part", ".partial") or not path.is_file():
continue
if not _is_sha_named(path.name):
continue
scanned += 1
sha = path.name[:_SHA256_HEX_LEN]
if sha in surviving_shas:
continue
try:
st = path.stat()
if st.st_mtime >= cutoff:
skipped_recent += 1
continue
size = st.st_size
if not dry_run:
path.unlink()
files += 1
freed_bytes += size
except OSError as exc:
files_failed += 1
log.warning("reclaim_orphaned_attachments: %s: %s", path, exc)
if not dry_run and (rows or files):
log.info(
"attachment reclaim: %d orphan row(s) deleted, %d blob(s) unlinked "
"(%d bytes), %d failed, partial=%s",
rows, files, freed_bytes, files_failed, partial,
)
return {
"rows": rows,
"files": files,
"bytes": freed_bytes,
"scanned": scanned,
"skipped_recent": skipped_recent,
"files_failed": files_failed,
"partial": partial,
}
+1 -1
View File
@@ -181,7 +181,7 @@ def _augment_cookies(platform: str, netscape: str) -> str:
"""Delegate to the platform's `augment_cookies` hook if one is
registered (subscribestar, hentaifoundry, etc. — see
`services/platforms/<name>.py`). No-op when the platform doesn't
register a hook (Patreon, DeviantArt). Centralizing the
register a hook (Patreon, Discord). Centralizing the
quirks-per-platform in the platforms package means adding a new
platform's cookie quirks doesn't require touching this file."""
info = PLATFORMS.get(platform)
+2 -3
View File
@@ -31,9 +31,8 @@ from .pixiv_ingester import PixivIngester
from .subscribestar_ingester import SubscribeStarIngester
# Platforms whose download + verify go through the native ingester rather than
# gallery-dl. gallery-dl still serves the rest (hentaifoundry, discord,
# deviantart — the latter slated for retirement, not migration) until they
# migrate too.
# gallery-dl. gallery-dl still serves the rest (hentaifoundry, discord) until
# they migrate too.
NATIVE_INGESTER_PLATFORMS = frozenset({"patreon", "subscribestar", "pixiv"})
# Mirrors patreon_resolver._CAMPAIGNS_URL — surfaced in resolution-failure
@@ -55,12 +55,6 @@ _PLATFORM_PATTERNS: list[tuple[str, re.Pattern[str]]] = [
r"^https?://(?:www\.)?hentai-foundry\.com/user/(?P<slug>[^/?#]+)",
re.IGNORECASE,
)),
("deviantart", re.compile(
r"^https?://(?:www\.)?deviantart\.com/"
r"(?!home$|watch\b|tag\b|browse\b)"
r"(?P<slug>[^/?#]+)/?$",
re.IGNORECASE,
)),
("pixiv", re.compile(
r"^https?://(?:www\.)?pixiv\.net/(?:en/)?users/(?P<slug>\d+)",
re.IGNORECASE,
+3 -11
View File
@@ -299,8 +299,9 @@ class GalleryDLService:
# (services/patreon_ingester.py), not gallery-dl.
PLATFORM_DEFAULTS = {
# subscribestar removed — native-ingester platform now (#71); pixiv
# removed likewise (#129). The remaining entries are the gallery-dl
# platforms not yet migrated.
# removed likewise (#129); deviantart removed at #3069 as a dropped
# platform, not a migrated one. The remaining entries are the
# gallery-dl platforms not yet migrated.
"hentaifoundry": {
"content_types": ["all"],
"directory": [],
@@ -316,15 +317,6 @@ class GalleryDLService:
"reactions": False,
"threads": True,
},
"deviantart": {
"content_types": ["all"],
"directory": [],
"filename": "{index:>03}_{title[:50]}.{extension}",
"flat": True,
"original": True,
"mature": True,
"metadata": True,
},
}
def __init__(
+51
View File
@@ -0,0 +1,51 @@
"""Bulk, idempotent writes to the ``image_tag`` association table.
Three writers attach tags to images in bulk: the WIP-title backfill
(`wip_title.apply_wip_image_tags`), the concept-head auto-apply sweep and the
system-tag auto-apply sweep (both in `ml/heads.py`). The two sweeps used to
issue ONE INSERT PER ROW from inside their per-image loop — fine in steady
state, but a first pass over a back-catalogue is tens of thousands of
individual round-trips (#3072). All three share this one chunked multi-row
insert now.
Sync only: every caller runs on a sync ``Session`` (the Celery task path). No
async service writes image_tag in bulk, so there is no async sibling to keep in
step — unlike `db_helpers.get_or_create`, which does have one.
"""
from __future__ import annotations
from sqlalchemy.dialects.postgresql import insert as pg_insert
from sqlalchemy.orm import Session
from ..models.tag import image_tag
# 5000 rows x 3 bound params = 15000, comfortably inside Postgres' 65535-param
# ceiling for a single statement. Raising this past ~21000 rows would exceed it.
INSERT_CHUNK = 5000
def insert_image_tags(
session: Session, rows: list[dict], *, chunk: int = INSERT_CHUNK
) -> None:
"""Attach ``rows`` to their images, skipping any tag already on one.
Each row is ``{"image_record_id": int, "tag_id": int, "source": str}``.
Does NOT commit — the caller owns the transaction.
ON CONFLICT DO NOTHING against the (image_record_id, tag_id) primary key,
so an existing tag keeps its ORIGINAL ``source``: re-running a sweep can
never re-stamp a tag the operator applied by hand as machine-applied.
Returns nothing on purpose. psycopg reports ``rowcount`` -1 for a multi-row
ON CONFLICT DO NOTHING insert (it runs via an executemany path), so a count
taken from the statement would be a lie rather than an approximation.
Callers that need an accurate count derive it themselves — see
`wip_title.apply_wip_image_tags`' pre-SELECT, and the sweeps' `skip` sets.
"""
for start in range(0, len(rows), chunk):
session.execute(
pg_insert(image_tag)
.values(rows[start:start + chunk])
.on_conflict_do_nothing(index_elements=["image_record_id", "tag_id"])
)
@@ -150,9 +150,7 @@ def refresh_character_prototypes(
"""Incrementally refresh the prototype store. `full=True` rebuilds every
character regardless of the gate/fingerprints (nightly reconcile). Returns
{skipped, rebuilt, removed}; commits."""
settings = session.execute(
select(MLSettings).where(MLSettings.id == 1)
).scalar_one()
settings = MLSettings.load_sync(session)
sig = _global_signature(session)
if not full and settings.ccip_ref_signature == sig:
return {"skipped": True, "rebuilt": 0, "removed": 0}
@@ -204,9 +202,7 @@ def retract_auto_applied_ccip(session: Session) -> int:
n_retracted."""
import numpy as np
settings = session.execute(
select(MLSettings).where(MLSettings.id == 1)
).scalar_one()
settings = MLSettings.load_sync(session)
if not settings.ccip_auto_apply_enabled:
return 0
thr = float(settings.ccip_auto_apply_threshold)
+85 -75
View File
@@ -23,6 +23,7 @@ from datetime import UTC, datetime
from typing import Any
from sqlalchemy import delete, exists, func, select
from sqlalchemy.dialects.postgresql import insert as pg_insert
from sqlalchemy.ext.asyncio import AsyncSession
from sqlalchemy.orm import Session
@@ -40,8 +41,10 @@ from ...models import (
TagSuggestionRejection,
)
from ...models.tag import CHROME_SYSTEM_TAGS, PROCESS_SYSTEM_TAGS, image_tag
from ..image_tag_apply import insert_image_tags
from .training_data import (
_AUTO_SOURCES,
_applied_or_rejected,
_auto_apply_point,
_hygiene_excluded_ids,
_ids_with_tag,
@@ -61,6 +64,14 @@ MIN_POSITIVES_FLOOR = 8 # hard floor; settings.head_min_positives can raise
_UNLABELED_POOL = 4000
_EXAMPLES_MIN = 8 # need at least this many embedded +/- to fit a head
# Auto-apply / match confidence operating range. Every graduated auto-apply or
# CCIP-match threshold the operator can set lives in this band, and the head
# precision target is clamped to it: below 0.5 "auto-apply" is meaningless, and
# 1.0 is unachievable so 0.999 is the ceiling. One source shared by the service
# clamp (_normalize_params) and the API validator (ml_admin._validate).
AUTO_APPLY_THRESHOLD_MIN = 0.5
AUTO_APPLY_THRESHOLD_MAX = 0.999
# Only these tag kinds get heads (the surfaced suggestion categories).
_HEAD_KINDS = (TagKind.general, TagKind.character)
# tag.kind -> the suggestion category the rail groups under.
@@ -78,6 +89,38 @@ _CATEGORY = {TagKind.general: "general", TagKind.character: "character"}
_SYSTEM_TAG_SUGGEST_FLOOR = 0.65
def _sigmoid(z, np):
"""Logistic sigmoid 1/(1+e^-z): the head score→probability transform. One home
for what was inlined at every scoring site (suggest, both sweeps, retract)."""
return 1.0 / (1.0 + np.exp(-z))
def _conflict_scores(Xn, Wc, bc, np):
"""The presentation conflict signal (#141): per row, the MAX content-head
probability and WHICH head produced it. Shared by the system-tag sweep's guard-2
and the soft-wip audit — both ask "does this ALSO look like real content?"."""
cprobs = _sigmoid(Xn @ Wc.T + bc, np)
return cprobs.max(axis=1), cprobs.argmax(axis=1)
def _insert_presentation_review(
session, *, image_record_id, tag_id, conflict_tag_id, conflict_score, mode,
):
"""Single-source the ring-loud PresentationReview row shape so the two writers
(system-tag sweep guard-2 + soft-wip audit) can't drift on columns or `mode` —
they share the (image_record_id, tag_id) composite PK, so a divergent `mode`
would be a silent first-writer-wins bug."""
session.execute(
pg_insert(PresentationReview)
.values(
image_record_id=image_record_id, tag_id=tag_id,
conflict_tag_id=conflict_tag_id, conflict_score=conflict_score,
mode=mode,
)
.on_conflict_do_nothing()
)
class HeadTrainingAlreadyRunning(Exception):
"""Raised by start_head_training_run when a run is already in flight."""
@@ -103,9 +146,7 @@ def start_head_training_run(session: Session, params: dict[str, Any]) -> int:
def _settings(session: Session) -> MLSettings:
return session.execute(
select(MLSettings).where(MLSettings.id == 1)
).scalar_one()
return MLSettings.load_sync(session)
def _normalize_params(session: Session, params: dict[str, Any] | None) -> dict[str, Any]:
@@ -124,7 +165,7 @@ def _normalize_params(session: Session, params: dict[str, Any] | None) -> dict[s
except (TypeError, ValueError):
cv_folds = DEFAULT_CV_FOLDS
try:
precision_target = min(max(float(params.get("precision_target", s.head_auto_apply_precision)), 0.5), 0.999)
precision_target = min(max(float(params.get("precision_target", s.head_auto_apply_precision)), AUTO_APPLY_THRESHOLD_MIN), AUTO_APPLY_THRESHOLD_MAX)
except (TypeError, ValueError):
precision_target = s.head_auto_apply_precision
return {
@@ -536,7 +577,7 @@ async def score_image(
norms[norms == 0] = 1.0
Xn = X / norms
Z = Xn @ heads["W"].T + heads["b"] # (B, H)
probs_bag = 1.0 / (1.0 + np.exp(-Z)) # (B, H)
probs_bag = _sigmoid(Z, np) # (B, H)
probs = probs_bag.max(axis=0) # (H,) best over the bag
# ARGMAX beside the max: WHICH bag row won each head → the region that grounds
# the tag (bag_meta[win]); None when the whole-image vector won (#1206).
@@ -614,9 +655,7 @@ async def ground_applied_tag(
async def _settings_async(session: AsyncSession) -> MLSettings:
return (
await session.execute(select(MLSettings).where(MLSettings.id == 1))
).scalar_one()
return await MLSettings.load(session)
# --- Earned auto-apply (sync, ml worker) ---------------------------------
@@ -687,7 +726,6 @@ def auto_apply_sweep(
embeddings in chunks; commits per chunk on a real run. Returns
{n_applied, concepts:[{tag_id,name,applied,scanned,threshold}]}."""
import numpy as np
from sqlalchemy.dialects.postgresql import insert as pg_insert
settings = _settings(session)
rows = _auto_apply_heads(
@@ -704,18 +742,7 @@ def auto_apply_sweep(
names = [r.name for r in rows]
# Skip images that already carry, or have rejected, each tag.
skip = {tid: set() for tid in tag_ids}
for tid in tag_ids:
for (iid,) in session.execute(
select(image_tag.c.image_record_id).where(image_tag.c.tag_id == tid)
):
skip[tid].add(iid)
for (iid,) in session.execute(
select(TagSuggestionRejection.image_record_id).where(
TagSuggestionRejection.tag_id == tid
)
):
skip[tid].add(iid)
skip = _applied_or_rejected(session, tag_ids)
applied = [0] * len(rows)
scanned = 0
@@ -729,8 +756,12 @@ def auto_apply_sweep(
if not cids:
continue
Xn = _l2norm(np.vstack([emb[i] for i in cids]).astype(np.float32), np)
probs = 1.0 / (1.0 + np.exp(-(Xn @ W.T + b))) # (N, H)
probs = _sigmoid(Xn @ W.T + b, np) # (N, H)
scanned += len(cids)
# Collected across every head, then written as ONE insert below. Was an
# insert per applied tag from inside this loop, which on a first sweep
# over a back-catalogue is tens of thousands of round-trips (#3072).
pending: list[dict] = []
for h in range(len(rows)):
tid = tag_ids[h]
for idx in np.where(probs[:, h] >= thr[h])[0]:
@@ -740,12 +771,12 @@ def auto_apply_sweep(
skip[tid].add(iid)
applied[h] += 1
if not dry_run:
session.execute(
pg_insert(image_tag)
.values(image_record_id=iid, tag_id=tid, source="head_auto")
.on_conflict_do_nothing()
)
pending.append({
"image_record_id": iid, "tag_id": tid,
"source": "head_auto",
})
if not dry_run:
insert_image_tags(session, pending)
session.commit()
run.last_progress_at = datetime.now(UTC)
session.commit()
@@ -840,7 +871,6 @@ def system_tag_auto_apply_sweep(
enabled flag is set. numpy-only (no sklearn). Returns {n_applied, n_flagged,
concepts}."""
import numpy as np
from sqlalchemy.dialects.postgresql import insert as pg_insert
cfg = _SWEEP_MODES[mode]
settings = _settings(session)
@@ -869,18 +899,7 @@ def system_tag_auto_apply_sweep(
valued = _valued_image_ids(session)
# Skip images that already carry, or have rejected, each presentation tag.
skip = {tid: set() for tid in pres_tag_ids}
for tid in pres_tag_ids:
for (iid,) in session.execute(
select(image_tag.c.image_record_id).where(image_tag.c.tag_id == tid)
):
skip[tid].add(iid)
for (iid,) in session.execute(
select(TagSuggestionRejection.image_record_id).where(
TagSuggestionRejection.tag_id == tid
)
):
skip[tid].add(iid)
skip = _applied_or_rejected(session, pres_tag_ids)
applied = [0] * len(pres)
n_flagged = 0
@@ -895,12 +914,15 @@ def system_tag_auto_apply_sweep(
if not cids:
continue
Xn = _l2norm(np.vstack([emb[i] for i in cids]).astype(np.float32), np)
probs = 1.0 / (1.0 + np.exp(-(Xn @ Wp.T + bp))) # (N, P)
probs = _sigmoid(Xn @ Wp.T + bp, np) # (N, P)
if Wc is not None:
cprobs = 1.0 / (1.0 + np.exp(-(Xn @ Wc.T + bc))) # (N, C)
max_c = cprobs.max(axis=1)
arg_c = cprobs.argmax(axis=1)
max_c, arg_c = _conflict_scores(Xn, Wc, bc, np) # (N,), (N,)
scanned += len(cids)
# Same batching as auto_apply_sweep (#3072): collect the chunk's rows
# and write them once, below. The PresentationReview rows stay per-row —
# they FK to image_record/tag, not to image_tag, so writing the tags
# after them is safe, and a flagged conflict is rare by construction.
pending: list[dict] = []
for p in range(len(pres)):
tid = pres_tag_ids[p]
for idx in np.where(probs[:, p] >= thr)[0]:
@@ -910,31 +932,25 @@ def system_tag_auto_apply_sweep(
skip[tid].add(iid)
applied[p] += 1
if not dry_run:
session.execute(
pg_insert(image_tag)
.values(
image_record_id=iid, tag_id=tid,
source=source,
)
.on_conflict_do_nothing()
)
pending.append({
"image_record_id": iid, "tag_id": tid,
"source": source,
})
# Guard 2: also looks like real content → still apply, but flag it
# for the review strip instead of silently marking (chrome hides,
# process stays visible — either way the operator gets a heads-up).
if Wc is not None and float(max_c[idx]) >= conflict_thr:
n_flagged += 1
if not dry_run:
session.execute(
pg_insert(PresentationReview)
.values(
image_record_id=iid, tag_id=tid,
conflict_tag_id=conf_tag_ids[int(arg_c[idx])],
conflict_score=float(max_c[idx]),
mode=mode,
)
.on_conflict_do_nothing()
_insert_presentation_review(
session,
image_record_id=iid, tag_id=tid,
conflict_tag_id=conf_tag_ids[int(arg_c[idx])],
conflict_score=float(max_c[idx]),
mode=mode,
)
if not dry_run:
insert_image_tags(session, pending)
session.commit()
concepts = [
@@ -956,7 +972,6 @@ def soft_wip_conflict_audit(session: Session, dry_run: bool = False) -> dict:
NOT remove the tag; the operator decides. No-op when there are no content heads.
numpy-only. Returns {n_scanned, n_flagged}."""
import numpy as np
from sqlalchemy.dialects.postgresql import insert as pg_insert
from ..wip_title import WIP_TITLE_SOFT_SOURCE, resolve_wip_tag_id
@@ -993,22 +1008,17 @@ def soft_wip_conflict_audit(session: Session, dry_run: bool = False) -> dict:
continue
scanned += len(cids)
Xn = _l2norm(np.vstack([emb[i] for i in cids]).astype(np.float32), np)
cprobs = 1.0 / (1.0 + np.exp(-(Xn @ Wc.T + bc)))
max_c = cprobs.max(axis=1)
arg_c = cprobs.argmax(axis=1)
max_c, arg_c = _conflict_scores(Xn, Wc, bc, np)
for k in range(len(cids)):
if float(max_c[k]) >= conflict_thr:
n_flagged += 1
if not dry_run:
session.execute(
pg_insert(PresentationReview)
.values(
image_record_id=cids[k], tag_id=wip_id,
conflict_tag_id=conf_tag_ids[int(arg_c[k])],
conflict_score=float(max_c[k]),
mode="process",
)
.on_conflict_do_nothing()
_insert_presentation_review(
session,
image_record_id=cids[k], tag_id=wip_id,
conflict_tag_id=conf_tag_ids[int(arg_c[k])],
conflict_score=float(max_c[k]),
mode="process",
)
if not dry_run:
session.commit()
@@ -1062,7 +1072,7 @@ def retract_auto_applied_heads(session: Session) -> int:
continue
Xn = _l2norm(np.vstack([emb[i] for i in cids]).astype(np.float32), np)
w = np.asarray(weights, dtype=np.float32)
probs = 1.0 / (1.0 + np.exp(-(Xn @ w + float(bias))))
probs = _sigmoid(Xn @ w + float(bias), np)
below = [cids[k] for k in np.where(probs < float(thr))[0]]
for iid in below:
session.execute(
+18
View File
@@ -94,6 +94,24 @@ def _rejected_ids(session: Session, tag_id: int) -> list[int]:
]
def _applied_or_rejected(session: Session, tag_ids) -> dict[int, set[int]]:
"""Per-tag skip set for the auto-apply sweeps: every image that ALREADY carries
the tag (ANY source — not just training positives) OR has rejected it. A sweep
never re-applies to these. Shared by auto_apply_sweep + system_tag_auto_apply_sweep
(heads.py) and scheduled_ccip_auto_apply (tasks/ml.py). Callers mutate the returned
sets in-place to also dedupe within a single run."""
skip: dict[int, set[int]] = {}
for tid in tag_ids:
ids = {
r[0] for r in session.execute(
select(image_tag.c.image_record_id).where(image_tag.c.tag_id == tid)
).all()
}
ids.update(_rejected_ids(session, tid))
skip[tid] = ids
return skip
def _sample_unlabeled(session: Session, exclude: set[int], limit: int) -> list[int]:
"""Random image ids (with an embedding) NOT carrying the tag. Concepts are
sparse, so an untagged image is almost always a true negative."""
+20 -36
View File
@@ -91,48 +91,46 @@ def _sync_lookup(vanity: str, cookies_path: str | None) -> str | None:
)
def _lookup_via_api(vanity: str, cookies_path: str | None) -> str | None:
def _campaigns_api_first(vanity: str, cookies_path: str | None) -> dict | None:
"""The first `data` object from Patreon's campaigns API filtered by vanity
(`?filter[vanity]=<vanity>&fields[campaign]=name`), or None on any failure
(network / non-200 / non-JSON / empty). The single request shape shared by
_lookup_via_api (plucks the campaign id) and resolve_display_name (plucks the
display name)."""
jar = _load_cookie_jar(cookies_path)
headers = {
"User-Agent": _USER_AGENT,
"Accept": "application/vnd.api+json",
}
params = {
"filter[vanity]": vanity,
"fields[campaign]": "name",
}
try:
resp = requests.get(
_CAMPAIGNS_URL,
params=params,
headers=headers,
params={"filter[vanity]": vanity, "fields[campaign]": "name"},
headers={"User-Agent": _USER_AGENT, "Accept": "application/vnd.api+json"},
cookies=jar,
timeout=_TIMEOUT_SECONDS,
)
except requests.RequestException as exc:
log.warning("Patreon campaigns API request failed for vanity=%s: %s", vanity, exc)
return None
if resp.status_code != 200:
log.warning(
"Patreon campaigns API returned HTTP %d for vanity=%s",
resp.status_code, vanity,
)
return None
try:
payload = resp.json()
except ValueError as exc:
log.warning("Patreon campaigns API returned non-JSON for vanity=%s: %s", vanity, exc)
return None
data = payload.get("data") if isinstance(payload, dict) else None
if not isinstance(data, list) or not data or not isinstance(data[0], dict):
return None
return data[0]
if not isinstance(payload, dict):
def _lookup_via_api(vanity: str, cookies_path: str | None) -> str | None:
first = _campaigns_api_first(vanity, cookies_path)
if first is None:
return None
data = payload.get("data")
if not isinstance(data, list) or not data:
return None
first = data[0] if isinstance(data[0], dict) else None
campaign_id = first.get("id") if first else None
campaign_id = first.get("id")
if not isinstance(campaign_id, str) or not campaign_id:
return None
log.info("Resolved Patreon vanity=%s → campaign_id=%s", vanity, campaign_id)
@@ -144,24 +142,10 @@ def resolve_display_name(vanity: str, cookies_path: str | None) -> str | None:
(`fields[campaign]=name`), used to name the Artist at add-time (#130). None
on any failure — the caller falls back to the vanity handle. Sync: call from
an executor."""
jar = _load_cookie_jar(cookies_path)
try:
resp = requests.get(
_CAMPAIGNS_URL,
params={"filter[vanity]": vanity, "fields[campaign]": "name"},
headers={"User-Agent": _USER_AGENT, "Accept": "application/vnd.api+json"},
cookies=jar,
timeout=_TIMEOUT_SECONDS,
)
if resp.status_code != 200:
return None
data = resp.json().get("data")
except (requests.RequestException, ValueError) as exc:
log.warning("Patreon name lookup failed for vanity=%s: %s", vanity, exc)
first = _campaigns_api_first(vanity, cookies_path)
if first is None:
return None
if not isinstance(data, list) or not data or not isinstance(data[0], dict):
return None
name = (data[0].get("attributes") or {}).get("name")
name = (first.get("attributes") or {}).get("name")
return name.strip() if isinstance(name, str) and name.strip() else None
+3 -4
View File
@@ -8,9 +8,10 @@ PLATFORMS below. Sidecar parsing, cookie materialization, and
Lifted from GallerySubscriber's
~/Nextcloud/Projects/GallerySubscriber/backend/app/api/platforms.py
and ~/.../extension/lib/platforms.js. Six platforms; auth_type and
and ~/.../extension/lib/platforms.js. Five platforms; auth_type and
URL patterns match GS exactly so the existing browser extension
hits FC unmodified.
hits FC unmodified. deviantart was dropped at #3069 (2026-08-27) —
FC downloaders are art-dedicated services only.
"""
from .base import (
@@ -18,7 +19,6 @@ from .base import (
DEFAULT_EXTERNAL_POST_ID_KEYS,
PlatformInfo,
)
from .deviantart import INFO as _DEVIANTART
from .discord import INFO as _DISCORD
from .hentaifoundry import INFO as _HENTAIFOUNDRY
from .patreon import INFO as _PATREON
@@ -33,7 +33,6 @@ PLATFORMS: dict[str, PlatformInfo] = {
_HENTAIFOUNDRY,
_DISCORD,
_PIXIV,
_DEVIANTART,
)
}
+1 -1
View File
@@ -63,7 +63,7 @@ class PlatformInfo:
# Synthesize a post permalink from sidecar data. Required when
# gallery-dl's `url` field is the file/CDN URL rather than the post
# permalink (subscribestar/pixiv/hf/discord). None = trust the bare
# `url` field (patreon, deviantart).
# `url` field (patreon).
derive_post_url: Callable[[dict], str | None] | None = None
# Post-process the materialized cookies.txt for gallery-dl. Used by
@@ -1,23 +0,0 @@
"""DeviantArt — no exercised quirks yet.
No operator-owned DeviantArt archive existed at the 2026-05-27 sidecar
audit, so we don't know yet whether DA's gallery-dl sidecars are
well-behaved or have their own quirks. When DA gets exercised for the
first time, add `derive_post_url` / `augment_cookies` here as needed.
"""
from .base import GD_DEFAULTS, PlatformInfo
INFO = PlatformInfo(
key="deviantart",
name="DeviantArt",
description="Download artwork from DeviantArt artists",
auth_type="cookies",
requires_auth=False,
url_pattern=r"^https?://(www\.)?deviantart\.com/",
url_examples=[
"https://www.deviantart.com/example-artist",
"https://www.deviantart.com/example-artist/gallery",
],
default_config={**GD_DEFAULTS, "content_types": ["gallery"]},
)
+2 -1
View File
@@ -24,6 +24,7 @@ from ..models import (
Post,
PostAttachment,
Source,
attachment_download_url,
)
from ..utils.html_sanitize import (
extract_img_srcs,
@@ -360,7 +361,7 @@ class PostFeedService:
"ext": att.ext,
"mime": att.mime,
"size_bytes": att.size_bytes,
"download_url": f"/api/attachments/{att.id}/download",
"download_url": attachment_download_url(att.id),
})
return out
+2 -1
View File
@@ -16,6 +16,7 @@ from ..models import (
Post,
PostAttachment,
Source,
attachment_download_url,
)
from ..utils.html_sanitize import sanitize_post_html
@@ -53,7 +54,7 @@ def _attachment_dict(a: PostAttachment) -> dict:
"original_filename": a.original_filename,
"size_bytes": a.size_bytes,
"ext": a.ext,
"download_url": f"/api/attachments/{a.id}/download",
"download_url": attachment_download_url(a.id),
}
+5 -9
View File
@@ -20,10 +20,10 @@ family gains one member.
import re
from sqlalchemy import select
from sqlalchemy.dialects.postgresql import insert as pg_insert
from sqlalchemy.orm import Session
from ..models.tag import WIP_SYSTEM_TAG, Tag, image_tag
from .image_tag_apply import insert_image_tags
# image_tag.source stamped on title-heuristic WIP tags — distinct from the other
# apply sources so provenance stays legible and a future undo can target only these.
@@ -113,13 +113,9 @@ def apply_wip_image_tags(
to_insert = [iid for iid in chunk if iid not in already]
if not to_insert:
continue
session.execute(
pg_insert(image_tag)
.values([
{"image_record_id": iid, "tag_id": tag_id, "source": source}
for iid in to_insert
])
.on_conflict_do_nothing(index_elements=["image_record_id", "tag_id"])
)
insert_image_tags(session, [
{"image_record_id": iid, "tag_id": tag_id, "source": source}
for iid in to_insert
])
inserted += len(to_insert)
return inserted
+28
View File
@@ -409,3 +409,31 @@ def rescan_series_suggestions_task(self, after_post_id: int = 0) -> dict:
)
rescan_series_suggestions_task.delay(summary["resume_after_id"])
return summary
@celery.task(
name="backend.app.tasks.admin.reclaim_orphaned_attachments_task",
bind=True,
autoretry_for=(OperationalError, DBAPIError),
retry_backoff=15, retry_backoff_max=180, max_retries=1,
# The service stops walking at its own 900s budget and reports partial, so
# these limits are the backstop for a wedged filesystem (NFS stall), not the
# expected exit. Comfortably above the budget so a normal run always returns
# its summary rather than being killed mid-walk.
soft_time_limit=1200, time_limit=1500, # 20 min / 25 min
)
def reclaim_orphaned_attachments_task(self, dry_run: bool = True) -> dict:
"""Reclaim unattributed PostAttachment rows and the store blobs nothing
references any more (#3068). dry_run (the default) returns the projection
without touching rows or files; apply deletes the orphan rows, then unlinks
every blob no surviving row references.
Defaults to the SAFE preview — unlike the other tasks here, whose apply is
reversible-ish or scoped; this one deletes files. Operator-triggered only,
never on a beat: an unattended sweep that unlinks blobs is not something to
run without someone reading the projection first."""
SessionLocal = _sync_session_factory()
with SessionLocal() as session:
return cleanup_service.reclaim_orphaned_attachments(
session, images_root=IMAGES_ROOT, dry_run=dry_run,
)
+51 -72
View File
@@ -173,6 +173,12 @@ TASK_STUCK_THRESHOLD_MINUTES: dict[str, int] = {
# task-name override beats the queue threshold whatever queue the row records
# (it recorded 'default' before the celery_signals fix → download). 65 = 60+5.
"backend.app.tasks.external.fetch_external_link": 65,
# Attachment reclaim walks the whole sha-addressed store; the service caps
# itself at a 900s budget and reports partial, but the task's hard limit is
# 25 min for a wedged filesystem (NFS stall). Same phantom-flag class as the
# external-fetch entry above — without an override a healthy in-flight walk
# is swept 'RecoverySweep' at the bare 5-min default. 30 = 25 + 5.
"backend.app.tasks.admin.reclaim_orphaned_attachments_task": 30,
}
@@ -776,89 +782,62 @@ def recover_stalled_library_audit_runs() -> int:
return recovered
def _recover_stalled_runs(model, *, stall_minutes: int, keep_runs: int, label: str) -> int:
"""Shared recovery + retention sweep for the head run-tracking tables
(HeadTrainingRun / HeadAutoApplyRun, which share the
status/last_progress_at/started_at/finished_at/error/id columns): flip 'running'
rows with no progress past `stall_minutes` to 'error', then prune to the last
`keep_runs` (rule 89). Returns the number recovered. NOTE the two other recover
tasks are deliberately NOT folded in — library-audit has no prune tail and
backup uses a single started_at cutoff."""
SessionLocal = _sync_session_factory()
now = datetime.now(UTC)
cutoff = now - timedelta(minutes=stall_minutes)
with SessionLocal() as session:
result = session.execute(
update(model)
.where(model.status == "running")
.where(func.coalesce(model.last_progress_at, model.started_at) < cutoff)
.values(
status="error", finished_at=now,
error=f"stranded by recovery sweep (no progress for {stall_minutes} min)",
)
)
keep = session.execute(
select(model.id).order_by(model.id.desc()).limit(keep_runs)
).scalars().all()
if keep:
session.execute(delete(model).where(model.id.not_in(keep)))
session.commit()
recovered = result.rowcount or 0
if recovered:
log.info("%s: recovered %d rows", label, recovered)
return recovered
@celery.task(name="backend.app.tasks.maintenance.recover_stalled_head_training_runs")
def recover_stalled_head_training_runs() -> int:
"""Flip HeadTrainingRun rows stuck in 'running' past the stall threshold to
'error', and prune old runs to the last HEAD_TRAINING_KEEP_RUNS (retention,
rule 89). Runs every 5 min on the maintenance lane; no-op when idle."""
SessionLocal = _sync_session_factory()
now = datetime.now(UTC)
cutoff = now - timedelta(minutes=HEAD_TRAINING_STALL_THRESHOLD_MINUTES)
with SessionLocal() as session:
result = session.execute(
update(HeadTrainingRun)
.where(HeadTrainingRun.status == "running")
.where(
func.coalesce(
HeadTrainingRun.last_progress_at, HeadTrainingRun.started_at
)
< cutoff
)
.values(
status="error", finished_at=now,
error=(
f"stranded by recovery sweep (no progress for "
f"{HEAD_TRAINING_STALL_THRESHOLD_MINUTES} min)"
),
)
)
keep = session.execute(
select(HeadTrainingRun.id).order_by(HeadTrainingRun.id.desc())
.limit(HEAD_TRAINING_KEEP_RUNS)
).scalars().all()
if keep:
session.execute(
delete(HeadTrainingRun).where(HeadTrainingRun.id.not_in(keep))
)
session.commit()
recovered = result.rowcount or 0
if recovered:
log.info(
"recover_stalled_head_training_runs: recovered %d rows", recovered
)
return recovered
return _recover_stalled_runs(
HeadTrainingRun,
stall_minutes=HEAD_TRAINING_STALL_THRESHOLD_MINUTES,
keep_runs=HEAD_TRAINING_KEEP_RUNS,
label="recover_stalled_head_training_runs",
)
@celery.task(name="backend.app.tasks.maintenance.recover_stalled_head_auto_apply_runs")
def recover_stalled_head_auto_apply_runs() -> int:
"""Flip stalled HeadAutoApplyRun 'running' rows to 'error' + prune to the
last HEAD_AUTO_APPLY_KEEP_RUNS (retention, rule 89). 5-min maintenance lane."""
SessionLocal = _sync_session_factory()
now = datetime.now(UTC)
cutoff = now - timedelta(minutes=HEAD_AUTO_APPLY_STALL_THRESHOLD_MINUTES)
with SessionLocal() as session:
result = session.execute(
update(HeadAutoApplyRun)
.where(HeadAutoApplyRun.status == "running")
.where(
func.coalesce(
HeadAutoApplyRun.last_progress_at, HeadAutoApplyRun.started_at
)
< cutoff
)
.values(
status="error", finished_at=now,
error=(
f"stranded by recovery sweep (no progress for "
f"{HEAD_AUTO_APPLY_STALL_THRESHOLD_MINUTES} min)"
),
)
)
keep = session.execute(
select(HeadAutoApplyRun.id).order_by(HeadAutoApplyRun.id.desc())
.limit(HEAD_AUTO_APPLY_KEEP_RUNS)
).scalars().all()
if keep:
session.execute(
delete(HeadAutoApplyRun).where(HeadAutoApplyRun.id.not_in(keep))
)
session.commit()
recovered = result.rowcount or 0
if recovered:
log.info(
"recover_stalled_head_auto_apply_runs: recovered %d rows", recovered
)
return recovered
return _recover_stalled_runs(
HeadAutoApplyRun,
stall_minutes=HEAD_AUTO_APPLY_STALL_THRESHOLD_MINUTES,
keep_runs=HEAD_AUTO_APPLY_KEEP_RUNS,
label="recover_stalled_head_auto_apply_runs",
)
# Keep ~6 months of daily head-metric snapshots (enough to see tuning trends).
+10 -30
View File
@@ -105,9 +105,7 @@ def embed_image(self, image_id: int) -> dict:
record = session.get(ImageRecord, image_id)
if record is None:
return {"status": "missing", "image_id": image_id}
settings = session.execute(
select(MLSettings).where(MLSettings.id == 1)
).scalar_one()
settings = MLSettings.load_sync(session)
src = Path(record.path)
is_vid = _is_video(src)
@@ -488,15 +486,10 @@ def scheduled_ccip_auto_apply() -> str:
from sqlalchemy import select as sa_select
from sqlalchemy.dialects.postgresql import insert as pg_insert
from ..models import ImageRegion, MLSettings, Tag, TagKind, TagSuggestionRejection
from ..models import ImageRegion, MLSettings, Tag, TagKind
from ..models.tag import image_tag
fig = ("face", "figure")
def _l2(m):
n = np.linalg.norm(m, axis=1, keepdims=True)
n[n == 0] = 1.0
return m / n
from ..services.ml.ccip import _FIGURE_KINDS
from ..services.ml.training_data import _applied_or_rejected, _l2norm
SessionLocal = _sync_session_factory()
with SessionLocal() as session:
@@ -521,7 +514,7 @@ def scheduled_ccip_auto_apply() -> str:
)
.join(Tag, Tag.id == image_tag.c.tag_id)
.where(Tag.kind == TagKind.character)
.where(ImageRegion.kind.in_(fig))
.where(ImageRegion.kind.in_(_FIGURE_KINDS))
.where(ImageRegion.ccip_embedding.is_not(None))
.where(ImageRegion.image_record_id.in_(single))
).all()
@@ -532,29 +525,16 @@ def scheduled_ccip_auto_apply() -> str:
for tid, vec in ref_rows:
by_char.setdefault(tid, []).append(vec)
ref_tags = list(by_char)
mats = [_l2(np.asarray(by_char[t], dtype=np.float32)) for t in ref_tags]
mats = [_l2norm(np.asarray(by_char[t], dtype=np.float32), np) for t in ref_tags]
allref = np.vstack(mats) # (total, 768)
seg = np.cumsum([0] + [len(m) for m in mats])[:-1] # per-char start
# Per character: images that already carry OR rejected the tag — skip.
skip = {t: set() for t in ref_tags}
for t in ref_tags:
for (iid,) in session.execute(
sa_select(image_tag.c.image_record_id).where(
image_tag.c.tag_id == t
)
):
skip[t].add(iid)
for (iid,) in session.execute(
sa_select(TagSuggestionRejection.image_record_id).where(
TagSuggestionRejection.tag_id == t
)
):
skip[t].add(iid)
skip = _applied_or_rejected(session, ref_tags)
img_ids = list(session.execute(
sa_select(ImageRegion.image_record_id)
.where(ImageRegion.kind.in_(fig), ImageRegion.ccip_embedding.is_not(None))
.where(ImageRegion.kind.in_(_FIGURE_KINDS), ImageRegion.ccip_embedding.is_not(None))
.distinct()
).scalars())
@@ -566,7 +546,7 @@ def scheduled_ccip_auto_apply() -> str:
sa_select(ImageRegion.image_record_id, ImageRegion.ccip_embedding)
.where(
ImageRegion.image_record_id.in_(chunk),
ImageRegion.kind.in_(fig),
ImageRegion.kind.in_(_FIGURE_KINDS),
ImageRegion.ccip_embedding.is_not(None),
)
).all()
@@ -574,7 +554,7 @@ def scheduled_ccip_auto_apply() -> str:
for iid, vec in rows:
by_img.setdefault(iid, []).append(vec)
for iid, vecs in by_img.items():
q = _l2(np.asarray(vecs, dtype=np.float32)) # (nq, 768)
q = _l2norm(np.asarray(vecs, dtype=np.float32), np) # (nq, 768)
colmax = (q @ allref.T).max(axis=0) # (total,)
charmax = np.maximum.reduceat(colmax, seg) # (n_chars,)
for ci in np.where(charmax >= thr)[0]:
+39 -3
View File
@@ -11,12 +11,26 @@ git.fabledsword.com/bvandeusen/ci-python:3.14
- python 3.14
- ruff (analyzer for `backend/`, `tests/`, `alembic/`)
- node (frontend job: `npm install` + vitest + vite build)
- docker CLI + buildx (`.forgejo/workflows/build.yml`: build-web, build-ml — Forgejo registry push)
- docker CLI + buildx (`.forgejo/workflows/build.yml`: build-web, build-ml — Fabled-Git registry push)
## Secondary runtime image
node:24-bookworm-slim — `.forgejo/workflows/extension.yml` only.
The extension lane is the one job that does NOT run on `ci-python:3.14`: it
needs a current Node for `web-ext` and vitest and nothing Python at all. Kept
on the upstream slim image rather than adding a Node toolchain to `ci-python`,
per `docs/process.md`'s "add deps to the image when used by >1 project".
## Per-job tool installs
- `pip install -r requirements.txt pytest pytest-asyncio` — in `backend-lint-and-test` and `integration` jobs
- `npm install --no-audit --no-fund` — in `frontend-build` job
- `npm install --no-audit --no-fund` — in `extension.yml`'s `lint` job (web-ext + vitest)
- `unzip` — in `extension.yml`'s "Verify XPI contents" step, installed via apt
only when absent (`node:24-bookworm-slim` may or may not carry it). Debian
package, ~2s. Not worth baking into a shared image for a single consumer, per
`docs/process.md`'s ">1 project" rule.
## Notes
@@ -26,12 +40,34 @@ git.fabledsword.com/bvandeusen/ci-python:3.14
"add deps to image when used by >1 project" rule: FC alone is one Python
project, so the deps live in `requirements.txt` and install per-job.
Reconsider when a second Fabled-family Python backend lands.
- Integration uses Forgejo Actions `services:` + socket-discovered bridge IPs
- Integration uses Fabled-Git Actions `services:` + socket-discovered bridge IPs
because `act_runner` (swarm-runner v0.6+) puts services on the default
bridge with no embedded DNS. The pattern is documented in the rulebook's
`forgejo.md` "CI philosophy" section and FC's `ci.yml` is the canonical
`fabled-git.md` "CI philosophy" section and FC's `ci.yml` is the canonical
example.
- No `package-lock.json` is tracked yet (FC's `feedback_no_local_runs`
memory bans `npm install` locally). Using `npm install` rather than
`npm ci` until a lockfile lands.
- No `imagemagick` / `pandoc` per-job installs needed.
- `extension/`'s vitest specs load `lib/*.js` by evaluating the real file as a
classic script (`test/helpers/loadLib.js`) rather than adding `module.exports`
shims to production code — the libs ship as `background.scripts`, not ES
modules, so the specs exercise exactly the bytes packaged into the XPI.
- **`extension/scripts/packaging.sh` is the single definition of what ships
inside the XPI.** Three consumers read from it rather than keeping their own
copy: web-ext's `--ignore-files` (`extension/package.json`), the `:(exclude)`
pathspec in `ci.yml`'s `extension-version` guard, and the `git log` pathspec
that derives the extension version. Three hand-kept copies of that one fact
is what allowed issue #2397.
- Jobs that derive the extension version check out with `fetch-depth: 0`. The
version is the commit TIME of the newest packaged-extension change (minutes
since 2020-01-01, per family rule 149 — never a commit count, which orders
by branch rather than by recency). A depth-1 clone sees one commit and
derives a wrong, too-low value rather than failing, so the full-history
checkout is load-bearing wherever `packaging.sh version` is called.
- Callers MUST `set -f` before substituting the script's output. Without it the
shell expands `test/**` against the working tree and silently narrows the
pattern to whatever files exist at that moment — a failure that looks like
nothing until dev files start appearing in the XPI. `test/version.spec.js`
asserts every `--ignore-files` consumer sets it, and that no consumer has
quietly reinstated a hardcoded list.
+3 -3
View File
@@ -1,9 +1,9 @@
# FabledCurator Firefox Extension
Self-hosted Firefox extension that pushes session cookies from supported
platforms (Patreon, SubscribeStar, Hentai-Foundry, Discord, Pixiv,
DeviantArt) into FabledCurator, and lets you add a creator as a Source
from their page in one click.
platforms (Patreon, SubscribeStar, Hentai-Foundry, Discord, Pixiv)
into FabledCurator, and lets you add a creator as a Source from their
page in one click.
## Install (operator)
+92 -30
View File
@@ -31,6 +31,68 @@ browser.runtime.onInstalled.addListener(() => ensureInitialized());
browser.runtime.onStartup.addListener(() => ensureInitialized());
ensureInitialized().catch(e => console.error('init failed:', e));
// ---- Extension self-update check (#1489) ----
// Installed per-instance from the operator's FC host, so Firefox's static
// update_url can't apply (each instance has a different host). Instead ask the
// configured backend for the latest published version and nudge the operator to
// reinstall the freshly-signed XPI — surfaced as a popup banner (on demand) and
// a toolbar badge (daily). /api/extension/manifest is public and returns
// {version, latest_url, sha256}; the XPI is served from the web root (not /api).
function versionIsNewer(candidate, current) {
// Dotted numeric compare so 1.0.10 > 1.0.9 (a plain string compare wouldn't).
const a = String(candidate).split('.').map(n => parseInt(n, 10) || 0);
const b = String(current).split('.').map(n => parseInt(n, 10) || 0);
for (let i = 0; i < Math.max(a.length, b.length); i++) {
if ((a[i] || 0) !== (b[i] || 0)) return (a[i] || 0) > (b[i] || 0);
}
return false;
}
async function checkForUpdateInfo() {
await ensureInitialized();
if (!api.isConfigured()) return { updateAvailable: false, configured: false };
let info;
try {
info = await api.getExtensionManifest();
} catch (e) {
return { updateAvailable: false, error: e.message };
}
const currentVersion = browser.runtime.getManifest().version;
const latestVersion = info && info.version ? info.version : null;
// latest_url is served from the web root, not the JSON API.
const base = api.webRoot();
return {
updateAvailable: !!latestVersion && versionIsNewer(latestVersion, currentVersion),
currentVersion,
latestVersion,
xpiUrl: info && info.latest_url ? `${base}${info.latest_url}` : null,
};
}
async function refreshUpdateBadge() {
let r;
try { r = await checkForUpdateInfo(); } catch { return; }
try {
await browser.action.setBadgeText({ text: r.updateAvailable ? '↑' : '' });
if (r.updateAvailable) {
await browser.action.setBadgeBackgroundColor({ color: '#F4BA7A' });
await browser.action.setTitle({ title: `FabledCurator — update available (v${r.latestVersion})` });
} else {
await browser.action.setTitle({ title: 'FabledCurator' });
}
} catch { /* action API unavailable — non-fatal */ }
}
// Daily proactive check (needs the "alarms" permission). create() is idempotent
// by name, so re-running it on each event-page load is safe.
browser.alarms.create('fc-update-check', { periodInMinutes: 24 * 60, delayInMinutes: 1 });
browser.alarms.onAlarm.addListener((alarm) => {
if (alarm.name === 'fc-update-check') refreshUpdateBadge();
});
browser.runtime.onStartup.addListener(() => refreshUpdateBadge());
browser.runtime.onInstalled.addListener(() => refreshUpdateBadge());
// ---- Discord token capture via webRequest ----
browser.webRequest.onBeforeSendHeaders.addListener(
@@ -148,6 +210,21 @@ browser.webRequest.onBeforeRedirect.addListener(
{ urls: ['https://app-api.pixiv.net/web/v1/users/auth/pixiv/callback*'] },
);
// Extract → verify → upload one cookie-auth platform. Returns a structured
// outcome so the two callers (EXPORT_COOKIES single, EXPORT_ALL_COOKIES) shape
// their own response + skip semantics. Verifies the captured cookies are
// actually live BEFORE uploading, so a confirmed-stale session doesn't overwrite
// good FC-side credentials; platforms with no verify config (v.ok === null) fall
// through to upload.
async function exportPlatformCookies(key) {
const cookies = await extractCookiesForPlatform(key);
if (cookies.length === 0) return { status: 'empty' };
const v = await verifyCookiesForPlatform(key);
if (v.ok === false) return { status: 'stale', reason: v.reason, cookieCount: cookies.length };
await api.uploadCredentials(key, 'cookies', toNetscapeFormat(cookies));
return { status: 'ok', cookieCount: cookies.length, verified: v.ok === true };
}
// ---- Message router ----
browser.runtime.onMessage.addListener(async (msg) => {
@@ -192,22 +269,14 @@ browser.runtime.onMessage.addListener(async (msg) => {
if (!platform) return { error: `Unknown platform: ${key}` };
try {
if (platform.authType === 'cookies') {
const cookies = await extractCookiesForPlatform(key);
if (cookies.length === 0) return { error: 'No cookies found — log in first.' };
// Verify the captured cookies are actually live BEFORE
// uploading. Skips upload on confirmed-stale sessions so we
// don't overwrite FC-side credentials with garbage. Platforms
// without a verify config (verify.ok === null) fall through
// to upload as before.
const v = await verifyCookiesForPlatform(key);
if (v.ok === false) {
const r = await exportPlatformCookies(key);
if (r.status === 'empty') return { error: 'No cookies found — log in first.' };
if (r.status === 'stale') {
return {
error: `Captured ${cookies.length} ${platform.name} cookies but they don't appear authenticated (${v.reason}). Log in again in this browser, then retry.`,
error: `Captured ${r.cookieCount} ${platform.name} cookies but they don't appear authenticated (${r.reason}). Log in again in this browser, then retry.`,
};
}
const data = toNetscapeFormat(cookies);
await api.uploadCredentials(key, 'cookies', data);
return { success: true, cookieCount: cookies.length, verified: v.ok === true };
return { success: true, cookieCount: r.cookieCount, verified: r.verified };
}
if (key === 'discord') {
if (!discordToken) return { error: 'Open discord.com to capture a token first.' };
@@ -235,18 +304,10 @@ browser.runtime.onMessage.addListener(async (msg) => {
continue;
}
try {
const cookies = await extractCookiesForPlatform(key);
if (cookies.length === 0) {
results[key] = { skipped: true, reason: 'no cookies' };
continue;
}
const v = await verifyCookiesForPlatform(key);
if (v.ok === false) {
results[key] = { error: `verify failed: ${v.reason}` };
continue;
}
await api.uploadCredentials(key, 'cookies', toNetscapeFormat(cookies));
results[key] = { success: true, cookieCount: cookies.length, verified: v.ok === true };
const r = await exportPlatformCookies(key);
if (r.status === 'empty') results[key] = { skipped: true, reason: 'no cookies' };
else if (r.status === 'stale') results[key] = { error: `verify failed: ${r.reason}` };
else results[key] = { success: true, cookieCount: r.cookieCount, verified: r.verified };
} catch (e) {
results[key] = { error: e.message };
}
@@ -283,11 +344,9 @@ browser.runtime.onMessage.addListener(async (msg) => {
}
case 'OPEN_ARTIST_PAGE': {
// apiUrl is configured with the /api suffix (see
// options/options.html placeholder); the SPA artist route is
// /artist/:slug, served from the same origin. Strip /api so the
// browser-level URL hits the Vue router, not the JSON API.
const base = (api.baseUrl || '').replace(/\/+$/, '').replace(/\/api$/, '');
// The SPA artist route (/artist/:slug) is served from the web root, not
// the JSON API — see api.webRoot().
const base = api.webRoot();
const slug = encodeURIComponent(msg.slug || '');
if (!base || !slug) return { error: 'apiUrl or slug missing' };
try {
@@ -298,6 +357,9 @@ browser.runtime.onMessage.addListener(async (msg) => {
}
}
case 'CHECK_UPDATE':
return await checkForUpdateInfo();
default:
return { error: `Unknown message type: ${msg.type}` };
}
+23 -1
View File
@@ -11,7 +11,10 @@ class FabledCuratorAPI {
async init() {
const cfg = await browser.storage.local.get(['apiUrl', 'apiKey']);
this.baseUrl = cfg.apiUrl || null;
// Normalize on READ, not just on save: configs stored before the options
// page started normalizing are missing the `/api` suffix, and this heals
// them without the operator having to reopen Settings.
this.baseUrl = normalizeApiUrl(cfg.apiUrl) || null;
this.apiKey = cfg.apiKey || null;
return this.isConfigured();
}
@@ -50,6 +53,13 @@ class FabledCuratorAPI {
} catch {
message = `HTTP ${response.status}: ${response.statusText}`;
}
// 404/405 from FC almost always means the request never reached the JSON
// API — it fell through to the SPA catch-all, which serves HTML on GET
// and rejects everything else. Say so, rather than making the operator
// decode "Method Not Allowed" on an endpoint that plainly allows POST.
if (response.status === 404 || response.status === 405) {
message += `${url} isn't the FC API. Check the FC URL in settings.`;
}
const err = new Error(message);
err.status = response.status;
throw err;
@@ -89,6 +99,18 @@ class FabledCuratorAPI {
const qs = new URLSearchParams({ url }).toString();
return this.request('GET', `/extension/probe?${qs}`);
}
// Latest published extension version on this instance — drives the in-app
// update prompt. Public endpoint (no key needed, but request() sends it
// harmlessly). Returns {version, xpi_url, latest_url, sha256}.
getExtensionManifest() {
return this.request('GET', '/extension/manifest');
}
// The web/SPA root: where the Vue router (artist pages) and the served XPI
// live, NOT the JSON API. Used by OPEN_ARTIST_PAGE + the self-update check.
webRoot() {
return webRootFromApiUrl(this.baseUrl);
}
// Connection test = the cheapest read with auth.
testConnection() {
-11
View File
@@ -68,16 +68,6 @@ const PLATFORMS = {
urlPattern: /^https?:\/\/(www\.)?pixiv\.net/,
note: 'Click to authenticate via OAuth',
},
deviantart: {
name: 'DeviantArt',
domains: ['.deviantart.com', 'www.deviantart.com', 'deviantart.com'],
authType: 'cookies',
color: '#05CC47',
urlPattern: /^https?:\/\/(www\.)?deviantart\.com/,
// DA's logged-in-only endpoints sit behind their internal _napi
// namespace which shifts; skipping verify until a stable check
// surfaces. Same posture as SubscribeStar.
},
};
/**
@@ -98,7 +88,6 @@ const PLATFORM_ARTIST_PATTERNS = {
patreon: /^https?:\/\/(www\.)?patreon\.com\/(?:cw\/|c\/)?(?!(?:home|search|messages|notifications|library|settings|posts)(?:[\/?#]|$))[^/?#]+/i,
subscribestar: /^https?:\/\/(www\.)?subscribestar\.(com|adult)\/(?!feed$|messages$|library$)[^/?#]+\/?$/i,
hentaifoundry: /^https?:\/\/(www\.)?hentai-foundry\.com\/user\/[^/?#]+/i,
deviantart: /^https?:\/\/(www\.)?deviantart\.com\/(?!home$|watch\b|tag\b|browse\b)[^/?#]+\/?$/i,
pixiv: /^https?:\/\/(www\.)?pixiv\.net\/(en\/)?users\/\d+/i,
};
+32
View File
@@ -0,0 +1,32 @@
/**
* Canonical FC endpoint derivation, shared by the background client and the
* options page so a URL entered either way behaves identically.
*
* FC serves two things on one origin: the JSON API under `/api`, and the Vue
* SPA from the root. `api.js` builds requests as `${baseUrl}/credentials`, so
* the stored base URL has to carry the `/api` suffix.
*/
/**
* Accept what an operator would naturally type — the instance root
* (`http://curator.example.com`) or the API root (`.../api`) — and return the
* API root either way.
*
* Worth normalizing rather than validating: a root-form URL doesn't fail
* loudly, it lands on the SPA catch-all, which answers `GET /credentials` with
* 200 HTML and rejects `POST /credentials` with 405. The operator sees a
* working Test Connection and a broken export.
*/
function normalizeApiUrl(raw) {
const trimmed = (raw || '').trim().replace(/\/+$/, '');
if (!trimmed) return '';
return /\/api$/i.test(trimmed) ? trimmed : `${trimmed}/api`;
}
/**
* The SPA root — where the Vue router (artist pages) and the served XPI live,
* NOT the JSON API. Accepts either input form, same as normalizeApiUrl.
*/
function webRootFromApiUrl(raw) {
return normalizeApiUrl(raw).replace(/\/api$/i, '');
}
+4 -5
View File
@@ -1,7 +1,7 @@
{
"manifest_version": 3,
"name": "FabledCurator",
"version": "1.0.8",
"version": "1.0.11",
"description": "Export cookies from supported platforms to FabledCurator and add creators as sources in one click.",
"browser_specific_settings": {
@@ -22,7 +22,8 @@
"tabs",
"activeTab",
"webRequest",
"webRequestBlocking"
"webRequestBlocking",
"alarms"
],
"host_permissions": [
@@ -32,7 +33,6 @@
"*://*.hentai-foundry.com/*",
"*://*.discord.com/*",
"*://*.pixiv.net/*",
"*://*.deviantart.com/*",
"*://app-api.pixiv.net/*",
"*://oauth.secure.pixiv.net/*",
"*://*/*"
@@ -45,7 +45,7 @@
},
"background": {
"scripts": ["lib/platforms.js", "lib/cookies.js", "lib/api.js", "background/background.js"]
"scripts": ["lib/platforms.js", "lib/cookies.js", "lib/url.js", "lib/api.js", "background/background.js"]
},
"options_ui": {
@@ -60,7 +60,6 @@
"*://*.subscribestar.com/*",
"*://*.subscribestar.adult/*",
"*://*.hentai-foundry.com/*",
"*://*.deviantart.com/*",
"*://*.pixiv.net/*"
],
"js": ["lib/platforms.js", "content/content-script.js"],
+7 -3
View File
@@ -21,9 +21,12 @@
<body>
<h1>FabledCurator extension</h1>
<label for="api-url">FC base URL</label>
<input id="api-url" type="url" placeholder="http://curator.example.com/api" />
<div class="hint">Find this on FC → Settings → Maintenance → Browser extension.</div>
<label for="api-url">FC instance URL</label>
<input id="api-url" type="url" placeholder="http://curator.example.com" />
<div class="hint">
Your FabledCurator address — with or without the trailing <code>/api</code>; both work.
Find it on FC → Settings → Maintenance → Browser extension.
</div>
<label for="api-key">Extension API key</label>
<input id="api-key" type="password" placeholder="paste from FC Settings card" />
@@ -36,6 +39,7 @@
<div id="status" class="status" style="display:none;"></div>
<script src="../lib/url.js"></script>
<script src="options.js"></script>
</body>
</html>
+23 -5
View File
@@ -8,7 +8,7 @@ document.addEventListener('DOMContentLoaded', async () => {
});
async function save() {
const apiUrl = document.getElementById('api-url').value.trim().replace(/\/+$/, '');
const apiUrl = normalizeApiUrl(document.getElementById('api-url').value);
const apiKey = document.getElementById('api-key').value.trim();
if (!apiUrl || !apiKey) {
showStatus('Both fields are required.', 'err');
@@ -16,11 +16,14 @@ async function save() {
}
await browser.storage.local.set({ apiUrl, apiKey });
await browser.storage.local.remove(['lastConnectionTest', 'lastConnectionStatus']);
showStatus('Saved.', 'ok');
// Show what was actually stored — the operator may have typed the instance
// root and it was normalized to the API root.
document.getElementById('api-url').value = apiUrl;
showStatus(`Saved — using ${apiUrl}`, 'ok');
}
async function test() {
const apiUrl = document.getElementById('api-url').value.trim().replace(/\/+$/, '');
const apiUrl = normalizeApiUrl(document.getElementById('api-url').value);
const apiKey = document.getElementById('api-key').value.trim();
if (!apiUrl || !apiKey) {
showStatus('Fill both fields first.', 'err');
@@ -31,8 +34,23 @@ async function test() {
method: 'GET',
headers: { 'X-Extension-Key': apiKey },
});
if (r.ok) showStatus(`Connected — HTTP ${r.status}.`, 'ok');
else showStatus(`HTTP ${r.status}: ${r.statusText}`, 'err');
if (!r.ok) {
showStatus(`HTTP ${r.status}: ${r.statusText}`, 'err');
return;
}
// A 200 is NOT sufficient. If the URL resolves to the Vue SPA instead of
// the JSON API, the catch-all route returns 200 with an HTML document —
// which used to report "Connected" on a config that could not POST at all.
const contentType = r.headers.get('content-type') || '';
if (!contentType.includes('json')) {
showStatus(
`${apiUrl} answered with ${contentType || 'no content-type'}, not JSON `
+ '— that looks like the FC web UI rather than its API.',
'err',
);
return;
}
showStatus(`Connected to ${apiUrl} — HTTP ${r.status}.`, 'ok');
} catch (e) {
showStatus(`Cannot reach ${apiUrl}: ${e.message}`, 'err');
}
+8 -5
View File
@@ -1,15 +1,18 @@
{
"name": "fabledcurator-extension",
"version": "1.0.8",
"version": "1.0.11",
"private": true,
"description": "Firefox extension for FabledCurator",
"comment_ignore_files": "The --ignore-files list comes from scripts/packaging.sh, the single source of truth shared with ci.yml's guard and the derived-version patch count. `set -f` is REQUIRED before the substitution: without it the shell globs `test/**` against the working tree and silently narrows the pattern to whatever files happen to exist.",
"scripts": {
"lint": "web-ext lint --source-dir=. --no-config-discovery --ignore-files package.json package-lock.json web-ext-artifacts node_modules README.md .gitignore",
"start": "web-ext run --source-dir=. --no-config-discovery --ignore-files package.json package-lock.json web-ext-artifacts node_modules README.md .gitignore --firefox=firefox",
"build": "web-ext build --source-dir=. --no-config-discovery --ignore-files package.json package-lock.json web-ext-artifacts node_modules README.md .gitignore --overwrite-dest",
"sign": "web-ext sign --source-dir=. --no-config-discovery --ignore-files package.json package-lock.json web-ext-artifacts node_modules README.md .gitignore --channel=unlisted --api-key=$WEB_EXT_API_KEY --api-secret=$WEB_EXT_API_SECRET"
"lint": "set -f; web-ext lint --source-dir=. --no-config-discovery --ignore-files $(sh scripts/packaging.sh ignore)",
"start": "set -f; web-ext run --source-dir=. --no-config-discovery --ignore-files $(sh scripts/packaging.sh ignore) --firefox=firefox",
"build": "set -f; web-ext build --source-dir=. --no-config-discovery --ignore-files $(sh scripts/packaging.sh ignore) --overwrite-dest",
"sign": "set -f; web-ext sign --source-dir=. --no-config-discovery --ignore-files $(sh scripts/packaging.sh ignore) --channel=unlisted --api-key=$WEB_EXT_API_KEY --api-secret=$WEB_EXT_API_SECRET",
"test:unit": "vitest run"
},
"devDependencies": {
"vitest": "^4.0.0",
"web-ext": "^10.0.0"
}
}
+11
View File
@@ -72,6 +72,17 @@ body {
.btn.block { display: block; width: 100%; margin-top: 8px; }
.btn.link { background: none; color: var(--on-surface-variant); padding: 4px; }
.btn.link:hover { color: var(--accent); }
.btn.small { padding: 6px 12px; font-size: 13px; }
/* In-app update prompt (accent-tinted so it reads as an actionable notice). */
.update-banner {
display: flex; align-items: center; gap: 10px;
margin: 10px 10px 0; padding: 10px 12px;
background: rgba(244, 186, 122, 0.12);
border: 1px solid rgba(244, 186, 122, 0.4);
border-radius: 6px;
}
#update-text { flex: 1; font-size: 13px; }
.source-row .play {
background: none; border: none; color: var(--on-surface-variant);
+5
View File
@@ -20,6 +20,11 @@
</section>
<section id="main-content" class="main hidden">
<div id="update-banner" class="update-banner hidden">
<span id="update-text"></span>
<button id="update-btn" class="btn primary small">Update</button>
</div>
<nav class="tabs">
<button class="tab active" data-tab="platforms">Platforms</button>
<button class="tab" data-tab="sources">Sources</button>
+33 -12
View File
@@ -2,6 +2,15 @@ document.addEventListener('DOMContentLoaded', init);
const CONNECTION_TEST_INTERVAL = 2 * 60 * 1000;
// A centered muted note div — the loading / empty state shared by the platform
// and sources lists.
function mutedNote(text) {
const d = document.createElement('div');
d.style.cssText = 'text-align:center;padding:18px;color:var(--on-surface-variant);';
d.textContent = text;
return d;
}
async function init() {
try {
const cfg = await browser.runtime.sendMessage({ type: 'GET_CONFIG' });
@@ -14,6 +23,7 @@ async function init() {
setupEventListeners();
showPlatformsLoading();
testConnectionIfNeeded();
checkForUpdate();
loadPlatformStatus().catch(e => showError(`Failed to load platforms: ${e.message}`));
} catch (e) {
showSetupRequired();
@@ -37,10 +47,7 @@ function showSetupRequired() {
function showPlatformsLoading() {
const c = document.getElementById('platforms-list');
c.textContent = '';
const d = document.createElement('div');
d.style.cssText = 'text-align:center;padding:18px;color:var(--on-surface-variant);';
d.textContent = 'Loading platforms…';
c.appendChild(d);
c.appendChild(mutedNote('Loading platforms…'));
}
async function testConnectionIfNeeded() {
@@ -63,6 +70,26 @@ function updateConnectionDot(connected) {
d.title = connected ? 'Connected to FabledCurator' : 'Disconnected';
}
// Nudge to reinstall when the configured instance publishes a newer signed XPI
// (the extension is self-hosted, so there's no Firefox auto-update). Never
// blocks the popup — a failed check just leaves the banner hidden.
async function checkForUpdate() {
try {
const r = await browser.runtime.sendMessage({ type: 'CHECK_UPDATE' });
if (r && r.updateAvailable && r.xpiUrl) showUpdateBanner(r);
} catch { /* non-fatal */ }
}
function showUpdateBanner(r) {
document.getElementById('update-text').textContent =
`Update available — v${r.latestVersion} (installed v${r.currentVersion})`;
// Opening the signed XPI triggers Firefox's native install prompt.
document.getElementById('update-btn').addEventListener('click', () => {
browser.tabs.create({ url: r.xpiUrl });
});
document.getElementById('update-banner').classList.remove('hidden');
}
async function loadPlatformStatus() {
const status = await browser.runtime.sendMessage({ type: 'GET_PLATFORM_STATUS' });
const c = document.getElementById('platforms-list');
@@ -162,10 +189,7 @@ async function exportAllCookies() {
async function loadSources() {
const c = document.getElementById('sources-list');
c.textContent = '';
const d = document.createElement('div');
d.style.cssText = 'text-align:center;padding:18px;color:var(--on-surface-variant);';
d.textContent = 'Loading sources…';
c.appendChild(d);
c.appendChild(mutedNote('Loading sources…'));
const r = await browser.runtime.sendMessage({ type: 'LIST_SOURCES' });
c.textContent = '';
if (r.error) {
@@ -176,10 +200,7 @@ async function loadSources() {
return;
}
if (!r.sources || r.sources.length === 0) {
const empty = document.createElement('div');
empty.style.cssText = 'text-align:center;padding:18px;color:var(--on-surface-variant);';
empty.textContent = 'No sources yet.';
c.appendChild(empty);
c.appendChild(mutedNote('No sources yet.'));
return;
}
for (const src of r.sources) c.appendChild(createSourceRow(src));
+130
View File
@@ -0,0 +1,130 @@
#!/bin/sh
# Single source of truth for "what ships inside the XPI", plus the version
# derived from it.
#
# Three consumers used to hand-maintain their own copy of this list, and
# keeping three copies of one fact in sync by hand is how issue #2397 happened:
#
# 1. web-ext's --ignore-files (extension/package.json's four scripts)
# 2. the :(exclude) pathspec (ci.yml's extension-version guard)
# 3. the git-log pathspec (the derived version, below)
#
# They now all read from here. POSIX sh only — CI's run shell is busybox.
#
# -f (no pathname expansion) is set for the whole script and is load-bearing:
# the lists below are iterated with deliberate word-splitting, and without -f
# the shell would also GLOB them, expanding `test/**` into whatever files
# happen to exist and corrupting the output. A caller's own `set -f` does not
# help here — this runs as a separate sh process and does not inherit it.
# Callers still need their own `set -f` for the substituted result; the two
# guards protect different expansions.
set -euf
# Paths under extension/ that are NOT packaged into the XPI.
#
# Split by whether git tracks them: node_modules and web-ext-artifacts are
# build/dependency output that never appears in a commit, so they belong in
# web-ext's ignore list but would be meaningless in a git pathspec.
#
# Directories need BOTH forms. `test/**` matches the files inside, but not the
# directory entry itself — web-ext writes an entry for the directory too, so
# with only the glob the XPI ends up carrying empty `test/` and `scripts/`
# entries (caught by the XPI-content check on 2026-08-03). The bare name alone
# is not enough either: minimatch's `test` does not match `test/url.spec.js`,
# so dropping the glob would ship the contents. Keep both.
NOT_PACKAGED_TRACKED='package.json package-lock.json README.md .gitignore vitest.config.js scripts scripts/** test test/**'
NOT_PACKAGED_BUILD='web-ext-artifacts node_modules'
usage() {
echo "usage: packaging.sh {ignore|pathspec|version|major-minor|patch}" >&2
exit 2
}
# web-ext --ignore-files values, space-separated.
#
# Callers MUST disable pathname expansion first (`set -f`), or the shell will
# glob `test/**` against the working tree before web-ext ever sees the pattern
# and silently narrow it to whatever happens to exist right now.
cmd_ignore() {
echo "$NOT_PACKAGED_TRACKED $NOT_PACKAGED_BUILD"
}
# git pathspec excluding the non-packaged tracked files, e.g.
# :(exclude)extension/package.json :(exclude)extension/test/**
# Same `set -f` requirement as above.
cmd_pathspec() {
for entry in $NOT_PACKAGED_TRACKED; do
printf ':(exclude)extension/%s ' "$entry"
done
echo
}
# MAJOR.MINOR stays hand-set in manifest.json — it's the part that carries
# deliberate meaning. Only the patch component is derived.
cmd_major_minor() {
root=$(git rev-parse --show-toplevel)
grep -E '"version"' "$root/extension/manifest.json" \
| head -1 \
| sed -E 's/.*"version"[[:space:]]*:[[:space:]]*"([0-9]+)\.([0-9]+).*/\1.\2/'
}
# 2020-01-01T00:00:00Z — the anchor for the derived patch component. Fixed
# forever; moving it would renumber every version downwards.
VERSION_EPOCH=1577836800
# Minutes since VERSION_EPOCH of the LATEST commit that touched a PACKAGED
# extension file.
#
# Time-derived, per family rule 149: an artifact's ordering key must never be a
# commit count. A count is per-branch — `dev` and `main` count different
# histories of the same code — so the moment BOTH channels publish, their
# versions order by which branch accumulated more commits rather than by which
# is newer. A squash-merge makes that permanent: main gains one commit where dev
# gained five, so dev climbs away from main and a dev install can never cross
# back. That is Roundtable's 2026-08-24 incident (`versionCode` was the branch's
# commit count) in a different repo. Measured here on 2026-08-27: main=23,
# dev=24 under the old formula — one apart, which is exactly how the inversion
# stays invisible until it strands somebody.
#
# Why the commit's time and not the build's:
# * MONOTONIC — max() over a set that only ever gains members. Verified
# across all 24 extension-touching commits: zero non-monotonic steps.
# * STABLE while the extension is unchanged, so an unchanged extension keeps
# its version, the ext-<version> signature cache still hits, and AMO is
# called once per extension CHANGE rather than once per push. Build-time
# minutes would re-sign on every push and never let two channels share a
# signature.
# * SHARED ACROSS CHANNELS — after a merge, `main` sees the same commit and
# derives the same number, so `:latest` reuses the signature `:dev` already
# produced for byte-identical code. Same code, same version, one signing.
# * REPRODUCIBLE — any checkout of a commit yields that commit's version.
#
# Requires real history: a depth-1 clone sees one commit and will derive a wrong
# (too low) value. Every consumer must check out with fetch-depth: 0.
cmd_patch() {
root=$(git rev-parse --show-toplevel)
# Unquoted on purpose: the pathspec must word-split into separate args.
# Globbing is already off script-wide (set -euf above).
# shellcheck disable=SC2046
ts=$(cd "$root" && git log --format=%ct HEAD -- extension/ $(cmd_pathspec) \
| sort -n | tail -1)
if [ -z "$ts" ]; then
echo "packaging.sh: no commit touches a packaged extension file" >&2
exit 1
fi
echo $(( (ts - VERSION_EPOCH) / 60 ))
}
cmd_version() {
echo "$(cmd_major_minor).$(cmd_patch)"
}
[ $# -ge 1 ] || usage
case "$1" in
ignore) cmd_ignore ;;
pathspec) cmd_pathspec ;;
version) cmd_version ;;
major-minor) cmd_major_minor ;;
patch) cmd_patch ;;
*) usage ;;
esac
+29
View File
@@ -0,0 +1,29 @@
import { readFileSync } from 'node:fs'
import { fileURLToPath } from 'node:url'
import path from 'node:path'
const LIB_DIR = path.join(path.dirname(fileURLToPath(import.meta.url)), '..', '..', 'lib')
/**
* Load an extension lib and hand back the globals it declares.
*
* The files under lib/ are CLASSIC scripts, not ES modules: manifest.json
* lists them in `background.scripts` and options.html pulls them in with a
* plain <script> tag, so they declare bare functions into a shared scope and
* export nothing. Rather than bolt a `module.exports` shim onto production
* code that would never run in the browser, evaluate the real file the same
* way the browser does — as a script body — and pick the declarations back out.
*
* This means the specs exercise the exact bytes that get packaged into the
* XPI. Only usable for libs that touch no browser APIs at load time
* (url.js, platforms.js); cookies.js and api.js reference `browser.*` and
* would need stubbing, which is why they aren't loaded this way.
*
* @param {string} filename e.g. 'url.js'
* @param {string[]} names declarations to return, e.g. ['normalizeApiUrl']
*/
export function loadLib(filename, names) {
const source = readFileSync(path.join(LIB_DIR, filename), 'utf8')
const factory = new Function(`${source}\nreturn { ${names.join(', ')} }`)
return factory()
}
+190
View File
@@ -0,0 +1,190 @@
import { describe, it, expect } from 'vitest'
import { readFileSync } from 'node:fs'
import { fileURLToPath } from 'node:url'
import path from 'node:path'
import { loadLib } from './helpers/loadLib.js'
const EXT_DIR = path.join(path.dirname(fileURLToPath(import.meta.url)), '..')
const manifest = JSON.parse(readFileSync(path.join(EXT_DIR, 'manifest.json'), 'utf8'))
const { getPlatformFromUrl, isArtistPage, PLATFORMS, PLATFORM_ARTIST_PATTERNS } = loadLib(
'platforms.js',
['getPlatformFromUrl', 'isArtistPage', 'PLATFORMS', 'PLATFORM_ARTIST_PATTERNS']
)
describe('getPlatformFromUrl', () => {
it('identifies each platform from a domain URL', () => {
expect(getPlatformFromUrl('https://www.patreon.com/Atole')).toBe('patreon')
expect(getPlatformFromUrl('https://subscribestar.adult/someone')).toBe('subscribestar')
expect(getPlatformFromUrl('https://www.hentai-foundry.com/user/someone')).toBe('hentaifoundry')
expect(getPlatformFromUrl('https://discord.com/channels/@me')).toBe('discord')
expect(getPlatformFromUrl('https://www.pixiv.net/en/users/123')).toBe('pixiv')
})
it('accepts http as well as https, with or without www', () => {
expect(getPlatformFromUrl('http://patreon.com/Atole')).toBe('patreon')
expect(getPlatformFromUrl('https://www.patreon.com/Atole')).toBe('patreon')
})
it('returns null for unrelated hosts', () => {
expect(getPlatformFromUrl('https://example.com/patreon.com')).toBe(null)
expect(getPlatformFromUrl('https://not-patreon.com/Atole')).toBe(null)
expect(getPlatformFromUrl('')).toBe(null)
})
it('returns null for deviantart, retired at #3069', () => {
// The 2026-07-05 product decision (FC downloaders = art-dedicated services
// only) left deviantart wired for seven weeks. Asserting the negative is
// what keeps a partial retirement from being re-completed by accident.
expect(getPlatformFromUrl('https://www.deviantart.com/someone')).toBe(null)
expect(PLATFORMS.deviantart).toBeUndefined()
expect(PLATFORM_ARTIST_PATTERNS.deviantart).toBeUndefined()
})
})
describe('isArtistPage', () => {
// Regression cases from issue #1485: the Add-to-FC button vanished once the
// operator SUBSCRIBED to a creator, because Patreon serves subscribed users
// the /cw/ ("creator workspace") URL and the pattern only matched the bare
// root. All three creator URL shapes must match, plus inner pages — the
// button matters most exactly when you're subscribed.
it('matches all three Patreon creator URL shapes', () => {
expect(isArtistPage('https://www.patreon.com/Atole', 'patreon')).toBe(true)
expect(isArtistPage('https://www.patreon.com/c/Atole', 'patreon')).toBe(true)
expect(isArtistPage('https://www.patreon.com/cw/Atole', 'patreon')).toBe(true)
})
it('matches Patreon creator inner pages', () => {
expect(isArtistPage('https://www.patreon.com/cw/Atole/posts', 'patreon')).toBe(true)
expect(isArtistPage('https://www.patreon.com/Atole/membership', 'patreon')).toBe(true)
})
it('excludes Patreon navigation pages that are not creators', () => {
for (const nav of ['home', 'search', 'messages', 'notifications', 'library', 'settings']) {
expect(isArtistPage(`https://www.patreon.com/${nav}`, 'patreon')).toBe(false)
expect(isArtistPage(`https://www.patreon.com/${nav}/anything`, 'patreon')).toBe(false)
}
})
it('matches SubscribeStar creator roots on both TLDs but not feed pages', () => {
expect(isArtistPage('https://subscribestar.adult/someone', 'subscribestar')).toBe(true)
expect(isArtistPage('https://subscribestar.com/someone', 'subscribestar')).toBe(true)
expect(isArtistPage('https://subscribestar.adult/feed', 'subscribestar')).toBe(false)
expect(isArtistPage('https://subscribestar.adult/messages', 'subscribestar')).toBe(false)
})
it('matches Hentai Foundry user pages only', () => {
expect(isArtistPage('https://www.hentai-foundry.com/user/someone', 'hentaifoundry')).toBe(true)
expect(isArtistPage('https://www.hentai-foundry.com/pictures/popular', 'hentaifoundry')).toBe(
false
)
})
it('matches Pixiv numeric user pages, with or without the /en/ prefix', () => {
expect(isArtistPage('https://www.pixiv.net/users/12345', 'pixiv')).toBe(true)
expect(isArtistPage('https://www.pixiv.net/en/users/12345', 'pixiv')).toBe(true)
expect(isArtistPage('https://www.pixiv.net/en/artworks/999', 'pixiv')).toBe(false)
})
it('returns false for a platform with no artist pattern (discord)', () => {
expect(isArtistPage('https://discord.com/channels/@me', 'discord')).toBe(false)
})
it('returns false for an unknown platform key', () => {
expect(isArtistPage('https://www.patreon.com/Atole', 'nope')).toBe(false)
})
})
describe('platform table integrity', () => {
it('gives every artist pattern a corresponding platform entry', () => {
// A pattern keyed to a platform that no longer exists is dead code that
// silently never fires; the reverse (a platform with no pattern) is the
// legitimate discord case, so only this direction is an error.
for (const key of Object.keys(PLATFORM_ARTIST_PATTERNS)) {
expect(Object.keys(PLATFORMS)).toContain(key)
}
})
it('gives every platform the fields the popup renders', () => {
for (const [key, platform] of Object.entries(PLATFORMS)) {
expect(platform.name, `${key}.name`).toBeTruthy()
expect(platform.color, `${key}.color`).toMatch(/^#[0-9A-Fa-f]{6}$/)
expect(['cookies', 'token'], `${key}.authType`).toContain(platform.authType)
expect(platform.urlPattern, `${key}.urlPattern`).toBeInstanceOf(RegExp)
expect(Array.isArray(platform.domains), `${key}.domains`).toBe(true)
expect(platform.domains.length, `${key}.domains`).toBeGreaterThan(0)
}
})
it('keeps every artist URL matched by its own platform pattern too', () => {
// isArtistPage is only ever consulted after getPlatformFromUrl resolves a
// key, so an artist pattern matching a URL its platform's urlPattern
// rejects would be unreachable.
const samples = {
patreon: 'https://www.patreon.com/cw/Atole',
subscribestar: 'https://subscribestar.adult/someone',
hentaifoundry: 'https://www.hentai-foundry.com/user/someone',
pixiv: 'https://www.pixiv.net/en/users/12345'
}
for (const [key, url] of Object.entries(samples)) {
expect(isArtistPage(url, key), `${key} artist pattern`).toBe(true)
expect(getPlatformFromUrl(url), `${key} urlPattern`).toBe(key)
}
})
})
describe('manifest.json agrees with the platform table', () => {
// #3069: deviantart was dropped from the product in July but survived in
// manifest.json until late August, because NOTHING tied the manifest's
// domain lists back to PLATFORMS. These two specs are that tie. Both
// directions matter: a stale match ships host access the product decided
// not to use, and a missing one silently kills the Add-to-FC button.
const matches = manifest.content_scripts[0].matches
// '*://*.patreon.com/*' -> '.patreon.com', the form PLATFORMS.domains uses.
const hostOf = (m) => m.replace(/^\*:\/\/\*/, '').replace(/\/\*$/, '')
it('injects the content script only on domains a platform claims', () => {
for (const m of matches) {
const host = hostOf(m)
const owner = Object.entries(PLATFORMS).find(
([, p]) => p.domains.includes(host)
)
expect(owner, `no platform claims content-script match "${m}"`).toBeTruthy()
// The content script exists to draw the Add-as-source button, so a
// platform with no artist pattern (discord) has no business here.
expect(
PLATFORM_ARTIST_PATTERNS[owner[0]],
`"${m}" injects for ${owner[0]}, which has no artist pattern`
).toBeTruthy()
}
})
it('injects on every platform that has an artist pattern', () => {
const covered = new Set(
matches
.map(hostOf)
.map((h) => Object.entries(PLATFORMS).find(([, p]) => p.domains.includes(h)))
.filter(Boolean)
.map(([key]) => key)
)
for (const key of Object.keys(PLATFORM_ARTIST_PATTERNS)) {
expect(covered, `${key} has an artist pattern but no content-script match`).toContain(key)
}
})
it('requests no host permission for a domain no platform claims', () => {
// '*://*/*' is the deliberate exception: FC is self-hosted at an arbitrary
// operator-chosen URL, so the extension cannot enumerate its own backend.
// Every OTHER entry is a platform domain and must still have an owner.
for (const h of manifest.host_permissions) {
if (h === '*://*/*') continue
const host = hostOf(h)
// pixiv's OAuth/API hosts are pixiv infrastructure, not creator pages,
// so they are matched by suffix rather than by the domains list.
const claimed = Object.values(PLATFORMS).some(
(p) => p.domains.includes(host) || p.domains.some((d) => host.endsWith(d))
)
expect(claimed, `host permission "${h}" belongs to no platform`).toBe(true)
}
})
})
+93
View File
@@ -0,0 +1,93 @@
import { describe, it, expect } from 'vitest'
import { loadLib } from './helpers/loadLib.js'
const { normalizeApiUrl, webRootFromApiUrl } = loadLib('url.js', [
'normalizeApiUrl',
'webRootFromApiUrl'
])
describe('normalizeApiUrl', () => {
// The bug this exists for (issue #2393): the instance root was accepted and
// stored verbatim, so every request went to /credentials instead of
// /api/credentials. That path is a Vue router route, so the SPA catch-all
// answered GET with 200 HTML and rejected POST with 405 — which read as a
// backend bug rather than a URL one.
it('appends /api to an instance root', () => {
expect(normalizeApiUrl('http://curator.traefik.internal')).toBe(
'http://curator.traefik.internal/api'
)
})
it('leaves an API root alone rather than doubling the suffix', () => {
expect(normalizeApiUrl('http://curator.traefik.internal/api')).toBe(
'http://curator.traefik.internal/api'
)
})
it('is idempotent', () => {
const once = normalizeApiUrl('http://curator.example.com')
expect(normalizeApiUrl(once)).toBe(once)
})
it('strips trailing slashes before deciding', () => {
expect(normalizeApiUrl('http://curator.example.com/')).toBe('http://curator.example.com/api')
expect(normalizeApiUrl('http://curator.example.com///')).toBe('http://curator.example.com/api')
expect(normalizeApiUrl('http://curator.example.com/api/')).toBe('http://curator.example.com/api')
})
it('trims surrounding whitespace (paste artifacts)', () => {
expect(normalizeApiUrl(' http://curator.example.com ')).toBe(
'http://curator.example.com/api'
)
})
it('matches the /api suffix case-insensitively', () => {
expect(normalizeApiUrl('http://curator.example.com/API')).toBe('http://curator.example.com/API')
})
it('returns empty string for empty/nullish input, never a bare "/api"', () => {
// isConfigured() gates on truthiness, so a bogus '/api' here would read as
// "configured" and produce a request against the options page's own origin.
expect(normalizeApiUrl('')).toBe('')
expect(normalizeApiUrl(' ')).toBe('')
expect(normalizeApiUrl(null)).toBe('')
expect(normalizeApiUrl(undefined)).toBe('')
})
it('does not treat a path merely containing "api" as the suffix', () => {
expect(normalizeApiUrl('http://curator.example.com/apiary')).toBe(
'http://curator.example.com/apiary/api'
)
})
it('preserves a subpath deployment', () => {
expect(normalizeApiUrl('http://host.internal/curator')).toBe('http://host.internal/curator/api')
})
})
describe('webRootFromApiUrl', () => {
// The SPA root, where the Vue router and the served XPI live. Used by
// OPEN_ARTIST_PAGE and the self-update check — NOT the JSON API.
it('strips the /api suffix', () => {
expect(webRootFromApiUrl('http://curator.example.com/api')).toBe('http://curator.example.com')
})
it('accepts an instance root unchanged', () => {
expect(webRootFromApiUrl('http://curator.example.com')).toBe('http://curator.example.com')
})
it('agrees with normalizeApiUrl in both directions', () => {
for (const input of ['http://curator.example.com', 'http://curator.example.com/api']) {
expect(normalizeApiUrl(webRootFromApiUrl(input))).toBe(normalizeApiUrl(input))
}
})
it('preserves a subpath deployment', () => {
expect(webRootFromApiUrl('http://host.internal/curator/api')).toBe('http://host.internal/curator')
})
it('returns empty string for empty/nullish input', () => {
expect(webRootFromApiUrl('')).toBe('')
expect(webRootFromApiUrl(null)).toBe('')
})
})
+112
View File
@@ -0,0 +1,112 @@
import { describe, it, expect } from 'vitest'
import { readFileSync } from 'node:fs'
import { execFileSync } from 'node:child_process'
import { fileURLToPath } from 'node:url'
import path from 'node:path'
const EXT_DIR = path.join(path.dirname(fileURLToPath(import.meta.url)), '..')
const read = (name) => JSON.parse(readFileSync(path.join(EXT_DIR, name), 'utf8'))
const readText = (...seg) => readFileSync(path.join(EXT_DIR, ...seg), 'utf8')
// Only the git-free subcommands are exercised here: `version`/`patch` shell out
// to git, and the extension lane runs on node:24-bookworm-slim which may not
// ship it. Those two are covered where git is guaranteed — ci.yml and build.yml
// run on ci-python.
const packaging = (cmd) =>
execFileSync('sh', [path.join(EXT_DIR, 'scripts', 'packaging.sh'), cmd], {
cwd: EXT_DIR,
encoding: 'utf8'
})
.trim()
.split(/\s+/)
.filter(Boolean)
describe('packaging.sh — the single definition of what ships', () => {
it('emits an ignore list and a pathspec that agree on the tracked files', () => {
const ignore = packaging('ignore')
const pathspec = packaging('pathspec').map((e) => e.replace(':(exclude)extension/', ''))
// Every git-excluded path must also be hidden from web-ext. The reverse is
// not required: node_modules and web-ext-artifacts are build output git
// never tracks, so they appear only in the ignore list.
for (const entry of pathspec) {
expect(ignore, `pathspec has "${entry}" but --ignore-files does not`).toContain(entry)
}
expect(pathspec.length).toBeGreaterThan(0)
expect(ignore).toContain('node_modules')
})
it('emits glob patterns literally, never expanded against the working tree', () => {
// The script iterates its lists with deliberate word-splitting, so it must
// run with pathname expansion off. Without that, invoking it from a cwd
// where test/ exists (exactly how ci.yml and vitest call it) expands
// `test/**` into the individual spec files, and the pathspec silently stops
// covering anything added later.
const pathspec = packaging('pathspec')
expect(pathspec).toContain(':(exclude)extension/test/**')
expect(pathspec).toContain(':(exclude)extension/scripts/**')
expect(pathspec.some((e) => e.includes('.spec.js'))).toBe(false)
expect(pathspec.some((e) => e.includes('helpers'))).toBe(false)
})
it('keeps its own scripts and specs out of the XPI', () => {
// Both are repo infrastructure. web-ext packages everything not ignored, so
// omitting either would ship dev tooling to users -- and `test/**` in
// particular only survives because callers `set -f` before substituting it.
const ignore = packaging('ignore')
expect(ignore).toContain('vitest.config.js')
// Both forms per directory. The glob covers the contents; the bare name
// covers the directory ENTRY, which web-ext writes separately — with only
// the glob, the XPI carries an empty `test/` and `scripts/`.
for (const dir of ['test', 'scripts']) {
expect(ignore, `${dir} contents`).toContain(`${dir}/**`)
expect(ignore, `${dir} directory entry`).toContain(dir)
}
})
})
describe('consumers delegate rather than keeping their own copy', () => {
// These assertions are the actual anti-regression value: it is easy for a
// future edit to "simplify" by inlining a literal list again, which silently
// reintroduces the drift that issue #2397 was about.
it('package.json derives --ignore-files from the script', () => {
for (const [name, script] of Object.entries(read('package.json').scripts)) {
if (!script.includes('--ignore-files')) continue
expect(script, `${name} should call packaging.sh`).toContain('scripts/packaging.sh ignore')
expect(script, `${name} must set -f before the substitution`).toMatch(/set -f;/)
}
})
it('ci.yml derives its pathspec from the script and hardcodes none', () => {
const ci = readText('..', '.forgejo', 'workflows', 'ci.yml')
expect(ci).toContain('extension/scripts/packaging.sh pathspec')
// A literal :(exclude)extension/... in the workflow means someone bypassed
// the shared definition.
expect(ci).not.toMatch(/:\(exclude\)extension\//)
})
})
describe('extension version', () => {
it('keeps manifest.json and package.json in lockstep', () => {
expect(read('manifest.json').version).toBe(read('package.json').version)
})
it('uses a plain dotted numeric version AMO will accept', () => {
expect(read('package.json').version).toMatch(/^\d+(\.\d+)*$/)
})
it('declares manifest v3', () => {
expect(read('manifest.json').manifest_version).toBe(3)
})
it('lists every background script that exists, in dependency order', () => {
// url.js must load BEFORE api.js: api.js calls normalizeApiUrl at
// init()-time, and these are classic scripts sharing one scope, so a
// reordering here is a runtime ReferenceError with no build-time signal.
const scripts = read('manifest.json').background.scripts
for (const rel of scripts) {
expect(() => readFileSync(path.join(EXT_DIR, rel)), `missing ${rel}`).not.toThrow()
}
expect(scripts.indexOf('lib/url.js')).toBeLessThan(scripts.indexOf('lib/api.js'))
})
})
+13
View File
@@ -0,0 +1,13 @@
import { defineConfig } from 'vitest/config'
// Mirrors frontend/vitest.config.js, minus the Vue plugin — the extension has
// no SFCs and mounts nothing. Pure-logic specs only, so `node` is enough; the
// libs under test are deliberately the ones with no browser-API surface (see
// test/helpers/loadLib.js).
export default defineConfig({
test: {
environment: 'node',
include: ['test/**/*.spec.js'],
passWithNoTests: true
}
})
@@ -50,14 +50,19 @@ const projected = ref(null)
const projectedCounts = computed(() => projected.value?.projected || null)
const modalDescription = computed(
() => projected.value
? `Artist “${props.artistName}” — `
+ `${projected.value.projected.images} images, `
+ `${projected.value.projected.sources} sources, `
+ `${Math.round(projected.value.projected.bytes_on_disk / 1_048_576)} MiB on disk`
: '',
)
// `posts` is named here, not left to the counts grid below it: an artist whose
// posts are body-only previews as `images: 0`, and a summary line that says
// only "0 images" reads as "this artist is empty" while the apply destroys
// every captured post body (#3067). Attachments stay in the grid — the grid
// renders every key, so this line carries only what changes the read.
const modalDescription = computed(() => {
const p = projectedCounts.value
return p
? `Artist “${props.artistName}” — ${p.images} images, `
+ `${p.posts} posts, ${p.sources} sources, `
+ `${Math.round(p.bytes_on_disk / 1_048_576)} MiB on disk`
: ''
})
async function onClick() {
loading.value = true
@@ -0,0 +1,56 @@
<!--
Canonical settings number field (DRY pass #161): a compact numeric v-text-field
with a built-in clamp to [min,max] on commit. Hand-rolled identically across the
ML settings cards (HeadsCard x6, CropProposersCard, VideoEmbeddingCard).
The clamp is the point: the cards previously sent Number(raw) straight to the
API, so an out-of-range value bounced off the API's 400 validator (only
TranslationCard clamped). This is now the single home for that clamp.
Binds `modelValue` (v-model) and emits `change` on blur/enter AFTER clamping, so
the parent's save reads the already-clamped value same as the prior
`v-model.number` + `@change=save` pattern.
-->
<template>
<v-text-field
:model-value="modelValue"
:label="label"
type="number"
:min="min"
:max="max"
:step="step"
:disabled="disabled"
:density="density" hide-details
:style="{ maxWidth }"
@update:model-value="v => emit('update:modelValue', v)"
@change="onCommit"
/>
</template>
<script setup>
const props = defineProps({
modelValue: { type: [Number, String], default: null },
label: { type: String, default: '' },
min: { type: [Number, String], default: null },
max: { type: [Number, String], default: null },
step: { type: [Number, String], default: 1 },
maxWidth: { type: String, default: '200px' },
density: { type: String, default: 'compact' },
disabled: { type: Boolean, default: false },
})
const emit = defineEmits(['update:modelValue', 'change'])
function onCommit() {
// On blur/enter: coerce to a number and clamp to [min,max] so an out-of-range
// value never reaches the API. props.modelValue reflects the latest keystroke
// (kept in sync by the passthrough above); re-emit the clamped number, then let
// the parent persist.
let n = Number(props.modelValue)
if (!Number.isNaN(n)) {
if (props.min !== null && props.min !== '') n = Math.max(Number(props.min), n)
if (props.max !== null && props.max !== '') n = Math.min(Number(props.max), n)
if (n !== Number(props.modelValue)) emit('update:modelValue', n)
}
emit('change')
}
</script>
@@ -0,0 +1,42 @@
<!--
Canonical settings toggle row (DRY pass #161): an accent icon + an uppercase
.fc-section-h label + a right-aligned switch. Hand-rolled identically in the
ML settings cards (HeadsCard x3, CropProposersCard, MLBackfillCard).
Two-way binds `modelValue` (so the parent switch state stays optimistic) AND
emits `change` with the new boolean, so the parent can persist + revert on
failure matching the prior `v-model` + `@update:model-value=handler` pattern.
-->
<template>
<div class="d-flex align-center mb-1" style="gap: 10px;">
<v-icon v-if="icon" size="18" :color="iconColor">{{ icon }}</v-icon>
<span class="fc-section-h">{{ label }}</span>
<v-switch
:model-value="modelValue"
:loading="loading"
:disabled="disabled"
hide-details density="compact" color="success" class="ml-auto"
@update:model-value="onSwitch"
/>
</div>
</template>
<script setup>
defineProps({
modelValue: { type: Boolean, default: false },
label: { type: String, default: '' },
icon: { type: String, default: '' },
// Icon tint. Default accent; pass null for the theme default (e.g. when a row
// is off). null (not undefined) so the default doesn't override it.
iconColor: { type: String, default: 'accent' },
loading: { type: Boolean, default: false },
disabled: { type: Boolean, default: false },
})
const emit = defineEmits(['update:modelValue', 'change'])
function onSwitch(v) {
const b = !!v
emit('update:modelValue', b)
emit('change', b)
}
</script>
@@ -139,9 +139,9 @@ function onThumbError() { thumbError.value = true }
position: absolute; top: 8px; left: 8px;
width: 22px; height: 22px; border-radius: 4px;
border: 2px solid rgba(232, 228, 216, 0.8);
background: rgba(20, 23, 26, 0.45);
background: rgba(var(--v-theme-background), 0.45);
display: grid; place-items: center;
color: #14171A; z-index: 11;
color: rgb(var(--v-theme-background)); z-index: 11;
}
.fc-gallery-item__checkbox.on {
background: rgb(var(--v-theme-accent));
@@ -152,7 +152,7 @@ function onThumbError() { thumbError.value = true }
min-width: 22px; height: 22px; padding: 0 5px;
border-radius: 11px;
background: rgb(var(--v-theme-accent));
color: #14171A; font-size: 12px; font-weight: 700;
color: rgb(var(--v-theme-background)); font-size: 12px; font-weight: 700;
display: grid; place-items: center; z-index: 11;
pointer-events: none;
}
@@ -160,7 +160,8 @@ function onThumbError() { thumbError.value = true }
position: absolute; left: 0; right: 0; bottom: 0;
padding: 14px 8px 6px;
background: linear-gradient(
to top, rgba(20, 23, 26, 0.78), rgba(20, 23, 26, 0)
to top, rgba(var(--v-theme-background), 0.78),
rgba(var(--v-theme-background), 0)
);
font-size: 12px; line-height: 1.2;
white-space: nowrap; overflow: hidden; text-overflow: ellipsis;
@@ -0,0 +1,127 @@
<template>
<!-- #3068: attachment reclamation. PostAttachment's FKs are both SET NULL, so
a deleted post or artist leaves the row behind; and the store is
sha-addressed, so one blob backs many rows and deleting a row never freed
its file. Nothing swept either. Preview first, then apply (destructive:
unlinks files). -->
<MaintenanceTile
icon="mdi-paperclip-off"
title="Reclaim orphaned attachments"
blurb="Remove attachment records belonging to nothing, and the files nothing references."
destructive
:open="applying || previewing"
>
<p class="text-body-2 mb-3">
Attachment records survive the post and artist they belonged to, and the
files behind them are shared between records — so a deleted record never
freed its file on its own. This finds records attributed to
<strong>neither</strong> a post nor an artist, and files in the attachment
store that <strong>no remaining record</strong> points at.
<strong>Preview</strong> first; <strong>Apply</strong> deletes those
records and unlinks those files. Files written in the last few hours are
always left alone, so an in-progress download is never caught mid-write.
</p>
<div class="d-flex align-center flex-wrap" style="gap: 12px;">
<v-btn
color="primary" variant="tonal" rounded="pill"
:loading="previewing" :disabled="applying" @click="preview"
>
<v-icon start>mdi-magnify</v-icon> Preview
</v-btn>
<v-btn
color="error" rounded="pill"
:loading="applying"
:disabled="previewing || !canApply"
@click="confirmOpen = true"
>
<v-icon start>mdi-paperclip-off</v-icon> Apply
</v-btn>
</div>
<v-alert
v-if="summary" :type="summaryType" variant="tonal" class="mt-4"
density="comfortable"
>
<span v-if="applied">
Deleted {{ summary.rows }} orphaned record(s) and unlinked
{{ summary.files }} file(s), reclaiming {{ humanBytes(summary.bytes) }}.
</span>
<span v-else-if="hasWork">
{{ summary.rows }} orphaned record(s) and {{ summary.files }}
unreferenced file(s) — {{ humanBytes(summary.bytes) }} reclaimable.
Click <strong>Apply</strong> to remove them.
</span>
<span v-else>Nothing to reclaim — every attachment is accounted for.</span>
<!-- Both of these change what the numbers MEAN, so they are stated
whenever they are non-zero rather than hidden in a tooltip. -->
<div v-if="summary.files_failed" class="mt-1 text-caption">
{{ summary.files_failed }} file(s) could not be read or removed — see
the worker log.
</div>
<div v-if="summary.partial" class="mt-1 text-caption">
Stopped early at the time limit; some of the store was not examined.
Run it again to continue.
</div>
</v-alert>
<QueueStatusBar queue="maintenance_long" queue-label="Maintenance" />
<v-dialog v-model="confirmOpen" max-width="440">
<v-card>
<v-card-title>Reclaim orphaned attachments?</v-card-title>
<v-card-text class="text-body-2">
This permanently deletes
<strong>{{ summary?.rows ?? 0 }}</strong> attachment record(s) and
unlinks <strong>{{ summary?.files ?? 0 }}</strong> file(s)
({{ humanBytes(summary?.bytes) }}). Only files that no remaining
record points at are removed, so nothing still attached to a post
is affected.
</v-card-text>
<v-card-actions>
<v-spacer />
<v-btn variant="text" @click="confirmOpen = false">Cancel</v-btn>
<v-btn color="error" @click="apply">Reclaim</v-btn>
</v-card-actions>
</v-card>
</v-dialog>
</MaintenanceTile>
</template>
<script setup>
import { computed, ref } from 'vue'
import { useMaintenanceTask } from '../../composables/useMaintenanceTask.js'
import { humanBytes } from '../../utils/bytes.js'
import MaintenanceTile from '../common/MaintenanceTile.vue'
import QueueStatusBar from './QueueStatusBar.vue'
const confirmOpen = ref(false)
// Walks the whole attachment store, so it can run for minutes on a large
// library — the service caps itself at 900s and reports `partial`. 150 polls
// × 2s ≈ 5m of foreground waiting; past that the composable hands off to the
// task dashboard rather than spinning forever.
const { previewing, applying, summary, applied, preview, apply: applyTask } = useMaintenanceTask({
endpoint: '/api/admin/maintenance/reclaim-attachments',
storageKey: 'fc.maint.reclaimAttachments',
appliedToast: 'Orphaned attachments reclaimed',
maxPolls: 150,
})
const hasWork = computed(
() => !!summary.value && (summary.value.rows > 0 || summary.value.files > 0),
)
const canApply = computed(() => hasWork.value && !applied.value)
const summaryType = computed(() => {
if (applied.value) return 'success'
return hasWork.value ? 'info' : 'success'
})
// The confirm dialog gates the destructive apply; close it, then run.
function apply () {
confirmOpen.value = false
applyTask()
}
</script>
@@ -9,7 +9,7 @@
<v-card-text>
<p class="fc-muted text-body-2">
Pushes session cookies from supported platforms
(patreon, subscribestar, hentaifoundry, discord, pixiv, deviantart)
(patreon, subscribestar, hentaifoundry, discord, pixiv)
into FabledCurator, and lets you add a creator as a source from
their page in one click.
</p>
@@ -15,28 +15,23 @@
</p>
<div v-for="p in proposers" :key="p.key" class="fc-proposer">
<div class="d-flex align-center mb-1" style="gap: 10px;">
<v-icon size="18" :color="p.on ? 'accent' : undefined">{{ p.icon }}</v-icon>
<span class="fc-section-h">{{ p.label }}</span>
<v-switch
v-model="p.on" :loading="busy" hide-details density="compact"
color="success" class="ml-auto"
@update:model-value="v => saveToggle(p, v)"
/>
</div>
<SettingToggleRow
v-model="p.on" :loading="busy" :icon="p.icon"
:icon-color="p.on ? 'accent' : null" :label="p.label"
@change="v => saveToggle(p, v)"
/>
<p class="fc-muted text-body-2 mb-2">{{ p.help }}</p>
<div class="d-flex flex-wrap mb-4" style="gap: 12px;">
<v-text-field
v-model="p.weights" label="Weights" density="compact" hide-details
style="min-width: 300px; flex: 1;" :disabled="busy || !p.on"
placeholder="name | URL | hf_repo::file"
@change="save({ [`detector_${p.key}_weights`]: p.weights })"
@change="saveField({ [`detector_${p.key}_weights`]: p.weights })"
/>
<v-text-field
v-model.number="p.conf" label="Confidence" type="number"
min="0" max="1" step="0.05" density="compact" hide-details
style="max-width: 140px;" :disabled="busy || !p.on"
@change="save({ [`detector_${p.key}_conf`]: Number(p.conf) })"
<SettingNumberField
v-model="p.conf" label="Confidence" :min="0" :max="1" :step="0.05"
max-width="140px" :disabled="busy || !p.on"
@change="saveField({ [`detector_${p.key}_conf`]: Number(p.conf) })"
/>
</div>
</div>
@@ -48,12 +43,12 @@
storage. Dedupe IoU drops near-duplicate crops before embedding.
</p>
<div class="d-flex flex-wrap" style="gap: 12px;">
<v-text-field
<SettingNumberField
v-for="c in caps" :key="c.key"
v-model.number="c.val" :label="c.label" type="number"
:min="c.min" :max="c.max" :step="c.step || 1" density="compact"
hide-details style="max-width: 165px;" :disabled="busy"
@change="save({ [c.key]: Number(c.val) })"
v-model="c.val" :label="c.label"
:min="c.min" :max="c.max" :step="c.step || 1"
max-width="165px" :disabled="busy"
@change="saveField({ [c.key]: Number(c.val) })"
/>
</div>
</div>
@@ -61,14 +56,16 @@
</template>
<script setup>
import { toast } from '../../utils/toast.js'
import { onMounted, ref } from 'vue'
import MaintenanceTile from '../common/MaintenanceTile.vue'
import SettingNumberField from '../common/SettingNumberField.vue'
import SettingToggleRow from '../common/SettingToggleRow.vue'
import { useSettingSave } from '../../composables/useSettingSave.js'
import { useMLStore } from '../../stores/ml.js'
const mlSettings = useMLStore()
const busy = ref(false)
const { busy, save } = useSettingSave(mlSettings.patchSettings)
const proposers = ref([])
const caps = ref([])
@@ -111,31 +108,20 @@ onMounted(async () => {
caps.value = CAP_DEFS.map(c => ({ ...c, val: s[c.key] ?? 0 }))
})
async function save(patch, revert) {
busy.value = true
try {
await mlSettings.patchSettings(patch)
toast({ text: 'Saved', type: 'success' })
} catch (e) {
if (revert) revert()
toast({ text: `Could not save: ${e.message}`, type: 'error' })
} finally {
busy.value = false
}
// Field @change → persist with a "Saved" confirmation. SettingNumberField has
// already clamped numeric values to their [min,max] before this fires.
function saveField(patch) {
save(patch, { successMessage: 'Saved' })
}
function saveToggle (p, v) {
async function saveToggle(p, v) {
// Revert the switch on failure so it never lies about the persisted state.
save({ [`detector_${p.key}_enabled`]: !!v }, () => { p.on = !v })
const ok = await save({ [`detector_${p.key}_enabled`]: !!v }, { successMessage: 'Saved' })
if (!ok) p.on = !v
}
</script>
<style scoped>
.fc-muted { color: rgb(var(--v-theme-on-surface-variant)); }
.fc-section-h {
font-size: 13px; font-weight: 700; letter-spacing: 0.03em;
text-transform: uppercase; color: rgb(var(--v-theme-on-surface));
}
.fc-proposer {
border-top: 1px solid rgb(var(--v-theme-surface-light)); padding-top: 14px;
}
@@ -112,7 +112,6 @@ async function onCommit() {
</script>
<style scoped>
.fc-muted { color: rgb(var(--v-theme-on-surface-variant)); }
.fc-code {
background: rgb(var(--v-theme-surface-light));
border-radius: 4px; padding: 2px 8px;
@@ -42,7 +42,7 @@
</tr>
</tbody>
</v-table>
<p v-else class="text-caption mt-3" style="opacity: 0.6;">
<p v-else class="text-caption mt-3 fc-muted">
No table statistics yet.
</p>
</MaintenanceTile>
@@ -35,7 +35,7 @@
All subscription sources healthy.
</p>
<p v-else class="text-body-2 mb-0">
<b class="fc-bad">{{ failing.length }}</b> failing source(s):
<b class="fc-weak">{{ failing.length }}</b> failing source(s):
<span class="fc-muted">{{ failingNames }}</span>
</p>
</v-card-text>
@@ -72,6 +72,4 @@ onUnmounted(() => { if (pollId) clearInterval(pollId) })
</script>
<style scoped>
.fc-muted { color: rgb(var(--v-theme-on-surface-variant)); }
.fc-bad { color: rgb(var(--v-theme-error)); }
</style>
@@ -102,6 +102,7 @@
import { computed, ref } from 'vue'
import { useMaintenanceTask } from '../../composables/useMaintenanceTask.js'
import { humanBytes } from '../../utils/bytes.js'
import MaintenanceTile from '../common/MaintenanceTile.vue'
import QueueStatusBar from './QueueStatusBar.vue'
@@ -122,14 +123,6 @@ const summaryType = computed(() => {
return summary.value && summary.value.matched > 0 ? 'info' : 'success'
})
function humanBytes (n) {
const b = Number(n || 0)
if (b >= 1 << 30) return (b / (1 << 30)).toFixed(1) + ' GB'
if (b >= 1 << 20) return (b / (1 << 20)).toFixed(1) + ' MB'
if (b >= 1 << 10) return (b / (1 << 10)).toFixed(1) + ' KB'
return b + ' B'
}
// The confirm dialog gates the destructive apply; close it, then run.
function apply () {
confirmOpen.value = false
@@ -25,7 +25,7 @@
<div class="fc-cell__l">done</div>
</div>
<div class="fc-cell">
<div class="fc-cell__n" :class="q.error ? 'fc-bad' : ''">{{ q.error }}</div>
<div class="fc-cell__n" :class="q.error ? 'fc-weak' : ''">{{ q.error }}</div>
<div class="fc-cell__l">errored</div>
</div>
</div>
@@ -95,7 +95,6 @@ onUnmounted(() => { if (pollId) clearInterval(pollId) })
</script>
<style scoped>
.fc-muted { color: rgb(var(--v-theme-on-surface-variant)); }
.fc-cells { display: flex; gap: 28px; }
.fc-cell__n {
font-size: 20px; font-weight: 700; line-height: 1.1;
@@ -105,6 +104,4 @@ onUnmounted(() => { if (pollId) clearInterval(pollId) })
font-size: 11px; text-transform: uppercase; letter-spacing: 0.04em;
color: rgb(var(--v-theme-on-surface-variant));
}
.fc-good { color: rgb(var(--v-theme-success)); }
.fc-bad { color: rgb(var(--v-theme-error)); }
</style>
@@ -367,11 +367,6 @@ async function onReprocess() {
</script>
<style scoped>
.fc-muted { color: rgb(var(--v-theme-on-surface-variant)); }
.fc-section-h {
font-size: 13px; font-weight: 700; letter-spacing: 0.03em;
text-transform: uppercase; color: rgb(var(--v-theme-on-surface));
}
.fc-token {
display: flex; align-items: center; gap: 4px;
background: rgb(var(--v-theme-surface-light)); border-radius: 6px;
@@ -390,6 +385,4 @@ async function onReprocess() {
font-size: 11px; text-transform: uppercase; letter-spacing: 0.04em;
color: rgb(var(--v-theme-on-surface-variant));
}
.fc-good { color: rgb(var(--v-theme-success)); }
.fc-weak { color: rgb(var(--v-theme-error)); }
</style>
@@ -155,11 +155,6 @@ async function onRecover(it) {
</script>
<style scoped>
.fc-muted { color: rgb(var(--v-theme-on-surface-variant)); }
.fc-section-h {
font-size: 13px; font-weight: 700; letter-spacing: 0.03em;
text-transform: uppercase; color: rgb(var(--v-theme-on-surface));
}
.fc-queue { display: flex; gap: 24px; }
.fc-q__n {
font-size: 20px; font-weight: 700; line-height: 1.1;
@@ -169,8 +164,6 @@ async function onRecover(it) {
font-size: 11px; text-transform: uppercase; letter-spacing: 0.04em;
color: rgb(var(--v-theme-on-surface-variant));
}
.fc-good { color: rgb(var(--v-theme-success)); }
.fc-weak { color: rgb(var(--v-theme-error)); }
.fc-defect {
display: flex; align-items: center; gap: 12px;
background: rgb(var(--v-theme-surface-light)); border-radius: 8px;
+60 -125
View File
@@ -95,14 +95,10 @@
<!-- Earned auto-apply -->
<div class="fc-auto mt-6">
<div class="d-flex align-center mb-1" style="gap: 10px;">
<v-icon size="18" color="accent">mdi-lightning-bolt</v-icon>
<span class="fc-section-h">Auto-apply</span>
<v-switch
v-model="autoEnabled" :loading="settingBusy" hide-details density="compact"
color="success" class="ml-auto" @update:model-value="onToggleAuto"
/>
</div>
<SettingToggleRow
v-model="autoEnabled" :loading="settingBusy"
icon="mdi-lightning-bolt" label="Auto-apply" @change="onToggleAuto"
/>
<p class="fc-muted text-body-2 mb-3">
Graduated heads (, with {{ autoMinPosInput }} examples) apply their tag
on their own where they clear {{ Math.round((autoPrecisionInput || 0) * 100) }}%
@@ -111,17 +107,14 @@
</p>
<div class="d-flex mb-3" style="gap: 12px;">
<v-text-field
v-model.number="autoPrecisionInput" label="Precision target"
type="number" min="0.5" max="0.999" step="0.01" density="compact"
hide-details style="max-width: 200px;" :disabled="settingBusy"
<SettingNumberField
v-model="autoPrecisionInput" label="Precision target"
:min="0.5" :max="0.999" :step="0.01" :disabled="settingBusy"
@change="onSaveSettings"
/>
<v-text-field
v-model.number="autoMinPosInput" label="Min examples to fire"
type="number" min="1" density="compact" hide-details
style="max-width: 200px;" :disabled="settingBusy"
@change="onSaveSettings"
<SettingNumberField
v-model="autoMinPosInput" label="Min examples to fire"
:min="1" :disabled="settingBusy" @change="onSaveSettings"
/>
</div>
@@ -161,15 +154,11 @@
<!-- Presentation chrome auto-hide (#141) -->
<div class="fc-auto mt-6">
<div class="d-flex align-center mb-1" style="gap: 10px;">
<v-icon size="18" color="accent">mdi-image-off-outline</v-icon>
<span class="fc-section-h">Hide presentation chrome</span>
<v-switch
v-model="presentationEnabled" :loading="settingBusy" hide-details
density="compact" color="success" class="ml-auto"
@update:model-value="onTogglePresentation"
/>
</div>
<SettingToggleRow
v-model="presentationEnabled" :loading="settingBusy"
icon="mdi-image-off-outline" label="Hide presentation chrome"
@change="onTogglePresentation"
/>
<p class="fc-muted text-body-2 mb-3">
Auto-hide <code>banner</code> chrome from the gallery once a head has
learned it ( {{ minPositives }} examples) and clears
@@ -180,16 +169,14 @@
tag), it's flagged for review instead of buried. Every auto-hide is reversible.
</p>
<div class="d-flex mb-3" style="gap: 12px;">
<v-text-field
v-model.number="presentationThresholdInput" label="Hide confidence"
type="number" min="0.5" max="0.999" step="0.01" density="compact"
hide-details style="max-width: 200px;" :disabled="settingBusy"
<SettingNumberField
v-model="presentationThresholdInput" label="Hide confidence"
:min="0.5" :max="0.999" :step="0.01" :disabled="settingBusy"
@change="onSavePresentation"
/>
<v-text-field
v-model.number="presentationConflictInput" label="Flag if content ≥"
type="number" min="0" max="1" step="0.05" density="compact"
hide-details style="max-width: 200px;" :disabled="settingBusy"
<SettingNumberField
v-model="presentationConflictInput" label="Flag if content ≥"
:min="0" :max="1" :step="0.05" :disabled="settingBusy"
@change="onSavePresentation"
/>
</div>
@@ -197,15 +184,11 @@
<!-- Process auto-tagging (#1464): wip / editor screenshot -->
<div class="fc-auto mt-6">
<div class="d-flex align-center mb-1" style="gap: 10px;">
<v-icon size="18" color="accent">mdi-progress-wrench</v-icon>
<span class="fc-section-h">Auto-tag work-in-progress</span>
<v-switch
v-model="processEnabled" :loading="settingBusy" hide-details
density="compact" color="success" class="ml-auto"
@update:model-value="onToggleProcess"
/>
</div>
<SettingToggleRow
v-model="processEnabled" :loading="settingBusy"
icon="mdi-progress-wrench" label="Auto-tag work-in-progress"
@change="onToggleProcess"
/>
<p class="fc-muted text-body-2 mb-3">
Auto-tag <code>wip</code> and <code>editor screenshot</code> process art
once a head has learned them (≥ {{ minPositives }} examples) and clears
@@ -217,16 +200,14 @@
manual tags, never its own guesses so it can't run away. Every tag reversible.
</p>
<div class="d-flex mb-3" style="gap: 12px;">
<v-text-field
v-model.number="processThresholdInput" label="Tag confidence"
type="number" min="0.5" max="0.999" step="0.01" density="compact"
hide-details style="max-width: 200px;" :disabled="settingBusy"
<SettingNumberField
v-model="processThresholdInput" label="Tag confidence"
:min="0.5" :max="0.999" :step="0.01" :disabled="settingBusy"
@change="onSaveProcess"
/>
<v-text-field
v-model.number="processConflictInput" label="Flag if content ≥"
type="number" min="0" max="1" step="0.05" density="compact"
hide-details style="max-width: 200px;" :disabled="settingBusy"
<SettingNumberField
v-model="processConflictInput" label="Flag if content ≥"
:min="0" :max="1" :step="0.05" :disabled="settingBusy"
@change="onSaveProcess"
/>
</div>
@@ -256,7 +237,7 @@
<td class="fc-r fc-mono">{{ c.n_auto_applied }}</td>
<td class="fc-r fc-mono">{{ c.n_misfires }}</td>
<td class="fc-r fc-mono" :class="rateClass(c.misfire_rate)">
{{ ratePct(c.misfire_rate) }}
{{ pct(c.misfire_rate) }}
</td>
<td class="fc-r fc-mono">{{ c.n_underfires }}</td>
</tr>
@@ -272,6 +253,9 @@ import { toast } from '../../utils/toast.js'
import { computed, onMounted, onUnmounted, ref } from 'vue'
import MaintenanceTile from '../common/MaintenanceTile.vue'
import SettingNumberField from '../common/SettingNumberField.vue'
import SettingToggleRow from '../common/SettingToggleRow.vue'
import { useSettingSave } from '../../composables/useSettingSave.js'
import { useHeadsStore } from '../../stores/heads.js'
import { useMLStore } from '../../stores/ml.js'
@@ -285,7 +269,9 @@ let pollTimer = null
const autoEnabled = ref(false)
const autoPrecisionInput = ref(0.97)
const autoMinPosInput = ref(30)
const settingBusy = ref(false)
// Shared settings-save flow (busy + toast + revert); `settingBusy` gates the
// toggles/fields, `save` returns ok/false for the optimistic-switch revert.
const { busy: settingBusy, save } = useSettingSave(mlSettings.patchSettings)
const autoBusy = ref(false)
const autoStatus = ref(null)
const metricsData = ref(null)
@@ -395,81 +381,39 @@ function startAutoPoll() {
function stopAutoPoll() { if (autoTimer) { clearInterval(autoTimer); autoTimer = null } }
async function onToggleAuto(val) {
settingBusy.value = true
try {
await mlSettings.patchSettings({ head_auto_apply_enabled: !!val })
toast({ text: val ? 'Auto-apply on' : 'Auto-apply off', type: 'success' })
} catch (e) {
autoEnabled.value = !val // revert the switch
toast({ text: `Could not update: ${e.message}`, type: 'error' })
} finally {
settingBusy.value = false
}
const ok = await save({ head_auto_apply_enabled: !!val },
{ successMessage: val ? 'Auto-apply on' : 'Auto-apply off', errorPrefix: 'Could not update' })
if (!ok) autoEnabled.value = !val // revert the switch
}
async function onSaveSettings() {
settingBusy.value = true
try {
await mlSettings.patchSettings({
head_auto_apply_precision: Number(autoPrecisionInput.value),
head_auto_apply_min_positives: Number(autoMinPosInput.value),
})
} catch (e) {
toast({ text: `Could not save: ${e.message}`, type: 'error' })
} finally {
settingBusy.value = false
}
await save({
head_auto_apply_precision: Number(autoPrecisionInput.value),
head_auto_apply_min_positives: Number(autoMinPosInput.value),
})
}
async function onTogglePresentation(val) {
settingBusy.value = true
try {
await mlSettings.patchSettings({ presentation_auto_apply_enabled: !!val })
toast({ text: val ? 'Chrome auto-hide on' : 'Chrome auto-hide off', type: 'success' })
} catch (e) {
presentationEnabled.value = !val // revert the switch
toast({ text: `Could not update: ${e.message}`, type: 'error' })
} finally {
settingBusy.value = false
}
const ok = await save({ presentation_auto_apply_enabled: !!val },
{ successMessage: val ? 'Chrome auto-hide on' : 'Chrome auto-hide off', errorPrefix: 'Could not update' })
if (!ok) presentationEnabled.value = !val // revert the switch
}
async function onSavePresentation() {
settingBusy.value = true
try {
await mlSettings.patchSettings({
presentation_auto_apply_threshold: Number(presentationThresholdInput.value),
presentation_conflict_threshold: Number(presentationConflictInput.value),
})
} catch (e) {
toast({ text: `Could not save: ${e.message}`, type: 'error' })
} finally {
settingBusy.value = false
}
await save({
presentation_auto_apply_threshold: Number(presentationThresholdInput.value),
presentation_conflict_threshold: Number(presentationConflictInput.value),
})
}
async function onToggleProcess(val) {
settingBusy.value = true
try {
await mlSettings.patchSettings({ process_auto_apply_enabled: !!val })
toast({ text: val ? 'WIP auto-tag on' : 'WIP auto-tag off', type: 'success' })
} catch (e) {
processEnabled.value = !val // revert the switch
toast({ text: `Could not update: ${e.message}`, type: 'error' })
} finally {
settingBusy.value = false
}
const ok = await save({ process_auto_apply_enabled: !!val },
{ successMessage: val ? 'WIP auto-tag on' : 'WIP auto-tag off', errorPrefix: 'Could not update' })
if (!ok) processEnabled.value = !val // revert the switch
}
async function onSaveProcess() {
settingBusy.value = true
try {
await mlSettings.patchSettings({
process_auto_apply_threshold: Number(processThresholdInput.value),
process_conflict_threshold: Number(processConflictInput.value),
})
} catch (e) {
toast({ text: `Could not save: ${e.message}`, type: 'error' })
} finally {
settingBusy.value = false
}
await save({
process_auto_apply_threshold: Number(processThresholdInput.value),
process_conflict_threshold: Number(processConflictInput.value),
})
}
function onPreview() { startSweep(true) }
function onApplyNow() { startSweep(false) }
@@ -495,7 +439,6 @@ function sweepConcepts(run) {
.sort((a, b) => b.applied - a.applied)
}
function sweepTotal(run) { return run?.n_applied ?? 0 }
function ratePct(x) { return x == null ? '—' : `${Math.round(x * 100)}%` }
function rateClass(x) {
if (x == null) return ''
if (x <= 0.03) return 'fc-good'
@@ -526,12 +469,6 @@ function relTime(iso) {
</script>
<style scoped>
.fc-muted { color: rgb(var(--v-theme-on-surface-variant)); }
.fc-section-h {
font-size: 13px; font-weight: 700; letter-spacing: 0.03em;
text-transform: uppercase; color: rgb(var(--v-theme-on-surface));
}
.fc-auto {
border-top: 1px solid rgb(var(--v-theme-surface-light)); padding-top: 16px;
}
@@ -588,7 +525,5 @@ function relTime(iso) {
background: rgb(var(--v-theme-surface-light));
padding: 1px 6px; border-radius: 999px;
}
.fc-good { color: rgb(var(--v-theme-success)); }
.fc-ok { color: rgb(var(--v-theme-on-surface)); }
.fc-weak { color: rgb(var(--v-theme-error)); }
</style>
@@ -39,13 +39,14 @@
import { toast } from '../../utils/toast.js'
import { onMounted, ref } from 'vue'
import { useMLStore } from '../../stores/ml.js'
import { useSettingSave } from '../../composables/useSettingSave.js'
import MaintenanceTile from '../common/MaintenanceTile.vue'
import QueueStatusBar from './QueueStatusBar.vue'
const store = useMLStore()
const { busy: saving, save } = useSettingSave(store.patchSettings)
const busy = ref(false)
const done = ref(false)
const enabled = ref(true)
const saving = ref(false)
onMounted(async () => {
try {
await store.loadSettings()
@@ -55,21 +56,12 @@ onMounted(async () => {
} catch { /* non-fatal */ }
})
async function onToggle() {
saving.value = true
try {
await store.patchSettings({ cpu_embed_enabled: enabled.value })
toast({
text: enabled.value
? 'CPU embedding on — imports queue embeds for the ml-worker'
: 'CPU embedding off — the GPU embed backfill owns whole-image embeds',
type: 'success',
})
} catch (e) {
toast({ text: `Could not save: ${e.message}`, type: 'error' })
enabled.value = !enabled.value
} finally {
saving.value = false
}
const ok = await save({ cpu_embed_enabled: enabled.value }, {
successMessage: enabled.value
? 'CPU embedding on — imports queue embeds for the ml-worker'
: 'CPU embedding off — the GPU embed backfill owns whole-image embeds',
})
if (!ok) enabled.value = !enabled.value
}
async function run() {
busy.value = true
@@ -80,5 +72,4 @@ async function run() {
</script>
<style scoped>
.fc-muted { color: rgb(var(--v-theme-on-surface-variant)); }
</style>
@@ -24,6 +24,7 @@
the CPU fallback.
</p>
<div class="fc-tile-stack">
<VideoEmbeddingCard />
<GpuAgentCard />
<GpuTriageCard />
<MLBackfillCard />
@@ -36,7 +37,6 @@
Suggestion thresholds, trained heads and tag aliases.
</p>
<div class="fc-tile-stack">
<MLThresholdSliders />
<CropProposersCard />
<HeadsCard />
<AliasTable />
@@ -77,7 +77,7 @@ import ArchiveReextractCard from './ArchiveReextractCard.vue'
import MissingFileRepairCard from './MissingFileRepairCard.vue'
import GpuTriageCard from './GpuTriageCard.vue'
import DbMaintenanceCard from './DbMaintenanceCard.vue'
import MLThresholdSliders from './MLThresholdSliders.vue'
import VideoEmbeddingCard from './VideoEmbeddingCard.vue'
import CropProposersCard from './CropProposersCard.vue'
import HeadsCard from './HeadsCard.vue'
import GpuAgentCard from './GpuAgentCard.vue'
@@ -78,6 +78,7 @@
import { computed, ref } from 'vue'
import { useMaintenanceTask } from '../../composables/useMaintenanceTask.js'
import { humanBytes } from '../../utils/bytes.js'
import MaintenanceTile from '../common/MaintenanceTile.vue'
import QueueStatusBar from './QueueStatusBar.vue'
@@ -98,14 +99,6 @@ const summaryType = computed(() => {
return summary.value && summary.value.redundant > 0 ? 'info' : 'success'
})
function humanBytes (n) {
const b = Number(n || 0)
if (b >= 1 << 30) return (b / (1 << 30)).toFixed(1) + ' GB'
if (b >= 1 << 20) return (b / (1 << 20)).toFixed(1) + ' MB'
if (b >= 1 << 10) return (b / (1 << 10)).toFixed(1) + ' KB'
return b + ' B'
}
// The confirm dialog gates the destructive apply; close it, then run.
function apply () {
confirmOpen.value = false
@@ -12,17 +12,17 @@
</div>
<v-row>
<v-col cols="12" sm="6">
<v-text-field
v-model.number="local.video_frame_interval_seconds"
label="Frame interval (s)" type="number" min="0.5" step="0.5"
density="comfortable" hide-details @change="save"
<SettingNumberField
v-model="local.video_frame_interval_seconds"
label="Frame interval (s)" :min="0.5" :step="0.5"
density="comfortable" max-width="none" @change="onSave"
/>
</v-col>
<v-col cols="12" sm="6">
<v-text-field
v-model.number="local.video_max_frames"
label="Max frames" type="number" min="1" step="1"
density="comfortable" hide-details @change="save"
<SettingNumberField
v-model="local.video_max_frames"
label="Max frames" :min="1" :step="1"
density="comfortable" max-width="none" @change="onSave"
/>
</v-col>
</v-row>
@@ -32,21 +32,23 @@
</template>
<script setup>
import { toast } from '../../utils/toast.js'
import { reactive, watch } from 'vue'
import { useMLStore } from '../../stores/ml.js'
import MaintenanceTile from '../common/MaintenanceTile.vue'
import SettingNumberField from '../common/SettingNumberField.vue'
import { useSettingSave } from '../../composables/useSettingSave.js'
const store = useMLStore()
const { save } = useSettingSave(store.patchSettings)
const local = reactive({})
watch(() => store.settings, (s) => { if (s) Object.assign(local, s) }, { immediate: true })
async function save() {
const patch = {
video_frame_interval_seconds: local.video_frame_interval_seconds,
video_max_frames: local.video_max_frames
}
try { await store.patchSettings(patch) }
catch (e) { toast({ text: e.message, type: 'error' }) }
// SettingNumberField clamps interval to 0.5 and max-frames to 1 before this
// fires, so an out-of-range value never reaches the API.
function onSave() {
save({
video_frame_interval_seconds: Number(local.video_frame_interval_seconds),
video_max_frames: Number(local.video_max_frames),
})
}
</script>
@@ -0,0 +1,33 @@
import { ref } from 'vue'
import { toast } from '../utils/toast.js'
// The shared "persist a settings patch" flow for the ML settings cards. Flips a
// busy flag, calls the store's patch (which rethrows on failure), toasts
// success/error, and returns true/false so a toggle handler can revert its
// optimistic switch on failure. Centralises the try/catch/toast the cards each
// hand-rolled (HeadsCard x6, CropProposersCard, MLBackfillCard) — and where the
// threshold-clamp drifted; the clamp now lives in <SettingNumberField>.
//
// Pass the store's patch fn, e.g. useSettingSave(ml.patchSettings).
export function useSettingSave(patchFn) {
const busy = ref(false)
// opts.successMessage — toast on success (toggles announce their new state;
// silent field-saves omit it). opts.errorPrefix — the failure toast prefix
// ("Could not save" default; toggles used "Could not update").
async function save(patch, { successMessage = '', errorPrefix = 'Could not save' } = {}) {
busy.value = true
try {
await patchFn(patch)
if (successMessage) toast({ text: successMessage, type: 'success' })
return true
} catch (e) {
toast({ text: `${errorPrefix}: ${e.message}`, type: 'error' })
return false
} finally {
busy.value = false
}
}
return { busy, save }
}
+19
View File
@@ -40,6 +40,25 @@
emits, so no specificity/reorder fight — no !important needed. */
.fc-muted { color: rgb(var(--v-theme-on-surface-variant)); }
/* Section sub-heading in settings cards (DRY pass #161): was redefined
identically in 4 cards, and TranslationCard used the class with NO local def
so its section headers rendered unstyled. Now one global utility. */
.fc-section-h {
font-size: 13px; font-weight: 700; letter-spacing: 0.03em;
text-transform: uppercase; color: rgb(var(--v-theme-on-surface));
}
/* Status text colours (DRY pass #161): fc-good = success, fc-weak = error,
consolidated from the GPU / heads cards. fc-ok is intentionally NOT global —
it means on-surface in HeadsCard but success in QueuesTable.
No `.fc-bad` (#3072): it was defined locally and identically in the Downloads
and GPU activity panels, and it is fc-weak under a second name — GpuAgentCard
and GpuActivityPanel were colouring the same "errored" count with different
class names. Both now use fc-weak. Reach for fc-weak, not a new synonym. */
.fc-good { color: rgb(var(--v-theme-success)); }
.fc-weak { color: rgb(var(--v-theme-error)); }
/* Vuetify 4 dropped its global CSS reset (normalisation moved into each
component). FC's layouts assumed the reset zeroed margins on text elements, so
restore just that — the "minimal reset" from the v4 upgrade guide — inside
+17
View File
@@ -0,0 +1,17 @@
// Human-readable byte sizes for maintenance summaries ("2.4 GB reclaimable").
//
// Promoted out of the cleanup cards, which had grown byte-identical private
// copies (VideoDedupCard, GatedPurgeCard) and were about to grow a third for
// the attachment reclaim. Binary units (1 KB = 1024 B) — these numbers come
// from st_size / SUM(size_bytes), so they describe disk, not marketing.
//
// NOT the same shape as the `formatBytes` helpers in SystemStatsCards,
// BackupRunsTable and PostCard — those differ in units, precision and
// zero-handling. Left alone deliberately rather than force-fitted here.
export function humanBytes (n) {
const b = Number(n || 0)
if (b >= 1 << 30) return (b / (1 << 30)).toFixed(1) + ' GB'
if (b >= 1 << 20) return (b / (1 << 20)).toFixed(1) + ' MB'
if (b >= 1 << 10) return (b / (1 << 10)).toFixed(1) + ' KB'
return b + ' B'
}
+5 -6
View File
@@ -1,8 +1,10 @@
// Single source of truth for platform → color + icon mapping. Used by
// PlatformChip and any other GS-style platform-tagged surface. The six
// PlatformChip and any other GS-style platform-tagged surface. The five
// platforms FC supports map 1:1 to the GS palette; unknown platforms fall
// back to grey + mdi-web. Operator-confirmed scope 2026-05-27. The ICONS key
// set is pinned against backend known_platform_keys() by
// back to grey + mdi-web — which is deliberately what a retired platform
// hits: a pre-#3069 deviantart source row still renders, as its raw key on
// a grey chip. Operator-confirmed scope 2026-05-27. The ICONS key set is
// pinned against backend known_platform_keys() by
// tests/test_fe_be_contract.py.
const ICONS = {
@@ -11,7 +13,6 @@ const ICONS = {
hentaifoundry: 'mdi-palette',
discord: 'mdi-discord',
pixiv: 'mdi-alpha-p-box',
deviantart: 'mdi-deviantart',
}
const COLORS = {
@@ -20,7 +21,6 @@ const COLORS = {
hentaifoundry: 'purple',
discord: 'indigo',
pixiv: 'blue',
deviantart: 'green',
}
const LABELS = {
@@ -29,7 +29,6 @@ const LABELS = {
hentaifoundry: 'HentaiFoundry',
discord: 'Discord',
pixiv: 'Pixiv',
deviantart: 'DeviantArt',
}
export function platformIcon(platform) {
+5 -2
View File
@@ -19,14 +19,16 @@
</section>
<section class="fc-section">
<h3 class="fc-section__title">Duplicates &amp; posts</h3>
<h3 class="fc-section__title">Duplicates &amp; leftovers</h3>
<p class="fc-section__hint">
Tidy post records, duplicates and locked-preview leftovers.
Tidy post records, duplicates, locked-preview leftovers and attachments
that outlived what they belonged to.
</p>
<div class="fc-tile-grid">
<PostMaintenanceCard />
<VideoDedupCard />
<GatedPurgeCard />
<AttachmentReclaimCard />
</div>
</section>
@@ -60,6 +62,7 @@ import SingleColorAuditCard from '../components/cleanup/SingleColorAuditCard.vue
import PostMaintenanceCard from '../components/settings/PostMaintenanceCard.vue'
import VideoDedupCard from '../components/settings/VideoDedupCard.vue'
import GatedPurgeCard from '../components/settings/GatedPurgeCard.vue'
import AttachmentReclaimCard from '../components/settings/AttachmentReclaimCard.vue'
import TagMaintenanceCard from '../components/settings/TagMaintenanceCard.vue'
import DangerZoneCard from '../components/settings/DangerZoneCard.vue'
</script>
-1
View File
@@ -284,7 +284,6 @@ onUnmounted(() => {
</script>
<style scoped>
.fc-muted { color: rgb(var(--v-theme-on-surface-variant)); }
/* Full-height workspace under the sticky top nav. --fc-nav-h is the nav's REAL
measured height (set by TopNav) — a hardcoded 64px here overflowed the
-1
View File
@@ -349,7 +349,6 @@ async function onDeleteTagConfirm() {
.fc-tags__sentinel {
display: flex; justify-content: center; padding: 32px 0; min-height: 60px;
}
.fc-muted { color: rgb(var(--v-theme-on-surface-variant)); }
.fc-merge-preview {
padding: 10px 12px;
border: 1px solid rgba(var(--v-theme-on-surface), 0.12);
+31
View File
@@ -677,3 +677,34 @@ async def test_reset_content_tagging_apply_requires_confirm_token(client, db):
)
assert resp.status_code == 200
assert (await resp.get_json())["deleted"] == 1
@pytest.mark.asyncio
async def test_trigger_reclaim_attachments_defaults_to_preview(client, monkeypatch):
"""Unlike the other maintenance triggers, this one's apply unlinks FILES —
so an empty body must mean preview, not apply."""
from backend.app.tasks import admin as admin_tasks
calls = []
monkeypatch.setattr(
admin_tasks.reclaim_orphaned_attachments_task, "delay", _fake_delay(calls)
)
resp = await client.post("/api/admin/maintenance/reclaim-attachments", json={})
assert resp.status_code == 202
assert (await resp.get_json())["task_id"] == "task-xyz"
assert calls[0][1] == {"dry_run": True}
@pytest.mark.asyncio
async def test_trigger_reclaim_attachments_threads_apply(client, monkeypatch):
from backend.app.tasks import admin as admin_tasks
calls = []
monkeypatch.setattr(
admin_tasks.reclaim_orphaned_attachments_task, "delay", _fake_delay(calls)
)
resp = await client.post(
"/api/admin/maintenance/reclaim-attachments", json={"dry_run": False},
)
assert resp.status_code == 202
assert calls[0][1] == {"dry_run": False}
+14 -1
View File
@@ -1,6 +1,6 @@
import pytest
from backend.app.models import Artist, PostAttachment
from backend.app.models import Artist, PostAttachment, attachment_download_url
pytestmark = pytest.mark.integration
@@ -33,3 +33,16 @@ async def test_download_streams_with_disposition(client, db, tmp_path):
async def test_download_404(client):
resp = await client.get("/api/attachments/999999/download")
assert resp.status_code == 404
@pytest.mark.asyncio
async def test_attachment_download_url_routes_to_the_download_endpoint(app):
"""The two serializers no longer hand-format this path (#3072) — but a
single definition is only worth having if it still matches the route. Pin
it by MATCHING against the real URL map rather than comparing to a literal:
a string equality test would pass just as happily after someone renamed the
route, which is the exact drift the helper exists to prevent."""
built = attachment_download_url(4242)
endpoint, args = app.url_map.bind("localhost").match(built)
assert endpoint == "attachments.download"
assert args == {"attachment_id": 4242}
+17 -1
View File
@@ -129,7 +129,6 @@ async def test_resolve_artist_name_dispatches_per_platform(db, monkeypatch):
("https://www.subscribestar.com/foobar", "subscribestar", "foobar"),
("https://subscribestar.adult/foobar", "subscribestar", "foobar"),
("https://www.hentai-foundry.com/user/Foo/profile", "hentaifoundry", "Foo"),
("https://www.deviantart.com/baz", "deviantart", "baz"),
("https://www.pixiv.net/users/12345", "pixiv", "12345"),
("https://www.pixiv.net/en/users/12345", "pixiv", "12345"),
])
@@ -160,6 +159,23 @@ async def test_quick_add_source_unknown_url_400(client, ext_key):
assert "known" in body
@pytest.mark.asyncio
async def test_quick_add_source_rejects_retired_deviantart(client, ext_key):
"""#3069: a DeviantArt creator URL used to derive cleanly. Now that the
platform is retired, the extension's own gate should never offer the
button — but a stale content script on an un-updated browser still can,
so the backend has to refuse it rather than create an unusable source."""
resp = await client.post(
"/api/extension/quick-add-source",
json={"url": "https://www.deviantart.com/baz"},
headers={"X-Extension-Key": ext_key},
)
assert resp.status_code == 400
body = await resp.get_json()
assert body["error"] == "unknown_platform"
assert "deviantart" not in body["known"]
@pytest.mark.asyncio
async def test_quick_add_source_invalid_url_400(client, ext_key):
resp = await client.post(
+3 -2
View File
@@ -6,16 +6,17 @@ pytestmark = pytest.mark.integration
@pytest.mark.asyncio
async def test_platforms_returns_gs_six(client):
async def test_platforms_returns_gs_five(client):
resp = await client.get("/api/platforms")
assert resp.status_code == 200
body = await resp.get_json()
platforms = body["platforms"]
assert set(platforms.keys()) == {
"patreon", "subscribestar", "hentaifoundry",
"discord", "pixiv", "deviantart",
"discord", "pixiv",
}
assert "fanbox" not in platforms
assert "deviantart" not in platforms # retired at #3069
@pytest.mark.asyncio
+1 -1
View File
@@ -154,7 +154,7 @@ async def test_list_platform_filter_excludes_no_source(db):
@pytest.mark.asyncio
async def test_list_platform_filter_excludes_wrong_platform(db):
a = await _seed_artist(db, "alice-wplat")
await _seed_source(db, a.id, "deviantart", "https://d/alice-wp")
await _seed_source(db, a.id, "discord", "https://d/alice-wp")
await db.commit()
page = await ArtistDirectoryService(db).list_artists(
+339
View File
@@ -5,6 +5,7 @@ side effects use tmp_path. Assertions on mutated rows use COLUMN
SELECTS per reference_async_coredml_test_assertions — never
re-read ORM attributes after a service mutates and re-fetches.
"""
import os
from datetime import UTC, datetime
import pytest
@@ -68,6 +69,8 @@ def test_project_artist_cascade_returns_zeroes_for_empty_artist(db_sync):
assert result["artist"]["slug"] == "empty"
assert result["projected"] == {
"images": 0,
"posts": 0,
"attachments": 0,
"sources": 0,
"thumbs": 0,
"import_tasks": 0,
@@ -92,6 +95,87 @@ def test_project_artist_cascade_counts_images_and_thumbs_and_bytes(db_sync, tmp_
assert result["projected"]["bytes_on_disk"] == 3500
def test_project_artist_cascade_counts_posts_and_attachments(db_sync, tmp_path):
"""The body-only artist: zero images, but posts and attachments that the
apply destroys. Previewing this as `images: 0` alone is what made a
content-only artist read as an empty one (#3067)."""
a = _make_artist(db_sync, slug="bodyonly")
p1 = Post(artist_id=a.id, external_post_id="bo-1", description="a body")
p2 = Post(artist_id=a.id, external_post_id="bo-2", description="another")
db_sync.add_all([p1, p2])
db_sync.flush()
db_sync.add(PostAttachment(
post_id=p1.id, artist_id=a.id, sha256="b0d1".ljust(64, "0"),
path="/store/b0d1/a.pdf", original_filename="a.pdf",
ext=".pdf", size_bytes=5,
))
# artist_id NULL, reachable only through its post — the second arm of
# _artist_attachments_conditions.
db_sync.add(PostAttachment(
post_id=p2.id, artist_id=None, sha256="b0d2".ljust(64, "0"),
path="/store/b0d2/b.pdf", original_filename="b.pdf",
ext=".pdf", size_bytes=5,
))
db_sync.commit()
projected = cleanup_service.project_artist_cascade(
db_sync, slug="bodyonly",
)["projected"]
assert projected["images"] == 0
assert projected["posts"] == 2
assert projected["attachments"] == 2
def test_artist_cascade_preview_matches_apply(db_sync, tmp_path):
"""Rule 93: the preview's numbers must be what the apply actually does.
Guards the drift directly rather than trusting that both halves happen to
use the same predicate — the preview and apply are asserted against each
other on one artist carrying all three row kinds.
"""
a = _make_artist(db_sync, slug="parity")
for i in range(3):
f = tmp_path / f"par{i}.jpg"
f.write_bytes(b"x")
_make_image(
db_sync, artist=a, path=str(f), sha256=f"{i:064x}", size=10,
)
posts = [
Post(artist_id=a.id, external_post_id=f"par-{i}") for i in range(4)
]
db_sync.add_all(posts)
db_sync.flush()
for i, p in enumerate(posts[:2]):
db_sync.add(PostAttachment(
post_id=p.id, artist_id=a.id, sha256=f"par{i}".ljust(64, "0"),
path=f"/store/par{i}/f.zip", original_filename="f.zip",
ext=".zip", size_bytes=9,
))
db_sync.commit()
artist_id = a.id
projected = cleanup_service.project_artist_cascade(
db_sync, slug="parity",
)["projected"]
summary = cleanup_service.delete_artist_cascade(
db_sync, artist_id=artist_id, images_root=tmp_path,
)["summary"]
assert projected["images"] == summary["images_deleted"] == 3
assert projected["posts"] == summary["posts_deleted"] == 4
assert projected["attachments"] == summary["attachments_deleted"] == 2
# And the apply really did remove them — a matching pair of numbers is
# worth nothing if neither half touched the DB.
assert db_sync.execute(
select(func.count(Post.id)).where(Post.artist_id == artist_id)
).scalar_one() == 0
assert db_sync.execute(
select(func.count(PostAttachment.id))
.where(PostAttachment.artist_id == artist_id)
).scalar_one() == 0
def test_project_artist_cascade_raises_on_unknown_slug(db_sync):
with pytest.raises(LookupError):
cleanup_service.project_artist_cascade(db_sync, slug="nope")
@@ -294,6 +378,88 @@ def test_delete_artist_cascade_idempotent_on_missing(db_sync, tmp_path):
assert result["summary"]["images_deleted"] == 0
def test_delete_artist_cascade_survives_same_sha_on_two_posts(db_sync, tmp_path):
"""Same file attached to two of the artist's posts must not abort the delete.
Left to the cascade this raises: artist delete CASCADEs to Post, which SET
NULLs post_attachment.post_id, and `uq_post_attachment_null_post_sha`
(sha256 alone, WHERE post_id IS NULL) then rejects the second row. That's an
ordinary shape — _capture_attachment writes one row per post over one
sha-addressed blob by design. Also covers the NULL-artist_id arm of the
delete predicate: the second row has no artist_id, only a post that does.
"""
a = _make_artist(db_sync, slug="casatt")
p1 = Post(artist_id=a.id, external_post_id="att-p1")
p2 = Post(artist_id=a.id, external_post_id="att-p2")
db_sync.add_all([p1, p2])
db_sync.flush()
shared_sha = "ca5a".ljust(64, "0")
db_sync.add(PostAttachment(
post_id=p1.id, artist_id=a.id, sha256=shared_sha,
path="/store/ca5a/bundle.zip", original_filename="bundle.zip",
ext=".zip", size_bytes=7,
))
db_sync.add(PostAttachment(
post_id=p2.id, artist_id=None, sha256=shared_sha,
path="/store/ca5a/bundle.zip", original_filename="bundle.zip",
ext=".zip", size_bytes=7,
))
db_sync.commit()
artist_id = a.id
result = cleanup_service.delete_artist_cascade(
db_sync, artist_id=artist_id, images_root=tmp_path,
)
assert result["summary"]["attachments_deleted"] == 2
assert db_sync.execute(
select(func.count(Artist.id)).where(Artist.id == artist_id)
).scalar_one() == 0
assert db_sync.execute(
select(func.count(PostAttachment.id))
.where(PostAttachment.sha256 == shared_sha)
).scalar_one() == 0
def test_delete_artist_cascade_keeps_unrelated_null_post_attachment(
db_sync, tmp_path,
):
"""A filesystem-import row (post_id NULL) sharing the sha is the other way
this collides — and it must SURVIVE: it belongs to no artist, so the
cascade has no claim on it."""
a = _make_artist(db_sync, slug="casorph")
p = Post(artist_id=a.id, external_post_id="orph-p1")
db_sync.add(p)
db_sync.flush()
sha = "0rfa".ljust(64, "0")
standalone = PostAttachment(
post_id=None, artist_id=None, sha256=sha,
path="/store/0rfa/manual.pdf", original_filename="manual.pdf",
ext=".pdf", size_bytes=3,
)
db_sync.add(standalone)
db_sync.add(PostAttachment(
post_id=p.id, artist_id=a.id, sha256=sha,
path="/store/0rfa/manual.pdf", original_filename="manual.pdf",
ext=".pdf", size_bytes=3,
))
db_sync.commit()
artist_id, standalone_id = a.id, standalone.id
result = cleanup_service.delete_artist_cascade(
db_sync, artist_id=artist_id, images_root=tmp_path,
)
assert result["summary"]["attachments_deleted"] == 1
surviving = db_sync.execute(
select(PostAttachment.id, PostAttachment.post_id)
.where(PostAttachment.sha256 == sha)
).all()
assert surviving == [(standalone_id, None)]
# --- delete_images --------------------------------------------------
@@ -916,3 +1082,176 @@ def test_reconcile_preserves_from_attachment_on_provenance_collision(db_sync, tm
.where(ImageProvenance.image_record_id == img_id)
).all()
assert rows == [(native_id, att_id)]
# --- reclaim_orphaned_attachments -----------------------------------
def _store_blob(root, sha, *, ext=".pdf", age_hours=48, data=b"blob"):
"""Write a file into the sha-addressed attachment store, aged past the
min-age guard by default."""
d = root / "attachments" / sha[:3]
d.mkdir(parents=True, exist_ok=True)
p = d / f"{sha}{ext}"
p.write_bytes(data)
old = datetime.now(UTC).timestamp() - age_hours * 3600
os.utime(p, (old, old))
return p
def _attachment(db_sync, *, sha, post=None, artist=None):
att = PostAttachment(
post_id=post.id if post else None,
artist_id=artist.id if artist else None,
sha256=sha, path=f"/store/{sha[:3]}/f.pdf",
original_filename="f.pdf", ext=".pdf", size_bytes=4,
)
db_sync.add(att)
db_sync.flush()
return att
def test_reclaim_attachments_dry_run_projects_without_mutating(db_sync, tmp_path):
a = _make_artist(db_sync, slug="recl-dry")
p = Post(artist_id=a.id, external_post_id="rd-1")
db_sync.add(p)
db_sync.flush()
kept_sha, orphan_sha = "aa11".ljust(64, "0"), "bb22".ljust(64, "0")
_attachment(db_sync, sha=kept_sha, post=p, artist=a)
_attachment(db_sync, sha=orphan_sha) # both FKs NULL → orphan
db_sync.commit()
kept_blob = _store_blob(tmp_path, kept_sha)
orphan_blob = _store_blob(tmp_path, orphan_sha)
result = cleanup_service.reclaim_orphaned_attachments(
db_sync, images_root=tmp_path, dry_run=True,
)
assert result["rows"] == 1
assert result["files"] == 1
assert result["bytes"] == orphan_blob.stat().st_size
# Nothing actually happened.
assert kept_blob.exists() and orphan_blob.exists()
assert db_sync.execute(
select(func.count(PostAttachment.id))
).scalar_one() == 2
def test_reclaim_attachments_apply_deletes_rows_and_unlinks_blobs(db_sync, tmp_path):
a = _make_artist(db_sync, slug="recl-apply")
p = Post(artist_id=a.id, external_post_id="ra-1")
db_sync.add(p)
db_sync.flush()
kept_sha, orphan_sha = "cc33".ljust(64, "0"), "dd44".ljust(64, "0")
_attachment(db_sync, sha=kept_sha, post=p, artist=a)
_attachment(db_sync, sha=orphan_sha)
db_sync.commit()
kept_blob = _store_blob(tmp_path, kept_sha)
orphan_blob = _store_blob(tmp_path, orphan_sha)
result = cleanup_service.reclaim_orphaned_attachments(
db_sync, images_root=tmp_path, dry_run=False,
)
assert result["rows"] == 1
assert result["files"] == 1
assert kept_blob.exists() # still referenced
assert not orphan_blob.exists() # nothing points at it any more
surviving = db_sync.execute(select(PostAttachment.sha256)).scalars().all()
assert surviving == [kept_sha]
def test_reclaim_attachments_preview_matches_apply(db_sync, tmp_path):
"""Rule 93 — the dry-run's numbers are what the apply does. The projection
has to negate the orphan predicate to be honest about blobs the delete is
about to free, so this is the assertion that catches getting that backwards.
"""
orphan_sha = "ee55".ljust(64, "0")
_attachment(db_sync, sha=orphan_sha)
db_sync.commit()
_store_blob(tmp_path, orphan_sha)
projected = cleanup_service.reclaim_orphaned_attachments(
db_sync, images_root=tmp_path, dry_run=True,
)
applied = cleanup_service.reclaim_orphaned_attachments(
db_sync, images_root=tmp_path, dry_run=False,
)
for key in ("rows", "files", "bytes"):
assert projected[key] == applied[key], key
assert applied["rows"] == 1 and applied["files"] == 1
def test_reclaim_attachments_keeps_shared_blob_while_any_row_remains(db_sync, tmp_path):
"""The refcount case this whole sweep exists for: one sha-addressed blob
backs several rows, so deleting SOME of them must not free the file."""
a = _make_artist(db_sync, slug="recl-shared")
p = Post(artist_id=a.id, external_post_id="rs-1")
db_sync.add(p)
db_sync.flush()
sha = "ff66".ljust(64, "0")
_attachment(db_sync, sha=sha, post=p, artist=a) # attributed — survives
_attachment(db_sync, sha=sha) # orphan — deleted
db_sync.commit()
blob = _store_blob(tmp_path, sha)
result = cleanup_service.reclaim_orphaned_attachments(
db_sync, images_root=tmp_path, dry_run=False,
)
assert result["rows"] == 1 # the orphan row went
assert result["files"] == 0 # the blob did NOT
assert blob.exists()
def test_reclaim_attachments_spares_filesystem_import_rows(db_sync, tmp_path):
"""post_id NULL with an artist_id is the deliberate filesystem-import shape
(importer._capture_attachment), not an orphan — it is still attributed."""
a = _make_artist(db_sync, slug="recl-fsimport")
sha = "1177".ljust(64, "0")
_attachment(db_sync, sha=sha, artist=a) # post NULL, artist set
db_sync.commit()
blob = _store_blob(tmp_path, sha)
result = cleanup_service.reclaim_orphaned_attachments(
db_sync, images_root=tmp_path, dry_run=False,
)
assert result["rows"] == 0
assert result["files"] == 0
assert blob.exists()
assert db_sync.execute(
select(func.count(PostAttachment.id))
).scalar_one() == 1
def test_reclaim_attachments_skips_recent_and_staging_files(db_sync, tmp_path):
"""A blob is written BEFORE its row commits, so a just-stored file with no
row is in-flight, not orphaned. `.partial` staging files belong to
cleanup_orphaned_temp_files and must be left alone either way."""
fresh_sha, staged_sha = "2288".ljust(64, "0"), "3399".ljust(64, "0")
fresh = _store_blob(tmp_path, fresh_sha, age_hours=0)
staged = _store_blob(tmp_path, staged_sha, ext=".pdf.partial")
db_sync.commit()
result = cleanup_service.reclaim_orphaned_attachments(
db_sync, images_root=tmp_path, dry_run=False,
)
assert result["files"] == 0
assert result["skipped_recent"] == 1
assert fresh.exists() and staged.exists()
def test_reclaim_attachments_ignores_non_sha_named_files(db_sync, tmp_path):
"""The walk must only judge files it can identify as store blobs — anything
else under the root is none of its business."""
d = tmp_path / "attachments" / "zzz"
d.mkdir(parents=True)
stray = d / "notes.txt"
stray.write_text("not a blob")
old = datetime.now(UTC).timestamp() - 48 * 3600
os.utime(stray, (old, old))
result = cleanup_service.reclaim_orphaned_attachments(
db_sync, images_root=tmp_path, dry_run=False,
)
assert result["files"] == 0
assert stray.exists()
+1 -1
View File
@@ -19,7 +19,7 @@ def test_native_platforms():
def test_gallery_dl_platforms_are_not_native():
# The platforms still served by gallery-dl must NOT route to the native
# ingester — guards an accidental over-broad migration.
for platform in ("hentaifoundry", "discord", "deviantart"):
for platform in ("hentaifoundry", "discord"):
assert uses_native_ingester(platform) is False
+119
View File
@@ -0,0 +1,119 @@
"""`insert_image_tags` — the shared bulk write behind the WIP-title backfill and
both auto-apply sweeps (#3072).
The sweeps previously issued one INSERT per applied tag from inside their
per-image loop; they now hand this helper a chunk's worth of rows. That makes
this function the single place three writers can be wrong at once, so it is
tested directly rather than only through its callers.
"""
import pytest
from sqlalchemy import select
from backend.app.models import ImageRecord, Tag, TagKind
from backend.app.models.tag import image_tag
from backend.app.services.image_tag_apply import insert_image_tags
pytestmark = pytest.mark.integration
_N = 0
def _img(db_sync):
global _N
_N += 1
rec = ImageRecord(
path=f"/images/ita/{_N}.jpg", sha256=f"a{_N:063d}",
size_bytes=1, mime="image/jpeg", width=1, height=1,
origin="imported_filesystem", integrity_status="unknown",
)
db_sync.add(rec)
db_sync.flush()
return rec
def _tag(db_sync, name):
t = Tag(name=name, kind=TagKind.general)
db_sync.add(t)
db_sync.flush()
return t
def _rows(db_sync, tag_id):
"""(image_record_id, source) pairs currently carrying `tag_id`."""
return dict(db_sync.execute(
select(image_tag.c.image_record_id, image_tag.c.source)
.where(image_tag.c.tag_id == tag_id)
).all())
def _row(image_record_id, tag_id, source):
return {
"image_record_id": image_record_id, "tag_id": tag_id, "source": source,
}
def test_inserts_every_row_in_one_call(db_sync):
t = _tag(db_sync, "ita-basic")
imgs = [_img(db_sync) for _ in range(3)]
insert_image_tags(
db_sync, [_row(i.id, t.id, "head_auto") for i in imgs]
)
assert _rows(db_sync, t.id) == {i.id: "head_auto" for i in imgs}
def test_spans_several_tags_in_a_single_call(db_sync):
"""The sweeps accumulate across ALL heads before flushing, so one call
carries rows for different tags. A per-tag implementation would drop all
but the first."""
t1, t2 = _tag(db_sync, "ita-multi-1"), _tag(db_sync, "ita-multi-2")
a, b = _img(db_sync), _img(db_sync)
insert_image_tags(db_sync, [
_row(a.id, t1.id, "head_auto"), _row(b.id, t1.id, "head_auto"),
_row(a.id, t2.id, "head_auto"),
])
assert _rows(db_sync, t1.id) == {a.id: "head_auto", b.id: "head_auto"}
assert _rows(db_sync, t2.id) == {a.id: "head_auto"}
def test_an_existing_tag_keeps_its_original_source(db_sync):
"""THE assertion this helper exists for. A sweep re-running over an image
the operator tagged by hand must not restamp it as machine-applied — that
would silently poison the head's own training data, which excludes the
auto sources. ON CONFLICT DO NOTHING, never DO UPDATE."""
t = _tag(db_sync, "ita-manual")
rec = _img(db_sync)
insert_image_tags(db_sync, [_row(rec.id, t.id, "manual")])
insert_image_tags(db_sync, [_row(rec.id, t.id, "head_auto")])
assert _rows(db_sync, t.id) == {rec.id: "manual"}
def test_a_repeat_within_one_call_does_not_raise(db_sync):
"""Two heads can both fire on the same (image, tag) inside one chunk. The
conflict is resolved by the statement, not by the caller de-duplicating."""
t = _tag(db_sync, "ita-dupe")
rec = _img(db_sync)
insert_image_tags(db_sync, [
_row(rec.id, t.id, "head_auto"), _row(rec.id, t.id, "head_auto"),
])
assert _rows(db_sync, t.id) == {rec.id: "head_auto"}
def test_more_rows_than_the_chunk_size_all_land(db_sync):
"""The chunk exists to stay under Postgres' 65535 bound-parameter ceiling.
Driven with a tiny chunk so the split is real rather than theoretical — at
the 5000 default no test would ever reach a second statement."""
t = _tag(db_sync, "ita-chunked")
imgs = [_img(db_sync) for _ in range(7)]
insert_image_tags(
db_sync, [_row(i.id, t.id, "head_auto") for i in imgs], chunk=2
)
assert _rows(db_sync, t.id) == {i.id: "head_auto" for i in imgs}
def test_no_rows_is_a_no_op(db_sync):
"""A dry-run chunk, or a chunk where every candidate was already skipped,
hands over an empty list. `.values([])` is a SQL error, so the empty case
must never reach the statement."""
insert_image_tags(db_sync, [])
+14
View File
@@ -778,3 +778,17 @@ def test_vacuum_analyze_runs_over_high_churn_tables():
result = vacuum_analyze.apply().get()
assert result["vacuumed"] == list(VACUUM_TABLES)
def test_reclaim_attachments_stuck_threshold_exceeds_hard_time_limit():
"""#883's invariant, applied to the attachment reclaim: a task whose stall
threshold is under its own hard limit gets phantom-flagged 'RecoverySweep'
while it is still healthily running."""
from backend.app.tasks.admin import reclaim_orphaned_attachments_task
from backend.app.tasks.maintenance import TASK_STUCK_THRESHOLD_MINUTES
hard_minutes = reclaim_orphaned_attachments_task.time_limit / 60
override = TASK_STUCK_THRESHOLD_MINUTES[
"backend.app.tasks.admin.reclaim_orphaned_attachments_task"
]
assert override >= hard_minutes
+59
View File
@@ -0,0 +1,59 @@
"""Shared ML helpers extracted in the DRY pass (milestone #161). These pin the
single sources the auto-apply sweeps now trust, so a future edit can't silently
drift them: `_applied_or_rejected` is the skip-set used by auto_apply_sweep,
system_tag_auto_apply_sweep (heads.py) and scheduled_ccip_auto_apply (tasks/ml.py);
`_sigmoid` is the head score→prob transform used at every scoring site."""
import pytest
from backend.app.models import ImageRecord, Tag, TagKind, TagSuggestionRejection
from backend.app.models.tag import image_tag
from backend.app.services.ml.training_data import _applied_or_rejected
def test_sigmoid_matches_naive_form():
import numpy as np
from backend.app.services.ml.heads import _sigmoid
z = np.array([-3.0, -0.5, 0.0, 1.5, 12.0], dtype=np.float32)
assert np.allclose(_sigmoid(z, np), 1.0 / (1.0 + np.exp(-z)))
assert float(_sigmoid(np.array([0.0]), np)[0]) == pytest.approx(0.5)
@pytest.mark.integration
def test_applied_or_rejected_unions_applied_any_source_and_rejected(db_sync):
a = Tag(name="dry-helper-a", kind=TagKind.general)
b = Tag(name="dry-helper-b", kind=TagKind.general)
db_sync.add_all([a, b])
db_sync.flush()
imgs = []
for i in range(5):
img = ImageRecord(
path=f"/images/dryhelp{i}.jpg", sha256=f"{i:064d}", size_bytes=1,
mime="image/jpeg", width=1, height=1, origin="imported_filesystem",
integrity_status="unknown", siglip_embedding=[0.0] * 1152,
)
db_sync.add(img)
imgs.append(img)
db_sync.flush()
# tag a: applied manually (img0), applied by an AUTO source (img1), rejected (img2).
db_sync.execute(image_tag.insert().values(
image_record_id=imgs[0].id, tag_id=a.id, source="manual"))
db_sync.execute(image_tag.insert().values(
image_record_id=imgs[1].id, tag_id=a.id, source="head_auto"))
db_sync.add(TagSuggestionRejection(image_record_id=imgs[2].id, tag_id=a.id))
# tag b: applied to img3 only.
db_sync.execute(image_tag.insert().values(
image_record_id=imgs[3].id, tag_id=b.id, source="manual"))
db_sync.flush()
skip = _applied_or_rejected(db_sync, [a.id, b.id])
# Applied-under-ANY-source (manual + head_auto) rejected, kept per-tag; the
# untouched image (img4) appears under neither tag.
assert skip[a.id] == {imgs[0].id, imgs[1].id, imgs[2].id}
assert skip[b.id] == {imgs[3].id}
assert imgs[4].id not in skip[a.id]
assert imgs[4].id not in skip[b.id]
+1 -1
View File
@@ -9,8 +9,8 @@ pytestmark = pytest.mark.integration
def test_non_serialized_platform_has_no_lock():
# gallery-dl platforms aren't capped — they get no lock at all.
assert platform_lock("deviantart", ttl_seconds=60) is None
assert platform_lock("hentaifoundry", ttl_seconds=60) is None
assert platform_lock("discord", ttl_seconds=60) is None
def test_subscribestar_is_serialized():
+11 -2
View File
@@ -11,10 +11,10 @@ from backend.app.services.platforms import (
)
def test_known_platform_keys_is_gs_six():
def test_known_platform_keys_is_gs_five():
assert known_platform_keys() == frozenset({
"patreon", "subscribestar", "hentaifoundry",
"discord", "pixiv", "deviantart",
"discord", "pixiv",
})
@@ -23,6 +23,15 @@ def test_fanbox_not_in_registry():
assert "fanbox" not in PLATFORMS
def test_deviantart_is_retired():
# #3069 executed the 2026-07-05 drop decision (FC downloaders = ART-
# DEDICATED services only). The registry is what /api/platforms, the
# source validator and the credential validator all read, so its absence
# here is what actually retires the platform everywhere else.
assert "deviantart" not in PLATFORMS
assert auth_type_for("deviantart") is None
def test_auth_type_for_known_and_unknown():
assert auth_type_for("patreon") == "cookies"
assert auth_type_for("discord") == "token"
+2 -2
View File
@@ -169,7 +169,7 @@ async def test_scroll_filters_by_artist(db):
async def test_scroll_filters_by_platform(db):
artist = await _seed_artist(db, "alice-platf")
src_p = await _seed_source(db, artist.id, "patreon", "https://p/alice-pp")
src_d = await _seed_source(db, artist.id, "deviantart", "https://d/alice-dd")
src_d = await _seed_source(db, artist.id, "discord", "https://d/alice-dd")
now = datetime.now(UTC)
pp = await _seed_post(db, src_p.id, external_id="PP", post_date=now)
await _seed_post(db, src_d.id, external_id="PD", post_date=now)
@@ -234,7 +234,7 @@ async def test_scroll_combined_artist_and_platform(db):
alice = await _seed_artist(db, "alice-combo")
bob = await _seed_artist(db, "bob-combo")
src_alice_patreon = await _seed_source(db, alice.id, "patreon", "https://p/alice-c")
src_alice_da = await _seed_source(db, alice.id, "deviantart", "https://d/alice-c")
src_alice_da = await _seed_source(db, alice.id, "discord", "https://d/alice-c")
src_bob_patreon = await _seed_source(db, bob.id, "patreon", "https://p/bob-c")
now = datetime.now(UTC)
target = await _seed_post(
+63
View File
@@ -0,0 +1,63 @@
"""recover_stalled_head_training_runs + recover_stalled_head_auto_apply_runs share
one helper (_recover_stalled_runs, DRY pass #161). These pin BOTH wrappers so the
shared source stays correct: a 'running' row with no progress past the stall
threshold flips to 'error'; a fresh 'running' row is left alone."""
from datetime import UTC, datetime, timedelta
import pytest
from sqlalchemy import select
from backend.app.models import HeadAutoApplyRun, HeadTrainingRun
pytestmark = pytest.mark.integration
def test_recover_stalled_head_training_runs_flips_stalled_keeps_fresh(db_sync):
from backend.app.tasks.maintenance import recover_stalled_head_training_runs
stale = HeadTrainingRun(
params={}, status="running",
last_progress_at=datetime.now(UTC) - timedelta(days=1),
)
fresh = HeadTrainingRun(
params={}, status="running", last_progress_at=datetime.now(UTC),
)
db_sync.add_all([stale, fresh])
db_sync.commit()
stale_id, fresh_id = stale.id, fresh.id
assert recover_stalled_head_training_runs.apply().get() == 1
db_sync.expire_all()
assert db_sync.execute(
select(HeadTrainingRun.status).where(HeadTrainingRun.id == stale_id)
).scalar_one() == "error"
assert db_sync.execute(
select(HeadTrainingRun.status).where(HeadTrainingRun.id == fresh_id)
).scalar_one() == "running"
def test_recover_stalled_head_auto_apply_runs_flips_stalled_keeps_fresh(db_sync):
from backend.app.tasks.maintenance import recover_stalled_head_auto_apply_runs
stale = HeadAutoApplyRun(
dry_run=False, params={}, status="running",
last_progress_at=datetime.now(UTC) - timedelta(days=1),
)
fresh = HeadAutoApplyRun(
dry_run=False, params={}, status="running",
last_progress_at=datetime.now(UTC),
)
db_sync.add_all([stale, fresh])
db_sync.commit()
stale_id, fresh_id = stale.id, fresh.id
assert recover_stalled_head_auto_apply_runs.apply().get() == 1
db_sync.expire_all()
assert db_sync.execute(
select(HeadAutoApplyRun.status).where(HeadAutoApplyRun.id == stale_id)
).scalar_one() == "error"
assert db_sync.execute(
select(HeadAutoApplyRun.status).where(HeadAutoApplyRun.id == fresh_id)
).scalar_one() == "running"
+324
View File
@@ -0,0 +1,324 @@
"""Layer-2 auto-refetch remediation — `services/refetch_service.py` (#3071).
This module was the only one under `backend/app/services/` with no test
file, which matters more than a coverage gap normally would: it runs
UNATTENDED off the recovery sweep (`tasks/maintenance.py`, gated by
FC_AUTO_REFETCH_CORRUPT) and it DELETES a file from disk before asking a
downloader for a fresh copy. The frontend cites it by name as the reason
the Import tab could be retired at all (`stores/import.js`: imports
"heal themselves").
`test_api_import_admin.py` already drives the happy path end-to-end
through `POST /api/import/tasks/<id>/refetch` — file deleted, task
flagged, one dispatch, second attempt a no-op. What it CANNOT reach is
the branching inside `resolve_refetch_source`, and it never proves the
negative that actually protects the operator's data: that a file whose
source does NOT resolve is still on disk afterwards. Its `no_source`
case points at a path that never existed, so nothing survives to check.
Those two things are what this module covers.
"""
import json
from pathlib import Path
import pytest
from sqlalchemy import select
from backend.app.models import Artist, ImportBatch, ImportTask, Source
from backend.app.services.refetch_service import (
attempt_refetch,
resolve_refetch_source,
)
pytestmark = pytest.mark.integration
# --- fixtures / helpers ----------------------------------------------------
@pytest.fixture
def import_root(tmp_path):
root = tmp_path / "import"
root.mkdir()
return root
def _media(import_root: Path, artist_dir: str, name: str = "post.jpg") -> Path:
"""A corrupt-import stand-in at import_root/<artist_dir>/<name>."""
d = import_root / artist_dir if artist_dir else import_root
d.mkdir(parents=True, exist_ok=True)
m = d / name
m.write_bytes(b"corrupt-bytes")
return m
def _sidecar(media: Path, payload) -> Path:
"""gallery-dl writes `<stem>.json` beside the media file."""
sc = media.with_suffix(".json")
sc.write_text(payload if isinstance(payload, str) else json.dumps(payload))
return sc
def _artist(session, name: str) -> Artist:
a = Artist(name=name, slug=name.lower())
session.add(a)
session.flush()
return a
def _source(session, artist, platform="patreon", url=None, enabled=True) -> Source:
s = Source(
artist_id=artist.id,
platform=platform,
url=url if url is not None else f"https://www.{platform}.com/{artist.slug}",
enabled=enabled,
config_overrides={},
)
session.add(s)
session.flush()
return s
def _task(session, media: Path, refetched: bool = False) -> ImportTask:
batch = ImportBatch(
triggered_by="manual", source_path=str(media.parent), scan_mode="quick",
)
session.add(batch)
session.flush()
t = ImportTask(
batch_id=batch.id, source_path=str(media), task_type="media",
status="failed", refetched=refetched,
)
session.add(t)
session.flush()
return t
@pytest.fixture
def no_dispatch(monkeypatch):
"""Capture download_source.delay instead of queueing a real re-check.
refetch_service imports the task lazily (inside attempt_refetch, to
dodge a tasks->services->tasks cycle), so patching the attribute on
the module is enough — the import resolves at call time.
"""
from backend.app.tasks import download as download_mod
calls = []
monkeypatch.setattr(download_mod.download_source, "delay", calls.append)
return calls
# --- resolve_refetch_source: what counts as re-pollable --------------------
def test_resolve_finds_enabled_source_matching_the_sidecar_platform(db_sync, import_root):
m = _media(import_root, "Alice")
_sidecar(m, {"category": "patreon", "post_id": 1})
artist = _artist(db_sync, "Alice")
src = _source(db_sync, artist)
assert resolve_refetch_source(db_sync, str(m), import_root).id == src.id
def test_resolve_skips_a_disabled_source(db_sync, import_root):
"""A disabled Source is not re-pollable: the operator turned it off,
and a sweep must not reach past that to delete their file."""
m = _media(import_root, "Alice")
_sidecar(m, {"category": "patreon"})
_source(db_sync, _artist(db_sync, "Alice"), enabled=False)
assert resolve_refetch_source(db_sync, str(m), import_root) is None
def test_resolve_rejects_a_synthetic_sidecar_anchor_url(db_sync, import_root):
"""`sidecar:<platform>:<slug>` is a bookkeeping anchor for files that
arrived on disk, not a feed. Re-polling it is impossible, so it must
not qualify — otherwise the file is deleted for a fetch that can
never happen."""
m = _media(import_root, "Alice")
_sidecar(m, {"category": "patreon"})
_source(db_sync, _artist(db_sync, "Alice"), url="sidecar:patreon:alice")
assert resolve_refetch_source(db_sync, str(m), import_root) is None
def test_resolve_needs_a_source_on_the_sidecars_own_platform(db_sync, import_root):
m = _media(import_root, "Alice")
_sidecar(m, {"category": "patreon"})
_source(db_sync, _artist(db_sync, "Alice"), platform="pixiv")
assert resolve_refetch_source(db_sync, str(m), import_root) is None
def test_resolve_picks_the_lowest_id_when_several_sources_qualify(db_sync, import_root):
m = _media(import_root, "Alice")
_sidecar(m, {"category": "patreon"})
artist = _artist(db_sync, "Alice")
first = _source(db_sync, artist, url="https://www.patreon.com/alice-one")
_source(db_sync, artist, url="https://www.patreon.com/alice-two")
# Deterministic choice, not "whichever the planner returned first" —
# the pick decides which downloader runs.
assert resolve_refetch_source(db_sync, str(m), import_root).id == first.id
def test_resolve_reads_the_gallery_dl_numbered_sidecar(db_sync, import_root):
"""gallery-dl prefixes media with `NN_` for in-post ordering but
writes the sidecar under the UNPREFIXED stem. Refetch resolves real
downloaded files, so it has to follow that convention."""
m = _media(import_root, "Alice", name="01_post.jpg")
(import_root / "Alice" / "post.json").write_text(json.dumps({"category": "patreon"}))
src = _source(db_sync, _artist(db_sync, "Alice"))
assert resolve_refetch_source(db_sync, str(m), import_root).id == src.id
# --- resolve_refetch_source: every way it declines -------------------------
def test_resolve_declines_without_a_sidecar(db_sync, import_root):
m = _media(import_root, "Alice")
_source(db_sync, _artist(db_sync, "Alice"))
assert resolve_refetch_source(db_sync, str(m), import_root) is None
def test_resolve_declines_on_unreadable_sidecar_json(db_sync, import_root):
m = _media(import_root, "Alice")
_sidecar(m, "{not valid json")
_source(db_sync, _artist(db_sync, "Alice"))
assert resolve_refetch_source(db_sync, str(m), import_root) is None
def test_resolve_declines_when_the_sidecar_is_not_an_object(db_sync, import_root):
# A bare JSON list parses fine but has no `category` to read.
m = _media(import_root, "Alice")
_sidecar(m, ["patreon"])
_source(db_sync, _artist(db_sync, "Alice"))
assert resolve_refetch_source(db_sync, str(m), import_root) is None
def test_resolve_declines_when_the_sidecar_names_no_platform(db_sync, import_root):
m = _media(import_root, "Alice")
_sidecar(m, {"post_id": 1})
_source(db_sync, _artist(db_sync, "Alice"))
assert resolve_refetch_source(db_sync, str(m), import_root) is None
def test_resolve_declines_for_a_file_sitting_directly_in_import_root(db_sync, import_root):
"""No artist folder means no artist bucket to resolve — the
filesystem-only drop case."""
m = _media(import_root, "")
_sidecar(m, {"category": "patreon"})
_source(db_sync, _artist(db_sync, "Alice"))
assert resolve_refetch_source(db_sync, str(m), import_root) is None
def test_resolve_declines_when_no_artist_row_matches_the_folder(db_sync, import_root):
m = _media(import_root, "Nobody")
_sidecar(m, {"category": "patreon"})
_source(db_sync, _artist(db_sync, "Alice"))
assert resolve_refetch_source(db_sync, str(m), import_root) is None
# --- attempt_refetch: the destructive half ---------------------------------
def test_attempt_refetch_deletes_the_file_and_queues_one_recheck(
db_sync, import_root, no_dispatch,
):
m = _media(import_root, "Alice")
_sidecar(m, {"category": "patreon"})
src = _source(db_sync, _artist(db_sync, "Alice"))
task = _task(db_sync, m)
result = attempt_refetch(db_sync, task, import_root)
assert result == {"status": "refetch_queued", "source_id": src.id}
assert not m.exists() # the bad copy is gone...
assert no_dispatch == [src.id] # ...and exactly one re-check was queued
assert task.refetched is True
def test_attempt_refetch_leaves_the_file_alone_when_nothing_resolves(
db_sync, import_root, no_dispatch,
):
"""THE assertion this module exists for. `no_source` is the common
case on a filesystem-only library, and the file on disk is then the
operator's ONLY copy — deleting it without a downloader that can
replace it destroys the thing the remediation was meant to repair.
The route-level `no_source` test cannot catch a regression here: its
path never existed, so an unconditional unlink would pass it.
"""
m = _media(import_root, "Alice")
_sidecar(m, {"category": "patreon"})
_source(db_sync, _artist(db_sync, "Alice"), enabled=False)
task = _task(db_sync, m)
assert attempt_refetch(db_sync, task, import_root) == {"status": "no_source"}
assert m.exists()
assert m.read_bytes() == b"corrupt-bytes"
assert no_dispatch == []
assert task.refetched is False # not consumed — a real fix can still run
def test_attempt_refetch_is_bounded_to_a_single_attempt(
db_sync, import_root, no_dispatch,
):
"""The `refetched` bound is what stops SOURCE-side corruption from
looping: re-downloading a file that is broken upstream returns the
same bytes forever. The check must come FIRST — a second call has to
leave the (re-downloaded) file untouched, not delete it again.
"""
m = _media(import_root, "Alice")
_sidecar(m, {"category": "patreon"})
_source(db_sync, _artist(db_sync, "Alice"))
task = _task(db_sync, m, refetched=True)
assert attempt_refetch(db_sync, task, import_root) == {"status": "already_refetched"}
assert m.exists()
assert no_dispatch == []
def test_attempt_refetch_proceeds_when_the_file_is_already_gone(
db_sync, import_root, no_dispatch,
):
"""`missing_ok=True`: the sweep races an operator who deleted the bad
file by hand. The re-check is still the right next move."""
m = _media(import_root, "Alice")
_sidecar(m, {"category": "patreon"})
src = _source(db_sync, _artist(db_sync, "Alice"))
task = _task(db_sync, m)
m.unlink()
assert attempt_refetch(db_sync, task, import_root)["status"] == "refetch_queued"
assert no_dispatch == [src.id]
def test_attempt_refetch_survives_an_unlink_failure(
db_sync, import_root, no_dispatch,
):
"""An unremovable path is logged and stepped over, not raised — this
runs unattended, and a raise would abort the whole recovery sweep for
every OTHER poison-pill row in the batch.
A directory standing where the media file should be produces a
genuine IsADirectoryError (an OSError) without patching pathlib, so
the handler is exercised rather than simulated.
"""
d = import_root / "Alice" / "post.jpg"
d.mkdir(parents=True)
(import_root / "Alice" / "post.json").write_text(json.dumps({"category": "patreon"}))
src = _source(db_sync, _artist(db_sync, "Alice"))
task = _task(db_sync, d)
assert attempt_refetch(db_sync, task, import_root)["status"] == "refetch_queued"
assert d.exists() # removal genuinely failed...
assert no_dispatch == [src.id] # ...and the sweep carried on anyway
assert db_sync.execute(
select(ImportTask.refetched).where(ImportTask.id == task.id)
).scalar_one() is True
+4 -2
View File
@@ -23,12 +23,14 @@ async def _artist(db, name="Alice"):
@pytest.mark.asyncio
async def test_known_platforms_is_gs_six(db):
async def test_known_platforms_is_gs_five(db):
assert KNOWN_PLATFORMS == frozenset({
"patreon", "subscribestar", "hentaifoundry",
"discord", "pixiv", "deviantart",
"discord", "pixiv",
})
assert "fanbox" not in KNOWN_PLATFORMS
# Retired at #3069 — a source can no longer be created on it.
assert "deviantart" not in KNOWN_PLATFORMS
@pytest.mark.asyncio