Compare commits

...
Author SHA1 Message Date
bvandeusenandClaude Opus 5.5 ef591edf10 feat: the extension picks a Discord source's posters from who posts in the channel, stored by id (#4488)
CI and images / lint (push) Successful in 3s
CI and images / extension-version (push) Successful in 3s
CI and images / extension-test (push) Successful in 18s
CI and images / frontend-build (push) Successful in 26s
CI and images / backend-lint-and-test (push) Successful in 32s
CI and images / integration (push) Successful in 2m24s
CI and images / build-agent (push) Successful in 5s
CI and images / sign-extension (push) Successful in 6m13s
CI and images / build-web (push) Successful in 1m38s
CI and images / smoke-web (push) Successful in 54s
CI and images / promote (push) Successful in 2s
On a Discord channel page the Add panel lists who posted in the newest 200
messages, the creator the server is named for (or owned by) ticked. On a
channel FabledCurator already follows, the chip opens the same list with Save
and Open artist. The list is kept as user ids, so a rename never quietly stops
it matching; the name each was picked under rides beside it for display, and
the Subscriptions dialog shows it.

- DiscordClient.recent_posters + rank_posters (owner / name-matches-server;
  bots never suggested; it suggests, the operator ticks).
- GET/POST /api/extension/discord/posters; quick-add takes discord_authors.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 09:47:27 -04:00
bvandeusenandClaude Opus 5.5 1efece0f21 feat: a Discord source can remove the posts it took from people outside its poster list (#4486)
CI and images / extension-version (push) Successful in 3s
CI and images / lint (push) Successful in 3s
CI and images / extension-test (push) Successful in 20s
CI and images / frontend-build (push) Successful in 21s
CI and images / backend-lint-and-test (push) Successful in 33s
CI and images / integration (push) Successful in 2m39s
CI and images / sign-extension (push) Successful in 3s
CI and images / build-agent (push) Successful in 6s
CI and images / build-web (push) Successful in 1m43s
CI and images / smoke-web (push) Successful in 54s
CI and images / promote (push) Successful in 2s
"Remove posts from other posters" on a Discord source with an "Only posts by"
list previews what goes, per poster, then deletes it behind a typed token:

- posts whose record names a poster not on the list (by id, username or
  display name); posts that record no poster are left alone and counted;
- images found only on those posts; one also on a kept post stays, and a
  synthetic drop's link never keeps one;
- their attachments, and every drop that absorbed one of them, so the grouping
  sweep regroups what remains.

Preview and apply share one predicate, and a parity test holds them to it.
The post record now also saves the poster's display name.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 08:34:06 -04:00
bvandeusenandClaude Opus 5.5 20aeac81df style: a source's backfill state chips are filled, so their text reads on the dark row
CI and images / lint (push) Successful in 2s
CI and images / extension-version (push) Successful in 2s
CI and images / extension-test (push) Successful in 19s
CI and images / frontend-build (push) Successful in 21s
CI and images / backend-lint-and-test (push) Successful in 31s
CI and images / integration (push) Successful in 2m23s
CI and images / sign-extension (push) Successful in 4s
CI and images / build-agent (push) Successful in 6s
CI and images / build-web (push) Successful in 1m54s
CI and images / smoke-web (push) Successful in 54s
CI and images / promote (push) Successful in 1s
The theme's success and info colours are dark; a tonal chip draws its text in
that colour, which vanished against the Subscriptions row.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 08:22:17 -04:00
bvandeusenandClaude Opus 5.5 d6d3184361 fix: the single-color filter no longer takes line art for a blank image (#4483)
CI and images / lint (push) Successful in 2s
CI and images / extension-version (push) Successful in 2s
CI and images / extension-test (push) Successful in 16s
CI and images / frontend-build (push) Successful in 21s
CI and images / backend-lint-and-test (push) Successful in 31s
CI and images / integration (push) Successful in 2m37s
CI and images / sign-extension (push) Successful in 2s
CI and images / build-agent (push) Successful in 5s
CI and images / build-web (push) Successful in 1m36s
CI and images / smoke-web (push) Successful in 56s
CI and images / promote (push) Successful in 2s
Five of Todding's Discord doodles (pencil lines on white, 3000px) were skipped
on import as "single color", so their posts showed text and no image. The
predicate sampled a 64px BILINEAR thumbnail, which blends thin strokes into
the paper, and 0.95 "one color" is below how white a doodle is.

- Sample 256x256 by NEAREST, so each sample is a real pixel.
- Default threshold 0.995: blank means essentially blank. Migration 0116 moves
  a stored 0.95 (the old default) with it; the settings slider now spans
  0.9-1 in 0.005 steps so the value is reachable.
- The Cleanup audit shares the predicate, so it stops flagging sketches too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 08:12:43 -04:00
bvandeusenandClaude Opus 5.5 29c22afcb7 fix: a Discord source takes only messages with an image attached, from the posters it names, and knows its channel's name (#4481)
CI and images / lint (push) Successful in 4s
CI and images / extension-version (push) Successful in 3s
CI and images / frontend-build (push) Successful in 21s
CI and images / extension-test (push) Successful in 20s
CI and images / backend-lint-and-test (push) Successful in 34s
CI and images / integration (push) Successful in 2m24s
CI and images / sign-extension (push) Successful in 3s
CI and images / build-agent (push) Successful in 7s
CI and images / build-web (push) Successful in 1m42s
CI and images / smoke-web (push) Successful in 59s
CI and images / promote (push) Successful in 2s
- A message with no image or video attachment is chat: extract_media returns
  nothing for it, so nothing downloads and no post record is written. A
  message that passes keeps every file, numbered as gallery-dl numbers them.
- `discord_authors` in a source's config limits the walk to those posters
  (id, username or display name); the edit dialog has a field for it and no
  longer drops config keys it has no field for.
- source.display_name (migration 0115), refreshed by every Discord walk, shows
  on Subscriptions in place of the two-id URL; a Discord post card names its
  channel from the record's `channel`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 07:03:47 -04:00
bvandeusenandClaude Opus 5.5 42a60e3e74 test: the patreon ingester tests run phase 3's seen-marking before counting post keys (#4436)
CI and images / lint (push) Successful in 4s
CI and images / extension-version (push) Successful in 4s
CI and images / extension-test (push) Successful in 21s
CI and images / frontend-build (push) Successful in 22s
CI and images / backend-lint-and-test (push) Successful in 33s
CI and images / integration (push) Successful in 2m24s
CI and images / sign-extension (push) Successful in 3s
CI and images / build-agent (push) Successful in 6s
CI and images / build-web (push) Successful in 1m42s
CI and images / smoke-web (push) Successful in 52s
CI and images / promote (push) Successful in 1s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-25 11:41:16 -04:00
bvandeusenandClaude Opus 5.5 621c2b8315 fix: a post record is marked seen only after it is upserted, and an hourly sweep dates the posts killed walks left undated (#4436)
CI and images / lint (push) Successful in 3s
CI and images / extension-version (push) Successful in 3s
CI and images / extension-test (push) Successful in 17s
CI and images / frontend-build (push) Successful in 27s
CI and images / backend-lint-and-test (push) Successful in 33s
CI and images / integration (push) Failing after 2m25s
CI and images / sign-extension (push) Skipped
CI and images / build-web (push) Skipped
CI and images / smoke-web (push) Skipped
CI and images / promote (push) Skipped
CI and images / build-agent (push) Skipped
ingest_core marked a post's record key seen when it wrote _post.json, but
the record reaches the database (and dates the post and its images) only
in phase 3. A walk killed before phase 3 left posts the ledger called
recorded and the database never dated; ticks early-out long before
reaching them again. Keys now join the media in mark_seen_after_import.

date_posts_from_records finds each undated native post's record under
its artist's folder and upserts it with the post's own source. Hourly on
maintenance_long; an empty query once everything is dated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-25 11:36:17 -04:00
bvandeusenandClaude Opus 5.5 ca802bd885 fix: a non-image file lands on the source it was downloaded for, and the undated shells it left are folded into their real posts (#4435)
CI and images / lint (push) Successful in 3s
CI and images / extension-version (push) Successful in 3s
CI and images / extension-test (push) Successful in 18s
CI and images / frontend-build (push) Successful in 22s
CI and images / backend-lint-and-test (push) Successful in 32s
CI and images / integration (push) Successful in 2m22s
CI and images / sign-extension (push) Successful in 3s
CI and images / build-agent (push) Successful in 5s
CI and images / build-web (push) Successful in 1m39s
CI and images / smoke-web (push) Successful in 54s
CI and images / promote (push) Successful in 1s
The attachment path looked the post's source up by (artist, platform),
taking the artist's lowest-id source. A Discord artist has one source per
channel, so every archive or pdf from a later channel became an undated
second post on the first channel's source. _post_for_sidecar now takes
the downloading source, as upsert_post_record and _apply_sidecar already
did. Migration 0114 re-points each shell's attachments to its dated twin
and deletes the shell.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-25 10:56:08 -04:00
bvandeusenandClaude Opus 5.5 e2eeb63115 style: sort the boot hook's imports (#4433)
CI and images / lint (push) Successful in 2s
CI and images / extension-version (push) Successful in 2s
CI and images / extension-test (push) Successful in 17s
CI and images / frontend-build (push) Successful in 23s
CI and images / backend-lint-and-test (push) Successful in 32s
CI and images / integration (push) Successful in 2m24s
CI and images / sign-extension (push) Successful in 3s
CI and images / build-agent (push) Successful in 6s
CI and images / build-web (push) Successful in 1m43s
CI and images / smoke-web (push) Successful in 54s
CI and images / promote (push) Successful in 1s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-25 10:36:32 -04:00
bvandeusenandClaude Opus 5.5 e704c70f32 fix: the download boot hook imports a model-only module, keeping the membership roster off the fetch path (#4433)
CI and images / lint (push) Failing after 4s
CI and images / extension-version (push) Successful in 4s
CI and images / extension-test (push) Successful in 18s
CI and images / frontend-build (push) Successful in 21s
CI and images / backend-lint-and-test (push) Successful in 32s
CI and images / integration (push) Successful in 2m22s
CI and images / sign-extension (push) Skipped
CI and images / build-web (push) Skipped
CI and images / smoke-web (push) Skipped
CI and images / promote (push) Skipped
CI and images / build-agent (push) Skipped
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-25 10:32:41 -04:00
bvandeusenandClaude Opus 5.5 0fad744bfb fix: a restart no longer strands downloads — the download lane closes its predecessor's runs as interrupted and frees the platform locks at boot (#4433)
CI and images / lint (push) Successful in 4s
CI and images / extension-version (push) Successful in 3s
CI and images / extension-test (push) Successful in 19s
CI and images / frontend-build (push) Successful in 21s
CI and images / backend-lint-and-test (push) Failing after 31s
CI and images / integration (push) Successful in 2m23s
CI and images / sign-extension (push) Skipped
CI and images / build-web (push) Skipped
CI and images / smoke-web (push) Skipped
CI and images / promote (push) Skipped
CI and images / build-agent (push) Skipped
A walk that outlives the 90s stop grace is SIGKILLed: its event never
finalizes and its platform lock is held for the 27-min TTL. The 30-min
sweep then errored every stranded event and bumped consecutive_failures,
backing sources off (and blocking backfills) as if the platform failed.

On worker_ready, the process consuming 'download' ends pre-boot
pending/running events as skipped with error_type 'interrupted', leaves
the source untouched so the next tick resumes it, and releases the
serialized platforms' locks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-25 10:29:05 -04:00
bvandeusenandClaude Opus 5.5 32874ca678 fix: the release page drops internal rule numbers and the miscounted "three" images
CI and images / lint (push) Successful in 2s
CI and images / extension-version (push) Successful in 2s
CI and images / extension-test (push) Successful in 17s
CI and images / frontend-build (push) Successful in 20s
CI and images / backend-lint-and-test (push) Successful in 32s
CI and images / integration (push) Successful in 2m21s
CI and images / sign-extension (push) Successful in 3s
CI and images / build-web (push) Successful in 6s
CI and images / build-agent (push) Successful in 6s
CI and images / smoke-web (push) Successful in 41s
CI and images / promote (push) Successful in 1s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-25 10:09:47 -04:00
bvandeusenandClaude Opus 5.5 edd2daa16a docs: the README says ML tagging ships off and fetches its weights when enabled
CI and images / lint (push) Successful in 3s
CI and images / extension-version (push) Successful in 4s
CI and images / extension-test (push) Successful in 20s
CI and images / frontend-build (push) Successful in 23s
CI and images / backend-lint-and-test (push) Successful in 33s
CI and images / integration (push) Successful in 2m25s
CI and images / sign-extension (push) Successful in 3s
CI and images / build-agent (push) Successful in 6s
CI and images / build-web (push) Successful in 6s
CI and images / smoke-web (push) Successful in 41s
CI and images / promote (push) Successful in 1s
The first-run note still said the ML worker downloads its model weights
on first boot. Milestone 422 moved that fetch to the moment the ML lane is
enabled under Settings → System, and the lane ships at 0 slots. The
overview paragraph is also the first release's notes (release_notes.py
reads it between the overview markers), so it now says the same.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-25 09:58:48 -04:00
bvandeusenandClaude Opus 5.5 23e062dd4a test: the stall-sweep tests read the thresholds they check, and stay clear of the new import value (#4432)
CI and images / lint (push) Successful in 3s
CI and images / extension-version (push) Successful in 4s
CI and images / extension-test (push) Successful in 18s
CI and images / frontend-build (push) Successful in 23s
CI and images / backend-lint-and-test (push) Successful in 33s
CI and images / integration (push) Successful in 2m21s
CI and images / sign-extension (push) Successful in 3s
CI and images / build-agent (push) Successful in 6s
CI and images / build-web (push) Successful in 1m39s
CI and images / smoke-web (push) Successful in 56s
CI and images / promote (push) Successful in 1s
Run 7505 failed test_recover_stalled_task_runs_ml_queue_uses_longer_threshold.
It restated the old 25-minute ml threshold as a 30-minute "stale" row,
which is now inside the 40-minute window. The test now reads the value
from QUEUE_STUCK_THRESHOLD_MINUTES.

The archive test's fast-import row was exactly 10 minutes old, which is
the new import threshold, so it passed only by the milliseconds between
seeding and sweeping. It is now 15 minutes old.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-25 09:51:42 -04:00
bvandeusenandClaude Opus 5.5 a360d69ee8 fix: task runs record the lane Celery really routes them to, and no healthy long job is swept as stalled (#4432)
CI and images / lint (push) Successful in 3s
CI and images / extension-version (push) Successful in 4s
CI and images / frontend-build (push) Successful in 25s
CI and images / extension-test (push) Successful in 28s
CI and images / backend-lint-and-test (push) Successful in 34s
CI and images / integration (push) Failing after 2m24s
CI and images / sign-extension (push) Skipped
CI and images / build-web (push) Skipped
CI and images / smoke-web (push) Skipped
CI and images / promote (push) Skipped
CI and images / build-agent (push) Skipped
celery_signals._queue_for was a hand-kept copy of task_routes and had
drifted:

- backup, admin and library_audit jobs, and backfill_phash, run on
  maintenance_long but were recorded as `maintenance`;
- translation and gpu_queue jobs were recorded as `default`, where the
  5-minute stall sweep failed healthy 35-minute translation runs.

It now asks the router, cached per task name.

A new guard test checks every registered task's hard time limit against
the stall threshold the sweep would use for it. It also caught these
sweeps, which failed healthy runs mid-flight and are fixed here:

- ml's scheduled sweeps (35 min, previously swept at 25);
- train_heads and apply_head_tags (65 min);
- import_media_file (6 min, previously swept at 5);
- the long lane, which now has its own 45-minute threshold.

UI changes:

- the admin job poller follows a job by celery_task_id, via a new filter
  on /runs, instead of by lane;
- the archive re-extract and missing-file repair cards show the long
  lane's backlog, where their jobs actually wait;
- the queue table lists maintenance_long.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-25 09:46:39 -04:00
bvandeusenandClaude Opus 5.5 dfd28a0aa6 fix(ci): the base refresh's lanes test main, the branch it publishes (#4430)
CI and images / lint (push) Successful in 3s
CI and images / extension-version (push) Successful in 3s
CI and images / extension-test (push) Successful in 20s
CI and images / frontend-build (push) Successful in 23s
CI and images / backend-lint-and-test (push) Successful in 31s
CI and images / integration (push) Successful in 2m24s
CI and images / sign-extension (push) Successful in 3s
CI and images / build-agent (push) Successful in 5s
CI and images / build-web (push) Successful in 5s
CI and images / smoke-web (push) Successful in 40s
CI and images / promote (push) Successful in 1s
A refresh publishes `:latest` from `main`, but the six lanes kept the
default checkout, which is the cron's triggering commit on dev. The gate
therefore tested dev's code and passed main's.

The new LANE_REF is `main` on a refresh and empty otherwise. Empty keeps
the checkout default, so push and PR runs are unchanged, including a PR's
merge ref, which the default reaches by ref rather than by SHA.

Every lane still runs on every trigger: no condition is added and no
lane can skip, so the `needs:` gate on each publishing job is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-25 09:40:08 -04:00
bvandeusenandClaude Opus 5.5 efcb548ebf fix: natively downloaded images take their post's date, not their download time (#4431)
CI and images / lint (push) Successful in 3s
CI and images / extension-version (push) Successful in 3s
CI and images / extension-test (push) Successful in 19s
CI and images / frontend-build (push) Successful in 22s
CI and images / backend-lint-and-test (push) Successful in 33s
CI and images / integration (push) Successful in 2m22s
CI and images / sign-extension (push) Successful in 3s
CI and images / build-agent (push) Successful in 6s
CI and images / build-web (push) Successful in 1m40s
CI and images / smoke-web (push) Successful in 55s
CI and images / promote (push) Successful in 2s
The native ingesters (Discord, Patreon, SubscribeStar) import a post's
media before its record. Only the record (`_post.json`) carries the
date: the Discord message timestamp, or the Patreon/SubscribeStar
published_at. So each image was linked to a post with no date yet, and
it kept its download time in both gallery date columns. The post itself
was dated correctly once the record landed, but nothing went back to
update its images.

- upsert_post_record now re-dates the images already linked to the
  post. `effective_date` becomes the primary post's date, and
  `earliest_post_date` the earliest dated post the image is in, which
  are the same rules _attach_provenance applies.
- Migration 0113 repairs the library from the posts. It writes only
  rows that differ.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-25 09:30:17 -04:00
bvandeusenandClaude Opus 5.5 4fa7975963 fix(ci): every publishing job builds the commit that fired the run, not the branch tip (#4427)
CI and images / lint (push) Successful in 2s
CI and images / extension-version (push) Successful in 2s
CI and images / extension-test (push) Successful in 18s
CI and images / frontend-build (push) Successful in 19s
CI and images / backend-lint-and-test (push) Successful in 32s
CI and images / integration (push) Successful in 2m21s
CI and images / sign-extension (push) Successful in 3s
CI and images / build-agent (push) Successful in 6s
CI and images / smoke-web (push) Successful in 41s
CI and images / promote (push) Successful in 1s
CI and images / build-web (push) Successful in 6s
BUILD_REF resolved to the branch name on ordinary runs, and each job's
checkout re-resolves a branch when that job starts. A push that lands
mid-run therefore moved the later jobs onto the new tip:

- run 7499 signed 423275a's extension;
- its build-web then checked out 83e1382, derived a version nobody had
  signed, and got a 404 on the download.

The guard failed closed. A job without such a guard would have published
a commit the run's lanes never tested.

The lanes already check out github.sha by default. BUILD_REF is now
github.sha on every trigger except the base refresh, which keeps `main`
and its branch guard. The publish therefore builds exactly what the
lanes passed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-25 09:13:41 -04:00
bvandeusenandClaude Opus 5.5 7b1570f2a5 feat: retire the sketch/doodle WIP title tier, and the review strip asks "Is this a WIP?" (milestone 430)
CI and images / lint (push) Successful in 2s
CI and images / extension-version (push) Successful in 3s
CI and images / extension-test (push) Successful in 20s
CI and images / frontend-build (push) Successful in 23s
CI and images / backend-lint-and-test (push) Successful in 32s
CI and images / integration (push) Successful in 2m24s
CI and images / sign-extension (push) Successful in 3s
CI and images / build-agent (push) Successful in 5s
CI and images / build-web (push) Successful in 1m35s
CI and images / smoke-web (push) Successful in 56s
CI and images / promote (push) Successful in 1s
The soft tier (#1474) tagged `wip` on any post titled sketch, doodle or
scribble: 6,096 of the library's 8,876 wip tags. Its conflict audit flagged
most of them, because finished art scores >= 0.5 on some content head,
which filled the Gallery strip with 2,086 cards. A "sketch" is usually
finished work, so the operator retired the tier. The artist's own
"WIP" / "work in progress" title rule and human wip tags stay.

- Removed: the soft matcher, source and prefilter; the importer and backfill
  soft branches; soft_wip_conflict_audit with its task and beat; the
  wip_soft_title_tagging_enabled setting, API field and toggle; and
  wip_title_soft from _AUTO_SOURCES.
- Migration 0112: a soft tag the operator confirmed, or kept from the strip,
  becomes `manual`. Every other soft tag is deleted, along with the open
  review cards whose tag is gone. The column is dropped.
- Review strip (#4424): each card asks "Is this a WIP?" (or banner, and so on),
  and the buttons read "Is a WIP" / "Is not a WIP". The content tag it also
  scored on is shown as the reason it was flagged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-25 08:20:28 -04:00
bvandeusenandClaude Opus 5.5 83e1382812 fix: extension updates install the new build — no 12h-cached "latest" XPI, and the popup's Update opens FC's install page
CI and images / lint (push) Successful in 3s
CI and images / extension-version (push) Successful in 3s
CI and images / frontend-build (push) Successful in 22s
CI and images / extension-test (push) Successful in 22s
CI and images / backend-lint-and-test (push) Successful in 33s
CI and images / integration (push) Successful in 2m23s
CI and images / build-agent (push) Successful in 5s
CI and images / sign-extension (push) Successful in 2m29s
CI and images / build-web (push) Successful in 1m43s
CI and images / smoke-web (push) Successful in 55s
CI and images / promote (push) Successful in 2s
Operator, 2026-09-25: "the extension update trigger from inside the
extension doesn't work and the manual update seems to not move it to the
most recent version or at least mark it the most recent."

- fabledcurator-latest.xpi was served with Quart's default
  `public, max-age=43200`: one URL whose bytes change every release, so a
  browser that had fetched it reinstalled the previous build for 12 hours
  (measured on the instance). It is now `no-cache` (the ETag keeps an
  unchanged file a 304); versioned XPIs are `immutable`.
- The web Settings card installs/downloads the VERSIONED xpi_url, which can
  only ever be that build's bytes.
- The popup's Update button did tabs.create() on the .xpi, which Firefox
  refuses (NS_ERROR_FAILURE on a 200: it only installs from a user click on
  a web page). It now opens FC's install card (/subscriptions?tab=settings).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-25 08:01:02 -04:00
79 changed files with 3007 additions and 439 deletions
+36 -7
View File
@@ -142,6 +142,16 @@ concurrency:
# Deriving it per job invites the two halves to disagree: sign-extension would
# derive dev's extension version while build-web bundled main's, and the
# release download would 404 on a version that exists perfectly well.
#
# On every other trigger it is the COMMIT that fired (`github.sha`), never the
# branch name. A branch is re-resolved by each job's checkout when that job
# starts, so a push landing mid-run moved the later jobs onto the new tip: run
# 7499 signed 423275a's extension, then build-web checked out 83e1382 (pushed
# while 7499 ran), derived a version nobody had signed, and 404'd on the
# download (#4427). The guard failed closed that time; a job without one would
# have published a commit the run's own lanes never tested — the lanes check
# out `github.sha` by default, so pinning here makes the publish build exactly
# what they passed.
# IS THIS A BASE REFRESH? Asked in five places and previously spelled five
# ways — `github.event_name == 'schedule'` in an `if:`, `$GITHUB_EVENT_NAME` in
# one shell, an `EVENT:` env passed into another, and a bare expression on
@@ -178,7 +188,16 @@ concurrency:
# where a step-level `if:` needs the answer before any shell runs.
env:
IS_REFRESH: ${{ (github.event_name == 'schedule' || format('{0}', github.event.inputs.refresh) == 'true') && 'true' || 'false' }}
BUILD_REF: ${{ (github.event_name == 'schedule' || format('{0}', github.event.inputs.refresh) == 'true') && 'main' || github.ref }}
BUILD_REF: ${{ (github.event_name == 'schedule' || format('{0}', github.event.inputs.refresh) == 'true') && 'main' || github.sha }}
# What the six LANES check out. Empty — the checkout default, the triggering
# commit (or a PR's merge ref) — on every trigger but the refresh, where it is
# `main`: the refresh publishes main, so the gate has to test main (#4430). It
# is not BUILD_REF itself because a pull_request run's `github.sha` is a merge
# commit the default checkout reaches through its ref, not by sha. Each job
# still resolves `main` when it starts, so a merge to main during the ~5 min of
# a Sunday-06:00 refresh could put the lanes and the build one commit apart;
# the build jobs' own guards assert the branch, not the commit.
LANE_REF: ${{ (github.event_name == 'schedule' || format('{0}', github.event.inputs.refresh) == 'true') && 'main' || '' }}
# Requires repo secret RELEASE_TOKEN — a Forgejo PAT with scopes:
# - write:package, read:package (for docker push to git.fabledsword.com)
@@ -219,6 +238,8 @@ jobs:
image: git.fabledsword.com/bvandeusen/ci-python:3.14
steps:
- uses: actions/checkout@v4
with:
ref: ${{ env.LANE_REF }}
- name: Ruff lint
# agent/ included so the GPU-agent is linted before its image is built
# (build.yml only `docker build`s it — this is where it gets checked).
@@ -265,6 +286,7 @@ jobs:
steps:
- uses: actions/checkout@v4
with:
ref: ${{ env.LANE_REF }}
# The derivation needs real history: a depth-1 clone sees one commit
# and produces a wrong, too-low value RATHER THAN FAILING. Checking
# that here is half the point of the lane.
@@ -320,6 +342,7 @@ jobs:
steps:
- uses: actions/checkout@v4
with:
ref: ${{ env.LANE_REF }}
# Full history for tests/test_artifact_identity.py, which derives
# each artifact's revision to check the identity scheme. On a
# depth-1 clone that derivation either fails or returns the tip sha
@@ -366,6 +389,8 @@ jobs:
working-directory: frontend
steps:
- uses: actions/checkout@v4
with:
ref: ${{ env.LANE_REF }}
# No package-lock.json is tracked yet (we don't run npm locally per
# feedback-no-local-runs). Using `npm install` instead of `npm ci`.
# If we want strict lockfile-based reproducibility later, commit a
@@ -395,6 +420,8 @@ jobs:
image: node:24-bookworm-slim
steps:
- uses: actions/checkout@v4
with:
ref: ${{ env.LANE_REF }}
# Not --no-save: vitest and web-ext are both real devDependencies now,
# and the suite needs vitest resolvable from node_modules.
- name: Install dev dependencies
@@ -497,6 +524,8 @@ jobs:
--health-retries 10
steps:
- uses: actions/checkout@v4
with:
ref: ${{ env.LANE_REF }}
- name: Integration suite (resolve service IPs, migrate, test)
run: |
set -eux
@@ -605,8 +634,8 @@ jobs:
- uses: actions/checkout@v4
with:
# Not the triggering ref — see the `env:` block at the top. On a
# scheduled refresh this is `main`; on everything else it is the ref
# that fired, so this is a no-op on every ordinary path.
# scheduled refresh this is `main`; on everything else it is the
# commit that fired, the same one the lanes above tested.
ref: ${{ env.BUILD_REF }}
# Full history is load-bearing, not a convenience: the version this
# job signs is derived from the commit TIME of the newest packaged
@@ -945,8 +974,8 @@ jobs:
- uses: actions/checkout@v4
with:
# Not the triggering ref — see the `env:` block at the top. On a
# scheduled refresh this is `main`; on everything else it is the ref
# that fired, so this is a no-op on every ordinary path.
# scheduled refresh this is `main`; on everything else it is the
# commit that fired, the same one the lanes above tested.
ref: ${{ env.BUILD_REF }}
# Full history: this job RE-DERIVES the extension version rather than
# being handed it, and a depth-1 clone derives a wrong, too-low value
@@ -2174,8 +2203,8 @@ jobs:
- uses: actions/checkout@v4
with:
# Not the triggering ref — see the `env:` block at the top. On a
# scheduled refresh this is `main`; on everything else it is the ref
# that fired, so this is a no-op on every ordinary path.
# scheduled refresh this is `main`; on everything else it is the
# commit that fired, the same one the lanes above tested.
ref: ${{ env.BUILD_REF }}
# Full history: this job derives its artifact's version from the
# commit its shipped files last changed in (milestone 313). A
+8 -4
View File
@@ -22,7 +22,8 @@ through afterwards.
- **ML tagging.** Runs image models in-container to suggest tags, group
characters, find near-duplicates and power similarity search. Suggestions are
reviewable — it proposes, you confirm, and it learns which proposals you keep
rejecting.
rejecting. It ships switched off: turn it on under Settings → System when
you want it, and it fetches its model weights then.
- **Deduplication and provenance.** Everything that arrives is hashed and
deduplicated by content, metadata sidecars are read wherever the source
writes them, and every file keeps a record of where it came from.
@@ -119,9 +120,12 @@ needs every credential entered again by hand.
A few other things are worth knowing about the first few minutes:
- **The ML worker downloads its model weights on first boot**, several GB from
HuggingFace into `./models`. Until that finishes, tagging is queued rather
than broken. It is idempotent — a restart resumes rather than refetches.
- **ML tagging starts switched off, and nothing is downloaded at boot.** Give
the ML lane a slot under **Settings → System** and it fetches its model
weights then — several GB from HuggingFace into `./models`, shown as a job
under **Settings → Activity** that you can watch and retry. Until it
finishes, tagging is queued rather than broken, and the fetch only takes
what is missing, so turning the lane off and on again does not refetch.
- **The gallery starts empty**, and that is the expected state. Add a creator
under **Subscriptions** and it fills as posts come down.
- **If you already have a library on disk**, there is no screen that imports
@@ -0,0 +1,78 @@
"""Retire the sketch/doodle WIP title tier — its tags, its review flags, its toggle.
Milestone 430, #4428. The soft tier (#1474) tagged `wip` on any post titled
sketch / doodle / scribble. Measured on the operator's library it was 6,096 of
8,876 wip tags, and its conflict audit filled the Gallery's review strip with
2,086 cards, because most finished art scores >= 0.5 on some content head. A
"sketch" is usually finished work, so the operator retired it (2026-09-25).
Data:
* A soft tag the operator stood behind is kept and relabelled `manual`: one they
confirmed (tag_positive_confirmation), or one whose review flag they resolved
while leaving the tag on (the strip's "Keep tag").
* Every other `wip_title_soft` row is deleted.
* Unresolved review flags whose tag is no longer on the image are deleted: the
question they ask no longer applies. That is the audit's cards, and any older
orphan the same way.
Then `import_settings.wip_soft_title_tagging_enabled` is dropped. The downgrade
restores the column only; deleted tags are not recreated.
Revision ID: 0112
Revises: 0111
Create Date: 2026-09-25
"""
import sqlalchemy as sa
from alembic import op
revision = "0112"
down_revision = "0111"
branch_labels = None
depends_on = None
def retire_soft_wip_tags(conn) -> None:
"""The data half, on a plain connection, so a test can run it directly."""
conn.execute(sa.text("""
UPDATE image_tag it SET source = 'manual'
WHERE it.source = 'wip_title_soft'
AND (
EXISTS (
SELECT 1 FROM tag_positive_confirmation c
WHERE c.image_record_id = it.image_record_id AND c.tag_id = it.tag_id
)
OR EXISTS (
SELECT 1 FROM presentation_review pr
WHERE pr.image_record_id = it.image_record_id AND pr.tag_id = it.tag_id
AND pr.resolved_at IS NOT NULL
)
)
"""))
conn.execute(sa.text("DELETE FROM image_tag WHERE source = 'wip_title_soft'"))
conn.execute(sa.text("""
DELETE FROM presentation_review pr
WHERE pr.resolved_at IS NULL
AND NOT EXISTS (
SELECT 1 FROM image_tag it
WHERE it.image_record_id = pr.image_record_id AND it.tag_id = pr.tag_id
)
"""))
def upgrade():
retire_soft_wip_tags(op.get_bind())
op.drop_column("import_settings", "wip_soft_title_tagging_enabled")
def downgrade():
op.add_column(
"import_settings",
sa.Column(
"wip_soft_title_tagging_enabled",
sa.Boolean(),
nullable=False,
server_default=sa.text("false"),
),
)
@@ -0,0 +1,60 @@
"""Re-date images whose post's date arrived after they were linked.
#4431. The native ingesters import a post's media before its record, and the
date travels in the record (`_post.json`). Every natively downloaded image was
therefore linked to an undated post and kept its download time in both gallery
date columns, while the post itself was dated correctly. The importer now
re-dates a post's images when its record lands; this repairs the images that
landed before that.
Both columns get back the rules the importer keeps:
* `effective_date` is the primary post's date (left alone when that post has
none, as the importer does);
* `earliest_post_date` is the earliest dated post the image is linked to.
Only rows that differ are written. The downgrade does nothing: the old values
were download times that no one chose.
Revision ID: 0113
Revises: 0112
Create Date: 2026-09-25
"""
import sqlalchemy as sa
from alembic import op
revision = "0113"
down_revision = "0112"
branch_labels = None
depends_on = None
def redate_images(conn) -> None:
"""The data step, on a plain connection, so a test can run it directly."""
conn.execute(sa.text("""
UPDATE image_record ir SET effective_date = p.post_date
FROM post p
WHERE p.id = ir.primary_post_id
AND p.post_date IS NOT NULL
AND ir.effective_date IS DISTINCT FROM p.post_date
"""))
conn.execute(sa.text("""
UPDATE image_record ir SET earliest_post_date = m.earliest
FROM (
SELECT ip.image_record_id, MIN(p.post_date) AS earliest
FROM image_provenance ip JOIN post p ON p.id = ip.post_id
WHERE p.post_date IS NOT NULL
GROUP BY ip.image_record_id
) m
WHERE m.image_record_id = ir.id
AND ir.earliest_post_date IS DISTINCT FROM m.earliest
"""))
def upgrade():
redate_images(op.get_bind())
def downgrade():
pass
@@ -0,0 +1,79 @@
"""Fold the undated shell posts a misfiled attachment created into the real post.
#4435. A non-image file (an archive, a pdf) downloaded for one of an artist's
Discord channels was filed under the artist's FIRST Discord source: the
attachment path looked the source up by (artist, platform), which takes the
lowest id. That created an undated, url-less post there holding only the
attachment, and the message's real post record then created the dated post
under the right source. The importer now uses the source it was downloading
for; this repairs the pairs it left.
A shell is folded only when all of this holds:
* it has no date and no url, and nothing synthesized it;
* another post of the same artist, platform and external id HAS a date;
* no image is linked to the shell (it held an attachment and nothing else).
Its attachments move to the dated post, dropping any the dated post already
has (same sha256, which the per-post unique forbids twice), and the shell is
deleted. A shell with no dated twin is left alone: which channel it belongs to
is not recorded anywhere but the file name. The downgrade does nothing.
Revision ID: 0114
Revises: 0113
Create Date: 2026-09-25
"""
import sqlalchemy as sa
from alembic import op
revision = "0114"
down_revision = "0113"
branch_labels = None
depends_on = None
_PAIRS = """
SELECT DISTINCT ON (shell.id) shell.id AS shell_id, real.id AS real_id
FROM post shell
JOIN source ss ON ss.id = shell.source_id
JOIN post real
ON real.artist_id = shell.artist_id
AND real.external_post_id = shell.external_post_id
AND real.id <> shell.id
AND real.post_date IS NOT NULL
JOIN source rs ON rs.id = real.source_id AND rs.platform = ss.platform
WHERE shell.post_date IS NULL
AND shell.post_url IS NULL
AND shell.synthesized_by IS NULL
AND NOT EXISTS (SELECT 1 FROM image_provenance ip WHERE ip.post_id = shell.id)
ORDER BY shell.id, real.id
"""
def fold_misfiled_attachment_posts(conn) -> int:
"""The data step, on a plain connection, so a test can run it directly.
Returns how many shells were folded."""
pairs = conn.execute(sa.text(_PAIRS)).all()
for shell_id, real_id in pairs:
conn.execute(sa.text("""
DELETE FROM post_attachment pa
WHERE pa.post_id = :shell
AND EXISTS (
SELECT 1 FROM post_attachment keep
WHERE keep.post_id = :real AND keep.sha256 = pa.sha256
)
"""), {"shell": shell_id, "real": real_id})
conn.execute(
sa.text("UPDATE post_attachment SET post_id = :real WHERE post_id = :shell"),
{"shell": shell_id, "real": real_id},
)
conn.execute(sa.text("DELETE FROM post WHERE id = :shell"), {"shell": shell_id})
return len(pairs)
def upgrade() -> None:
fold_misfiled_attachment_posts(op.get_bind())
def downgrade() -> None:
pass
@@ -0,0 +1,27 @@
"""source.display_name — what a source is called on its platform.
#4481. A Discord source's URL is a server id and a channel id, so the
Subscriptions page could only show two numbers. The Discord ingester already
loads the server and channel names to walk them; it now keeps them here,
refreshed on every walk. NULL until a walk has read it.
Revision ID: 0115
Revises: 0114
Create Date: 2026-09-28
"""
import sqlalchemy as sa
from alembic import op
revision = "0115"
down_revision = "0114"
branch_labels = None
depends_on = None
def upgrade() -> None:
op.add_column("source", sa.Column("display_name", sa.Text(), nullable=True))
def downgrade() -> None:
op.drop_column("source", "display_name")
@@ -0,0 +1,34 @@
"""The single-color filter's default becomes near-total: 0.95 → 0.995.
#4483. At 0.95 the filter rejected line art: a pencil doodle on white is
mostly white, and five of Todding's Discord doodles were skipped on import as
"single color", leaving their posts with text and no image. The predicate now
samples real pixels instead of a blurred thumbnail, and the default only
calls an image blank when it essentially is one.
A stored 0.95 is the old default, so it moves with it. Any other value was
chosen by hand and is left alone.
Revision ID: 0116
Revises: 0115
Create Date: 2026-09-28
"""
from alembic import op
revision = "0116"
down_revision = "0115"
branch_labels = None
depends_on = None
def upgrade() -> None:
op.alter_column("import_settings", "single_color_threshold", server_default="0.995")
op.execute(
"UPDATE import_settings SET single_color_threshold = 0.995 "
"WHERE single_color_threshold = 0.95"
)
def downgrade() -> None:
op.alter_column("import_settings", "single_color_threshold", server_default="0.95")
+60
View File
@@ -19,8 +19,10 @@ from ..models import AppSetting
from ..services.extension_service import (
ExtensionService,
InvalidUrlError,
PosterLookupError,
UnknownArtistError,
UnknownPlatformError,
UnknownSourceError,
)
from ..services.source_service import KNOWN_PLATFORMS
from ._responses import error_response as _bad
@@ -121,6 +123,10 @@ async def quick_add_source():
# Patreon is canon: adding a Patreon source to an existing artist can take
# the creator's Patreon display name (name only; the slug never moves).
use_platform_name = body.get("use_platform_name") is True
# #4488: a Discord add can carry the posters picked in the panel.
discord_authors = body.get("discord_authors")
if discord_authors is not None and not _valid_authors(discord_authors):
return _bad("invalid_body", detail="discord_authors must be a list of {id, name}")
from .credentials import _get_crypto
@@ -133,6 +139,7 @@ async def quick_add_source():
result = await ExtensionService(session, _get_crypto()).quick_add_source(
url, artist_id=artist_id, artist_name=artist_name,
use_platform_name=use_platform_name,
discord_authors=discord_authors,
)
except UnknownArtistError as exc:
return _bad("not_found", detail=str(exc), status=404)
@@ -147,6 +154,59 @@ async def quick_add_source():
return jsonify(result), (201 if result["created_source"] else 200)
def _valid_authors(value) -> bool:
"""[{id, name}]: the id is what the list keys on, the name only shows it."""
return isinstance(value, list) and all(
isinstance(a, dict) and isinstance(a.get("id"), str | int)
and not isinstance(a.get("id"), bool)
and (a.get("name") is None or isinstance(a.get("name"), str))
for a in value
)
@extension_bp.route("/discord/posters", methods=["GET"])
async def discord_posters():
"""Who posts in the Discord channel `url` shows, lately, with the likely
creator marked (#4488). Reads the channel with the stored token; writes
nothing."""
url = (request.args.get("url") or "").strip()
if not url:
return _bad("invalid_body", detail="url query parameter is required")
from .credentials import _get_crypto
async with get_session() as session:
if not await _ext_key_required(session):
return _bad("unauthorized", status=401)
try:
result = await ExtensionService(session, _get_crypto()).discord_posters(url)
except (InvalidUrlError, UnknownPlatformError) as exc:
return _bad("invalid_url", detail=str(exc))
except PosterLookupError as exc:
return _bad("posters_unavailable", detail=str(exc), status=409)
return jsonify(result)
@extension_bp.route("/discord/posters", methods=["POST"])
async def set_discord_posters():
"""Set the poster list of the source following `url`'s channel."""
body = await request.get_json(silent=True)
if not isinstance(body, dict) or not isinstance(body.get("url"), str):
return _bad("invalid_body", detail="url is required")
authors = body.get("authors")
if not _valid_authors(authors):
return _bad("invalid_body", detail="authors must be a list of {id, name}")
async with get_session() as session:
if not await _ext_key_required(session):
return _bad("unauthorized", status=401)
try:
result = await ExtensionService(session).set_discord_posters(body["url"], authors)
except (InvalidUrlError, UnknownPlatformError, PosterLookupError) as exc:
return _bad("invalid_url", detail=str(exc))
except UnknownSourceError as exc:
return _bad("not_found", detail=str(exc), status=404)
return jsonify(result)
def _read_manifest_sync() -> dict | None:
"""All the filesystem-touching work for /api/extension/manifest,
in a sync helper so the async route can dispatch it via
-7
View File
@@ -57,7 +57,6 @@ _EDITABLE_FIELDS = (
"translation_target_lang",
"translation_min_confidence",
"wip_title_tagging_enabled",
"wip_soft_title_tagging_enabled",
)
# Per-host external-download toggles — all plain booleans, validated uniformly.
@@ -196,12 +195,6 @@ async def update_import_settings():
return jsonify(
{"error": "wip_title_tagging_enabled must be a boolean"}
), 400
if "wip_soft_title_tagging_enabled" in body and not isinstance(
body["wip_soft_title_tagging_enabled"], bool
):
return jsonify(
{"error": "wip_soft_title_tagging_enabled must be a boolean"}
), 400
async with get_session() as session:
row = await ImportSettings.load(session)
+49
View File
@@ -1,10 +1,13 @@
"""FC-3a: CRUD over Source rows. FC-3c adds POST /<id>/check."""
from pathlib import Path
from quart import Blueprint, jsonify, request
from sqlalchemy import func, select
from ..extensions import get_session
from ..models import DownloadEvent, MembershipSync, PlatformMembership, Source
from ..services import discord_poster_cleanup as poster_cleanup
from ..services.artist_membership_service import ArtistMembershipService
from ..services.artist_membership_service import rescan as membership_rescan
from ..services.artist_service import ArtistService
@@ -463,3 +466,49 @@ async def adopt_membership():
return jsonify({"already_tracked": exc.existing_id})
artist_id = artist.id
return jsonify({"source_id": record.id, "artist_id": artist_id}), 201
# -- Discord: remove posts by people outside the poster list (#4486) --------------
_IMAGES_ROOT = Path("/images")
async def _poster_cleanup_preview(source_id: int):
async with get_session() as session:
return await session.run_sync(
lambda s: poster_cleanup.preview(s, source_id=source_id)
)
@sources_bp.route("/<int:source_id>/discord/other-posters", methods=["GET"])
async def other_posters_preview(source_id: int):
"""What removing other posters' posts would delete. Nothing is touched."""
try:
projection = await _poster_cleanup_preview(source_id)
except LookupError:
return _bad("not_found", status=404)
except poster_cleanup.PosterCleanupError as exc:
return _bad("not_applicable", detail=str(exc), status=409)
projection["confirm_token"] = poster_cleanup.confirm_token(projection)
return jsonify(projection)
@sources_bp.route("/<int:source_id>/discord/other-posters/remove", methods=["POST"])
async def other_posters_remove(source_id: int):
"""Delete them. `confirm` must be the token of the preview the operator saw:
a poster list edited since then changes the set, and the token with it."""
body = await request.get_json(silent=True) or {}
try:
projection = await _poster_cleanup_preview(source_id)
except LookupError:
return _bad("not_found", status=404)
except poster_cleanup.PosterCleanupError as exc:
return _bad("not_applicable", detail=str(exc), status=409)
expected = poster_cleanup.confirm_token(projection)
if body.get("confirm") != expected:
return _bad("confirm_mismatch", expected=expected)
async with get_session() as session:
result = await session.run_sync(
lambda s: poster_cleanup.apply(s, source_id=source_id, images_root=_IMAGES_ROOT)
)
return jsonify(result)
+5
View File
@@ -152,6 +152,8 @@ async def list_runs():
queue=<name> filter to one queue
status=<status> filter to one status (running/ok/error/timeout/retry)
task=<substr> case-insensitive substring match on task_name
celery_task_id=<id> exactly one run — how a page follows a job it
started without having to know its lane
limit=<int> default 50, max 200
before_id=<int> cursor for keyset pagination
@@ -167,6 +169,7 @@ async def list_runs():
queue = request.args.get("queue")
status = request.args.get("status")
task = request.args.get("task")
celery_task_id = request.args.get("celery_task_id")
before_id_raw = request.args.get("before_id")
before_id = int(before_id_raw) if before_id_raw else None
@@ -176,6 +179,8 @@ async def list_runs():
stmt = stmt.where(TaskRun.queue == queue)
if status:
stmt = stmt.where(TaskRun.status == status)
if celery_task_id:
stmt = stmt.where(TaskRun.celery_task_id == celery_task_id)
if task:
# Task names contain literal underscores (download_source,
# vacuum_analyze) — escape LIKE wildcards so a search for
+7 -5
View File
@@ -69,6 +69,8 @@ def make_celery() -> Celery:
# up behind it (2026-09-24: 7 waiting, "all workers busy for 18
# minutes"). An exact name wins over the glob above.
"backend.app.tasks.maintenance.backfill_phash": {"queue": "maintenance_long"},
# Walks a folder tree per artist with undated posts (#4436).
"backend.app.tasks.maintenance.date_posts_from_records": {"queue": "maintenance_long"},
"backend.app.tasks.backup.*": {"queue": "maintenance_long"},
"backend.app.tasks.admin.*": {"queue": "maintenance_long"},
"backend.app.tasks.library_audit.*": {"queue": "maintenance_long"},
@@ -156,6 +158,11 @@ def make_celery() -> Celery:
"schedule": 86400.0, # daily — sweep .part/.partial left by a
# download/import killed mid-write (graceful-shutdown fallout)
},
"date-posts-from-records-hourly": {
"task": "backend.app.tasks.maintenance.date_posts_from_records",
"schedule": 3600.0, # an empty query once every native post is
# dated; otherwise upserts the _post.json a killed walk left (#4436)
},
"backfill-phash-daily": {
"task": "backend.app.tasks.maintenance.backfill_phash",
"schedule": 86400.0, # daily — NULL-only, so a no-op once the
@@ -223,11 +230,6 @@ def make_celery() -> Celery:
"schedule": 86400.0, # auto-tag wip/editor process art (#1464);
# no-op unless process_auto_apply_enabled (opt-in)
},
"soft-wip-conflict-audit-daily": {
"task": "backend.app.tasks.ml.scheduled_soft_wip_conflict_audit",
"schedule": 86400.0, # flag ring-loud soft-WIP (sketch/doodle) tags
# for review (#1474); no-op with no content heads
},
"prune-presentation-reviews-daily": {
"task": "backend.app.tasks.ml.prune_presentation_reviews",
"schedule": 86400.0, # retention: drop resolved review flags >30d
+67 -33
View File
@@ -18,11 +18,18 @@ dark for that interval. Monitoring NEVER breaks the thing it's
monitoring.
"""
import functools
import logging
from datetime import UTC, datetime
from celery.exceptions import SoftTimeLimitExceeded
from celery.signals import task_failure, task_postrun, task_prerun, task_retry
from celery.signals import (
task_failure,
task_postrun,
task_prerun,
task_retry,
worker_ready,
)
from .models import TaskRun
from .tasks._sync_engine import sync_session_factory
@@ -53,42 +60,29 @@ _INT32_MIN = -2_147_483_648
def _queue_for(task) -> str:
"""Reverse the task→queue routing from celery_app.task_routes.
Keep in sync if task_routes is reordered.
"""The queue Celery routes this task to — asked of the router itself.
Audit 2026-06-02: backup/admin/library_audit prefixes were
missing here even though task_routes sent all three to
'maintenance'. The TaskRun.queue column then lied for those
rows (claimed 'default') so per-queue dashboard filters and
per-queue threshold overrides silently missed them.
This was a hand-kept copy of `celery_app.task_routes`, and it drifted
twice (the 2026-06-02 audit, then #4432). Long-lane jobs were recorded as
`maintenance`, and translation and gpu_queue runs as `default`, where the
5-minute stall sweep failed healthy 35-minute translation runs. The router
answers from the same table the broker uses, so the two cannot disagree.
"""
name = getattr(task, "name", "") or ""
if name.startswith("backend.app.tasks.import_file."):
return "import"
if name.startswith("backend.app.tasks.ml."):
return "ml"
if name.startswith("backend.app.tasks.thumbnail."):
return "thumbnail"
if name.startswith((
"backend.app.tasks.download.",
# External file-host fetches share the download lane (celery_app
# routes external.* → download). Mirror it here or TaskRun.queue
# lies 'default' for them, so per-queue dashboard filters and the
# per-queue threshold override miss them — the same gap the
# 2026-06-02 audit fixed for backup/admin/library_audit.
"backend.app.tasks.external.",
)):
return "download"
if name.startswith("backend.app.tasks.scan."):
return "scan"
if name.startswith((
"backend.app.tasks.maintenance.",
"backend.app.tasks.backup.",
"backend.app.tasks.admin.",
"backend.app.tasks.library_audit.",
)):
return "maintenance"
app = getattr(task, "app", None)
if app is None:
from .celery_app import celery as app
return _routed_queue(app, name)
@functools.lru_cache(maxsize=1024)
def _routed_queue(app, name: str) -> str:
try:
queue = app.amqp.router.route({}, name).get("queue")
except Exception: # noqa: BLE001 — monitoring never breaks the task
log.warning("task_run: could not resolve the queue for %s", name)
return "default"
return getattr(queue, "name", None) or (queue if isinstance(queue, str) else "default")
def _target_id_from_args(args) -> int | None:
@@ -225,3 +219,43 @@ def _on_retry(sender=None, request=None, reason=None, einfo=None, **_):
error_message=str(reason) if reason else None,
retry_count=getattr(request, "retries", 0),
)
def _consumed_queues(consumer) -> set[str]:
"""The queue names this worker process consumes (its `-Q`), or empty when
they can't be read — which makes the boot hook below a no-op, not a guess."""
try:
return {q.name for q in consumer.task_consumer.queues}
except Exception: # noqa: BLE001 — shape varies across celery versions
return set()
@worker_ready.connect
def _on_worker_ready(sender=None, **_):
"""The download lane clears what its previous process left behind (#4433).
A restart SIGKILLs any walk that outlives the stop grace, so its event is
never finalized and its platform lock is never released. Both used to wait
out timers — a 30-min stall sweep that then blamed the source, and a 27-min
lock TTL that stalled every other source on the platform. Only the process
consuming `download` does this; the other lanes booting beside it must not.
Best-effort: a failure here is logged and the worker still starts.
"""
if "download" not in _consumed_queues(sender):
return
booted_at = datetime.now(UTC)
try:
from .services.download_recovery import interrupt_orphaned_download_events
from .services.platform_lock import release_all_platform_locks
with sync_session_factory()() as session:
closed = interrupt_orphaned_download_events(session, booted_at=booted_at)
session.commit()
released = release_all_platform_locks()
if closed or released:
log.info(
"download lane boot: closed %d orphaned download event(s) as "
"interrupted, released %d platform lock(s)", closed, released,
)
except Exception: # noqa: BLE001 — never block the worker from starting
log.exception("download lane boot recovery failed")
+20 -2
View File
@@ -46,6 +46,14 @@ async def serve_extension(filename: str):
The application/x-xpinstall MIME tells Firefox to show its native
install prompt instead of downloading the file as a blob.
Caching differs by name, and has to. A versioned name is one build's bytes
forever, so it can be cached for good. `fabledcurator-latest.xpi` is ONE
URL whose bytes change on every release, and Quart's default for a file
is `public, max-age=43200`: a browser that fetched it once reused those
bytes for 12 hours, so "install the latest" quietly reinstalled the
previous build (operator-flagged 2026-09-25). It is `no-cache` — the ETag
still makes an unchanged file a cheap 304.
"""
if not _XPI_NAME_RE.fullmatch(filename):
abort(404)
@@ -56,10 +64,11 @@ async def serve_extension(filename: str):
if not xpis:
abort(404)
latest = xpis[-1]
return await send_file(
resp = await send_file(
latest, mimetype="application/x-xpinstall",
attachment_filename=latest.name,
)
return _cache(resp, "no-cache")
target = (XPI_DIR / filename).resolve()
try:
target.relative_to(XPI_DIR)
@@ -67,10 +76,19 @@ async def serve_extension(filename: str):
abort(404)
if not target.is_file():
abort(404)
return await send_file(
resp = await send_file(
target, mimetype="application/x-xpinstall",
attachment_filename=filename,
)
return _cache(resp, "public, max-age=31536000, immutable")
def _cache(resp, policy: str):
"""Set the XPI's Cache-Control, dropping the Expires send_file adds so the
two can never disagree."""
resp.headers["Cache-Control"] = policy
resp.headers.pop("Expires", None)
return resp
@frontend_bp.route("/")
+1 -8
View File
@@ -39,7 +39,7 @@ class ImportSettings(Base):
transparency_threshold: Mapped[float] = mapped_column(Float, nullable=False, default=0.9, server_default="0.9")
skip_single_color: Mapped[bool] = mapped_column(Boolean, nullable=False, default=False, server_default="false")
single_color_threshold: Mapped[float] = mapped_column(Float, nullable=False, default=0.95, server_default="0.95")
single_color_threshold: Mapped[float] = mapped_column(Float, nullable=False, default=0.995, server_default="0.995")
single_color_tolerance: Mapped[int] = mapped_column(Integer, nullable=False, default=30, server_default="30")
# Hamming distance over a 256-bit pHash (utils.phash, hash_size=16). The
@@ -238,13 +238,6 @@ class ImportSettings(Base):
wip_title_tagging_enabled: Mapped[bool] = mapped_column(
Boolean, nullable=False, default=True, server_default="true",
)
# Soft WIP title tier (#1474): also tag sketch/doodle/scribble titles, but with
# a PROVISIONAL source (`wip_title_soft`) that never trains the head, since these
# are lower-precision (a finished "sketch" isn't WIP). OFF by default — a lower-
# precision tier is opt-in (the ring-loud audit surfaces false positives).
wip_soft_title_tagging_enabled: Mapped[bool] = mapped_column(
Boolean, nullable=False, default=False, server_default="false",
)
@classmethod
async def load(cls, session) -> ImportSettings:
+5
View File
@@ -46,6 +46,11 @@ class Source(Base):
enabled: Mapped[bool] = mapped_column(Boolean, nullable=False, default=True, server_default="true")
config_overrides: Mapped[dict | None] = mapped_column(JSON, nullable=True)
# alembic 0115: what the source is called on its platform, where the URL
# alone doesn't say — a Discord link is two numbers. The ingester that
# walks it refreshes it every walk, so a rename reads as the new name.
# NULL until a walk has read it, and on platforms that don't write it.
display_name: Mapped[str | None] = mapped_column(Text, nullable=True)
last_checked_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True), nullable=True)
last_error: Mapped[str | None] = mapped_column(Text, nullable=True)
+12 -3
View File
@@ -4,11 +4,19 @@ predicate for BOTH surfaces: FC-Cleanup's retroactive audit and — since
2026-07-02 — the import-side filter (Importer._single_color_hit /
SkipReason.single_color), so what the audit flags and what the import
skips can never disagree.
It is meant to catch the blank: a placeholder, an error tile, a solid fill.
Line art is the case it must not catch (#4483): a pencil doodle on white can be
well over 95% white even at full size, and a smoothing downsample blends its strokes
into the paper until it measures as blank. So the sample is taken by NEAREST
(each sampled pixel is a real pixel, strokes keep their contrast), at 256px
so thin strokes are still hit, and the default threshold is near-total
(0.995) — a few percent of ink is a drawing, not an empty image.
"""
from PIL import Image
_THUMB_SIZE = (64, 64)
_THUMB_SIZE = (256, 256)
def evaluate(
@@ -20,7 +28,8 @@ def evaluate(
"""True iff the fraction of pixels within `tolerance` (Euclidean RGB
distance) of the dominant color exceeds `threshold`.
Downsamples to 64x64 for speed (~4ms regardless of source size).
Samples 256x256 by NEAREST (see the module docstring for why not a
smoothing resample).
Alpha channels are stripped; only RGB is considered. Animated images
use frame 0 (PIL's default after Image.open without seek).
"""
@@ -30,7 +39,7 @@ def evaluate(
elif im.mode not in ("RGB", "L"):
im = im.convert("RGB")
if im.size != _THUMB_SIZE:
im = im.resize(_THUMB_SIZE, Image.Resampling.BILINEAR)
im = im.resize(_THUMB_SIZE, Image.Resampling.NEAREST)
pixels = list(im.getdata())
if not pixels:
return False
+200 -10
View File
@@ -32,6 +32,13 @@ Two deliberate departures, both about the walk order, neither about content:
the walk. gallery-dl skips only nested channels; one private thread the
token cannot read would otherwise stop every channel after it.
And one about content: a message is taken only if it has an image or video
ATTACHMENT (`has_visual_attachment`). Discord is where creators chat as well
as post, and the operator wants the art, not the conversation (2026-09-28): a
text line, a lone archive or PSD, a pasted link's preview or a Tenor GIF is
chat. A message that passes keeps every file gallery-dl would take, numbered
as gallery-dl numbers them, so the on-disk names still match.
FC runs on a plain-HTTP homelab; nothing here uses a secure-context Web API.
"""
@@ -68,6 +75,9 @@ _THREADS_BATCH = 25
# it is honoured when present; 60s is the fallback and the cap.
_MAX_429_RETRIES = 4
_429_WAIT_SECONDS = 60.0
# How far back the poster picker looks: two pages. Enough to see who posts in
# a channel, few enough to answer while the operator waits (#4488).
_POSTER_SCAN = 200
# https://discord.com/developers/docs/resources/message#message-object-message-types
# DEFAULT, REPLY, CHAT_INPUT_COMMAND — the ones that carry user content.
@@ -79,6 +89,13 @@ _FORUM = frozenset({15, 16}) # forum, media: threads only
_CATEGORY = 4
_SERVER_WALK = _TEXT | _FORUM
_EMBED_TYPES = frozenset({"image", "gifv", "video"})
# What makes a message art rather than chat: an attached file FC imports as an
# image or video. Mirrors importer.ALL_EXTS (not imported: that module is the
# import pipeline, and this one is a network client).
_VISUAL_EXTS = frozenset({
"png", "jpg", "jpeg", "gif", "webp", "bmp", "tif", "tiff",
"mp4", "mov", "avi", "mkv", "webm", "m4v", "wmv", "flv",
})
_URL_RE = re.compile(
r"^(?:https?://)?(?:www\.|ptb\.|canary\.)?discord(?:app)?\.com/channels/"
@@ -163,6 +180,82 @@ def message_text(message: dict) -> str:
return "\n".join(p for p in parts if p)
def _message_and_snapshots(message: dict) -> list[dict]:
"""The message itself, then each forwarded message it carries."""
return [message] + [
(s or {}).get("message") or {}
for s in message.get("message_snapshots") or []
if ((s or {}).get("message") or {}).get("type", 0) in MESSAGE_TYPES
]
def _is_visual(attachment: dict) -> bool:
ctype = (attachment.get("content_type") or "").lower()
if ctype.startswith(("image/", "video/")):
return True
name = attachment.get("filename") or attachment.get("url") or ""
return nameext_from_url(name)[1] in _VISUAL_EXTS
def has_visual_attachment(message: dict) -> bool:
"""Does this message, or a message it forwards, have an image or video
attached? The one test for "art, not chat" — see the module docstring."""
return any(
att.get("url") and _is_visual(att)
for snap in _message_and_snapshots(message)
for att in snap.get("attachments") or []
)
_SQUASH = re.compile(r"[\W_]+", re.UNICODE)
# Words a server name wraps around its creator's: "Todding's Server",
# "The Official Todding Discord".
_SERVER_WORDS = frozenset({"server", "discord", "the", "official", "community", "s"})
def _squash(text: str | None) -> str:
return _SQUASH.sub("", (text or "").lower())
def _server_tokens(server_name: str | None) -> list[str]:
name = (server_name or "").lower().replace("'s", " ").replace("\u2019s", " ")
words = [w for w in _SQUASH.split(name) if w and w not in _SERVER_WORDS]
tokens = [w for w in words if len(w) >= 3]
joined = "".join(words)
if len(joined) >= 3 and joined not in tokens:
tokens.append(joined)
return tokens
def rank_posters(server_name: str | None, owner_id, posters: list[dict]) -> list[dict]:
"""Mark who is probably the creator a server is named for, and order the
list for picking: suggested first, then by images posted, then messages.
Two signals, each a reason the picker shows beside the name:
- they own the server;
- their username or display name matches the server's name, once
"'s Server", "Official", "Discord" and punctuation are taken off
("Todding's Server" and "todding").
Neither signal alone is proof, which is why this suggests and never
decides: the operator ticks the list. Bots are never suggested."""
tokens = _server_tokens(server_name)
owner = str(owner_id) if owner_id else None
ranked = []
for p in posters:
reasons = []
if owner and p.get("id") == owner:
reasons.append("owns the server")
names = [n for n in (_squash(p.get("username")), _squash(p.get("global_name"))) if len(n) >= 3]
if any(t in n or n in t for t in tokens for n in names):
reasons.append("name matches the server")
ranked.append({**p, "reasons": reasons, "suggested": False})
best = max((len(r["reasons"]) for r in ranked if not r.get("bot")), default=0)
for r in ranked:
r["suggested"] = best > 0 and not r.get("bot") and len(r["reasons"]) == best
ranked.sort(key=lambda r: (not r["suggested"], -r.get("images", 0), -r.get("messages", 0)))
return ranked
@dataclass
class MediaItem:
"""One file of a Discord message. `filename`/`extension` are gallery-dl's
@@ -212,6 +305,8 @@ class DiscordClient:
self._server: dict = {}
self._channels: dict[str, dict] = {}
self._skip_feed = False
# Whose messages the walk takes (`only_from`); empty takes everyone's.
self._authors: frozenset[str] = frozenset()
# -- request -----------------------------------------------------------
@@ -373,6 +468,43 @@ class DiscordClient:
if meta["channel_type"] in _SERVER_WALK:
yield from self._feeds(meta["channel_id"], safe=True)
def only_from(self, authors) -> None:
"""Take only messages posted by these people: user ids, usernames or
display names, any case. A creator's server is full of other members
posting their own pictures; the source is subscribed to the creator.
Empty or None takes everyone's, as before."""
self._authors = frozenset(
str(a).strip().lower() for a in authors or () if str(a).strip()
)
def _author_wanted(self, message: dict) -> bool:
if not self._authors:
return True
author = message.get("author") or {}
names = (author.get("id"), author.get("username"), author.get("global_name"))
return any(str(n).lower() in self._authors for n in names if n)
def source_label(self, url: str) -> str | None:
"""What the source is called in Discord — `Server · #channel`, a
thread as `#parent › thread`, a whole server by its name — from the
metadata the walk loaded. None when the walk never got that far."""
try:
_, channel_id = parse_source_url(url)
except DiscordAPIError:
return None
server = self._server.get("server") or None
if channel_id is None:
return server
meta = self._channels.get(channel_id) or {}
name = meta.get("channel")
if not name:
return None
if name != "DMs":
name = f"#{name}"
if meta.get("is_thread") and meta.get("parent"):
name = f"#{meta['parent']} › {meta['channel']}"
return f"{server} · {name}" if server else name
def skip_feed(self) -> None:
"""Optional core seam (#4413): end the current channel and go on to the
next one. A tick's early-out means THIS channel has nothing new, not
@@ -440,6 +572,8 @@ class DiscordClient:
for message in messages:
if message.get("type") not in MESSAGE_TYPES:
continue
if not self._author_wanted(message):
continue
message["_meta"] = meta
yield message, meta, page_cursor
if self._skip_feed:
@@ -454,13 +588,14 @@ class DiscordClient:
def extract_media(post: dict, included: dict | None = None) -> list[MediaItem]:
"""gallery-dl's file list for one message: attachments, then the first
of video/image/thumbnail `proxy_url` of each file-bearing embed, then
the same for every forwarded snapshot; numbered from 1 across them."""
the same for every forwarded snapshot; numbered from 1 across them.
Empty for a message with no image or video attached: that is chat,
whatever else it carries (`has_visual_attachment`)."""
if not has_visual_attachment(post):
return []
mid = str(post.get("id") or "")
snapshots = [post] + [
(s or {}).get("message") or {}
for s in post.get("message_snapshots") or []
if ((s or {}).get("message") or {}).get("type", 0) in MESSAGE_TYPES
]
snapshots = _message_and_snapshots(post)
found: list[tuple[str, str, str | None]] = []
for snap in snapshots:
for att in snap.get("attachments") or []:
@@ -498,10 +633,10 @@ class DiscordClient:
"""`(message:<id>, <id>)` — gates the message record through the seen
ledger, like `post:<id>` on the other platforms.
None for a message with no files. gallery-dl wrote a sidecar only
beside a file, so a text-only chat line never became a post, and the
drop grouping (discord_grouping) is built on that: a channel's chatter
recorded as posts would bury the drops it exists to surface."""
None for a message `extract_media` takes nothing from, i.e. one with
no image or video attached. The drop grouping (discord_grouping) is
built on that: a channel's chatter recorded as posts would bury the
drops it exists to surface."""
mid = post.get("id")
mid = str(mid) if mid is not None else ""
if not mid or not cls.extract_media(post):
@@ -528,6 +663,61 @@ class DiscordClient:
pass
return out
def recent_posters(
self, server_id: str | None, channel_id: str, *, max_messages: int = _POSTER_SCAN,
) -> dict:
"""Who has posted lately in a channel, for picking a source's poster
list (#4488): the newest `max_messages` messages, tallied per author.
Deliberately shallow — it answers "who posts here", not "who ever did".
Every poster comes back with their stable id, which is what the list
stores: a username can change on a whim, an id can't."""
server: dict = {}
if server_id:
try:
server = self._get(f"/guilds/{server_id}") or {}
except DiscordAPIError:
server = {}
tally: dict[str, dict] = {}
before = None
seen = 0
while seen < max_messages:
page = self._get(
f"/channels/{channel_id}/messages",
{"limit": min(_MESSAGES_BATCH, max_messages - seen), "before": before},
)
if not isinstance(page, list) or not page:
break
for message in page:
seen += 1
if message.get("type") not in MESSAGE_TYPES:
continue
author = message.get("author") or {}
aid = str(author.get("id") or "")
if not aid:
continue
row = tally.setdefault(aid, {
"id": aid,
"username": author.get("username"),
"global_name": author.get("global_name"),
"bot": bool(author.get("bot")),
"messages": 0,
"images": 0,
})
row["messages"] += 1
if has_visual_attachment(message):
row["images"] += 1
if len(page) < _MESSAGES_BATCH:
break
before = str(page[-1]["id"])
return {
"server": server.get("name") or None,
"owner_id": str(server["owner_id"]) if server.get("owner_id") else None,
"scanned": seen,
"posters": rank_posters(
server.get("name"), server.get("owner_id"), list(tally.values()),
),
}
def verify_auth(self, url: str) -> tuple[bool | None, str]:
"""Is the token valid, and can it see what the source names?"""
try:
@@ -195,6 +195,7 @@ class DiscordDownloader(BaseNativeDownloader):
"parent": meta.get("parent"),
"is_thread": meta.get("is_thread"),
"author": author.get("username"),
"author_name": author.get("global_name"),
"author_id": author.get("id"),
"message": body,
"date": post.get("timestamp"),
+54 -1
View File
@@ -13,6 +13,11 @@ Two things differ from the cookie platforms:
something; a Discord drop is routinely files and nothing else, so on
Discord that is an ordinary backfill, not a broken parser.
And two things only a Discord source has: `discord_authors` in its
config_overrides limits it to the creator's own messages (a server's other
members post pictures too), and it learns its own name — the server and
channel it walks — into `source.display_name`.
`campaign_id` is the source URL (a server, channel, thread or category link).
FC runs on a plain-HTTP homelab; nothing here uses a secure-context Web API.
"""
@@ -20,15 +25,23 @@ FC runs on a plain-HTTP homelab; nothing here uses a secure-context Web API.
from __future__ import annotations
import asyncio
import logging
from collections.abc import Callable
from pathlib import Path
from ..models import DiscordFailedMedia, DiscordSeenMedia
from sqlalchemy import select, update
from ..models import DiscordFailedMedia, DiscordSeenMedia, Source
from .discord_client import DiscordAPIError, DiscordClient, MediaItem
from .discord_downloader import DiscordDownloader
from .ingest_core import Ingester
log = logging.getLogger(__name__)
_LEDGER_KEY_MAX = 128
# `config_overrides` key: the people whose messages this source takes (user
# ids, usernames or display names). Absent or empty takes everyone's.
AUTHORS_KEY = "discord_authors"
def _ledger_key(media: MediaItem) -> str:
@@ -74,6 +87,46 @@ class DiscordIngester(Ingester):
body_canary=False,
)
def run(self, **kwargs):
"""The core walk, bracketed by the two things only a Discord source
has: whose messages it takes, read before the walk, and the name of
what it walks, written after it (the walk is what loads that name)."""
source_id = kwargs["source_id"]
self.client.only_from(self._source_authors(source_id))
try:
return super().run(**kwargs)
finally:
self._record_display_name(source_id, kwargs["campaign_id"])
def _source_authors(self, source_id: int) -> list:
if self.session_factory is None:
return []
with self.session_factory() as session:
overrides = session.execute(
select(Source.config_overrides).where(Source.id == source_id)
).scalar_one_or_none() or {}
authors = overrides.get(AUTHORS_KEY) or []
return authors if isinstance(authors, list) else []
def _record_display_name(self, source_id: int, url: str) -> None:
"""Refreshed on every walk, never written once: a renamed channel
should read as its new name. A name the walk couldn't read leaves the
stored one alone, and a failure here never fails the walk."""
label = self.client.source_label(url)
if not label or self.session_factory is None:
return
try:
with self.session_factory() as session:
session.execute(
update(Source)
.where(Source.id == source_id)
.where(Source.display_name.is_distinct_from(label))
.values(display_name=label)
)
session.commit()
except Exception as exc: # a name is decoration — never fail the walk
log.warning("Discord: couldn't record source %s's name: %s", source_id, exc)
async def verify_discord_credential(url: str, auth_token: str | None) -> tuple[bool | None, str]:
"""The uniform `(ok, message)` probe: is the token valid, and can its
@@ -0,0 +1,243 @@
"""Remove a Discord source's posts by people outside its poster list (#4486).
A creator's server is full of other members posting their own pictures. The
poster list (`discord_authors`, #4481) stops a walk taking them; this removes
the ones taken before the list existed. The operator sets the list first, then
previews what would go, then applies.
## What goes
A post of this source whose record names its poster (`author_id` / `author`,
written by `DiscordDownloader.write_post_record`) and names someone NOT on the
list. Two kinds of post are never touched:
* a post whose record names no poster — a message recorded before the native
ingester, whose poster FC cannot tell. Counted as `unknown`, never guessed;
* a synthetic drop post (`synthesized_by`) — FC authored it, and it is handled
below as a consequence, not matched as a poster's post.
An image goes when every REAL post it belongs to is going. A link to a
synthetic drop does not keep an image: the drop only references its members'
images. An image also on a kept post — the creator re-posting a piece someone
else shared — stays.
A drop that absorbed a removed message is deleted outright. Deleting a drop is
its undo (discord_grouping's honesty rule): its remaining members return to
the feed and the grouping sweep regroups them on its next run, so no half-
rebuilt drop keeps text or thumbnails from someone it no longer contains.
Preview and apply spread the same predicates (rule 93, snippet #3087).
"""
from __future__ import annotations
import logging
from pathlib import Path
from sqlalchemy import and_, delete, exists, func, or_, select
from sqlalchemy.orm import Session, aliased
from ..models import ImageProvenance, ImageRecord, Post, PostAttachment, Source
from .cleanup_service import delete_images
from .discord_ingester import AUTHORS_KEY
log = logging.getLogger(__name__)
PLATFORM = "discord"
# The record keys that name a message's poster. `author_name` (the display
# name) is only on records written after 2026-09-28.
_POSTER_KEYS = ("author_id", "author", "author_name")
class PosterCleanupError(ValueError):
"""The source can't be cleaned this way (not Discord, or no list)."""
def source_authors(source: Source) -> list[str]:
"""The source's poster list, lowercased; empty means everyone's wanted."""
authors = (source.config_overrides or {}).get(AUTHORS_KEY) or []
if not isinstance(authors, list):
return []
return sorted({str(a).strip().lower() for a in authors if str(a).strip()})
def _poster(key: str):
return func.lower(Post.raw_metadata[key].as_string())
def _has_poster():
return or_(*(Post.raw_metadata[k].as_string().isnot(None) for k in ("author_id", "author")))
def _real_post_conditions(source_id: int) -> list:
return [Post.source_id == source_id, Post.synthesized_by.is_(None)]
def _named_poster_conditions(source_id: int) -> list:
"""A real post of this source whose record says who posted it."""
return [*_real_post_conditions(source_id), _has_poster()]
def _other_poster_post_conditions(source_id: int, authors: list[str]) -> list:
"""The posts that go: a named poster none of whose names is on the list."""
return [
*_named_poster_conditions(source_id),
*(func.coalesce(_poster(k), "").notin_(authors) for k in _POSTER_KEYS),
]
def _doomed_post_ids(source_id: int, authors: list[str]):
# correlate(None): used inside queries that are themselves over `post` (the
# drops, the delete), where auto-correlation would bind this to the outer
# row instead of scanning the table.
return (
select(Post.id)
.where(*_other_poster_post_conditions(source_id, authors))
.correlate(None)
)
def _image_conditions(source_id: int, authors: list[str]) -> list:
"""Images that belong to a removed post and to no kept real post."""
doomed = _doomed_post_ids(source_id, authors)
keeper = aliased(Post)
on_doomed = or_(
ImageRecord.primary_post_id.in_(doomed),
exists().where(
ImageProvenance.image_record_id == ImageRecord.id,
ImageProvenance.post_id.in_(doomed),
),
)
on_kept = or_(
and_(
ImageRecord.primary_post_id.isnot(None),
ImageRecord.primary_post_id.notin_(doomed),
),
exists().where(
ImageProvenance.image_record_id == ImageRecord.id,
ImageProvenance.post_id == keeper.id,
keeper.synthesized_by.is_(None),
keeper.id.notin_(doomed),
),
)
return [on_doomed, ~on_kept]
def _drop_conditions(source_id: int, authors: list[str]) -> list:
"""The synthetic drops that absorbed a removed message."""
member = aliased(Post)
return [
Post.source_id == source_id,
Post.synthesized_by.isnot(None),
exists().where(
member.absorbed_by_post_id == Post.id,
member.id.in_(_doomed_post_ids(source_id, authors)),
),
]
def _attachment_conditions(source_id: int, authors: list[str]) -> list:
return [PostAttachment.post_id.in_(_doomed_post_ids(source_id, authors))]
def _load(session: Session, source_id: int) -> tuple[Source, list[str]]:
source = session.get(Source, source_id)
if source is None:
raise LookupError(source_id)
if source.platform != PLATFORM:
raise PosterCleanupError("Only a Discord source has posters to filter by.")
authors = source_authors(source)
if not authors:
raise PosterCleanupError(
"Set this source's 'Only posts by' list first — with no list, "
"every poster is wanted and nothing would be removed."
)
return source, authors
def _count(session: Session, stmt) -> int:
return session.execute(stmt).scalar_one()
def preview(session: Session, *, source_id: int) -> dict:
"""What `apply` would remove, per poster, without touching anything."""
_, authors = _load(session, source_id)
who = func.coalesce(
Post.raw_metadata["author_name"].as_string(),
Post.raw_metadata["author"].as_string(),
Post.raw_metadata["author_id"].as_string(),
)
rows = session.execute(
select(who, func.count(Post.id))
.where(*_other_poster_post_conditions(source_id, authors))
.group_by(who)
.order_by(func.count(Post.id).desc())
).all()
posters = [{"poster": name, "posts": n} for name, n in rows]
unknown = _count(session, select(func.count(Post.id)).where(
*_real_post_conditions(source_id), ~_has_poster(),
))
return {
"authors": authors,
"posters": posters,
"posts": sum(p["posts"] for p in posters),
"images": _count(session, select(func.count(ImageRecord.id)).where(
*_image_conditions(source_id, authors))),
"attachments": _count(session, select(func.count(PostAttachment.id)).where(
*_attachment_conditions(source_id, authors))),
"drops": _count(session, select(func.count(Post.id)).where(
*_drop_conditions(source_id, authors))),
"unknown_posts": unknown,
}
def confirm_token(projection: dict) -> str:
"""What the operator's apply must echo: the preview it was shown, so a
list edited between preview and apply can't delete an unseen set."""
return f"remove-{projection['posts']}-posts-{projection['images']}-images"
def apply(session: Session, *, source_id: int, images_root: Path) -> dict:
"""Remove them. Counts are taken before the deletes they describe — once
the rows are gone there is nothing left to count (snippet #3087 note 2)."""
projection = preview(session, source_id=source_id)
_, authors = _load(session, source_id)
image_ids = session.execute(
select(ImageRecord.id).where(*_image_conditions(source_id, authors))
).scalars().all()
# Drops first-class: found BEFORE their members go, since the link that
# finds them is the member's absorbed_by_post_id.
drop_ids = session.execute(
select(Post.id).where(*_drop_conditions(source_id, authors))
).scalars().all()
deleted = delete_images(session, image_ids=list(image_ids), images_root=images_root)
# Attachments before posts: post_attachment.post_id is SET NULL, and a
# NULL-post attachment collides on its partial unique index (see
# cleanup_service.delete_artist_cascade).
attachments = session.execute(
delete(PostAttachment).where(*_attachment_conditions(source_id, authors))
).rowcount or 0
posts = session.execute(
Post.__table__.delete().where(Post.id.in_(_doomed_post_ids(source_id, authors)))
).rowcount or 0
drops = 0
if drop_ids:
drops = session.execute(
Post.__table__.delete().where(Post.id.in_(drop_ids))
).rowcount or 0
session.commit()
log.info(
"discord poster cleanup (source %s, keeping %s): %d post(s), %d image(s), "
"%d attachment(s), %d drop(s) removed",
source_id, authors, posts, deleted["images_deleted"], attachments, drops,
)
return {
**projection,
"posts_deleted": posts,
"images_deleted": deleted["images_deleted"],
"files_deleted": deleted["files_deleted"],
"attachments_deleted": attachments,
"drops_deleted": drops,
}
+57
View File
@@ -0,0 +1,57 @@
"""Closing the download runs a worker restart orphaned (#4433).
Its own module, importing nothing but the model, because its caller is the
worker boot hook in `celery_signals` — which every download task imports via
`celery_app`. Living in `tasks.maintenance` put the whole maintenance import
graph, the membership roster included, on the fetch path, which
`test_gated_reason` forbids.
"""
from __future__ import annotations
from datetime import UTC, datetime
from sqlalchemy import literal, update
from sqlalchemy.dialects.postgresql import JSONB
from ..models import DownloadEvent
DOWNLOAD_INTERRUPTED_MESSAGE = (
"interrupted by a worker restart — the next check picks it up where it left off"
)
def interrupt_orphaned_download_events(session, *, booted_at: datetime) -> int:
"""Close the download events a restart orphaned, without blaming the source.
Called when the download lane comes up (#4433). Anything still
pending/running from before this boot belongs to the previous process:
a walk that outlived the 90s stop grace was SIGKILLed, and a queued or
serialize-deferred task is held unacked until Redis redelivers it about an
hour later. Left alone, the 30-min stall sweep would error each one and
bump `consecutive_failures`, backing the source off as if the platform had
failed.
Instead they end as `skipped` (terminal, not a failure) and the source is
not touched: `last_checked_at` keeps its old value, so the next tick finds
it due and the walk resumes from its checkpoint. A redelivered message
that arrives later finds no pending event and opens a fresh one.
An event promoted to running after the boot has `started_at` reset to its
real start (download_service), so it is never caught here. Does NOT commit.
"""
now = datetime.now(UTC)
result = session.execute(
update(DownloadEvent)
.where(DownloadEvent.status.in_(["pending", "running"]))
.where(DownloadEvent.started_at < booted_at)
.values(
status="skipped",
finished_at=now,
error=DOWNLOAD_INTERRUPTED_MESSAGE,
metadata_=DownloadEvent.metadata_.op("||")(
literal({"error_type": "interrupted"}, JSONB)
),
)
.returning(DownloadEvent.id)
)
return len(result.all())
+109 -1
View File
@@ -23,6 +23,12 @@ log = logging.getLogger(__name__)
# The probe runs while the chip is drawing; names that take longer are skipped.
_NAME_LOOKUP_SECONDS = 6.0
# The poster picker waits on two pages of a channel's messages (#4488).
_POSTER_LOOKUP_SECONDS = 20.0
# A source's poster list (discord_ingester.AUTHORS_KEY) holds user ids; this
# keeps the name each id was picked under, for showing the list to a person.
AUTHOR_LABELS_KEY = "discord_author_labels"
class UnknownPlatformError(Exception):
@@ -37,6 +43,15 @@ class UnknownArtistError(Exception):
"""quick-add named an `artist_id` that does not exist."""
class PosterLookupError(Exception):
"""The Discord poster picker couldn't read the channel (no token, no
access, not a channel page, or Discord was slow)."""
class UnknownSourceError(Exception):
"""No source follows this Discord channel or its server."""
# Mirrored byte-for-byte from extension/lib/platforms.js
# PLATFORM_ARTIST_PATTERNS. Keep these two copies in sync by hand —
# reviewers catch drift.
@@ -110,6 +125,7 @@ class ExtensionService:
artist_id: int | None = None,
artist_name: str | None = None,
use_platform_name: bool = False,
discord_authors: list[dict] | None = None,
) -> dict:
"""Add `url` as a source. `artist_id` connects it to an existing
artist, `artist_name` to that artist (created if new); with neither,
@@ -119,7 +135,10 @@ class ExtensionService:
name is canon: a Patreon source added to an existing artist renames
that artist to the creator's Patreon display name. Name only — the
slug, and every path keyed off it, never moves (#130). Ignored on every
other platform, and when the name can't be read."""
other platform, and when the name can't be read.
`discord_authors` ([{id, name}], #4488) sets a Discord source's poster
list in the same step, on a new source or one already there."""
platform, raw_slug = self._derive(url)
url = canonical_source_url(platform, url, raw_slug)
renamed_from = None
@@ -133,6 +152,8 @@ class ExtensionService:
artist = (await self.session.execute(
select(Artist).where(Artist.id == existing.artist_id)
)).scalar_one()
if platform == DISCORD and discord_authors is not None:
await self._apply_posters(existing, discord_authors)
return self._shape(existing, artist, created_source=False, created_artist=False)
if artist_id is not None:
@@ -154,11 +175,98 @@ class ExtensionService:
source, created_source = await self._find_or_create_source(
artist_id=artist.id, platform=platform, url=url,
)
if platform == DISCORD and discord_authors is not None:
await self._apply_posters(source, discord_authors)
shaped = self._shape(source, artist, created_source, created_artist)
if renamed_from is not None:
shaped["renamed_from"] = renamed_from
return shaped
# -- Discord poster picker (#4488) ------------------------------------------
def _discord_page(self, url: str) -> tuple[str, str]:
platform, raw_slug = self._derive(url)
if platform != DISCORD:
raise InvalidUrlError("not a Discord page")
server_id, _, channel_id = raw_slug.partition("/")
if not channel_id:
raise PosterLookupError("Open a channel to see who posts in it.")
return server_id, channel_id
async def _source_for_page(self, server_id: str, channel_id: str) -> Source | None:
"""The source this channel's posts land on: the channel's own, else a
whole-server source that walks it."""
base = f"https://discord.com/channels/{server_id}"
return (
await self._existing_source(DISCORD, f"{base}/{channel_id}")
or await self._existing_source(DISCORD, base)
)
async def discord_posters(self, url: str) -> dict:
"""Who posts in the channel `url` shows, newest `_POSTER_SCAN`
messages deep, with the likely creator marked — and, when a source
already follows it, the poster list that source has now."""
server_id, channel_id = self._discord_page(url)
import asyncio
from .credential_service import CredentialService
from .discord_client import DiscordAPIError, DiscordClient
token = None
if self._crypto is not None:
token = await CredentialService(self.session, self._crypto).get_token(DISCORD)
if not token:
raise PosterLookupError("No Discord token is saved in FabledCurator.")
client = DiscordClient(token, max_retries=0)
loop = asyncio.get_running_loop()
try:
found = await asyncio.wait_for(
loop.run_in_executor(None, client.recent_posters, server_id, channel_id),
timeout=_POSTER_LOOKUP_SECONDS,
)
except TimeoutError as exc:
raise PosterLookupError("Discord took too long to answer; try again.") from exc
except DiscordAPIError as exc:
raise PosterLookupError(f"Couldn't read this channel: {exc}") from exc
source = await self._source_for_page(server_id, channel_id)
co = (source.config_overrides or {}) if source is not None else {}
from .discord_ingester import AUTHORS_KEY
return {
**found,
"source_id": source.id if source is not None else None,
"selected": list(co.get(AUTHORS_KEY) or []),
# So someone on the list who hasn't posted lately still shows by
# name, and can be unticked.
"selected_labels": dict(co.get(AUTHOR_LABELS_KEY) or {}),
}
async def set_discord_posters(self, url: str, authors: list[dict]) -> dict:
"""Store the picked posters on the source that follows `url`'s channel."""
server_id, channel_id = self._discord_page(url)
source = await self._source_for_page(server_id, channel_id)
if source is None:
raise UnknownSourceError("No source follows this channel yet.")
return await self._apply_posters(source, authors)
async def _apply_posters(self, source: Source, authors: list[dict]) -> dict:
"""Ids are what the list keys on — a username can change on a whim,
an id can't — and the name each was picked under rides beside it for
showing the list to a person. An empty pick clears the list, which
takes everyone's posts again."""
from .discord_ingester import AUTHORS_KEY
picked = [a for a in authors if str(a.get("id") or "").strip()]
co = dict(source.config_overrides or {})
if picked:
co[AUTHORS_KEY] = [str(a["id"]).strip() for a in picked]
co[AUTHOR_LABELS_KEY] = {
str(a["id"]).strip(): str(a.get("name") or a["id"]) for a in picked
}
else:
co.pop(AUTHORS_KEY, None)
co.pop(AUTHOR_LABELS_KEY, None)
source.config_overrides = co or None
await self.session.commit()
return {"source_id": source.id, "discord_authors": co.get(AUTHORS_KEY, [])}
async def _adopt_patreon_name(self, artist, raw_slug: str, url: str) -> str | None:
"""Rename `artist` to the Patreon display name; the old name when it
changed, else None. Unreadable name → no rename, never the handle."""
+64 -19
View File
@@ -49,10 +49,8 @@ from .audits import single_color
from .link_extract import extract_external_links
from .thumbnailer import Thumbnailer
from .wip_title import (
WIP_TITLE_SOFT_SOURCE,
WIP_TITLE_SOURCE,
apply_wip_image_tags,
matches_soft_wip_title,
matches_wip_title,
resolve_wip_tag_id,
)
@@ -428,11 +426,19 @@ class Importer:
return self._upsert_artist(name) if name else None
def _post_for_sidecar(
self, source: Path, artist: Artist | None
self, source: Path, artist: Artist | None,
*, source_row: Source | None = None,
) -> Post | None:
"""If a sidecar sits next to `source`, ensure its Source+Post
exist (idempotent) and return the Post — so attachments can link
to the same Post the per-member _apply_sidecar will reuse."""
to the same Post the per-member _apply_sidecar will reuse.
`source_row` is the subscription being downloaded, when there is one,
and wins over the (artist, platform) lookup — as it does in
`upsert_post_record`. The lookup takes the artist's FIRST source on the
platform, which is right only while an artist has one: a Discord artist
has one per channel, and every non-image file from a later channel was
filed under the first as an undated second post (#4435)."""
sc = find_sidecar(source)
if sc is None or artist is None:
return None
@@ -444,6 +450,9 @@ class Importer:
log.warning("sidecar parse failed for %s: %s", sc, exc)
return None
sd = parse_sidecar(data)
if source_row is not None:
src = source_row
else:
platform = sd.platform or "unknown"
src = self._lookup_source_for_sidecar(
artist_id=artist.id, platform=platform,
@@ -545,7 +554,7 @@ class Importer:
# nothing silently vanishes, matching extract_archive's
# fail-soft contract.
artist_use = artist if artist is not None else self._resolve_artist(source)
post = self._post_for_sidecar(source, artist_use)
post = self._post_for_sidecar(source, artist_use, source_row=source_row)
self._capture_attachment(
source, post=post, artist=artist_use, resolved=True,
)
@@ -554,7 +563,7 @@ class Importer:
return ImportResult(status="attached", error=reason)
artist_use = artist if artist is not None else self._resolve_artist(source)
post = self._post_for_sidecar(source, artist_use)
post = self._post_for_sidecar(source, artist_use, source_row=source_row)
member_ids: list[int] = []
# Every member image touched (new + superseded + deduped), so the
# from_attachment_id stamp below covers files that already existed in the
@@ -1040,9 +1049,7 @@ class Importer:
removal sticks. The existing catalogue is covered separately by the
operator-triggered backfill sweep. Gated by the settings toggle, and
best-effort: any failure is logged, never allowed to fail the import."""
hard_on = self.settings.wip_title_tagging_enabled
soft_on = self.settings.wip_soft_title_tagging_enabled
if not (hard_on or soft_on):
if not self.settings.wip_title_tagging_enabled:
return
if record.primary_post_id is None:
return
@@ -1050,21 +1057,14 @@ class Importer:
title = self.session.execute(
select(Post.post_title).where(Post.id == record.primary_post_id)
).scalar_one_or_none()
# HARD tier ("WIP"/"work in progress") wins — higher precision, and it
# trains the head; SOFT (sketch/doodle, #1474) is the provisional fallback
# that never trains (source wip_title_soft).
if hard_on and matches_wip_title(title):
source = WIP_TITLE_SOURCE
elif soft_on and matches_soft_wip_title(title):
source = WIP_TITLE_SOFT_SOURCE
else:
if not matches_wip_title(title):
return
if self._wip_tag_id is _UNSET:
self._wip_tag_id = resolve_wip_tag_id(self.session)
if self._wip_tag_id is None:
return
apply_wip_image_tags(
self.session, [record.id], self._wip_tag_id, source=source
self.session, [record.id], self._wip_tag_id, source=WIP_TITLE_SOURCE
)
except Exception as exc: # noqa: BLE001 — a tag must never fail an import
log.warning(
@@ -1141,9 +1141,51 @@ class Importer:
if post.artist_id is None:
post.artist_id = artist.id
self._apply_post_fields(post, sd)
self._redate_post_images(post)
self.session.commit()
return True
def _redate_post_images(self, post: Post) -> None:
"""Carry a post's date onto the images already linked to it (#4431).
The native ingesters import a post's media BEFORE its record: the
per-media sidecar holds only the image identity (post-first, #856), and
the date arrives with `_post.json`. So `_attach_provenance` links each
image to a post that has no date yet, and the image keeps its download
time. This runs when the record lands, and applies the same two rules
`_attach_provenance` applies: `effective_date` is the PRIMARY post's
date, and `earliest_post_date` is the earliest date across every post
the image is linked to. Only rows that differ are written."""
if post.post_date is None:
return
self.session.flush()
self.session.execute(
update(ImageRecord)
.where(ImageRecord.primary_post_id == post.id)
.where(ImageRecord.effective_date.is_distinct_from(post.post_date))
.values(effective_date=post.post_date)
.execution_options(synchronize_session=False)
)
linked = select(ImageProvenance.image_record_id).where(
ImageProvenance.post_id == post.id
)
earliest = (
select(func.min(Post.post_date))
.select_from(ImageProvenance)
.join(Post, Post.id == ImageProvenance.post_id)
.where(ImageProvenance.image_record_id == ImageRecord.id)
.where(Post.post_date.is_not(None))
.correlate(ImageRecord)
.scalar_subquery()
)
self.session.execute(
update(ImageRecord)
.where(ImageRecord.id.in_(linked))
.where(ImageRecord.earliest_post_date.is_distinct_from(earliest))
.values(earliest_post_date=earliest)
.execution_options(synchronize_session=False)
)
def attach_in_place(
self,
path: Path,
@@ -1187,7 +1229,10 @@ class Importer:
path, artist=artist, source_row=source,
)
if not is_supported(path):
post = self._post_for_sidecar(path, artist) if artist else None
post = (
self._post_for_sidecar(path, artist, source_row=source)
if artist else None
)
return self._capture_attachment(
path, post=post, artist=artist, resolved=True,
)
+9 -2
View File
@@ -289,6 +289,11 @@ class Ingester:
# Media handed to phase 3 for import. Marked seen by phase 3 once the
# import has run (`mark_seen_after_import`), not here — see there.
fetched: list[tuple[str, str]] = []
# Post-record keys written this walk. Marked with the media, after
# phase 3 has upserted the records — marking them at write time left a
# walk killed before phase 3 with posts the ledger calls recorded that
# the database never dated, and no later tick walks back to them (#4436).
recorded: list[tuple[str, str]] = []
downloaded = 0
errors = 0
quarantined = 0
@@ -342,7 +347,9 @@ class Ingester:
written_paths=written,
post_record_paths=list(post_records),
relink_source_paths=list(relink),
mark_seen_after_import=lambda: self._mark_seen(source_id, fetched),
mark_seen_after_import=lambda: self._mark_seen(
source_id, fetched + recorded,
),
stdout="\n".join(log_lines),
stderr="",
return_code=return_code,
@@ -511,7 +518,7 @@ class Ingester:
posts_with_body += 1
if rec.path is not None:
post_records.append(str(rec.path))
self._mark_seen(source_id, [(pkey, ppid)])
recorded.append((pkey, ppid))
# Per-post handling line in the run stdout (the existing
# "Raw stdout" panel) — the downloader already read the
# post; we only format its outcome here. post_type beside
+5 -68
View File
@@ -97,8 +97,8 @@ def _sigmoid(z, np):
def _conflict_scores(Xn, Wc, bc, np):
"""The presentation conflict signal (#141): per row, the MAX content-head
probability and WHICH head produced it. Shared by the system-tag sweep's guard-2
and the soft-wip audit — both ask "does this ALSO look like real content?"."""
probability and WHICH head produced it — the system-tag sweep's guard-2 asks
"does this ALSO look like real content?"."""
cprobs = _sigmoid(Xn @ Wc.T + bc, np)
return cprobs.max(axis=1), cprobs.argmax(axis=1)
@@ -106,10 +106,9 @@ def _conflict_scores(Xn, Wc, bc, np):
def _insert_presentation_review(
session, *, image_record_id, tag_id, conflict_tag_id, conflict_score, mode,
):
"""Single-source the ring-loud PresentationReview row shape so the two writers
(system-tag sweep guard-2 + soft-wip audit) can't drift on columns or `mode` —
they share the (image_record_id, tag_id) composite PK, so a divergent `mode`
would be a silent first-writer-wins bug."""
"""Single-source the ring-loud PresentationReview row shape, so every writer of
the (image_record_id, tag_id) composite PK agrees on columns and `mode` — a
divergent `mode` would be a silent first-writer-wins bug."""
session.execute(
pg_insert(PresentationReview)
.values(
@@ -963,68 +962,6 @@ def system_tag_auto_apply_sweep(
}
def soft_wip_conflict_audit(session: Session, dry_run: bool = False) -> dict:
"""Ring-loud audit for the SOFT WIP-title cohort (#1474). Images auto-tagged
`wip` from a low-precision sketch/doodle title (source='wip_title_soft') that ALSO
score >= the process conflict threshold on a content head are probably FINISHED
art mis-tagged as process — flag them (PresentationReview, mode='process') so the
review strip surfaces them ("also looks like <X>", Keep tag / Remove tag). Does
NOT remove the tag; the operator decides. No-op when there are no content heads.
numpy-only. Returns {n_scanned, n_flagged}."""
import numpy as np
from ..wip_title import WIP_TITLE_SOFT_SOURCE, resolve_wip_tag_id
settings = _settings(session)
ver = settings.embedder_model_version
conflict_thr = float(settings.process_conflict_threshold)
conf = _conflict_heads(session, ver)
wip_id = resolve_wip_tag_id(session)
if not conf or wip_id is None:
return {"n_scanned": 0, "n_flagged": 0}
Wc = np.vstack([np.asarray(r.weights, dtype=np.float32) for r in conf])
bc = np.asarray([r.bias for r in conf], dtype=np.float32)
conf_tag_ids = [r.tag_id for r in conf]
soft_ids = [iid for (iid,) in session.execute(
select(image_tag.c.image_record_id)
.where(image_tag.c.tag_id == wip_id)
.where(image_tag.c.source == WIP_TITLE_SOFT_SOURCE)
)]
# Skip images already flagged for this tag (idempotent re-runs).
flagged = {iid for (iid,) in session.execute(
select(PresentationReview.image_record_id)
.where(PresentationReview.tag_id == wip_id)
)}
soft_ids = [i for i in soft_ids if i not in flagged]
n_flagged = 0
scanned = 0
for start in range(0, len(soft_ids), _AUTO_APPLY_CHUNK):
chunk = soft_ids[start:start + _AUTO_APPLY_CHUNK]
emb = _load_embeddings(session, chunk)
cids = [i for i in chunk if i in emb]
if not cids:
continue
scanned += len(cids)
Xn = _l2norm(np.vstack([emb[i] for i in cids]).astype(np.float32), np)
max_c, arg_c = _conflict_scores(Xn, Wc, bc, np)
for k in range(len(cids)):
if float(max_c[k]) >= conflict_thr:
n_flagged += 1
if not dry_run:
_insert_presentation_review(
session,
image_record_id=cids[k], tag_id=wip_id,
conflict_tag_id=conf_tag_ids[int(arg_c[k])],
conflict_score=float(max_c[k]),
mode="process",
)
if not dry_run:
session.commit()
return {"n_scanned": scanned, "n_flagged": n_flagged}
def retract_auto_applied_heads(session: Session) -> int:
"""Soft auto-apply (milestone 139): re-score every standing source='head_auto'
tag against its CURRENT head and REMOVE the ones now BELOW the head's
-3
View File
@@ -32,11 +32,8 @@ from ...models.tag import image_tag
# `process_auto` (#1464): wip/editor screenshot applied by the process sweep are
# ALSO provisional — the head must learn only from title (`wip_title`) + manual
# labels, never its own auto-applied output, or it would runaway (operator 2026-07-12).
# `wip_title_soft` (#1474): the soft title tier (sketch/doodle) is LOW-precision, so
# it's provisional too — a finished piece titled "sketch" must not train the wip head.
_AUTO_SOURCES = (
"head_auto", "ccip_auto", "ml_auto", "presentation_auto", "process_auto",
"wip_title_soft",
)
+19
View File
@@ -60,3 +60,22 @@ def platform_lock(platform: str, *, ttl_seconds: int):
except redis.RedisError as exc: # pragma: no cover - broker outage
log.warning("platform_lock unavailable for %s: %s", platform, exc)
return None
def release_all_platform_locks() -> int:
"""Drop every serialized platform's lock. Returns how many were held.
Only for the download lane's boot (#4433). A worker that restarts mid-walk
is SIGKILLed past its stop grace, so its `finally` never releases the lock,
and the TTL keeps every other source on that platform bouncing for up to
27 minutes after the new worker is ready. At boot no walk of ours can be
running, so a held lock names a dead one. Assumes one download consumer —
the only shape FC deploys; a second replica booting would free a live
walk's lock (not corrupt it: the walk runs on, a second walk may overlap it).
"""
try:
client = _redis()
return int(client.delete(*(f"{_LOCK_PREFIX}{p}" for p in SERIALIZED_PLATFORMS)))
except redis.RedisError as exc: # pragma: no cover - broker outage
log.warning("could not release platform locks at boot: %s", exc)
return 0
+14
View File
@@ -479,6 +479,10 @@ class PostFeedService:
# the UI uses this to explain why a post it linked to is not in the
# stream.
"absorbed_by_post_id": post.absorbed_by_post_id,
# #4481. The Discord channel a message was posted in: both the native
# post record and gallery-dl's sidecars carry it as `channel`. None
# on every other platform, and on a Discord post that never said.
"channel": _discord_channel(post, source),
"artist": {"id": artist.id, "name": artist.name, "slug": artist.slug},
"source": (
{"id": source.id, "platform": source.platform}
@@ -488,3 +492,13 @@ class PostFeedService:
"thumbnails_more": thumbs_entry["more"],
"attachments": atts_map.get(post.id, []),
}
def _discord_channel(post: Post, source: Source | None) -> str | None:
if source is None or source.platform != "discord":
return None
raw = post.raw_metadata if isinstance(post.raw_metadata, dict) else {}
channel = raw.get("channel")
if not isinstance(channel, str):
return None
return channel.strip() or None
@@ -0,0 +1,86 @@
"""Date the posts whose record reached the disk but never the database (#4436).
A native walk writes each post's record (`_post.json`, or Discord's
`<day>_<message id>_post.json`) as it goes, and phase 3 upserts those records
after the walk — which is when the post, and through it its images, get their
date. Until #4436 the walk marked a record's post key seen at write time, so a
walk killed before phase 3 (a restart, a stall) left the post undated with the
ledger saying it was done, and no later tick walked back that far to fix it.
The record files are still on disk. This finds each undated native post's
record under its artist's folder and upserts it with the post's OWN source —
never the (artist, platform) lookup, which picks the artist's first source and
would misfile a Discord channel's post (#4435). Idempotent: a post that is
already dated is never looked at, and upserting a record twice changes nothing.
"""
from __future__ import annotations
import json
import logging
from collections import defaultdict
from pathlib import Path
from sqlalchemy import select
from sqlalchemy.orm import Session
from ..models import Artist, Post, Source
from ..utils.sidecar import parse_sidecar
from .download_backends import NATIVE_INGESTER_PLATFORMS
log = logging.getLogger(__name__)
def _is_record(path: Path) -> bool:
return path.name == "_post.json" or path.name.endswith("_post.json")
def _record_id(path: Path) -> str | None:
try:
data = json.loads(path.read_text("utf-8"))
except (OSError, ValueError):
return None
if not isinstance(data, dict):
return None
return parse_sidecar(data).external_post_id
def date_posts_from_records(session: Session, importer, images_root: Path) -> dict:
"""Upsert the on-disk record of every undated native post. Returns counts."""
rows = session.execute(
select(Post.external_post_id, Post.source_id, Post.artist_id, Source.platform)
.join(Source, Source.id == Post.source_id)
.where(
Post.post_date.is_(None),
Post.synthesized_by.is_(None),
Source.platform.in_(NATIVE_INGESTER_PLATFORMS),
)
).all()
wanted: dict[tuple[int, str], dict[str, int]] = defaultdict(dict)
for epid, source_id, artist_id, platform in rows:
wanted[(artist_id, platform)][epid] = source_id
summary = {"undated": len(rows), "dated": 0, "no_record": 0}
for (artist_id, platform), by_epid in wanted.items():
artist = session.get(Artist, artist_id)
root = Path(images_root) / artist.slug / platform if artist else None
if root is None or not root.is_dir():
summary["no_record"] += len(by_epid)
continue
found = 0
for record in root.rglob("*post.json"):
if not _is_record(record):
continue
epid = _record_id(record)
source_id = by_epid.pop(epid, None) if epid else None
if source_id is None:
continue
source = session.get(Source, source_id)
if importer.upsert_post_record(record, artist=artist, source=source):
found += 1
if not by_epid:
break
summary["dated"] += found
summary["no_record"] += len(by_epid)
if summary["undated"]:
log.info("date_posts_from_records: %s", summary)
return summary
+5
View File
@@ -68,6 +68,9 @@ class SourceRecord:
artist_slug: str
platform: str
url: str
# alembic 0115: the name the platform gives it, where the URL is opaque
# (a Discord link is two ids). None until a walk has read it.
display_name: str | None
enabled: bool
config_overrides: dict | None
last_checked_at: str | None
@@ -110,6 +113,7 @@ class SourceRecord:
"artist_slug": self.artist_slug,
"platform": self.platform,
"url": self.url,
"display_name": self.display_name,
"enabled": self.enabled,
"config_overrides": self.config_overrides,
"last_checked_at": self.last_checked_at,
@@ -301,6 +305,7 @@ class SourceService:
artist_slug=artist.slug,
platform=source.platform,
url=source.url,
display_name=source.display_name,
enabled=source.enabled,
config_overrides=source.config_overrides,
last_checked_at=source.last_checked_at.isoformat() if source.last_checked_at else None,
+6 -26
View File
@@ -27,13 +27,10 @@ from .image_tag_apply import insert_image_tags
# image_tag.source stamped on title-heuristic WIP tags — distinct from the other
# apply sources so provenance stays legible and a future undo can target only these.
# HARD tier ("WIP"/"work in progress") is high-precision → trains the wip head.
# Only the artist's own "WIP"/"work in progress" label counts — high-precision, so it
# trains the wip head. A sketch/doodle/scribble tier (#1474) was retired in milestone
# 430: a "sketch" is usually finished art, and its 6k tags flooded the review strip.
WIP_TITLE_SOURCE = "wip_title"
# SOFT tier (sketch/doodle/scribble, #1474) is LOWER-precision — a finished "sketch"
# is often not WIP. This source is PROVISIONAL (in training_data._AUTO_SOURCES) so it
# NEVER trains the wip head; a soft-tagged image that also looks like real content is
# surfaced by the ring-loud audit for review.
WIP_TITLE_SOFT_SOURCE = "wip_title_soft"
# A standalone "WIP" / "W.I.P" token, or the phrase "work in progress"
# (space/underscore/hyphen separated). The letter-boundary lookarounds are what
@@ -45,20 +42,10 @@ _WIP_RE = re.compile(
re.IGNORECASE,
)
# Soft tier: sketch / doodle / scribble (+ plurals), letter-boundary anchored so
# "sketchbook" / "kadoodle" don't trip it. Deliberately conservative — recall is
# secondary because the soft source doesn't train the head and the ring-loud audit
# catches false positives.
_SOFT_WIP_RE = re.compile(
r"(?<![A-Za-z])(?:sketch|sketches|doodle|doodles|scribble|scribbles)(?![A-Za-z])",
re.IGNORECASE,
)
# Coarse SQL prefilters for the backfill sweep — narrow the post scan to rows that
# Coarse SQL prefilter for the backfill sweep — narrows the post scan to rows that
# COULD match before the precise regex confirms. Case-insensitive ILIKE patterns.
# Each MUST stay a SUPERSET of its regex or the sweep would silently miss posts.
# It MUST stay a SUPERSET of the regex or the sweep would silently miss posts.
WIP_TITLE_SQL_PREFILTER = ("%wip%", "%work%progress%")
SOFT_WIP_TITLE_SQL_PREFILTER = ("%sketch%", "%doodle%", "%scribble%")
# Chunk bulk inserts so a large sweep can't blow past psycopg's 65535-parameter
# ceiling (3 params/row → ~21k rows max; 5k stays comfortably under).
@@ -66,19 +53,12 @@ _INSERT_CHUNK = 5000
def matches_wip_title(title: str | None) -> bool:
"""True when a post title explicitly marks it work-in-progress (HARD tier)."""
"""True when a post title explicitly marks it work-in-progress."""
if not title:
return False
return _WIP_RE.search(title) is not None
def matches_soft_wip_title(title: str | None) -> bool:
"""True when a title carries a SOFT WIP cue (sketch/doodle/scribble, #1474)."""
if not title:
return False
return _SOFT_WIP_RE.search(title) is not None
def resolve_wip_tag_id(session: Session) -> int | None:
"""The seeded ``wip`` system tag's id (migration 0075), or None if absent."""
return session.execute(
+47 -21
View File
@@ -137,7 +137,13 @@ IMPORT_BATCH_KEEP_DAYS = 30
# (the import queue itself stays at the 5-min default for single
# files); time_limit=2100.
QUEUE_STUCK_THRESHOLD_MINUTES: dict[str, int] = {
"ml": 25,
# ml: the scheduled auto-apply sweeps and refresh_character_prototypes run
# to a 35-min hard limit (2100s); 25 swept them mid-run (#4432). The two
# 65-min jobs have their own entries below.
"ml": 40,
# import: import_media_file's hard limit is 6 min (360s), one past the
# 5-min default this queue fell to (#4432).
"import": 10,
# download_source legitimately walks 5-25 min (Patreon/gallery-dl
# deep creators); its hard time_limit is DOWNLOAD_HARD_TIME_LIMIT
# (1500s = 25m). The 5-min default flagged healthy in-flight walks as
@@ -154,6 +160,12 @@ QUEUE_STUCK_THRESHOLD_MINUTES: dict[str, int] = {
# overrides below cover the outliers (backups, library audit).
"maintenance": 75,
"scan": 75,
# The long lane (#4432). Until TaskRun.queue asked the router, nothing was
# recorded here: these runs read as `maintenance` (75) or, for
# translation, `default` (5 — which failed healthy 35-min runs). The
# longest task without its own entry below is the admin family at a
# 40-min hard limit; 45 = 40 + 5.
"maintenance_long": 45,
}
TASK_STUCK_THRESHOLD_MINUTES: dict[str, int] = {
"backend.app.tasks.import_file.import_archive_file": 40,
@@ -179,6 +191,10 @@ TASK_STUCK_THRESHOLD_MINUTES: dict[str, int] = {
# external-fetch entry above — without an override a healthy in-flight walk
# is swept 'RecoverySweep' at the bare 5-min default. 30 = 25 + 5.
"backend.app.tasks.admin.reclaim_orphaned_attachments_task": 30,
# Head training and the manual head apply run to 65 min (3900s) — past the
# ml queue's threshold (#4432). 70 = 65 + 5.
"backend.app.tasks.ml.train_heads": 70,
"backend.app.tasks.ml.apply_head_tags": 70,
}
@@ -715,6 +731,32 @@ def recover_stalled_download_events() -> int:
return events_recovered
@celery.task(
name="backend.app.tasks.maintenance.date_posts_from_records",
soft_time_limit=1500,
time_limit=1800,
)
def date_posts_from_records() -> dict:
"""Date undated native posts from the records their walk left on disk
(#4436). Hourly and self-limiting: once every post is dated it is one
empty query."""
from ..services.importer import Importer
from ..services.post_record_repair import date_posts_from_records as _repair
from ..services.thumbnailer import Thumbnailer
images_root = IMAGES_ROOT
SessionLocal = _sync_session_factory()
with SessionLocal() as session:
importer = Importer(
session=session,
images_root=images_root,
import_root=images_root,
thumbnailer=Thumbnailer(images_root=images_root),
settings=ImportSettings.load_sync(session),
)
return _repair(session, importer, images_root)
@celery.task(name="backend.app.tasks.maintenance.recover_stalled_backup_runs")
def recover_stalled_backup_runs() -> int:
"""Flip BackupRun rows stuck in running/restoring past the hard limit
@@ -1041,8 +1083,7 @@ def cleanup_old_download_events() -> int:
def _backfill_wip_tier(session, tag_id, prefilter, matcher, source) -> int:
"""One keyset-paginated pass over posts whose title matches a WIP tier, applying
`tag_id` (stamped `source`) to their images. Shared by the hard + soft tiers
(#1458 / #1474). Coarse `prefilter` (ILIKE superset) narrows the scan; the precise
`tag_id` (stamped `source`) to their images (#1458). Coarse `prefilter` (ILIKE superset) narrows the scan; the precise
`matcher` confirms. Idempotent-additive (ON CONFLICT DO NOTHING). Returns the row
count newly applied."""
from ..models import Post
@@ -1082,26 +1123,17 @@ def _backfill_wip_tier(session, tag_id, prefilter, matcher, source) -> int:
)
def backfill_wip_title_tags() -> int:
"""Scan EXISTING posts for WIP titles and apply the `wip` system tag to their
images — the operator-triggered back-catalogue catch-up (task #1458 hard tier +
#1474 soft tier). New imports are tagged live by the importer; this covers the
existing library.
HARD tier ("WIP"/"work in progress") always runs (the operator triggered the
scan); the SOFT tier (sketch/doodle, provisional source) runs only when
wip_soft_title_tagging_enabled, AFTER hard so a title matching both keeps the
trained hard tag (ON CONFLICT DO NOTHING). Keyset-paginated, restart-safe.
images — the operator-triggered back-catalogue catch-up (task #1458). New
imports are tagged live by the importer; this covers the existing library.
Keyset-paginated, restart-safe.
Deliberately NOT scheduled as a beat: a periodic re-run would re-apply to matching
posts and silently undo a manual WIP removal, so it stays an explicit operator
action (Settings → "Scan existing posts for WIP titles"). Returns rows applied.
"""
from ..models import ImportSettings
from ..services.wip_title import (
SOFT_WIP_TITLE_SQL_PREFILTER,
WIP_TITLE_SOFT_SOURCE,
WIP_TITLE_SOURCE,
WIP_TITLE_SQL_PREFILTER,
matches_soft_wip_title,
matches_wip_title,
resolve_wip_tag_id,
)
@@ -1114,16 +1146,10 @@ def backfill_wip_title_tags() -> int:
"backfill_wip_title_tags: no `wip` system tag present; nothing to do"
)
return 0
settings = ImportSettings.load_sync(session)
applied = _backfill_wip_tier(
session, tag_id, WIP_TITLE_SQL_PREFILTER, matches_wip_title,
WIP_TITLE_SOURCE,
)
if settings.wip_soft_title_tagging_enabled:
applied += _backfill_wip_tier(
session, tag_id, SOFT_WIP_TITLE_SQL_PREFILTER, matches_soft_wip_title,
WIP_TITLE_SOFT_SOURCE,
)
if applied:
log.info("backfill_wip_title_tags: applied wip to %d image(s)", applied)
return applied
-18
View File
@@ -620,24 +620,6 @@ def scheduled_process_auto_apply() -> str:
return f"applied={result['n_applied']} flagged={result['n_flagged']}"
@celery.task(
name="backend.app.tasks.ml.scheduled_soft_wip_conflict_audit",
soft_time_limit=1800, time_limit=2100,
)
def scheduled_soft_wip_conflict_audit() -> str:
"""Ring-loud audit over the SOFT WIP-title cohort (#1474) — flag sketch/doodle
auto-tags that ALSO look like real content for review. No-op when there are no
content heads; idempotent (already-flagged images skipped). Runs regardless of
the process-sweep toggle, since soft-title tags come from the importer, not that
sweep. Wall-clock bounded by the task time limits."""
from ..services.ml.heads import soft_wip_conflict_audit
SessionLocal = _sync_session_factory()
with SessionLocal() as session:
result = soft_wip_conflict_audit(session)
return f"scanned={result['n_scanned']} flagged={result['n_flagged']}"
@celery.task(name="backend.app.tasks.ml.prune_presentation_reviews")
def prune_presentation_reviews() -> str:
"""Retention (rule 89): drop RESOLVED presentation-review flags older than 30
+4 -1
View File
@@ -180,7 +180,10 @@ per `docs/process.md`'s "add deps to the image when used by >1 project".
BUILD_REF` that every checkout in the file takes, rather than per job —
otherwise `sign-extension` would derive dev's extension version while
`build-web` bundled main's, and the release download would 404 on a version
that exists perfectly well. Every job then ASSERTS its checkout is `main`
that exists perfectly well. On every other trigger `BUILD_REF` is the
triggering COMMIT (`github.sha`), not the branch: a branch is re-resolved
per job, so a push landing mid-run used to move the publishing jobs onto a
commit the run's lanes never tested (run 7499, #4427). Every job then ASSERTS its checkout is `main`
before doing anything, because `env` inside `with:` is not a context this
runner is known to evaluate — if it silently resolved to empty, checkout
would fall back to the triggering ref and the refresh would publish dev's
+21 -1
View File
@@ -88,7 +88,12 @@ async function checkForUpdateInfo() {
currentVersion,
latestVersion,
channel,
xpiUrl: info && info.latest_url ? `${base}${info.latest_url}` : null,
// Where the Update button sends the operator: FC's own install card, not
// the XPI. Firefox refuses an add-on install whose navigation an extension
// started (tabs.create on the .xpi dies with NS_ERROR_FAILURE — operator-
// flagged 2026-09-25); it accepts one from a user click on a web page,
// which is exactly what the card's Install button is.
installPageUrl: base ? `${base}/subscriptions?tab=settings` : null,
};
}
@@ -271,11 +276,26 @@ browser.runtime.onMessage.addListener(async (msg) => {
artistId: msg.artistId ?? null,
artistName: msg.artistName ?? null,
usePlatformName: msg.usePlatformName === true,
discordAuthors: Array.isArray(msg.discordAuthors) ? msg.discordAuthors : null,
});
} catch (e) {
return { error: e.message };
}
case 'DISCORD_POSTERS':
try {
return await api.getDiscordPosters(msg.url);
} catch (e) {
return { error: e.message };
}
case 'SET_DISCORD_POSTERS':
try {
return await api.setDiscordPosters(msg.url, msg.authors || []);
} catch (e) {
return { error: e.message };
}
case 'SEARCH_ARTISTS':
try {
return { artists: await api.searchArtists(msg.q || '') };
+4
View File
@@ -100,3 +100,7 @@
/* The panel's own [hidden] — its rows are display:flex, which beats the UA's. */
.fc-panel [hidden] { display: none !important; }
.fc-panel__rename { margin-top: 8px; font-size: 13px; }
/* Poster picker (#4488): who posts in the channel, the likely creator ticked. */
.fc-panel__posters { max-height: 220px; overflow-y: auto; }
.fc-panel__poster { flex-wrap: wrap; }
.fc-panel__poster-meta { flex-basis: 100%; padding-left: 21px; font-size: 11px; color: rgb(170, 166, 156); }
+127
View File
@@ -81,6 +81,13 @@
const probe = currentProbe;
if (probe?.state === 'source_match') {
// A followed Discord channel opens its poster list (#4488); the artist
// page is one click further, in that panel.
if (probe.platform === 'discord' && probe.discord?.channel_id) {
if (document.getElementById('fc-add-panel')) closePanel();
else openPostersPanel(btn, probe);
return;
}
await openArtist(btn, probe.artist?.slug);
return;
}
@@ -158,6 +165,120 @@
document.getElementById('fc-add-panel')?.remove();
}
// ---- Poster picker (#4488) ----
// Who posted in this channel lately, read by FabledCurator with its Discord
// token, the creator the server is named for (or owned by) ticked. The
// source keeps their ids: a username can change on a whim, an id can't, so
// the list never quietly stops matching. Names are other people's text —
// createElement only.
function posterSection() {
const list = el('div', { class: 'fc-panel__posters' }, [
el('div', { class: 'fc-panel__hint', text: 'Reading who posts here…' }),
]);
const node = el('div', {}, [
el('div', { class: 'fc-panel__label', text: 'Only posts by' }),
list,
]);
const state = { data: null, picked: new Set() };
function render() {
const posters = withSelected(state.data);
if (!posters.length) {
list.replaceChildren(el('div', { class: 'fc-panel__hint', text: 'No one has posted here lately.' }));
return;
}
const rows = posters.map((p) => {
const id = String(p.id);
const box = el('input', { type: 'checkbox', checked: state.picked.has(id) });
box.addEventListener('change', () => {
if (box.checked) state.picked.add(id);
else state.picked.delete(id);
hint.textContent = pickHint();
});
const bits = [`${p.images} image${p.images === 1 ? '' : 's'}`, `${p.messages} message${p.messages === 1 ? '' : 's'}`];
if (p.reasons?.length) bits.unshift(p.reasons.join(', '));
return el('label', { class: 'fc-panel__radio fc-panel__poster' }, [
box,
el('span', { text: posterLabel(p) }),
el('span', { class: 'fc-panel__poster-meta', text: bits.join(' · ') }),
]);
});
const hint = el('div', { class: 'fc-panel__hint', text: pickHint() });
list.replaceChildren(...rows, hint);
}
function pickHint() {
return state.picked.size
? 'Only the ticked people’s posts are taken.'
: 'Nothing ticked: everyone’s posts are taken.';
}
async function load() {
let r;
try {
r = await browser.runtime.sendMessage({ type: 'DISCORD_POSTERS', url: window.location.href });
} catch (e) {
r = { error: e.message };
}
if (r?.error) {
list.replaceChildren(el('div', { class: 'fc-panel__hint', text: `Couldn’t list posters — ${r.error}` }));
return;
}
state.data = r;
state.picked = initialPosterPick(r);
render();
}
// null until the list has loaded: an add made before then leaves the
// source's list as it is instead of clearing it.
const request = () => (state.data ? postersRequest(state.data.posters, state.picked) : null);
return { node, load, request };
}
function openPostersPanel(btn, probe) {
closePanel();
const d = probe.discord || {};
const posters = posterSection();
const saveBtn = el('button', { class: 'fc-panel__btn fc-panel__btn--primary', text: 'Save' });
const artistBtn = el('button', { class: 'fc-panel__btn', text: `Open ${probe.artist?.name || 'artist'}` });
const cancelBtn = el('button', { class: 'fc-panel__btn', text: 'Close' });
const panel = el('div', { id: 'fc-add-panel', class: 'fc-panel' }, [
el('div', { class: 'fc-panel__title', text: 'Discord source' }),
el('div', { class: 'fc-panel__sub', text: d.channel_id ? channelLabel(d) : serverLabel(d) }),
posters.node,
el('div', { class: 'fc-panel__actions' }, [artistBtn, cancelBtn, saveBtn]),
]);
panel.addEventListener('keydown', (e) => {
e.stopPropagation();
if (e.key === 'Escape') closePanel();
});
cancelBtn.addEventListener('click', closePanel);
artistBtn.addEventListener('click', () => { closePanel(); openArtist(btn, probe.artist?.slug); });
saveBtn.addEventListener('click', async () => {
const authors = posters.request();
if (!authors) return;
saveBtn.disabled = true;
try {
const r = await browser.runtime.sendMessage({
type: 'SET_DISCORD_POSTERS', url: window.location.href, authors,
});
if (r?.error) {
showToast(`Error: ${r.error}`, 'error');
return;
}
showToast(authors.length ? `Only posts by ${authors.map((a) => a.name || a.id).join(', ')}` : 'Taking everyone’s posts', 'success');
closePanel();
} catch (e) {
showToast(`Error: ${e.message}`, 'error');
} finally {
saveBtn.disabled = false;
}
});
document.body.appendChild(panel);
posters.load();
}
async function openAddPanel(btn, probe) {
closePanel();
const platformName = PLATFORMS[probe.platform]?.name || probe.platform;
@@ -184,6 +305,8 @@
const d = probe.discord || {};
const discord = probe.platform === 'discord';
const choice = panelDefaults(probe, window.location.href);
// #4488: on a channel page, pick whose posts the new source takes.
const posters = discord && d.channel_id ? posterSection() : null;
const scopeRow = (value, label, disabled) => {
const input = el('input', {
@@ -234,6 +357,7 @@
el('div', { class: 'fc-panel__label', text: 'Artist' }),
el('div', { class: 'fc-panel__combo' }, [nameInput, results]),
renameRow,
...(posters ? [posters.node] : []),
hint,
el('div', { class: 'fc-panel__actions' }, [cancelBtn, addBtn]),
]);
@@ -392,6 +516,8 @@
addBtn.addEventListener('click', async () => {
const req = addRequest(choice);
if (!req) return;
const authors = posters?.request();
if (authors) req.discordAuthors = authors;
const btn = document.getElementById('fc-add-source-btn');
addBtn.disabled = true;
const ok = await add(btn || addBtn, req);
@@ -402,6 +528,7 @@
document.body.appendChild(panel);
refresh();
nameInput.focus();
posters?.load();
nameInput.select();
// Search what the field opens with — the server's name, usually — so an
// artist it already matches is picked before the operator types anything.
+13 -1
View File
@@ -106,8 +106,11 @@ class FabledCuratorAPI {
// artist from the URL. A Discord channel always sends one.
// usePlatformName: a Patreon source joining an existing artist renames it
// to the Patreon display name (Patreon is canon; name only, never the slug).
quickAddSource(url, { artistId = null, artistName = null, usePlatformName = false } = {}) {
// discordAuthors ([{id, name}], #4488): a Discord source's poster list, set
// in the same step.
quickAddSource(url, { artistId = null, artistName = null, usePlatformName = false, discordAuthors = null } = {}) {
const body = { url };
if (discordAuthors) body.discord_authors = discordAuthors;
if (artistId != null) body.artist_id = artistId;
else if (artistName) body.artist_name = artistName;
if (usePlatformName) body.use_platform_name = true;
@@ -123,6 +126,15 @@ class FabledCuratorAPI {
const qs = new URLSearchParams(params).toString();
return this.request('GET', `/extension/probe?${qs}`);
}
// #4488: who posts in the Discord channel `url` shows, lately, with the
// likely creator marked; and setting the list on the source that follows it.
getDiscordPosters(url) {
const qs = new URLSearchParams({ url }).toString();
return this.request('GET', `/extension/discord/posters?${qs}`);
}
setDiscordPosters(url, authors) {
return this.request('POST', '/extension/discord/posters', { url, authors });
}
// Latest published extension version on this instance — drives the in-app
// update prompt. Public endpoint (no key needed, but request() sends it
// harmlessly). Returns {version, xpi_url, latest_url, sha256}.
+46
View File
@@ -144,3 +144,49 @@ function inlineCompletion(typed, results) {
if (!t) return null;
return (results || []).find((a) => a.name.length > t.length && a.name.toLowerCase().startsWith(t)) || null;
}
// ---- Discord poster picker (#4488) ----
// /api/extension/discord/posters lists who posted in the channel lately, with
// the creator the server is named for (or owned by) marked `suggested`. The
// picker ticks posters; the source keeps their ids, since a username can
// change on a whim and an id can't.
/** "Todding (@todding)" — the display name, then the handle when it differs. */
function posterLabel(p) {
const shown = p?.global_name || p?.username || p?.id || '';
const handle = p?.username && p.username !== shown ? ` (@${p.username})` : '';
return `${shown}${handle}`;
}
/** What the picker starts with: the source's current list when it has one,
* else the suggested creator(s). A set of ids. */
function initialPosterPick(data) {
const current = (data?.selected || []).map(String);
if (current.length) return new Set(current);
return new Set((data?.posters || []).filter((p) => p.suggested).map((p) => String(p.id)));
}
/** The checklist's rows: who posted lately, then anyone already on the
* source's list who hasn't — by the name they were picked under — so they
* can still be unticked. */
function withSelected(data) {
const posters = [...(data?.posters || [])];
const shown = new Set(posters.map((p) => String(p.id)));
const labels = data?.selected_labels || {};
for (const id of (data?.selected || []).map(String)) {
if (shown.has(id)) continue;
posters.push({ id, global_name: labels[id] || null, username: null,
messages: 0, images: 0, reasons: ['on the list, not posting lately'] });
}
return posters;
}
/** The picked posters as the API takes them: [{id, name}], in list order.
* An id already on the source but no longer posting lately is kept. */
function postersRequest(posters, picked) {
const byId = new Map((posters || []).map((p) => [String(p.id), p]));
return [...picked].map((id) => ({
id,
name: byId.has(id) ? (byId.get(id).global_name || byId.get(id).username || id) : null,
}));
}
+8 -4
View File
@@ -76,7 +76,7 @@ function updateConnectionDot(connected) {
async function checkForUpdate() {
try {
const r = await browser.runtime.sendMessage({ type: 'CHECK_UPDATE' });
if (r && r.updateAvailable && r.xpiUrl) showUpdateBanner(r);
if (r && r.updateAvailable && r.installPageUrl) showUpdateBanner(r);
} catch { /* non-fatal */ }
}
@@ -86,10 +86,14 @@ function showUpdateBanner(r) {
// exactly as it did before the field existed.
const channel = r.channel ? ` (${r.channel})` : '';
document.getElementById('update-text').textContent =
`Update available${channel} — v${r.latestVersion} (installed v${r.currentVersion})`;
// Opening the signed XPI triggers Firefox's native install prompt.
`Update available${channel} — v${r.latestVersion} (installed v${r.currentVersion}). ` +
'Opens FabledCurator — click “Install Firefox extension” there.';
// Opens FC's install card rather than the XPI: Firefox only installs an
// add-on from a user click on a web page, never from a tab an extension
// opened on the .xpi itself.
document.getElementById('update-btn').addEventListener('click', () => {
browser.tabs.create({ url: r.xpiUrl });
browser.tabs.create({ url: r.installPageUrl });
window.close();
});
document.getElementById('update-banner').classList.remove('hidden');
}
+43
View File
@@ -198,3 +198,46 @@ describe('Add panel on Patreon and SubscribeStar', () => {
expect(c.artistName).toBe('Tamada')
})
})
const { posterLabel, initialPosterPick, postersRequest, withSelected } = loadLib('chip.js', [
'posterLabel', 'initialPosterPick', 'postersRequest', 'withSelected',
])
describe('the poster checklist rows', () => {
it('adds someone on the list who has not posted lately, by their saved name', () => {
const rows = withSelected({
posters: [{ id: '7', username: 'todding', global_name: 'Todding' }],
selected: ['7', '42'],
selected_labels: { 42: 'Old Name' },
})
expect(rows.map((r) => r.id)).toEqual(['7', '42'])
expect(posterLabel(rows[1])).toBe('Old Name')
expect(rows[1].reasons).toEqual(['on the list, not posting lately'])
})
})
describe('the Discord poster picker (#4488)', () => {
const posters = [
{ id: '7', username: 'todding', global_name: 'Todding', suggested: true },
{ id: '8', username: 'jakeboii', global_name: 'Jake Boii', suggested: false },
{ id: '9', username: 'plain', global_name: null, suggested: false },
]
it('labels a poster by display name, then handle', () => {
expect(posterLabel(posters[0])).toBe('Todding (@todding)')
expect(posterLabel(posters[2])).toBe('plain')
})
it('starts from the source list when it has one, else the suggestion', () => {
expect([...initialPosterPick({ posters, selected: [] })]).toEqual(['7'])
expect([...initialPosterPick({ posters, selected: ['8'] })]).toEqual(['8'])
expect([...initialPosterPick({ posters: [] })]).toEqual([])
})
it('sends ids with the names they were picked under, keeping unknown ids', () => {
expect(postersRequest(posters, new Set(['7', '42']))).toEqual([
{ id: '7', name: 'Todding' },
{ id: '42', name: null },
])
})
})
@@ -101,7 +101,7 @@ import MaintenanceTile from '../common/MaintenanceTile.vue'
import { useCleanupStore } from '../../stores/cleanup.js'
const store = useCleanupStore()
const threshold = ref(0.95)
const threshold = ref(0.995)
const tolerance = ref(30)
const audit = ref(null)
const busy = ref(false)
@@ -2,14 +2,14 @@
<!-- System-tag auto-applies (chrome hides / process WIP tags) that ALSO looked
like real content — surfaced PROACTIVELY atop the gallery whenever there's
something to review (NOT gated on the Show-hidden toggle, so misfires can't
go unnoticed), most-concerning first, with keep / remove (#141, #1464).
go unnoticed), most-concerning first (#141, #1464). Each card asks one
question — is this image a <tag>? — and its buttons answer it (#4424).
Renders nothing when there's nothing to review. -->
<section v-if="items.length" class="fc-review" aria-label="Auto-tagged images to review">
<div class="fc-review__head">
<v-icon size="18" color="warning">mdi-alert-outline</v-icon>
<span class="fc-review__title">
{{ items.length }} auto-tagged {{ items.length === 1 ? 'image' : 'images' }}
may be real content — review
{{ items.length }} {{ items.length === 1 ? 'auto-tag' : 'auto-tags' }} to check
</span>
</div>
<div class="fc-review__cards">
@@ -21,22 +21,23 @@
class="fc-review-card__thumb" loading="lazy"
>
<div class="fc-review-card__body">
<div class="fc-review-card__question">{{ question(it) }}</div>
<div
class="fc-review-card__conflict"
:title="`Scored ${Math.round(it.conflict_score * 100)}% on “${it.conflict_name || 'a content tag'}”`"
class="fc-review-card__reason"
:title="reasonTitle(it)"
>
also looks like <strong>{{ it.conflict_name || 'content' }}</strong>
also {{ Math.round(it.conflict_score * 100) }}%
<strong>{{ it.conflict_name || 'content' }}</strong>
</div>
<div class="fc-review-card__tag">{{ tagLine(it) }}</div>
<div class="fc-review-card__acts">
<button
type="button" class="fc-review-btn fc-review-btn--keep"
:disabled="busy.includes(keyOf(it))" @click="resolve(it, 'keep')"
>{{ keepLabel(it) }}</button>
>Is {{ withArticle(it) }}</button>
<button
type="button" class="fc-review-btn fc-review-btn--unhide"
:disabled="busy.includes(keyOf(it))" @click="resolve(it, 'unhide')"
>{{ removeLabel(it) }}</button>
>Is not {{ withArticle(it) }}</button>
</div>
</div>
</div>
@@ -55,11 +56,18 @@ const items = ref([])
const busy = ref([])
function keyOf(it) { return `${it.image_id}:${it.tag_id}` }
// Chrome flags hide the image (keep-hidden / un-hide); process flags leave it
// visible and just tagged (keep-tag / remove-tag). Same endpoints, different words.
function tagLine(it) { return (it.mode === 'process' ? 'auto-tagged ' : 'hidden as ') + it.tag_name }
function keepLabel(it) { return it.mode === 'process' ? 'Keep tag' : 'Keep hidden' }
function removeLabel(it) { return it.mode === 'process' ? 'Remove tag' : 'Un-hide' }
// The card asks whether the image IS the auto-applied system tag, and the buttons
// answer that (operator, #4424: "is a <tag>" / "is not a <tag>"). "Is" keeps the
// tag ('keep'); "Is not" removes it, un-hiding a chrome image ('unhide'). The
// content tag it also scored on is the reason it was flagged, not the question.
function noun(it) { return it.tag_name === 'wip' ? 'WIP' : it.tag_name }
function withArticle(it) { return (/^[aeiou]/i.test(noun(it)) ? 'an ' : 'a ') + noun(it) }
function question(it) { return `Is this ${withArticle(it)}?` }
function reasonTitle(it) {
const hidden = it.mode === 'process' ? '' : ' It is hidden from the gallery until you answer.'
return `Auto-tagged “${it.tag_name}”, but it also scored ${Math.round(it.conflict_score * 100)}% `
+ `on “${it.conflict_name || 'a content tag'}”, so it may be finished art.${hidden}`
}
async function load() {
// Fetched unconditionally on mount — the strip prompts for pending misfires
@@ -77,14 +85,11 @@ async function resolve(it, action) {
await api.post(`/api/gallery/hidden-review/${it.image_id}/${it.tag_id}/${action}`)
items.value = items.value.filter((x) => keyOf(x) !== k)
if (action === 'unhide') {
const verb = it.mode === 'process' ? 'Removed' : 'Un-hidden'
toast({ text: `${verb} — “${it.tag_name}” removed; it'll train the head`, type: 'success' })
const shown = it.mode === 'process' ? '' : ', back in the gallery'
toast({ text: `Not ${withArticle(it)} — “${it.tag_name}” removed${shown}; the tagger learns from it`, type: 'success' })
}
} catch (e) {
toast({
text: `Could not ${action === 'keep' ? 'keep hidden' : 'un-hide'}: ${e.message}`,
type: 'error',
})
toast({ text: `Could not save your answer: ${e.message}`, type: 'error' })
} finally {
busy.value = busy.value.filter((x) => x !== k)
}
@@ -112,7 +117,7 @@ onMounted(load)
display: flex; gap: 10px; overflow-x: auto; padding-bottom: 4px;
}
.fc-review-card {
flex: 0 0 auto; width: 150px;
flex: 0 0 auto; width: 170px;
display: flex; flex-direction: column;
border: 1px solid rgb(var(--v-theme-surface-light));
border-radius: 6px; overflow: hidden;
@@ -123,16 +128,16 @@ onMounted(load)
background: rgb(var(--v-theme-surface-light));
}
.fc-review-card__body { padding: 6px 8px; }
.fc-review-card__conflict {
font-size: 11px; color: rgb(var(--v-theme-on-surface));
.fc-review-card__question {
font-size: 12px; font-weight: 600; color: rgb(var(--v-theme-on-surface));
overflow: hidden; text-overflow: ellipsis; white-space: nowrap;
}
.fc-review-card__conflict strong { color: rgb(var(--v-theme-warning)); }
.fc-review-card__tag {
.fc-review-card__reason {
font-size: 10px; color: rgb(var(--v-theme-on-surface-variant));
margin: 1px 0 6px;
overflow: hidden; text-overflow: ellipsis; white-space: nowrap;
}
.fc-review-card__reason strong { color: rgb(var(--v-theme-warning)); font-weight: 600; }
.fc-review-card__acts { display: flex; gap: 4px; }
.fc-review-btn {
flex: 1; font-size: 11px; padding: 3px 4px; border-radius: 4px;
@@ -26,6 +26,7 @@
:to="{ name: 'artist', params: { slug: post.artist.slug } }"
class="fc-post-card__artist"
>{{ post.artist.name }}</RouterLink>
<span v-if="post.channel" class="fc-post-card__meta">#{{ post.channel }}</span>
<span class="fc-post-card__date" :title="absoluteDate">{{ relativeDate }}</span>
<span v-if="totalImages" class="fc-post-card__meta">
· {{ totalImages }} image{{ totalImages === 1 ? '' : 's' }}
@@ -17,7 +17,7 @@
<v-icon start>mdi-folder-zip-outline</v-icon> Re-extract archives now
</v-btn>
<span v-if="queued" class="ml-3 text-caption text-success">Queued ✓</span>
<QueueStatusBar queue="maintenance" queue-label="Maintenance" />
<QueueStatusBar queue="maintenance_long" queue-label="Long maintenance" />
</MaintenanceTile>
</template>
@@ -49,16 +49,20 @@
sometimes triggered nothing instead of the install dialog
(operator-flagged 2026-05-26). No `download` attribute —
that would force a save dialog instead of install. -->
<!-- The VERSIONED file, not the `latest` alias: a versioned URL can
only ever be this build's bytes, so a browser cache can't hand
back the previous build (operator-flagged 2026-09-25: a cached
alias reinstalled the old version and the update never took). -->
<v-btn
v-if="isFirefox"
color="accent" variant="flat" rounded="pill"
prepend-icon="mdi-firefox"
:href="manifest.latest_url"
:href="manifest.xpi_url"
>Install Firefox extension</v-btn>
<v-btn
variant="outlined" rounded="pill"
:href="manifest.latest_url" download
:href="manifest.xpi_url" download
prepend-icon="mdi-download"
>Download XPI</v-btn>
@@ -77,7 +77,7 @@
<v-col cols="12" sm="6">
<v-slider
v-model="local.single_color_threshold" label="Single-color threshold"
min="0.5" max="1" step="0.05" thumb-label hide-details color="accent"
min="0.9" max="1" step="0.005" thumb-label hide-details color="accent"
:disabled="!local.skip_single_color" @end="save"
/>
</v-col>
@@ -103,17 +103,6 @@
the Explore browse. Applies to new imports; run the scan below to catch
posts already in your library.
</div>
<v-switch
v-model="local.wip_soft_title_tagging_enabled"
label="Also tag “sketch” / “doodle” titles (lower precision)"
density="compact" hide-details color="primary" @change="save"
/>
<div class="fc-help mb-3">
Extends the above to softer cues (<code>sketch</code>, <code>doodle</code>,
<code>scribble</code>). These stay <strong>visible</strong> and never train
the tagging model — a daily audit flags any that actually look like finished
art for review. Off by default.
</div>
<v-btn
variant="tonal" color="primary" size="small"
:loading="store.wipScanBusy" prepend-icon="mdi-magnify"
@@ -153,10 +142,9 @@ const PHASH_TICKS = { 0: 'Exact', 12: 'Strict', 24: 'Default', 48: 'Loose' }
const local = reactive({
min_width: 0, min_height: 0,
skip_transparent: false, transparency_threshold: 0.9,
skip_single_color: false, single_color_threshold: 0.95,
skip_single_color: false, single_color_threshold: 0.995,
phash_threshold: 24,
wip_title_tagging_enabled: true,
wip_soft_title_tagging_enabled: false,
})
watch(() => store.settings, (s) => { if (s) Object.assign(local, s) }, { immediate: true })
@@ -18,7 +18,7 @@
<v-icon start>mdi-file-remove-outline</v-icon> Repair missing-file records
</v-btn>
<span v-if="queued" class="ml-3 text-caption text-success">Queued ✓</span>
<QueueStatusBar queue="maintenance" queue-label="Maintenance" />
<QueueStatusBar queue="maintenance_long" queue-label="Long maintenance" />
</MaintenanceTile>
</template>
@@ -41,7 +41,7 @@ const props = defineProps({
const QUEUE_NAMES = [
'default', 'import', 'thumbnail', 'ml',
'download', 'scan', 'maintenance',
'download', 'scan', 'maintenance', 'maintenance_long',
]
function formatDepth(name) {
@@ -0,0 +1,135 @@
<template>
<!-- #4486. A Discord source's posts by people outside its "Only posts by"
list, taken before the list existed. Preview first, always: nothing is
deleted until the operator has seen the per-poster breakdown and typed
the token that preview returned. -->
<v-dialog :model-value="modelValue" max-width="560"
@update:model-value="$emit('update:modelValue', $event)">
<v-card>
<v-card-title>Remove posts from other posters</v-card-title>
<v-card-text>
<div class="text-caption mb-2">{{ source.display_name || source.url }}</div>
<v-progress-linear v-if="loading" indeterminate color="accent" class="mb-3" />
<v-alert v-else-if="error" type="warning" variant="tonal" density="compact">
{{ error }}
</v-alert>
<template v-else-if="result">
<v-alert type="success" variant="tonal" density="compact">
Removed {{ result.posts_deleted }} post(s), {{ result.images_deleted }} image(s),
{{ result.attachments_deleted }} attachment(s) and {{ result.drops_deleted }}
grouped post(s). Grouping rebuilds what remains on its next sweep.
</v-alert>
</template>
<template v-else-if="preview">
<p class="mb-2">
Keeping posts by <strong>{{ preview.authors.join(', ') }}</strong>.
</p>
<p v-if="!preview.posts" class="fc-dim">
Nothing to remove: every post with a known poster is by someone on the list.
</p>
<template v-else>
<v-table density="compact" class="mb-3">
<thead><tr><th>Poster</th><th class="text-right">Posts</th></tr></thead>
<tbody>
<tr v-for="p in preview.posters" :key="p.poster">
<td>{{ p.poster }}</td>
<td class="text-right">{{ p.posts }}</td>
</tr>
</tbody>
</v-table>
<p class="mb-1">
{{ preview.posts }} post(s), {{ preview.images }} image(s) found only on them,
{{ preview.attachments }} attachment(s).
</p>
<p v-if="preview.drops" class="mb-1">
{{ preview.drops }} grouped post(s) that include them are ungrouped; the
rest regroup on the next sweep.
</p>
<p class="fc-dim">
Images that also belong to a kept post stay. If a name you want to keep
appears above, add it to the source's "Only posts by" list first.
</p>
</template>
<p v-if="preview.unknown_posts" class="fc-dim mt-2">
{{ preview.unknown_posts }} older post(s) don't record who posted them and
are left alone.
</p>
</template>
</v-card-text>
<v-card-actions>
<v-spacer />
<v-btn variant="text" @click="$emit('update:modelValue', false)">
{{ result ? 'Close' : 'Cancel' }}
</v-btn>
<v-btn
v-if="preview && preview.posts && !result"
color="error" variant="flat" :loading="busy"
@click="confirmOpen = true"
>Remove {{ preview.posts }} post(s)</v-btn>
</v-card-actions>
</v-card>
<DestructiveConfirmModal
v-if="preview"
v-model="confirmOpen"
action="delete" kind="other posters' posts" tier="C"
:expected-token-override="preview.confirm_token"
:projected-counts="{ posts: preview.posts, images: preview.images,
attachments: preview.attachments, grouped: preview.drops }"
@confirm="onConfirm"
/>
</v-dialog>
</template>
<script setup>
import { ref, watch } from 'vue'
import { useSourcesStore } from '../../stores/sources.js'
import DestructiveConfirmModal from '../modal/DestructiveConfirmModal.vue'
const props = defineProps({
modelValue: { type: Boolean, default: false },
source: { type: Object, required: true },
})
defineEmits(['update:modelValue'])
const store = useSourcesStore()
const loading = ref(false)
const busy = ref(false)
const error = ref('')
const preview = ref(null)
const result = ref(null)
const confirmOpen = ref(false)
function message(e) {
return e?.body?.detail || e?.body?.error || e?.message || String(e)
}
watch(() => props.modelValue, async (open) => {
if (!open) return
preview.value = null
result.value = null
error.value = ''
loading.value = true
try {
preview.value = await store.previewOtherPosters(props.source.id)
} catch (e) {
error.value = message(e)
} finally {
loading.value = false
}
})
async function onConfirm(token) {
busy.value = true
try {
result.value = await store.removeOtherPosters(props.source.id, token)
} catch (e) {
error.value = message(e)
} finally {
busy.value = false
}
}
</script>
<style scoped>
.fc-dim { color: rgb(var(--v-theme-on-surface-variant)); }
</style>
@@ -50,6 +50,17 @@
library — without re-downloading media
</v-list-item-subtitle>
</v-list-item>
<v-list-item
v-if="hasPosterList && !running"
prepend-icon="mdi-account-remove-outline"
@click="cleanupOpen = true"
>
<v-list-item-title>Remove posts from other posters</v-list-item-title>
<v-list-item-subtitle>
Preview, then delete the posts taken from people outside this
source's "Only posts by" list
</v-list-item-subtitle>
</v-list-item>
<v-divider v-if="isNative && !running" />
<v-list-item
base-color="error"
@@ -59,12 +70,14 @@
<v-list-item-title>Remove source</v-list-item-title>
</v-list-item>
</KebabMenu>
<PosterCleanupDialog v-if="hasPosterList" v-model="cleanupOpen" :source="source" />
</div>
</template>
<script setup>
import { computed } from 'vue'
import { computed, ref } from 'vue'
import KebabMenu from '../common/KebabMenu.vue'
import PosterCleanupDialog from './PosterCleanupDialog.vue'
const props = defineProps({
source: { type: Object, required: true },
@@ -80,6 +93,10 @@ const recapturing = computed(() => !!props.source.backfill_recapture)
// which those are (`native_ingester`); a copied list here went stale when
// Discord moved over.
const isNative = computed(() => !!props.source.native_ingester)
// #4486: only a Discord source with an "Only posts by" list has other posters.
const hasPosterList = computed(() => props.source.platform === 'discord'
&& (props.source.config_overrides?.discord_authors || []).length > 0)
const cleanupOpen = ref(false)
</script>
<style scoped>
@@ -16,7 +16,7 @@
<a
:href="source.url" target="_blank" rel="noopener"
class="fc-source-card__url" @click.stop
>{{ source.url }}</a>
>{{ source.display_name || source.url }}</a>
<v-btn
icon="mdi-pencil" size="x-small" variant="text"
@click.stop="$emit('edit', source)"
@@ -39,18 +39,18 @@
</v-chip>
<v-chip
v-else-if="source.backfill_state === 'running'"
size="x-small" color="info" variant="tonal" label
size="x-small" color="info" variant="flat" label
>{{ source.backfill_bypass_seen ? 'Recovering' : (source.backfill_recapture ? 'Recapturing' : 'Backfilling')
}}{{ source.backfill_posts
? ` · ${source.backfill_posts} posts`
: (source.backfill_chunks ? ` (${source.backfill_chunks})` : '') }}</v-chip>
<v-chip
v-else-if="source.backfill_state === 'complete'"
size="x-small" color="success" variant="tonal" label
size="x-small" color="success" variant="flat" label
>Backfilled</v-chip>
<v-chip
v-else-if="source.backfill_state === 'stalled'"
size="x-small" color="warning" variant="tonal" label
size="x-small" color="warning" variant="flat" label
>Stalled</v-chip>
</div>
@@ -44,6 +44,17 @@
v-model="structuredSince" label="Skip posts older than (YYYY-MM-DD)"
placeholder="2024-01-01" hide-details class="mt-2"
/>
<!-- A creator's server is full of other members posting their own
pictures; the source is subscribed to the creator (#4481). -->
<v-text-field
v-if="platform === 'discord'"
v-model="structuredAuthors" label="Only posts by (Discord names or ids)"
placeholder="Todding" hint="Comma-separated. Empty takes everyone's."
persistent-hint class="mt-2"
/>
<div v-if="platform === 'discord' && authorNames.length" class="text-caption mt-2">
{{ authorNames.join(', ') }}
</div>
<p class="text-caption mt-2" style="opacity: 0.75">
More per-platform fields land here over time. Use Advanced JSON for everything else.
</p>
@@ -98,6 +109,33 @@ const urlError = ref('')
const configTab = ref('structured')
const structuredVideos = ref(true)
const structuredSince = ref('')
const structuredAuthors = ref('')
// Keys the structured view has no field for, carried through its saves so
// switching tabs never drops what the JSON view set.
const otherConfig = ref({})
const STRUCTURED_KEYS = ['videos', 'since', 'discord_authors']
// The extension's poster picker stores ids (#4488) and the name each was
// picked under, so a list of numbers still says who it means.
const authorNames = computed(() => {
const labels = otherConfig.value.discord_author_labels || {}
return splitAuthors(structuredAuthors.value)
.map(a => (labels[a] ? `${labels[a]} (${a})` : null))
.filter(Boolean)
})
function splitAuthors(txt) {
return (txt || '').split(',').map(s => s.trim()).filter(Boolean)
}
function takeConfig(co) {
structuredVideos.value = co.videos !== false
structuredSince.value = co.since ?? ''
structuredAuthors.value = Array.isArray(co.discord_authors) ? co.discord_authors.join(', ') : ''
otherConfig.value = Object.fromEntries(
Object.entries(co).filter(([k]) => !STRUCTURED_KEYS.includes(k)),
)
}
const jsonText = ref('{}')
const jsonError = ref('')
@@ -105,13 +143,15 @@ const busy = ref(false)
// Sync config_overrides between the two views.
const config = computed(() => {
const out = {}
const out = { ...otherConfig.value }
if (!structuredVideos.value) out.videos = false
if (structuredSince.value) out.since = structuredSince.value
const authors = splitAuthors(structuredAuthors.value)
if (authors.length) out.discord_authors = authors
return out
})
watch([structuredVideos, structuredSince], () => {
watch([structuredVideos, structuredSince, structuredAuthors], () => {
if (configTab.value === 'structured') {
jsonText.value = JSON.stringify(config.value, null, 2)
jsonError.value = ''
@@ -128,8 +168,7 @@ watch(jsonText, (txt) => {
}
jsonError.value = ''
// Reflect recognized keys into the structured view.
structuredVideos.value = parsed.videos !== false
structuredSince.value = parsed.since ?? ''
takeConfig(parsed)
} catch {
jsonError.value = 'Invalid JSON'
}
@@ -146,14 +185,13 @@ watch(() => props.modelValue, async (open) => {
url.value = props.source.url
enabled.value = props.source.enabled
const co = props.source.config_overrides || {}
structuredVideos.value = co.videos !== false
structuredSince.value = co.since ?? ''
takeConfig(co)
jsonText.value = JSON.stringify(co, null, 2)
artistChoice.value = { id: props.source.artist_id, name: props.source.artist_name }
} else {
platform.value = platformsStore.list[0]?.key || 'patreon'
url.value = ''; enabled.value = true
structuredVideos.value = true; structuredSince.value = ''
takeConfig({})
jsonText.value = '{}'
artistChoice.value = props.initialArtist
? { id: props.initialArtist.id, name: props.initialArtist.name }
@@ -10,7 +10,7 @@
<div class="fc-source-row__url-wrap">
<a :href="source.url" target="_blank" rel="noopener" class="fc-source-row__url"
@click.stop>
{{ source.url }}
{{ source.display_name || source.url }}
</a>
<!-- Edit sits next to the source identity (operator-requested), not in
the action cluster where it was easy to fat-finger Remove. -->
@@ -72,7 +72,7 @@
worked, the content just isn't ours. -->
<v-chip
v-else-if="source.error_type === 'tier_limited'"
size="x-small" color="info" variant="tonal" label
size="x-small" color="info" variant="flat" label
prepend-icon="mdi-lock-outline"
>{{ source.tier_gated_count ? `${source.tier_gated_count} gated` : 'No access' }}
<v-tooltip activator="parent" location="top" max-width="480">
@@ -81,18 +81,18 @@
</v-chip>
<v-chip
v-else-if="source.backfill_state === 'running'"
size="x-small" color="info" variant="tonal" label
size="x-small" color="info" variant="flat" label
>{{ source.backfill_bypass_seen ? 'Recovering' : (source.backfill_recapture ? 'Recapturing' : 'Backfilling')
}}{{ source.backfill_posts
? ` · ${source.backfill_posts} posts`
: (source.backfill_chunks ? ` (${source.backfill_chunks})` : '') }}</v-chip>
<v-chip
v-else-if="source.backfill_state === 'complete'"
size="x-small" color="success" variant="tonal" label
size="x-small" color="success" variant="flat" label
>Backfilled</v-chip>
<v-chip
v-else-if="source.backfill_state === 'stalled'"
size="x-small" color="warning" variant="tonal" label
size="x-small" color="warning" variant="flat" label
>Stalled</v-chip>
<span v-else class="fc-source-row__zero">0</span>
</td>
@@ -108,7 +108,7 @@
v-if="item.singleSource"
:href="item.singleSource.url" target="_blank" rel="noopener"
class="fc-subs__sub-url" @click.stop
>{{ item.singleSource.url }}</a>
>{{ item.singleSource.display_name || item.singleSource.url }}</a>
</template>
<template #item.platforms="{ item }">
@@ -517,6 +517,7 @@ const filteredGroups = computed(() => {
g.sources.some(
(s) =>
(s.url || '').toLowerCase().includes(q) ||
(s.display_name || '').toLowerCase().includes(q) ||
(s.platform || '').toLowerCase().includes(q),
),
)
+4 -2
View File
@@ -119,7 +119,7 @@ export const useAdminStore = defineStore('admin', () => {
// --- Task progress polling (taps FC-3i activity dashboard) --------
/**
* Polls /api/system/activity/runs?queue=maintenance every 3s,
* Polls /api/system/activity/runs?celery_task_id=<id> every 3s,
* resolves when a task_run row with the given celery task_id
* reaches a terminal status (ok / error / timeout). Returns the
* row. Times out after 30 min by default.
@@ -129,7 +129,9 @@ export const useAdminStore = defineStore('admin', () => {
while (Date.now() < deadline) {
const body = await api.get(
'/api/system/activity/runs',
{ params: { queue: 'maintenance', limit: 20 } },
// By id, not by lane: these jobs run on `maintenance_long`, and a
// lane filter here is one more copy of the routing table (#4432).
{ params: { celery_task_id: taskId, limit: 1 } },
)
const row = (body.runs || []).find(r => r.celery_task_id === taskId)
if (row && ['ok', 'error', 'timeout'].includes(row.status)) {
+2 -2
View File
@@ -12,7 +12,7 @@ export const useCleanupStore = defineStore('cleanup', () => {
min_width: 0,
min_height: 0,
transparency_threshold: 0.9,
single_color_threshold: 0.95,
single_color_threshold: 0.995,
single_color_tolerance: 30,
})
@@ -24,7 +24,7 @@ export const useCleanupStore = defineStore('cleanup', () => {
min_width: s.min_width ?? 0,
min_height: s.min_height ?? 0,
transparency_threshold: s.transparency_threshold ?? 0.9,
single_color_threshold: s.single_color_threshold ?? 0.95,
single_color_threshold: s.single_color_threshold ?? 0.995,
single_color_tolerance: s.single_color_tolerance ?? 30,
}
}
+10
View File
@@ -149,6 +149,15 @@ export const useSourcesStore = defineStore('sources', () => {
return body
}
// #4486: a Discord source's posts by people outside its poster list.
// The preview touches nothing and hands back the token the remove must echo.
async function previewOtherPosters(id) {
return api.get(`/api/sources/${id}/discord/other-posters`)
}
async function removeOtherPosters(id, confirm) {
return api.post(`/api/sources/${id}/discord/other-posters/remove`, { body: { confirm } })
}
function sourcesByArtistGrouped() {
// returns [{artist: {id,name,slug}, sources: [...]}, ...]
const arr = byArtist.value.get(null) ?? []
@@ -179,6 +188,7 @@ export const useSourcesStore = defineStore('sources', () => {
stopBackfill,
recoverSource,
recaptureSource,
previewOtherPosters, removeOtherPosters,
findOrCreateArtist, autocompleteArtist, reassign,
loadScheduleStatus,
sourcesByArtistGrouped,
+8 -8
View File
@@ -210,8 +210,8 @@ def render(
)
parts.append(
f"Built from `{short}`. The rollback unit is the immutable `:c-` tag "
f"(rule 145) — these three move together:\n\n```\n"
f"Built from `{short}`. To roll back to this release, pull these "
f"immutable `:c-` tags — the images move together:\n\n```\n"
+ "\n".join(f"{image}:c-{short}" for image in IMAGES)
+ "\n```"
)
@@ -221,10 +221,10 @@ def render(
# truncated to MAX_COMMITS, which is 200 lines of internal build-out
# presented to someone who has never seen this project.
parts.append(
"---\n\n_First release under rule 148's `vYYYY.MM.DD.HHMM` shape, so "
"there is no predecessor to diff against and no changelog to derive. "
"The description above is README.md's, quoted at publish time. Later "
"releases carry the commits since the previous one._"
"---\n\n_The first release, so there is no earlier one to diff "
"against and no changelog to derive. The description above is "
"README.md's, quoted at publish time. Later releases carry the "
"commits since the previous one._"
)
return "\n\n".join(parts)
@@ -257,8 +257,8 @@ def cross_checks(tag: str, sha: str) -> list[str]:
if not RULE_148.match(tag):
notes.append(
f"`{tag}` is not rule 148's `vYYYY.MM.DD.HHMM` shape. Published "
f"anyway — the old `v26.*` tags predate the rule."
f"`{tag}` is not the `vYYYY.MM.DD.HHMM` release-tag shape. "
f"Published anyway — the old `v26.*` tags predate it."
)
else:
derived = artifact_version("web")
+22
View File
@@ -779,6 +779,28 @@ async def test_serve_extension_latest_returns_most_recent_xpi(
assert data == b"new"
@pytest.mark.asyncio
async def test_the_latest_alias_is_never_served_from_a_stale_cache(
client, monkeypatch, tmp_path,
):
"""One URL whose bytes change every release: a cached copy reinstalls the
previous build (operator-flagged 2026-09-25, when it was max-age=43200)."""
(tmp_path / "fabledcurator-1.0.1.xpi").write_bytes(b"new")
monkeypatch.setattr(frontend_module, "XPI_DIR", tmp_path)
resp = await client.get("/extension/fabledcurator-latest.xpi")
assert resp.headers["Cache-Control"] == "no-cache"
assert "Expires" not in resp.headers
@pytest.mark.asyncio
async def test_a_versioned_xpi_is_cached_for_good(client, monkeypatch, tmp_path):
"""A versioned name is one build's bytes forever."""
(tmp_path / "fabledcurator-1.0.1.xpi").write_bytes(b"new")
monkeypatch.setattr(frontend_module, "XPI_DIR", tmp_path)
resp = await client.get("/extension/fabledcurator-1.0.1.xpi")
assert "immutable" in resp.headers["Cache-Control"]
@pytest.mark.asyncio
async def test_serve_extension_latest_404_when_dir_empty(client, monkeypatch, tmp_path):
monkeypatch.setattr(frontend_module, "XPI_DIR", tmp_path)
+8
View File
@@ -295,3 +295,11 @@ async def test_failures_only_within_24h_window(client, _seed_failures):
body = await resp.get_json()
ids = {r["error_type"] for r in body["recent"]}
assert "OldError" not in ids
@pytest.mark.asyncio
async def test_runs_filter_by_celery_task_id(client, _seed_runs):
# How a page follows a job it started, without knowing its lane (#4432).
resp = await client.get("/api/system/activity/runs?celery_task_id=tid-3")
body = await resp.get_json()
assert [r["celery_task_id"] for r in body["runs"]] == ["tid-3"]
+45
View File
@@ -40,3 +40,48 @@ def test_single_color_evaluate_handles_rgba_input():
# Alpha channel should be ignored — only RGB matters for the rule.
im = Image.new("RGBA", (50, 50), (100, 100, 100, 128))
assert single_color.evaluate(im, threshold=0.9, tolerance=10) is True
# -- line art is a drawing, not a blank (#4483) ---------------------------------
_T = 0.995 # the default since alembic 0116
_TOL = 30
def _doodle(size=(1500, 2000), lines=12):
"""Pencil line art on white: a few thin dark strokes plus faint
construction lines — a couple of percent ink, the rest paper."""
from PIL import ImageDraw
im = Image.new("RGB", size, (255, 255, 255))
draw = ImageDraw.Draw(im)
w, h = size
for i in range(lines):
x = (i + 1) * w // (lines + 1)
draw.line([(x, 0), (w - x, h)], fill=(70, 60, 60), width=3)
draw.line([(0, (i + 1) * h // (lines + 1)), (w, h - x)], fill=(235, 235, 235), width=2)
return im
def test_line_art_on_white_is_not_single_color():
"""Five of Todding's Discord doodles were skipped as blank: a smoothing
64px thumbnail blended the strokes into the paper."""
assert single_color.evaluate(_doodle(), threshold=_T, tolerance=_TOL) is False
def test_a_blank_page_with_a_small_mark_is_still_single_color():
from PIL import ImageDraw
im = Image.new("RGB", (1000, 1000), (250, 250, 250))
ImageDraw.Draw(im).rectangle([480, 480, 520, 520], fill=(40, 40, 40))
assert single_color.evaluate(im, threshold=_T, tolerance=_TOL) is True
def test_a_noisy_solid_fill_is_still_single_color():
"""JPEG noise on a placeholder stays inside the tolerance."""
im = Image.new("RGB", (600, 600), (30, 90, 160))
for i in range(0, 600 * 600, 7):
x, y = i % 600, i // 600
d = (i % 21) - 10
im.putpixel((x, y), (30 + d, 90 + d, 160 + d))
assert single_color.evaluate(im, threshold=_T, tolerance=_TOL) is True
+55 -11
View File
@@ -2,6 +2,8 @@
`maintenance_long` lane so a 30-min backup or a multi-chunk audit can't starve
the quick recovery sweeps / vacuum on the concurrency-1 `maintenance` lane
(operator-flagged 2026-06-07)."""
import pytest
from backend.app.celery_app import celery
@@ -20,23 +22,65 @@ def test_quick_maintenance_stays_on_maintenance():
assert routes["backend.app.tasks.maintenance.*"]["queue"] == "maintenance"
def test_queue_for_mirrors_external_to_download():
"""celery_signals._queue_for is a hand-maintained mirror of task_routes
that stamps TaskRun.queue. external.* routes to the download lane, so the
mirror must agree — else TaskRun.queue lies 'default' for external fetches
and per-queue dashboard filters / threshold overrides miss them
(operator-flagged 2026-06-17)."""
@pytest.mark.parametrize(("name", "queue"), [
("backend.app.tasks.external.fetch_external_link", "download"),
# The rows #4432 found recorded on the wrong lane:
("backend.app.tasks.translation.translate_posts", "maintenance_long"),
("backend.app.tasks.gpu_queue.enqueue_gpu_backfill", "maintenance"),
("backend.app.tasks.maintenance.backfill_phash", "maintenance_long"),
("backend.app.tasks.maintenance.date_posts_from_records", "maintenance_long"),
("backend.app.tasks.admin.normalize_tags_task", "maintenance_long"),
("backend.app.tasks.backup.backup_db_task", "maintenance_long"),
("backend.app.tasks.maintenance.vacuum_analyze", "maintenance"),
("backend.app.tasks.not_routed.anything", "default"),
])
def test_task_run_records_the_queue_the_router_sends_to(name, queue):
"""TaskRun.queue comes from the router, not a copy of task_routes (#4432):
the stall sweep's per-queue thresholds and the System activity filters both
key off it."""
from backend.app.celery_signals import _queue_for
class _T:
name = "backend.app.tasks.external.fetch_external_link"
pass
assert _queue_for(_T()) == "download"
assert (
celery.conf.task_routes["backend.app.tasks.external.*"]["queue"]
== "download"
t = _T()
t.name = name
t.app = celery
assert _queue_for(t) == queue
def test_no_task_outlives_its_stall_threshold():
"""A task whose hard time limit is longer than the stall sweep's threshold
for it gets failed 'RecoverySweep' while it is still healthy — the class
#4432 found on the ml, import and long-maintenance lanes. The threshold is
resolved the way recover_stalled_task_runs resolves it: task-name override,
then queue, then the default."""
from backend.app.celery_signals import _queue_for
from backend.app.tasks.maintenance import (
QUEUE_STUCK_THRESHOLD_MINUTES,
STUCK_THRESHOLD_MINUTES,
TASK_STUCK_THRESHOLD_MINUTES,
)
celery.loader.import_default_modules()
checked = 0
too_short = []
for name, task in sorted(celery.tasks.items()):
if name.startswith("celery."):
continue
limit = getattr(task, "time_limit", None)
if not limit:
continue
threshold = TASK_STUCK_THRESHOLD_MINUTES.get(
name,
QUEUE_STUCK_THRESHOLD_MINUTES.get(_queue_for(task), STUCK_THRESHOLD_MINUTES),
)
checked += 1
if limit / 60 > threshold:
too_short.append(f"{name}: limit {limit / 60:.0f} min > sweep {threshold} min")
assert checked > 20, "the task registry did not load — this guard checked nothing"
assert not too_short, "\n".join(too_short)
def test_backfill_phash_runs_on_the_long_lane():
"""It lives in maintenance.py, so the quick-lane glob matches it too —
+148 -4
View File
@@ -153,9 +153,11 @@ def test_an_embeds_identity_ignores_its_signature():
def embed(sig):
return {"type": "image", "image": {"proxy_url": f"https://media/p/x.png?ex={sig}"}}
one = DiscordClient.extract_media(_msg(9, embeds=[embed("a")]))[0].media_id
two = DiscordClient.extract_media(_msg(9, embeds=[embed("b")]))[0].media_id
assert one == two and len(one) <= 33
art = [{"url": "https://cdn/a/1.png"}]
one = DiscordClient.extract_media(_msg(9, attachments=art, embeds=[embed("a")]))[1]
two = DiscordClient.extract_media(_msg(9, attachments=art, embeds=[embed("b")]))[1]
assert one.kind == two.kind == "embed"
assert one.media_id == two.media_id and len(one.media_id) <= 33
def test_post_seams():
@@ -165,11 +167,52 @@ def test_post_seams():
def test_a_text_only_message_is_not_a_post():
"""gallery-dl never made one: chat lines would bury the drops."""
"""Chat lines would bury the drops."""
assert DiscordClient.post_record_key(_msg(6, content="brb")) is None
assert DiscordClient.post_meta(_msg(1))["date"].startswith("2026-09-20")
def test_a_message_without_an_attached_image_is_chat():
"""Operator 2026-09-28: only content with an image attached. A lone
archive, a link preview or a Tenor GIF takes nothing and records nothing."""
chat = [
_msg(1, content="the stash", attachments=[
{"url": "https://cdn/a/Links_Stash.rar", "content_type": "application/x-rar"},
]),
_msg(2, content="look", embeds=[
{"type": "image", "image": {"proxy_url": "https://media/p/x.png"}},
]),
_msg(3, content="lol", embeds=[
{"type": "gifv", "video": {"proxy_url": "https://media/p/t.mp4"}},
]),
_msg(4, content="wip", attachments=[{"url": "https://cdn/a/wip.psd"}]),
]
for m in chat:
assert DiscordClient.extract_media(m) == []
assert DiscordClient.post_record_key(m) is None
def test_an_attached_image_takes_every_file_numbered_as_gallery_dl_does():
"""The gate decides WHETHER a message is taken, never which of its files:
the rar beside the image keeps its number, so on-disk names still match."""
m = _msg(7, attachments=[
{"url": "https://cdn/a/pack.rar"},
{"url": "https://cdn/a/noext", "content_type": "image/png"},
])
items = DiscordClient.extract_media(m)
assert [(i.num, i.filename) for i in items] == [(1, "pack"), (2, "noext")]
assert DiscordClient.post_record_key(m) == ("message:7", "7")
video = _msg(8, attachments=[{"url": "https://cdn/a/clip.MP4?ex=1"}])
assert len(DiscordClient.extract_media(video)) == 1
def test_a_forwarded_image_counts_as_attached():
m = _msg(9, message_snapshots=[
{"message": {"type": 0, "attachments": [{"url": "https://cdn/a/fwd.png"}]}},
])
assert DiscordClient.post_record_key(m) == ("message:9", "9")
# -- the walk --------------------------------------------------------------------
def test_a_channel_pages_newest_first_and_skips_system_messages():
@@ -206,6 +249,50 @@ def test_messages_carry_server_and_channel_metadata():
assert meta["parent"] == "Art"
def _one_channel(messages):
return {
"/guilds/1": _ok({"id": "1", "name": "Todding's Server"}),
"/guilds/1/channels": _ok([
{"id": "4", "type": 4, "name": "Todding"},
{"id": "2", "type": 0, "name": "banana-land", "parent_id": "4"},
]),
("/channels/2/messages", None): _ok(messages),
("/channels/2/threads/search", 0): _ok({"threads": []}),
}
def test_only_from_takes_the_creators_messages_and_no_one_elses():
"""Operator 2026-09-28: in the creator's server, another member posting a
meme is not the creator's art. Matched by id, username or display name."""
creator = {"id": "7", "username": "todding", "global_name": "Todding"}
other = {"id": "8", "username": "jakeboii", "global_name": "Jake Boii"}
msgs = [_msg(3, author=creator), _msg(2, author=other), _msg(1, author=creator)]
url = "https://discord.com/channels/1/2"
everyone = _client(_one_channel([dict(m) for m in msgs]))
assert [mid for mid, _ in _ids(everyone, url)] == ["3", "2", "1"]
for who in (["Todding"], ["TODDING "], ["7"]):
client = _client(_one_channel([dict(m) for m in msgs]))
client.only_from(who)
assert [mid for mid, _ in _ids(client, url)] == ["3", "1"]
client = _client(_one_channel([dict(m) for m in msgs]))
client.only_from([])
assert len(_ids(client, url)) == 3
def test_the_source_label_names_the_server_and_channel_the_walk_read():
client = _client(_one_channel([]))
url = "https://discord.com/channels/1/2"
assert client.source_label(url) is None # nothing walked yet
list(client.iter_posts(url))
assert client.source_label(url) == "Todding's Server · #banana-land"
assert client.source_label("https://discord.com/channels/1") == "Todding's Server"
client._channels["9"] = {"channel": "wip", "is_thread": True, "parent": "banana-land"}
assert client.source_label("https://discord.com/channels/1/9") == (
"Todding's Server · #banana-land › wip"
)
def test_a_server_walks_text_then_threads_newest_created_first_and_skips_private():
routes = {
"/guilds/1": _ok({"id": "1", "name": "S"}),
@@ -337,3 +424,60 @@ def test_the_adapter_authenticates_with_the_token_and_keys_by_identity(tmp_path)
_msg(9, attachments=[{"id": "300", "url": "https://cdn/a/1.png"}])
)
assert _ledger_key(media) == "9:300"
# -- the poster picker (#4488) -------------------------------------------------
def _person(pid, username, global_name=None, **extra):
return {"id": pid, "username": username, "global_name": global_name,
"messages": 1, "images": 0, **extra}
def test_the_creator_a_server_is_named_for_is_suggested():
ranked = dc.rank_posters("Todding's Server", "99", [
_person("8", "jakeboii", "Jake Boii", images=3),
_person("7", "todding", "Todding", images=1),
])
assert ranked[0]["id"] == "7" and ranked[0]["suggested"] is True
assert ranked[0]["reasons"] == ["name matches the server"]
assert ranked[1]["suggested"] is False
def test_owning_the_server_and_matching_its_name_outranks_either_alone():
ranked = dc.rank_posters("The Official Todding Discord", "7", [
_person("5", "toddingfan"),
_person("7", "t0dd", "Todding"),
])
assert [r["id"] for r in ranked if r["suggested"]] == ["7"]
assert ranked[0]["reasons"] == ["owns the server", "name matches the server"]
def test_no_signal_suggests_no_one_and_bots_never():
assert not any(r["suggested"] for r in dc.rank_posters("Art Club", None, [
_person("1", "alice"), _person("2", "bob"),
]))
ranked = dc.rank_posters("Todding's Server", None, [_person("3", "toddingbot", bot=True)])
assert ranked[0]["suggested"] is False
def test_recent_posters_tallies_a_shallow_window_by_id():
creator = {"id": "7", "username": "todding", "global_name": "Todding"}
other = {"id": "8", "username": "jakeboii", "global_name": "Jake Boii"}
art = [{"url": "https://cdn/a/1.png"}]
page = [_msg(i, author=creator, attachments=art) for i in range(300, 240, -1)]
page += [_msg(i, author=other) for i in range(240, 200, -1)]
routes = {
"/guilds/1": _ok({"id": "1", "name": "Todding's Server", "owner_id": "7"}),
("/channels/2/messages", None): _ok(page),
("/channels/2/messages", "201"): _ok([_msg(150, author=other)]),
}
client = _client(routes)
got = client.recent_posters("1", "2", max_messages=150)
assert got["server"] == "Todding's Server" and got["owner_id"] == "7"
assert got["scanned"] == 101
first = got["posters"][0]
assert (first["id"], first["messages"], first["images"], first["suggested"]) == ("7", 60, 60, True)
assert got["posters"][1]["messages"] == 41
# It asked for no more than the window: 100, then the 50 left of 150.
limits = [p.get("limit") for e, p in client._session.calls if e.endswith("/messages")]
assert limits == [100, 50]
+131
View File
@@ -0,0 +1,131 @@
"""Removing a Discord source's posts by people outside its poster list (#4486).
Real Postgres (db_sync): the predicates read `raw_metadata->>'author'`, which
only the database can evaluate. The test that matters runs preview and apply
on one fixture and asserts they agree AND that the rows went (snippet #3087).
"""
import pytest
from sqlalchemy import func, select
from backend.app.models import (
Artist,
ImageProvenance,
ImageRecord,
Post,
PostAttachment,
Source,
)
from backend.app.services import discord_poster_cleanup as cleanup
pytestmark = pytest.mark.integration
def _image(db, artist, tmp_path, name, post):
f = tmp_path / f"{name}.png"
f.write_bytes(b"x")
img = ImageRecord(
artist_id=artist.id, path=str(f), sha256=name.ljust(64, "0"),
size_bytes=1, mime="image/png", origin="downloaded", primary_post_id=post.id,
)
db.add(img)
db.flush()
return img
def _post(db, source, mid, raw, **extra):
p = Post(
source_id=source.id, artist_id=source.artist_id, external_post_id=mid,
raw_metadata=raw, **extra,
)
db.add(p)
db.flush()
return p
def _fixture(db, tmp_path, authors=("Todding",)):
"""Todding's post, two of Jake's (one sharing an image with Todding's), an
old post with no poster, and a drop that absorbed Todding's and Jake's."""
artist = Artist(name="Todding", slug="todding-cleanup")
db.add(artist)
db.flush()
src = Source(
artist_id=artist.id, platform="discord",
url="https://discord.com/channels/1/2",
config_overrides={"discord_authors": list(authors)} if authors else None,
)
db.add(src)
db.flush()
todd = _post(db, src, "10", {"author": "todding", "author_name": "Todding", "author_id": "1"})
jake1 = _post(db, src, "11", {"author": "jakeboii", "author_id": "8"})
jake2 = _post(db, src, "12", {"author": "jakeboii", "author_id": "8"})
old = _post(db, src, "13", {"category": "discord"})
drop = _post(db, src, "fc-drop:10", None, synthesized_by="discord_drop")
todd.absorbed_by_post_id = drop.id
jake1.absorbed_by_post_id = drop.id
art = _image(db, artist, tmp_path, "art", todd)
meme = _image(db, artist, tmp_path, "meme", jake1)
shared = _image(db, artist, tmp_path, "shared", jake2)
legacy = _image(db, artist, tmp_path, "legacy", old)
db.add_all([
ImageProvenance(image_record_id=shared.id, post_id=todd.id, source_id=src.id),
ImageProvenance(image_record_id=art.id, post_id=drop.id, source_id=src.id),
ImageProvenance(image_record_id=meme.id, post_id=drop.id, source_id=src.id),
PostAttachment(
post_id=jake2.id, artist_id=artist.id, sha256="z" * 64,
path="/store/z/f.zip", original_filename="f.zip", ext=".zip", size_bytes=1,
),
])
db.commit()
return src, {
"todd": todd.id, "jake1": jake1.id, "jake2": jake2.id, "old": old.id,
"drop": drop.id, "art": art.id, "meme": meme.id, "shared": shared.id,
"legacy": legacy.id,
}
def test_preview_matches_apply_and_the_rows_go(db_sync, tmp_path):
src, ids = _fixture(db_sync, tmp_path)
projected = cleanup.preview(db_sync, source_id=src.id)
assert projected["authors"] == ["todding"]
assert projected["posters"] == [{"poster": "jakeboii", "posts": 2}]
assert projected["unknown_posts"] == 1
done = cleanup.apply(db_sync, source_id=src.id, images_root=tmp_path)
assert projected["posts"] == done["posts_deleted"] == 2
assert projected["images"] == done["images_deleted"] == 1 # the meme only
assert projected["attachments"] == done["attachments_deleted"] == 1
assert projected["drops"] == done["drops_deleted"] == 1
post_ids = set(db_sync.execute(select(Post.id).where(Post.source_id == src.id)).scalars())
assert post_ids == {ids["todd"], ids["old"]}
image_ids = set(db_sync.execute(
select(ImageRecord.id).where(ImageRecord.id.in_(
[ids["art"], ids["meme"], ids["shared"], ids["legacy"]]))
).scalars())
# The meme was only Jake's. The shared piece is also on Todding's post, and
# the drop's link to the art was never what kept it.
assert image_ids == {ids["art"], ids["shared"], ids["legacy"]}
assert not (tmp_path / "meme.png").exists()
# Todding's message is back in the feed, for the sweep to regroup.
assert db_sync.get(Post, ids["todd"]).absorbed_by_post_id is None
assert db_sync.execute(select(func.count(PostAttachment.id)).where(
PostAttachment.sha256 == "z" * 64)).scalar_one() == 0
def test_a_poster_matches_by_display_name_or_id_too(db_sync, tmp_path):
src, _ = _fixture(db_sync, tmp_path, authors=("8",))
projected = cleanup.preview(db_sync, source_id=src.id)
assert projected["posters"] == [{"poster": "Todding", "posts": 1}]
def test_no_list_means_nothing_to_remove(db_sync, tmp_path):
src, _ = _fixture(db_sync, tmp_path, authors=())
with pytest.raises(cleanup.PosterCleanupError):
cleanup.preview(db_sync, source_id=src.id)
def test_the_token_names_what_the_preview_showed():
assert cleanup.confirm_token({"posts": 2, "images": 1}) == "remove-2-posts-1-images"
+71
View File
@@ -0,0 +1,71 @@
"""#4433: only the process consuming `download` clears a restart's leftovers.
The consolidated container boots four lanes side by side. If the scheduler or
ml lane ran this too, a lane restarting on its own (supervisord restarts a
crashed program) would close the events of walks that are still running.
"""
from __future__ import annotations
from types import SimpleNamespace
import pytest
from backend.app import celery_signals
def _consumer(*queues: str):
return SimpleNamespace(
task_consumer=SimpleNamespace(queues=[SimpleNamespace(name=q) for q in queues])
)
@pytest.fixture
def calls(monkeypatch):
seen: list[str] = []
class _Session:
def __enter__(self):
return self
def __exit__(self, *exc):
return False
def commit(self):
seen.append("commit")
monkeypatch.setattr(celery_signals, "sync_session_factory", lambda: _Session)
monkeypatch.setattr(
"backend.app.services.download_recovery.interrupt_orphaned_download_events",
lambda session, *, booted_at: seen.append("interrupt") or 0,
)
monkeypatch.setattr(
"backend.app.services.platform_lock.release_all_platform_locks",
lambda: seen.append("release") or 0,
)
return seen
def test_the_download_lane_clears_orphans_and_locks(calls):
celery_signals._on_worker_ready(sender=_consumer("default", "download", "import"))
assert calls == ["interrupt", "commit", "release"]
@pytest.mark.parametrize("queues", [("maintenance", "scan"), ("ml",), ("maintenance_long",)])
def test_other_lanes_leave_downloads_alone(calls, queues):
celery_signals._on_worker_ready(sender=_consumer(*queues))
assert calls == []
def test_unreadable_queues_do_nothing_rather_than_guess(calls):
celery_signals._on_worker_ready(sender=SimpleNamespace())
assert calls == []
def test_a_failure_does_not_stop_the_worker_starting(monkeypatch, calls):
def boom(session, *, booted_at):
raise RuntimeError("db down")
monkeypatch.setattr(
"backend.app.services.download_recovery.interrupt_orphaned_download_events", boom
)
celery_signals._on_worker_ready(sender=_consumer("download")) # must not raise
@@ -0,0 +1,106 @@
"""Migration 0114 (#4435): the undated shell post a misfiled attachment made on
the artist's first Discord source is folded into the real, dated post."""
import importlib.util
from datetime import UTC, datetime
from pathlib import Path
import pytest
from sqlalchemy import select
from backend.app.models import (
Artist,
ImageProvenance,
Post,
PostAttachment,
Source,
)
from tests.factories import make_image as _img
pytestmark = pytest.mark.integration
_MIGRATION = (
Path(__file__).resolve().parents[1]
/ "alembic" / "versions" / "0114_fold_misfiled_attachment_posts.py"
)
def _fold():
spec = importlib.util.spec_from_file_location("m0114", _MIGRATION)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
return mod.fold_misfiled_attachment_posts
def _source(db, artist, channel):
s = Source(
artist_id=artist.id, platform="discord",
url=f"https://discord.com/channels/1/{channel}",
)
db.add(s)
db.flush()
return s
def _post(db, artist, source, epid, when=None):
p = Post(
artist_id=artist.id, source_id=source.id, external_post_id=epid,
post_date=when,
post_url=f"https://discord.com/channels/1/x/{epid}" if when else None,
)
db.add(p)
db.flush()
return p
def _attach(db, post, sha, name):
db.add(PostAttachment(
post_id=post.id, artist_id=post.artist_id, sha256=sha,
path=f"/att/{sha}", original_filename=name, ext=".rar", size_bytes=1,
))
db.flush()
def test_shells_fold_into_their_dated_twin_and_nothing_else_moves(db_sync):
sent = datetime(2024, 12, 26, 1, 42, tzinfo=UTC)
artist = Artist(name="Yellow", slug="yellow")
db_sync.add(artist)
db_sync.flush()
first = _source(db_sync, artist, 100)
second = _source(db_sync, artist, 200)
# The #4435 shape: shell on the first source, real post on the second.
shell = _post(db_sync, artist, first, "555")
real = _post(db_sync, artist, second, "555", sent)
_attach(db_sync, shell, "a" * 64, "pack.rar")
# A shell whose attachment the real post already has: dropped, not doubled.
shell2 = _post(db_sync, artist, first, "556")
real2 = _post(db_sync, artist, second, "556", sent)
_attach(db_sync, shell2, "b" * 64, "same.rar")
_attach(db_sync, real2, "b" * 64, "same.rar")
# Left alone: no dated twin, and an undated post that holds an image.
lonely = _post(db_sync, artist, first, "557")
_attach(db_sync, lonely, "c" * 64, "lonely.rar")
with_image = _post(db_sync, artist, first, "558")
_post(db_sync, artist, second, "558", sent)
img = _img(db_sync, "d" * 64)
db_sync.add(ImageProvenance(image_record_id=img.id, post_id=with_image.id))
db_sync.flush()
shell_id, shell2_id = shell.id, shell2.id
folded = _fold()(db_sync.connection())
db_sync.expire_all()
assert folded == 2
assert db_sync.get(Post, shell_id) is None
assert db_sync.get(Post, shell2_id) is None
owners = dict(db_sync.execute(
select(PostAttachment.original_filename, PostAttachment.post_id)
).all())
assert owners["pack.rar"] == real.id
assert owners["lonely.rar"] == lonely.id
kept = db_sync.execute(
select(PostAttachment.id).where(PostAttachment.post_id == real2.id)
).scalars().all()
assert len(kept) == 1
assert db_sync.get(Post, lonely.id) is not None
assert db_sync.get(Post, with_image.id) is not None
+45
View File
@@ -435,3 +435,48 @@ def test_attach_in_place_non_media_routes_to_attachment(importer, db_sync):
).scalar_one()
assert row.ext == ".txt"
assert row.artist_id == artist.id
def test_a_non_media_file_lands_on_the_source_being_downloaded(importer, db_sync):
"""#4435: a Discord artist has one source per channel. A non-media file
downloaded for the SECOND channel was filed under the artist's first
Discord source (the (artist, platform) lookup takes the lowest id), as an
undated shell post; the real post record then made a second post under
the right source. The attachment must follow the source it was fetched for."""
from backend.app.models import Post, PostAttachment, Source
images_root = importer.images_root
artist = Artist(name="Yara", slug="yara")
db_sync.add(artist)
db_sync.flush()
first = Source(
artist_id=artist.id, platform="discord",
url="https://discord.com/channels/1/100",
)
second = Source(
artist_id=artist.id, platform="discord",
url="https://discord.com/channels/1/200",
)
db_sync.add_all([first, second])
db_sync.flush()
rar = images_root / "yara" / "discord" / "rewards" / "20241226_555_01_pack.rar"
rar.parent.mkdir(parents=True, exist_ok=True)
rar.write_bytes(b"not a real archive, so it is kept as an attachment")
rar.with_suffix(rar.suffix + ".json").write_text(
'{"category": "discord", "id": "555", "message_id": "555"}'
)
result = importer.attach_in_place(rar, artist=artist, source=second)
assert result.status == "attached"
owner = db_sync.execute(
select(Post.source_id, Post.external_post_id)
.join(PostAttachment, PostAttachment.post_id == Post.id)
.where(PostAttachment.original_filename == rar.name)
).one()
assert owner.source_id == second.id
assert owner.external_post_id == "555"
assert db_sync.execute(
select(func.count()).select_from(Post).where(Post.source_id == first.id)
).scalar_one() == 0
+70 -14
View File
@@ -335,26 +335,30 @@ def test_recover_stalled_task_runs_ml_queue_uses_longer_threshold(db_sync):
"""ml-queue tasks (embed_image video branch) legitimately run
past the default 5-min threshold. The sweep must NOT flag an
ml-queue task that's only been running 10 min — the override
threshold (25 min via QUEUE_STUCK_THRESHOLD_MINUTES) protects
in-flight video tagging. Operator-flagged 2026-05-28 after
image 6288 (mp4) was marked failed at the 5-min tick mid-run."""
threshold (QUEUE_STUCK_THRESHOLD_MINUTES["ml"]) protects in-flight
video tagging. Operator-flagged 2026-05-28 after image 6288 (mp4)
was marked failed at the 5-min tick mid-run. Read from the table
rather than restated: the value moved 25 -> 40 in #4432."""
from sqlalchemy import select
from backend.app.models import TaskRun
from backend.app.tasks.maintenance import recover_stalled_task_runs
from backend.app.tasks.maintenance import (
QUEUE_STUCK_THRESHOLD_MINUTES,
recover_stalled_task_runs,
)
ml_threshold = QUEUE_STUCK_THRESHOLD_MINUTES["ml"]
now = datetime.now(UTC)
# 10-min-old ml-queue row: stale by the default 5-min rule but
# fresh by the 25-min ml override. Must survive the sweep.
# fresh by the ml override. Must survive the sweep.
ml_fresh_id = _make_task_run(
db_sync, status="running", queue="ml",
started_at=now - timedelta(minutes=10),
)
# 30-min-old ml-queue row: past even the ml override. Must be
# flagged.
# Past even the ml override. Must be flagged.
ml_stale_id = _make_task_run(
db_sync, status="running", queue="ml",
started_at=now - timedelta(minutes=30),
started_at=now - timedelta(minutes=ml_threshold + 5),
)
db_sync.commit()
@@ -431,9 +435,10 @@ def test_download_stuck_threshold_exceeds_hard_time_limit():
def test_recover_stalled_task_runs_archive_task_uses_longer_threshold(db_sync):
"""import_archive_file shares the 'import' queue with fast
single-file import_media_file, so it gets a per-task-name override
(40 min) while the import queue stays at the 5-min default. A
10-min-old archive task-run must survive; a 50-min-old one is
flagged. Operator-flagged 2026-05-28."""
(40 min) while the import queue keeps its short threshold (10 min
since #4432; import_media_file's hard limit is 6). A 10-min-old
archive task-run must survive; a 50-min-old one is flagged.
Operator-flagged 2026-05-28."""
from sqlalchemy import select
from backend.app.models import TaskRun
@@ -441,12 +446,12 @@ def test_recover_stalled_task_runs_archive_task_uses_longer_threshold(db_sync):
archive_name = "backend.app.tasks.import_file.import_archive_file"
now = datetime.now(UTC)
# Fast single-file import on the same queue, 10 min old → flagged
# by the default 5-min rule.
# Fast single-file import on the same queue, 15 min old → flagged
# by the import queue's 10-min threshold.
media_id = _make_task_run(
db_sync, status="running", queue="import",
task_name="backend.app.tasks.import_file.import_media_file",
started_at=now - timedelta(minutes=10),
started_at=now - timedelta(minutes=15),
)
# Archive on the same queue, 10 min old → survives (40-min override).
archive_fresh_id = _make_task_run(
@@ -713,6 +718,57 @@ def test_recover_stalled_download_skips_fresh(db_sync):
assert failures == 0
def test_download_lane_boot_closes_pre_boot_events_without_blaming_the_source(db_sync):
"""#4433: a restart SIGKILLs a walk past its stop grace and strands queued
ones. At the download lane's boot, everything pending/running from before
it ends as `skipped`/interrupted — and the source is left as it was, so it
is due on the next tick instead of backed off as a failure."""
from sqlalchemy import select
from backend.app.models import DownloadEvent, Source
from backend.app.services.download_recovery import (
DOWNLOAD_INTERRUPTED_MESSAGE,
interrupt_orphaned_download_events,
)
sid = _make_source(db_sync, slug="rebooted")
booted_at = datetime.now(UTC)
before = booted_at - timedelta(minutes=5)
walking = DownloadEvent(
source_id=sid, status="running", started_at=before,
metadata_={"live": {"downloaded": 3}},
)
queued = DownloadEvent(source_id=sid, status="pending", started_at=before)
finished = DownloadEvent(source_id=sid, status="ok", started_at=before)
# Promoted to running after the boot: download_service resets started_at.
fresh = DownloadEvent(
source_id=sid, status="running", started_at=booted_at + timedelta(seconds=5),
)
db_sync.add_all([walking, queued, finished, fresh])
db_sync.commit()
closed = interrupt_orphaned_download_events(db_sync, booted_at=booted_at)
db_sync.commit()
assert closed == 2
db_sync.expire_all()
for ev in (walking, queued):
assert ev.status == "skipped"
assert ev.error == DOWNLOAD_INTERRUPTED_MESSAGE
assert ev.finished_at is not None
assert ev.metadata_["error_type"] == "interrupted"
assert walking.metadata_["live"] == {"downloaded": 3}
assert finished.status == "ok"
assert fresh.status == "running"
src = db_sync.execute(
select(Source.consecutive_failures, Source.last_error, Source.last_checked_at)
.where(Source.id == sid)
).one()
assert src.consecutive_failures == 0
assert src.last_error is None
assert src.last_checked_at is None
def test_recover_stalled_download_flips_stale_pending(db_sync):
"""A 2-hour-old pending event flips to error AND the source is bumped
(consecutive_failures, last_error, last_checked_at) so the next scan
+16 -7
View File
@@ -239,8 +239,8 @@ async def test_tick_downloads_unseen_and_marks_seen(source_id, sync_engine, tmp_
# plan #704: structured run_stats carry the real counts.
assert result.run_stats["downloaded_count"] == 2
assert result.posts_processed == 1
# The media wait for phase 3 to import them; only the post key is in yet.
assert _count_ledger(sync_engine, source_id) == 1
# The media and the post record both wait for phase 3 to import them (#4436).
assert _count_ledger(sync_engine, source_id) == 0
result.mark_seen_after_import()
# 2 media keys + 1 synthetic post key (body/links recaptured per post).
assert _count_ledger(sync_engine, source_id) == 3
@@ -648,8 +648,10 @@ async def test_recovery_tier2_disk_still_skips(source_id, sync_engine, tmp_path)
assert result.files_downloaded == 0
assert downloader.download_calls == 0
assert result.written_paths == []
# Disk-skip reconciles the media key + the synthetic post key (recovery
# recaptures the body/links per post) = 2.
# Disk-skip reconciles the media key at once; the synthetic post key
# (recovery recaptures the body/links per post) waits for phase 3 (#4436).
assert _count_ledger(sync_engine, source_id) == 1
result.mark_seen_after_import()
assert _count_ledger(sync_engine, source_id) == 2
@@ -943,8 +945,10 @@ async def test_a_run_that_dies_before_import_leaves_its_media_unmarked(
_FakeDownloader(tmp_path))
ing.run(source_id=source_id, campaign_id="c1", artist_slug="ingest",
url="https://patreon.com/ingest", mode="tick")
# Phase 3 never ran, so `mark_seen_after_import` never did: only the post key.
assert _count_ledger(sync_engine, source_id) == 1
# Phase 3 never ran, so `mark_seen_after_import` never did: neither the
# media nor the post record is marked (#4436 — the record, marked at write
# time, left the post undated for good once the run died before phase 3).
assert _count_ledger(sync_engine, source_id) == 0
# The next walk finds the file on disk with no record, and imports it.
ing2 = _ingester(sync_engine, tmp_path, _FakeClient([(None, [("p1", [m1])])]),
@@ -954,6 +958,7 @@ async def test_a_run_that_dies_before_import_leaves_its_media_unmarked(
assert result.written_paths == [str(tmp_path / "p1_1.jpg")]
assert result.files_downloaded == 0 # not fetched again
assert "on disk but never imported: p1_1.jpg" in result.stdout
assert len(result.post_record_paths) == 1 # the record is written again
result.mark_seen_after_import()
assert _count_ledger(sync_engine, source_id) == 2
@@ -1200,7 +1205,10 @@ async def test_tick_captures_media_less_post_once(source_id, sync_engine, tmp_pa
assert result.success is True
assert len(result.post_record_paths) == 1
assert downloader.post_records == 1
# The synthetic `post:ptext` key was marked seen (gates re-capture).
# The synthetic `post:ptext` key is marked once phase 3 has upserted the
# record (#4436), and then gates re-capture.
assert _count_ledger(sync_engine, source_id) == 0
result.mark_seen_after_import()
assert _count_ledger(sync_engine, source_id) == 1
# Second walk: already recorded → gated, no re-write, no new ledger row.
@@ -1540,6 +1548,7 @@ async def test_revisits_do_not_feed_the_body_drift_canary(
url="https://patreon.com/ingest", mode="tick", revisit_days=30,
)
assert first.success is True
first.mark_seen_after_import() # phase 3 ran
# Second walk: every post is a revisit, and every body comes back empty.
client2 = _FakeClient([(None, posts)], published=published, empty_body=True)
+21
View File
@@ -283,6 +283,27 @@ async def test_scroll_item_shape_minimal(db):
assert "description_full" not in item
@pytest.mark.asyncio
async def test_a_discord_post_names_its_channel(db):
"""#4481: the card says which channel a message came from. Read from the
record's `channel`; never on another platform, even with the same key."""
artist = await _seed_artist(db, "todding-ch")
dsrc = await _seed_source(db, artist.id, "discord", "https://discord.com/channels/1/2")
psrc = await _seed_source(db, artist.id, "patreon", "https://p/todding-ch")
now = datetime.now(UTC)
d = await _seed_post(db, dsrc.id, external_id="D1", post_date=now)
d.raw_metadata = {"category": "discord", "channel": "banana-land"}
blank = await _seed_post(db, dsrc.id, external_id="D2", post_date=now)
blank.raw_metadata = {"category": "discord", "channel": " "}
p = await _seed_post(db, psrc.id, external_id="P1", post_date=now)
p.raw_metadata = {"channel": "not-discord"}
await db.commit()
page = await PostFeedService(db).scroll(cursor=None, limit=10, artist_id=artist.id)
by_id = {it["external_post_id"]: it["channel"] for it in page["items"]}
assert by_id == {"D1": "banana-land", "D2": None, "P1": None}
@pytest.mark.asyncio
async def test_scroll_surfaces_translation_fields(db):
# #143: a translated post exposes the translated title/description + source
+3 -39
View File
@@ -126,53 +126,17 @@ def test_process_sweep_flags_conflict_with_process_mode(db_sync):
def test_process_auto_source_never_trains_head(db_sync):
# The runaway break: provisional wip tags (process sweep 'process_auto', soft
# title 'wip_title_soft') are NOT training positives; a HARD title-heuristic /
# manual one IS. So the head learns only from trusted labels, never its own
# output or the low-precision sketch/doodle tier (#1464 + #1474).
# The runaway break: provisional wip tags (process sweep 'process_auto') are NOT
# training positives; a title-heuristic / manual one IS. So the head learns only
# from trusted labels, never its own output (#1464).
wip = _system_tag(db_sync, "wip")
auto_img = _img(db_sync, "f" * 64, _emb(0))
soft_img = _img(db_sync, "9" * 64, _emb(2))
title_img = _img(db_sync, "0" * 64, _emb(1))
db_sync.execute(image_tag.insert().values(
image_record_id=auto_img.id, tag_id=wip.id, source="process_auto"))
db_sync.execute(image_tag.insert().values(
image_record_id=soft_img.id, tag_id=wip.id, source="wip_title_soft"))
db_sync.execute(image_tag.insert().values(
image_record_id=title_img.id, tag_id=wip.id, source="wip_title"))
db_sync.commit()
positives = set(_ids_with_tag(db_sync, wip.id))
assert title_img.id in positives # trusted HARD label trains the head
assert auto_img.id not in positives # its own auto-applied output does NOT
assert soft_img.id not in positives # low-precision soft tier does NOT
def test_soft_wip_conflict_audit_flags_ring_loud(db_sync):
# A soft-tagged image (sketch/doodle title) that ALSO scores high on a content
# head is probably finished art mis-tagged — flagged for review; a quiet one is not.
from backend.app.services.ml.heads import soft_wip_conflict_audit
s = db_sync.execute(select(MLSettings).where(MLSettings.id == 1)).scalar_one()
s.process_conflict_threshold = 0.6
wip = _system_tag(db_sync, "wip")
content = Tag(name="looksreal", kind=TagKind.general)
db_sync.add(content)
db_sync.flush()
_head(db_sync, content.id, 0, weight=1.0) # sigmoid(1)=0.73 > 0.6 conflict
ring = _img(db_sync, "1" * 64, _emb(0)) # scores on the content head
quiet = _img(db_sync, "2" * 64, _emb(5)) # orthogonal → 0.5 < 0.6
for img in (ring, quiet):
db_sync.execute(image_tag.insert().values(
image_record_id=img.id, tag_id=wip.id, source="wip_title_soft"))
db_sync.commit()
res = soft_wip_conflict_audit(db_sync)
assert res["n_flagged"] == 1
flag = db_sync.execute(
select(PresentationReview).where(PresentationReview.image_record_id == ring.id)
).scalar_one()
assert flag.mode == "process"
assert flag.conflict_tag_id == content.id
assert db_sync.execute(
select(PresentationReview).where(PresentationReview.image_record_id == quiet.id)
).scalar_one_or_none() is None
+70
View File
@@ -0,0 +1,70 @@
"""Migration 0113 (#4431): images linked to a post before its date arrived get
the post's date back. Runs the migration's data step against real rows."""
import importlib.util
from datetime import UTC, datetime
from pathlib import Path
import pytest
from backend.app.models import Artist, ImageProvenance, ImageRecord, Post
from tests.factories import make_image as _img
pytestmark = pytest.mark.integration
_MIGRATION = (
Path(__file__).resolve().parents[1]
/ "alembic" / "versions" / "0113_redate_native_images.py"
)
def _redate():
spec = importlib.util.spec_from_file_location("m0113", _MIGRATION)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
return mod.redate_images
def _post(db, artist, epid, when):
p = Post(artist_id=artist.id, external_post_id=epid, post_date=when)
db.add(p)
db.flush()
return p
def _link(db, img, post, *, primary):
db.add(ImageProvenance(image_record_id=img.id, post_id=post.id))
if primary:
img.primary_post_id = post.id
db.flush()
def test_redate_images_from_their_posts(db_sync):
sent = datetime(2024, 3, 1, 18, 30, tzinfo=UTC)
earlier = datetime(2023, 1, 5, tzinfo=UTC)
artist = Artist(name="Alice", slug="alice")
db_sync.add(artist)
db_sync.flush()
undated = _post(db_sync, artist, "u1", None)
dated = _post(db_sync, artist, "d1", sent)
repost = _post(db_sync, artist, "r1", earlier)
stale = _img(db_sync, "a" * 64) # primary post dated, image not
reposted = _img(db_sync, "b" * 64) # also in an earlier post
orphan = _img(db_sync, "c" * 64) # only an undated post
_link(db_sync, stale, dated, primary=True)
_link(db_sync, reposted, dated, primary=True)
_link(db_sync, reposted, repost, primary=False)
_link(db_sync, orphan, undated, primary=True)
orphan_before = orphan.effective_date
_redate()(db_sync.connection())
db_sync.expire_all()
stale = db_sync.get(ImageRecord, stale.id)
reposted = db_sync.get(ImageRecord, reposted.id)
orphan = db_sync.get(ImageRecord, orphan.id)
assert stale.effective_date == sent
assert stale.earliest_post_date == sent
assert reposted.effective_date == sent # the primary post's date
assert reposted.earliest_post_date == earlier # the earliest post's date
assert orphan.effective_date == orphan_before # no date to take
+10
View File
@@ -14,6 +14,7 @@ re-implementation of it.
"""
from __future__ import annotations
import re
import subprocess
from pathlib import Path
@@ -162,6 +163,15 @@ def test_the_first_release_describes_the_product_instead_of_diffing(shaped_histo
assert not [ln for ln in body.split("\n") if ln.startswith("- work landing")]
@pytest.mark.parametrize("tag", ["v2026.08.28.2208", "v2026.08.29.1000"])
def test_the_release_page_cites_no_internal_rule_numbers(shaped_history, tag):
"""The release page is read by strangers. "rule 145" names a record in the
operator's own notes, which a reader cannot open — say what the rule means
instead. Covers the first-release overview and the changelog body."""
body = body_of(notes(tag, cwd=shaped_history))
assert not re.search(r"\brule\s+\d+", body, re.IGNORECASE), body
def test_the_overview_is_readmes_words_not_a_second_copy(shaped_history):
"""Two hand-maintained descriptions of one product drift and nothing
catches it. The release page quotes README.md so there is one source."""
+90
View File
@@ -0,0 +1,90 @@
"""Migration 0112 (milestone 430, #4428): retiring the sketch/doodle WIP title tier.
Runs the migration's data step against real rows. The soft tags the operator stood
behind (confirmed, or kept through the review strip) survive as `manual`; the rest go,
and so do the unresolved review cards whose tag went with them. Hard title tags and
their flags are untouched.
"""
import importlib.util
from datetime import UTC, datetime
from pathlib import Path
import pytest
from sqlalchemy import select
from backend.app.models import PresentationReview, TagPositiveConfirmation
from backend.app.models.tag import image_tag
from backend.app.services.wip_title import resolve_wip_tag_id
from tests.factories import make_image as _img
pytestmark = pytest.mark.integration
_MIGRATION = (
Path(__file__).resolve().parents[1]
/ "alembic" / "versions" / "0112_retire_soft_wip_title.py"
)
def _retire():
spec = importlib.util.spec_from_file_location("m0112", _MIGRATION)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
return mod.retire_soft_wip_tags
def _source(db, image_id, tag_id):
return db.execute(
select(image_tag.c.source)
.where(image_tag.c.image_record_id == image_id)
.where(image_tag.c.tag_id == tag_id)
).scalar_one_or_none()
def _review(db, image_id, tag_id):
return db.execute(
select(PresentationReview)
.where(PresentationReview.image_record_id == image_id)
.where(PresentationReview.tag_id == tag_id)
).scalar_one_or_none()
def _flag(db, image_id, tag_id, *, resolved):
db.add(PresentationReview(
image_record_id=image_id, tag_id=tag_id, conflict_score=0.8, mode="process",
resolved_at=datetime.now(UTC) if resolved else None,
))
def test_retire_soft_wip_keeps_human_judged_and_clears_the_rest(db_sync):
wip = resolve_wip_tag_id(db_sync)
plain = _img(db_sync, "a" * 64) # soft, never looked at → removed
confirmed = _img(db_sync, "b" * 64) # soft + confirmed → kept as manual
kept = _img(db_sync, "c" * 64) # soft + "Keep tag" in the strip → manual
pending = _img(db_sync, "d" * 64) # soft + open card → tag and card removed
hard = _img(db_sync, "e" * 64) # hard title tag + open card → untouched
for img in (plain, confirmed, kept, pending):
db_sync.execute(image_tag.insert().values(
image_record_id=img.id, tag_id=wip, source="wip_title_soft"))
db_sync.execute(image_tag.insert().values(
image_record_id=hard.id, tag_id=wip, source="wip_title"))
db_sync.add(TagPositiveConfirmation(image_record_id=confirmed.id, tag_id=wip))
_flag(db_sync, kept.id, wip, resolved=True)
_flag(db_sync, pending.id, wip, resolved=False)
_flag(db_sync, hard.id, wip, resolved=False)
db_sync.flush()
_retire()(db_sync.connection())
db_sync.expire_all()
assert _source(db_sync, plain.id, wip) is None
assert _source(db_sync, confirmed.id, wip) == "manual"
assert _source(db_sync, kept.id, wip) == "manual"
assert _source(db_sync, pending.id, wip) is None
assert _source(db_sync, hard.id, wip) == "wip_title"
assert _review(db_sync, pending.id, wip) is None
assert _review(db_sync, kept.id, wip) is not None # resolved history stays
assert _review(db_sync, hard.id, wip) is not None # its tag is still on
assert db_sync.execute(
select(image_tag.c.image_record_id).where(image_tag.c.source == "wip_title_soft")
).first() is None
+81
View File
@@ -382,3 +382,84 @@ def test_external_links_not_duplicated_on_reimport(importer, import_layout):
assert importer.session.execute(
select(func.count()).select_from(ExternalLink)
).scalar_one() == 1
def test_post_record_redates_images_linked_before_it(importer, import_layout):
"""#4431: the native ingesters import a message's media before its record,
and only the record carries the date. The images start on their download
time; when the record lands they take the post's date."""
import_root, _ = import_layout
artist = Artist(name="Alice", slug="alice")
importer.session.add(artist)
importer.session.flush()
m = import_root / "Alice" / "20240301_123_01_art.jpg"
_split(m, "v")
_sidecar(m, {"category": "discord", "message_id": "123"})
r = importer.import_one(m)
assert r.status == "imported"
rec = importer.session.get(ImageRecord, r.image_id)
post = importer.session.execute(select(Post)).scalar_one()
assert post.post_date is None
download_time = rec.effective_date
sc = import_root / "Alice" / "20240301_123_post.json"
sc.write_text(json.dumps({
"category": "discord", "message_id": "123", "message": "",
"date": "2024-03-01T18:30:00.000000+00:00",
}))
assert importer.upsert_post_record(sc, artist=artist) is True
importer.session.expire_all()
rec = importer.session.get(ImageRecord, r.image_id)
post = importer.session.execute(select(Post)).scalar_one()
assert post.post_date is not None
assert post.post_date != download_time
assert rec.effective_date == post.post_date
assert rec.earliest_post_date == post.post_date
def test_an_undated_post_is_dated_from_the_record_its_walk_left_on_disk(
importer, import_layout,
):
"""#4436: a walk killed before phase 3 wrote the message's record to disk
but never upserted it, and marked it seen, so no later tick fixed it. The
repair finds the record under the artist's folder and dates the post — and
its images — under the post's own source."""
from backend.app.services.post_record_repair import date_posts_from_records
import_root, images_root = import_layout
artist = Artist(name="Alice", slug="alice")
importer.session.add(artist)
importer.session.flush()
importer.session.add(Source(
artist_id=artist.id, platform="discord",
url="https://discord.com/channels/1/100",
))
importer.session.flush()
m = import_root / "Alice" / "20240301_123_01_art.jpg"
_split(m, "v")
_sidecar(m, {"category": "discord", "message_id": "123"})
r = importer.import_one(m)
assert r.status == "imported"
post = importer.session.execute(select(Post)).scalar_one()
assert post.post_date is None and post.source_id is not None
channel = images_root / "alice" / "discord" / "rewards"
channel.mkdir(parents=True)
(channel / "20240301_123_post.json").write_text(json.dumps({
"category": "discord", "message_id": "123", "message": "",
"date": "2024-03-01T18:30:00.000000+00:00",
}))
(channel / "20240301_999_post.json").write_text("not json") # skipped, not fatal
summary = date_posts_from_records(importer.session, importer, images_root)
assert summary == {"undated": 1, "dated": 1, "no_record": 0}
importer.session.expire_all()
post = importer.session.execute(select(Post)).scalar_one()
rec = importer.session.get(ImageRecord, r.image_id)
assert post.post_date is not None
assert rec.earliest_post_date == post.post_date
# Nothing left undated: the next sweep is an empty query.
assert date_posts_from_records(importer.session, importer, images_root) == {
"undated": 0, "dated": 0, "no_record": 0,
}
+1 -27
View File
@@ -8,7 +8,7 @@ negative cases (substrings like ``swipe`` / ``wiped``) are the load-bearing ones
"""
import pytest
from backend.app.services.wip_title import matches_soft_wip_title, matches_wip_title
from backend.app.services.wip_title import matches_wip_title
@pytest.mark.parametrize("title", [
@@ -48,29 +48,3 @@ def test_matches_positive(title):
])
def test_matches_negative(title):
assert matches_wip_title(title) is False
@pytest.mark.parametrize("title", [
"quick sketch",
"Nami sketch",
"morning doodle",
"some doodles",
"sketches from today",
"a little scribble",
"SKETCH",
])
def test_soft_matches_positive(title):
assert matches_soft_wip_title(title) is True
@pytest.mark.parametrize("title", [
None,
"",
"sketchbook tour", # 'sketch' inside sketchbook — must NOT match
"kadoodle mascot", # 'doodle' mid-word
"the final piece",
"WIP", # a HARD cue is not a SOFT cue
"prescribed colours", # 'scrib' inside prescribed — must NOT match
])
def test_soft_matches_negative(title):
assert matches_soft_wip_title(title) is False
-12
View File
@@ -10,7 +10,6 @@ from backend.app.celery_app import celery
from backend.app.models import Artist, ImageProvenance, ImageRecord, Post, Source
from backend.app.models.tag import image_tag
from backend.app.services.wip_title import (
WIP_TITLE_SOFT_SOURCE,
apply_wip_image_tags,
resolve_wip_tag_id,
)
@@ -108,14 +107,3 @@ def test_backfill_tags_only_wip_titled_posts(db_sync):
# Idempotent: a second sweep finds the tag already present and applies nothing.
assert backfill_wip_title_tags.apply().get() == 0
def test_apply_soft_source_stamps_wip_title_soft(db_sync):
# The soft tier (#1474) stamps a distinct provisional source.
tag_id = resolve_wip_tag_id(db_sync)
rec = _img(db_sync)
db_sync.commit()
assert apply_wip_image_tags(
db_sync, [rec.id], tag_id, source=WIP_TITLE_SOFT_SOURCE
) == 1
assert _wip_source(db_sync, rec.id, tag_id) == "wip_title_soft"