Compare commits

...
Author SHA1 Message Date
bvandeusenandClaude Opus 5 c2f9e9cc08 docs: stop claiming pixiv, and stop claiming everything gallery-dl supports (406 step 4)
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 2s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 24s
CI / backend-lint-and-test (push) Failing after 53s
extension / lint (push) Successful in 22s
Build images / build-ml (push) Successful in 2m10s
CI / integration (push) Failing after 2m33s
Build images / sign-extension (push) Successful in 5m33s
Build images / build-web (push) Successful in 1m7s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
Ships with the switch-off rather than with the code removal: a doc that promises a platform the product refuses is the Install and Public Surface area's characteristic defect.

pixiv comes out of README (twice), SECURITY.md (twice), .env.example and the compose header. The stored-credential warnings now name Patreon and SubscribeStar - the accounts that usually carry a payment method.

One correction beyond pixiv. README said FabledCurator follows creators on Patreon, SubscribeStar, Pixiv "and anything gallery-dl supports". That was already false: a platform not in the registry is rejected, however capable gallery-dl is. It now names the real set, which rule 171 records: Patreon, SubscribeStar, Discord and HentaiFoundry.

The 3422 docs guards still hold - the key path and bootstrap variable are untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SHQB1YukL3VyvMK8rcbmV9
2026-09-13 11:45:49 -04:00
bvandeusenandClaude Opus 5 e3fd8c67d4 feat: switch pixiv off — unregistered, unreachable, and refused at dispatch (406 phase 1)
Milestone 406 retires pixiv (rule 171) in two phases at the operator's explicit ask: switch it off, then later delete its code. This is the switch-off. Steps 2 and 3 ship together because each is a half-state of the other: unregistered but still in the extension, pixiv creator pages would offer a button the backend then refuses.

Reachability removed, never gated (rule 22 - no flag, no `if platform == "pixiv"`):
- platforms registry: pixiv unregistered, so /api/platforms, the source validator and quick-add all refuse it through their existing unknown-platform paths.
- NATIVE_INGESTER_PLATFORMS: pixiv removed.
- extension_service: pixiv's quick-add URL pattern removed (the Python half of the JS mirror).
- extension: pixiv's host permissions, content-script match, platform entry and artist pattern removed; popup's pixiv branches removed; and the whole pixiv PKCE OAuth flow cut out of background.js. That last one could not wait for phase 2 - a webRequest listener on a host the manifest no longer grants is at best dead and at worst a startup failure for the entire background script. On startup the extension now also removes any pixiv refresh token a browser still holds in storage, for the same reason as the server-side credential cleanup (3980).
- frontend: the extension card stops listing pixiv; SourceActions' copy of the native list drops it. platformColor keeps rendering a pixiv key so existing pixiv posts do not look broken.

The guard, and why a registry change alone was not enough. A source outlives its platform: the live instance still had one ENABLED pixiv source (step 1). Tracing it: the scheduler only selects enabled rows and every platform lookup uses .get(), so a disabled row is inert - but re-enabling it and pressing Check would have routed pixiv, no longer native, straight into the gallery-dl branch, which still has a pixiv extractor. And a worker can pick up a still-enabled row before a deploy's migration runs. So run_download and verify_source_credential - the two functions every download and credential probe pass through - now refuse any platform not in the registry: an unsupported_url failure for downloads, and an inconclusive (None, not False) verify, since nothing was probed so nothing was rejected. Generic by registration, so it covers deviantart's leftovers too. Positive-controlled: a supported gallery-dl platform must still reach gallery-dl, or a guard that refused everything would pass (rule 167).

Migration 0097 disables sources on retired platforms (pixiv, deviantart) and clears their failure state exactly as disabling through the app does (1285), so the stale row stops being scheduled and stops showing as failing. Nothing is deleted: removing a source can collide with uq_post_artist_external_id_null_source on real data, which is phase 2's step 6 to check. No post or image is touched.

Tests: the known-platform lists drop pixiv and gain retirement assertions beside deviantart's; pixiv's positive extension cases become negative guards; the pixiv sidecar post-URL test is deleted with the behaviour it tested; quick-add rejects a pixiv URL. The pixiv client/downloader/ingester suites stay - that code stays until phase 2.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SHQB1YukL3VyvMK8rcbmV9
2026-09-13 11:45:49 -04:00
bvandeusenandClaude Opus 5 0835da8a91 fix: give the roster's column zip an explicit strict=False (387 D1, B905)
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 6s
CI / frontend-build (push) Successful in 22s
CI / backend-lint-and-test (push) Successful in 32s
Build images / build-web (push) Successful in 58s
Build images / smoke-web (push) Skipped
CI / integration (push) Successful in 2m5s
Build images / build-ml (push) Successful in 2m46s
Build images / promote (push) Skipped
ef91fcf failed ruff's B905 lane on one zip(labels, cells) without strict=. Tests, integration and the frontend were already green on that SHA.

strict=False is the deliberate side, not the quiet one. strict=True raises a bare ValueError - not SubscribeStarDriftError - and would fail the whole roster sync over a column mismatch in `details`, which nothing reads yet. That would take down reconciliation and the gated-post reasons over a cosmetic markup change, while creator identity (id, slug) never depended on the columns at all.

But a shifted column would mislabel details silently (a price filed under "discord"), so a count mismatch now logs a canary warning, mirroring the feed parser's existing parse canary: diagnosable from the worker log, never fatal.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SHQB1YukL3VyvMK8rcbmV9
2026-09-13 10:56:18 -04:00
bvandeusenandClaude Opus 5 ef91fcfd26 feat: SubscribeStar joins the membership roster (387 D1)
CI / lint (push) Failing after 2s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 3s
Build images / build-agent (push) Successful in 6s
CI / frontend-build (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 32s
Build images / build-web (push) Successful in 55s
Build images / smoke-web (push) Skipped
Build images / build-ml (push) Successful in 1m41s
Build images / promote (push) Skipped
CI / integration (push) Successful in 2m12s
The second platform through the seam note 3970 contracted, characterized first from a live capture of the account's /subscriptions page (note 3989). The capture lives in the gitignored captures dir; the committed fixture is hand-built with invented values and was verified tag-for-tag against it - card wrappers, both table heads, and every distinct row shape - before any code depended on it.

What the page is, and the three decisions it forced:

The table IS the status. SubscribeStar has no per-row status word: a creator is either in the active_subscriptions card or the cancelled_subscriptions one. The card's data-identifier is stored verbatim as Membership.status and mapped in MEMBERSHIP_STATUS, keyed on the identifier rather than the table class because the cancelled table's class names the same list differently (for-unsubscribed_users).

The creator's numeric data-user-id is the key, not the slug. A slug re-keys when a creator renames; the old row stops appearing; and a disappearance is exactly what reconciliation reads as a lapse. Keyed on the slug, a rename would have told a paying subscriber they had cancelled. The slug rides as vanity, where the identity join already looks for a handle.

Price is kept as text, never parsed into amount_cents. A bare $ names no currency and a page price is not proven to be the charge - 3970 finding 4. Tier names live behind a per-row modal and are not fetched.

Refusals, because SubscribeStar offers nothing like Patreon's meta.pagination.total and every conclusion downstream is drawn from absence. The parser raises when: the active card is missing (auth error on a login/age wall, drift otherwise); a row lacks a numeric creator id or a creator link; anything renders after a card's table; or the page carries a page= link. Both cards are paginatable (app#embed_pagination) and the captured account was too small to show what pagination looks like, so possible pagination is a roster FC cannot prove complete. A loud error on a larger account beats a quiet half-list. A missing cancelled card is not drift, and a creator in both tables is reported once, as active.

Fetched from subscribestar.adult, not the .art the capture came from: FC's requests never clear the .art age wall with the 18+ cookie (1259, 1284). Whether /subscriptions on .adult authenticates exactly as .art did in the browser is untested - if not, the sweep records a visible error and C6 shows its unavailable rung.

The seam leak D1 found. Note 3970 promised a second platform would be one builders line plus the client method. The sweep instead called current_user_id() on every client, which only Patreon's has, so SubscribeStar would have raised AttributeError on the first sweep. roster_user_id probes it with getattr, the same way the sweep already probes iter_memberships.

Two existing tests were passing for the wrong reason and now can fail:
- "a platform that has never been characterised says nothing" named SubscribeStar, and stayed green only because active_patron is not a SubscribeStar word. Now uses hentaifoundry, with a positive SubscribeStar test beside it.
- the freshness test gave SubscribeStar a Patreon word, so the vocabulary excluded it and deleting the freshness gate outright would have left it green. It now uses cancelled_subscriptions, making the gate the only thing that excludes it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SHQB1YukL3VyvMK8rcbmV9
2026-09-13 10:53:01 -04:00
bvandeusenandClaude Opus 5 529d4bff57 test: tie the install docs to the code they quote (3422 follow-up)
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 4s
Build images / sign-extension (push) Successful in 4s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 7s
Build images / build-web (push) Successful in 5s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / frontend-build (push) Successful in 24s
CI / backend-lint-and-test (push) Successful in 38s
CI / integration (push) Successful in 2m11s
3422 was already fixed. Commit 86abaf0 applied option 1 in full and it is on main: .env.example carries the bootstrap section with the backup warning, README explains the refusal and why it is deliberate, docker-compose forwards the variable, and the milestone-362 smoke gate that FOUND the bug now sets it (build.yml:1159) and passes. The issue's premise - "CURATOR_BOOTSTRAP_NEW_KEY appears nowhere outside backend/" - is stale.

What was left is the dependency that fix created. README.md and .env.example now both print the literal error text, the literal key path and the variable name, because a stranger greps for the string their terminal showed them. That is the right call and it means two user-facing files now depend on this module's wording with nothing connecting them - the install surface's characteristic defect, one rename away from a README that sends strangers to a path that does not exist.

Three guards, all presence checks on both sides. An absence check against prose would pass for the wrong reason the moment a sentence were reworded (snippet 3352):

  - the raised message still contains the sentence README reproduces, the variable both docs say to set, and the restore-rather-than-mint alternative the whole refusal rests on;
  - both docs still name _CREDENTIAL_KEY_PATH and the variable, read from the code rather than retyped, so a rename fails here;
  - compose still forwards the variable - without that line the docs' "set it in .env" is silently inert and fails identically to not setting it.

Option 2 (mint when the credential table is empty) is deliberately NOT done. The issue's own guidance is "(1) now, (2) if the friction proves annoying", and the friction has not been reported. Worth recording that its predicate checks out exactly: the Fernet key protects Credential.encrypted_blob and nothing else - no other Fernet user exists - so "no credential rows means nothing can be made undecryptable" is provable rather than probable. The cost is placement: create_app() is sync and constructs the key before any engine exists, so the check cannot live where the failure is. entrypoint.sh, which already runs alembic against the DB, is the natural seam. Only the web role is affected; the Celery roles build the key lazily inside tasks.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SHQB1YukL3VyvMK8rcbmV9
2026-09-12 21:31:21 -04:00
bvandeusenandClaude Opus 5 eb6e0df858 feat: the empty front door offers to find what you already subscribe to (387 C6)
CI / lint (push) Successful in 2s
Build images / sign-extension (push) Successful in 5s
CI / extension-version (push) Successful in 2s
Build images / build-agent (push) Successful in 8s
Build images / build-ml (push) Successful in 8s
CI / frontend-build (push) Successful in 28s
CI / backend-lint-and-test (push) Successful in 39s
Build images / build-web (push) Successful in 1m16s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / integration (push) Successful in 2m45s
B4's on-ramp read "add a credential, add a source" - the manual path, and the one that makes a new installer retype creators they have already told Patreon about. With C4 in place the app can just look them up. This is the step where the milestone's original framing actually lands on screen.

Four rungs, chosen by ONE predicate rather than four independent v-ifs. The rungs are mutually exclusive by construction and the rung shown matches what is actually POSSIBLE:

  no credential            -> Add a credential. Discovery is not offered, because a button that cannot work is worse than its absence.
  credential, never synced -> Find what you already subscribe to. Queues the C3 sweep.
  synced, unmatched > 0    -> the count, linking into C4's bucket 1.
  synced, nothing unmatched-> the manual add, because there is genuinely nothing to discover.

The fifth state is rule 164's. A sweep that has been ATTEMPTED and never succeeded reads as "couldn't reach patreon", with the error type and the manual path still open - not a spinner, not a crash, not a retry button that will fail identically. An install with no outbound network lands here.

That is deliberately narrower than "there is an error": a roster that synced once and failed since is NOT unavailable. It has a roster, just an ageing one, and C4's freshness gate already handles that. Collapsing the two would hide a usable roster behind an error banner.

Still offer, never auto-add. An empty front door is exactly where "just add all thirty" is most tempting and most wrong - thirty backfills on first boot - so the rung carries the count and sends them to the picker.

No new machinery: three existing stores (credentials, membershipSync, membershipReconcile). Reconcile is fetched only once something has synced, since bucket 1's count is meaningless before that. Every load swallows its failure, because this screen renders on an install that can reach nothing.

Ten tests on top of B4's, including the mutual-exclusion property asserted directly. Its phrases are each pinned to a single template line and both sides whitespace-normalised - a phrase spanning a line break would never match, and a mutual-exclusion check whose phrases never match passes vacuously (rule 167).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SHQB1YukL3VyvMK8rcbmV9
2026-09-12 20:19:42 -04:00
bvandeusenandClaude Opus 5 4533e036ac refactor: the membership seam's contract type is the seam's, not Patreon's (387 C7)
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 5s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 29s
CI / backend-lint-and-test (push) Successful in 33s
Build images / build-web (push) Successful in 1m16s
Build images / smoke-web (push) Skipped
Build images / build-ml (push) Successful in 2m36s
Build images / promote (push) Skipped
CI / integration (push) Successful in 3m11s
Membership moves from patreon_client to native_ingest_common, beside PostRecordOutcome, for exactly the reason that one lives there: it is the seam's contract rather than the first platform's. Left where it was, D1 would have had to import the shape it implements from the module of the platform it is being mirrored FROM - which inverts the dependency and is how a seam advertised as portable quietly stays Patreon-shaped.

Found by C7's own pass, which is the point of running C7 before D1 rather than writing it up afterwards: this is invisible while there is only one implementer and load-bearing the moment there are two.

No behaviour change. Three files, no shim (rule 122): patreon_client imports it, the dataclass keeps its docstring, and the test imports from the seam's home. The docstring gains what the contract owes a second platform - that a missing field supplies the empty answer and never a guess: no tiers -> [], no pledge -> None (absent stays distinguishable from zero, since "free" and "we don't know" are different answers), no vanity -> None with identity falling back to the URL tail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SHQB1YukL3VyvMK8rcbmV9
2026-09-12 20:05:11 -04:00
bvandeusenandClaude Opus 5 6f5ea5d1d3 fix: the provenance panel hotlinked the CDN for images already on disk (3965)
CI / lint (push) Successful in 8s
CI / extension-version (push) Successful in 9s
Build images / sign-extension (push) Successful in 9s
Build images / build-agent (push) Successful in 17s
CI / frontend-build (push) Successful in 35s
CI / backend-lint-and-test (push) Successful in 1m18s
Build images / build-web (push) Successful in 1m22s
Build images / smoke-web (push) Skipped
Build images / build-ml (push) Successful in 2m19s
Build images / promote (push) Skipped
CI / integration (push) Successful in 2m38s
Two surfaces render a post's HTML body. get_post did _localize_inline_images(sanitize_post_html(...)); provenance_service._post_dict called sanitize_post_html alone. So opening the Provenance panel fetched images from Patreon's CDN for files FC had already downloaded - the archive reaching out to the platform to display what it had archived, which is the thing 830 Phase 2 set out to stop. The same bodies also break when a CDN URL expires or a post is removed, while the identical local copy sits unused.

The fix is the shape, not the call. Those were two separately-callable halves and only the first looked mandatory, so a second caller was always going to do half of it. render_post_body in the new services/post_body.py is the whole pipeline in one call, and sanitizing-without-localizing is no longer a reachable operation. _localize_inline_images moves there verbatim; post_feed_service loses five imports that went with it.

_post_dict becomes async and takes the session. Both call sites are already inside async methods, so for_image's list comprehension awaits per entry - fine, because localization issues ZERO queries for a body with no inline <img>, which is most of them. Recorded that early exit in the module so the loop isn't "optimized" into a batch without a measurement.

Deliberately NOT fixed: provenance still names the same columns url/title/date where the feed says post_url/post_title/post_date, and its description_translated is full text where the feed truncates to DESCRIPTION_LIMIT. That is 3965's wider half - a breaking payload change for ProvenancePanel with no second reason to spend it today. Noted in _post_dict's docstring so the next reader knows it was seen and left.

Four regression tests on the provenance path, covering both entry points plus the two refusals the feed already pins: an uncaptured image stays hotlinked (a broken local path is worse than an intact remote one), and a filehash owned by another artist never leaks in. Reverting render_post_body to a bare sanitize fails the first two.

Recorded as snippet 3968.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SHQB1YukL3VyvMK8rcbmV9
2026-09-12 19:57:00 -04:00
bvandeusenandClaude Opus 5 2862fadcb1 fix: the roster guard's import walk resolved package __init__ imports wrongly (387 C5)
CI / lint (push) Successful in 2s
Build images / sign-extension (push) Successful in 3s
CI / extension-version (push) Successful in 1s
Build images / build-agent (push) Successful in 6s
Build images / build-ml (push) Successful in 6s
CI / frontend-build (push) Successful in 30s
CI / backend-lint-and-test (push) Successful in 32s
Build images / build-web (push) Successful in 7s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / integration (push) Successful in 2m37s
Two of C5's three structural tests errored in CI with ValueError: PosixPath('.') has an empty name. The walk special-cased a package's __init__.py, dropping the __init__ component before computing what `from .` refers to — which made `api/__init__.py`'s `from . import health` resolve to the app root instead of to `api`, and `celery_app.py`'s `from . import celery_signals` resolve to the empty string, which is what actually crashed.

The special case was never needed: `parts[:-1]` already gives the CONTAINING package for both forms, because `services/foo.py` drops `foo` to leave `services` and `api/__init__.py` drops `__init__` to leave `api` — exactly what `from .` means inside each. The level slice is now clamped at 0 as well; an import climbing past backend/app left the tree, and the unclamped negative index wrapped and resolved to the wrong module rather than to nothing.

A module directly under backend/app doing `from . import x` still yields no package prefix, and there the alias alone IS the dotted name - handled explicitly rather than by falling through into a Path built from an empty string.

The two positive controls earned their place immediately: they are what failed. Without them the walk would have resolved almost nothing and test_no_fetch_path_can_read_the_roster would have passed on a broken walker, reading as coverage.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SHQB1YukL3VyvMK8rcbmV9
2026-09-11 23:02:18 -04:00
bvandeusenandClaude Opus 5 aa765f0a72 feat: say why the posts are invisible, without ever deciding they are (387 C5)
Build images / promote (push) Skipped
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 21s
CI / backend-lint-and-test (push) Failing after 32s
Build images / build-web (push) Successful in 58s
Build images / smoke-web (push) Skipped
Build images / build-ml (push) Successful in 1m47s
CI / integration (push) Successful in 2m16s
A3 made a tier-gated source say "47 posts you can't see". The roster turns that into a reason: the membership ended, or the tier doesn't reach these posts, or it's a free follow. Rendered under A3's count in the health tooltip, quieter than the count it explains.

FREE is a fourth case the step didn't enumerate, and it earns its own sentence. has_paid_access collapses "former patron" and "current free follower" to the same False, so deriving the reason from that boolean would tell a free follower "you're not a patron any more" - a false statement about a state they were never in. gated_reason reads the status axis first, calling has_paid_access with is_free_member forced off, then splits on the free flag.

Silence is the default, and there are four ways into it: campaign absent from the roster, roster stale, platform never swept, status word not yet characterised. All four send null and the count stands alone. The frontend has no fallback sentence either - a default would turn "we don't know why" into a reason, which is the one thing this step must not do.

The line that must not be crossed is pinned structurally rather than by inspection: test_no_fetch_path_can_read_the_roster walks the transitive first-party imports from the fetch roots and asserts the roster is unreachable. FC runs no local verification (rule 85), so a guard cannot be falsified by hand before it lands - it carries two positive controls instead, proving the walker finds roster imports that ARE there, one direct and one through a hop, so the real assertion can never pass merely because the walk resolved nothing.

C4's identity loop moved to membership_roster.pair_sources_with_memberships when C5 became its second caller; two copies would let the Subscriptions row and the reconciliation card disagree about which creator a source IS. Three test files were each building PlatformMembership rows with their own drifting helper - consolidated into tests/roster_builders.py, same family as issue 3109.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SHQB1YukL3VyvMK8rcbmV9
2026-09-11 22:59:29 -04:00
bvandeusenandClaude Opus 5 de11c14448 refactor: one relative-time formatter, not three
CI / lint (push) Successful in 2s
CI / extension-version (push) Successful in 2s
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 6s
Build images / build-ml (push) Successful in 6s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 31s
Build images / build-web (push) Successful in 1m26s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / integration (push) Successful in 2m11s
MembershipRosterCard had grown its own ago(iso) helper, a near-copy of utils/date.js::formatRelative. C4's card was about to become a third copy before a hook caught it. Both now use the shared helper.

The hand-rolled copy was also slightly wrong in ways the shared one is not: it floored everything under a minute to '1m ago', and would have rendered NaNm ago for a null timestamp had a caller ever reached it without a v-if guard. Only sub-minute output changes, which no spec exercises - membershipRosterCard.spec.js seeds its rows at exactly 1h and 9d, where both helpers agree, and asserts on literal phrases rather than time strings.

Recorded formatRelative as snippet 3959 so the next component is offered it instead of deriving a fourth copy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-11 22:48:24 -04:00
bvandeusenandClaude Opus 5 005680f234 fix: one source edit no longer wipes the campaign id and the backfill position
Build images / sign-extension (push) Successful in 2s
Build images / build-agent (push) Successful in 6s
CI / lint (push) Successful in 1s
CI / extension-version (push) Successful in 2s
CI / backend-lint-and-test (push) Successful in 39s
CI / frontend-build (push) Successful in 28s
Build images / build-web (push) Successful in 1m10s
Build images / smoke-web (push) Skipped
Build images / build-ml (push) Successful in 3m4s
Build images / promote (push) Skipped
CI / integration (push) Successful in 2m22s
Source.config_overrides carries two unrelated things under one column: the operator's per-source download settings, and state FC writes for itself. update() treated the whole column as operator-owned and assigned it wholesale, so a dialog save discarded patreon_campaign_id and the entire #693 backfill state machine. Not a hand-edited-JSON edge case: SourceFormDialog's structured tab rebuilds the object from two fields, so saving without touching anything was enough.

_merged_config now merges: the operator's keys replace wholesale (removing a key must still remove it), FC's keys survive and are applied LAST so a stale echoed cursor cannot roll a walk backwards. App-managed is _-prefixed or *_campaign_id, both matching data already on disk, so no migration.

Preserving the id exposed a bug the wipe was MASKING: nothing cleared it when a source's URL changed, and patreon_resolver reads that cache before attempting any lookup. A repointed source would have resolved the old creator forever and 387 C4 would have reported a confident wrong match. So update() drops *_campaign_id when the URL actually changes - but keeps the backfill cursor, which the walk's own stall guard validates and which is expensive to rebuild.

test_update_changes_fields was DOCUMENTING the bug: it asserted config_overrides == {videos: False} on a source whose create() had armed _backfill_state, so it could only pass because the state had been destroyed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-11 21:50:28 -04:00
bvandeusenandClaude Opus 5 fc136006b7 feat: what you pay for, against what FC actually follows (387 C4)
CI / lint (push) Successful in 2s
Build images / sign-extension (push) Successful in 3s
CI / extension-version (push) Successful in 2s
Build images / build-agent (push) Successful in 5s
CI / frontend-build (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 32s
Build images / build-web (push) Successful in 1m18s
Build images / smoke-web (push) Skipped
Build images / build-ml (push) Successful in 2m11s
Build images / promote (push) Skipped
CI / integration (push) Successful in 3m4s
Reconciliation in Subscriptions, asymmetric on purpose. Subscriptions FC does not follow get a per-row add; sources the roster cannot account for are REPORT ONLY (operator decision) and link to the list on the same page. No one-click disable, so no disabled-reason column and no migration.

The membership<->source join lands as a SHARED resolver in membership_roster, not inline here: E4 now uses it as its negative check, so the two features cannot give different answers to 'is this membership already tracked?'. Keys on the exact cached campaign id FIRST and the URL handle only as fallback, because the id is written only after a source has been walked once. Read via any <platform>_campaign_id override rather than naming Patreon's, per rule 169.

The report-only bucket is gated on roster freshness and carries a per-row basis, so 'your membership says former patron', 'we know this id and it is absent', and 'we only have a handle' stay three different sentences. has_paid_access None never reads as lapsed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-11 21:38:03 -04:00
bvandeusenandClaude Opus 5 f8614d437d fix: a self-contradicting test fixture, and import order (E4 follow-up)
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
CI / extension-version (push) Successful in 2s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 34s
Build images / build-web (push) Successful in 1m13s
Build images / smoke-web (push) Skipped
Build images / build-ml (push) Successful in 2m15s
Build images / promote (push) Skipped
CI / integration (push) Successful in 2m46s
Two failures on 51e78a3, neither in the shipped logic.

**The integration failure was a TEST bug, not a code bug.**
`test_a_weak_name_needs_the_declaration` set `display_name` to something
deliberately weak but left `_membership`'s DEFAULT vanity, which matched the
artist slug exactly. So the name signal was legitimately 1.0 and the matcher
was right to propose — my assertion of 0 was asserting the wrong scenario.
Both identity fields now have to be weak for the test to mean what it says,
and the arithmetic was checked before pushing: 0.39 without the declaration,
0.74 with it.

Worth keeping: a fixture whose fields disagree with each other will pass or
fail for reasons unrelated to the property under test, and this one was one
default away from silently testing nothing.

**The lint failure was import order** — `artist_membership_service` sorts
before `credential_*` and I inserted it after. Second isort slip this session
from patching an import block with a script rather than reading it back; the
repo has a rule about exactly this (#102).

Everything else passed on that SHA: 1256 tests, the frontend suite, and
migration 0096.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-11 08:02:08 -04:00
bvandeusenandClaude Opus 5 51e78a329b feat: offer the creator you already track as the one you subscribe to (388 E4)
CI / lint (push) Failing after 3s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 8s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 31s
Build images / build-ml (push) Successful in 2m23s
Build images / build-web (push) Successful in 1m25s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / integration (push) Failing after 2m31s
**The verification the step asked for came back "not the schema".**
`Source.artist_id` is a plain FK so many sources per artist already works;
`POST /api/sources` already takes an `artist_id`; the add-source dialog already
has an artist autocomplete that attaches to an EXISTING artist; and
`SourceService.reassign` already moves a source between artists WITH post and
image re-attribution. A sweep for one-source-per-artist assumptions found only
`func.count()` calls — the opposite of assuming one.

So no parallel association table was built for a relationship the schema
already expresses (rule 28). What was missing is FC OFFERING the link, and that
is all this adds.

**Accepting adds a SOURCE. It never merges two artists.** That asymmetry sets
the whole posture: adding a source is trivially undone, while a wrong merge
silently mixes two creators' work and corrupts tagging, series and provenance
downstream with nothing left to tell them apart by. A test asserts the artist
count is unchanged by accepting.

The weights encode the judgement rather than a code path doing it — name 0.65,
declared 0.35, cut at 0.60 — so that:

* an EXACT name match alone proposes (same slug on both sides is strong, and
  demanding corroboration would propose almost nothing);
* a CONTAINMENT match alone does not ("art" sits inside "artgirl"), and short
  slugs are excluded from containment entirely because a 3-character slug is
  inside a great many longer ones;
* the declaration ALONE never proposes, because a creator may link another
  creator's Patreon and a link is not a claim of identity.

A guard test pins all three against WEIGHTS directly and says not to fix a
failure by moving the numbers.

Two corrections carried forward from earlier steps rather than rediscovered:

* The declaration is NOT read from `ExternalLink`. `SUPPORTED_HOSTS` is file
  hosts only and `host_for()` returns None for patreon.com, so no row is ever
  written for one — the same trap that caught E5 for Discord invites. It reads
  the raw body, because these links live in an `href` and `html_to_plain`
  discards attributes.
* `vanity` is not a column: C1 modelled the roster before any platform was
  characterised, which is exactly what `details` exists for. `vanity_or_none()`
  reads it from there and falls back to the URL's last segment, so a row
  written before the field was understood still resolves.

Two fixes during the writing. `accept()` first created a bare `Source()`,
skipping the platform/URL validation, duplicate check and #693 backfill-arming
that a hand-added source gets — a second, quieter way to create a source is how
two paths drift until one is subtly broken; it now goes through
`SourceService.create`. And the candidate query used a bare `exists().where()`,
which has no FROM to correlate against; now `select(...).exists()`.

Chained onto the roster sweep rather than given its own beat entry: a
suggestion can only be as good as the roster behind it, so any other cadence
would just propose from staler data.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-11 07:56:16 -04:00
bvandeusenandClaude Opus 5 61abd0007c fix: stdlib imports split by a stray blank line (isort I001)
CI / extension-version (push) Successful in 3s
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 5s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 30s
CI / backend-lint-and-test (push) Successful in 35s
Build images / build-web (push) Successful in 1m34s
Build images / smoke-web (push) Skipped
CI / integration (push) Successful in 2m42s
Build images / build-ml (push) Successful in 3m1s
Build images / promote (push) Skipped
My scripted patch inserted the new stdlib from-imports after `import logging`
with a blank line between them, which isort reads as a group boundary.

Ruff-only failure — every test passed on 751e7dd, including the full
integration suite and migration 0095.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-10 23:46:02 -04:00
bvandeusenandClaude Opus 5 751e7ddb9f feat: the membership sweep, and the state that makes its failures readable (387 C3)
CI / lint (push) Failing after 3s
Build images / sign-extension (push) Successful in 4s
CI / extension-version (push) Successful in 2s
Build images / build-agent (push) Successful in 8s
CI / frontend-build (push) Successful in 28s
CI / backend-lint-and-test (push) Successful in 38s
Build images / build-web (push) Successful in 1m28s
Build images / smoke-web (push) Skipped
Build images / build-ml (push) Successful in 2m35s
Build images / promote (push) Skipped
CI / integration (push) Successful in 2m55s
A daily sweep that walks each platform's roster into `platform_membership`.
Daily because memberships change on a BILLING cycle, not a download cadence.

**`membership_sync` is the part that earns its keep.** Without it three very
different situations are one indistinguishable state — the account subscribes
to nothing, the sweep never ran, the sweep failed — and all three leave zero
rows in `platform_membership`. "You are tracking 12 sources you no longer
subscribe to" is correct in the first case and an invitation to cancel things
the operator is actively paying for in the other two. So C4 gates its
CONCLUSIONS on `last_success_at`, not merely its display, and `roster_is_fresh`
is computed server-side so no caller can forget to.

Two timestamps rather than one: `last_attempt_at` moves every run,
`last_success_at` only on a clean walk. The gap between them is the signal —
a sweep hammering a broken credential every day must not look healthy because
it ran recently, and there is a test for exactly that.

Rejected shortcuts, both tempting: `MAX(platform_membership.last_seen_at)`
cannot tell "synced fine, found nothing" from "never synced"; `task_run` is
worse, since its retention prunes ok rows after 24h and a sweep that last
succeeded three days ago would leave no trace at all.

**The fetch completes before anything is written.** That ordering is the safety
property: a walk that dies mid-pagination writes nothing, so a failure can
never leave a roster half this week's and half last week's. `touch_membership`
never deletes, so a failure cannot empty the roster either — but "intact"
should mean intact, not merely non-empty.

Rule 89's four, each where it actually lives: recovery is "run it again"
(upsert, no deletes); retention is C1's age-out-never-delete, because
disappearing IS the signal; the wall-clock deadline is per-platform and
distinct from the per-REQUEST timeout the client already has (rule 156 — a
paginated roster answering every page slowly-but-within-timeout would never
trip that one and would sit on a worker indefinitely); duration comes from the
existing TaskRun signal plumbing.

**A bug caught in review, not production:** the broad `except Exception` would
have swallowed Celery's SoftTimeLimitExceeded — which is an ORDINARY Exception
subclass, not a BaseException — letting the sweep run past the soft limit into
the hard one, where it is SIGKILLed mid-transaction. A sweep that cannot be
stopped is worse than one that fails. Now re-raised explicitly, with a test
that also asserts SoftTimeLimitExceeded is still an Exception, so the re-raise
cannot quietly become dead code.

Rule 164 is why this ships with UI rather than backend-only: a roster that
never synced must be VISIBLE as such. The card says "never synced" in words and
states no count at all — rendering it as 0 is the precise conflation the whole
step exists to prevent — while a real zero behind a real sync is reported as
zero, because that one IS an answer. Pinned in both directions.

Three independent gates decide whether a platform is swept — registered here,
client exposes `iter_memberships`, credential exists — each silent, so adding
SubscribeStar (D1) is one line and nothing else. A missing credential is not an
error: recording a failure would light up the UI for a feature never enabled.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-10 23:40:48 -04:00
bvandeusenandClaude Opus 5 afcde8e457 feat: PatreonClient.iter_memberships — the roster seam (milestone 387 step C2)
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 5s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 24s
CI / backend-lint-and-test (push) Successful in 58s
Build images / build-web (push) Successful in 1m16s
Build images / smoke-web (push) Skipped
CI / integration (push) Successful in 2m25s
Build images / build-ml (push) Successful in 2m45s
Build images / promote (push) Skipped
Built on C0's real capture (Scribe note #3886), not on API docs — gallery-dl
has no membership extractor and Patreon's public v2 API is the CREATOR surface
behind OAuth, so the rule-130 reference had to be a characterized response.

**The request is deliberately minimal, and that is a privacy decision.** The
browser's own include set pulls `latest_pledge.card`, and those card resources
come back carrying the ACCOUNT HOLDER'S EMAIL in `merchant_name`; `address` is
in there too. Copying the query string wholesale is the obvious move and would
have FC fetching payment PII it has no use for and can only mishandle. We ask
for `include=campaign,reward` and nothing else, and a test asserts on the
params actually sent so nobody widens it back.

**We do not send `filter[membership_type]`.** The browser sends the six buckets
its settings page displays, which excludes lapsed memberships — and a
DISAPPEARANCE is precisely the signal the roster exists to read. Filtering here
would manufacture the event C4 acts on.

Two corrections the capture forced, both now in code:

* **The filter vocabulary is not the status vocabulary.** I had read the six
  filter words off a screenshot and was about to write them into
  MEMBERSHIP_STATUS as the enum. The body shows `patron_status` carrying
  `former_patron` — absent from that filter — on a row the filter selected as
  `free_member`. So the map is taught exactly the two OBSERVED values, and
  `declined_patron` stays out despite looking obviously right: believing the
  filter is the mistake that was just caught.
* **Free membership is a boolean, not a status.** `has_paid_access` gains an
  `is_free_member` axis, because `active_patron` alone would report a free
  follower as a paying patron and C4 would never offer to clean it up. Honest
  limit, stated in the docstring: the capture has no ACTIVE free member, so it
  shows the separation is possible, not that it occurs.

C1's tripwire test did its job — it was written to fail the moment anyone
populated the status map, and updating it here IS the confirmation step, done
with the capture rather than ahead of it.

`_fetch`'s retry/backoff/auth-vs-drift/Retry-After logic is extracted to a
shared `_request` so the roster rides the same path rather than growing a
second copy — two copies would drift, and the half that drifted would be the
one that only runs daily. Every error message and log line renders
byte-identically for the posts path, so the existing tests pin the refactor.

Pagination is driven by `page[offset]` against `meta.pagination.total`, never
by `links`: the response's own `links.first` is built WITHOUT the `/api/`
prefix the request uses, so following it would hit the web page. An empty page
is terminal regardless of what the total claims, so a server reporting more
rows than it hands over cannot spin the walk forever.

Drift is stricter here than on the posts path, on purpose: a missing
`meta.pagination.total` raises rather than returning a short list, because a
truncated roster reads downstream as "you cancelled those" — the worst wrong
answer this feature can give.

`current_user_id()` is marked INFERRED, not characterized: C0 captured
/api/members, not /api/current_user, so it relies only on the JSON:API envelope
this API demonstrably uses elsewhere, and raises drift rather than returning
something plausible if that is wrong.

The fixture is derived from the real capture with every piece of account data
replaced (the raw capture stays gitignored). Six members, each earning its
place: a former patron with a null pledge, an active patron with no tier, an
annual cadence, a previous_pledge whose included resource has no
`relationships` key at all, and a reward priced in CAD beside a USD charge —
the trap that makes reading `reward.amount_cents` report a number the operator
was never charged. A leak check caught a free-membership-subscription id and
six real campaign launch timestamps before any of it was staged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-10 22:26:01 -04:00
bvandeusenandClaude Opus 5 533a1ce674 chore: a gitignored home for raw platform captures (milestone 387 C0)
CI / extension-version (push) Successful in 3s
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 7s
Build images / build-ml (push) Successful in 8s
Build images / build-web (push) Successful in 6s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / frontend-build (push) Successful in 24s
CI / backend-lint-and-test (push) Successful in 32s
CI / integration (push) Successful in 2m12s
C0 characterized Patreon's `/api/members` from a live capture of the
operator's own session. That capture is worth keeping — re-capturing means
re-authenticating by hand, and it is the ground truth the characterization
(Scribe note #3886) gets re-checked against when a platform's shape is
suspected to have drifted.

It cannot be committed. It carries the operator's creator list, pledge
amounts, and — inside the `card` resources the web app's include set pulls —
the account's own email address. So: a directory that is ignored wholesale
rather than by filename, so the next capture is covered by this rule instead
of needing a line somebody has to remember to add.

Two details that are the point rather than incidental:

* The ignore is written as `captures/*` plus a negation for README.md, NOT as
  `captures/`. Git does not descend into an excluded DIRECTORY, so a negation
  for a file inside one never takes effect — the README would have been
  silently ignored along with everything else, and the convention would not
  have survived a fresh clone.
* The README states plainly that SANITIZED fixtures belong in git, elsewhere
  under tests/fixtures/. The raw capture exists to derive those from and to
  re-check against; it is not the thing tests should load.

The ignore rule landed before the capture file did, deliberately: a payload
with an email address in it should never be sitting in the working tree
un-ignored, however briefly.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-10 21:07:48 -04:00
bvandeusenandClaude Opus 5 059e2128ee fix(ci): stop telling readers the runner can't run upload-artifact@v4+
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 7s
Build images / build-ml (push) Successful in 7s
Build images / build-web (push) Successful in 7s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / frontend-build (push) Successful in 24s
extension / lint (push) Successful in 25s
CI / backend-lint-and-test (push) Successful in 32s
CI / integration (push) Successful in 2m17s
build.yml and baseline.yml both explained a missing artifact step by
saying act_runner cannot run actions/upload-artifact@v4+ (and baseline.yml
said ci-requirements.md records that, which it does not). The runner is
now gitea/runner 3.x, which edits the action's GHES refusal out of its
bundle, and stock v4+ is proven working on this forge (Scribe spike #3843).

Comments only. Neither lane gains an upload step: build-web still reads
the signed XPI from the release asset, and the baseline candidate is
still printed to the log, because those remain the better channels. The
comments now say why for the right reason.

Scribe snippet #2271, milestone 395.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DwoKYuw3qJmUUYsJeNherB
2026-09-10 17:43:36 -04:00
bvandeusenandClaude Opus 5 d0b0458d27 feat: the announcement link, on the card and in a review queue (388 E5)
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 5s
Build images / build-agent (push) Successful in 8s
Build images / build-ml (push) Successful in 8s
CI / frontend-build (push) Successful in 22s
CI / backend-lint-and-test (push) Successful in 31s
Build images / build-web (push) Successful in 1m1s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / integration (push) Successful in 2m11s
Rule 27 — E5's other half. The matcher can propose; this is where the operator
decides, and where an accepted link actually shows up.

**The review queue** (Settings → Ingestion & filters). Each proposal shows the
per-signal breakdown, not just the total: "why did it suggest this" is the
question the operator actually has, and a lone percentage cannot answer it. So
a row reads "72% · timing 95% · says so 60%", and the copy states outright that
a pair always needs two reasons — which is the property that stops a busy
posting day from producing false pairs.

The empty state says so explicitly. Nothing proposed is the EXPECTED state most
of the time, and an empty queue that looks like a failure invites turning the
threshold down until it produces noise.

**On the card**, both directions, and accepted links only: the teaser gets "The
full set is in Discord", the drop gets "Announced on Patreon". A pending
proposal is a question for the review queue, never a claim to render beside the
artwork — that distinction is the whole confirm-only design, so it is asserted
in the backend (only `linked` rows reach the payload) and again here.

One detail worth the comment it carries: the link's target is `{ query: {
post_id } }` with no name or path. In vue-router that means "the current route
with these query params", so it works identically from Latest and from Browse —
and, more usefully, the card never reaches for `useRoute()`, which it has no
other reason to know about and which is not available when it is mounted in a
test without a router.

Backend CI on 235393c was green: all 13 E5 tests and migration 0094.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-10 11:53:17 -04:00
bvandeusenandClaude Opus 5 235393c08b feat: link the Patreon teaser to the Discord drop it announced (388 step E5)
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 25s
CI / backend-lint-and-test (push) Successful in 32s
Build images / build-web (push) Successful in 1m3s
Build images / smoke-web (push) Skipped
Build images / build-ml (push) Successful in 1m53s
Build images / promote (push) Skipped
CI / integration (push) Successful in 2m12s
The point of the milestone rather than its tail. Two of the operator's artists
post a deliberately cropped fragment on Patreon to signal that the real thing
has landed in their Discord; this proposes those pairs.

Confirm-only, following the FC-6.3 series matcher. A wrongly-asserted
association tells the operator two different pieces are one, which is strictly
worse than no link: no link leaves them where they already were, a wrong one
actively misinforms and then propagates into whatever reads it. So the
matcher's job is a SHORT list worth reading, not a long list worth trusting.

**The threshold sits above every single signal weight, and that is the
design.** Proximity is 0.55, declaration 0.45, the cut 0.60 — so neither
signal can carry a pair alone. That makes "time proximity alone is never
sufficient" an arithmetic property rather than an aspiration: on a busy day an
artist posts several times, and a matcher that could pair on proximity alone
would turn every one of those days into false pairs until the review queue got
abandoned. A guard test asserts the relationship against WEIGHTS directly, so
it survives any refactor of the scorer, and says in its own failure message not
to fix it by lowering the assertion.

**Crop-to-source matching is HELD, on the plan's instruction** — real work with
real false-positive risk, worth building only once signals 1 and 2 are shown
insufficient against the operator's actual artists. Worth stating: a naive
whole-image SigLIP similarity is NOT that signal. A cropped teaser and its full
version are precisely the pair a whole-image comparison handles worst, so
adding one as a "bonus" would mostly add noise while looking like progress.

Two premises in the plan corrected in the building:

* **E4 is not actually a prerequisite.** A Patreon Source and a Discord Source
  the operator has added under one Artist already share `Post.artist_id`, and
  the synthetic grouping inherits it. E4 EXTENDS this to creators FC has to
  learn the association for; it is not needed to represent one FC was told.
  Same-artist is then a hard filter, not a scored signal — two different
  creators posting minutes apart is a coincidence, not evidence.
* **`link_extract` cannot supply the declaration signal.** It exists, but
  `SUPPORTED_HOSTS` is file hosts only and `host_for()` returns None for a
  Discord URL, so no ExternalLink row is ever written for one. The signal
  reads the post body directly instead.

And a bug my own test would have caught: `declared_signal` stripped the HTML
before looking for an invite, but `html_to_plain` discards attributes and
these creators put the invite in an anchor's `href` — so the strongest form of
the signal was being thrown away, leaving only whatever the link text said.
The invite now matches the raw body; the bare mention still matches stripped
text, so `\bdiscord\b` is tested against prose rather than against markup.

Dismissed rows are kept, not deleted: the row is what remembers the rejection,
and re-proposing a rejected pair on every scan is the one behaviour that makes
a review queue get ignored. Both FKs CASCADE, so E3's one-DELETE reversal
cannot leave a proposal pointing at a post that no longer exists.

Only ACCEPTED links reach the post payload. A pending proposal is a question
for the review queue, not a claim to render beside the artwork.

UI (rule 27) follows in the next commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-10 11:47:28 -04:00
bvandeusenandClaude Opus 5 ba96ecfb2d fix: the disabled sweep's shape assertion pinned the pre-E3 payload
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 5s
Build images / build-ml (push) Successful in 9s
Build images / build-agent (push) Successful in 9s
Build images / build-web (push) Successful in 6s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / frontend-build (push) Successful in 20s
CI / backend-lint-and-test (push) Successful in 32s
CI / integration (push) Successful in 2m3s
My own E2 test asserted `sweep`'s disabled return by exact equality, and E3
added `images_joined` to it — a rule 90 miss on a consumer I wrote an hour
earlier. Every E3 test passed; this was the only failure (1 failed, 1202
passed).

Fixed by extending the assertion, NOT by loosening it to a subset check. The
exactness is the point: a disabled sweep reports a complete zeroed shape
rather than a shorter one, so a caller can read any counter unconditionally,
and this assertion is what notices when a new counter skips that path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-10 11:35:53 -04:00
bvandeusenandClaude Opus 5 1e45e2c56c feat: an open grouping — a later drop joins its post (milestone 388 step E3)
CI / extension-version (push) Successful in 3s
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 22s
CI / backend-lint-and-test (push) Successful in 31s
Build images / build-web (push) Successful in 1m6s
Build images / smoke-web (push) Skipped
Build images / build-ml (push) Successful in 1m59s
Build images / promote (push) Skipped
CI / integration (push) Failing after 2m7s
A synthetic post is no longer sealed at creation. A creator who adds two more
variants the next day extends the existing post, its body grows with the new
messages, and no rival post appears. That is what makes chat capture read as
content trickling in rather than as a stream of separate arrivals.

The sweep now runs two passes per source and the ORDER is load-bearing: offer
new messages to still-open groups BEFORE founding new ones, because whichever
runs first claims a message.

E3's three named problems, each answered rather than discovered later:

**Bridging.** A candidate near two groups joins NEITHER. Nearest-wins would
silently make an arbitrary choice between two posts the operator may already
have seen; merging them is worse still, because a merge rewrites history and
anything pointing at the absorbed post dangles. Leaving it to found its own
group is the recoverable failure. AMBIGUITY_MARGIN is a module constant and
deliberately not a setting — it is not a quality dial anyone would tune toward
a better feed, and exposing it would invite turning it to zero, which is
exactly the silent arbitrary choice it prevents.

**Re-surfacing without thrashing.** A grouping has two dates, and which one
orders the feed is a real decision, so the feed orders by neither directly.
Ordering by when the drop STARTED buries a group that grows a week later under
a week of other posts — defeating the point of keeping it open. Ordering by
every growth lets a group gaining one image a day live permanently at the top,
so chat out-competes authored posts for the front page — the opposite of "post
pacing stays front and centre". Instead `resurfaced_at` moves only when growth
clears BOTH a minimum-images bar and a cooldown, so a drip-feed updates in
place and a genuine second wave resurfaces exactly once. It is NULL on every
ordinary post, so the sort key COALESCEs through it without moving anything
that is not a grouping.

**Reopening forever.** Groups close after a quiet period — artists reuse
characters for years, and a group left open indefinitely will eventually
absorb something it shouldn't. Openness is DERIVED, not stored: a group is
open if it grew (or started) within the window. Lowering the setting closes
old groups and raising it reopens them, with nothing to repair either way; a
stored closed_at would have needed a sweep to set it and a repair path to ever
change the policy.

Rule 89 is satisfied structurally rather than by a parallel mechanism:
celery_signals writes a TaskRun for every task, which already supplies
duration, the 5-minute stalled-run recovery, and retention pruning. What this
step owed on top of that was a wall-clock limit (present) and idempotence —
re-running the joiner adds nothing, asserted directly rather than left to the
unique (image, post) constraint to catch.

Two bugs fixed in the writing, one of which my own test would have hit:

* `assign_to_group` sorted bare (distance, Post) tuples, which falls through
  to comparing Posts when two distances tie — and a perfectly symmetric
  bridge, the exact case the function exists for, would have raised TypeError
  instead of declining to choose. Now keyed on the distance alone.
* The cursor was still built from `post_date or downloaded_at` while the
  ORDER BY had gained `resurfaced_at`. Two expressions that disagree at a page
  boundary don't error, they silently skip or repeat rows; both sites now go
  through one `_post_sort_value`, and a test pages through one row at a time
  to prove the walk matches the whole list.

Image linking is now one shared helper rather than written twice, because
creation and joining would otherwise be free to drift on exactly the detail
(which post owns the image) that makes a grouping reversible.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-10 11:30:17 -04:00
bvandeusenandClaude Opus 5 7071c87cd6 feat: a grouped post says so, and the operator can tune the grouping (388 E2)
CI / lint (push) Successful in 2s
CI / extension-version (push) Successful in 2s
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 6s
CI / frontend-build (push) Successful in 22s
CI / backend-lint-and-test (push) Successful in 31s
Build images / build-web (push) Successful in 1m5s
Build images / smoke-web (push) Skipped
Build images / build-ml (push) Successful in 1m52s
Build images / promote (push) Skipped
CI / integration (push) Successful in 2m7s
Rule 27 — E2's other half. The backend can author posts; this is what makes
that visible and adjustable.

**The honesty marker.** A chip on every synthetic post's card: "grouped by
FabledCurator", titled with what it was built from ("Grouped from 4 Discord
messages"). This chip is the only thing standing between "FC assembled this"
and the card reading as something the artist authored, so it keys off nothing
but the flag, and it states the member count rather than just disclosing that
grouping happened — a claim you can check beats a claim you're asked to trust.

A synthetic post has no title on purpose (inventing one is the one place this
feature could put words in a creator's mouth), and the untitled fallback would
otherwise have printed the internal key: "Post fc-drop:99887766". It now names
the post for what it is. There's a test for that specifically.

**The tuning card.** Ingestion & filters gets a Discord-drop-grouping tile:
the switch, the distance cut and the drop window, each with the sentence that
tells the operator which one to reach for. The window's copy says outright
that it is the setting doing most of the work — without it, everything an
artist ever drew of one character collapses into a single post.

Both directions are pinned in postCard.spec.js, including the one that
actually matters: an ordinary post is never marked. Also covered — a post dict
composed before these fields existed degrades to unmarked rather than throwing
on `synthesis.message_count`.

Also fixes the ruff UP017 that failed the lint job on 73eeb7a (timezone.utc →
datetime.UTC, the convention everywhere else in this repo). The integration
suite on that SHA was green: all 12 grouping tests passed and migration 0092
applied cleanly.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-10 11:17:56 -04:00
bvandeusenandClaude Opus 5 73eeb7a377 feat: FC authors the post that Discord never wrote (milestone 388 step E2)
CI / lint (push) Failing after 3s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 5s
Build images / build-agent (push) Successful in 9s
CI / frontend-build (push) Successful in 25s
CI / backend-lint-and-test (push) Successful in 31s
Build images / build-web (push) Successful in 1m5s
Build images / smoke-web (push) Skipped
Build images / build-ml (push) Successful in 1m53s
Build images / promote (push) Skipped
CI / integration (push) Successful in 2m5s
Discord is a delivery channel, not a publisher. One message is not one post,
and today every message lands as its own `post` row, so chat lines compete
with authored work for the same surface. Rather than demote them into a
second-class feed, FC now writes the post itself: one row per DROP, its
images the drop's images, its body the messages' text in arrival order.

Synthesising a `Post` (rather than inventing a parallel entity) is the whole
point — the result is post-shaped by construction, so feed, provenance,
translation, attachments and series keep working on it unchanged.

The predicate is three axes ANDed, and the time one does the real work:

    same source  AND  cosine distance <= threshold  AND  no gap > window

Similarity alone over-groups, and that is the failure that would make this
useless: any two pieces of the same character by the same artist sit close in
SigLIP space, so a cosine-only rule collapses a month of one character into a
single "post". Two details inside the predicate are load-bearing —

* distance is measured to the group's SEED, never to the previous member,
  because chaining lets a group DRIFT: twenty small steps walk from one piece
  to a completely different one, each hop individually within threshold;
* the window is measured between CONSECUTIVE messages, not from the first, so
  an artist trickling variants out over an evening stays one drop.

Why a post-import sweep and not part of ingest. The obvious alternative was to
migrate Discord to the native post-first ingester (#1266) and group at capture
time. That cannot work: the grouping signal is `siglip_embedding`, which is
produced asynchronously AFTER import (tasks/ml.py, the GPU backfill), so at
capture time there is nothing to group on. Grouping is necessarily something
that happens once the vectors catch up — hence a re-runnable sweep that skips
what it cannot yet place, and an hourly (not daily) cadence.

The honesty rule, enforced in the schema. `post.synthesized_by` names the
grouper; `synthesis_details` records the members, the count, and the
thresholds AS THEY WERE (they are operator-tunable, so without that "why did
it group these" is unanswerable a month later). Member posts are absorbed, not
destroyed — they remain the images' true origin and the audit trail — and
`absorbed_by_post_id` is ON DELETE SET NULL, so deleting a synthetic post
releases its members back into the feed in one DELETE with no repair step.
`post_title` stays NULL deliberately: a synthesised title is the one place
this could put words in a creator's mouth.

Two guards the first draft would have failed:

* the per-run cap took the lowest post IDs, not the oldest posts — DISTINCT ON
  forces its own ORDER BY, so the sort now happens outside the subquery;
* a cap landing mid-drop would have published a truncated group claiming to be
  a whole drop, so the last group is left for the next run.

And one vacuous test caught before it shipped: the support vector perturbed a
single component of an all-ones vector, moving it ~1e-6, so every distance
assertion passed regardless of what the predicate did. `_vec` now builds a
unit vector at a stated angle, where distance is exactly 1 - cos(delta) —
rule 167, a guard has to be able to fail.

UI (rule 27) follows in the next commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-10 11:15:14 -04:00
bvandeusenandClaude Opus 5 4fe792b61c feat: platform_membership — the learned roster of what the account pays for (milestone 387 step C1)
CI / extension-version (push) Successful in 3s
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 8s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 39s
Build images / build-web (push) Successful in 1m14s
Build images / smoke-web (push) Skipped
Build images / build-ml (push) Successful in 2m18s
Build images / promote (push) Skipped
CI / integration (push) Successful in 2m21s
FC knows which creators it was TOLD to follow and nothing about which
ones the operator is subscribed to. Those two sets drift both ways and
neither drift is currently visible: a subscription FC doesn't track is
content the operator believes they're archiving and aren't, and a
source walked after the subscription lapsed is requests spent on a wall
reported as a creator gone quiet.

Sibling of service_seen (milestone 365) and the same insight — an
absence is only observable against a record of presence. touch_membership
reuses the recorded touch_service shape (snippet 3447): upsert rather
than read-modify-write, first_seen_at deliberately outside the update
set because it's the one field that makes a later DISAPPEARANCE
readable as a lapse rather than as a creator we never knew.

Nothing populates it yet, and that's the intended intermediate state.
The sweep (C3) needs a client seam (C2) that needs Patreon's real
response characterised from a captured sample (C0), which needs the
operator's browser session. The table's SHAPE doesn't wait on that,
because it's deliberately free-form exactly where C0's findings would
otherwise dictate a column.

status is an unconstrained String holding the PLATFORM's own word, not
a normalised FC value. Rule 36 considered and declined, same reasoning
service_seen.kind records: the vocabulary isn't ours to invent, and
picking a lowest-common-denominator enum before any platform has been
characterised would bake a guess into the schema. The service owns the
whitelist and the mapping; the column owns the evidence.

MEMBERSHIP_STATUS ships EMPTY, guarded by a test that fails if anyone
adds an entry — every one must come from a characterised response, not
from API docs. That's rule 130 at the one place it's easiest to break,
and the failure message says so.

has_paid_access returns None, never False, for a word it hasn't been
taught. The difference is load-bearing: False means the operator lost
access, which C4 turns into an offer to disable the source, so
asserting it from an unrecognised word would tell them to cancel a
subscription they're still paying for.

Retention decided here rather than deferred (rule 89): a membership
that stops appearing is aged out on time, never deleted on absence —
deleting would destroy the signal at the moment it became interesting.

Rule 90 check, done on the right thing this time: the per-test TRUNCATE
teardown derives its table list from Base.metadata.sorted_tables, so
the new table is picked up automatically; test_models asserts a subset,
so it doesn't break. 0091 follows 0090 on the collapsed baseline.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-10 10:58:25 -04:00
bvandeusenandClaude Opus 5 6b19012bb6 feat: the empty front door is the install's first screen (milestone 387 step B4)
CI / lint (push) Successful in 5s
CI / extension-version (push) Successful in 6s
CI / frontend-build (push) Successful in 26s
CI / backend-lint-and-test (push) Successful in 36s
CI / integration (push) Successful in 2m3s
Build images / sign-extension (push) Successful in 5s
Build images / build-agent (push) Successful in 9s
Build images / build-ml (push) Successful in 9s
Build images / build-web (push) Successful in 6s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
The front door is now a feed, and a blank feed implies things should be
here in a way a blank masonry does not. On a fresh install this is the
first screen anyone sees — including someone who is not the operator,
which is what milestone 328 is making possible.

Tells the two empties apart, which is the point. "No sources yet" gets
the on-ramp; "sources configured, nothing landed yet" gets told that
the first check takes a while and pointed at Downloads. Telling someone
to add a source when they already have three and are mid-backfill reads
as the app not knowing its own state.

Needs total_sources on schedule-status to distinguish them —
deliberately not auto_sources, which counts only what is on a schedule,
so a source with auto_check off would have read as "nothing
configured". Both exact-shape assertions updated in THIS change rather
than after CI caught them, which is the lesson from B3's red push.

An absent status falls back to the on-ramp on purpose: it is merely
redundant to an established operator, whereas "see what's running"
shown to someone with nothing configured is a dead end.

A filtered miss is deliberately NOT the onboarding case — the operator
has posts, they just narrowed past them. Showing a fresh-install
on-ramp there would tell someone with a full library to go set it up.

This is where logo.svg lands, as the operator asked. It earns its place
on a first-run screen and not on a populated feed, and gives the
on-ramp something to compose around instead of prose plus two buttons.
Large: the mark stops reading below ~48px, which is why the 22px nav
slot has a different one. Pinned by test so a later tidy-up cannot
quietly shrink it to a glyph.

Also extracts mountWithStore into the shared test support module.
Writing the second spec created exactly the copy-paste that open issue
3109 tracks for the backend row factories, so it is consolidated now
rather than at copy three, and recorded as snippet 3829 with the two
traps it does NOT solve — named slots rendering nothing, and
components that fetch on mount.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-09 23:10:06 -04:00
bvandeusenandClaude Opus 5 3f8306f705 fix: two exact-shape assertions pinned the old schedule-status payload
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 5s
Build images / build-ml (push) Successful in 9s
Build images / build-agent (push) Successful in 10s
CI / frontend-build (push) Successful in 24s
CI / backend-lint-and-test (push) Successful in 33s
Build images / build-web (push) Successful in 1m14s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / integration (push) Successful in 2m9s
B3 added failing_sources / no_access_sources to schedule-status and
broke test_schedule_status_shape and test_summary_returns_rollup_shape,
both of which assert the payload's EXACT key set. My own tests passed;
these two did not, and integration caught it.

A rule 90 miss: I grepped for the predicates I changed and not for
consumers of the response shape. The shape is the thing I actually
changed.

Both keys added to the assertions rather than loosening them to a
subset check — an exact-set assertion is what catches a key being
renamed out from under a consumer, which is precisely the value these
two tests just demonstrated.

The other readers (PipelineStatusChip, SchedulerStatusBar) pull
individual keys, so they were unaffected. SchedulerStatusBar's prop
comment documented the old shape and is corrected here; a comment that
lies about a contract is worth the same as a doc that does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-09 23:01:08 -04:00
bvandeusenandClaude Opus 5 a708f5e9db feat: the front door says whether ingestion is working (milestone 387 step B3)
CI / extension-version (push) Successful in 4s
CI / lint (push) Successful in 4s
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 9s
CI / frontend-build (push) Successful in 27s
CI / backend-lint-and-test (push) Successful in 34s
Build images / build-web (push) Successful in 1m17s
Build images / smoke-web (push) Skipped
CI / integration (push) Failing after 2m8s
Build images / build-ml (push) Successful in 2m18s
Build images / promote (push) Skipped
The step phase A was building toward. A1 made the gated count true, A2
made it a durable state, A3 made it visible in Subscriptions — but
Subscriptions is where you go once you already suspect something. This
is the line that reaches someone who wasn't looking.

A thin grey strip above the feed, front door only: last check, sources
failing, sources you can't see. Only the actionable items take a
colour, and nothing renders at zero — a permanent "0 failing" trains
you to skip the line, which would hide the real number when it appears.

Two predicates, defined once. The ribbon counts and the surfaces it
links to have to agree on what "failing" and "no access" MEAN, or the
ribbon says 3 and the card shows 4. They live in db_helpers, which
exists for exactly this reason (its docstring: divergent copies are how
the race bugs crept in). Not in source_service, because
scheduler_service needs them too and source_service already imports
scheduler_service — the other direction is a cycle.

Counting deliberately spans all ENABLED sources rather than the
auto_check subset scheduler_status already walks: a source erroring on
a manual-only artist is still erroring. Disabled sources count for
nothing, which is what makes issue 1285 the real escape hatch for a sub
you stopped paying for.

Extends the existing schedule-status endpoint rather than adding a
parallel aggregate — the store already fetches it. Two scalar COUNTs.

The status filter is now URL-addressable, which it had to be for the
ribbon's links to land anywhere: a count that drops you on an
unfiltered list makes the reader redo the filtering the ribbon just
did. Mirrors how artistFilter already reads from route.query.

Front-door-only via a route prop, not a route.name check, so the view
doesn't need to know what it's mounted as and the router states the
intent in one place. Inside Browse's Posts tab you're looking FOR
something and the hub is one click away.

The fetch is swallowed on mount by design (rule 164): this is an aside,
and the feed must render whether or not the status call succeeds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-09 22:55:14 -04:00
bvandeusenandClaude Opus 5 ecd72015a7 feat: the front door answers "what arrived?" (milestone 387 step B1)
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Successful in 3s
Build images / build-ml (push) Successful in 8s
Build images / build-agent (push) Successful in 9s
CI / extension-version (push) Successful in 3s
CI / frontend-build (push) Successful in 28s
CI / backend-lint-and-test (push) Successful in 35s
Build images / build-web (push) Successful in 1m19s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / integration (push) Successful in 2m47s
Moves FRONT_DOOR from /showcase to /latest. Showcase is a random
TABLESAMPLE — lean-back, and it can never tell you anything is wrong.
The feed is the only view where a failing source surfaces on its own,
as a creator who has gone quiet. The per-artist "new since last visit"
badges (#597) were a workaround for this view not being the door.

Showcase is demoted to a nav entry, not removed. Nothing is being
replaced, so rule 22's delete-the-legacy-path does not apply.

Deviation from the filed plan, deliberate: the step said write a new
LatestView.vue. Rejected — PostsView is ALREADY a self-contained feed
(own container, own store, infinite scroll, filters, deep-link
anchoring, empty state), and Browse only ever wrapped it in a tab
strip. A new view would have duplicated 231 working lines to gain
nothing. Mounting PostsView directly at its own route IS the whole
difference the promotion was after: a door you arrive at, not a hub
you navigate out of. Rule 28.

Backend untouched, as scoped — PostFeedService.scroll already does
cursor-paginated newest-first.

Two things this shook loose:

PostsView's deep-link "All posts" button was a hard `{ name: 'posts' }`,
which redirects into Browse. Correct while the view only ever rendered
inside Browse's tab; from the front door it would have yanked the
operator sideways into a different surface. Now returns to the current
route minus post_id, so Browse keeps its tab and any active scope.

The README claimed "A Showcase front page". That block is the SOURCE
the release notes quote (scripts/release_notes.py product_overview),
not a generated copy, so it is fixed here — a document contradicting
the code is the characteristic defect of the public-surface area.

No stickyChrome on the route: unlike Browse/Gallery/Settings this view
has no sticky sub-header for the nav to butt against.

The router spec pinned FRONT_DOOR to /showcase and now pins /latest,
plus that Showcase stayed reachable and in the nav — the demotion is
asserted, not assumed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-09 22:46:04 -04:00
bvandeusenandClaude Opus 5 7ca6ee0666 feat: new curator brand mark — traced logo + redrawn glyph
CI / lint (push) Successful in 4s
Build images / sign-extension (push) Successful in 4s
CI / extension-version (push) Successful in 3s
Build images / build-ml (push) Successful in 9s
Build images / build-agent (push) Successful in 10s
CI / frontend-build (push) Successful in 24s
CI / backend-lint-and-test (push) Successful in 35s
Build images / build-web (push) Successful in 1m50s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / integration (push) Successful in 3m5s
Replaces the placeholder folder glyph with the new logo: a white-gloved
hand presenting a framed work, which says what the app is for far
better than the old mark did.

Two assets, because one cannot serve both jobs. Measured, not assumed:
rendered at 22px the full logo is unreadable mush, so favicon.svg stays
a separate, much simpler mark.

logo.svg — traced from the source raster with potrace, then repainted
from the design tokens. Segmentation notes, since this is the part that
is easy to get wrong on a re-do: hue does NOT separate the glove from
the frame's highlights (both sit near 38 degrees) and neither does
saturation alone. The split is a connected-component fill seeded inside
the cuff, dilated first so it can cross the dark outline strokes that
cut the fingertips off from the palm.

The source plate was a warm brown (#1B1105), not the app's cool
obsidian (#14171A) — side by side it read as a logo sitting on its own
warmer card. It is dropped entirely: the mark is transparent and the
frame interior shows whatever surface hosts it.

The source gold was #AA7E39, which is within a couple of points of
accent.curator #A87338 — so the mark now shares one colour with
nav-active text and the wordmark rather than nearly sharing it. The
glove goes to text.parchment for the same reason.

favicon.svg — hand-drawn rather than traced. At 16px a traced mark
carries hundreds of wobble nodes that read as fuzz and can never be
tidied. Ring + frame + star merged into a blob at that size, so it
keeps two elements: the frame and the star. Frame over a plain ring
because it carries the meaning, and it is the full logo's own
centrepiece; the glove, cufflink, finials and sparkle rays are
deliberately absent rather than drawn and lost.

The favicon keeps its obsidian plate so the tab icon is self-contained
against any browser chrome; on the nav that plate is invisible because
it matches --fc-chrome-rgb exactly. logo.svg has no plate at all.

Both files carry a comment explaining why they are shaped this way, so
the next edit does not undo the reasoning. The old favicon is one
revert away in history.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-09 22:20:00 -04:00
bvandeusenandClaude Opus 5 6bb18050a4 feat: no-access is visible per source, and findable (milestone 387 step A3)
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 3s
Build images / build-agent (push) Successful in 9s
CI / frontend-build (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 33s
CI / integration (push) Successful in 2m40s
Build images / build-ml (push) Successful in 2m47s
Build images / build-web (push) Successful in 1m35s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
A3 of milestone 387, completing phase A. A1 made the count true, A2
made it a durable state; this makes it something the operator can see
without going looking.

Turned out smaller than filed, because A2 revealed why the existing
`tier_limited` palette entry in FailingSourcesCard had never rendered:
the chip was being cleared by the same successful run that produced it.
The colour was already chosen.

Where it surfaces:

- SourceHealthDot gains a `no-access` grade. Deliberately its own grade
  rather than folded into healthy (which hides it) or warning (which
  sends the operator hunting for a break that isn't there). A source
  with real failures still grades as failing whether or not it is also
  gated.
- SourceRow gets an info-coloured lock chip in the status cell, which
  was empty for these sources — they have zero failures. Placed ahead
  of the backfill states: "we can't see this creator" is the more
  useful thing to say than which walk phase it is in, and unlike those
  it does not resolve on its own.
- A "No access" status filter, deliberately separate from "Has errors".
  Without it a gated source is invisible in a long list, because it
  correctly stays out of the failing rollup.

Left OUT of NeedsAttentionCard on purpose. That card's only affordance
is Retry, and you cannot retry your way into a subscription tier —
issue 1285 already gives the real escape hatch, since disabling a
source clears its state. Nothing structural needed changing: the card
is fed by consecutive_failures > 0, which a tier-limited source never
has.

The count lives on the download event, not the source, so `list()`
joins it in with one DISTINCT ON query — selecting the run_stats
sub-object rather than whole metadata blobs, which carry up to 500KB of
truncated stdout each. Scoped to tier-gated rows only, so a healthy
library issues no extra query at all. Absent stays None rather than 0,
and both UI surfaces phrase the state without a number when it is
missing instead of printing a fabricated zero.

Also covers A1's live gated count, which shipped untested, and extends
the mount helper with slot stubs: SourceHealthDot puts the dot in a
NAMED slot, and unresolved Vuetify components render default slots
only — so those assertions would have found an empty wrapper and
passed vacuously.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-09 21:26:48 -04:00
bvandeusenandClaude Opus 5 7751715b83 feat: a paywalled creator is no longer indistinguishable from a silent one (milestone 387 step A2)
CI / extension-version (push) Successful in 5s
Build images / sign-extension (push) Successful in 5s
CI / lint (push) Successful in 6s
Build images / build-agent (push) Successful in 12s
CI / frontend-build (push) Successful in 36s
CI / backend-lint-and-test (push) Successful in 42s
Build images / build-web (push) Successful in 1m24s
Build images / smoke-web (push) Skipped
Build images / build-ml (push) Successful in 2m30s
Build images / promote (push) Skipped
CI / integration (push) Successful in 2m34s
A2 of milestone 387. A1 made the gated-post count true; this makes it
mean something.

A native walk that reached the bottom returned `error_type=None`
whether the creator had posted nothing or every post sat behind a tier
we don't hold. `source.error_type` stayed NULL and the source read
healthy and quiet. gallery-dl has classified this as TIER_LIMITED since
the paywall-as-"needs attention" complaint; the native path never did.

Two things had to move that the plan didn't foresee, both found by
reading the consumers rather than by testing afterwards:

The backfill lifecycle's completion test required `error_type is None`.
Returning TIER_LIMITED naively would have dropped a fully-paywalled
backfill into the not-finished branch — zero downloads means no
progress, two strikes marks it "stalled" — so the creator we can see
least would become the one we re-walk most. `walk_completed` now admits
informational classes.

`_update_source_health` only stamps `error_type` on status "error" and
CLEARS it on "ok". Since TIER_LIMITED is a success, the chip was wiped
by the very run that produced it — which is why FailingSourcesCard's
`tier_limited` palette entry has never been reachable. An "ok" run now
keeps an informational class while failures stay 0 and last_error stays
clear: the run did not fail and must not earn a backoff.

Deviation from the plan, deliberate: the filed step said classify only
when `downloaded == 0`. gallery-dl doesn't condition on that, and
diverging the two backends over the same concept is what rule 169
forbids — so the native path mirrors it. "There is content here you
aren't paying for" is equally true in a week we also got the cheap
posts. Pinned by a test, since the stricter rule looks more correct.

The predicate, the wording and the completion test are defined once in
gallery_dl.py and spread into both backends (snippet 3087), rather than
re-derived per half — which is exactly how they drifted apart before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-09 15:40:19 -04:00
bvandeusenandClaude Opus 5 173f4b00aa fix: the native path never reported tier-gated posts (milestone 387 step A1)
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 4s
CI / frontend-build (push) Successful in 26s
CI / backend-lint-and-test (push) Successful in 55s
CI / integration (push) Successful in 2m4s
Build images / sign-extension (push) Successful in 11s
Build images / build-agent (push) Successful in 2m41s
Build images / build-web (push) Successful in 2m3s
Build images / smoke-web (push) Skipped
Build images / build-ml (push) Successful in 5m40s
Build images / promote (push) Skipped
`make_run_stats` has always declared `tier_gated_count`, and
DownloadDetailModal has always rendered it. gallery-dl populated it;
`ingest_core._result` did not — it built run_stats with six keys and let
the seventh default to 0, while the very same walk counted gated posts
into `gated_skipped` and spent the number on a log line.

So on Patreon, SubscribeStar and pixiv — the three platforms we now own —
the Downloads modal read "Tier-gated: 0" for a walk that skipped N
paywalled posts. A creator we've lost access to was indistinguishable
from a creator who stopped posting. Migrating Patreon off gallery-dl is
what dropped the signal.

Pass the count through, and tick it in the live-progress payload too, so
a long backfill on an inaccessible creator explains itself while it runs
rather than only at finalization. ActiveDownloadsPanel renders it only
when non-zero, coloured 'info' to match the severity FailingSourcesCard
already assigns tier_limited — this is not a failure.

Tests assert the run_stats key the UI actually reads rather than the
ingester's internal counter, so the guard tracks the property and not a
name. Falsification is structural: `make_run_stats` defaults the key to
0, so both assertions fail against the pre-fix call.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
2026-09-09 12:34:08 -04:00
bvandeusenandClaude Opus 5 ad8392b790 fix: system health is a Settings tab, not a page only the dot reached
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 5s
Build images / build-ml (push) Successful in 6s
Build images / build-agent (push) Successful in 8s
CI / frontend-build (push) Successful in 22s
CI / backend-lint-and-test (push) Successful in 35s
Build images / build-web (push) Successful in 54s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / integration (push) Successful in 1m47s
The surface shipped at /system with no nav entry, reachable only by
clicking the health dot beside the brand — a target you have to already
suspect something is wrong to go looking for. Operator-flagged: it needs
a path someone can walk to.

Settings is where you go to ask the instance about itself, so the view
becomes a tab there, beside Activity — Activity answers "what is the
queue doing", System answers "is anything left to do it".

- SystemView.vue moves to components/settings/SystemHealthTab.vue; the
  content is unchanged apart from shedding its own container and h1.
- SettingsView adopts useTabQuery (the composable Browse and
  Subscriptions already use) so a tab can be linked TO. The health dot
  now points at ?tab=system, and /system redirects there so the previous
  build's link and any bookmark still land.
- The tab drops its own 10s poll. v-window keeps a visited item mounted
  rather than destroyed, so that timer would have gone on firing behind
  Maintenance — and TopNav already polls the same store every 15s for
  the dot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 19:45:54 -04:00
bvandeusenandClaude Opus 5 5084ba666b feat: the dot beside the brand now means the whole stack (milestone 365 step 4)
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Successful in 5s
CI / extension-version (push) Successful in 5s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 9s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 31s
Build images / build-web (push) Successful in 55s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / integration (push) Successful in 1m43s
The ask was a surface AND a path. The path is the part that was missing —
everything that could answer "is it running" lived inside Settings, which you
only open once you already suspect something.

**Re-used the indicator that already existed rather than adding a fourth.**
There were three partial surfaces: TopNav's health dot, PipelineStatusChip's
pulse, and the Settings Activity tab. None answered "is every part alive", and
a fourth would have made the question harder to answer, not easier.

TopNav's dot read /api/health — a no-DB liveness check proving only that the
WEB container is serving. Green there while a worker was dead is exactly what
it looked like, and a green dot beside the product name gets read as
"everything is fine". It now reflects the whole-stack verdict, and it is a
link: the place someone already looks when they suspect something is now also
the way to the detail.

The tooltip names the actual problem. "Scheduler has not checked in for 6 min"
sends someone somewhere; "something is unhealthy" sends them hunting.

/system is deliberately NOT in the nav row — TopNav builds that from routes
with a meta.title, and a sixth top-level tab for a page visited twice a year
costs more attention than it returns. It is reached from the dot.

The page lists every learned part with its state as a sentence rather than a
chip, and prints the staleness thresholds it was judged by, taken from the
endpoint so the UI keeps no second copy of them. PipelineStatusChip still
hand-rolls its own 3-minute scheduler window; that is now a duplicate of a
threshold the server owns, and worth collapsing once this has been watched
working.

The stores stay separate on purpose: system.js is "can I reach the API",
systemActivity.js is "what is the pipeline doing", systemHealth.js is "is
anything broken". Running and alive fail independently — an idle stack with a
dead worker looks identical to a healthy one on every activity surface, which
is the whole reason this milestone exists.

Not yet verified against a real stopped service. Rule 12 keeps a local stack
out of it, and frontend CI has no Vue type-check or visual regression, so this
needs an operator look rather than a green lane.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 17:20:55 -04:00
bvandeusenandClaude Opus 5 fe4e0f2b71 feat: /api/system/health — one verdict for the whole stack (milestone 365 step 3)
CI / lint (push) Successful in 5s
Build images / sign-extension (push) Successful in 5s
Build images / build-agent (push) Successful in 8s
CI / extension-version (push) Successful in 4s
CI / frontend-build (push) Successful in 20s
CI / backend-lint-and-test (push) Successful in 35s
Build images / build-web (push) Successful in 1m1s
Build images / smoke-web (push) Skipped
CI / integration (push) Successful in 1m51s
Build images / build-ml (push) Successful in 1m59s
Build images / promote (push) Skipped
The single endpoint the nav indicator and the System page will both read.
Composing a verdict is this module's job, not the UI's.

Two kinds of part, answered differently. LEARNED — celery roles and the GPU
agent, out of service_seen, where the question is "how long since it checked
in" and the answer can be "it has not". PROBED — Postgres and Redis, always
expected, never learned, because a last-seen for them would be actively
misleading: that Redis answered thirty seconds ago says nothing about now.

**The endpoint must never fail because something it checks has failed.** That
inversion is easy to write by accident and it destroys the feature exactly
when it is needed — a 500 when Redis is down instead of `redis: down`. Every
probe is wrapped, every wait carries a deadline (rule 156), and the roster
refresh swallows its own errors. The worst case is a part reported `unknown`,
which is a true statement about the system.

Postgres is probed first and gates the rest, because if it is unreachable
nothing else can be read — and "the database is down" is the most useful
single thing this can ever say.

The staleness thresholds are the design risk, not the code, and they are
deliberately generous: 90s to doubt, 300s to disbelieve. The constraint is a
deploy rather than a crash — `docker compose up -d` rolls start-first, so a
role is briefly served by two containers and then by neither while the old one
drains. Thresholds tight enough to catch a crash in seconds would paint the
page red on every update, and an alarm that cries wolf on every deploy is one
nobody reads. Tune down only after watching a real deploy pass through. The
numbers ship in the response so the UI can explain a `stale` without keeping a
second copy of them.

States are described in sentences rather than left as chips: "Scheduler has
not checked in for 6 min — treat it as stopped" is what someone needs at the
moment they are deciding whether to go and open Portainer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 17:17:22 -04:00
bvandeusenandClaude Opus 5 dc8af8b1a7 feat: a learned roster, so a stopped part is observable (milestone 365 steps 1-2)
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 30s
Build images / build-web (push) Successful in 55s
Build images / smoke-web (push) Skipped
Build images / build-ml (push) Successful in 1m45s
Build images / promote (push) Skipped
CI / integration (push) Successful in 1m49s
Nothing in FabledCurator knew what was SUPPOSED to be running. `celery
inspect` reports the workers that ANSWER, so a dead worker was a shorter list
rather than a red light, and grep for any notion of expected services returned
nothing. That is why Portainer was the only place an operator could see it:
Portainer knows the intended set.

`service_seen` is the memory that makes an absence observable — every part
that has checked in, and when it last did.

**Keyed on the queue set, not the worker hostname.** Celery's worker names
here are `celery@<container id>`, minted fresh on every deploy. Keyed on those,
this table would record a death and a birth every time the stack updates — and
a status page that goes red on every deploy is a status page nobody reads,
which is worse than not having one. CELERY_QUEUES is assigned per role in
compose and survives container replacement, so it is the stable identity. Two
replicas of a role are therefore ONE row, which is right: the question is
whether the role is served, not how many containers exist.

The GPU agent is keyed on agent_id, the identity its lease protocol already
uses. gpu.py received it on both lease and heartbeat and threw it away — an
idle agent with nothing to lease left no trace and was indistinguishable from
one switched off a week ago. Now recorded on the calls that were already
happening.

**Who observes, corrected from the plan.** The plan said "record from the
existing inspect path", which would only run when someone opened the Activity
tab. Two other candidates and why they lost:

- A beat sweep. If the scheduler dies the sweep stops, every row goes stale,
  and the page says everything is down when one thing is. An alarm that cannot
  distinguish "a part died" from "the observer died" is worse than none.
- A background task in web. hypercorn runs --workers 4, so that is four
  concurrent inspect loops per container, forever.

Taken instead: refresh on demand, rate-limited by the newest last_seen_at that
every process can already see. The observer is then the thing serving the page
— if web is down you get a browser error, not a confidently green page — and
it self-limits with no coordination, since a race costs one redundant inspect
that writes identical values.

Migration 0090 is the first written on the collapsed baseline (milestone 328),
so it is also the first evidence the chain steps FORWARD from 0089 rather than
merely reproducing the schema. No secondary indexes: one row per moving part
means every read is a handful of rows, and #3301 is the record of what
speculative indexes cost.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 17:15:46 -04:00
bvandeusenandClaude Opus 5 131237143b Revert "test: force the smoke gate to fail, to watch it block a publish"
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 6s
Build images / build-ml (push) Successful in 8s
Build images / build-agent (push) Successful in 9s
CI / frontend-build (push) Successful in 22s
extension / lint (push) Successful in 21s
Build images / build-web (push) Successful in 6s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / backend-lint-and-test (push) Successful in 34s
CI / integration (push) Successful in 2m0s
extension / lint (pull_request) Successful in 20s
The gate held. Run 5320, dispatched with the forced failure in place:

  build-web    success   (candidate published)
  build-ml     success
  build-agent  success
  smoke-web    FAILED
  promote      skipped
  run          failure

And the three channel tags did not move:

  fabledcurator        33d3d8332f74 -> 33d3d8332f74
  fabledcurator-ml     e94a5435cb45 -> e94a5435cb45
  fabledcurator-agent  bae27d34d811 -> bae27d34d811

So a refresh that breaks something now leaves :latest naming the build that
works, which is the property milestone 362 exists to establish. The rejected
candidate is still published under :refresh-candidate, so whoever reads the
red job on Monday can pull the exact image that failed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 16:23:49 -04:00
bvandeusenandClaude Opus 5 59d27ef76e test: force the smoke gate to fail, to watch it block a publish
CI / lint (push) Successful in 4s
Build images / sign-extension (push) Successful in 4s
CI / extension-version (push) Successful in 3s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 8s
CI / frontend-build (push) Successful in 19s
extension / lint (push) Successful in 18s
Build images / build-web (push) Successful in 7s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / backend-lint-and-test (push) Successful in 31s
CI / integration (push) Successful in 1m57s
TEMPORARY, reverted in the next commit. Milestone 362's verification section
requires the gate to be seen rejecting a build — a gate nobody has watched
reject anything is a gate nobody knows is wired up. Every real check passes,
so the rejection has to be forced.

Under test is the job dependency, not the assertions: a failed smoke-web must
skip the promote job, and the three :latest tags must still name the digests
they named before the run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 16:22:01 -04:00
bvandeusenandClaude Opus 5 f630e50e75 ci: the refresh publishes only what the gate passed (#3265 milestone step 4)
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 5s
Build images / build-ml (push) Successful in 8s
Build images / build-agent (push) Successful in 9s
Build images / build-web (push) Successful in 6s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
extension / lint (push) Successful in 19s
CI / frontend-build (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 34s
CI / integration (push) Successful in 1m53s
The gate reported a verdict nothing consulted. Now it decides.

The promote moved out of the three build jobs into its own `promote` job,
because the verdict cannot exist until build-web has finished and the promote
used to run inside it. `needs: [build-web, build-ml, build-agent, smoke-web]`
is the whole mechanism: a failed smoke skips the promote, so a refresh that
broke something leaves :latest naming the build that works. "The refresh
failed" and "production is broken" must not be the same event.

A SKIPPED smoke also skips it, and that is the case that matters most. On run
5290 the gate silently skipped itself — job-level `if:` cannot read the env
context — and a design where only a FAILED gate blocks would have published
unverified images while reporting success. Not running is not the same as
passing, and today produced two separate bugs of exactly that shape (#3414,
and the smoke-web skip).

All three images now promote together or not at all. They are one stack:
build.yml already refuses to publish a :dev web image beside a stale :dev ml
because the mismatch only surfaces as a runtime failure, and a refresh that
published ml while withholding web would be that same trap reached through the
gate. Stated plainly in the job comment: the gate covers web only, so ml and
agent are held to web's verdict rather than their own. That is the
conservative direction, not equivalent evidence, and should not be read as if
it were.

Three near-identical promote steps collapsed into one loop. A partial failure
now says which images moved and that the state is inconsistent, rather than
leaving that to be inferred — the promote is idempotent and the candidates are
still published, so the instruction is simply to re-run.

Also removed the now-dead `promote` output from the ml and agent reuse steps.
Only build-web's is read (as outputs.candidate); two more copies nothing
consults is the kind of thing that reads as load-bearing a year later.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 16:21:27 -04:00
bvandeusenandClaude Opus 5 86abaf0b94 docs: a new install could not start, and nothing told anyone why (#3422)
CI / lint (push) Successful in 4s
Build images / sign-extension (push) Successful in 4s
CI / extension-version (push) Successful in 3s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 8s
Build images / build-web (push) Successful in 6s
Build images / smoke-web (push) Skipped
CI / frontend-build (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 31s
CI / integration (push) Successful in 1m46s
The install path milestone 328 wrote produces a web container that exits on
boot. entrypoint.sh runs alembic, then app construction raises:

  MissingCredentialKey: Fernet key file not found at
  /images/secrets/credential_key.b64. For first-time setup, set
  CURATOR_BOOTSTRAP_NEW_KEY=1.

That variable appeared in no README, no .env.example and no compose file —
only in backend/. So a stranger following the documented steps got an app
that does not start and an error with no context. Found by the milestone-362
smoke gate on its first real run (#3422).

The product behaviour stays exactly as it is. credential_crypto refuses to
mint a key because the 2026-06-02 audit found a partial restore — database
back, ./images/secrets lost — silently generating a fresh one and producing a
healthy-looking instance where every authenticated download failed AUTH_ERROR.
Failing fast is right; not saying so is the bug.

So: .env.example carries the variable in its own FIRST BOOT ONLY section with
the reasoning and an instruction to delete the line afterwards, and README's
First run leads with it, because "the app will not start" belongs before "the
ML worker downloads weights". Both say to back up ./images/secrets/ alongside
the database, which is the part that costs real data if it is learned late.

**compose had to change too, and this is the part that would have shipped a
second broken instruction.** A variable in `.env` is only used for ${...}
interpolation — it does not reach the container unless the service names it.
Telling people to set it in .env, without that, would have documented a step
that does nothing. Added to the shared app_env anchor, defaulted to empty so
the refusal still stands for everyone who has not opted in.

Not taken: auto-bootstrapping when the credential table is empty, which would
remove the manual step entirely and keep the audit's protection for restores.
That is the better product and it is a code change with a predicate that has
to be exactly right; this is the smallest correct fix, and #3422 stays open
for the other one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 16:19:28 -04:00
bvandeusenandClaude Opus 5 4815040d74 ci: the smoke gate found a real one on its first run — and had two bugs of its own
Build images / sign-extension (push) Successful in 6s
CI / lint (push) Successful in 6s
CI / extension-version (push) Successful in 4s
Build images / build-ml (push) Successful in 11s
Build images / build-agent (push) Successful in 13s
extension / lint (push) Successful in 24s
CI / frontend-build (push) Successful in 28s
CI / backend-lint-and-test (push) Successful in 35s
Build images / build-web (push) Successful in 8s
Build images / smoke-web (push) Skipped
CI / integration (push) Successful in 2m5s
Run 5296 was `smoke-web`'s first genuine execution. Checks 1 and 2 passed:
alembic built the schema from empty inside the image, all five apt binaries
resolved, and the application's own Thumbnailer produced JPEG, PNG-with-alpha,
WebP and an ffmpeg video frame against the image's libraries. Check 3 failed,
and the trap's log dump said exactly why:

  MissingCredentialKey: Fernet key file not found at
  /images/secrets/credential_key.b64. For first-time setup, set
  CURATOR_BOOTSTRAP_NEW_KEY=1.

That is the product being right. credential_crypto refuses to mint a key
unless someone opts in, because the 2026-06-02 audit found a partial restore
(DB back, /images/secrets/ lost) silently generating a fresh one and leaving a
working-looking system where every authenticated download failed AUTH_ERROR.

It is also a first-run blocker for milestone 328, filed as #3422: the variable
appears in no README, no .env.example and no compose file, so the install path
that milestone just finished writing produces a container that exits on boot.
Not fixed here — the fix trades safety against friction and is the operator's
call.

Two defects in the gate itself, both surfaced by the same run:

- A throwaway CI instance IS first-time setup, so it now passes
  CURATOR_BOOTSTRAP_NEW_KEY=1. The check was asserting a condition no fresh
  container can satisfy.

- The health loop polled a dead container for 3m35s. Docker had already
  recycled its IP, so the replies were a baffling mix of connection-refused
  and 5s timeouts from whatever took the address next. It now checks
  `.State.Running` each iteration and fails immediately with the container's
  log. The trap had the real answer the whole time; this stops burying it
  under four minutes of noise.

Also corrected a message claiming a 120s budget: 60 iterations of up to 5s
connect plus 2s sleep is nearer seven minutes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 16:00:18 -04:00
bvandeusenandClaude Opus 5 81b7b6f308 ci: smoke-web never ran — a job's if: cannot read the env context
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 4s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 8s
CI / frontend-build (push) Successful in 22s
extension / lint (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 32s
Build images / build-web (push) Successful in 8s
Build images / smoke-web (push) Skipped
CI / integration (push) Successful in 1m51s
Run 5290 dispatched a refresh. Everything worked: the guard fired, the build
published the candidate, the promote pointed :latest at it. And `smoke-web`
reported conclusion "skipped", with no steps and no log.

Its condition was `if: env.IS_REFRESH == 'true'`. The env context is available
to STEP conditions and step bodies but never to a job's own `if:`, and an
unresolvable context there evaluates to empty rather than erroring. So the
gate skipped itself, silently, on the one run that existed to exercise it.

Second silent-skip of this family today, after #3414. Same shape both times:
something evaluated false, nothing failed, and the run reported success. It is
worth naming the pattern — on this pipeline, "green" and "ran" are different
claims, and the steps' own conclusions are the only place the difference shows.

Fixed by keying off a job output rather than re-deriving the trigger:
build-web now exposes the reuse step's `promote` decision as `outputs.candidate`
and smoke-web consumes it. That is better than duplicating the expression:
it is the same single decision the build, the XPI download and the promote all
take already — build.yml's own "one decision drives everything downstream" —
and it asserts the thing smoke-web actually depends on, that a candidate was
published, rather than restating the reason one would be.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 15:53:40 -04:00
bvandeusenandClaude Opus 5 bfa9fd678b ci: smoke the refreshed image against real Postgres and Redis
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 3s
CI / extension-version (push) Successful in 4s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 9s
Build images / build-web (push) Successful in 7s
Build images / smoke-web (push) Skipped
extension / lint (push) Successful in 20s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 32s
CI / integration (push) Successful in 1m48s
extension / lint (pull_request) Successful in 20s
Milestone 362 step 3. This is the gate the weekly base refresh never had.

`ci.yml` cannot be that gate, and the reason matters more than the fix. Its
lanes run on ci-python:3.14 and install requirements.txt — a base refresh
changes neither, so all five stay green through a bump that breaks the product.
What a refresh re-resolves is the Dockerfile's apt layer:

    ffmpeg unar libpq5 postgresql-client zstd megatools
    libjpeg62-turbo libwebp7 libpng16-16 ca-certificates

Unpinned, every build, and nothing else in this repo looks at it. That line is
the dependency creep; it is also precisely what the test suite structurally
cannot observe, since the suite never runs inside the image and the image
carries no tests and no pytest.

So `smoke-web` runs the CANDIDATE IMAGE against real service containers:

  1. `alembic upgrade head` on an empty database — the image's own libpq and
     psycopg, and the same call entrypoint.sh makes before it serves anything,
     so a failure here is a failure to boot.
  2. The apt binaries, then the application's own `Thumbnailer` — JPEG, PNG
     with alpha, WebP, and a video frame through ffmpeg. `Thumbnailer` needs no
     database and no app context, so the check exercises real product code
     rather than a proxy for it. `ffmpeg -version` exiting 0 would pass while a
     codec removal broke every thumbnail in the library.
  3. The web role boots and answers /api/health.

Every failure names the package it implicates. This fires on a Sunday,
unattended, about a change nobody made deliberately — "assertion failed" a week
later teaches nobody anything.

The script is piped over stdin rather than bind-mounted: the workspace is a
docker volume belonging to the job's own container, so a host bind of $PWD does
not resolve for a sibling. Container logs are dumped only on failure, and the
trap re-exits with the real status rather than the status of `docker rm`.

Deliberately NOT gating the promote yet — that is step 4. Landing the gate and
the thing it gates together would mean the first time anyone saw this job run
would also be the first time it could stop a publish.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 15:45:10 -04:00
bvandeusenandClaude Opus 5 24a2b70a5a ci: a boolean input never equals the string 'true'
Build images / sign-extension (push) Successful in 4s
Build images / build-ml (push) Successful in 6s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 23s
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 2s
CI / backend-lint-and-test (push) Successful in 35s
Build images / build-web (push) Successful in 8s
CI / integration (push) Successful in 1m48s
The refresh lever did not work, and the way it did not work is the point.

Run 5270 dispatched with refresh=true. Its log:

  expression '(github.event_name == 'schedule'
               || github.event.inputs.refresh == 'true') && 'true' || 'false''
    evaluated to '%!t(string=false)'
  trigger: event=workflow_dispatch IS_REFRESH='false' BUILD_REF='refs/heads/dev'
  trigger: raw inputs refresh='true' force_build='false'

The input arrived as true and the comparison still said false. `type: boolean`
delivers a real boolean, and GitHub expression semantics cast operands to
numbers when their types differ — so `true == 'true'` compares 1 against NaN.
My comment on the previous commit asserted the opposite, that Forgejo delivers
inputs as strings, and asserted it without checking.

The run went GREEN with every step skipped, because a refresh that evaluates
false is indistinguishable from an ordinary push. A lever that silently does
nothing is worse than no lever: it would have been trusted.

Normalised through format(), which is representation-independent — a boolean
true and a string 'true' both render 'true'. That is also why force_build was
never bitten: it passes its raw value into an env var and compares in the
shell, where everything is a string already. format() buys the same thing at
expression level, which is where a step `if:` needs the answer.

The diagnostic from the previous commit stays. It is what turned this from a
guess into a measurement, and it is the only thing that would catch the same
class of failure next time.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 14:59:48 -04:00
bvandeusenandClaude Opus 5 2c88ad3efb ci: report the raw and normalised trigger values
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
CI / extension-version (push) Successful in 2s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 7s
CI / backend-lint-and-test (push) Successful in 31s
Build images / build-web (push) Successful in 9s
extension / lint (push) Successful in 16s
CI / frontend-build (push) Successful in 22s
CI / integration (push) Successful in 2m39s
The refresh dispatch on run 5265 went green with every step skipped: the
main-only guard did not fire, checkout took dev, and the reuse step read
IS_REFRESH as false. So both workflow-level expressions evaluated false while
the identical accessor works for force_build, which compares its value in the
shell rather than in an expression.

That is a guess until it is measured, and the failure is silent by
construction — a refresh that evaluates false behaves exactly like an ordinary
push and reports success. This prints the raw input beside the normalised
value in the step that already exists to say what a run derived, so the two
disagreeing is visible rather than inferred.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 14:58:20 -04:00
bvandeusenandClaude Opus 5 bfc4f9cec9 ci: one fact for "is this a base refresh", and a lever to trigger one
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / build-agent (push) Successful in 7s
Build images / build-ml (push) Successful in 7s
Build images / build-web (push) Successful in 8s
CI / frontend-build (push) Successful in 20s
extension / lint (push) Successful in 18s
CI / backend-lint-and-test (push) Successful in 37s
CI / integration (push) Successful in 1m45s
extension / lint (pull_request) Successful in 24s
Milestone 362, enabling step 2's verification and everything after it.

The weekly refresh was testable once a week. That is not a cadence anything
can be developed against, and milestone 362's whole point is a gate — which
has to be watched rejecting something before anyone can believe it is wired
up. So `refresh` joins `force_build` as a dispatch input, on the same
reasoning that added that one (#3252: confirm #3190 was gone rather than wait
for it to recur).

Adding it meant confronting that "is this a refresh?" was asked in five
places and spelled five ways: `github.event_name == 'schedule'` in an `if:`,
`$GITHUB_EVENT_NAME` in one shell, an `EVENT:` env passed into another, and a
bare expression on `pull:`. Five spellings of one fact is how half of them
come to disagree once somebody adds a sixth trigger — which is precisely what
this commit is. So it is derived once at the top, next to BUILD_REF, which
already exists for exactly this reason on exactly this question.

String comparison, not boolean: Forgejo delivers dispatch inputs as strings,
so `inputs.refresh` is 'true'/'false' and `&&` on it would read the string
'false' as truthy.

**A constraint this makes visible, which pre-dates it.** A refresh checks out
`main` (BUILD_REF) while running the workflow definition from the branch that
triggered it — the cron registers from the default branch. So dev's workflow
builds main's source, and dev's workflow cannot depend on anything main's
tree does not have yet. It does now: the reuse step calls `artifacts.sh
epoch`, which lands on main with this batch. Until then a refresh dispatch
fails loudly at that call, which is the right failure — the alternative is
tolerating a missing epoch and silently rebuilding #3265 into every refresh.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 14:41:33 -04:00
bvandeusenandClaude Opus 5 b590d25f8f ci: the scheduled refresh builds a candidate, then names the channel
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 2s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 8s
Build images / build-web (push) Successful in 6s
CI / frontend-build (push) Successful in 24s
extension / lint (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 32s
CI / integration (push) Successful in 1m50s
Milestone 362 step 2. Structural: it creates a moment between "built" and
"published" for step 3's gate to occupy. No behaviour change.

A refresh rebuilds against freshly resolved base images, and the web image's
runtime is a line of UNPINNED Debian packages — ffmpeg, libjpeg62-turbo,
libpq5, megatools — re-resolved on every build. Nothing in ci.yml can see
that: its lanes run on ci-python:3.14 and install requirements.txt, and a base
bump changes neither. So refreshed bytes need proving before :latest names
them, and proving needs somewhere to stand.

On a push nothing changes: build_ref IS channel_ref, promote is false, and the
build writes the channel tag directly the way it always has. On the schedule
the build writes :refresh-candidate — one moving ref per image, overwritten in
place, holding a build nobody is told to pull. That is the shape rule 145
already allows for :buildcache, not the per-build tag family 318 withdrew.

Both values are decided in the reuse step beside `hit`, because that step
already owns "what does this job do" (build.yml's own rule, at the force
branch). A promote condition derived somewhere else could disagree with the
tag the build actually wrote.

**The promote is a manifest PUT, not `imagetools create`.** That distinction is
the whole risk in this change. `imagetools create` wraps its source in an
index, and an indexed channel tag is the one thing this pipeline cannot
survive: `.Image.Config.Labels` does not resolve through an index, so the
fc.revision the reuse check reads back would come up empty, every later push
would miss and rebuild, and nothing would go red. That is #3183, observed on
run 4751 — reuse worked exactly once and the only symptom was the bill. The
repoint step already excludes its own source tag for this reason; a promote
that re-introduced the wrap through another door would undo that care.

A manifest PUT is what "make this tag name that image" means at the registry:
same bytes, same media type, identical digest, no layer transfer. It reads the
result back and fails if the tag does not name what was just written — a PUT
that 2xx'd and landed something else is exactly the silent-and-plausible
failure this pipeline keeps producing. Every call carries a deadline (rule
156); a registry that stops answering must fail the step, not hang the weekly
refresh until the job times out.

Promote is UNCONDITIONAL today, deliberately. Gating it before the gate exists
would leave the refresh building something and publishing nothing for as long
as this milestone takes. Step 4 wraps it in the smoke suite's verdict.

Not yet verified on the refresh path — that needs a scheduled run, and the
lever to trigger one on demand is the next commit. This one is verified by the
push path being untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 14:40:19 -04:00
bvandeusenandClaude Opus 5 635138b0d1 ci: pin the build clock to the commit, so an unchanged refresh publishes nothing
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / build-ml (push) Successful in 6s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 20s
extension / lint (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 33s
Build images / build-web (push) Successful in 56s
CI / integration (push) Successful in 1m50s
Milestone 362 step 1, closing #3265's root cause.

The weekly base refresh rewrote all three `:latest` tags on 2026-08-30 with
nothing changed in any of them. Not a cache miss — run 4934's log shows every
content step CACHED and both bases resolved to unchanged pinned digests.
buildkit stamps the image config with the wall clock of the build, so
identical layers get republished under a new config blob and therefore a new
manifest digest.

The cost is not storage, it is meaning: `:latest` moved on a calendar, so a
digest change stopped being evidence that anything was different. That is the
one thing a digest is any use for, and it is load-bearing here — the reuse
check, the `:c-<sha>` rollback story and any future redeploy signal all rest
on it.

SOURCE_DATE_EPOCH normalises `created` and the history timestamps, so the same
source produces the same config bytes and the same digest, and pushing it is a
registry no-op.

The value is routed through artifacts.sh's existing `newest()` rather than
taken from git separately. `revision`, `version` and now `epoch` are three
fields of ONE lookup, so they cannot drift into naming different commits — a
divergence that would stamp an image reproducibly against one commit while it
reported being another, with both values looking perfectly well-formed. Note
#3127 §2 is the record of what a second clock costs; this adds a view, not a
clock.

Also corrected: the build step comment and ci-requirements.md both described
the churn as current behaviour with the fix as a "likely" future. They now
describe what the file does.

Tests pin the property the fix depends on, not the fix: epoch is the same
commit version names, in both renderings including the extension's unpadded
one, and it does not move between two calls on one checkout. A future
refactor that gave epoch its own `git log` would pass every other test in
that file.

Not yet verified end to end — proving it needs two consecutive refreshes to
land on the same digest, which is the next thing, and is the step #3265 exists
because nobody did last time.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 13:46:16 -04:00
bvandeusenandClaude Opus 5 c0370069e0 release: the first release describes the product; it has nothing to diff against
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 5s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 8s
Build images / build-web (push) Successful in 7s
CI / frontend-build (push) Successful in 22s
CI / backend-lint-and-test (push) Successful in 36s
CI / integration (push) Successful in 1m42s
Step 7 needs a release that reads as "what is FabledCurator and how do I run
it". What the script would actually have published is "changes since
v26.06.04.0" over 533 commits, truncated to 200 — a release page whose first
screen is the internal build-out that milestone 328 exists to stop shipping,
addressed to a reader who has never seen this project.

Two causes, fixed separately.

**A pre-convention tag is history, not a predecessor.** The 28 `v26.*` tags
were kept when their releases were deleted, so `--match v*` walks ancestry
straight back to one of them. Reachable is not comparable: nobody has run
v26.06.04.0 and its release page no longer exists to compare against. The
match is now `v[0-9][0-9][0-9][0-9].*` — rule 148's shape, which is exactly
the set of tags naming a release a reader could have been running.

**With that narrowed, the first rule-148 tag reaches no predecessor**, and the
old fallback — diff against the whole history — is worse than the problem it
replaced. A release with no predecessor now renders the product overview and
no commit list at all.

The overview is READ OUT OF README.md between `<!-- overview:start -->` and
`<!-- overview:end -->`, not written into the script, for the same reason the
changelog is derived: two hand-maintained descriptions of one product drift
and nothing ever catches it. The release page and the repo front page are one
source. Missing markers are reported as a note and publish anyway, on
cross_checks()'s reasoning — the release is still the useful object.

Every later release goes back to being a changelog, which is what note #3127
§5 says a release is for. MAX_COMMITS still guards the case it now guards:
two real releases far enough apart that the list stops being readable.

Also corrected while marking up the README: "Importing — ingests an existing
library from disk" was still advertising the folder-import feature that
3590c47 documented as deliberately retired. Replaced with what FC actually
does with what arrives — content-hash dedup, sidecar metadata, provenance.

Tests: the two that asserted the old no-predecessor behaviour are rewritten
rather than left; synthetic repos now carry their own copy of the script,
since the overview resolves relative to `__file__` (correct in production,
where release.yml checks out the tag) and would otherwise have every fixture
silently quoting FabledCurator's real README.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 12:28:48 -04:00
bvandeusenandClaude Opus 5 3590c478f5 docs+ci: folder import stays retired, and fix a readiness probe that never probed
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
CI / extension-version (push) Successful in 3s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 13s
Build images / build-web (push) Successful in 7s
CI / frontend-build (push) Successful in 22s
extension / lint (push) Successful in 29s
CI / backend-lint-and-test (push) Successful in 49s
CI / integration (push) Successful in 1m57s
extension / lint (pull_request) Successful in 22s
Two unrelated things, both found while closing out milestone 328.

**Folder import (#3367).** The operator's call, this session: the
import-from-file surface was abandoned on purpose and is not coming back —
"it has its own complexities that we didn't need." The README and the
compose comment both described the missing button as a rough edge with a
tracking issue, which promised a fix that is not coming. Both now say the
retirement is the decision, name Subscriptions as the supported way to fill
a new install, and describe /api/import/trigger as an unsupported escape
hatch for anyone who wants to script one.

**The CI readiness probe.** ci.yml's integration job and baseline.yml both
waited for Postgres with `(echo > /dev/tcp/$PG_IP/5432)`. Those steps run
under `sh -e` — act's default shell — where /dev/tcp is not a magic path but
a filename that does not exist. The probe could therefore never succeed: run
18035, a GREEN run, spends 05:20:53 → 05:22:53 in that loop and exits it by
exhaustion, not by connecting. Every integration run has been paying a flat
120s for a check that established nothing, and proceeding regardless.

Replaced with a socket connect in python (present in the image, no package
needed), and exhausting the budget is now a named failure instead of a
silent fall-through — rule 156.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 00:51:25 -04:00
bvandeusenandClaude Opus 5 8a4af589f1 docs: write the install path for someone who is not the operator (#3271)
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 9s
Build images / build-ml (push) Successful in 29s
Build images / build-web (push) Successful in 23s
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 4s
CI / frontend-build (push) Successful in 20s
CI / backend-lint-and-test (push) Successful in 30s
CI / integration (push) Successful in 3m42s
FC has no login — no User model, no session auth, nothing. That was a
deliberate call for a single-operator tool and it stays (operator, this
session), but it was nowhere in the docs, and the app stores live Patreon /
SubscribeStar / Pixiv session cookies on accounts that carry a payment
method. Anyone standing this up from the README could reasonably have put it
behind a TLS-terminating proxy and considered it handled.

So the no-auth posture is now stated three times, in the three places someone
decides where to bind the port: README has a "Before you expose it" section
above the install instructions, .env.example explains why there is no auth
variable in it, and the compose header says it before the first service.

SECURITY.md claimed the opposite. It listed "a multi-user sharing ACL —
instances can be shared" among the things worth protecting; there are no
accounts to share between. That was rule 47 applied to a codebase that does
not implement it, and it would have told a researcher FC holds a boundary it
does not. Replaced with the real posture, including that TLS without an
authenticating layer in front changes nothing.

Also corrected, all of it stale rather than wrong-at-the-time:

- EXTENSION_API_KEY was dead config. config.py read it into a field nothing
  consumed; the real key is generated into app_setting on first use and
  managed in the UI. Removed from config.py, compose and .env.example.
- .env.example pointed at docs/superpowers/specs/… — there is no docs/ dir —
  and described the extension key as "lands in FC-3", closed 2026-05-21.
- The /import mount comment described an FC-5 ImageRepo migration run from
  "Settings → Maintenance → Legacy migration", a surface with no frontend.
- README said the extension installs from Settings → Maintenance. It is on
  Subscriptions → Settings.

README is now split: running FC above the line, developing FC below it, with
requirements, first run, the extension, upgrading and troubleshooting on the
running side. First run documents the one real gap it found — a new installer
with a library on disk has no button to import it, only POST
/api/import/trigger, because the manual-scan UI was retired 2026-07-02 when
that stopped mattering for an established install. Filed as #3367.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017QHszn9H8VBvx5Ke8x1hvw
2026-09-02 00:16:23 -04:00
130 changed files with 14234 additions and 530 deletions
+97 -18
View File
@@ -1,24 +1,103 @@
# Database
DB_USER=fabledcurator
DB_PASSWORD=changeme_use_a_real_password
DB_HOST=postgres
DB_PORT=5432
DB_NAME=fabledcurator
# FabledCurator configuration.
#
# Copy to `.env` and edit before your first production start:
#
# cp .env.example .env
#
# Only the two values under CHANGE THESE actually need your attention. The
# rest have working defaults baked into docker-compose.yml and are listed
# here so you know they exist, not because you have to set them.
#
# Almost nothing else lives here on purpose. FabledCurator is configured from
# its own Settings UI, backed by the database — no restart, no YAML. If you
# are looking for where to set an import path, a download schedule or an ML
# threshold, it is in the app, not in this file.
# Redis / Celery
CELERY_BROKER_URL=redis://redis:6379/0
CELERY_RESULT_BACKEND=redis://redis:6379/0
# App
# Generate with: openssl rand -hex 32
SECRET_KEY=changeme_32_byte_hex_secret
# ---------------------------------------------------------------------------
# CHANGE THESE
# ---------------------------------------------------------------------------
# Extension API key — used in FC-3, lands later but reserved now
# Generate with: openssl rand -hex 32
EXTENSION_API_KEY=
# The Postgres password. docker-compose.yml falls back to a published default
# (`fabledcurator_dev`) so that `docker compose up` works with no config at
# all — which is exactly why you must not leave it at that on a real install.
# It is the credential protecting your stored platform session cookies.
DB_PASSWORD=
# Logging
# Sets Quart's app.secret_key. Today it signs nothing: FabledCurator has no
# login and uses no session cookies, so no value here is protecting anything
# right now. Set it anyway. It is required at boot rather than defaulted so
# that the day something session-backed does land, no instance is already
# running on a value published in this file.
#
# openssl rand -hex 32
SECRET_KEY=
# ---------------------------------------------------------------------------
# FIRST BOOT ONLY — then delete this line
# ---------------------------------------------------------------------------
# FabledCurator encrypts your stored platform credentials with a Fernet key it
# keeps at /images/secrets/credential_key.b64 — inside the ./images bind mount,
# so it outlives the container. On a brand-new install that file does not exist
# yet, and the app REFUSES TO START rather than quietly create one:
#
# MissingCredentialKey: Fernet key file not found at
# /images/secrets/credential_key.b64
#
# That refusal is deliberate. Auto-creating a key is indistinguishable from the
# disaster case — a restore that brought the database back but lost
# ./images/secrets — and there it would mint a key that cannot decrypt anything,
# leaving an instance that looks healthy while every paywalled download fails.
# So the choice is yours to make explicitly, once.
#
# Set this for your first `up`, watch the container come up, then DELETE THE
# LINE. Leaving it set disarms the protection permanently, on an instance that
# by then has credentials worth protecting.
#
# BACK UP ./images/secrets/ ALONGSIDE YOUR DATABASE. The key is the only thing
# that can read your stored credentials; a database restored without it needs
# every credential re-entered by hand.
CURATOR_BOOTSTRAP_NEW_KEY=1
# ---------------------------------------------------------------------------
# Optional — defaults are fine
# ---------------------------------------------------------------------------
# Host port the UI is published on. The container always listens on 8080;
# this is only the left-hand side of the port mapping.
PORT=8080
# DEBUG | INFO | WARNING | ERROR
LOG_LEVEL=INFO
# Deployment posture: plain HTTP (no TLS in the app; reverse proxy if needed)
# See docs/superpowers/specs/2026-05-13-fabledcurator-merge-design.md §2.1
# Postgres identity. Change these only if you are pointing at a database you
# manage yourself — the bundled postgres service is created with whatever is
# set here, so changing them after the first start will not rename anything.
DB_USER=fabledcurator
DB_NAME=fabledcurator
# Set by docker-compose.yml to reach the bundled services. Override only when
# running Postgres or Redis outside this stack.
# DB_HOST=postgres
# DB_PORT=5432
# CELERY_BROKER_URL=redis://redis:6379/0
# CELERY_RESULT_BACKEND=redis://redis:6379/0
# ---------------------------------------------------------------------------
# There is no authentication variable here, and that is not an omission
# ---------------------------------------------------------------------------
#
# FabledCurator has no login, no accounts and no permission model. Anything
# that can reach PORT is an administrator and can read the platform session
# cookies the app stores for Patreon and SubscribeStar.
#
# Bind it to a trusted network. See "Before you expose it" in README.md and
# the deployment posture section of SECURITY.md.
#
# The Firefox extension's API key is NOT configured here — it is generated
# automatically on first use and shown under Settings → Maintenance, where you
# can also rotate it.
+17 -5
View File
@@ -87,10 +87,21 @@ jobs:
test -n "$PG_IP"
echo "PG_CONTAINER=$PG" >> "$GITHUB_ENV"
echo "DB_HOST=$PG_IP" >> "$GITHUB_ENV"
# Socket probe in python, not bash's /dev/tcp — these steps run under
# `sh -e`, where that path does not exist. Same fix and same reasoning
# as ci.yml's integration job; see the comment there.
pg_ready=""
for i in $(seq 1 60); do
(echo > "/dev/tcp/$PG_IP/5432") >/dev/null 2>&1 && break
if python -c "import socket,sys; s=socket.socket(); s.settimeout(2); sys.exit(0 if s.connect_ex(('$PG_IP', 5432)) == 0 else 1)"; then
pg_ready=1
break
fi
sleep 2
done
if [ -z "$pg_ready" ]; then
echo "postgres at $PG_IP:5432 did not accept a connection within 120s"
exit 1
fi
if command -v uv >/dev/null 2>&1; then
uv pip install --system -r requirements.txt
else
@@ -161,10 +172,11 @@ jobs:
mkdir -p /tmp/versions_held
mv alembic/versions/*.py /tmp/versions_held/ 2>/dev/null || true
DB_NAME=fc_gen alembic revision --autogenerate -m "baseline" || true
# Printed rather than uploaded: ci-requirements.md records that this
# runner cannot do actions/upload-artifact@v4+, and the repo dropped
# the action entirely in 2026-05, so the job log is the retrieval
# channel actually proven here.
# Printed rather than uploaded: the repo dropped actions/upload-artifact
# in 2026-05, when the runner could not run v4+, and the job log is the
# retrieval channel this job has proven. (gitea/runner 3.x runs stock
# upload-artifact now — Scribe snippet #2271 — so an artifact is an
# option if the log ever stops being enough.)
#
# base64, not the raw file. A plain `cat` of the ~33KB candidate was
# TRUNCATED MID-LINE by the runner on run 4964 — it stopped inside
+571 -64
View File
@@ -44,6 +44,10 @@ on:
description: 'Rebuild every image even if the published revision matches'
type: boolean
default: false
refresh:
description: 'Behave as the weekly base refresh: build main against fresh bases, publish through the candidate tag'
type: boolean
default: false
# The base-image refresh (milestone 326 step 4, #3154).
#
@@ -72,8 +76,43 @@ on:
# Deriving it per job invites the two halves to disagree: sign-extension would
# derive dev's extension version while build-web bundled main's, and the
# release download would 404 on a version that exists perfectly well.
# IS THIS A BASE REFRESH? Asked in five places and previously spelled five
# ways — `github.event_name == 'schedule'` in an `if:`, `$GITHUB_EVENT_NAME` in
# one shell, an `EVENT:` env passed into another, and a bare expression on
# `pull:`. Five spellings of one fact is how half of them come to disagree
# after somebody adds a sixth trigger.
#
# The `refresh` dispatch input is here so this path can be EXERCISED. A weekly
# cron is otherwise testable once a week, which is not a cadence anything can
# be developed against — the same reason `force_build` exists (#3252, added to
# confirm #3190 was gone rather than wait for it to recur). It is also what
# makes the milestone-362 gate verifiable at all: a gate has to be watched
# rejecting something before anyone can believe it is wired up.
#
# The input is normalised through `format()` before it is compared, and that
# is not defensive styling — the direct comparison is WRONG and fails silently.
#
# `type: boolean` delivers a real boolean, and GitHub expression semantics cast
# operands to numbers when their types differ: `true == 'true'` compares 1
# against NaN and is FALSE. Measured on run 5270, whose own log says it —
#
# expression '(github.event_name == 'schedule'
# || github.event.inputs.refresh == 'true') && 'true' || 'false''
# evaluated to '%!t(string=false)'
# trigger: raw inputs refresh='true'
#
# — the input arrived as `true` and the expression still said false. The run
# then went green with every step skipped, because a refresh that evaluates
# false behaves exactly like an ordinary push. That is the whole hazard: the
# failure has no symptom.
#
# `force_build` never hit this because it never compares in an expression. It
# passes the raw value into an env var and tests it in the shell, where
# everything is already a string. `format('{0}', x)` buys the same thing here,
# where a step-level `if:` needs the answer before any shell runs.
env:
BUILD_REF: ${{ github.event_name == 'schedule' && 'main' || github.ref }}
IS_REFRESH: ${{ (github.event_name == 'schedule' || format('{0}', github.event.inputs.refresh) == 'true') && 'true' || 'false' }}
BUILD_REF: ${{ (github.event_name == 'schedule' || format('{0}', github.event.inputs.refresh) == 'true') && 'main' || github.ref }}
# Requires repo secret RELEASE_TOKEN — a Forgejo PAT with scopes:
# - write:package, read:package (for docker push to git.fabledsword.com)
@@ -143,7 +182,7 @@ jobs:
# evaluate — this file already gates steps on it — so the guard cannot
# be disabled by the same uncertainty it exists to cover.
- name: Guard — a scheduled run must have checked out main
if: github.event_name == 'schedule'
if: env.IS_REFRESH == 'true'
run: |
set -eu
BRANCH=$(git rev-parse --abbrev-ref HEAD)
@@ -406,13 +445,26 @@ jobs:
trap - EXIT
echo "Uploaded fabledcurator-$VERSION.xpi to ext-$VERSION release"
# No actions/upload-artifact step: Forgejo Actions (and our
# act_runner) doesn't support upload-artifact@v4+ (GHES limitation
# surfaced 2026-05-26). Instead build-web reads the signed XPI
# straight from the ext-<version> Forgejo release we just uploaded
# to. Same source of truth; no double-store.
# No actions/upload-artifact step: build-web reads the signed XPI
# straight from the ext-<version> Forgejo release we just uploaded to.
# Same source of truth; no double-store. The step was dropped 2026-05-26
# because act_runner could not run upload-artifact@v4+; gitea/runner 3.x
# can (Scribe snippet #2271), but the release asset stays the better
# channel for a file build-web needs on every run.
build-web:
# Consumed by smoke-web's job-level `if:`. It cannot read `env` — the env
# context is available to STEP `if:` and step bodies, never to a job's own
# condition, and an unresolvable context there is empty rather than an
# error. `smoke-web` skipped silently on run 5290 for exactly that reason.
#
# Keying off the reuse step's own output is better than re-deriving the
# trigger anyway: it is the same single decision the build, the XPI
# download and the promote all take (build.yml's "one decision drives
# everything downstream"), and it says the thing smoke-web actually needs
# to know — a candidate was published — rather than restating why.
outputs:
candidate: ${{ steps.reuse.outputs.promote }}
# A plain `needs` — no `always()`. That expression existed to let a
# SKIPPED sign-extension through on a tag push while still blocking a
# FAILED one. With no tag trigger, sign-extension always runs, so the
@@ -437,7 +489,7 @@ jobs:
# See sign-extension's copy for why this guard exists.
- name: Guard — a scheduled run must have checked out main
if: github.event_name == 'schedule'
if: env.IS_REFRESH == 'true'
run: |
set -eu
BRANCH=$(git rev-parse --abbrev-ref HEAD)
@@ -470,8 +522,18 @@ jobs:
# the PREVIOUS XPI while the freshly signed one is orphaned (#3156).
# * dev and main derive the same values for the same source.
- name: Report the derived artifact version
env:
# Diagnostic for the trigger normalisation. `refresh` is reported RAW
# as well as normalised, because the two disagreeing is the whole
# failure mode: a dispatch input whose type does not compare the way
# the expression assumes evaluates to false silently, and the only
# symptom is a refresh that quietly behaves like an ordinary push.
RAW_REFRESH: ${{ github.event.inputs.refresh }}
RAW_FORCE: ${{ github.event.inputs.force_build }}
run: |
set -u
echo "trigger: event=$GITHUB_EVENT_NAME IS_REFRESH='${IS_REFRESH:-<unset>}' BUILD_REF='${BUILD_REF:-<unset>}'"
echo "trigger: raw inputs refresh='${RAW_REFRESH:-<unset>}' force_build='${RAW_FORCE:-<unset>}'"
A=web
V=$(sh scripts/artifacts.sh version "$A" 2>&1 || echo UNAVAILABLE)
R=$(sh scripts/artifacts.sh revision "$A" 2>&1 || echo UNAVAILABLE)
@@ -528,7 +590,7 @@ jobs:
# Checked BEFORE the ref test, not after: a scheduled run's
# GITHUB_REF is the default branch (dev), so the main test would
# never fire on it.
if [ "${GITHUB_EVENT_NAME:-}" = "schedule" ]; then
if [ "${IS_REFRESH:-}" = "true" ]; then
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator:latest" >> "$GITHUB_OUTPUT"
echo "channel=main" >> "$GITHUB_OUTPUT"
elif [ "${GITHUB_REF##*/}" = "main" ]; then
@@ -628,7 +690,6 @@ jobs:
# A scheduled refresh has to bypass reuse by construction: it
# rebuilds the SAME source, so fc.revision always matches and the
# check would skip every refresh there has ever been.
EVENT: ${{ github.event_name }}
run: |
set -eu
DERIVED=$(sh scripts/artifacts.sh revision web)
@@ -638,11 +699,60 @@ jobs:
# adds no variability the reuse check would have to account for.
echo "version=$(sh scripts/artifacts.sh version web)" >> "$GITHUB_OUTPUT"
# The build clock, pinned to the same commit (#3265). Without it
# buildkit stamps the image config with the wall clock of the build,
# so identical layers republish under a new config blob and the
# channel tag gets a new manifest digest for no reason. Derived from
# `newest()` like revision and version, so all three name one commit
# and cannot drift apart.
echo "epoch=$(sh scripts/artifacts.sh epoch web)" >> "$GITHUB_OUTPUT"
# The moving tag for this channel. Which tag we ask IS the channel —
# that is why the revision needs no -main/-dev qualifier any more.
if [ "$CHANNEL" = "main" ]; then T=latest; else T=dev; fi
echo "channel_ref=$IMAGE:$T" >> "$GITHUB_OUTPUT"
# WHERE THE BUILD PUBLISHES, which is not always the channel — and
# whether the channel then has to be written separately.
#
# On a push the build writes the channel tag directly: the bytes came
# from a commit, and a commit is the thing CI tests. Nothing to hold
# it behind.
#
# On the scheduled refresh it writes a CANDIDATE tag instead. A
# refresh rebuilds against freshly resolved base images, and the web
# image's runtime is a line of UNPINNED Debian packages (ffmpeg,
# libjpeg62-turbo, libpq5, megatools…) re-resolved on every build.
# Nothing in ci.yml can see that: its lanes run on ci-python:3.14 and
# install requirements.txt, and a base bump changes neither. So
# refreshed bytes have to be proven before :latest names them, and
# proving needs a moment between "built" and "published" to occupy.
# This is that moment; :latest goes on naming the build that works
# until something says otherwise.
#
# `:refresh-candidate` is one moving ref per image, overwritten in
# place, holding a build nobody is told to pull — the shape rule 145
# already allows for :buildcache, not the per-build tag family that
# milestone 318 withdrew.
#
# Decided HERE, beside `hit`, for the reason the force/schedule
# branch below gives: one step decides what this job does. A
# condition derived independently could disagree with the tag the
# build actually wrote.
#
# build-web additionally exposes this as `outputs.candidate`, which is
# what gates the `promote` job — a job's `if:` cannot read `env`, and
# one flag is enough because all three derive it from the same
# IS_REFRESH. ml and agent do not re-emit it; a second copy nothing
# reads is the kind of thing that later reads as load-bearing.
if [ "${IS_REFRESH:-}" = "true" ]; then
echo "build_ref=$IMAGE:refresh-candidate" >> "$GITHUB_OUTPUT"
echo "promote=true" >> "$GITHUB_OUTPUT"
else
echo "build_ref=$IMAGE:$T" >> "$GITHUB_OUTPUT"
echo "promote=false" >> "$GITHUB_OUTPUT"
fi
# Compare VALUES, never exit codes. Measured on buildx v0.36.1
# (run 4732): a missing key returns an empty string and exits 0, so
# branching on the exit code would read "no label yet" as success.
@@ -675,7 +785,7 @@ jobs:
if [ "${FORCE:-false}" = "true" ]; then
echo "hit=false" >> "$GITHUB_OUTPUT"
echo "reuse: force_build set — building regardless"
elif [ "${EVENT:-}" = "schedule" ]; then
elif [ "${IS_REFRESH:-}" = "true" ]; then
echo "hit=false" >> "$GITHUB_OUTPUT"
echo "reuse: scheduled base refresh — building regardless"
elif [ -n "$PUBLISHED" ] && [ "$PUBLISHED" = "$DERIVED" ]; then
@@ -764,6 +874,12 @@ jobs:
- name: Build and push web image
if: steps.reuse.outputs.hit != 'true'
# Read by buildx out of the ENVIRONMENT, not passed as a build-arg —
# it normalises the image config's `created` field and the history
# timestamps rather than being consumed by the Dockerfile. See #3265
# and the reuse step's `epoch` output.
env:
SOURCE_DATE_EPOCH: ${{ steps.reuse.outputs.epoch }}
uses: docker/build-push-action@v5
with:
context: .
@@ -776,20 +892,17 @@ jobs:
# invalidates, and the image genuinely rebuilds.
#
# MEASURED on the first real fire, run 4934 (#3265): when the base
# did NOT move, the build is ~13s and every content step reports
# CACHED — but the channel tag STILL gets a new manifest digest.
# buildkit mints a fresh image config each run, so identical layers
# are republished under a new config blob. All three images moved
# that way on 2026-08-30 with nothing whatsoever changed in them.
# did NOT move, the build was ~13s with every content step CACHED —
# and the channel tag STILL got a new manifest digest, because
# buildkit stamps a fresh image config per run and republishes the
# identical layers under it. All three images moved that way on
# 2026-08-30 with nothing whatsoever changed in them.
#
# So a refresh currently rewrites :latest every Sunday whether or
# not there is anything new in it, and :c-<sha> is handed a new
# manifest to diverge from on the same cadence. Layers are shared,
# so the storage cost is a config blob; the cost that matters is
# that a digest change no longer MEANS anything. Tracked in #3265 —
# the likely fix is a deterministic SOURCE_DATE_EPOCH, which would
# make "same source, same bytes" true and turn the no-op case into
# a genuine no-op.
# SOURCE_DATE_EPOCH (below) is the fix: pinned to the commit the
# content came from, the config is byte-identical across runs, so
# the manifest digest is too and the push is a registry no-op. A
# digest change means the content changed again, which is the only
# thing a digest is any use for.
#
# What `pull` does NOT catch either: a Debian package update inside
# the `apt-get install` layer while the base tag itself stands
@@ -799,14 +912,14 @@ jobs:
# churn #3265 is about.
#
# Only on the schedule. An ordinary push wants the cached base.
pull: ${{ github.event_name == 'schedule' }}
pull: ${{ env.IS_REFRESH == 'true' }}
# ONE tag, the channel's. Every other tag is written by the step
# below, registry-side. buildx here pushes the first tag to the
# registry and then re-pushes the rest through the DOCKER driver,
# out of a local image store a registry-direct build never filled —
# #3190, which cost `main` its :c-<sha> on 2026-08-29 while :latest
# published perfectly well.
tags: ${{ steps.reuse.outputs.channel_ref }}
tags: ${{ steps.reuse.outputs.build_ref }}
# The reuse key. Read back off the channel tag on the next push to
# decide whether that push needs to build at all, so this is not
# decoration — an unstamped image is one that will always rebuild.
@@ -937,6 +1050,282 @@ jobs:
docker buildx imagetools create $ARGS "$SOURCE"
echo "repointed from $SOURCE:$ARGS"
# Does the image a refresh just built still work?
#
# This is the gate the base refresh never had. `ci.yml` cannot be it: its
# lanes run on ci-python:3.14 and install requirements.txt, and a base bump
# changes neither — all five stay green through a refresh that breaks the
# product. What a refresh re-resolves is the Dockerfile's apt layer (ffmpeg,
# unar, libpq5, postgresql-client, zstd, megatools, libjpeg62-turbo,
# libwebp7, libpng16-16), unpinned, every build.
#
# So this runs the CANDIDATE IMAGE, against real Postgres and Redis. Not the
# source tree, and not a static inspection: `ffmpeg -version` exiting 0 would
# pass while a codec removal broke every thumbnail in the library.
#
# Refresh-only. On a push the bytes came from a commit, and a commit is what
# ci.yml already tests.
#
# Reports a verdict; it does not yet gate the promote (milestone 362 step 4).
# Landing the gate and the thing it gates in one change would mean the first
# time anyone saw this job run would also be the first time it could stop a
# publish.
smoke-web:
needs: [build-web]
if: needs.build-web.outputs.candidate == 'true'
runs-on: python-ci
container:
image: git.fabledsword.com/bvandeusen/ci-python:3.14
env:
DB_USER: fabledcurator
DB_PASSWORD: ci_smoke
DB_PORT: "5432"
DB_NAME: fabledcurator_smoke
SECRET_KEY: ci_smoke_placeholder
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator
services:
postgres:
image: pgvector/pgvector:pg16
env:
POSTGRES_USER: fabledcurator
POSTGRES_PASSWORD: ci_smoke
POSTGRES_DB: fabledcurator_smoke
options: >-
--health-cmd "pg_isready -U fabledcurator"
--health-interval 10s
--health-timeout 5s
--health-retries 10
redis:
image: redis:7-alpine
options: >-
--health-cmd "redis-cli ping"
--health-interval 10s
--health-timeout 5s
--health-retries 10
steps:
- uses: actions/checkout@v4
with:
# The same ref the image was built from, so the smoke script matches
# the code inside the candidate.
ref: ${{ env.BUILD_REF }}
- name: Smoke the candidate image
env:
TOKEN: ${{ secrets.RELEASE_TOKEN }}
ACTOR: ${{ github.actor }}
run: |
set -eux
# Service discovery mirrors ci.yml's integration lane: these jobs run
# in a container against a mounted docker socket, so the services are
# SIBLINGS reachable by IP, not by hostname.
PG=$(docker ps --filter "name=smoke" --filter "ancestor=pgvector/pgvector:pg16" -q | head -n1)
RD=$(docker ps --filter "name=smoke" --filter "ancestor=redis:7-alpine" -q | head -n1)
test -n "$PG" && test -n "$RD"
PG_IP=$(docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' "$PG")
RD_IP=$(docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' "$RD")
test -n "$PG_IP" && test -n "$RD_IP"
# Socket probe in python, not bash's /dev/tcp — these steps run under
# `sh -e`, where that path does not exist. Same fix and reasoning as
# ci.yml's integration job; see the comment there.
pg_ready=""
for i in $(seq 1 60); do
if python -c "import socket,sys; s=socket.socket(); s.settimeout(2); sys.exit(0 if s.connect_ex(('$PG_IP', 5432)) == 0 else 1)"; then
pg_ready=1
break
fi
sleep 2
done
if [ -z "$pg_ready" ]; then
echo "postgres at $PG_IP:5432 did not accept a connection within 120s"
exit 1
fi
echo "$TOKEN" | docker login git.fabledsword.com -u "$ACTOR" --password-stdin
CANDIDATE="$IMAGE:refresh-candidate"
docker pull "$CANDIDATE"
ENVOPTS="-e DB_USER=$DB_USER -e DB_PASSWORD=$DB_PASSWORD -e DB_HOST=$PG_IP"
ENVOPTS="$ENVOPTS -e DB_PORT=5432 -e DB_NAME=$DB_NAME -e SECRET_KEY=$SECRET_KEY"
ENVOPTS="$ENVOPTS -e CELERY_BROKER_URL=redis://$RD_IP:6379/0"
ENVOPTS="$ENVOPTS -e CELERY_RESULT_BACKEND=redis://$RD_IP:6379/0"
# A throwaway CI instance IS first-time setup, which is the one case
# credential_crypto allows a key to be minted in. Without it the web
# role refuses to boot — deliberately, since silently generating a
# key on a restored-DB-but-lost-secrets deployment would leave every
# Credential row undecryptable (the 2026-06-02 audit). Discovered by
# this job on its first real run; see #3422 for the fact that no
# user-facing file mentions this variable at all.
ENVOPTS="$ENVOPTS -e CURATOR_BOOTSTRAP_NEW_KEY=1"
# 1. The schema builds from empty, using the image's OWN libpq and
# psycopg. This is the same call entrypoint.sh makes before it
# serves anything, so a failure here is a failure to boot.
echo "smoke: alembic upgrade head"
docker run --rm $ENVOPTS "$CANDIDATE" alembic upgrade head
# 2. The apt layer's binaries and the app's own thumbnail path, run
# inside the image. Piped over stdin rather than bind-mounted: the
# workspace is a docker VOLUME belonging to this job's container,
# so a host bind of $PWD would not resolve for a sibling.
echo "smoke: image-internal checks"
docker run --rm -i $ENVOPTS "$CANDIDATE" shell -c 'python3 -' < scripts/smoke_image.py
# 3. It actually serves. `docker run -d` then poll the container's own
# IP — no port publishing, because the job container reaches
# siblings directly and a published port would collide with
# whatever else the runner is hosting.
echo "smoke: web boots and answers /api/health"
CID=$(docker run -d $ENVOPTS "$CANDIDATE" web)
# Clean up the container however this ends, and dump its log ONLY
# on failure — a boot that never answers must fail with the reason
# visible rather than as a bare timeout (rule 156), while a green run
# has nothing to say. `exit $rc` preserves the real status, which a
# trap that ends on a successful `docker rm` would otherwise mask.
trap 'rc=$?; [ $rc -eq 0 ] || docker logs "$CID" 2>&1 | tail -40; docker rm -f "$CID" >/dev/null 2>&1 || true; exit $rc' EXIT
WEB_IP=$(docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' "$CID")
test -n "$WEB_IP"
healthy=""
for i in $(seq 1 60); do
if curl -fsS --max-time 5 "http://$WEB_IP:8080/api/health" >/dev/null 2>&1; then
healthy=1
break
fi
# A container that has EXITED will never answer, so stop asking.
# Without this the loop spent 3m35s polling a dead container on
# this job's first run, and — because docker recycles the IP — got
# a confusing mix of connection-refused and 5s timeouts from
# whatever took the address next. The trap's log dump had the real
# answer the whole time; this just stops burying it.
if [ "$(docker inspect -f '{{.State.Running}}' "$CID" 2>/dev/null)" != "true" ]; then
echo "smoke: FAILED — the web container exited during boot." >&2
echo "smoke: its log follows; entrypoint runs alembic BEFORE" >&2
echo "smoke: serving, so a startup exception lands here." >&2
exit 1
fi
sleep 2
done
if [ -z "$healthy" ]; then
# 60 iterations of (up to 5s connect + 2s sleep) — up to ~7min, not
# the 120s an earlier version of this message claimed.
echo "smoke: FAILED — web is running but never answered" >&2
echo "smoke: /api/health. It is up, so look at hypercorn and the" >&2
echo "smoke: python base rather than at startup." >&2
exit 1
fi
curl -fsS --max-time 5 "http://$WEB_IP:8080/api/health"
echo
echo "smoke: all checks passed against $CANDIDATE"
# Move the channel tags — the whole point of the gate.
#
# Lives in its own job because the verdict it depends on cannot exist until
# after build-web has finished, and the promote used to run INSIDE build-web.
#
# `needs` on smoke-web is the gate. A failed smoke skips this job, so a
# refresh that broke something leaves :latest naming the build that works —
# "the refresh failed" and "production is broken" must not be the same event.
# A SKIPPED smoke also skips this job, which is the behaviour that matters
# most: on run 5290 the gate silently skipped itself, and a design where only
# a FAILED gate blocks would have published unverified images while reporting
# success. Not running is not the same as passing.
#
# All three images promote TOGETHER, or none do. They are one stack: build.yml
# already refuses to publish a :dev web image beside a stale :dev ml, because
# the mismatch only shows up as a runtime failure. A refresh that published ml
# and withheld web would be that same trap, arrived at through the gate.
#
# The gate covers the web image only (milestone 362 step 3 scoped it there),
# so ml and agent are being held to web's verdict rather than their own. That
# is deliberate and it is the conservative direction — they ship together, so
# the weakest evidence should govern all three — but it is not the same as
# having smoked them, and it should not be read as if it were.
promote:
needs: [build-web, build-ml, build-agent, smoke-web]
# Only a refresh publishes through a candidate; a push writes its channel
# tag directly from the build. Reads the same reuse-step decision the build
# took, via a job output — a job's `if:` cannot see the `env` context.
if: needs.build-web.outputs.candidate == 'true'
runs-on: python-ci
container:
image: git.fabledsword.com/bvandeusen/ci-python:3.14
steps:
- name: Point the channel tags at the smoked candidates
env:
TOKEN: ${{ secrets.RELEASE_TOKEN }}
ACTOR: ${{ github.actor }}
run: |
set -eu
# `latest` is not a guess: a refresh always builds `main` (BUILD_REF),
# and the "must have checked out main" guard in every build job fails
# the run if that did not hold. So the channel is main's.
TAG=latest
FAILED=""
for NAME in fabledcurator fabledcurator-ml fabledcurator-agent; do
REPO="bvandeusen/$NAME"
echo "promote: $REPO"
# Registry auth is its own token exchange — `docker login`
# authenticates the docker client, not curl. Deadline on every call
# (rule 156): a registry that stops answering must fail this step,
# not hang the weekly refresh until the job times out.
BEARER=$(curl -fsS --max-time 30 -u "$ACTOR:$TOKEN" \
"https://git.fabledsword.com/v2/token?scope=repository:$REPO:pull,push&service=git.fabledsword.com" \
| python3 -c 'import sys,json; print(json.load(sys.stdin)["token"])')
# Ask for the IMAGE manifest media types only. Offering the index
# types too would let the registry hand back an index if one ever
# existed at this tag, and we would faithfully copy the thing this
# whole approach exists to avoid creating.
ACCEPT='application/vnd.oci.image.manifest.v1+json, application/vnd.docker.distribution.manifest.v2+json'
CT=$(curl -fsS --max-time 60 -o manifest.json -D headers.txt \
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
"https://git.fabledsword.com/v2/$REPO/manifests/refresh-candidate" \
&& tr -d '\r' < headers.txt | awk -F': ' '/^[Cc]ontent-[Tt]ype:/{print $2}')
test -n "$CT"
SRC=$(tr -d '\r' < headers.txt | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
echo "promote: candidate $SRC ($CT)"
# NOT `imagetools create`. That wraps its source in an INDEX, and
# `.Image.Config.Labels` does not resolve through one — the
# fc.revision the reuse check reads off the channel tag would come
# back empty, every later push would miss and rebuild, and nothing
# would go red (#3183, run 4751). A manifest PUT is what "make this
# tag name that image" means at the registry: same bytes, same media
# type, same digest, no layer transfer.
curl -fsS --max-time 120 -X PUT \
-H "Authorization: Bearer $BEARER" -H "Content-Type: $CT" \
--data-binary @manifest.json \
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG"
# Read it back. A PUT that returned 2xx but landed something else is
# exactly the silent-and-plausible failure this pipeline keeps
# producing, and the check costs one request.
NOW=$(curl -fsS --max-time 30 -o /dev/null -D - \
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG" \
| tr -d '\r' | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
if [ "$NOW" != "$SRC" ]; then
echo "promote: FAILED — $NAME:$TAG is $NOW, expected $SRC" >&2
FAILED="$FAILED $NAME"
continue
fi
echo "promote: $NAME:$TAG now names $NOW"
done
if [ -n "$FAILED" ]; then
echo "" >&2
echo "promote: FAILED for:$FAILED" >&2
echo "promote: the channel tags are now INCONSISTENT — some images" >&2
echo "promote: moved and some did not. Re-run this refresh; the" >&2
echo "promote: candidates are still published and the promote is" >&2
echo "promote: idempotent." >&2
exit 1
fi
echo "promote: all three channel tags moved"
build-ml:
runs-on: python-ci
container:
@@ -957,7 +1346,7 @@ jobs:
# See sign-extension's copy for why this guard exists.
- name: Guard — a scheduled run must have checked out main
if: github.event_name == 'schedule'
if: env.IS_REFRESH == 'true'
run: |
set -eu
BRANCH=$(git rev-parse --abbrev-ref HEAD)
@@ -990,8 +1379,18 @@ jobs:
# the PREVIOUS XPI while the freshly signed one is orphaned (#3156).
# * dev and main derive the same values for the same source.
- name: Report the derived artifact version
env:
# Diagnostic for the trigger normalisation. `refresh` is reported RAW
# as well as normalised, because the two disagreeing is the whole
# failure mode: a dispatch input whose type does not compare the way
# the expression assumes evaluates to false silently, and the only
# symptom is a refresh that quietly behaves like an ordinary push.
RAW_REFRESH: ${{ github.event.inputs.refresh }}
RAW_FORCE: ${{ github.event.inputs.force_build }}
run: |
set -u
echo "trigger: event=$GITHUB_EVENT_NAME IS_REFRESH='${IS_REFRESH:-<unset>}' BUILD_REF='${BUILD_REF:-<unset>}'"
echo "trigger: raw inputs refresh='${RAW_REFRESH:-<unset>}' force_build='${RAW_FORCE:-<unset>}'"
A=ml
V=$(sh scripts/artifacts.sh version "$A" 2>&1 || echo UNAVAILABLE)
R=$(sh scripts/artifacts.sh revision "$A" 2>&1 || echo UNAVAILABLE)
@@ -1008,7 +1407,7 @@ jobs:
SHORT_SHA=$(printf '%s' "$GITHUB_SHA" | cut -c1-7)
# Mirrors build-web's tag list and its schedule handling; see
# the comments there.
if [ "${GITHUB_EVENT_NAME:-}" = "schedule" ]; then
if [ "${IS_REFRESH:-}" = "true" ]; then
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator-ml:latest" >> "$GITHUB_OUTPUT"
echo "channel=main" >> "$GITHUB_OUTPUT"
elif [ "${GITHUB_REF##*/}" = "main" ]; then
@@ -1091,17 +1490,63 @@ jobs:
# A scheduled refresh has to bypass reuse by construction: it
# rebuilds the SAME source, so fc.revision always matches and the
# check would skip every refresh there has ever been.
EVENT: ${{ github.event_name }}
run: |
set -eu
DERIVED=$(sh scripts/artifacts.sh revision ml)
echo "revision=$DERIVED" >> "$GITHUB_OUTPUT"
# The build clock, pinned to the same commit (#3265). Without it
# buildkit stamps the image config with the wall clock of the build,
# so identical layers republish under a new config blob and the
# channel tag gets a new manifest digest for no reason. Derived from
# `newest()` like revision and version, so all three name one commit
# and cannot drift apart.
echo "epoch=$(sh scripts/artifacts.sh epoch ml)" >> "$GITHUB_OUTPUT"
# The moving tag for this channel. Which tag we ask IS the channel —
# that is why the revision needs no -main/-dev qualifier any more.
if [ "$CHANNEL" = "main" ]; then T=latest; else T=dev; fi
echo "channel_ref=$IMAGE:$T" >> "$GITHUB_OUTPUT"
# WHERE THE BUILD PUBLISHES, which is not always the channel — and
# whether the channel then has to be written separately.
#
# On a push the build writes the channel tag directly: the bytes came
# from a commit, and a commit is the thing CI tests. Nothing to hold
# it behind.
#
# On the scheduled refresh it writes a CANDIDATE tag instead. A
# refresh rebuilds against freshly resolved base images, and the web
# image's runtime is a line of UNPINNED Debian packages (ffmpeg,
# libjpeg62-turbo, libpq5, megatools…) re-resolved on every build.
# Nothing in ci.yml can see that: its lanes run on ci-python:3.14 and
# install requirements.txt, and a base bump changes neither. So
# refreshed bytes have to be proven before :latest names them, and
# proving needs a moment between "built" and "published" to occupy.
# This is that moment; :latest goes on naming the build that works
# until something says otherwise.
#
# `:refresh-candidate` is one moving ref per image, overwritten in
# place, holding a build nobody is told to pull — the shape rule 145
# already allows for :buildcache, not the per-build tag family that
# milestone 318 withdrew.
#
# Decided HERE, beside `hit`, for the reason the force/schedule
# branch below gives: one step decides what this job does. A
# condition derived independently could disagree with the tag the
# build actually wrote.
#
# build-web additionally exposes this as `outputs.candidate`, which is
# what gates the `promote` job — a job's `if:` cannot read `env`, and
# one flag is enough because all three derive it from the same
# IS_REFRESH. ml and agent do not re-emit it; a second copy nothing
# reads is the kind of thing that later reads as load-bearing.
if [ "${IS_REFRESH:-}" = "true" ]; then
echo "build_ref=$IMAGE:refresh-candidate" >> "$GITHUB_OUTPUT"
else
echo "build_ref=$IMAGE:$T" >> "$GITHUB_OUTPUT"
fi
# Compare VALUES, never exit codes. Measured on buildx v0.36.1
# (run 4732): a missing key returns an empty string and exits 0, so
# branching on the exit code would read "no label yet" as success.
@@ -1134,7 +1579,7 @@ jobs:
if [ "${FORCE:-false}" = "true" ]; then
echo "hit=false" >> "$GITHUB_OUTPUT"
echo "reuse: force_build set — building regardless"
elif [ "${EVENT:-}" = "schedule" ]; then
elif [ "${IS_REFRESH:-}" = "true" ]; then
echo "hit=false" >> "$GITHUB_OUTPUT"
echo "reuse: scheduled base refresh — building regardless"
elif [ -n "$PUBLISHED" ] && [ "$PUBLISHED" = "$DERIVED" ]; then
@@ -1147,6 +1592,12 @@ jobs:
- name: Build and push ml image
if: steps.reuse.outputs.hit != 'true'
# Read by buildx out of the ENVIRONMENT, not passed as a build-arg —
# it normalises the image config's `created` field and the history
# timestamps rather than being consumed by the Dockerfile. See #3265
# and the reuse step's `epoch` output.
env:
SOURCE_DATE_EPOCH: ${{ steps.reuse.outputs.epoch }}
uses: docker/build-push-action@v5
with:
context: .
@@ -1159,20 +1610,17 @@ jobs:
# invalidates, and the image genuinely rebuilds.
#
# MEASURED on the first real fire, run 4934 (#3265): when the base
# did NOT move, the build is ~13s and every content step reports
# CACHED — but the channel tag STILL gets a new manifest digest.
# buildkit mints a fresh image config each run, so identical layers
# are republished under a new config blob. All three images moved
# that way on 2026-08-30 with nothing whatsoever changed in them.
# did NOT move, the build was ~13s with every content step CACHED —
# and the channel tag STILL got a new manifest digest, because
# buildkit stamps a fresh image config per run and republishes the
# identical layers under it. All three images moved that way on
# 2026-08-30 with nothing whatsoever changed in them.
#
# So a refresh currently rewrites :latest every Sunday whether or
# not there is anything new in it, and :c-<sha> is handed a new
# manifest to diverge from on the same cadence. Layers are shared,
# so the storage cost is a config blob; the cost that matters is
# that a digest change no longer MEANS anything. Tracked in #3265 —
# the likely fix is a deterministic SOURCE_DATE_EPOCH, which would
# make "same source, same bytes" true and turn the no-op case into
# a genuine no-op.
# SOURCE_DATE_EPOCH (below) is the fix: pinned to the commit the
# content came from, the config is byte-identical across runs, so
# the manifest digest is too and the push is a registry no-op. A
# digest change means the content changed again, which is the only
# thing a digest is any use for.
#
# What `pull` does NOT catch either: a Debian package update inside
# the `apt-get install` layer while the base tag itself stands
@@ -1182,14 +1630,14 @@ jobs:
# churn #3265 is about.
#
# Only on the schedule. An ordinary push wants the cached base.
pull: ${{ github.event_name == 'schedule' }}
pull: ${{ env.IS_REFRESH == 'true' }}
# ONE tag, the channel's. Every other tag is written by the step
# below, registry-side. buildx here pushes the first tag to the
# registry and then re-pushes the rest through the DOCKER driver,
# out of a local image store a registry-direct build never filled —
# #3190, which cost `main` its :c-<sha> on 2026-08-29 while :latest
# published perfectly well.
tags: ${{ steps.reuse.outputs.channel_ref }}
tags: ${{ steps.reuse.outputs.build_ref }}
# The reuse key. Read back off the channel tag on the next push to
# decide whether that push needs to build at all, so this is not
# decoration — an unstamped image is one that will always rebuild.
@@ -1331,7 +1779,7 @@ jobs:
# See sign-extension's copy for why this guard exists.
- name: Guard — a scheduled run must have checked out main
if: github.event_name == 'schedule'
if: env.IS_REFRESH == 'true'
run: |
set -eu
BRANCH=$(git rev-parse --abbrev-ref HEAD)
@@ -1364,8 +1812,18 @@ jobs:
# the PREVIOUS XPI while the freshly signed one is orphaned (#3156).
# * dev and main derive the same values for the same source.
- name: Report the derived artifact version
env:
# Diagnostic for the trigger normalisation. `refresh` is reported RAW
# as well as normalised, because the two disagreeing is the whole
# failure mode: a dispatch input whose type does not compare the way
# the expression assumes evaluates to false silently, and the only
# symptom is a refresh that quietly behaves like an ordinary push.
RAW_REFRESH: ${{ github.event.inputs.refresh }}
RAW_FORCE: ${{ github.event.inputs.force_build }}
run: |
set -u
echo "trigger: event=$GITHUB_EVENT_NAME IS_REFRESH='${IS_REFRESH:-<unset>}' BUILD_REF='${BUILD_REF:-<unset>}'"
echo "trigger: raw inputs refresh='${RAW_REFRESH:-<unset>}' force_build='${RAW_FORCE:-<unset>}'"
A=agent
V=$(sh scripts/artifacts.sh version "$A" 2>&1 || echo UNAVAILABLE)
R=$(sh scripts/artifacts.sh revision "$A" 2>&1 || echo UNAVAILABLE)
@@ -1377,7 +1835,7 @@ jobs:
SHORT_SHA=$(printf '%s' "$GITHUB_SHA" | cut -c1-7)
# Mirrors build-web's tag list and its schedule handling; see
# the comments there.
if [ "${GITHUB_EVENT_NAME:-}" = "schedule" ]; then
if [ "${IS_REFRESH:-}" = "true" ]; then
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator-agent:latest" >> "$GITHUB_OUTPUT"
echo "channel=main" >> "$GITHUB_OUTPUT"
elif [ "${GITHUB_REF##*/}" = "main" ]; then
@@ -1460,17 +1918,63 @@ jobs:
# A scheduled refresh has to bypass reuse by construction: it
# rebuilds the SAME source, so fc.revision always matches and the
# check would skip every refresh there has ever been.
EVENT: ${{ github.event_name }}
run: |
set -eu
DERIVED=$(sh scripts/artifacts.sh revision agent)
echo "revision=$DERIVED" >> "$GITHUB_OUTPUT"
# The build clock, pinned to the same commit (#3265). Without it
# buildkit stamps the image config with the wall clock of the build,
# so identical layers republish under a new config blob and the
# channel tag gets a new manifest digest for no reason. Derived from
# `newest()` like revision and version, so all three name one commit
# and cannot drift apart.
echo "epoch=$(sh scripts/artifacts.sh epoch agent)" >> "$GITHUB_OUTPUT"
# The moving tag for this channel. Which tag we ask IS the channel —
# that is why the revision needs no -main/-dev qualifier any more.
if [ "$CHANNEL" = "main" ]; then T=latest; else T=dev; fi
echo "channel_ref=$IMAGE:$T" >> "$GITHUB_OUTPUT"
# WHERE THE BUILD PUBLISHES, which is not always the channel — and
# whether the channel then has to be written separately.
#
# On a push the build writes the channel tag directly: the bytes came
# from a commit, and a commit is the thing CI tests. Nothing to hold
# it behind.
#
# On the scheduled refresh it writes a CANDIDATE tag instead. A
# refresh rebuilds against freshly resolved base images, and the web
# image's runtime is a line of UNPINNED Debian packages (ffmpeg,
# libjpeg62-turbo, libpq5, megatools…) re-resolved on every build.
# Nothing in ci.yml can see that: its lanes run on ci-python:3.14 and
# install requirements.txt, and a base bump changes neither. So
# refreshed bytes have to be proven before :latest names them, and
# proving needs a moment between "built" and "published" to occupy.
# This is that moment; :latest goes on naming the build that works
# until something says otherwise.
#
# `:refresh-candidate` is one moving ref per image, overwritten in
# place, holding a build nobody is told to pull — the shape rule 145
# already allows for :buildcache, not the per-build tag family that
# milestone 318 withdrew.
#
# Decided HERE, beside `hit`, for the reason the force/schedule
# branch below gives: one step decides what this job does. A
# condition derived independently could disagree with the tag the
# build actually wrote.
#
# build-web additionally exposes this as `outputs.candidate`, which is
# what gates the `promote` job — a job's `if:` cannot read `env`, and
# one flag is enough because all three derive it from the same
# IS_REFRESH. ml and agent do not re-emit it; a second copy nothing
# reads is the kind of thing that later reads as load-bearing.
if [ "${IS_REFRESH:-}" = "true" ]; then
echo "build_ref=$IMAGE:refresh-candidate" >> "$GITHUB_OUTPUT"
else
echo "build_ref=$IMAGE:$T" >> "$GITHUB_OUTPUT"
fi
# Compare VALUES, never exit codes. Measured on buildx v0.36.1
# (run 4732): a missing key returns an empty string and exits 0, so
# branching on the exit code would read "no label yet" as success.
@@ -1503,7 +2007,7 @@ jobs:
if [ "${FORCE:-false}" = "true" ]; then
echo "hit=false" >> "$GITHUB_OUTPUT"
echo "reuse: force_build set — building regardless"
elif [ "${EVENT:-}" = "schedule" ]; then
elif [ "${IS_REFRESH:-}" = "true" ]; then
echo "hit=false" >> "$GITHUB_OUTPUT"
echo "reuse: scheduled base refresh — building regardless"
elif [ -n "$PUBLISHED" ] && [ "$PUBLISHED" = "$DERIVED" ]; then
@@ -1516,6 +2020,12 @@ jobs:
- name: Build and push agent image
if: steps.reuse.outputs.hit != 'true'
# Read by buildx out of the ENVIRONMENT, not passed as a build-arg —
# it normalises the image config's `created` field and the history
# timestamps rather than being consumed by the Dockerfile. See #3265
# and the reuse step's `epoch` output.
env:
SOURCE_DATE_EPOCH: ${{ steps.reuse.outputs.epoch }}
uses: docker/build-push-action@v5
with:
context: agent
@@ -1528,20 +2038,17 @@ jobs:
# invalidates, and the image genuinely rebuilds.
#
# MEASURED on the first real fire, run 4934 (#3265): when the base
# did NOT move, the build is ~13s and every content step reports
# CACHED — but the channel tag STILL gets a new manifest digest.
# buildkit mints a fresh image config each run, so identical layers
# are republished under a new config blob. All three images moved
# that way on 2026-08-30 with nothing whatsoever changed in them.
# did NOT move, the build was ~13s with every content step CACHED —
# and the channel tag STILL got a new manifest digest, because
# buildkit stamps a fresh image config per run and republishes the
# identical layers under it. All three images moved that way on
# 2026-08-30 with nothing whatsoever changed in them.
#
# So a refresh currently rewrites :latest every Sunday whether or
# not there is anything new in it, and :c-<sha> is handed a new
# manifest to diverge from on the same cadence. Layers are shared,
# so the storage cost is a config blob; the cost that matters is
# that a digest change no longer MEANS anything. Tracked in #3265 —
# the likely fix is a deterministic SOURCE_DATE_EPOCH, which would
# make "same source, same bytes" true and turn the no-op case into
# a genuine no-op.
# SOURCE_DATE_EPOCH (below) is the fix: pinned to the commit the
# content came from, the config is byte-identical across runs, so
# the manifest digest is too and the push is a registry no-op. A
# digest change means the content changed again, which is the only
# thing a digest is any use for.
#
# What `pull` does NOT catch either: a Debian package update inside
# the `apt-get install` layer while the base tag itself stands
@@ -1551,14 +2058,14 @@ jobs:
# churn #3265 is about.
#
# Only on the schedule. An ordinary push wants the cached base.
pull: ${{ github.event_name == 'schedule' }}
pull: ${{ env.IS_REFRESH == 'true' }}
# ONE tag, the channel's. Every other tag is written by the step
# below, registry-side. buildx here pushes the first tag to the
# registry and then re-pushes the rest through the DOCKER driver,
# out of a local image store a registry-direct build never filled —
# #3190, which cost `main` its :c-<sha> on 2026-08-29 while :latest
# published perfectly well.
tags: ${{ steps.reuse.outputs.channel_ref }}
tags: ${{ steps.reuse.outputs.build_ref }}
# The reuse key. Read back off the channel tag on the next push to
# decide whether that push needs to build at all, so this is not
# decoration — an unstamped image is one that will always rebuild.
+18 -1
View File
@@ -255,10 +255,27 @@ jobs:
export DB_HOST="$PG_IP"
export CELERY_BROKER_URL="redis://$RD_IP:6379/0"
export CELERY_RESULT_BACKEND="redis://$RD_IP:6379/0"
# These steps run under `sh -e`, not bash, so bash's /dev/tcp magic
# path does not exist here — the probe this loop used to run could
# never succeed and simply burned the full 120s on every run, green
# or red, then continued without having established anything. Python
# is in the image and needs no installed package for a socket
# connect, so it is the probe. Exhausting the budget is now a named
# failure rather than a silent fall-through (rule 156): if Postgres
# is genuinely not up, that is what the log should say, instead of
# whatever the first query happens to raise two minutes later.
pg_ready=""
for i in $(seq 1 60); do
(echo > "/dev/tcp/$PG_IP/5432") >/dev/null 2>&1 && break
if python -c "import socket,sys; s=socket.socket(); s.settimeout(2); sys.exit(0 if s.connect_ex(('$PG_IP', 5432)) == 0 else 1)"; then
pg_ready=1
break
fi
sleep 2
done
if [ -z "$pg_ready" ]; then
echo "postgres at $PG_IP:5432 did not accept a connection within 120s"
exit 1
fi
if command -v uv >/dev/null 2>&1; then
uv pip install --system -r requirements.txt pytest pytest-asyncio
else
+19
View File
@@ -70,3 +70,22 @@ alembic/versions/__pycache__/
*.sqlite
*.sqlite-journal
.superpowers/
# Raw platform captures (milestone 387 C0 and successors). These are real
# authenticated API responses taken from the operator's own account, so they
# carry account data — creator lists, pledge amounts, and (in Patreon's case)
# the account email inside the `card` resources. They are kept locally because
# re-capturing means re-authenticating by hand, and they are the ground truth a
# characterization gets re-checked against.
#
# The whole directory is ignored, not one filename, so a future capture is
# covered by this rule instead of needing a new line somebody has to remember.
#
# SANITIZED fixtures derived from these DO belong in git — put them somewhere
# else (tests/fixtures/, not here), with the account data stripped.
# Ignore the CONTENTS, not the directory: git does not descend into an
# excluded directory, so a negation for a file inside one never takes effect.
# Writing it this way lets README.md be committed while everything else here
# stays out.
tests/fixtures/captures/*
!tests/fixtures/captures/README.md
+235 -45
View File
@@ -1,14 +1,233 @@
<img src="frontend/public/logo.svg" alt="" width="132" align="right" />
# FabledCurator
Self-hosted media curation — gallery, ML tagging, and subscription-driven downloading in one app. Part of the FabledSword family.
<!-- overview:start -->
Self-hosted media curation — a gallery, ML auto-tagging, and subscription-driven
downloading in one application. Part of the FabledSword family.
Combines what was [ImageRepo](https://git.fabledsword.com/bvandeusen/ImageRepo) (gallery, ML, importer) and [GallerySubscriber](https://git.fabledsword.com/bvandeusen/GallerySubscriber) (gallery-dl wrapper, subscriptions, credential capture) into a single product.
## What it does
## Status
You point it at creators you follow. It downloads what they post, files it,
tags it, and gives you something better than a folder full of images to look
through afterwards.
In production. `main` is continuously deployed — every merge to `main` builds
and publishes `:latest` images, so whatever is on `main` is what is running.
Day-to-day work happens on `dev`, which publishes `:dev` images.
- **Gallery and browsing.** Images, videos and multi-page works, organised by
artist, tag, post and series. A newest-first feed of what just arrived as the
front page, a random Showcase, a filterable gallery, a similarity-driven
Explore view, and a page-turning reader for series.
- **Subscriptions.** Follows creators on Patreon, SubscribeStar, Discord and
HentaiFoundry, on a schedule. Handles paywalled posts using your own
logged-in session.
- **ML tagging.** Runs image models in-container to suggest tags, group
characters, find near-duplicates and power similarity search. Suggestions are
reviewable — it proposes, you confirm, and it learns which proposals you keep
rejecting.
- **Deduplication and provenance.** Everything that arrives is hashed and
deduplicated by content, metadata sidecars are read wherever the source
writes them, and every file keeps a record of where it came from.
- **Maintenance.** Backups, library audits, thumbnail and embedding backfills,
orphan cleanup — all from the UI, all as background jobs you can watch.
Everything is configured from the Settings UI and stored in the database. There
is no config file to edit beyond a handful of bootstrap environment variables.
<!-- overview:end -->
## Before you expose it
**FabledCurator has no login.** There are no user accounts, no passwords and no
permission model. Anything that can reach the port is an administrator.
That matters more here than it would in most self-hosted apps, because of what
this one stores: **live platform session cookies for Patreon and
SubscribeStar** — accounts that usually have a payment method attached. Whoever reaches
the port can read them, alongside your entire library.
So:
- Bind it to a LAN, a VPN, or a tunnel you control.
- Do not port-forward it. Do not put it on a public hostname.
- A reverse proxy that adds TLS but no authentication **does not help**. If you
want it reachable from outside, put an authenticating proxy in front of it —
a forward-auth provider, HTTP basic auth, an identity-aware tunnel — and treat
that layer as the only thing standing between the internet and your accounts.
This is a deliberate design decision for a single-operator tool on a trusted
network, not a bug and not an oversight. It is stated here because it decides
how you are allowed to deploy it. [SECURITY.md](SECURITY.md) covers the rest of
the threat model.
## Requirements
- **Docker** with Compose v2.
- **~4 GB RAM** for the app, plus whatever Postgres needs for your library size.
- **Disk** for your media, plus several GB for ML model weights.
- **No GPU required.** The ML worker runs on CPU — tagging and embedding are
slower, and that is the whole difference. A GPU is only involved if you
separately run the optional agent (below), which is a different machine's job.
## Install
```bash
git clone https://git.fabledsword.com/bvandeusen/FabledCurator.git
cd FabledCurator
cp .env.example .env
$EDITOR .env # set DB_PASSWORD and SECRET_KEY
docker compose -f docker-compose.yml up -d
```
Then open <http://localhost:8080>.
**The `-f docker-compose.yml` is required, not decoration.** Compose
auto-merges `docker-compose.override.yml` when you leave it off, and that
override builds the images locally from source — the contributor path, not
yours. Naming the file explicitly skips the override and pulls the published
`:latest` images, which is the stable channel built from `main`.
If you forget it, the symptom is a long build instead of a quick pull.
## First run
The database schema is created automatically on first start — the web container
runs its migrations before serving. Nothing to initialise by hand.
**One thing does need a deliberate act, and the app will not start without it.**
FabledCurator encrypts your stored platform credentials with a key it keeps at
`./images/secrets/credential_key.b64`. On a brand-new install that file does not
exist, and rather than quietly creating one the app stops:
```
MissingCredentialKey: Fernet key file not found at /images/secrets/credential_key.b64
```
Set `CURATOR_BOOTSTRAP_NEW_KEY=1` in your `.env` for the first `up`, then delete
the line once the container is running. `.env.example` ships it with that
instruction attached.
The refusal is deliberate, and worth understanding rather than working around:
auto-creating a key is indistinguishable from the disaster case — a restore that
brought the database back but lost `./images/secrets` — where it would mint a key
that cannot decrypt anything, leaving an instance that looks healthy while every
paywalled download fails. Making you say so once, on an empty install, is the
price of that not happening silently later.
**Which means: back up `./images/secrets/` alongside your database.** It is the
only thing that can read your stored credentials. A database restored without it
needs every credential entered again by hand.
A few other things are worth knowing about the first few minutes:
- **The ML worker downloads its model weights on first boot**, several GB from
HuggingFace into `./models`. Until that finishes, tagging is queued rather
than broken. It is idempotent — a restart resumes rather than refetches.
- **The gallery starts empty**, and that is the expected state. Add a creator
under **Subscriptions** and it fills as posts come down.
- **If you already have a library on disk**, there is no screen that imports
it, and there is not going to be one. Folder ingestion had a UI until July
2026; it was retired once posts began arriving entirely through
subscriptions and the browser extension, and the decision to leave it
retired is deliberate — the folder path carries complexity the product does
not need in order to do its job. The supported way to fill a new install is
to add the creators you follow under **Subscriptions** and let it pull.
The `/api/import/trigger` endpoint is still wired up for anyone who wants to
script a one-off against a folder mounted at `./import`, and its progress
shows under **Settings → Activity**. Treat it as an unsupported escape
hatch rather than a feature: nothing in the UI drives it and nothing else
in this README depends on it.
- **To download from a paywalled account**, FabledCurator needs that account's
session — see the browser extension below. Without one it can still fetch
public posts.
- **Check Settings → Overview** to confirm the workers are alive. Every long
operation in FabledCurator is a background job, so if the queues are not
running, the UI will look like it is ignoring you rather than like it is
broken.
## The browser extension
A Firefox extension does two jobs: it hands your logged-in platform sessions to
FabledCurator so it can download on your behalf, and it adds a creator as a
subscription in one click from their page.
It ships **inside the web image** — there is no add-on store listing to find.
Go to **Subscriptions → Settings**, find the *Browser extension* card, and click
**Install Firefox extension**. The XPI is Mozilla-signed, so Firefox installs it
like any other add-on; the button serves it directly rather than making you
download and side-load a file.
It pairs with your instance using an API key generated automatically on first
use. The bar directly under that card shows the key and can rotate it.
See [extension/README.md](extension/README.md) for what it does in detail.
## The GPU agent
Optional, and separate. If you have a desktop with a graphics card, you can run
an agent on it that leases ML jobs from FabledCurator over HTTP, does them on
the GPU, and hands the results back. It never touches the database or Redis, so
it is safe to run somewhere the rest of the stack is not.
Run it for a burst of tagging, stop it to get your card back. It deploys from
`agent/docker-compose.yml`, not the main stack — see
[agent/README.md](agent/README.md).
## Upgrading
```bash
docker compose -f docker-compose.yml pull
docker compose -f docker-compose.yml up -d
```
Migrations run automatically on start. Take a database backup first — Settings →
Maintenance has one — because the schema moves forward and does not move back.
## Deployment posture
FabledCurator is built to run inside a homelab over plain HTTP. It does not
generate certificates, redirect to HTTPS, or set HSTS. If you want TLS,
terminate it at your reverse proxy. See [Before you expose it](#before-you-expose-it)
for why TLS alone is not enough.
## Troubleshooting
**The UI loads but nothing ever finishes.** The web container is up and the
workers are not. `docker compose -f docker-compose.yml ps` — check `worker`,
`scheduler` and `ml-worker` are healthy, not restarting.
**`docker compose up` started building instead of pulling.** You left off
`-f docker-compose.yml`, so the dev override took over. See [Install](#install).
**Downloads fail with an auth error.** The stored session for that platform has
expired. Re-capture it with the extension; sessions do not last forever.
**Which build am I running?** The foot of Settings shows a version and a
channel, and `/api/health` returns the same two fields. There are no version
tags on the images, so this is the authoritative answer.
---
# Developing FabledCurator
Everything below is about working on FabledCurator rather than running it. If
you are installing it, you are done — see [CONTRIBUTING.md](CONTRIBUTING.md) if
you want to send a patch.
## Status and channels
In production. `main` is continuously deployed — every merge builds and
publishes `:latest`, so whatever is on `main` is what is running. Day-to-day
work happens on `dev`, which publishes `:dev`.
For local development, the dev override handles everything:
```bash
docker compose up -d # note: no -f, so the override applies
```
That builds the images from source, turns on DEBUG logging, and exposes
Postgres and Redis on the host. No `.env` required.
## Versions and tags
@@ -47,48 +266,10 @@ Five deployable pieces, built by `.forgejo/workflows/build.yml`:
| --- | --- | --- | --- |
| **Web / workers** | `Dockerfile` | `fabledcurator` | Quart API + the built Vue SPA in one image. `entrypoint.sh` picks the role: `web`, `worker`, `scheduler`. The `maintenance-long` service is a second `worker` pinned to the long-running maintenance queue. |
| **ML worker** | `Dockerfile.ml` | `fabledcurator-ml` | Same app, plus `requirements-ml.txt` — tagging and embedding models that run in-container. |
| **GPU agent** | `agent/Dockerfile` | `fabledcurator-agent` | Optional desktop-GPU worker (`agent/`). Leases jobs over **HTTP only** — never touches the database or Redis. Run it for a burst, stop it to reclaim the card. See `agent/README.md`. |
| **GPU agent** | `agent/Dockerfile` | `fabledcurator-agent` | Optional desktop-GPU worker (`agent/`). Leases jobs over **HTTP only** — never touches the database or Redis. See `agent/README.md`. |
| **Firefox extension** | `extension/` | signed XPI | MV3 extension: pushes platform session cookies into FC and adds a creator as a Source in one click. AMO-signed on both `dev` and `main` (one signature per extension change, shared by the two channels), bundled into that channel's web image and served from Settings → Maintenance. See `extension/README.md`. |
| **Data** | — | `pgvector/pgvector:pg16`, `redis:7-alpine` | Postgres with pgvector for embeddings; Redis as the Celery broker. |
## Quick start
For local development and testing, just:
```bash
docker compose up -d
# UI: http://localhost:8080
```
That uses sane dev defaults baked into `docker-compose.yml` and the dev
override (`docker-compose.override.yml`, auto-merged) — local builds, DEBUG
logging, exposed Postgres + Redis ports on the host. No `.env` required.
For a production-like deployment, override the dev defaults via shell env
or a `.env` file (see `.env.example` for the variable names) and use:
```bash
docker compose -f docker-compose.yml up -d
# (skips the dev override, so containers pull published :latest images)
```
`-f` is doing real work there: it tells Compose to use *only* that file, which
skips `docker-compose.override.yml` and its local builds. What you get is the
`:latest` images — the stable channel, built from `main`. This is the install
path, and it is the one to use if you are running FabledCurator rather than
working on it.
`:dev` is the other channel: rebuilt from the `dev` branch several times a day,
bleeding edge, no stability promise. Nothing in this repo points an installer at
it, and nothing should.
The GPU agent is deployed separately, on the machine with the card —
`agent/docker-compose.yml`, not this stack.
## Deployment posture
FabledCurator is designed to run inside a self-hosted homelab environment over plain HTTP. If you want TLS, terminate it at your reverse proxy. The app does not generate certificates, redirect to HTTPS, or set HSTS.
## CI / Forgejo setup
Four workflows: `ci.yml` (lint, extension-version check, backend unit tests,
@@ -118,6 +299,15 @@ source, so `main` finds `dev`'s signature already cached and makes no second AMO
call. That cache is why signing must be one-shot — AMO rejects a re-signed
version.
## History
FabledCurator combines what was
[ImageRepo](https://git.fabledsword.com/bvandeusen/ImageRepo) (gallery, ML,
importer) and
[GallerySubscriber](https://git.fabledsword.com/bvandeusen/GallerySubscriber)
(gallery-dl wrapper, subscriptions, credential capture) into a single product.
Both are superseded; neither is maintained.
## License
**GNU Affero General Public License v3.0** — see [LICENSE](LICENSE).
+29 -15
View File
@@ -26,35 +26,49 @@ FabledCurator is self-hosted and holds things worth stating plainly, because
they shape what counts as a serious bug here:
- **Platform credentials.** The app captures and stores session cookies for
third-party subscription sites (Patreon, SubscribeStar, Pixiv) so it can
third-party subscription sites (Patreon, SubscribeStar) so it can
download on the operator's behalf. These are live credentials for accounts
that usually carry a payment method. Anything that discloses them, decrypts
them, or lets one user of a shared instance read another's is high severity.
- **An extension API key.** The Firefox extension authenticates to the backend
with a shared key. Anything that leaks it or lets it be bypassed is a way in.
- **A multi-user sharing ACL.** Instances can be shared. A bug that lets one
account see content another has not shared is an access-control failure, not
a cosmetic one.
- **No authentication of its own.** This is the most important thing on this
page. FabledCurator has no login, no user accounts and no permission model —
there is no `User` table and no session auth anywhere in the backend. Every
HTTP client that can reach the port is the administrator, with full read and
write access to everything above, including the stored platform credentials.
Access control is entirely the operator's job, done at the network layer.
Reports that an unauthenticated caller can reach an endpoint are therefore
describing the design; reports that something *crosses the network boundary
the operator drew* — an SSRF, a request forgery that rides a browser the
operator already has open, a path that leaks state to an origin the operator
did not authorise — are in scope and are serious.
- **Arbitrary media from the internet.** Downloaded files are decoded, hashed,
thumbnailed and fed to ML models. Anything that turns a hostile file into
code execution is in scope.
## Deployment posture — read this before reporting
FabledCurator is designed to run **inside a private network, over plain HTTP**.
It does not terminate TLS, redirect to HTTPS, or set HSTS; if you want
transport security, terminate it at your reverse proxy. This is a documented
design decision, not an oversight.
FabledCurator is designed to run **inside a private network, over plain HTTP,
reachable only by its operator**. It does not terminate TLS, redirect to
HTTPS, or set HSTS; if you want transport security, terminate it at your
reverse proxy. It also does not authenticate anyone — see above. These are
documented design decisions, not oversights.
Reports that reduce to "the application is served over HTTP" or "there is no
HSTS header" describe that decision rather than a vulnerability. Reports that
an authenticated operator can cause the software to do something destructive
are usually also by design — the operator is the administrator of their own
instance.
Putting this on the public internet, with or without TLS, hands whoever finds
it your Patreon and SubscribeStar sessions. A reverse proxy that adds
TLS but not an authentication layer does not change that.
Reports that reduce to "the application is served over HTTP", "there is no
HSTS header", or "the API needs no credentials" describe those decisions
rather than vulnerabilities. Reports that the operator can cause the software
to do something destructive are usually also by design — the operator is the
administrator of their own instance.
What remains in scope is everything that crosses a boundary the software is
supposed to hold: between one user and another, between an unauthenticated
visitor and any of it, and between untrusted downloaded content and the host.
actually supposed to hold: between untrusted downloaded content and the host,
between a third-party origin and an operator's open browser session, and
between the credentials at rest and anything that is not the operator.
## Supported versions
+64
View File
@@ -0,0 +1,64 @@
"""service_seen — the learned roster that makes a stopped part observable.
Milestone 365. Nothing in FabledCurator knew what was SUPPOSED to be running:
`celery inspect` reports the workers that answer, so a dead worker was a
shorter list rather than a red light, and the only surface that could tell an
operator otherwise was Portainer. This table is the memory that turns an
absence into something the app can see.
Keyed on the queue set for a celery role and on agent_id for the GPU agent —
NOT on the celery worker name, which here is `celery@<container id>` and is
minted fresh on every deploy. See the model docstring for why that choice is
the whole design.
## First migration on the collapsed baseline
0089 is the single generated baseline that replaced revisions 0001..0089
(milestone 328). This is the first revision written on top of it, so it is
also the first evidence that the chain steps forward from the collapse rather
than merely reproducing the schema — which nothing had demonstrated yet.
An existing install is at 0089 because it ran the real 0089; a fresh one is at
0089 because it ran the baseline. Both arrive here identically, which was the
property the collapse was designed around.
Revision ID: 0090
Revises: 0089
Create Date: 2026-09-02
"""
from typing import Sequence, Union
import sqlalchemy as sa
from alembic import op
revision: str = "0090"
down_revision: Union[str, None] = "0089"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
op.create_table(
"service_seen",
sa.Column("key", sa.String(length=128), nullable=False),
sa.Column("kind", sa.String(length=16), nullable=False),
sa.Column("display_name", sa.String(length=64), nullable=False),
sa.Column(
"first_seen_at", sa.DateTime(timezone=True),
server_default=sa.text("now()"), nullable=False,
),
sa.Column(
"last_seen_at", sa.DateTime(timezone=True),
server_default=sa.text("now()"), nullable=False,
),
sa.Column("details", sa.JSON(), nullable=False),
sa.PrimaryKeyConstraint("key", name=op.f("pk_service_seen")),
)
# No secondary indexes, deliberately: one row per moving part means every
# read is a handful of rows and an index would be write cost buying
# nothing (#3301 removed seven of exactly that shape).
def downgrade() -> None:
op.drop_table("service_seen")
@@ -0,0 +1,81 @@
"""platform_membership — the learned roster of what the account actually pays for.
Milestone 387, phase C. FC knows which creators it was told to follow and
nothing about which ones the operator is subscribed to; this table is the
memory that makes the drift in both directions observable. See the model
docstring for why the roster is learned rather than looked up live, and why
`status` holds the platform's own word rather than a normalised FC value.
## Nothing populates this yet, on purpose
The sweep that fills it (C3) depends on a client seam (C2) that depends on
characterising Patreon's real membership response from a captured sample (C0),
which needs the operator's authenticated browser session. The table's SHAPE
does not wait on that: it is deliberately free-form where C0's findings would
otherwise dictate a column — `status` is an unconstrained String and `details`
keeps the raw payload — so no capture can invalidate what is created here.
An empty table is the correct intermediate state. It is not dead code: C5 reads
it to explain a tier-limited source, and C4 reads it to reconcile.
Revision ID: 0091
Revises: 0090
Create Date: 2026-09-10
"""
from typing import Sequence, Union
import sqlalchemy as sa
from alembic import op
revision: str = "0091"
down_revision: Union[str, None] = "0090"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
op.create_table(
"platform_membership",
sa.Column("id", sa.Integer(), nullable=False),
sa.Column("platform", sa.String(length=64), nullable=False),
# Text, not a bounded String: an opaque upstream identifier we do not
# mint, and guessing a ceiling for one is how a walk dies on a silent
# truncation.
sa.Column("external_campaign_id", sa.Text(), nullable=False),
sa.Column("display_name", sa.Text(), nullable=True),
sa.Column("url", sa.Text(), nullable=True),
# No CHECK, deliberately (rule 36 considered and declined): the
# vocabulary is each platform's own and is not ours to fix before C0
# has characterised even one of them. The service owns the whitelist.
sa.Column("status", sa.String(length=32), nullable=True),
sa.Column("tier_names", sa.JSON(), nullable=True),
sa.Column("amount_cents", sa.Integer(), nullable=True),
sa.Column("currency", sa.String(length=8), nullable=True),
sa.Column(
"first_seen_at", sa.DateTime(timezone=True),
server_default=sa.text("now()"), nullable=False,
),
sa.Column(
"last_seen_at", sa.DateTime(timezone=True),
server_default=sa.text("now()"), nullable=False,
),
sa.Column("details", sa.JSON(), nullable=False),
sa.PrimaryKeyConstraint("id", name=op.f("pk_platform_membership")),
# The upsert's conflict target. Named explicitly because
# touch_membership references it by name in ON CONFLICT — an
# autogenerated name would make that call break on a rename nobody
# connected to it.
sa.UniqueConstraint(
"platform", "external_campaign_id",
name="uq_platform_membership_platform_campaign",
),
)
# No secondary indexes. This table holds one row per subscription — tens,
# not millions — so every query against it is a short scan and an index
# would be write cost buying nothing (#3301 removed seven of that shape).
# The unique constraint above already backs the only lookup that matters.
def downgrade() -> None:
op.drop_table("platform_membership")
+92
View File
@@ -0,0 +1,92 @@
"""Synthetic posts — FC authors a post for content that arrived as chat.
Milestone 388, step E2. Discord is a delivery channel, not a publisher: one
message is not one post, and today every message becomes its own `post` row
competing with authored work for the same surface. This adds the three columns
that let FC group a creator's variant drop into a post it wrote itself, while
keeping that fact visible and the grouping reversible.
## Why a flag and a back-pointer rather than a separate table
A synthetic post has to BE a post — same row, same columns — or every existing
surface (feed, provenance, translation, attachments, series) would need a
second code path for it. `synthesized_by` marks the ones FC authored;
`absorbed_by_post_id` points a member message-post at the post that replaced
it in the feed. The members are not deleted: they remain the images' true
origin, and destroying them would make the grouping un-auditable at exactly
the moment somebody wants to check it.
Reversal is one DELETE. `absorbed_by_post_id` is ON DELETE SET NULL, so
removing a synthetic post releases its members and they return to the feed
unaided.
Revision ID: 0092
Revises: 0091
Create Date: 2026-09-10
"""
from typing import Sequence, Union
import sqlalchemy as sa
from alembic import op
revision: str = "0092"
down_revision: Union[str, None] = "0091"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
# No CHECK on synthesized_by (rule 36 considered and declined): there is one
# grouper today and a second would be a new VALUE, not a new invariant —
# matching source.error_type and service_seen.kind.
op.add_column("post", sa.Column("synthesized_by", sa.String(length=32), nullable=True))
op.add_column("post", sa.Column("synthesis_details", sa.JSON(), nullable=True))
op.add_column(
"post", sa.Column("absorbed_by_post_id", sa.Integer(), nullable=True),
)
op.create_index(
op.f("ix_post_absorbed_by_post_id"), "post", ["absorbed_by_post_id"],
)
# SET NULL, not CASCADE: deleting the synthetic post must RELEASE its
# members, never take them with it. The members are the real capture.
op.create_foreign_key(
"fk_post_absorbed_by_post_id_post", "post", "post",
["absorbed_by_post_id"], ["id"], ondelete="SET NULL",
)
# Grouping tunables. Every one of these is operator-facing (project rule
# 25) because the quality bar here is a judgement call no test can settle:
# too greedy merges distinct pieces, too shy leaves a drop scattered.
op.add_column(
"ml_settings",
sa.Column(
"discord_grouping_enabled", sa.Boolean(),
server_default="true", nullable=False,
),
)
op.add_column(
"ml_settings",
sa.Column(
"discord_group_max_distance", sa.Float(),
server_default=sa.text("0.10"), nullable=False,
),
)
op.add_column(
"ml_settings",
sa.Column(
"discord_group_window_minutes", sa.Float(),
server_default=sa.text("60"), nullable=False,
),
)
def downgrade() -> None:
op.drop_column("ml_settings", "discord_group_window_minutes")
op.drop_column("ml_settings", "discord_group_max_distance")
op.drop_column("ml_settings", "discord_grouping_enabled")
op.drop_constraint("fk_post_absorbed_by_post_id_post", "post", type_="foreignkey")
op.drop_index(op.f("ix_post_absorbed_by_post_id"), table_name="post")
op.drop_column("post", "absorbed_by_post_id")
op.drop_column("post", "synthesis_details")
op.drop_column("post", "synthesized_by")
+86
View File
@@ -0,0 +1,86 @@
"""An open grouping — a synthetic post that a later drop can still join.
Milestone 388, step E3. E2's synthetic post was sealed at creation: a creator
who added two more variants the next day started a second post. These two
columns let the group stay open and absorb the follow-up, without the post
either freezing or thrashing the feed.
## Why openness is derived rather than stored
There is no `closed_at` here on purpose. A group is open if it grew (or
started) within `ml_settings.discord_group_close_after_hours`, so openness is a
comparison rather than a state — which means lowering the setting closes old
groups and raising it reopens them, with nothing to repair either way. A stored
flag would need its own sweep to set it and its own repair path to ever change
the policy, for no gain.
## Why `resurfaced_at` is separate from `last_grew_at`
They answer different questions. `last_grew_at` is when the group last
absorbed something — it decides how long the group stays joinable and is what
the card shows. `resurfaced_at` is the FEED POSITION, advanced only when the
anti-thrash rule fires, so a group that gains one image a day updates in place
while a genuine second wave moves once. Folding them together would make every
addition a bump, which is the annoyance this step exists to avoid.
Both are NULL on every ordinary post, so the feed's sort key can COALESCE
through `resurfaced_at` without moving anything that is not a grouping.
Revision ID: 0093
Revises: 0092
Create Date: 2026-09-10
"""
from typing import Sequence, Union
import sqlalchemy as sa
from alembic import op
revision: str = "0093"
down_revision: Union[str, None] = "0092"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
op.add_column(
"post", sa.Column("last_grew_at", sa.DateTime(timezone=True), nullable=True),
)
op.add_column(
"post", sa.Column("resurfaced_at", sa.DateTime(timezone=True), nullable=True),
)
# No index on either. The feed already sorts on an unindexed
# COALESCE(post_date, downloaded_at) expression, so adding resurfaced_at to
# that COALESCE changes nothing about how the query plans — and inventing a
# functional index here would be guessing at the fix for a cost nobody has
# measured. Measuring it is step B2's job.
op.add_column(
"ml_settings",
sa.Column(
"discord_group_close_after_hours", sa.Float(),
server_default=sa.text("168"), nullable=False,
),
)
op.add_column(
"ml_settings",
sa.Column(
"discord_group_resurface_min_images", sa.Integer(),
server_default="2", nullable=False,
),
)
op.add_column(
"ml_settings",
sa.Column(
"discord_group_resurface_cooldown_hours", sa.Float(),
server_default=sa.text("24"), nullable=False,
),
)
def downgrade() -> None:
op.drop_column("ml_settings", "discord_group_resurface_cooldown_hours")
op.drop_column("ml_settings", "discord_group_resurface_min_images")
op.drop_column("ml_settings", "discord_group_close_after_hours")
op.drop_column("post", "resurfaced_at")
op.drop_column("post", "last_grew_at")
+124
View File
@@ -0,0 +1,124 @@
"""post_association — "this Patreon post announced that Discord drop".
Milestone 388, step E5.
Two of the operator's artists post a deliberately cropped fragment on Patreon
to signal that the real thing has landed in their Discord. This table holds the
proposed and accepted links between the announcement and the drop.
Directional and confirm-only. The pair is asymmetric (the teaser announces the
drop, not the reverse), the two posts are never merged (the creator published
twice, deliberately — flattening that hides the behaviour being modelled), and
nothing is linked until the operator accepts, following the FC-6.3 series
matcher. A wrongly-asserted association tells them two different pieces are
one, which is worse than no link at all.
Dismissed rows are KEPT. The row is what remembers the rejection, and
re-proposing a rejected pair on every scan is what makes a review queue get
ignored.
Revision ID: 0094
Revises: 0093
Create Date: 2026-09-10
"""
from typing import Sequence, Union
import sqlalchemy as sa
from alembic import op
revision: str = "0094"
down_revision: Union[str, None] = "0093"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
op.create_table(
"post_association",
sa.Column("id", sa.Integer(), nullable=False),
sa.Column("announcement_post_id", sa.Integer(), nullable=False),
sa.Column("payload_post_id", sa.Integer(), nullable=False),
sa.Column("score", sa.Float(), nullable=False),
sa.Column("signals", sa.JSON(), nullable=True),
# No CHECK on status (rule 36 considered and declined), matching
# series_suggestion.status — the same review-queue vocabulary, and the
# same check-existing-enums lesson.
sa.Column(
"status", sa.String(length=16), server_default="pending", nullable=False,
),
sa.Column(
"created_at", sa.DateTime(timezone=True),
server_default=sa.text("now()"), nullable=False,
),
sa.Column(
"updated_at", sa.DateTime(timezone=True),
server_default=sa.text("now()"), nullable=False,
),
sa.PrimaryKeyConstraint("id", name=op.f("pk_post_association")),
# CASCADE on both sides: an association to a post that no longer exists
# is not a fact worth keeping, and E3's reversal path (delete the
# grouping) must not leave a dangling proposal behind.
sa.ForeignKeyConstraint(
["announcement_post_id"], ["post.id"], ondelete="CASCADE",
name=op.f("fk_post_association_announcement_post_id_post"),
),
sa.ForeignKeyConstraint(
["payload_post_id"], ["post.id"], ondelete="CASCADE",
name=op.f("fk_post_association_payload_post_id_post"),
),
sa.UniqueConstraint(
"announcement_post_id", "payload_post_id",
name="uq_post_association_pair",
),
)
op.create_index(
op.f("ix_post_association_announcement_post_id"),
"post_association", ["announcement_post_id"],
)
op.create_index(
op.f("ix_post_association_payload_post_id"),
"post_association", ["payload_post_id"],
)
op.create_index(
op.f("ix_post_association_status"), "post_association", ["status"],
)
op.add_column(
"import_settings",
sa.Column(
"discord_link_enabled", sa.Boolean(),
server_default="true", nullable=False,
),
)
# 0.60 sits ABOVE the largest single signal weight on purpose — see
# post_association_service.WEIGHTS. That is what makes "time proximity
# alone is never sufficient" arithmetic rather than aspirational.
op.add_column(
"import_settings",
sa.Column(
"discord_link_threshold", sa.Float(),
server_default="0.60", nullable=False,
),
)
op.add_column(
"import_settings",
sa.Column(
"discord_link_window_hours", sa.Float(),
server_default="24", nullable=False,
),
)
def downgrade() -> None:
op.drop_column("import_settings", "discord_link_window_hours")
op.drop_column("import_settings", "discord_link_threshold")
op.drop_column("import_settings", "discord_link_enabled")
op.drop_index(op.f("ix_post_association_status"), table_name="post_association")
op.drop_index(
op.f("ix_post_association_payload_post_id"), table_name="post_association",
)
op.drop_index(
op.f("ix_post_association_announcement_post_id"), table_name="post_association",
)
op.drop_table("post_association")
+63
View File
@@ -0,0 +1,63 @@
"""membership_sync — whether the roster actually synced, and when.
Milestone 387, step C3.
`platform_membership` (0091) records what was SEEN. This records whether
looking happened at all — a different fact, and the one that makes an empty
roster readable.
Without it, three situations collapse into one: the account subscribes to
nothing, the sweep never ran, or the sweep failed. All three leave zero rows
in `platform_membership`. "You are tracking 12 sources you no longer subscribe
to" is correct in the first case and an invitation to cancel things the
operator is actively paying for in the other two, which is why C4 gates its
CONCLUSIONS on `last_success_at` rather than merely displaying it.
Two timestamps rather than one, deliberately: `last_attempt_at` moves every
run, `last_success_at` only on a clean walk, and the gap between them is what
lets the UI say "last synced 3 days ago, tried 20 minutes ago, failing".
Revision ID: 0095
Revises: 0094
Create Date: 2026-09-11
"""
from typing import Sequence, Union
import sqlalchemy as sa
from alembic import op
revision: str = "0095"
down_revision: Union[str, None] = "0094"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
op.create_table(
"membership_sync",
sa.Column("id", sa.Integer(), nullable=False),
sa.Column("platform", sa.String(length=64), nullable=False),
sa.Column("last_attempt_at", sa.DateTime(timezone=True), nullable=True),
sa.Column("last_success_at", sa.DateTime(timezone=True), nullable=True),
sa.Column("last_count", sa.Integer(), nullable=True),
# No CHECK: this carries an exception class name, and the vocabulary is
# whatever the client raises — same call as source.error_type.
sa.Column("last_error_type", sa.String(length=64), nullable=True),
sa.Column("last_error_message", sa.Text(), nullable=True),
sa.Column(
"updated_at", sa.DateTime(timezone=True),
server_default=sa.text("now()"), nullable=False,
),
sa.PrimaryKeyConstraint("id", name=op.f("pk_membership_sync")),
# The upsert's conflict target, named explicitly because the service
# references it by name in ON CONFLICT.
sa.UniqueConstraint("platform", name="uq_membership_sync_platform"),
)
# No secondary indexes: one row per platform, so every read is a short scan
# and an index would be write cost buying nothing (#3301 removed seven of
# that shape). Same reasoning as platform_membership in 0091.
def downgrade() -> None:
op.drop_table("membership_sync")
@@ -0,0 +1,106 @@
"""artist_membership_suggestion — proposing that a creator and a membership match.
Milestone 388, step E4.
## What this migration deliberately does NOT add
No association table between Artist and Source, and no schema change to either.
E4's first job was to verify what was actually missing, and the answer was
neither the model nor the flows: `Source.artist_id` is a plain FK so many
sources per artist already works, `POST /api/sources` already takes an
`artist_id`, the add-source dialog already attaches to an EXISTING artist, and
`SourceService.reassign` already moves a source between artists with post and
image re-attribution. Building a parallel association table for a relationship
the schema already expresses would have been the mistake rule 28 names.
What was missing is the SUGGESTION, and that is all this table holds.
Accepting a suggestion adds a SOURCE under the existing artist — it never
merges two artists. Adding a source is trivially undone; a wrong merge silently
mixes two creators' work and corrupts tagging, series and provenance with
nothing left to separate them by.
Revision ID: 0096
Revises: 0095
Create Date: 2026-09-11
"""
from typing import Sequence, Union
import sqlalchemy as sa
from alembic import op
revision: str = "0096"
down_revision: Union[str, None] = "0095"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
op.create_table(
"artist_membership_suggestion",
sa.Column("id", sa.Integer(), nullable=False),
sa.Column("platform_membership_id", sa.Integer(), nullable=False),
sa.Column("artist_id", sa.Integer(), nullable=False),
sa.Column("score", sa.Float(), nullable=False),
sa.Column("signals", sa.JSON(), nullable=True),
# No CHECK on status (rule 36 considered and declined), matching
# series_suggestion and post_association — the same review-queue
# vocabulary and the same check-existing-enums lesson.
sa.Column(
"status", sa.String(length=16), server_default="pending", nullable=False,
),
sa.Column(
"created_at", sa.DateTime(timezone=True),
server_default=sa.text("now()"), nullable=False,
),
sa.Column(
"updated_at", sa.DateTime(timezone=True),
server_default=sa.text("now()"), nullable=False,
),
sa.PrimaryKeyConstraint("id", name=op.f("pk_artist_membership_suggestion")),
# CASCADE both ways: a suggestion about a membership or an artist that
# no longer exists is not a fact worth keeping, and a dangling proposal
# would render as a broken row in the review queue.
sa.ForeignKeyConstraint(
["platform_membership_id"], ["platform_membership.id"],
ondelete="CASCADE",
name=op.f("fk_artist_membership_suggestion_membership"),
),
sa.ForeignKeyConstraint(
["artist_id"], ["artist.id"], ondelete="CASCADE",
name=op.f("fk_artist_membership_suggestion_artist_id_artist"),
),
sa.UniqueConstraint(
"platform_membership_id", "artist_id",
name="uq_artist_membership_suggestion_pair",
),
)
op.create_index(
op.f("ix_artist_membership_suggestion_platform_membership_id"),
"artist_membership_suggestion", ["platform_membership_id"],
)
op.create_index(
op.f("ix_artist_membership_suggestion_artist_id"),
"artist_membership_suggestion", ["artist_id"],
)
op.create_index(
op.f("ix_artist_membership_suggestion_status"),
"artist_membership_suggestion", ["status"],
)
def downgrade() -> None:
op.drop_index(
op.f("ix_artist_membership_suggestion_status"),
table_name="artist_membership_suggestion",
)
op.drop_index(
op.f("ix_artist_membership_suggestion_artist_id"),
table_name="artist_membership_suggestion",
)
op.drop_index(
op.f("ix_artist_membership_suggestion_platform_membership_id"),
table_name="artist_membership_suggestion",
)
op.drop_table("artist_membership_suggestion")
@@ -0,0 +1,68 @@
"""Disable sources on retired platforms, so the scheduler stops selecting them.
Milestone #406, phase 1 (switch pixiv off). Rule #171 records the scope decision.
## Why this is a migration and not a button
The live instance had one pixiv source still ENABLED when pixiv was retired
(read 2026-09-13, step 1) even though the operator believed it gone. Unregistering
a platform removes it from code; it does not touch the `source` rows that name it.
Left enabled, that row keeps being picked by the scheduler every interval, and
`download_backends` now refuses it with `unsupported_url` — forever, as a
climbing failure count on a source the operator has already given up.
A migration reaches the live instance on deploy without depending on anyone
finding the row and clicking it. The `run_download` guard is what makes a stale
enabled row SAFE; this is what makes it QUIET.
## Deliberately NOT done here
- **No rows are deleted.** Deleting a source sets its posts' `source_id` to NULL
(FK `ON DELETE SET NULL`), and `uq_post_artist_external_id_null_source` can
reject that if a source-less copy of one of those posts already exists. That
needs checking against real data first, which is phase 2's job (step 6). A
disable cannot collide with anything.
- **No posts or images are touched.** The art stays.
- **deviantart is included** because #3069 retired it and nothing disabled its
rows either. The read found none, so for it this is a no-op — written anyway,
so the statement names every retired platform rather than just the latest one.
## Hardcoded platform names
A migration is a record of one event, frozen in time, so it names the platforms
it acted on rather than importing today's registry — the registry will keep
changing and this revision must not.
Revision ID: 0097
Revises: 0096
Create Date: 2026-09-13
"""
from typing import Sequence, Union
from alembic import op
revision: str = "0097"
down_revision: Union[str, None] = "0096"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
# Clears the failure state the same way `SourceService.update` does when a
# source is disabled through the app (issue #1285), so a retired source
# does not keep showing as failing after it stops being polled. A disable
# done here and one done by clicking must leave identical rows.
op.execute(
"UPDATE source SET enabled = false, last_error = NULL, "
"error_type = NULL, consecutive_failures = 0 "
"WHERE enabled AND platform IN ('pixiv', 'deviantart')"
)
def downgrade() -> None:
# Irreversible by design: which of these rows were enabled before is not
# recorded, and re-enabling every retired-platform source would resume
# polling services the product no longer supports. Rule #22 owes no
# migration story back to a dropped platform.
pass
+2
View File
@@ -38,6 +38,7 @@ def all_blueprints() -> list[Blueprint]:
from .suggestions import suggestions_bp
from .system_activity import system_activity_bp
from .system_backup import system_backup_bp
from .system_health import system_health_bp
from .tags import tags_bp
from .thumbnails import thumbnails_bp
return [
@@ -51,6 +52,7 @@ def all_blueprints() -> list[Blueprint]:
showcase_bp,
settings_bp,
system_activity_bp,
system_health_bp,
system_backup_bp,
admin_bp,
cleanup_bp,
+20
View File
@@ -21,6 +21,7 @@ from ..services.gallery_service import image_url
from ..services.ml.gpu_jobs import GpuJobService, error_dedupe_statements
from ..services.ml.gpu_triage import classify_reason, recover_defective_image
from ..services.ml.regions import RegionService
from ..services.service_roster import touch_service
gpu_bp = Blueprint("gpu", __name__, url_prefix="/api/gpu")
@@ -256,6 +257,18 @@ async def lease():
if not await _agent_authed(session):
return jsonify({"error": "unauthorized"}), 401
jobs = await GpuJobService(session).lease(agent_id, batch_size=batch)
# The agent cannot be polled — it is HTTP-only and pulls from here, so
# web never dials it. A lease IS the check-in, and until milestone 365
# it was thrown away: an agent sitting idle with nothing to lease left
# no trace at all and was indistinguishable from one switched off a
# week ago. Recorded on the call that was already happening.
await touch_service(
session,
key=f"agent:{agent_id}",
kind="agent",
display_name="GPU agent" if agent_id == "agent" else f"GPU agent ({agent_id})",
details={"agent_id": agent_id, "last_call": "lease", "leased": len(jobs)},
)
ml = await MLSettings.load(session)
# image rows for url/mime in one shot
ids = [j.image_record_id for j in jobs]
@@ -329,6 +342,13 @@ async def heartbeat():
if not await _agent_authed(session):
return jsonify({"error": "unauthorized"}), 401
n = await GpuJobService(session).heartbeat(agent_id, job_ids)
await touch_service(
session,
key=f"agent:{agent_id}",
kind="agent",
display_name="GPU agent" if agent_id == "agent" else f"GPU agent ({agent_id})",
details={"agent_id": agent_id, "last_call": "heartbeat", "extended": n},
)
await session.commit()
return jsonify({"extended": n})
+29
View File
@@ -48,6 +48,17 @@ _EDITABLE = (
"process_conflict_threshold",
"embedder_model_name",
"embedder_model_version",
# Discord drop grouping (#388 E2). Operator-facing because the quality bar
# is a judgement no test can settle: too greedy merges distinct pieces, too
# shy leaves a drop scattered.
"discord_grouping_enabled",
"discord_group_max_distance",
"discord_group_window_minutes",
# E3: how long a grouping stays open, and the anti-thrash rule that keeps
# a growing one from monopolising the feed.
"discord_group_close_after_hours",
"discord_group_resurface_min_images",
"discord_group_resurface_cooldown_hours",
*_DETECTOR_FIELDS,
)
@@ -148,6 +159,24 @@ def _validate(p: dict) -> str | None:
return f"process_auto_apply_threshold must be between {AUTO_APPLY_THRESHOLD_MIN} and {AUTO_APPLY_THRESHOLD_MAX}"
if not (0.0 <= float(p["process_conflict_threshold"]) <= 1.0):
return "process_conflict_threshold must be between 0 and 1"
# Discord drop grouping (#388 E2). max_distance is a cosine DISTANCE, so
# unlike the *_threshold family above it is not on the auto-apply scale:
# 0 is identical and 1 is unrelated, and both ends are legal. The upper
# bound is 1.0 rather than AUTO_APPLY_THRESHOLD_MAX for that reason.
if not (0.0 <= float(p["discord_group_max_distance"]) <= 1.0):
return "discord_group_max_distance must be between 0 and 1"
if float(p["discord_group_window_minutes"]) <= 0:
return "discord_group_window_minutes must be > 0"
# A group must stay open at least as long as the drop window it was cut
# with, or the joiner could never reach a message the grouper deferred —
# the two would fight, and the symptom (drops that never grow) would look
# like the predicate failing rather than a settings contradiction.
if float(p["discord_group_close_after_hours"]) * 60 < float(p["discord_group_window_minutes"]):
return "discord_group_close_after_hours must be at least the drop window"
if int(p["discord_group_resurface_min_images"]) < 1:
return "discord_group_resurface_min_images must be >= 1"
if float(p["discord_group_resurface_cooldown_hours"]) < 0:
return "discord_group_resurface_cooldown_hours must be >= 0"
# Embedder model swap (#1190): both must be non-empty. Changing them means a
# different embedding space — the operator must re-embed + retrain after.
for key in ("embedder_model_name", "embedder_model_version"):
+45
View File
@@ -5,6 +5,8 @@ from quart import Blueprint, jsonify, request
from ..extensions import get_session
from ..models import ImportSettings, Post
from ..services import interpreter_client as ic
from ..services.post_association_service import PostAssociationService
from ..services.post_association_service import rescan as association_rescan
from ..services.post_feed_service import PostFeedService
from ..services.source_service import KNOWN_PLATFORMS
from ..utils.text import html_to_plain
@@ -165,3 +167,46 @@ async def set_translation_override(post_id: int):
"translated_source_lang": post.translated_source_lang,
"applied": applied,
})
# --- #388 E5: the announcement review queue -------------------------------
#
# Confirm-only, following the series-suggestion routes (api/tags.py). Nothing
# here links anything on its own: the matcher proposes, the operator decides.
@posts_bp.route("/associations", methods=["GET"])
async def list_associations():
async with get_session() as session:
return jsonify({"items": await PostAssociationService(session).list_pending()})
@posts_bp.route("/associations/<int:association_id>/accept", methods=["POST"])
async def accept_association(association_id: int):
async with get_session() as session:
result = await PostAssociationService(session).accept(association_id)
if result is None:
return _bad("association not found", 404)
await session.commit()
return jsonify(result)
@posts_bp.route("/associations/<int:association_id>/dismiss", methods=["POST"])
async def dismiss_association(association_id: int):
async with get_session() as session:
result = await PostAssociationService(session).dismiss(association_id)
if result is None:
return _bad("association not found", 404)
await session.commit()
return jsonify(result)
@posts_bp.route("/associations/rescan", methods=["POST"])
async def rescan_associations():
"""Manual re-scan. The beat sweep only looks at recent posts (a pair has to
be within the window to exist at all); this is the button for a first run
over a library that predates the feature."""
async with get_session() as session:
result = await association_rescan(session)
await session.commit()
return jsonify(result)
+20
View File
@@ -39,6 +39,10 @@ _EDITABLE_FIELDS = (
"download_failure_warning_threshold",
"series_suggest_enabled",
"series_suggest_threshold",
# #388 E5 — the announcement matcher (Patreon teaser ↔ Discord drop).
"discord_link_enabled",
"discord_link_threshold",
"discord_link_window_hours",
"extdl_mega_enabled",
"extdl_gdrive_enabled",
"extdl_mediafire_enabled",
@@ -150,6 +154,22 @@ async def update_import_settings():
return jsonify(
{"error": "series_suggest_threshold must be a number in [0, 1]"}
), 400
if "discord_link_enabled" in body and not isinstance(
body["discord_link_enabled"], bool
):
return jsonify({"error": "discord_link_enabled must be a boolean"}), 400
if "discord_link_threshold" in body:
v = body["discord_link_threshold"]
if not isinstance(v, (int, float)) or isinstance(v, bool) or v < 0 or v > 1:
return jsonify(
{"error": "discord_link_threshold must be a number in [0, 1]"}
), 400
if "discord_link_window_hours" in body:
v = body["discord_link_window_hours"]
if not isinstance(v, (int, float)) or isinstance(v, bool) or v <= 0:
return jsonify(
{"error": "discord_link_window_hours must be a positive number"}
), 400
if "wip_title_tagging_enabled" in body and not isinstance(
body["wip_title_tagging_enabled"], bool
):
+167 -2
View File
@@ -1,10 +1,15 @@
"""FC-3a: CRUD over Source rows. FC-3c adds POST /<id>/check."""
from quart import Blueprint, jsonify, request
from sqlalchemy import select
from sqlalchemy import func, select
from ..extensions import get_session
from ..models import DownloadEvent, Source
from ..models import DownloadEvent, MembershipSync, PlatformMembership, Source
from ..services.artist_membership_service import ArtistMembershipService
from ..services.artist_membership_service import rescan as membership_rescan
from ..services.artist_service import ArtistService
from ..services.membership_reconcile import reconcile_all
from ..services.membership_roster import roster_is_fresh, source_for_membership
from ..services.scheduler_service import active_platform_cooldowns, scheduler_status
from ..services.source_service import (
KNOWN_PLATFORMS,
@@ -288,3 +293,163 @@ async def check_source(source_id: int):
download_source.delay(source_id)
return jsonify({"download_event_id": event_id, "status": "pending"}), 202
# --- #387 C3: the membership roster's sync state --------------------------
#
# Rule 164's visibility requirement lives here. A roster that failed to sync,
# or never has, must be DISTINGUISHABLE from an account that subscribes to
# nothing — otherwise the reconciliation this unlocks would tell the operator
# to cancel sources they are actively paying for.
@sources_bp.route("/membership-sync", methods=["GET"])
async def membership_sync_status():
async with get_session() as session:
rows = (await session.execute(select(MembershipSync))).scalars().all()
counts = dict(
(await session.execute(
select(PlatformMembership.platform, func.count())
.group_by(PlatformMembership.platform)
)).all()
)
return jsonify({"platforms": [
{
"platform": r.platform,
"last_attempt_at": r.last_attempt_at.isoformat() if r.last_attempt_at else None,
# NULL here means NEVER, and the UI must say so in words. Rendering
# it as 0 or as "-" is the exact conflation this endpoint exists to
# prevent.
"last_success_at": r.last_success_at.isoformat() if r.last_success_at else None,
"last_count": r.last_count,
"last_error_type": r.last_error_type,
"last_error_message": r.last_error_message,
# Whether a CONCLUSION may be drawn from this roster — not merely
# whether it looks recent. C4 gates on this, and it is computed
# server-side so no caller can forget to.
"fresh": roster_is_fresh(r),
"known_memberships": counts.get(r.platform, 0),
}
for r in sorted(rows, key=lambda r: r.platform)
]})
@sources_bp.route("/membership-sync", methods=["POST"])
async def trigger_membership_sync():
"""Run the roster sweep now.
The beat schedule runs daily, which is right for a billing-cycle fact but
far too slow when the operator has just connected a credential and wants to
see whether it works. Queued rather than run inline: it crosses the network
to an external service and the request path is not where that belongs.
"""
from ..tasks.maintenance import sync_memberships
sync_memberships.delay()
return jsonify({"queued": True})
# --- #388 E4: creator/membership suggestions ------------------------------
#
# Confirm-only. Accepting ADDS A SOURCE under the existing artist — it never
# merges two artists, because adding a source is trivially undone and a wrong
# merge silently mixes two creators' work with nothing left to separate them by.
@sources_bp.route("/membership-suggestions", methods=["GET"])
async def list_membership_suggestions():
async with get_session() as session:
return jsonify({"items": await ArtistMembershipService(session).list_pending()})
@sources_bp.route("/membership-suggestions/<int:sid>/accept", methods=["POST"])
async def accept_membership_suggestion(sid: int):
async with get_session() as session:
result = await ArtistMembershipService(session).accept(sid)
if result is None:
return _bad("suggestion_not_found", status=404)
await session.commit()
return jsonify(result)
@sources_bp.route("/membership-suggestions/<int:sid>/dismiss", methods=["POST"])
async def dismiss_membership_suggestion(sid: int):
async with get_session() as session:
result = await ArtistMembershipService(session).dismiss(sid)
if result is None:
return _bad("suggestion_not_found", status=404)
await session.commit()
return jsonify(result)
@sources_bp.route("/membership-suggestions/rescan", methods=["POST"])
async def rescan_membership_suggestions():
async with get_session() as session:
result = await membership_rescan(session)
await session.commit()
return jsonify(result)
# --- #387 C4: reconciling the roster against what FC actually tracks -------
#
# Asymmetric on purpose. The "you subscribe but FC doesn't follow it" direction
# carries a per-row action, because adding a source is the reversible half. The
# "FC follows it but your roster doesn't show it" direction is REPORT ONLY by
# the operator's decision (2026-09-11): it says what it sees and links to the
# Subscriptions row, and offers no one-click disable.
@sources_bp.route("/reconciliation", methods=["GET"])
async def reconciliation():
async with get_session() as session:
return jsonify(await reconcile_all(session))
@sources_bp.route("/reconciliation/adopt", methods=["POST"])
async def adopt_membership():
"""Start tracking a creator the roster says the account already pays for.
One row, one click, never a sweep side effect: adding a source commits disk,
worker time and rate budget, and unwinding it means deleting files.
"""
body = await request.get_json()
if not isinstance(body, dict):
return _bad("invalid_body", status=400)
membership_id = body.get("membership_id")
if not isinstance(membership_id, int):
return _bad("membership_id_required", status=400)
async with get_session() as session:
membership = await session.get(PlatformMembership, membership_id)
if membership is None:
return _bad("membership_not_found", status=404)
if not membership.url:
return _bad("membership_has_no_url", status=400)
existing = await source_for_membership(session, membership)
if existing is not None:
# The operator got there by another route between the page load and
# the click. That is them being ahead of us, not an error.
return jsonify({"already_tracked": existing.id})
# The sweep already captured the creator's real display name, so the
# artist gets its true name with NO lookup on the request path. Task
# #1293 asked for `resolve_display_name` here; the roster satisfies that
# concern earlier in the pipeline than #1293 expected, which also keeps
# this route off the network entirely (rule 164). The vanity is the
# fallback, never the preferred value.
name = membership.display_name or membership.vanity_or_none()
if not name:
return _bad("membership_has_no_name", status=400)
artist, _created = await ArtistService(session).find_or_create(name)
try:
record = await SourceService(session).create(
artist_id=artist.id,
platform=membership.platform,
url=membership.url,
)
except DuplicateSourceError as exc:
return jsonify({"already_tracked": exc.existing_id})
artist_id = artist.id
return jsonify({"source_id": record.id, "artist_id": artist_id}), 201
+192
View File
@@ -0,0 +1,192 @@
"""Is every part of FabledCurator running? One verdict, one endpoint.
Milestone 365. The nav indicator and the System page both read this and
nothing else — composing a verdict is this module's job, not the UI's.
## Two kinds of part, answered two different ways
**Learned** — celery roles and the GPU agent, from `service_seen`. The
question is "how long since it checked in", and these are the parts that can
be ABSENT, which is the whole point: `celery inspect` alone reports presence,
so a dead worker is a shorter list rather than a red light.
**Probed live** — Postgres and Redis. Always expected, never learned, and a
last-seen for them would be actively misleading: that Redis answered thirty
seconds ago says nothing about now.
## This endpoint must never fail because something it checks has failed
The inversion is easy to write by accident and it destroys the feature exactly
when it is needed — a 500 when Redis is down, instead of `redis: down`. Every
probe is wrapped, every wait has a deadline (rule 156), and the roster refresh
swallows its own errors. The worst case is a part reported `unknown`, which is
a true statement.
"""
from __future__ import annotations
import asyncio
import logging
import time
from datetime import UTC, datetime
from quart import Blueprint, jsonify
from sqlalchemy import select, text
from ..config import get_config
from ..extensions import get_session
from ..models import ServiceSeen
from ..services.service_roster import refresh_if_stale
log = logging.getLogger(__name__)
system_health_bp = Blueprint("system_health", __name__, url_prefix="/api/system")
# How long a learned part may go quiet before it is doubted, then disbelieved.
#
# These are deliberately generous, and the reason is a deploy rather than a
# worker: `docker compose up -d` rolls start-first, so a role is briefly served
# by two containers and then by neither while the old one drains. Thresholds
# tight enough to catch a crash in seconds would paint the page red every time
# the stack is updated, and an alarm that cries wolf on every deploy is one
# nobody reads. Tune down only after watching a real deploy pass through.
STALE_AFTER_SECONDS = 90
DOWN_AFTER_SECONDS = 300
# Probes cross a process boundary, so they carry deadlines. A hung Postgres
# must make this endpoint say "postgres: down", not hang alongside it.
PROBE_TIMEOUT_SECONDS = 2.0
_OK, _STALE, _DOWN, _UNKNOWN = "ok", "stale", "down", "unknown"
# Worst-first, so an overall verdict is just the max.
_SEVERITY = {_OK: 0, _UNKNOWN: 1, _STALE: 2, _DOWN: 3}
def _age_state(age_seconds: float) -> str:
if age_seconds >= DOWN_AFTER_SECONDS:
return _DOWN
if age_seconds >= STALE_AFTER_SECONDS:
return _STALE
return _OK
def _describe_learned(name: str, state: str, age: float, details: dict) -> str:
"""Say what the state MEANS. A red chip tells an operator less than a
sentence does at the moment they are deciding whether to go and look."""
if state == _OK:
replicas = details.get("replicas")
if replicas and replicas > 1:
return f"{name} is running ({replicas} replicas)"
return f"{name} is running"
mins = int(age // 60)
ago = f"{mins} min" if mins else f"{int(age)}s"
if state == _STALE:
return f"{name} has not checked in for {ago}"
return f"{name} has not checked in for {ago} — treat it as stopped"
async def _probe_postgres(session) -> dict:
started = time.monotonic()
try:
await asyncio.wait_for(
session.execute(text("SELECT 1")), timeout=PROBE_TIMEOUT_SECONDS
)
except Exception as exc: # noqa: BLE001 — a probe reports, it never raises
return {
"key": "postgres", "kind": "datastore", "name": "PostgreSQL",
"state": _DOWN, "detail": f"not answering: {type(exc).__name__}",
}
return {
"key": "postgres", "kind": "datastore", "name": "PostgreSQL", "state": _OK,
"detail": "answering", "latency_ms": round((time.monotonic() - started) * 1000, 1),
}
def _ping_redis_sync() -> None:
import redis # local import; mirrors system_activity's pattern
client = redis.Redis.from_url(
get_config().celery_broker_url,
socket_connect_timeout=PROBE_TIMEOUT_SECONDS,
socket_timeout=PROBE_TIMEOUT_SECONDS,
)
client.ping()
async def _probe_redis() -> dict:
started = time.monotonic()
try:
await asyncio.wait_for(
asyncio.to_thread(_ping_redis_sync), timeout=PROBE_TIMEOUT_SECONDS * 2
)
except Exception as exc: # noqa: BLE001
return {
"key": "redis", "kind": "datastore", "name": "Redis",
"state": _DOWN,
"detail": f"not answering: {type(exc).__name__} — queues and workers "
f"cannot be reached either",
}
return {
"key": "redis", "kind": "datastore", "name": "Redis", "state": _OK,
"detail": "answering", "latency_ms": round((time.monotonic() - started) * 1000, 1),
}
@system_health_bp.route("/health", methods=["GET"])
async def system_health():
"""Every part, its state, and one overall verdict.
Response: {overall, parts: [{key, kind, name, state, detail, last_seen_at,
…}], checked_at}
"""
parts: list[dict] = []
now = datetime.now(UTC)
async with get_session() as session:
# Postgres first, and if it is unreachable nothing else can be read —
# say so rather than failing, because "the database is down" is the
# single most useful thing this endpoint can ever report.
pg = await _probe_postgres(session)
parts.append(pg)
if pg["state"] == _OK:
# Rate-limited inside; see service_roster on why the web process
# is the right observer.
try:
await refresh_if_stale(session)
await session.commit()
except Exception: # noqa: BLE001
log.warning("system health: roster refresh failed", exc_info=True)
rows = (
await session.execute(select(ServiceSeen).order_by(ServiceSeen.display_name))
).scalars().all()
for row in rows:
age = (now - row.last_seen_at).total_seconds()
state = _age_state(age)
parts.append({
"key": row.key,
"kind": row.kind,
"name": row.display_name,
"state": state,
"detail": _describe_learned(row.display_name, state, age, row.details or {}),
"last_seen_at": row.last_seen_at.isoformat(),
"first_seen_at": row.first_seen_at.isoformat(),
**{k: v for k, v in (row.details or {}).items() if k != "agent_id"},
})
parts.append(await _probe_redis())
overall = max((p["state"] for p in parts), key=lambda s: _SEVERITY[s], default=_UNKNOWN)
return jsonify({
"overall": overall,
"parts": sorted(parts, key=lambda p: (-_SEVERITY[p["state"]], p["name"])),
"checked_at": now.isoformat(),
# So the UI can explain a `stale` without hard-coding the same numbers
# in a second place.
"thresholds": {
"stale_after_seconds": STALE_AFTER_SECONDS,
"down_after_seconds": DOWN_AFTER_SECONDS,
},
})
+20
View File
@@ -200,6 +200,26 @@ def make_celery() -> Celery:
"task": "backend.app.tasks.maintenance.snapshot_head_metrics",
"schedule": 86400.0,
},
"group-discord-drops-hourly": {
"task": "backend.app.tasks.maintenance.group_discord_drops",
"schedule": 3600.0, # hourly. Not daily: the grouping signal is
# the SigLIP embedding, which lands asynchronously AFTER import
# (#388 E2), so this sweep is what picks up a drop once its
# vectors have caught up. No-op unless discord_grouping_enabled.
},
"match-post-associations-hourly": {
"task": "backend.app.tasks.maintenance.match_post_associations",
"schedule": 3600.0, # hourly, and AFTER the grouper's own cadence
# by construction: a pair cannot be proposed until the drop it
# points at exists as a grouping (#388 E5). No-op unless
# discord_link_enabled.
},
"sync-memberships-daily": {
"task": "backend.app.tasks.maintenance.sync_memberships",
"schedule": 86400.0, # daily — memberships change on a BILLING
# cycle, not a download cadence (#387 C3). No-op per platform
# when the client lacks the seam or no credential exists.
},
"integrity-verify-weekly": {
"task": "backend.app.tasks.maintenance.verify_integrity",
"schedule": 604800.0, # weekly
+4 -2
View File
@@ -16,8 +16,11 @@ class Config:
celery_broker_url: str
celery_result_backend: str
# Sets Quart's app.secret_key. Nothing signs a cookie today (FC has no
# login and no session use), so this currently protects nothing — it is
# required rather than defaulted so that the day something session-backed
# does land, no instance is already running on a value we published.
secret_key: str
extension_api_key: str # used by the Firefox extension; lands in FC-3 but read here
log_level: str
@property
@@ -47,6 +50,5 @@ def get_config() -> Config:
celery_broker_url=os.environ.get("CELERY_BROKER_URL", "redis://redis:6379/0"),
celery_result_backend=os.environ.get("CELERY_RESULT_BACKEND", "redis://redis:6379/0"),
secret_key=os.environ["SECRET_KEY"],
extension_api_key=os.environ.get("EXTENSION_API_KEY", ""),
log_level=os.environ.get("LOG_LEVEL", "INFO"),
)
+10
View File
@@ -2,6 +2,7 @@
from .app_setting import AppSetting
from .artist import Artist
from .artist_membership_suggestion import ArtistMembershipSuggestion
from .artist_visit import ArtistVisit
from .backup_run import BackupRun
from .base import Base
@@ -21,17 +22,21 @@ from .import_batch import ImportBatch
from .import_settings import ImportSettings
from .import_task import ImportTask
from .library_audit_run import LibraryAuditRun
from .membership_sync import MembershipSync
from .ml_settings import MLSettings
from .patreon_failed_media import PatreonFailedMedia
from .patreon_seen_media import PatreonSeenMedia
from .pixiv_failed_media import PixivFailedMedia
from .pixiv_seen_media import PixivSeenMedia
from .platform_membership import PlatformMembership
from .post import Post
from .post_association import PostAssociation
from .post_attachment import PostAttachment, attachment_download_url
from .presentation_review import PresentationReview
from .series_chapter import SeriesChapter
from .series_page import SeriesPage
from .series_suggestion import SeriesSuggestion
from .service_seen import ServiceSeen
from .source import Source
from .subscribestar_failed_media import SubscribeStarFailedMedia
from .subscribestar_seen_media import SubscribeStarSeenMedia
@@ -46,6 +51,7 @@ __all__ = [
"Base",
"AppSetting",
"Artist",
"ArtistMembershipSuggestion",
"ArtistVisit",
"BackupRun",
"Source",
@@ -57,12 +63,15 @@ __all__ = [
"SubscribeStarFailedMedia",
"SubscribeStarSeenMedia",
"Post",
"PostAssociation",
"PostAttachment",
"attachment_download_url",
"PresentationReview",
"SeriesChapter",
"SeriesPage",
"SeriesSuggestion",
"PlatformMembership",
"ServiceSeen",
"ImageRecord",
"ImageProvenance",
"ImageRegion",
@@ -76,6 +85,7 @@ __all__ = [
"ImportTask",
"ImportSettings",
"LibraryAuditRun",
"MembershipSync",
"MLSettings",
"HeadAutoApplyRun",
"HeadMetric",
@@ -0,0 +1,83 @@
"""artist_membership_suggestion — "this creator and that membership are the same".
Milestone 388, step E4.
## What was NOT needed here
E4's first job was to check what is actually missing, and the answer was: not
the schema, and not the flows. `Source.artist_id` is a plain FK, so many
sources per artist is already the data model; `POST /api/sources` already takes
an `artist_id`; the add-source dialog already has an artist autocomplete that
attaches to an EXISTING artist; and `SourceService.reassign` already moves a
source between artists WITH post and image re-attribution. A sweep for
one-source-per-artist assumptions found only `func.count()` calls, which are
the opposite of assuming one.
So no parallel association table was built for a relationship the schema
already expresses (rule 28). What was missing is the SUGGESTION — FC proposing
the link from the roster instead of waiting to be told.
## Confirm-only, and what "accept" actually does
Accepting adds a SOURCE for the membership's platform under the artist that
already has the other channel. It does NOT merge two artists. That distinction
is the whole safety margin: adding a source is trivially undone, whereas a
wrong artist merge silently mixes two creators' work and corrupts tagging,
series and provenance downstream — with nothing left to tell them apart by.
Dismissed rows are kept, not deleted, for the same reason as every other review
queue here: the row is what remembers the rejection, and re-proposing a
rejected pair on every scan is what makes a queue get ignored.
"""
from datetime import datetime
from sqlalchemy import (
JSON,
DateTime,
Float,
ForeignKey,
Integer,
String,
UniqueConstraint,
func,
)
from sqlalchemy.orm import Mapped, mapped_column
from .base import Base
class ArtistMembershipSuggestion(Base):
__tablename__ = "artist_membership_suggestion"
__table_args__ = (
UniqueConstraint(
"platform_membership_id", "artist_id",
name="uq_artist_membership_suggestion_pair",
),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True)
platform_membership_id: Mapped[int] = mapped_column(
ForeignKey("platform_membership.id", ondelete="CASCADE"),
nullable=False, index=True,
)
artist_id: Mapped[int] = mapped_column(
ForeignKey("artist.id", ondelete="CASCADE"), nullable=False, index=True
)
score: Mapped[float] = mapped_column(Float, nullable=False)
# Per-signal strengths as scored. Without it, "why was this suggested" is
# unanswerable the moment a weight or the threshold moves.
signals: Mapped[dict | None] = mapped_column(JSON, nullable=True)
# pending | linked | dismissed. Plain String, no CHECK — same call as
# series_suggestion.status and post_association.status.
status: Mapped[str] = mapped_column(
String(16), nullable=False, server_default="pending", index=True
)
created_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now()
)
updated_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False,
server_default=func.now(), onupdate=func.now(),
)
+27
View File
@@ -97,6 +97,33 @@ class ImportSettings(Base):
server_default="0.5",
)
# Milestone 388 E5 — the announcement matcher: "this Patreon post announced
# that Discord drop". Lives here rather than in MLSettings, with the series
# matcher it is modelled on, because it runs no inference: the signals are
# time proximity and whether the post says so.
discord_link_enabled: Mapped[bool] = mapped_column(
Boolean, nullable=False, default=True,
server_default="true",
)
# The weighted-score cut-off. 0.60 is not arbitrary: it is deliberately set
# ABOVE the largest single signal weight, which is what makes "time
# proximity alone must never be sufficient" an ARITHMETIC property rather
# than a hope. On a busy day an artist posts several times; if proximity
# could carry a pair by itself, every one of those days would produce false
# pairs and the review queue would be abandoned. See
# post_association_service.WEIGHTS — a guard test pins the relationship.
discord_link_threshold: Mapped[float] = mapped_column(
Float, nullable=False, default=0.60,
server_default="0.60",
)
# How far apart the announcement and the drop may be. The Patreon post
# exists IN ORDER TO announce the drop, so they are minutes-to-hours apart;
# a day is generous and still excludes "same week".
discord_link_window_hours: Mapped[float] = mapped_column(
Float, nullable=False, default=24.0,
server_default="24",
)
# #830 off-platform file-host downloads — per-host enable lever (default on,
# rule #26). Column names are extdl_<host>_enabled so the worker reads them
# via getattr(settings, f"extdl_{host}_enabled", True).
+77
View File
@@ -0,0 +1,77 @@
"""membership_sync — did the roster actually sync, and when.
Milestone 387, step C3.
`platform_membership` records what was SEEN. This records whether looking
happened at all, and that is a different fact — the one that makes an empty
roster readable.
## Why this table has to exist
Without it, three very different situations are one indistinguishable state:
* the account genuinely subscribes to nothing,
* the sweep has never run,
* the sweep ran and failed.
All three produce zero rows in `platform_membership`. Telling the operator
"you are tracking 12 sources you do not subscribe to" is correct in the first
case and catastrophic in the other two — it is an invitation to cancel things
they are actively paying for. C4 must therefore gate its CONCLUSIONS on
`last_success_at`, not merely display it.
`MAX(platform_membership.last_seen_at)` was the tempting shortcut and does not
work: it cannot distinguish "synced fine, found nothing" from "never synced".
`task_run` was the other candidate and is worse — its retention prunes ok rows
after 24h, so a sweep that last succeeded three days ago would leave no trace
at all.
## Separate attempt and success timestamps, deliberately
`last_attempt_at` moves every run; `last_success_at` moves only on a clean
walk. The GAP between them is the staleness signal, and keeping them apart is
what lets the UI say "last synced 3 days ago, last tried 20 minutes ago,
failing" — which is a different message from either half alone.
"""
from datetime import datetime
from sqlalchemy import DateTime, Integer, String, Text, UniqueConstraint, func
from sqlalchemy.orm import Mapped, mapped_column
from .base import Base
class MembershipSync(Base):
__tablename__ = "membership_sync"
__table_args__ = (
UniqueConstraint("platform", name="uq_membership_sync_platform"),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True)
platform: Mapped[str] = mapped_column(String(64), nullable=False)
# Moves on EVERY run, success or not — so "we are trying" is visible even
# while "we are succeeding" is not.
last_attempt_at: Mapped[datetime | None] = mapped_column(
DateTime(timezone=True), nullable=True
)
# Moves only on a COMPLETE walk. This is the freshness signal C4 gates its
# conclusions on; NULL means never — which must never be rendered as zero.
last_success_at: Mapped[datetime | None] = mapped_column(
DateTime(timezone=True), nullable=True
)
# How many memberships the last SUCCESSFUL walk saw. Paired with
# last_success_at so "0" is only ever readable as a real zero.
last_count: Mapped[int | None] = mapped_column(Integer, nullable=True)
# Cleared on success. Plain String, no CHECK — this carries an exception
# class name (PatreonAuthError, PatreonDriftError, ...) and the vocabulary
# is whatever the client raises, exactly as source.error_type works.
last_error_type: Mapped[str | None] = mapped_column(String(64), nullable=True)
last_error_message: Mapped[str | None] = mapped_column(Text, nullable=True)
updated_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False,
server_default=func.now(), onupdate=func.now(),
)
+63
View File
@@ -252,6 +252,69 @@ class MLSettings(Base):
Integer, nullable=False, default=64,
server_default="64",
)
# -- Discord drop grouping (milestone 388) -----------------------------
# FC authors a post out of a creator's variant drop. The predicate is three
# axes ANDed together, and the time one does the real work: SIMILARITY
# ALONE OVER-GROUPS. Any two pieces of the same character by the same
# artist sit close in SigLIP space, so a cosine-only rule collapses a month
# of one character into a single "post". What makes a variant set a set is
# that it was dropped TOGETHER.
discord_grouping_enabled: Mapped[bool] = mapped_column(
# ON by default, matching the operator's standing opt-OUT preference for
# automatic behaviour (2026-06-29, recorded on the head/ccip auto-apply
# switches). Safe to default on because the act is reversible by one
# DELETE: removing a synthetic post un-absorbs its members.
Boolean, nullable=False, default=True,
server_default="true",
)
# Cosine DISTANCE, not similarity — this is the units gallery_service's
# `cosine_distance` already speaks, and converting at the query site is a
# step to get backwards. Lower = stricter. 0.10 is deliberately TIGHT: the
# two failure modes are not symmetric. Grouping too shy leaves a drop
# scattered, which is visible and fixable by raising this; grouping too
# greedy merges distinct pieces into a post that claims they belong
# together, which is the failure that would discredit the feature.
discord_group_max_distance: Mapped[float] = mapped_column(
Float, nullable=False, default=0.10,
server_default=text("0.10"),
)
# The gap that ENDS a drop, measured between CONSECUTIVE messages rather
# than from the first — an artist trickling variants out over an evening is
# one drop, and a window anchored on the first message would cut it in half
# at an arbitrary point.
discord_group_window_minutes: Mapped[float] = mapped_column(
Float, nullable=False, default=60.0,
server_default=text("60"),
)
# How long a synthetic post keeps accepting new members (#388 E3). This is
# NOT the drop window above: the window cuts one sweep's messages into
# drops, this decides how long a finished drop can still be REJOINED when a
# creator adds variants days later. A week by default — long enough for the
# "and here is the alt outfit" follow-up that motivated the feature, short
# enough that a group does not still be open when the same character comes
# round again months later and gets absorbed by mistake.
#
# Openness is DERIVED from this, not stored: a group is open if it grew (or
# started) within this period. So lowering it closes old groups and raising
# it reopens them, which is comprehensible and reversible — the alternative,
# a stored closed_at, would need its own repair path to ever change.
discord_group_close_after_hours: Mapped[float] = mapped_column(
Float, nullable=False, default=168.0,
server_default=text("168"),
)
# Anti-thrash (#388 E3). An updated post SHOULD be visible — that is the
# point of keeping it open — but a group gaining one image a day must not
# monopolise the feed. Growth smaller than this never moves the post, and
# no group moves more than once per cooldown, so a drip-feed updates in
# place while a real second wave resurfaces exactly once.
discord_group_resurface_min_images: Mapped[int] = mapped_column(
Integer, nullable=False, default=2,
server_default="2",
)
discord_group_resurface_cooldown_hours: Mapped[float] = mapped_column(
Float, nullable=False, default=24.0,
server_default=text("24"),
)
updated_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now()
)
+152
View File
@@ -0,0 +1,152 @@
"""platform_membership — the learned roster of what the account actually pays for.
Milestone 387, phase C. FabledCurator knows which creators it has been TOLD to
follow (`source`), and nothing about which ones the operator is actually
subscribed to. Those two sets drift in both directions and the app cannot
currently see either drift:
* A subscription the operator pays for that FC does not track is content they
believe they are archiving and are not.
* A source FC keeps walking after the subscription lapsed is requests spent on
a wall, reported as a creator who has gone quiet.
This table is the memory that makes both visible — every membership the account
has been observed to hold, and when it was last seen.
## Why a learned roster rather than a live lookup
Same reasoning as `service_seen` (milestone 365), and the same shape: an
absence is only observable against a record of presence. A membership that
stops appearing in a sweep is the signal — "you were subscribed to this, now
you aren't" — and there is nowhere to read that from a live call, because a
live call returns what IS, never what stopped being.
It also means the reconciliation surface keeps working when Patreon is
unreachable, degraded to a stale roster with a visible age rather than an empty
page (rule 164).
## Roster truth, NOT per-post truth
The single most important thing about this table: `tier_names` says which tiers
the account holds. It does **not** say which posts those tiers unlock. A
creator can gate a post behind an access rule that maps onto no tier name at
all.
`current_user_can_view` — read per post by `patreon_client.post_is_gated` — is
the authoritative signal, and phase A already turned it into a durable
per-source state. This roster EXPLAINS that state ("you are no longer a patron"
vs "your tier doesn't cover these posts"). It must never be used to decide
whether to fetch something. Getting that backwards would make FC silently stop
fetching content the operator is paying for, which is the worst failure
available in this milestone.
## status is a plain String, and deliberately the platform's own word
Not a Postgres ENUM, not CHECK-gated — matching `service_seen.kind`,
`gpu_job.status` and `source.error_type`. Two reasons, and the first is the
real one:
1. **The vocabulary is not ours to invent.** Patreon says `active_patron` /
`former_patron` / `declined_patron`; SubscribeStar and FANBOX will say
something else. Storing each platform's own word verbatim and mapping to
FC's meaning at the READ site keeps this table a record of what was
observed rather than a lossy translation of it. A lowest-common-denominator
enum picked before any platform has been characterised (step C0) would be a
guess baked into the schema.
2. A constraint swap per new value (rule 36) would be cost with no invariant
behind it, exactly as `service_seen.kind` records.
The service layer owns the whitelist and the mapping; the column owns the
evidence.
## Retention: aged out, never deleted on disappearance
A membership that stops appearing in a sweep is NOT removed. Its disappearance
is the fact the reconciliation surface reads, and deleting the row would
destroy the signal at the moment it became interesting. `last_seen_at` is what
makes "gone" decidable, and a retention policy ages rows out on time rather
than on absence.
"""
from datetime import datetime
from sqlalchemy import JSON, DateTime, Integer, String, Text, UniqueConstraint, func
from sqlalchemy.orm import Mapped, mapped_column
from .base import Base
class PlatformMembership(Base):
__tablename__ = "platform_membership"
__table_args__ = (
# The natural key the sweep's upsert conflicts on. Named explicitly
# because `touch_membership` references it by name in ON CONFLICT.
UniqueConstraint(
"platform", "external_campaign_id",
name="uq_platform_membership_platform_campaign",
),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True)
platform: Mapped[str] = mapped_column(String(64), nullable=False)
# The platform's own id for the thing subscribed to — a Patreon campaign
# id, whatever SubscribeStar and FANBOX call theirs. Text rather than a
# bounded String: these are opaque upstream identifiers and guessing a
# ceiling for a value we do not mint is how a walk dies on a truncation.
external_campaign_id: Mapped[str] = mapped_column(Text, nullable=False)
# For the reconciliation UI, and for matching against Source.url — the
# vanity/URL is what the two sides actually have in common.
display_name: Mapped[str | None] = mapped_column(Text, nullable=True)
url: Mapped[str | None] = mapped_column(Text, nullable=True)
# The platform's own word. See the module docstring — this is evidence,
# not a normalised FC status.
status: Mapped[str | None] = mapped_column(String(32), nullable=True)
# Nullable throughout: a free follow has no tier and no money attached, and
# a platform may not expose an amount at all. Absent must stay
# distinguishable from zero — "free" and "we don't know" are different
# answers to "what is this costing".
tier_names: Mapped[list | None] = mapped_column(JSON, nullable=True)
amount_cents: Mapped[int | None] = mapped_column(Integer, nullable=True)
currency: Mapped[str | None] = mapped_column(String(8), nullable=True)
# NEVER updated after insert. The one field that answers "has this ever
# been true", which is what makes a disappearance readable rather than
# indistinguishable from never having existed. `touch_membership`
# deliberately excludes it from the ON CONFLICT update set.
first_seen_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now(),
)
last_seen_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now(),
)
# The raw membership as the platform returned it, so a later question can
# be answered without re-fetching — and so a field we did not think to
# model is not lost. Displayed and never queried, like service_seen.details.
details: Mapped[dict] = mapped_column(JSON, nullable=False, default=dict)
def vanity_or_none(self) -> str | None:
"""The platform's URL slug for this creator, if it can be known.
NOT a column, and that is C1's design working as intended rather than
an omission: the roster was modelled before any platform had been
characterised, so `details` exists precisely to carry the fields we did
not know to model. The vanity turned out to be one of them (#3886), and
it is reachable without a migration.
Falls back to the URL's last segment, which is what a vanity IS on
every platform seen so far — but only as a fallback, because the
platform's own word for it is the better answer when present.
"""
campaign = (self.details or {}).get("campaign") or {}
vanity = campaign.get("vanity")
if isinstance(vanity, str) and vanity:
return vanity
if self.url:
tail = self.url.rstrip("/").rsplit("/", 1)[-1]
return tail or None
return None
+58
View File
@@ -102,3 +102,61 @@ class Post(Base):
downloaded_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now()
)
# -- Synthetic posts (milestone 388). ----------------------------------
# Discord is a delivery CHANNEL, not a publisher: one message is not one
# post. So FC authors the post itself, grouping a creator's variant drop
# into a single row (services/discord_grouping.py).
#
# NULL for every post a creator actually wrote — which is all of them until
# a grouper runs. Non-NULL names the grouper that authored this row, and is
# the ONE flag the UI keys off to say so. The honesty rule is the whole
# point: a synthetic post must never present itself as authored, and a
# column that is absent-or-a-name makes "was this us?" answerable from the
# row rather than inferred from its shape.
#
# Plain String, no CHECK (rule 36 considered and declined) — same reasoning
# as source.error_type and service_seen.kind. There is exactly one grouper
# today; a second would be a value, not an invariant.
synthesized_by: Mapped[str | None] = mapped_column(String(32), nullable=True)
# What it was built from, so the operator can audit a grouping FC invented:
# member post ids, message count, and the thresholds in force when the
# decision was made. That last part matters — the thresholds are operator-
# tunable, so "why did it group these" is unanswerable a month later
# without recording the values that produced it.
synthesis_details: Mapped[dict | None] = mapped_column(JSON, nullable=True)
# Set on a MEMBER post, pointing at the synthetic post that absorbed it.
# The feed hides absorbed posts (they are the chat lines the synthetic post
# replaced); every other surface still reaches them by id, because they
# remain the image's true origin and the grouping has to be inspectable.
#
# Self-FK, ON DELETE SET NULL: deleting a synthetic post un-absorbs its
# members and they return to the feed on their own. That is the reversal
# path, and it is one DELETE — nothing to undo by hand.
absorbed_by_post_id: Mapped[int | None] = mapped_column(
ForeignKey("post.id", ondelete="SET NULL"), nullable=True, index=True
)
# -- An OPEN grouping (milestone 388 E3) -------------------------------
# A synthetic post is not sealed at creation: a creator who adds two more
# variants the next day extends the existing post rather than starting a
# new one. These two columns are what make that possible without the post
# either freezing or thrashing the feed.
#
# `last_grew_at` is when the group last absorbed something. It answers two
# questions: how long the group stays JOINABLE (a group closes after a
# quiet period — artists reuse characters for years, and a group left open
# forever will eventually absorb something it shouldn't), and what the card
# shows as "updated N ago". NULL means it has never grown since creation.
last_grew_at: Mapped[datetime | None] = mapped_column(
DateTime(timezone=True), nullable=True
)
# The feed position, and ONLY set when the anti-thrash rule fires — see
# discord_grouping.should_resurface. A group that gains one image a day
# must not sit permanently at the top of the feed, so growth updates the
# post without necessarily moving it; a genuine second wave moves it once.
#
# NULL on every ordinary post, which is why the feed's sort key can
# COALESCE through it without changing where anything else lands.
resurfaced_at: Mapped[datetime | None] = mapped_column(
DateTime(timezone=True), nullable=True
)
+100
View File
@@ -0,0 +1,100 @@
"""PostAssociation — "this Patreon post announced that Discord drop".
Milestone 388, step E5, and the point of the milestone rather than its tail.
Two of the operator's artists post a deliberately CROPPED fragment on Patreon
to signal that the real thing has landed in their Discord. The Patreon post is
the announcement; the Discord grouping (milestone 388 E2) is the payload. This
row is the link between them.
## Directional, and NOT a merge
`announcement` → `payload` is asymmetric on purpose. The teaser announces the
drop; the drop does not announce the teaser, and a symmetric "related posts"
edge would lose the only thing that makes the pair interesting.
Nor are the two collapsed into one post. The creator published twice,
deliberately, on two platforms with different audiences — flattening that
hides the very behaviour being modelled, and would destroy the operator's
ability to see that the Patreon post is a teaser at all.
## Confirm-only, following FC-6.3 (task 737)
`status` starts at `pending` and nothing is linked until the operator accepts.
A wrongly-asserted association tells them two different pieces are one, which
is worse than no link: no link leaves them where they already are, a wrong one
actively misinforms. Same reason the series matcher writes to a review queue
instead of filing posts on its own.
`status` is a plain String, no CHECK — matching `series_suggestion.status`,
which records the same check-existing-enums lesson.
## Why pHash could not do this, and the correction matters
The original plan claimed this link was already sitting in `image_provenance`
via pHash dedup. It is not. `compute_phash` is `imagehash.phash` at
`hash_size=8` — a DCT hash over the WHOLE image, robust to rescaling and
recompression but NOT to cropping, because a crop changes the global
signature. Cross-platform provenance still links straight re-posts; it does
nothing for a cropped teaser and its full version, which is precisely the pair
the operator described. Hence a scored proposal rather than a lookup.
"""
from datetime import datetime
from sqlalchemy import (
JSON,
DateTime,
Float,
ForeignKey,
Integer,
String,
UniqueConstraint,
func,
)
from sqlalchemy.orm import Mapped, mapped_column
from .base import Base
class PostAssociation(Base):
__tablename__ = "post_association"
__table_args__ = (
UniqueConstraint(
"announcement_post_id", "payload_post_id",
name="uq_post_association_pair",
),
)
id: Mapped[int] = mapped_column(Integer, primary_key=True)
# The teaser — a real post the creator wrote (Patreon, today).
announcement_post_id: Mapped[int] = mapped_column(
ForeignKey("post.id", ondelete="CASCADE"), nullable=False, index=True
)
# What it announced — a synthetic Discord grouping, today. CASCADE on both
# sides: an association to a post that no longer exists is not a fact worth
# keeping, and E3's reversal path (delete the grouping) should not leave a
# dangling proposal behind.
payload_post_id: Mapped[int] = mapped_column(
ForeignKey("post.id", ondelete="CASCADE"), nullable=False, index=True
)
score: Mapped[float] = mapped_column(Float, nullable=False)
# Per-signal strengths as scored, so a proposal stays explicable after the
# weights or the threshold are tuned. Without it, "why was this suggested"
# is unanswerable the moment anything moves.
signals: Mapped[dict | None] = mapped_column(JSON, nullable=True)
# pending | linked | dismissed. A DISMISSED row is kept, not deleted — it
# is what stops the matcher proposing the same rejected pair on every
# subsequent scan, which is the behaviour that makes a review queue
# unusable.
status: Mapped[str] = mapped_column(
String(16), nullable=False, server_default="pending", index=True
)
created_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now()
)
updated_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False,
server_default=func.now(), onupdate=func.now(),
)
+88
View File
@@ -0,0 +1,88 @@
"""service_seen — the learned roster of FabledCurator's own moving parts.
Nothing else in this application knows what is SUPPOSED to be running.
`celery inspect` reports the workers that answer, so a stopped worker is a
shorter list rather than a red light, and Postgres and Redis have no
representation at all. That is why the only place an operator could see a
dead service was Portainer, which knows the intended set (milestone 365).
This table is the memory that makes an absence observable: every part that
has ever checked in, and when it last did. A row that stops advancing is a
part that stopped.
## Why the key is not the hostname
`_read_workers_sync()` returns celery's worker names, which here are
`celery@<container id>`. Those are minted fresh on every deploy. Keyed on
them, this table would record a death and a birth every time the stack is
updated — and a status page that goes red on every deploy is a status page
nobody reads, which is worse than not having one.
So a celery role is keyed on its **queue set**, which is assigned per role in
docker-compose.yml (`CELERY_QUEUES`) and survives container replacement:
default,import,thumbnail,download -> worker
maintenance,scan -> scheduler (celery worker --beat)
ml -> ml-worker
Two replicas of one role share a queue set and are therefore ONE row — which
is right, because the question being answered is "is that role being served",
not "how many containers exist". The replica count and their hostnames go in
`details`, where they can change without the identity changing.
The GPU agent is keyed on its `agent_id`, the identity its lease protocol
already uses (`api/gpu.py`).
## What is NOT in here
Postgres and Redis. They are always expected and never learned, and a
last-seen for them would be actively misleading — that one answered thirty
seconds ago says nothing about now. They are probed live at request time.
## kind
Plain `String`, not a Postgres ENUM and not CHECK-gated, matching
`gpu_job.status` and `backup_run.status`. The value set here is expected to
grow as parts are added, and a constraint swap per new kind (rule 36) would
be cost with no invariant behind it.
celery — a worker role, keyed on its queue set
agent — a GPU agent, keyed on its agent_id
"""
from datetime import datetime
from sqlalchemy import JSON, DateTime, String, func
from sqlalchemy.orm import Mapped, mapped_column
from .base import Base
class ServiceSeen(Base):
__tablename__ = "service_seen"
# No indexes beyond the primary key, deliberately. This table holds one row
# per moving part — a handful, forever — so every query against it is a
# full read of a few rows and an index would be write cost buying nothing
# (the lesson of #3301, which removed seven redundant ones).
key: Mapped[str] = mapped_column(String(128), primary_key=True)
kind: Mapped[str] = mapped_column(String(16), nullable=False)
# What to call it in the UI. Derived from the queue set where it is
# recognised, and falling back to the raw queue list where it is not — a
# deployment that slices its queues differently should still show something
# true rather than a name this code invented for it.
display_name: Mapped[str] = mapped_column(String(64), nullable=False)
first_seen_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now(),
)
last_seen_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now(),
)
# The parts that change without changing identity: replica hostnames,
# active task counts, the queues actually being served. Kept as a blob
# because it is displayed and never queried — giving it columns would
# invite filtering on it, which is what the activity endpoints are for.
details: Mapped[dict] = mapped_column(JSON, nullable=False, default=dict)
@@ -0,0 +1,337 @@
"""Proposing that a creator FC tracks and a membership it found are the same.
Milestone 388, step E4. An instance of the confirm-only matcher shape
(snippet #3842), and a sibling of `post_association_service`.
## What E4 turned out NOT to need
The step's own first instruction was to verify before building, and the
verification said: not the schema, not the flows. `Source.artist_id` is a plain
FK so many sources per artist already works; `POST /api/sources` already takes
an `artist_id`; the add-source dialog already has an artist autocomplete that
attaches to an EXISTING artist; `SourceService.reassign` already moves a source
between artists WITH post and image re-attribution; and a sweep for
one-source-per-artist assumptions found only `func.count()` calls, which are
the opposite of assuming one.
So the association a Discord source and a Patreon source share is already
expressible today. What was missing is FC OFFERING it.
## Accept adds a SOURCE — it never merges artists
The asymmetry that sets the whole posture: adding a source is trivially undone.
A wrong artist merge silently mixes two creators' work and corrupts tagging,
series and provenance downstream, with nothing left to tell the two apart by.
So the accepted action is "add the missing channel to this artist", and merging
is not offered at all.
## The signals
1. **Name.** The roster's `display_name` and `vanity`, slugified, against the
artist's `slug`. Graded rather than boolean — an exact match is strong
evidence, a containment match is a hint.
2. **Declared.** A post already under this artist whose body links to
`patreon.com/<vanity>` for this exact membership. A creator pointing at
their own Patreon from their own Discord is close to a statement.
Signal 2 is NOT read from `ExternalLink`, and that correction is worth keeping:
`link_extract.SUPPORTED_HOSTS` is file hosts only (mega/gdrive/mediafire/
dropbox/pixeldrain) and `host_for()` returns None for patreon.com, so no
`ExternalLink` row is ever written for one. The same trap already caught E5 for
Discord invites.
## Weights, and what they make impossible
name 0.65 · declared 0.35, cut at 0.60
Chosen so the arithmetic encodes the judgement rather than a code path doing it:
* an EXACT name match alone (0.65) proposes — same slug on both sides is
strong, and requiring corroboration would mean proposing almost nothing;
* a CONTAINMENT name match alone (0.6 * 0.65 = 0.39) does not — "art" inside
"artgirl" is a coincidence generator, and it needs the declaration;
* the declaration ALONE (0.35) never proposes, at any setting at or above
0.60 — a creator may link another creator's Patreon, and a link is not a
claim of identity.
A guard test pins all three against WEIGHTS directly, so they survive a
refactor of the scorer.
"""
from __future__ import annotations
import logging
import re
from sqlalchemy import func, select
from sqlalchemy.ext.asyncio import AsyncSession
from ..models import (
Artist,
ArtistMembershipSuggestion,
PlatformMembership,
Post,
Source,
)
from ..utils.slug import slugify
from ..utils.text import html_to_plain
from .membership_roster import source_for_membership
log = logging.getLogger(__name__)
WEIGHTS = {"name": 0.65, "declared": 0.35}
DEFAULT_THRESHOLD = 0.60
NAME_EXACT = 1.0
# Containment is a hint, not a match: "art" sits inside "artgirl", and slugs
# are short enough that coincidental containment is common.
NAME_CONTAINS = 0.6
# Below this many characters, containment is noise rather than signal — a
# 3-character slug is inside a great many longer ones.
_MIN_CONTAINMENT_LEN = 5
MAX_CANDIDATES = 25
def name_signal(membership: PlatformMembership, artist: Artist) -> float:
"""Graded slug agreement between a membership and an artist.
Both the display name and the vanity are tried, because creators routinely
differ between the two ("Team Melon Collie" vs "MelonCollieStudios") and
either may be the one the operator typed when they created the artist.
"""
artist_slug = slugify(artist.name or "") if artist.name else ""
if not artist_slug or artist_slug == "untitled":
return 0.0
candidates = {
slugify(v) for v in (membership.display_name, membership.vanity_or_none())
if v
}
candidates.discard("untitled")
if not candidates:
return 0.0
if artist_slug in candidates:
return NAME_EXACT
for c in candidates:
if len(c) < _MIN_CONTAINMENT_LEN or len(artist_slug) < _MIN_CONTAINMENT_LEN:
continue
if c in artist_slug or artist_slug in c:
return NAME_CONTAINS
return 0.0
def declared_signal(body: str | None, vanity: str | None) -> float:
"""Does this post body point at THIS membership's Patreon page?
Matched against the RAW body, not the stripped text: these links live in an
anchor's `href`, and `html_to_plain` discards attributes — the same trap
that caught E5's invite detection. The stripped text is checked too, for
bodies that paste the URL as plain text.
"""
if not body or not vanity:
return 0.0
pattern = re.compile(
r"patreon\.com/(?:c/|cw/|checkout/)?" + re.escape(vanity) + r"\b", re.I
)
if pattern.search(body):
return 1.0
return 1.0 if pattern.search(html_to_plain(body) or "") else 0.0
def weighted_score(signals: dict) -> float:
return round(sum(WEIGHTS[k] * signals.get(k, 0.0) for k in WEIGHTS), 4)
class ArtistMembershipService:
def __init__(self, session: AsyncSession):
self.session = session
async def _decided(self, membership_id: int) -> set[int]:
"""Artists already proposed for this membership, in ANY status.
Dismissed included: the row is what remembers the rejection, and
re-proposing a rejected pair every scan is what makes a queue ignored.
"""
rows = (await self.session.execute(
select(ArtistMembershipSuggestion.artist_id).where(
ArtistMembershipSuggestion.platform_membership_id == membership_id
)
)).scalars().all()
return set(rows)
async def _candidate_artists(self, membership: PlatformMembership) -> list[Artist]:
"""Artists that have SOME source but none for this membership's platform.
A hard filter, not a scored signal. An artist FC already tracks on this
platform needs no suggestion — the link exists — and an artist with no
sources at all is not a creator FC is following through another channel,
which is the whole case this step is about.
"""
# `select(...).exists()` rather than a bare `exists().where(...)`: the
# latter has no FROM to correlate against and does not reliably render.
has_any = select(Source.id).where(Source.artist_id == Artist.id).exists()
has_this = (
select(Source.id)
.where(
Source.artist_id == Artist.id,
Source.platform == membership.platform,
)
.exists()
)
return (await self.session.execute(
select(Artist).where(has_any, ~has_this).limit(MAX_CANDIDATES)
)).scalars().all()
async def _declared_for(self, artist_id: int, vanity: str | None) -> float:
if not vanity:
return 0.0
# Bounded scan: the newest posts are where a creator's current links
# live, and an unbounded body scan per (artist, membership) pair would
# be the expensive part of this sweep.
bodies = (await self.session.execute(
select(Post.description)
.where(Post.artist_id == artist_id, Post.description.is_not(None))
.order_by(func.coalesce(Post.post_date, Post.downloaded_at).desc())
.limit(50)
)).scalars().all()
for body in bodies:
if declared_signal(body, vanity) > 0:
return 1.0
return 0.0
async def match_membership(
self, membership_id: int, *, threshold: float = DEFAULT_THRESHOLD,
) -> int:
membership = await self.session.get(PlatformMembership, membership_id)
if membership is None:
return 0
# The shared identity join (C4), used here as the NEGATIVE check. A
# membership FC already has a source for is tracked — whoever it happens
# to be filed under — and proposing it to some OTHER artist would be
# exactly the wrong link this service exists to avoid making.
# `_candidate_artists` only knows whether a GIVEN artist has a source on
# the platform, which cannot see a source sitting under someone else.
if await source_for_membership(self.session, membership) is not None:
return 0
already = await self._decided(membership_id)
made = 0
for artist in await self._candidate_artists(membership):
if artist.id in already:
continue
signals = {
"name": name_signal(membership, artist),
"declared": await self._declared_for(
artist.id, membership.vanity_or_none()
),
}
score = weighted_score(signals)
if score < threshold:
continue
self.session.add(ArtistMembershipSuggestion(
platform_membership_id=membership.id,
artist_id=artist.id,
score=score,
signals=signals,
status="pending",
))
made += 1
return made
async def list_pending(self) -> list[dict]:
rows = (await self.session.execute(
select(ArtistMembershipSuggestion, PlatformMembership, Artist)
.join(
PlatformMembership,
PlatformMembership.id
== ArtistMembershipSuggestion.platform_membership_id,
)
.join(Artist, Artist.id == ArtistMembershipSuggestion.artist_id)
.where(ArtistMembershipSuggestion.status == "pending")
.order_by(
ArtistMembershipSuggestion.score.desc(),
ArtistMembershipSuggestion.id.desc(),
)
)).all()
return [
{
"id": s.id,
"score": s.score,
"signals": s.signals,
"artist": {"id": a.id, "name": a.name, "slug": a.slug},
"membership": {
"id": m.id,
"platform": m.platform,
"display_name": m.display_name,
"url": m.url,
},
}
for s, m, a in rows
]
async def accept(self, suggestion_id: int) -> dict | None:
"""Add the missing channel to the artist. NEVER merges two artists.
Returns the created source's id, or `already_linked` when a source for
that platform appeared between the proposal and the click — which is
not an error, it is the operator having done it by hand.
"""
s = await self.session.get(ArtistMembershipSuggestion, suggestion_id)
if s is None:
return None
membership = await self.session.get(PlatformMembership, s.platform_membership_id)
if membership is None or not membership.url:
return None
existing = (await self.session.execute(
select(Source.id).where(
Source.artist_id == s.artist_id,
Source.platform == membership.platform,
)
)).scalars().first()
if existing is not None:
s.status = "linked"
return {"id": s.id, "status": s.status, "already_linked": existing}
# Through SourceService, NOT a bare Source() insert. It carries the
# platform/URL validation, the duplicate check and the #693
# backfill-arming that a hand-added source gets — building a second,
# quieter way to create a source is how the two drift until one of them
# is subtly broken (rule 28: repurpose the existing surface).
from .source_service import DuplicateSourceError, SourceService
try:
record = await SourceService(self.session).create(
artist_id=s.artist_id,
platform=membership.platform,
url=membership.url,
)
except DuplicateSourceError as exc:
# The same URL already exists for this artist — the operator got
# there first by a different route. Not an error.
s.status = "linked"
return {"id": s.id, "status": s.status, "already_linked": exc.existing_id}
s.status = "linked"
return {"id": s.id, "status": s.status, "source_id": record.id}
async def dismiss(self, suggestion_id: int) -> dict | None:
s = await self.session.get(ArtistMembershipSuggestion, suggestion_id)
if s is None:
return None
# Kept, not deleted — the row is what remembers the rejection.
s.status = "dismissed"
return {"id": s.id, "status": s.status}
async def rescan(session: AsyncSession, *, threshold: float = DEFAULT_THRESHOLD) -> dict:
"""Offer every known membership to the artists FC already tracks."""
ids = (await session.execute(select(PlatformMembership.id))).scalars().all()
svc = ArtistMembershipService(session)
proposed = 0
for mid in ids:
proposed += await svc.match_membership(mid, threshold=threshold)
log.info(
"artist/membership matcher: scanned %d membership(s), proposed %d pair(s)",
len(ids), proposed,
)
return {"scanned": len(ids), "proposed": proposed}
+31
View File
@@ -20,6 +20,9 @@ from sqlalchemy import Select
from sqlalchemy.exc import IntegrityError
from sqlalchemy.ext.asyncio import AsyncSession
from ..models import Source
from .gallery_dl import ErrorType
async def get_or_create[T](
session: AsyncSession,
@@ -50,3 +53,31 @@ async def get_or_create[T](
except IntegrityError:
await sp.rollback()
return (await session.execute(select_stmt)).scalar_one(), False
# --- shared Source health predicates ----------------------------------------
#
# The subscriptions rollup, the front-door status ribbon and the list endpoint
# all have to agree on what "failing" and "no access" MEAN, or the ribbon says
# 3 and the card it links to shows 4. Same reasoning as get_or_create above:
# divergent copies of one predicate are how the drift creeps in. Defined here
# rather than in source_service because scheduler_service needs them too, and
# source_service already imports scheduler_service (the other direction would
# be a cycle).
def failing_sources_clause():
"""A source is FAILING when its runs are actually erroring.
Deliberately not `last_error IS NOT NULL` — a tier-limited source clears
last_error and keeps a chip, and must never be counted as broken.
"""
return Source.consecutive_failures > 0
def no_access_sources_clause():
"""A source we can't see the content of: the walk works, the tier doesn't
grant it (#874 / milestone #387 phase A). Not a failure — kept separate
from failing_sources_clause on purpose, and the two are disjoint because
an informational class only ever rides an otherwise-OK run."""
return Source.error_type == ErrorType.TIER_LIMITED
+655
View File
@@ -0,0 +1,655 @@
"""Discord drop grouping — FC authors the post that Discord never wrote.
Milestone 388, step E2.
Discord is a delivery CHANNEL, not a publisher. A creator drops a set of
near-variants — the same piece with different hair colour, accessories, an
outfit swap — across a handful of messages, and today each of those messages
lands as its own `post` row, so chat lines compete with authored work for the
same surface. The fix is not to demote them into a second-class feed; it is to
let FC write the post: one row per DROP, its images the drop's images, its body
the messages' text in arrival order.
The result is post-shaped by construction, which is the entire reason to
synthesise a `Post` rather than invent a parallel entity — feed, provenance,
translation, attachments and series all keep working on it unchanged.
## The predicate: three axes, ANDed, and the time one does the real work
**Similarity alone over-groups, and that is the failure that would make this
useless.** Any two pieces of the same character by the same artist sit close in
SigLIP space; a cosine-only rule collapses a month of one character into a
single "post". What makes a variant set a set is that it was dropped TOGETHER.
same source AND cosine distance <= threshold AND no gap > window
Two details in there are load-bearing:
* **Distance is measured to the group's SEED, never to the previous member.**
Chaining to the previous member lets a group DRIFT: twenty small steps walk
from one piece to a completely different one, each hop individually within
threshold. Anchoring on the seed bounds the whole group to one neighbourhood.
* **The window is measured between CONSECUTIVE messages, not from the first.**
An artist trickling variants out over an evening is one drop; a window
anchored on the first message would cut it in half at an arbitrary point.
## Why this is a post-import sweep and not part of ingest
The obvious alternative was to migrate Discord to the native post-first
ingester (#1266) and group at capture time. **That cannot work**, and the
reason is worth recording: the grouping signal is `siglip_embedding`, which is
produced ASYNCHRONOUSLY after import (`tasks/ml.py`, the GPU queue backfill).
At capture time the embedding does not exist yet, so an ingester has nothing to
group on. Grouping is necessarily something that happens once the vectors have
caught up — which also means this sweep must be re-runnable and must simply
skip what it cannot yet place. It does: a post whose image has no embedding is
left alone and picked up on a later run.
## The honesty rule
A synthetic post must never pretend an artist authored it. It carries
`synthesized_by`, records what it was built from in `synthesis_details`
(members, count, and the thresholds in force at the time), and leaves its
member posts intact and reachable. Deleting the synthetic post releases the
members back into the feed — one DELETE, no repair step. FC invented this
grouping; the operator has to be able to see that, inspect it, and undo it.
"""
from __future__ import annotations
import logging
import math
from dataclasses import dataclass, field
from datetime import UTC, datetime, timedelta
from sqlalchemy import Select, func, select, update
from sqlalchemy.dialects.postgresql import insert as pg_insert
from sqlalchemy.ext.asyncio import AsyncSession
from ..models import ImageProvenance, ImageRecord, MLSettings, Post, Source
log = logging.getLogger(__name__)
# The value that lands in `post.synthesized_by`. One grouper today; a second
# would be another value here, which is exactly why the column has no CHECK.
DROP_GROUPER = "discord_drop"
PLATFORM = "discord"
# Ceiling on member posts examined per source per run. A first sweep over an
# established library would otherwise pull every Discord message's 1152-float
# vector into memory at once. The sweep is re-runnable and works oldest-first,
# so a backlog simply drains over successive runs rather than needing one
# heroic pass.
MAX_CANDIDATES_PER_SOURCE = 500
@dataclass
class DropGroup:
"""One drop: the member posts, in arrival order, that will become a post."""
member_ids: list[int] = field(default_factory=list)
seed: list[float] | None = None
last_at: datetime | None = None
def cosine_distance(a, b) -> float:
"""Cosine distance between two embeddings, in the same units pgvector's
`cosine_distance` operator returns (0 = identical, 1 = orthogonal).
Computed in Python rather than SQL because the comparison is against a
group seed held in a loop, not against a column — and pure arithmetic keeps
numpy off this path entirely. pgvector may hand back a numpy array or a
list depending on driver version, so both are coerced.
"""
va = [float(x) for x in a]
vb = [float(x) for x in b]
# strict=True: two embeddings of different length is a corrupted row or a
# model swap that skipped the re-embed, and silently truncating to the
# shorter one would score it as a near match.
dot = sum(x * y for x, y in zip(va, vb, strict=True))
na = math.sqrt(sum(x * x for x in va))
nb = math.sqrt(sum(y * y for y in vb))
if na == 0.0 or nb == 0.0:
# A zero vector has no direction, so no meaningful distance. Return the
# maximum so it can never pull anything into a group.
return 1.0
return 1.0 - (dot / (na * nb))
def _candidate_stmt(source_id: int, *, not_after: datetime) -> Select:
"""Ungrouped Discord message-posts, one representative image each, OLDEST
FIRST — which is the order `build_groups` requires.
DISTINCT ON the post picks the lowest-id embedded image as that post's
representative: a Discord message carrying several attachments is still one
point in the drop, and comparing every attachment would let one incidental
image drag an unrelated message into the group.
The DISTINCT ON is wrapped in a subquery rather than ordered directly,
because Postgres requires a DISTINCT ON query's ORDER BY to LEAD with the
distinct expression — so the inner query must sort by `post.id`, which is
insertion order and not arrival order at all once a backfill has imported
anything out of sequence. Sorting outside is what makes the caller's LIMIT
take the OLDEST candidates instead of the lowest-numbered ones.
"""
sort_key = func.coalesce(Post.post_date, Post.downloaded_at)
inner = (
select(
Post.id.label("post_id"),
sort_key.label("occurred_at"),
ImageRecord.siglip_embedding.label("embedding"),
)
.join(ImageRecord, ImageRecord.primary_post_id == Post.id)
.where(
Post.source_id == source_id,
# Never absorb a post FC wrote, and never re-absorb one already
# taken — both would build groups out of groups.
Post.synthesized_by.is_(None),
Post.absorbed_by_post_id.is_(None),
ImageRecord.siglip_embedding.is_not(None),
sort_key <= not_after,
)
.distinct(Post.id)
.order_by(Post.id, ImageRecord.id)
.subquery()
)
return (
select(inner.c.post_id, inner.c.occurred_at, inner.c.embedding)
.order_by(inner.c.occurred_at, inner.c.post_id)
)
def build_groups(
rows: list[tuple[int, datetime, list[float]]],
*,
max_distance: float,
window: timedelta,
) -> list[DropGroup]:
"""Walk candidates in arrival order and cut them into drops.
`rows` must be sorted oldest-first — the whole predicate is about
adjacency in time, so an unsorted input would silently produce nonsense
rather than fail.
"""
groups: list[DropGroup] = []
current: DropGroup | None = None
for post_id, occurred_at, embedding in rows:
if current is not None:
gap_ok = occurred_at - current.last_at <= window
# Distance to the SEED, not to the previous member — see the module
# docstring on drift.
near = cosine_distance(current.seed, embedding) <= max_distance
if gap_ok and near:
current.member_ids.append(post_id)
current.last_at = occurred_at
continue
groups.append(current)
current = DropGroup(
member_ids=[post_id], seed=embedding, last_at=occurred_at,
)
if current is not None:
groups.append(current)
return groups
async def _synthesize(
session: AsyncSession,
*,
source: Source,
group: DropGroup,
max_distance: float,
window_minutes: float,
) -> Post | None:
"""Write one synthetic post for `group` and absorb its members."""
members = (await session.execute(
select(Post)
.where(Post.id.in_(group.member_ids))
.order_by(func.coalesce(Post.post_date, Post.downloaded_at), Post.id)
)).scalars().all()
if not members:
return None
first = members[0]
# Deterministic key, so a re-run cannot mint a second post for the same
# drop: the unique (source_id, external_post_id) constraint would reject it
# even if the member filter somehow let the drop through twice.
external_id = f"fc-drop:{first.external_post_id}"[:128]
# The messages' own text, in arrival order, IS the post's body — that is
# what the operator asked for and it is the only text a drop has. Blank
# messages (an attachment with no caption) contribute nothing rather than a
# run of empty lines.
body = "\n\n".join(m.description.strip() for m in members if m.description and m.description.strip())
post = Post(
source_id=source.id,
artist_id=source.artist_id,
external_post_id=external_id,
# post_title stays NULL DELIBERATELY. A synthesised title is the one
# place this feature could accidentally put words in a creator's mouth;
# the UI labels the row from `synthesized_by` instead, which cannot be
# mistaken for something the artist wrote.
post_title=None,
post_url=first.post_url,
post_date=first.post_date or first.downloaded_at,
description=body or None,
synthesized_by=DROP_GROUPER,
synthesis_details={
"member_post_ids": [m.id for m in members],
"message_count": len(members),
# The thresholds AS THEY WERE. They are operator-tunable, so
# without this "why did it group these" is unanswerable later.
"max_distance": max_distance,
"window_minutes": window_minutes,
"grouped_at": datetime.now(UTC).isoformat(),
# Growth accumulated since the post last moved in the feed (E3).
# Seeded here so the joiner never meets its absence — creation IS
# a surfacing, so the count starts at zero.
"images_since_surface": 0,
},
)
session.add(post)
await session.flush()
await session.execute(
update(Post)
.where(Post.id.in_([m.id for m in members]))
.values(absorbed_by_post_id=post.id)
)
await _link_member_images(
session, post_id=post.id, member_ids=[m.id for m in members],
source_id=source.id,
)
return post
async def _link_member_images(
session: AsyncSession, *, post_id: int, member_ids: list[int], source_id: int | None,
) -> int:
"""Attach every member's images to the synthetic post. Returns how many.
The feed and detail views already union provenance with `primary_post_id`
(post_feed_service._thumbnails_for), so this alone makes the drop's images
show up under the post FC wrote — no second render path.
`primary_post_id` is deliberately NOT rewritten: the message post remains
the image's true origin, and the synthetic post is an ADDITIONAL claim on
it, which is what keeps the grouping reversible.
Shared by creation and by the E3 joiner rather than written twice, because
the two would otherwise be free to drift on exactly the detail (which post
owns the image) that makes a grouping reversible.
"""
if not member_ids:
return 0
image_rows = (await session.execute(
select(ImageRecord.id).where(ImageRecord.primary_post_id.in_(member_ids))
)).scalars().all()
if not image_rows:
return 0
await session.execute(
pg_insert(ImageProvenance)
.values([
{"image_record_id": iid, "post_id": post_id, "source_id": source_id}
for iid in image_rows
])
# (image, post) is unique. This is a BACKSTOP against a re-run that
# raced itself, not the correctness argument: callers only ever pass
# members that were unabsorbed a moment ago, so a conflict here means
# concurrency, not a logic error.
.on_conflict_do_nothing(constraint="uq_image_provenance_image_post")
)
return len(image_rows)
async def group_source(
session: AsyncSession,
source: Source,
*,
max_distance: float,
window_minutes: float,
now: datetime | None = None,
) -> int:
"""Group one Discord source's ungrouped messages. Returns posts created."""
window = timedelta(minutes=window_minutes)
now = now or datetime.now(UTC)
# Leave the most recent window alone: a drop that is still arriving would
# otherwise be cut in half by whichever sweep happened to land mid-drop,
# and the second half would become a separate post claiming to be its own
# drop. Waiting one window costs nothing (the sweep re-runs) and is the E2
# side of "keep the grouping open"; E3 handles the harder case where a
# matching drop resumes after the gap has already passed.
rows = (await session.execute(
_candidate_stmt(source.id, not_after=now - window)
.limit(MAX_CANDIDATES_PER_SOURCE)
)).all()
if not rows:
return 0
groups = build_groups(
[(pid, occurred, emb) for pid, occurred, emb in rows],
max_distance=max_distance, window=window,
)
if len(rows) == MAX_CANDIDATES_PER_SOURCE and len(groups) > 1:
# The cap may have fallen INSIDE the last drop, and synthesising a
# truncated group would publish a post that claims to be the whole drop
# while the rest of it sits one row past the limit. Leave it for the
# next run, which starts from the same place and sees the remainder.
# Guarded on len > 1 so a single oversized group is not dropped
# forever — it would make no progress at all.
groups = groups[:-1]
created = 0
for group in groups:
post = await _synthesize(
session, source=source, group=group,
max_distance=max_distance, window_minutes=window_minutes,
)
if post is not None:
created += 1
return created
# ---------------------------------------------------------------------------
# E3: an open grouping — a later drop joins its post and updates it.
# ---------------------------------------------------------------------------
#
# A synthetic post is not sealed at creation. A creator who adds two more
# variants the next day extends the existing post rather than starting a new
# one, and its body grows with the new messages. That is what makes chat
# capture read as content TRICKLING IN rather than as a stream of separate
# arrivals.
#
# Three hard problems, each answered deliberately below: bridging (a candidate
# near two groups), re-surfacing without thrashing the feed, and groups that
# stay open forever.
# How much closer the nearest group must be than the runner-up before a
# candidate is assigned to it at all.
#
# NOT a setting, deliberately. It is not a quality dial the operator would tune
# toward a better feed — it expresses "these two are too close to call", and
# exposing it would invite turning it to zero, which is precisely the silent
# arbitrary choice it exists to prevent. When a candidate is genuinely between
# two groups the recoverable answer is to leave it out and let it start its
# own; the unrecoverable one is to merge, because a merge rewrites history —
# two posts the operator may already have seen become one, and anything
# pointing at the absorbed post dangles.
AMBIGUITY_MARGIN = 0.02
def should_resurface(
*,
images_since_surface: int,
last_surface_at: datetime,
now: datetime,
min_images: int,
cooldown: timedelta,
) -> bool:
"""Has this group grown enough, and waited long enough, to move in the feed?
An updated post SHOULD be visible — that is the point of keeping it open —
but a group gaining one image a day must not sit permanently at the top.
Both conditions have to hold: enough new images that the update is worth an
interruption, and enough time since the last one that a steady drip cannot
chain bumps together.
"""
if images_since_surface < min_images:
return False
return now - last_surface_at >= cooldown
async def _group_seed(session: AsyncSession, post_id: int) -> list[float] | None:
"""The embedding a group is measured against — its FIRST member's image.
Derived rather than stored, and derived by the same definition `build_groups`
used (earliest member, lowest-id embedded image). Storing it at creation
would have meant a backfill for groups already written and two definitions
free to disagree; this way there is one.
"""
sort_key = func.coalesce(Post.post_date, Post.downloaded_at)
return (await session.execute(
select(ImageRecord.siglip_embedding)
.join(Post, ImageRecord.primary_post_id == Post.id)
.where(
Post.absorbed_by_post_id == post_id,
ImageRecord.siglip_embedding.is_not(None),
)
.order_by(sort_key, Post.id, ImageRecord.id)
.limit(1)
)).scalar_one_or_none()
async def open_groups(
session: AsyncSession, source_id: int, *, now: datetime, close_after: timedelta,
) -> list[tuple[Post, list[float]]]:
"""This source's synthetic posts that are still accepting members.
Openness is DERIVED, not stored: a group is open if it grew — or started —
within `close_after`. A group left open forever would eventually absorb
something it shouldn't, because artists reuse characters for years; a
stored `closed_at` would need a sweep to set it and a repair path to ever
change the policy. This way the policy IS the query.
"""
cutoff = now - close_after
posts = (await session.execute(
select(Post).where(
Post.source_id == source_id,
Post.synthesized_by == DROP_GROUPER,
# A synthetic post that was itself absorbed is not a thing today
# (nothing absorbs one), but joining into one would nest groups.
Post.absorbed_by_post_id.is_(None),
func.coalesce(Post.last_grew_at, Post.post_date, Post.downloaded_at) >= cutoff,
)
)).scalars().all()
out: list[tuple[Post, list[float]]] = []
for post in posts:
seed = await _group_seed(session, post.id)
if seed is not None:
out.append((post, seed))
return out
def assign_to_group(
embedding: list[float],
groups: list[tuple[Post, list[float]]],
*,
max_distance: float,
) -> Post | None:
"""Pick the one group this image belongs to, or None to leave it alone.
Returns None in two cases that mean different things and are deliberately
treated the same: nothing is close enough (so E2 will start a new group
from it), or two groups are BOTH close and too near each other to choose
between (so E2 will start a new group from it). The second is the bridging
case, and letting it start its own post is the recoverable failure —
merging two existing posts is not.
"""
# key= on the distance ALONE. Sorting bare tuples falls through to the
# second element when two distances tie, and `Post` has no ordering — so a
# perfectly symmetric bridge (the exact case this function exists for)
# would raise TypeError instead of declining to choose.
scored = sorted(
((cosine_distance(seed, embedding), post) for post, seed in groups),
key=lambda pair: pair[0],
)
within = [(d, p) for d, p in scored if d <= max_distance]
if not within:
return None
if len(within) >= 2 and (within[1][0] - within[0][0]) < AMBIGUITY_MARGIN:
return None
return within[0][1]
async def _absorb_into(
session: AsyncSession,
*,
group: Post,
member_ids: list[int],
source_id: int | None,
now: datetime,
min_images: int,
cooldown: timedelta,
) -> int:
"""Extend an existing synthetic post with new members. Returns images added."""
members = (await session.execute(
select(Post)
.where(Post.id.in_(member_ids))
.order_by(func.coalesce(Post.post_date, Post.downloaded_at), Post.id)
)).scalars().all()
if not members:
return 0
added_images = await _link_member_images(
session, post_id=group.id, member_ids=[m.id for m in members],
source_id=source_id,
)
await session.execute(
update(Post)
.where(Post.id.in_([m.id for m in members]))
.values(absorbed_by_post_id=group.id)
)
# The new messages' text joins the body, in arrival order, exactly as at
# creation — the post's body is the drop's text and the drop just grew.
new_text = "\n\n".join(
m.description.strip() for m in members if m.description and m.description.strip()
)
if new_text:
group.description = f"{group.description}\n\n{new_text}" if group.description else new_text
details = dict(group.synthesis_details or {})
existing_ids = list(details.get("member_post_ids") or [])
details["member_post_ids"] = existing_ids + [
m.id for m in members if m.id not in existing_ids
]
details["message_count"] = len(details["member_post_ids"])
since = int(details.get("images_since_surface") or 0) + added_images
details["last_grew_at"] = now.isoformat()
# The feed position moves only when the anti-thrash rule fires. Measured
# from the last time the post actually MOVED (resurfaced_at), falling back
# to when the drop started — creation is itself a surfacing.
last_surface = group.resurfaced_at or group.post_date or group.downloaded_at
if should_resurface(
images_since_surface=since, last_surface_at=last_surface, now=now,
min_images=min_images, cooldown=cooldown,
):
group.resurfaced_at = now
since = 0
details["images_since_surface"] = since
group.synthesis_details = details
group.last_grew_at = now
return added_images
async def join_open_groups(
session: AsyncSession,
source: Source,
*,
max_distance: float,
window_minutes: float,
close_after_hours: float,
resurface_min_images: int,
resurface_cooldown_hours: float,
now: datetime | None = None,
) -> int:
"""Offer this source's ungrouped messages to its open groups.
Runs BEFORE `group_source` in the sweep: a message that belongs to an
existing drop must join it rather than found a rival post, and whichever
runs first wins that message.
"""
now = now or datetime.now(UTC)
window = timedelta(minutes=window_minutes)
groups = await open_groups(
session, source.id, now=now, close_after=timedelta(hours=close_after_hours),
)
if not groups:
return 0
# Same quarantine as E2: a drop still arriving is left for the next run.
rows = (await session.execute(
_candidate_stmt(source.id, not_after=now - window)
.limit(MAX_CANDIDATES_PER_SOURCE)
)).all()
if not rows:
return 0
claimed: dict[int, list[int]] = {}
for post_id, _occurred_at, embedding in rows:
target = assign_to_group(embedding, groups, max_distance=max_distance)
if target is not None:
claimed.setdefault(target.id, []).append(post_id)
by_id = {post.id: post for post, _seed in groups}
joined = 0
for group_id, member_ids in claimed.items():
joined += await _absorb_into(
session, group=by_id[group_id], member_ids=member_ids,
source_id=source.id, now=now,
min_images=resurface_min_images,
cooldown=timedelta(hours=resurface_cooldown_hours),
)
return joined
async def sweep(session: AsyncSession, *, now: datetime | None = None) -> dict:
"""Group every enabled Discord source. No-op when the switch is off.
Two passes per source, and the ORDER is load-bearing: offer new messages to
the groups that are still open (E3) BEFORE founding new ones (E2). Whichever
runs first claims a message, and a variant that belongs to yesterday's drop
must extend that post rather than found a rival to it.
"""
settings = await MLSettings.load(session)
if not settings.discord_grouping_enabled:
return {
"enabled": False, "sources": 0, "posts_created": 0, "images_joined": 0,
}
sources = (await session.execute(
select(Source).where(
Source.platform == PLATFORM,
Source.enabled.is_(True),
)
)).scalars().all()
max_distance = float(settings.discord_group_max_distance)
window_minutes = float(settings.discord_group_window_minutes)
created = 0
joined = 0
for source in sources:
joined += await join_open_groups(
session, source,
max_distance=max_distance,
window_minutes=window_minutes,
close_after_hours=float(settings.discord_group_close_after_hours),
resurface_min_images=int(settings.discord_group_resurface_min_images),
resurface_cooldown_hours=float(
settings.discord_group_resurface_cooldown_hours
),
now=now,
)
created += await group_source(
session, source,
max_distance=max_distance,
window_minutes=window_minutes,
now=now,
)
log.info(
"discord drop grouping: %d source(s), %d synthetic post(s) created, "
"%d image(s) joined to open groups",
len(sources), created, joined,
)
return {
"enabled": True, "sources": len(sources),
"posts_created": created, "images_joined": joined,
}
+34 -2
View File
@@ -28,12 +28,32 @@ from .patreon_ingester import PatreonIngester
from .patreon_resolver import extract_vanity, resolve_campaign_id_for_source
from .pixiv_client import user_id_from_url
from .pixiv_ingester import PixivIngester
from .platforms import known_platform_keys
from .subscribestar_ingester import SubscribeStarIngester
# Platforms whose download + verify go through the native ingester rather than
# gallery-dl. gallery-dl still serves the rest (hentaifoundry, discord) until
# they migrate too.
NATIVE_INGESTER_PLATFORMS = frozenset({"patreon", "subscribestar", "pixiv"})
# they migrate too. pixiv left this set when it was retired (milestone #406).
NATIVE_INGESTER_PLATFORMS = frozenset({"patreon", "subscribestar"})
def _unsupported_platform_message(platform: str) -> str | None:
"""Why `platform` may not be downloaded or verified, or None if it may.
A source can outlive its platform. Retiring one (DeviantArt #3069, pixiv
#406) unregisters it, but its `Source` rows — and the `enabled` flag on
them — are data, and data survives a deploy. So this refuses at the two
functions every download and every credential probe pass through, instead
of trusting the scheduler's `enabled` filter and every future caller to
agree.
Without it a retired platform does not fail: it falls through to the
gallery-dl branch, which is precisely where a platform lands once it is no
longer native — and gallery-dl still has an extractor for it.
"""
if platform in known_platform_keys():
return None
return f"{platform!r} is not a supported platform (retired or unknown)"
# Mirrors patreon_resolver._CAMPAIGNS_URL — surfaced in resolution-failure
# messages so the operator sees the exact lookup endpoint that was hit.
@@ -80,6 +100,13 @@ async def run_download(
backfill state machine and owns phase 3.
"""
platform = ctx["platform"]
refusal = _unsupported_platform_message(platform)
if refusal is not None:
return DownloadResult(
success=False, url=ctx["url"], artist_slug=ctx["artist_slug"],
platform=platform,
error_type=ErrorType.UNSUPPORTED_URL, error_message=refusal,
), None
if uses_native_ingester(platform):
return await _run_native_ingester(
ctx, source_config, mode, gdl, sync_session_factory
@@ -217,6 +244,11 @@ async def verify_source_credential(
network / nothing to test). Callers don't branch on platform — they call
this and render the result.
"""
refusal = _unsupported_platform_message(platform)
if refusal is not None:
# Inconclusive rather than False: nothing was probed, so nothing was
# rejected. False would tell the operator their credential is bad.
return None, refusal
if uses_native_ingester(platform):
# Native ingester platforms verify via their own lightweight auth probe.
# SubscribeStar's probe takes the creator URL directly; Patreon's
+20 -7
View File
@@ -34,7 +34,9 @@ from .gallery_dl import (
GalleryDLService,
SourceConfig,
extract_errors_warnings,
is_informational,
truncate_log,
walk_completed,
)
from .importer import Importer
from .platforms import auth_type_for
@@ -552,11 +554,13 @@ class DownloadService:
# page is no longer double-counted. new_overrides (read fresh above)
# carries the ingester's committed value forward untouched.
completed = (
dl_result.success
and dl_result.error_type is None
and dl_result.return_code == 0
)
# Shared with the result path so the two halves can't disagree about
# what "finished" means. Note it admits an INFORMATIONAL error_type: a
# fully-paywalled creator's backfill really did reach the bottom, and
# treating it as unfinished would re-walk that wall every chunk until
# the stall counter tripped — the creator we can see least becoming the
# one we fetch most.
completed = walk_completed(dl_result)
if completed:
new_overrides["_backfill_state"] = "complete"
new_overrides.pop("_backfill_cursor", None)
@@ -625,8 +629,17 @@ class DownloadService:
if status == "ok":
source.consecutive_failures = 0
source.last_error = None
# alembic 0032 — clear the failure-class chip on success.
source.error_type = None
# alembic 0032 — clear the failure-class chip on success, EXCEPT an
# informational class. tier_limited rides an otherwise-successful
# run: failures stay 0 and last_error stays clear (the run did not
# fail and must not earn a backoff), but "there is content here we
# aren't allowed to see" is a durable fact about the SOURCE, not
# about this run. Clearing it here is what left FailingSourcesCard's
# `tier_limited` palette entry unreachable — the chip was wiped by
# the very success that produced it.
source.error_type = (
error_type if is_informational(error_type) else None
)
elif status == "error":
source.consecutive_failures = (source.consecutive_failures or 0) + 1
source.last_error = error_message
@@ -55,10 +55,6 @@ _PLATFORM_PATTERNS: list[tuple[str, re.Pattern[str]]] = [
r"^https?://(?:www\.)?hentai-foundry\.com/user/(?P<slug>[^/?#]+)",
re.IGNORECASE,
)),
("pixiv", re.compile(
r"^https?://(?:www\.)?pixiv\.net/(?:en/)?users/(?P<slug>\d+)",
re.IGNORECASE,
)),
]
+72 -6
View File
@@ -261,6 +261,72 @@ def make_run_stats(
}
# --- tier-gated classification, shared by BOTH backends ---------------------
#
# These three live together because the native ingester and the gallery-dl
# subprocess must reach the same verdict from the same number. They did not:
# gallery-dl classified TIER_LIMITED while ingest_core counted gated posts and
# threw the count away, so the platforms FC owns reported a paywalled creator as
# a silent one (#874 follow-up). One predicate, spread into both, rather than
# the condition re-derived per backend.
def classify_tier_gated(tier_gated_count: int) -> ErrorType | None:
"""TIER_LIMITED when a walk saw tier-gated posts and nothing else failed.
Deliberately NOT conditioned on `downloaded == 0`. A creator whose top-tier
posts we cannot see is tier-limited even in a week we did get their cheaper
ones — the fact the operator needs ("there is content here you are not
paying for") is true either way. gallery-dl has classified it this way since
the paywall-as-"needs attention" complaint (see `_categorize_error`), and the
native path now matches rather than inventing a stricter rule.
Callers must apply this only AFTER the real error categories (auth, rate
limit, drift, …) have had their turn; tier-gating is the weakest signal and
must never mask a genuine failure.
"""
return ErrorType.TIER_LIMITED if tier_gated_count else None
def tier_gated_message(count: int) -> str:
"""The one wording for the tier-gated verdict, so the two backends can't
describe the same state differently in the Logs UI."""
return (
f"Subscription tier does not grant access to "
f"{count} post{'s' if count != 1 else ''}"
)
# `Source.error_type` doubles as the failure-class chip, and a status of "ok"
# CLEARS it (alembic 0032). TIER_LIMITED breaks that assumption: it rides an
# otherwise-successful run, so without an exemption the chip is wiped the moment
# it is set and `FailingSourcesCard`'s `tier_limited` palette entry can never
# render. Informational classes are the exemption — they describe the source,
# not a failure of the run.
INFORMATIONAL_ERROR_TYPES = frozenset({ErrorType.TIER_LIMITED.value})
def is_informational(error_type) -> bool:
"""True for a class that reports a state rather than a failure. Accepts an
ErrorType or the plain string persisted on Source.error_type."""
return error_type is not None and str(error_type) in INFORMATIONAL_ERROR_TYPES
def walk_completed(result: DownloadResult) -> bool:
"""Did this walk reach the bottom cleanly?
The backfill lifecycle's completion test. An informational error_type still
counts as complete: a fully-paywalled creator's backfill DID finish, and
treating it as unfinished re-walks the same wall until the stall counter
trips — the creator we can see least becoming the one we fetch most.
"""
return (
result.success
and result.return_code == 0
and (result.error_type is None or is_informational(result.error_type))
)
class GalleryDLService:
"""Service for executing gallery-dl downloads."""
@@ -531,12 +597,12 @@ class GalleryDLService:
line for line in combined.split("\n")
if "][warning]" in line and "not allowed to view post" in line
]
if tier_gated_lines:
count = len(tier_gated_lines)
return (
ErrorType.TIER_LIMITED,
f"Subscription tier does not grant access to {count} post{'s' if count != 1 else ''}",
)
# Same predicate + wording the native path uses, so the two backends
# can't drift on what counts as tier-gated or how it reads.
count = len(tier_gated_lines)
gated = classify_tier_gated(count)
if gated is not None:
return (gated, tier_gated_message(count))
# Partial-success: the subprocess exited non-zero (typically because
# the wall-clock timeout fired mid-walk), but it had downloaded ≥1
+42 -10
View File
@@ -35,7 +35,13 @@ from collections.abc import Callable
from sqlalchemy import delete, func, select, text
from sqlalchemy.dialects.postgresql import insert as pg_insert
from .gallery_dl import DownloadResult, ErrorType, make_run_stats
from .gallery_dl import (
DownloadResult,
ErrorType,
classify_tier_gated,
make_run_stats,
tier_gated_message,
)
from .native_ingest_common import NativeAuthError, NativeDriftError
log = logging.getLogger(__name__)
@@ -245,6 +251,12 @@ class Ingester:
per_item_failures=errors,
quarantined_count=quarantined,
dead_lettered_count=dead_lettered,
# #874 follow-up: the native path counted gated posts but
# never reported them, so DownloadDetailModal's "Tier-gated"
# field read 0 on every native walk while gallery-dl's read
# true. A paywalled creator was indistinguishable from a
# silent one.
tier_gated_count=gated_skipped,
),
)
@@ -464,6 +476,10 @@ class Ingester:
"errors": errors,
"quarantined": quarantined,
"posts": posts_processed,
# Ticks during the walk, not only at finalization: a
# deep backfill on a creator we've lost access to is
# otherwise a long run of zeros with no explanation.
"gated": gated_skipped,
})
if early_out:
@@ -564,17 +580,33 @@ class Ingester:
error_type=ErrorType.API_DRIFT, error_message=msg,
)
# Normal success: reached the bottom, or a tick that early-outed. rc 0 +
# error_type None is REQUIRED for a backfill/recovery walk that reached
# the bottom to be marked COMPLETE by
# download_service._apply_backfill_lifecycle — so we return None even
# when downloaded == 0 (a re-confirming walk that found nothing new still
# completed). success=True maps to status "ok" regardless. A tick that
# early-outed also returns here; ticks never set backfill state so the
# lifecycle is a no-op for them.
# Normal success: reached the bottom, or a tick that early-outed. A
# zero-download walk still returns success here — a re-confirming walk
# that found nothing new genuinely completed. A tick that early-outed
# also lands here; ticks never set backfill state so the lifecycle is a
# no-op for them.
#
# success=True and return_code=0 are load-bearing, not cosmetic. They
# are what make this a COMPLETE walk for
# download_service._apply_backfill_lifecycle (via walk_completed) and
# what map it to status "ok", so a walk that fetched nothing doesn't
# accrue consecutive_failures or a backoff it hasn't earned.
#
# #874 follow-up: "nothing new" and "everything sat behind a tier you
# don't hold" are different facts, and returning None for both made a
# paywalled creator indistinguishable from a silent one. TIER_LIMITED is
# classified LAST — every real failure has already returned above —
# because tier-gating is the weakest signal and must never mask a
# genuine error. It is informational, so walk_completed still counts
# this walk as finished (see that predicate for why re-walking a
# paywalled creator forever is the bug being avoided).
gated_error = classify_tier_gated(gated_skipped)
return _result(
success=True, return_code=0,
error_type=None, error_message=None,
error_type=gated_error,
error_message=(
tier_gated_message(gated_skipped) if gated_error else None
),
)
# -- failure mapping (adapter overrides) -------------------------------
@@ -0,0 +1,235 @@
"""Reconciling the learned roster against the sources FC actually tracks.
Milestone 387, step C4. The step the operator asked for; C0-C3 are what make it
trustworthy enough to act on.
## The buckets
1. `subscribed_not_tracked` — you pay for this and FC does not follow it. The
adoption win, and the only bucket carrying an action.
2. `tracked_not_subscribed` — FC follows this and the roster does not show you
paying for it. REPORT ONLY, by the operator's decision (2026-09-11): it says
what it sees and links to the existing Subscriptions row, and offers no
one-click disable.
3. `matched` — the healthy set. Counted, not listed loudly.
4. `unidentified` — sources this join cannot speak to at all. Reported as
exactly that, because the alternative is filing them under a verdict.
## Why absence is the dangerous direction
Bucket 1 is safe to be wrong about: the cost of offering a source the operator
does not want is one ignored row. Bucket 2 is not. It is computed from an
ABSENCE — no membership matched — and three different things produce that
absence: the subscription genuinely lapsed, the sweep failed, or the creator
renamed and this source has never been walked so no exact id was ever cached.
Two guards follow from that, and they are the substance of this module:
* the whole bucket is gated on `roster_is_fresh`, so a failed or never-run sweep
yields an empty list rather than a confident accusation (C3 built the state
this reads);
* every row carries the BASIS for its claim, so "your membership says former
patron" and "we know this creator's id and it is not in your roster" and "we
only have a URL handle to go on" are three different sentences rather than one
overconfident one.
`has_paid_access` returning None is honoured throughout: unknown is never
rendered as lapsed. That is the whole reason it returns a tri-state.
"""
from __future__ import annotations
from datetime import datetime
from sqlalchemy import select
from sqlalchemy.ext.asyncio import AsyncSession
from ..models import Artist, MembershipSync, PlatformMembership, Source
from .membership_roster import (
get_sync_state,
has_paid_access,
identity_keys_for_source,
pair_sources_with_memberships,
roster_is_fresh,
url_tail,
)
# Why a source appears in `tracked_not_subscribed`. Ordered strongest first —
# the UI renders a different sentence per basis, because collapsing them into
# one would make the weakest claim sound like the strongest.
BASIS_LAPSED = "lapsed" # a matched membership says access ended
BASIS_ABSENT_EXACT = "absent_exact" # exact id known, not in a fresh roster
BASIS_ABSENT_HANDLE = "absent_handle" # only a URL handle to go on
def _membership_row(m: PlatformMembership) -> dict:
return {
"id": m.id,
"platform": m.platform,
"external_campaign_id": m.external_campaign_id,
"display_name": m.display_name or m.vanity_or_none(),
"url": m.url,
"vanity": m.vanity_or_none(),
"status": m.status,
"tier_names": m.tier_names,
"amount_cents": m.amount_cents,
"currency": m.currency,
"paid_access": has_paid_access(
m.platform, m.status,
is_free_member=bool((m.details or {}).get("is_free_member")),
),
}
def _source_row(source: Source, artist: Artist) -> dict:
return {
"id": source.id,
"platform": source.platform,
"url": source.url,
"enabled": source.enabled,
"artist": {"id": artist.id, "name": artist.name, "slug": artist.slug},
}
async def reconcile(
session: AsyncSession, *, platform: str, now: datetime | None = None,
) -> dict:
"""Sort one platform's memberships and sources into the four buckets.
Always returns the COMPLETE shape, including when the roster is not fresh —
a caller reading `len(result["tracked_not_subscribed"])` must not have to
check which keys exist first. `fresh` is what says whether the emptiness
means anything.
"""
state = await get_sync_state(session, platform)
fresh = roster_is_fresh(state, now=now)
memberships = (await session.execute(
select(PlatformMembership).where(PlatformMembership.platform == platform)
)).scalars().all()
rows = (await session.execute(
select(Source, Artist)
.join(Artist, Artist.id == Source.artist_id)
.where(Source.platform == platform)
)).all()
# The join itself lives in `membership_roster` beside `match_kind`, so C5's
# gated-reason annotation pairs sources with memberships by exactly the same
# rule this card sorts them by. Two copies would let the Subscriptions row
# and this card disagree about which creator a source IS.
pairs = pair_sources_with_memberships([s for s, _a in rows], memberships)
matched_membership_ids = {m.id for m, _kind in pairs.values()}
subscribed_not_tracked = []
for m in memberships:
if m.id in matched_membership_ids:
continue
paid = has_paid_access(
m.platform, m.status,
is_free_member=bool((m.details or {}).get("is_free_member")),
)
# A membership FC knows has ENDED is not an adoption opportunity —
# adding it would start a walk that can only fetch what is already
# public. Unknown (None) is still offered: the operator can judge it,
# and refusing to show it would hide a real subscription behind a word
# this code has not been taught.
if paid is False:
continue
subscribed_not_tracked.append(_membership_row(m))
tracked_not_subscribed = []
matched = []
unidentified = []
for source, artist in rows:
pair = pairs.get(source.id)
if pair is not None:
m, kind = pair
paid = has_paid_access(
m.platform, m.status,
is_free_member=bool((m.details or {}).get("is_free_member")),
)
if paid is False:
if not source.enabled:
# Already off. Reporting a source the operator has already
# stopped following is noise, not a finding.
continue
row = _source_row(source, artist)
row["basis"] = BASIS_LAPSED
row["matched_by"] = kind
row["membership"] = _membership_row(m)
tracked_not_subscribed.append(row)
else:
row = _source_row(source, artist)
row["matched_by"] = kind
row["membership"] = _membership_row(m)
matched.append(row)
continue
# No membership matched. Whether that MEANS anything depends entirely on
# how well this source can be identified at all.
has_exact = bool(identity_keys_for_source(source))
if not has_exact and url_tail(source.url) is None:
# Nothing to match on — a sidecar anchor or a URL with no handle.
# Reported as unidentified rather than silently dropped, so the
# counts add up to the source list the operator can see.
unidentified.append(_source_row(source, artist))
continue
if not source.enabled:
# Already off. Telling the operator to stop following something they
# have stopped following is noise, not a finding.
continue
row = _source_row(source, artist)
row["basis"] = BASIS_ABSENT_EXACT if has_exact else BASIS_ABSENT_HANDLE
row["matched_by"] = None
row["membership"] = None
tracked_not_subscribed.append(row)
# THE GATE. Everything above computed the bucket; this decides whether it may
# be shown. A stale or never-run roster makes every absence meaningless, and
# an absence rendered as a verdict is how this feature would tell the
# operator to cancel something they are still paying for.
if not fresh:
tracked_not_subscribed = []
return {
"platform": platform,
"fresh": fresh,
# How many sources exist on this platform at all. The UI needs it to
# decide whether an untrustworthy roster is worth mentioning: with no
# sources here there is nothing to reconcile, and a stale-roster warning
# would be noise on an install that simply has not started yet (that
# empty-install case is C6's, not this card's).
"tracked_total": len(rows),
"last_success_at": (
state.last_success_at.isoformat()
if state is not None and state.last_success_at else None
),
"subscribed_not_tracked": subscribed_not_tracked,
"tracked_not_subscribed": tracked_not_subscribed,
"matched": matched,
"unidentified": unidentified,
}
async def reconcile_all(session: AsyncSession, now: datetime | None = None) -> dict:
"""Every platform the roster knows about, in one payload for the UI.
The platform list is the UNION of platforms with memberships and platforms
with sync state, not just the former. A sweep that has never succeeded has
recorded zero memberships, and deriving the list from memberships alone
would drop exactly that platform from the payload — making a broken
credential indistinguishable from a platform FC was never asked about. That
distinction is the whole reason C3 records sync state.
"""
with_memberships = (await session.execute(
select(PlatformMembership.platform).distinct()
)).scalars().all()
with_state = (await session.execute(
select(MembershipSync.platform)
)).scalars().all()
platforms = set(with_memberships) | set(with_state)
return {
"platforms": [
await reconcile(session, platform=p, now=now) for p in sorted(platforms)
]
}
+519
View File
@@ -0,0 +1,519 @@
"""The learned membership roster: what the account actually subscribes to.
Milestone 387, phase C. Sibling of `service_roster` (milestone 365) and built
on the same insight — an absence is only observable against a record of
presence. There, a stopped worker; here, a subscription that lapsed.
## Nothing calls this yet
`touch_membership` is written before its caller because the caller (the sweep,
C3) needs a client seam (C2) that needs Patreon's real response characterised
from a captured sample (C0), and that capture needs the operator's browser
session. The write side does not depend on any of it: an upsert keyed on
(platform, external_campaign_id) is the same regardless of what the payload
turns out to look like, and `details` carries whatever C0 finds.
## Why the whitelist lives here and not in the column
`platform_membership.status` is an unconstrained String holding the PLATFORM's
own word — `active_patron`, not some normalised FC value. The mapping from
those words to FC's meaning is a read-site concern and belongs in code that can
be corrected without a migration, because the vocabulary comes from whatever
each platform says and will be discovered per platform rather than designed up
front. `MEMBERSHIP_STATUS` below is a place for that knowledge to accumulate as
platforms are characterised; it is deliberately empty of guesses today.
"""
from __future__ import annotations
import logging
from collections.abc import Awaitable, Callable
from datetime import UTC, datetime, timedelta
from sqlalchemy import func, select
from sqlalchemy.dialects.postgresql import insert as pg_insert
from sqlalchemy.ext.asyncio import AsyncSession
from ..models import MembershipSync, PlatformMembership, Source
log = logging.getLogger(__name__)
# Platform word -> whether the account currently has paid access.
#
# Every entry here must come from a CHARACTERISED response, never from API docs
# or a plausible guess — project rule 130, and inventing a status before seeing
# it in a real payload is exactly the failure it names.
#
# patreon: from a live capture of the operator's own session, 2026-09-10
# (Scribe note #3886). Only two values were OBSERVED in `patron_status` and
# only those two are here.
#
# `declined_patron` is deliberately ABSENT even though it looks obviously
# right. It appears in the request's `filter[membership_type]`, and the capture
# proved that filter is NOT the same vocabulary as the attribute — a row
# selected by the filter as `free_member` came back with
# `patron_status: former_patron`, a word the filter does not contain. Reading
# the filter as an enum is the specific mistake the capture caught; adding
# `declined_patron` on the strength of it would be repeating that mistake one
# step later.
#
# Unknown words are NOT an error: an unrecognised status means the roster
# records evidence it cannot yet interpret, which is a better state than
# dropping the row or asserting a meaning for it.
#
# subscribestar: from a live capture of the account's /subscriptions page,
# 2026-09-13 (Scribe note #3989). SubscribeStar gives NO per-row status word —
# a membership's state is which of two tables it sits in — so the "word" stored
# is the table card's own `data-identifier`, verbatim. Those two identifiers are
# the whole vocabulary; there is nothing further to characterise later.
MEMBERSHIP_STATUS: dict[str, dict[str, bool]] = {
"patreon": {
"active_patron": True,
"former_patron": False,
},
"subscribestar": {
"active_subscriptions": True,
"cancelled_subscriptions": False,
},
}
def has_paid_access(
platform: str, status: str | None, *, is_free_member: bool = False,
) -> bool | None:
"""Does this membership mean the account currently PAYS for access?
Returns None for a status this code has not been taught, which callers must
treat as "unknown" rather than as False. The difference matters: False says
the operator has lost access, and asserting that from an unrecognised word
would tell them to cancel a source they are still paying for.
`is_free_member` is a second axis, not a status, and that is Patreon's
design rather than ours: the capture shows a free follow expressed as a
boolean alongside `patron_status`, so a "current" membership can still be
one nobody is paying for. Taking status alone would report a free follower
as a paying patron, and C4 would then never offer to clean it up.
(Honest limit: the capture contains no ACTIVE free member, so it cannot
demonstrate the two axes coming apart. The separation is what the payload's
shape says; the sample only shows it is possible, not that it happens.)
"""
if status is None:
return None
known = MEMBERSHIP_STATUS.get(platform, {}).get(status)
if known is None:
return None
if not known:
return False
return not is_free_member
async def touch_membership(
session: AsyncSession,
*,
platform: str,
external_campaign_id: str,
display_name: str | None = None,
url: str | None = None,
status: str | None = None,
tier_names: list | None = None,
amount_cents: int | None = None,
currency: str | None = None,
details: dict | None = None,
) -> None:
"""Record that this membership was observed just now.
Upsert rather than read-modify-write, for the same reason as
`service_roster.touch_service`: a sweep may overlap its own previous run,
and the last writer is simply the most recent sighting.
`first_seen_at` is deliberately NOT in the update set. It is the one field
that answers "has this ever been true", which is what makes a membership's
later DISAPPEARANCE readable as a lapse rather than indistinguishable from
a creator FC never knew about. Every other column is last-writer-wins,
including status — a membership that goes from active to former must move.
"""
stmt = pg_insert(PlatformMembership).values(
platform=platform,
external_campaign_id=external_campaign_id,
display_name=display_name,
url=url,
status=status,
tier_names=tier_names,
amount_cents=amount_cents,
currency=currency,
details=details or {},
)
stmt = stmt.on_conflict_do_update(
constraint="uq_platform_membership_platform_campaign",
set_={
"display_name": stmt.excluded.display_name,
"url": stmt.excluded.url,
"status": stmt.excluded.status,
"tier_names": stmt.excluded.tier_names,
"amount_cents": stmt.excluded.amount_cents,
"currency": stmt.excluded.currency,
"details": stmt.excluded.details,
"last_seen_at": func.now(),
},
)
await session.execute(stmt)
# ---------------------------------------------------------------------------
# The sweep, and the state that makes its failures readable (#387 C3)
# ---------------------------------------------------------------------------
#
# How long a successful sync stays trustworthy. Beyond this the roster is
# STALE, and C4 must refuse to draw conclusions from it — "you are tracking 12
# sources you no longer subscribe to", computed from a roster that stopped
# syncing a week ago, is an invitation to cancel things the operator is still
# paying for.
#
# Generous relative to the daily cadence: a few missed runs are a blip, not a
# reason to stop trusting a roster that changes on a billing cycle.
ROSTER_STALE_AFTER = timedelta(days=3)
async def get_sync_state(session: AsyncSession, platform: str) -> MembershipSync | None:
return (await session.execute(
select(MembershipSync).where(MembershipSync.platform == platform)
)).scalar_one_or_none()
def roster_is_fresh(state: MembershipSync | None, *, now: datetime | None = None) -> bool:
"""May a caller draw CONCLUSIONS from this roster?
False for never-synced and for stale, and those are deliberately the same
answer here even though the UI must tell them apart: both mean the roster
is not evidence. The asymmetry that matters is that `False` never means
"you subscribe to nothing" — it means "we do not know", and a caller that
cannot represent "we do not know" must not be asking this question.
"""
if state is None or state.last_success_at is None:
return False
now = now or datetime.now(UTC)
return (now - state.last_success_at) <= ROSTER_STALE_AFTER
async def _record_sync(session: AsyncSession, platform: str, **values) -> None:
stmt = pg_insert(MembershipSync).values(platform=platform, **values)
await session.execute(stmt.on_conflict_do_update(
constraint="uq_membership_sync_platform",
set_={**values, "updated_at": func.now()},
))
def roster_user_id(client) -> str | None:
"""The account id a client's roster walk needs, if that client needs one.
Patreon's members endpoint filters on the account's own user id, so the
sweep has to resolve it first. SubscribeStar's /subscriptions page is simply
the logged-in account's, with nothing to resolve. Probed with `getattr`,
the same way the sweep probes `iter_memberships` itself (rule #169), rather
than called unconditionally.
Calling `current_user_id()` unconditionally was the one place the membership
seam was still Patreon-shaped: note #3970 promised a second platform would be
one `builders` line plus the client method, and D1 found the sweep would
instead have crashed on the first client without that method.
"""
resolve = getattr(client, "current_user_id", None)
return resolve() if resolve is not None else None
async def sync_platform(
session: AsyncSession,
*,
platform: str,
fetch: Callable[[], Awaitable[list]],
now: datetime | None = None,
) -> dict:
"""Walk one platform's roster and record what happened.
`fetch` is injected rather than built here so the error-to-state mapping —
the part with the consequences — is testable without a credential, and so
this service needs to know nothing about how any particular client is
constructed.
THE FETCH COMPLETES BEFORE ANYTHING IS WRITTEN. That ordering is the whole
safety property: a walk that dies half way through pagination writes
nothing, so a failure can never leave a roster that is partly this week's
and partly last week's. (`touch_membership` never deletes, so a failure
cannot empty the roster either — but "intact" should mean intact, not
merely non-empty.)
Returns a summary dict; never raises for a platform failure, because one
platform failing must not abort the others.
"""
now = now or datetime.now(UTC)
await _record_sync(session, platform, last_attempt_at=now)
await session.commit()
try:
memberships = await fetch()
except Exception as exc: # noqa: BLE001 - deliberately broad, see below
# Broad on purpose: a sweep is a background job, and ANY escape here
# kills the run for every other platform too. The exception's class
# name is recorded so the distinction the client drew (auth vs drift
# vs transport) survives into the UI, which is where it is actionable.
#
# EXCEPT the worker asking us to stop. Celery raises its soft time
# limit as an ordinary Exception subclass, so a broad catch swallows
# the shutdown request and lets the sweep run on into the HARD limit,
# where it is SIGKILLed mid-transaction. A sweep that cannot be stopped
# is worse than one that fails. (KeyboardInterrupt and SystemExit are
# BaseException and pass through this clause already.)
from celery.exceptions import SoftTimeLimitExceeded
if isinstance(exc, SoftTimeLimitExceeded):
raise
await session.rollback()
await _record_sync(
session, platform,
last_error_type=type(exc).__name__,
last_error_message=str(exc)[:2000],
)
await session.commit()
log.warning("membership sync failed for %s: %s", platform, exc)
return {"platform": platform, "ok": False, "error": type(exc).__name__}
for m in memberships:
await touch_membership(
session,
platform=platform,
external_campaign_id=m.campaign_id,
display_name=m.display_name,
url=m.url,
status=m.status,
tier_names=m.tier_names or None,
amount_cents=m.amount_cents,
currency=m.currency,
details={**(m.details or {}), "is_free_member": m.is_free_member},
)
await _record_sync(
session, platform,
last_success_at=now,
last_count=len(memberships),
# Cleared on success — a stale error beside a fresh success would read
# as "still broken" forever.
last_error_type=None,
last_error_message=None,
)
await session.commit()
log.info("membership sync ok for %s: %d membership(s)", platform, len(memberships))
return {"platform": platform, "ok": True, "count": len(memberships)}
# ---------------------------------------------------------------------------
# Membership <-> Source identity (#387 C4)
# ---------------------------------------------------------------------------
#
# "Is this membership already tracked?" is asked by TWO features — C4's
# reconciliation buckets and E4's creator suggestions — and it lives here, once,
# on purpose. Built inline in C4 it would have looked finished while leaving E4
# matching on name similarity alone, so the two would answer the same question
# differently and only one of them would be right.
#
# E4 and C4 use it from opposite sides: E4 as the NEGATIVE check (propose only
# where nothing matches) and C4 as the join itself.
# Any platform that caches its creator id does so under this suffix; see
# `download_service._phase3_persist`, which writes `patreon_campaign_id`.
_CAMPAIGN_KEY_SUFFIX = "_campaign_id"
def identity_keys_for_source(source: Source) -> set[str]:
"""Every platform-side creator id cached on this source.
Reads ANY `<platform>_campaign_id` override rather than naming Patreon's,
so a second platform participates by caching its id under the same suffix —
no registry, no `if platform ==` branch (rule 169). A source that has never
been walked has cached nothing and simply contributes no exact key, which is
what makes the handle fallback below necessary rather than merely tolerated.
"""
keys = set()
for name, value in (source.config_overrides or {}).items():
if name.endswith(_CAMPAIGN_KEY_SUFFIX) and isinstance(value, str) and value:
keys.add(value)
return keys
def url_tail(url: str | None) -> str | None:
"""The creator handle at the end of a source URL, lowercased.
Deliberately the same derivation as `PlatformMembership.vanity_or_none`'s
own fallback, so both sides of the comparison reduce a URL to a handle the
same way. Query strings and fragments are stripped first; Patreon's `/c/`
and `/cw/` forms both end in the vanity, so they need no special case (the
missing-`/c/` regex is what broke creator detection in #1485).
Returns None for the pre-0030 `sidecar:` synthetic anchors, which are not
feeds and must never match anything.
"""
if not url or url.startswith("sidecar:"):
return None
cleaned = url.split("?", 1)[0].split("#", 1)[0]
tail = cleaned.rstrip("/").rsplit("/", 1)[-1]
return tail.lower() or None
def match_kind(source: Source, membership: PlatformMembership) -> str | None:
"""How this source and this membership are known to be the same creator.
Returns "campaign" for an exact platform-id match, "vanity" for agreeing URL
handles, or None for no evidence.
THE ORDER MUST NOT BE INVERTED. The campaign id is exact and the handle is
not, but the id is only written AFTER a source has been walked at least once
— so checking the handle first would let a stale or renamed URL outvote the
authoritative id on every source FC has actually polled.
"""
if source.platform != membership.platform:
return None
if membership.external_campaign_id in identity_keys_for_source(source):
return "campaign"
vanity = membership.vanity_or_none()
tail = url_tail(source.url)
if vanity and tail and vanity.strip().lower() == tail:
return "vanity"
return None
def pair_sources_with_memberships(
sources: list[Source], memberships: list[PlatformMembership],
) -> dict[int, tuple[PlatformMembership, str]]:
"""Source id -> the membership it is the same creator as, and how we know.
Extracted from C4's reconcile loop when C5 became its second caller. It is
a nested loop rather than a SQL join because the match is a predicate over
a JSON blob and a derived URL handle, neither of which is indexable, and
both sides are tens of rows on any real library. Keeping it in Python means
ONE definition of identity (`match_kind`) instead of a second one in SQL
that could drift from it.
First match wins, which is `match_kind`'s ordering doing its job: a source
with a cached campaign id can only pair with the membership holding that
id, so an ambiguous handle never outvotes it.
"""
pairs: dict[int, tuple[PlatformMembership, str]] = {}
for source in sources:
for m in memberships:
kind = match_kind(source, m)
if kind:
pairs[source.id] = (m, kind)
break
return pairs
# ---------------------------------------------------------------------------
# Why the posts are invisible (#387 C5)
# ---------------------------------------------------------------------------
#
# A3 made a tier-gated source say "47 posts you can't see". These are the words
# the roster is allowed to add to that count — and ONLY to that count.
#
# THE LINE: the roster ANNOTATES the gated flag, it never produces it.
# `current_user_can_view` (read per post by `patreon_client.post_is_gated`) is
# the authoritative per-post signal, and entitled-tier data cannot stand in for
# it — a creator can gate a post behind an access rule that maps onto no tier
# name at all. So nothing here may suppress a download, skip a walk, or decide
# a post is inaccessible. It explains a skip that ALREADY happened. Getting
# that backwards would make FC silently stop fetching content the operator is
# paying for, which is the worst failure available in this milestone.
# `test_no_fetch_path_can_read_the_roster` pins that structurally.
GATED_LAPSED = "lapsed" # the membership ended — resubscribe, or disable
GATED_TIER = "tier" # paying, but this tier doesn't reach these posts
GATED_FREE = "free" # a current FREE follow — nobody is paying for access
def gated_reason(
platform: str, status: str | None, *, is_free_member: bool = False,
) -> str | None:
"""Why a tier-gated source's posts are out of reach, if the roster knows.
None means "no words beyond the count" and is the answer for every case
where the roster is not evidence: a status this code has not been taught,
and (at the call site) a campaign absent from the roster or a roster too
stale to trust. Absence is not evidence — the same discipline as
`test_post_is_gated_only_on_explicit_false`.
`is_free_member` is read AFTER the status axis, not folded into it, which
is why `has_paid_access` is called here with it forced off. The two axes
are independent in Patreon's payload, and collapsing them loses a real
distinction: a current free follower has not lost anything, so telling them
"you're not a patron any more" would be a false sentence about a state they
were never in.
"""
by_status = has_paid_access(platform, status, is_free_member=False)
if by_status is None:
return None
if not by_status:
return GATED_LAPSED
return GATED_FREE if is_free_member else GATED_TIER
async def gated_reasons_for_sources(
session: AsyncSession, sources: list[Source], *, now: datetime | None = None,
) -> dict[int, str]:
"""The reason word for each of these sources, where the roster has one.
Callers pass ONLY the sources already known to be tier-gated: the question
"why can't I see these posts" is meaningless for a source whose posts are
all visible, and asking it anyway would put roster data on rows that have
no gated state for it to annotate.
Sources with no entry in the result get A3's bare count, which is the
correct degraded rendering for all three of: platform never swept, roster
stale, campaign not in the roster.
"""
if not sources:
return {}
platforms = {s.platform for s in sources}
# Per platform, because freshness is per platform: a working Patreon sweep
# must not lend its credibility to a SubscribeStar roster that has never
# run. Same gate as C4's `tracked_not_subscribed`, for the same reason.
fresh = {
p for p in platforms
if roster_is_fresh(await get_sync_state(session, p), now=now)
}
if not fresh:
return {}
memberships = (await session.execute(
select(PlatformMembership).where(
PlatformMembership.platform.in_(sorted(fresh))
)
)).scalars().all()
pairs = pair_sources_with_memberships(
[s for s in sources if s.platform in fresh], memberships,
)
reasons: dict[int, str] = {}
for source_id, (m, _kind) in pairs.items():
reason = gated_reason(
m.platform, m.status,
is_free_member=bool((m.details or {}).get("is_free_member")),
)
if reason is not None:
reasons[source_id] = reason
return reasons
async def source_for_membership(
session: AsyncSession, membership: PlatformMembership,
) -> Source | None:
"""The source FC already tracks for this membership, if there is one.
Scoped to the membership's own platform, so a creator tracked on Discord and
subscribed to on Patreon does not read as already-tracked — that pairing is
E4's suggestion to make, not an identity.
"""
rows = (await session.execute(
select(Source).where(Source.platform == membership.platform)
)).scalars().all()
for source in rows:
if match_kind(source, membership):
return source
return None
@@ -211,6 +211,59 @@ class PostRecordOutcome:
body_chars: int
# -- membership roster seam (shared dataclass, #387 C2/C7) -----------------
@dataclass
class Membership:
"""One membership the ACCOUNT holds, as the roster needs it (#387 C2).
Lives HERE rather than in the platform module that first produced it, for
the same reason `PostRecordOutcome` does: it is the seam's contract, not
Patreon's. C7 moved it — while it sat in `patreon_client` a second platform
would have had to import its contract from the first platform's module,
which inverts the dependency and is how a "portable" seam quietly becomes
Patreon-shaped.
Deliberately not a raw upstream row: the sweep should not have to know that
a tier lives behind a JSON:API `reward` relationship, and
`platform_membership` should not gain columns because one platform shapes
things a certain way.
`status` carries the PLATFORM's own word, verbatim and unmapped
(`active_patron`, `former_patron`, ...). Deciding what it means is the read
site's job — `membership_roster.has_paid_access` — precisely so an
unrecognised word records as evidence rather than as a decision.
`is_free_member` is SEPARATE from status and must stay that way. Patreon
expresses a free follow as this boolean rather than as a status value, so
"does the account pay for this" is `status == "active_patron" and not
is_free_member` — a question the status string alone cannot answer. NOTE:
the C0 capture contains no ACTIVE free member, so the two fields are
perfectly correlated in that sample; the separation is what the schema
says, not something the sample proves.
A platform that lacks a field supplies the empty answer, never a guess:
no tiers -> `[]`, no pledge -> `amount_cents=None` (absent stays
distinguishable from zero — "free" and "we don't know" are different
answers), no vanity -> None and identity falls back to the URL tail.
"""
campaign_id: str
display_name: str | None
url: str | None
vanity: str | None
status: str | None
is_free_member: bool
tier_names: list[str]
amount_cents: int | None
currency: str | None
# Everything the roster did not model, kept so a later question can be
# answered without another authenticated round-trip. Scoped to the
# membership's own attributes plus the creator's — never the raw page,
# which is where the card/address resources live.
details: dict
# -- base downloader (shared fetch/validate plumbing) ----------------------
class BaseNativeDownloader:
+220 -19
View File
@@ -14,6 +14,18 @@ the later step can drive it:
- extract_media(post, included_index) → list[MediaItem]
- parse_cursor_from_url(url) → cursor
Milestone 387 added a SECOND read path on the same session: the membership
roster — what the ACCOUNT subscribes to, as opposed to what one creator has
posted.
- iter_memberships(user_id) → Iterator[Membership]
- current_user_id() → str
It is an OPTIONAL seam by construction, probed with
`getattr(client, "iter_memberships", None)` exactly as `post_is_gated` already
is. A client that does not implement it (Discord, HentaiFoundry) makes the
whole feature invisible for that platform — no flag, no config row, no
"unsupported" branch to keep alive.
Drift detection is loud on purpose: Patreon ships JSON:API and the shapes we
depend on (top-level `data`, media resources carrying `file_name`/`url`) are
the contract. If a response comes back as an HTML login page or a media
@@ -41,6 +53,7 @@ from ..utils.paths import filehash_from_url
from ..utils.prosemirror import post_body_html
from .native_ingest_common import (
_MAX_429_RETRIES,
Membership,
NativeAuthError,
NativeDriftError,
NativeIngestError,
@@ -52,8 +65,34 @@ from .native_ingest_common import (
log = logging.getLogger(__name__)
_POSTS_URL = "https://www.patreon.com/api/posts"
_MEMBERS_URL = "https://www.patreon.com/api/members"
_CURRENT_USER_URL = "https://www.patreon.com/api/current_user"
_TIMEOUT_SECONDS = 30.0
# --- membership roster contract (#387 C2) ---------------------------------
# Characterized from a real capture of the operator's own session — Scribe note
# #3886. NOT from Patreon's public v2 API, which is the CREATOR api behind
# OAuth scopes and a different surface entirely (project rule 130).
#
# DELIBERATELY MINIMAL, and that is a privacy decision rather than a
# performance one. The web app's own include set pulls `latest_pledge.card`
# and `address`; the card resources come back carrying the ACCOUNT HOLDER'S
# EMAIL in `merchant_name`. Copying the browser's query string wholesale — the
# obvious move — would have FC fetching payment PII it has no use for and can
# only mishandle. We ask for the creator and the tier, and nothing else.
_MEMBERS_INCLUDE = "campaign,reward"
_FIELDS_MEMBER = (
"patron_status,is_free_member,is_gifted,pledge_amount_cents,currency,"
"pledge_cadence,next_charge_date,access_expires_at"
)
_FIELDS_MEMBERS_CAMPAIGN = "name,url,vanity,is_active"
_FIELDS_REWARD = "title"
# The browser sends 1000. Whether a server-side ceiling applies below that is
# untested (note #3886, open question 4), so page conservatively: a wrong guess
# costs one extra request, and the paging loop is driven by meta.pagination
# rather than by this number.
_MEMBERS_PAGE_COUNT = 200
# JSON:API request contract (observed from real traffic — see module plan).
_INCLUDE = (
"campaign,access_rules,attachments,attachments_media,audio,images,media,"
@@ -182,20 +221,27 @@ class PatreonClient:
params["page[cursor]"] = cursor
return params
def _fetch(self, campaign_id: str, cursor: str | None) -> dict:
def _request(self, url: str, params: dict[str, str], *, what: str, scope: str) -> dict:
"""One paced, retried, error-classified GET returning parsed JSON.
Extracted from `_fetch` so the membership endpoint (#387 C2) rides the
SAME request path rather than growing a second copy of the 429 backoff,
the auth-vs-drift classification and the Retry-After plumbing. Two
copies of this would drift, and the half that drifted would be the one
that only runs once a day.
`what` / `scope` only shape the messages ("posts"/"campaign_id=123"),
so a failure still says which call failed and against what.
"""
if self._request_sleep > 0:
time.sleep(self._request_sleep) # pace the API endpoint
attempt = 0
while True:
try:
resp = self._session.get(
_POSTS_URL,
params=self._params(campaign_id, cursor),
timeout=_TIMEOUT_SECONDS,
)
resp = self._session.get(url, params=params, timeout=_TIMEOUT_SECONDS)
except requests.RequestException as exc:
raise PatreonAPIError(
f"Patreon posts request failed (campaign_id={campaign_id}): {exc}"
f"Patreon {what} request failed ({scope}): {exc}"
) from exc
# Transient rate-limit: back off and retry rather than failing the
@@ -205,8 +251,8 @@ class PatreonClient:
attempt += 1
delay = retry_after_seconds(resp, attempt)
log.warning(
"Patreon 429 (campaign_id=%s) — backing off %.1fs (retry %d/%d)",
campaign_id, delay, attempt, self._max_retries,
"Patreon 429 (%s) — backing off %.1fs (retry %d/%d)",
scope, delay, attempt, self._max_retries,
)
time.sleep(delay)
continue
@@ -216,9 +262,8 @@ class PatreonClient:
# Auth rejected — expired/missing cookies or an insufficient tier.
# Actionable as "rotate credentials", so it's auth, not drift/http.
raise PatreonAuthError(
f"Patreon posts API returned HTTP {resp.status_code} — auth "
f"rejected (cookies expired or tier insufficient; "
f"campaign_id={campaign_id})",
f"Patreon {what} API returned HTTP {resp.status_code} — auth "
f"rejected (cookies expired or tier insufficient; {scope})",
status_code=resp.status_code,
)
if resp.status_code != 200:
@@ -234,24 +279,27 @@ class PatreonClient:
except (TypeError, ValueError):
retry_after = None
raise PatreonAPIError(
f"Patreon posts API returned HTTP {resp.status_code} "
f"(campaign_id={campaign_id})",
f"Patreon {what} API returned HTTP {resp.status_code} ({scope})",
status_code=resp.status_code,
retry_after=retry_after,
)
try:
payload = resp.json()
return resp.json()
except ValueError as exc:
# A non-JSON body here is almost always the HTML login/challenge
# page served when cookies are missing/expired — that is an AUTH
# failure (rotate cookies), not API drift (update the ingester) and
# not a transient network error.
raise PatreonAuthError(
"Patreon posts API returned a non-JSON response (likely an "
f"HTML login/challenge page — session expired; "
f"campaign_id={campaign_id}): {exc}"
f"Patreon {what} API returned a non-JSON response (likely an "
f"HTML login/challenge page — session expired; {scope}): {exc}"
) from exc
return payload
def _fetch(self, campaign_id: str, cursor: str | None) -> dict:
return self._request(
_POSTS_URL, self._params(campaign_id, cursor),
what="posts", scope=f"campaign_id={campaign_id}",
)
# -- parsing -----------------------------------------------------------
@@ -510,6 +558,159 @@ class PatreonClient:
return
current_cursor = next_cursor
# -- membership roster (#387 C2) ---------------------------------------
def current_user_id(self) -> str:
"""The signed-in account's own numeric user id.
Needed because `/api/members` is filtered by `filter[user_id]` — the
endpoint answers "who are the members of X", and the account asking
about ITSELF still has to say so.
INFERRED, NOT CHARACTERIZED. C0 captured `/api/members`, not this; what
is relied on here is only the JSON:API envelope (`data.id`), which this
same API demonstrably uses everywhere else. If that inference is wrong
it raises drift rather than returning something plausible — which is
the right failure, because the alternative is a confidently empty
roster and an empty roster means "cancel everything" to C4.
"""
payload = self._request(
_CURRENT_USER_URL, {"json-api-version": "1.0"},
what="current_user", scope="self",
)
data = (payload or {}).get("data")
if not isinstance(data, dict) or not data.get("id"):
raise PatreonDriftError(
"Patreon current_user response had no data.id — cannot scope "
"the membership roster to this account"
)
return str(data["id"])
def _members_params(self, user_id: str | None, offset: int) -> dict[str, str]:
params = {
"include": _MEMBERS_INCLUDE,
"fields[member]": _FIELDS_MEMBER,
"fields[campaign]": _FIELDS_MEMBERS_CAMPAIGN,
"fields[reward]": _FIELDS_REWARD,
"page[offset]": str(offset),
"page[count]": str(_MEMBERS_PAGE_COUNT),
"json-api-version": "1.0",
"json-api-use-default-includes": "false",
}
if user_id:
params["filter[user_id]"] = user_id
# NOTE: `filter[membership_type]` is deliberately NOT sent. The browser
# sends the six values its settings page wants to show, and the capture
# proves that list is NOT the same vocabulary as the `patron_status`
# attribute — a row selected as `free_member` came back with
# `patron_status: former_patron`, a word absent from the filter. Sending
# no filter asks for everything the endpoint will give, which is what a
# roster wants: a membership that DISAPPEARS is the signal C4 reads, and
# a filter tuned for a UI that hides lapses would manufacture exactly
# that disappearance. (Note #3886, open question 1.)
return params
@staticmethod
def _validate_members_response(response: dict) -> None:
"""Drift checks specific to the roster.
Stricter than the posts path about pagination on purpose: `iter_posts`
can treat a missing `links.next` as "that was the last page", but here
a missing total is indistinguishable from a truncated page — and a
roster that silently stops half way reads downstream as "you cancelled
those", which is the worst wrong answer this feature can give.
"""
PatreonClient._validate_response(response)
meta = response.get("meta")
if not isinstance(meta, dict):
raise PatreonDriftError("Patreon members response missing 'meta'")
pagination = meta.get("pagination")
if not isinstance(pagination, dict) or "total" not in pagination:
raise PatreonDriftError(
"Patreon members response missing meta.pagination.total — "
"cannot tell a complete roster from a truncated one"
)
def _membership(self, member: dict, index: dict) -> Membership:
attrs = member.get("attributes") or {}
if "patron_status" not in attrs:
raise PatreonDriftError(
"Patreon member resource has no patron_status attribute"
)
campaign_ids = self._related_ids(member, "campaign")
if not campaign_ids:
raise PatreonDriftError(
"Patreon member resource has no campaign relationship — a "
"membership we cannot attribute to a creator is not usable"
)
campaign_id = campaign_ids[0]
campaign = index.get(("campaign", campaign_id)) or {}
# A member has at most one reward, and `reward.data` is legitimately
# null — an active patron with no tier. Absence is a fact about the
# membership, not a parse failure.
tier_names: list[str] = []
for reward_id in self._related_ids(member, "reward"):
title = (index.get(("reward", reward_id)) or {}).get("title")
if title:
tier_names.append(str(title))
return Membership(
campaign_id=campaign_id,
display_name=campaign.get("name"),
url=campaign.get("url"),
vanity=campaign.get("vanity"),
status=attrs.get("patron_status"),
# Default False, not None: the attribute is always present in the
# capture, and treating a missing one as "free" would understate
# access rather than overstate it.
is_free_member=bool(attrs.get("is_free_member")),
tier_names=tier_names,
# The MEMBER's amount, never the reward's. `reward.amount_cents` is
# the creator's list price in the CREATOR's currency (the capture
# has CAD, DKK and EUR rewards sitting on USD pledges), so reading
# it would report a number the operator has never been charged.
amount_cents=attrs.get("pledge_amount_cents"),
currency=attrs.get("currency"),
details={"member": attrs, "campaign": campaign},
)
def iter_memberships(self, user_id: str | None = None) -> Iterator[Membership]:
"""Yield every membership the account holds.
Pages on `page[offset]`/`page[count]` against `meta.pagination.total` —
NOT on `links`. The response's own `links.first` is built without the
`/api/` prefix the request uses, so following it verbatim would hit the
web page instead of the API (note #3886).
`user_id` omitted means the `filter[user_id]` parameter is omitted.
Whether the endpoint then defaults to self is UNTESTED — pass
`current_user_id()` unless you are deliberately probing that.
"""
user_id = user_id or None
offset = 0
seen = 0
while True:
response = self._request(
_MEMBERS_URL, self._members_params(user_id, offset),
what="members", scope="membership roster",
)
self._validate_members_response(response)
index = self._transform(response)
rows = [m for m in (response.get("data") or []) if isinstance(m, dict)]
for member in rows:
yield self._membership(member, index)
seen += len(rows)
total = int(response["meta"]["pagination"]["total"] or 0)
# An empty page terminates regardless of what `total` claims. Trusting
# the total alone would spin forever against a server that reports
# more rows than it will hand over.
if not rows or seen >= total:
return
offset += len(rows)
# -- detail (full body enrichment) -------------------------------------
def fetch_post_detail_content(self, post_id: str) -> str | None:
+5 -3
View File
@@ -11,7 +11,11 @@ Lifted from GallerySubscriber's
and ~/.../extension/lib/platforms.js. Five platforms; auth_type and
URL patterns match GS exactly so the existing browser extension
hits FC unmodified. deviantart was dropped at #3069 (2026-08-27) —
FC downloaders are art-dedicated services only.
FC downloaders are art-dedicated services only. pixiv was retired at
milestone #406 (2026-09-13, rule #171): unregistered here first, which
switches it off everywhere this registry is consulted; `pixiv.py` and the
pixiv client/downloader/ingester stay in the tree, uncalled, until the
milestone's phase 2 deletes them.
"""
from .base import (
@@ -22,7 +26,6 @@ from .base import (
from .discord import INFO as _DISCORD
from .hentaifoundry import INFO as _HENTAIFOUNDRY
from .patreon import INFO as _PATREON
from .pixiv import INFO as _PIXIV
from .subscribestar import INFO as _SUBSCRIBESTAR
PLATFORMS: dict[str, PlatformInfo] = {
@@ -32,7 +35,6 @@ PLATFORMS: dict[str, PlatformInfo] = {
_SUBSCRIBESTAR,
_HENTAIFOUNDRY,
_DISCORD,
_PIXIV,
)
}
@@ -0,0 +1,308 @@
"""The announcement matcher: which Patreon post announced which Discord drop.
Milestone 388, step E5.
Two of the operator's artists post a deliberately CROPPED fragment on Patreon
to signal that the real thing has landed in their Discord. This service
proposes those pairs, and proposes only — the operator accepts or dismisses,
following the FC-6.3 series matcher (task 737) rather than linking on its own.
## Why confirm-only is not caution for its own sake
A wrongly-asserted association tells the operator that two different pieces are
one. That is strictly worse than no link at all: no link leaves them exactly
where they already were, a wrong one actively misinforms and then propagates
into whatever reads the association. So the matcher's job is to make a SHORT
list worth reading, not a long list worth trusting.
## Signals, and the one deliberately NOT built
1. **Time proximity.** The Patreon post exists in order to announce the drop,
so the two are minutes-to-hours apart. Nearly free, and strong.
2. **The post says so.** These announcements routinely name Discord or carry
an invite link, which is close to a declaration.
3. **Crop-to-source matching is HELD, on the plan's own instruction** — it is
real work with real false-positive risk, and it is only worth building once
1 and 2 are shown to be insufficient against the operator's actual artists.
Nothing here should be read as evidence it is unnecessary; it is deferred,
and the thing that would justify it is an empty review queue on a pair the
operator can see with their own eyes.
Note also that a naive whole-image SigLIP similarity is NOT that signal. A
cropped teaser and its full version are exactly the pair a whole-image
comparison handles worst, so adding one as a "bonus" would mostly add noise
while looking like progress.
## Creator identity comes free, so E4 is not actually a prerequisite
The plan listed E4 (creator identity across the two channels) as a dependency.
It is not one for the pairs that matter today: a Patreon `Source` and a Discord
`Source` the operator has added under the same `Artist` already share
`Post.artist_id`, and the synthetic grouping inherits it (E2). E4 EXTENDS this
to creators whose association FC has to learn rather than being told; it is not
needed to represent an association FC already knows.
## A correction the plan carried, worth restating
`link_extract.py` does exist and does capture off-platform links — but only
file hosts (`SUPPORTED_HOSTS` is mega/gdrive/mediafire/dropbox/pixeldrain).
`host_for()` returns None for a Discord URL, so no `ExternalLink` row is ever
written for one. The declaration signal therefore reads the post body itself
rather than the extracted-links table the plan assumed it could use.
"""
from __future__ import annotations
import logging
import re
from datetime import UTC, datetime, timedelta
from sqlalchemy import func, or_, select
from sqlalchemy.ext.asyncio import AsyncSession
from ..models import ImportSettings, Post, PostAssociation
from ..utils.text import html_to_plain
from .discord_grouping import DROP_GROUPER
log = logging.getLogger(__name__)
# Additive weights, summing to 1.0. Kept as constants rather than settings —
# the sensitivity knob that matters is the threshold, and per-signal weights
# are an over-tune (same call as series_match_service.WEIGHTS).
#
# THE RELATIONSHIP TO THE THRESHOLD IS THE DESIGN. No single weight may reach
# the default threshold, which is what makes "time proximity alone is never
# enough" arithmetic rather than aspirational: on a busy day an artist posts
# several times, and a matcher that could pair on proximity alone would turn
# every busy day into false pairs. A guard test pins this.
WEIGHTS = {"proximity": 0.55, "declared": 0.45}
# A Discord INVITE in the body is close to a declaration; the bare word is
# weaker but still meaningful, because these posts are short and on-topic.
_INVITE = re.compile(r"discord\.(?:gg|com/invite)/", re.I)
_MENTION = re.compile(r"\bdiscord\b", re.I)
DECLARED_INVITE = 1.0
DECLARED_MENTION = 0.6
MAX_CANDIDATES = 25
def proximity_signal(gap: timedelta, window: timedelta) -> float:
"""1.0 when the two posts are simultaneous, decaying linearly to 0 at the
window's edge. Linear rather than a step, so a pair an hour outside a
hand-tuned window degrades instead of vanishing."""
if window <= timedelta(0):
return 0.0
seconds = abs(gap.total_seconds())
if seconds >= window.total_seconds():
return 0.0
return round(1.0 - (seconds / window.total_seconds()), 4)
def declared_signal(description: str | None) -> float:
"""Does the announcement say, in its own body, that this is about Discord?
The invite is matched against the RAW body and the bare mention against the
stripped text, which is not fussiness — post bodies are HTML, and these
creators put the invite in an anchor's `href`. `html_to_plain` discards
attributes, so stripping first would have thrown away the strongest form of
the signal and left only whatever the link text happened to say.
The mention still reads stripped text, so `\bdiscord\b` is matched against
prose rather than against markup and URLs, where it would fire on any link
that merely passes through a discord domain.
"""
if not description:
return 0.0
if _INVITE.search(description):
return DECLARED_INVITE
if _MENTION.search(html_to_plain(description) or ""):
return DECLARED_MENTION
return 0.0
def weighted_score(signals: dict) -> float:
return round(sum(WEIGHTS[k] * signals.get(k, 0.0) for k in WEIGHTS), 4)
def _post_time(post: Post) -> datetime:
return post.post_date or post.downloaded_at
class PostAssociationService:
def __init__(self, session: AsyncSession):
self.session = session
async def _decided(self, announcement_id: int) -> set[int]:
"""Payload posts already proposed for this announcement, in ANY status.
Dismissed pairs are included deliberately: re-proposing a pair the
operator has already rejected on every subsequent scan is the single
behaviour that makes a review queue get ignored.
"""
rows = (await self.session.execute(
select(PostAssociation.payload_post_id)
.where(PostAssociation.announcement_post_id == announcement_id)
)).scalars().all()
return set(rows)
async def _candidate_groups(
self, announcement: Post, *, window: timedelta,
) -> list[Post]:
"""Synthetic Discord groupings by the SAME artist, inside the window.
Same-artist is the identity signal and it is free (see the module
docstring on E4). It is also a hard filter rather than a scored one:
two different creators posting minutes apart is a coincidence, not
evidence, and letting it score at all would mean a busy hour across the
library could out-vote everything else.
"""
at = _post_time(announcement)
sort_key = func.coalesce(Post.post_date, Post.downloaded_at)
return (await self.session.execute(
select(Post)
.where(
Post.artist_id == announcement.artist_id,
Post.synthesized_by == DROP_GROUPER,
Post.id != announcement.id,
sort_key >= at - window,
sort_key <= at + window,
)
.order_by(sort_key)
.limit(MAX_CANDIDATES)
)).scalars().all()
async def match_post(
self, announcement_id: int, *, threshold: float, window_hours: float,
) -> int:
"""Score one announcement against nearby groupings. Returns proposals made."""
announcement = await self.session.get(Post, announcement_id)
if announcement is None or announcement.synthesized_by is not None:
# A synthetic post cannot announce anything — FC wrote it.
return 0
window = timedelta(hours=window_hours)
declared = declared_signal(announcement.description)
already = await self._decided(announcement_id)
made = 0
for group in await self._candidate_groups(announcement, window=window):
if group.id in already:
continue
signals = {
"proximity": proximity_signal(
_post_time(group) - _post_time(announcement), window,
),
"declared": declared,
}
score = weighted_score(signals)
if score < threshold:
continue
self.session.add(PostAssociation(
announcement_post_id=announcement.id,
payload_post_id=group.id,
score=score,
signals=signals,
status="pending",
))
made += 1
return made
async def list_pending(self) -> list[dict]:
rows = (await self.session.execute(
select(PostAssociation)
.where(PostAssociation.status == "pending")
.order_by(PostAssociation.score.desc(), PostAssociation.id.desc())
)).scalars().all()
return [
{
"id": a.id,
"announcement_post_id": a.announcement_post_id,
"payload_post_id": a.payload_post_id,
"score": a.score,
"signals": a.signals,
}
for a in rows
]
async def accept(self, association_id: int) -> dict | None:
a = await self.session.get(PostAssociation, association_id)
if a is None:
return None
a.status = "linked"
return {"id": a.id, "status": a.status}
async def dismiss(self, association_id: int) -> dict | None:
a = await self.session.get(PostAssociation, association_id)
if a is None:
return None
# Kept, not deleted — the row is what remembers the rejection.
a.status = "dismissed"
return {"id": a.id, "status": a.status}
async def linked_for(self, post_ids: list[int]) -> dict[int, list[dict]]:
"""Accepted links touching these posts, keyed by post id, BOTH ways.
A post is either end of the relationship, and each end wants the other
one: the teaser wants "the full set is over here", the grouping wants
"this is what announced me". One query, both directions.
"""
if not post_ids:
return {}
rows = (await self.session.execute(
select(PostAssociation).where(
PostAssociation.status == "linked",
or_(
PostAssociation.announcement_post_id.in_(post_ids),
PostAssociation.payload_post_id.in_(post_ids),
),
)
)).scalars().all()
out: dict[int, list[dict]] = {}
for a in rows:
if a.announcement_post_id in post_ids:
out.setdefault(a.announcement_post_id, []).append(
{"role": "announces", "post_id": a.payload_post_id, "id": a.id}
)
if a.payload_post_id in post_ids:
out.setdefault(a.payload_post_id, []).append(
{"role": "announced_by", "post_id": a.announcement_post_id, "id": a.id}
)
return out
async def rescan(session: AsyncSession, *, now: datetime | None = None) -> dict:
"""Score every recent non-synthetic post against nearby groupings."""
settings = await ImportSettings.load(session)
if not settings.discord_link_enabled:
return {"enabled": False, "scanned": 0, "proposed": 0}
now = now or datetime.now(UTC)
window_hours = float(settings.discord_link_window_hours)
# Only look at announcements that could still have a partner in range —
# a full-library rescan is the manual button's job, not the sweep's.
horizon = now - timedelta(hours=window_hours * 2)
sort_key = func.coalesce(Post.post_date, Post.downloaded_at)
ids = (await session.execute(
select(Post.id).where(
Post.synthesized_by.is_(None),
Post.absorbed_by_post_id.is_(None),
sort_key >= horizon,
)
)).scalars().all()
svc = PostAssociationService(session)
proposed = 0
for pid in ids:
proposed += await svc.match_post(
pid,
threshold=float(settings.discord_link_threshold),
window_hours=window_hours,
)
log.info(
"discord announcement matcher: scanned %d post(s), proposed %d pair(s)",
len(ids), proposed,
)
return {"enabled": True, "scanned": len(ids), "proposed": proposed}
+93
View File
@@ -0,0 +1,93 @@
"""Rendering a post's stored HTML body for display — sanitized, and pointing
at our own copies of the images rather than the platform's.
## Why this is one function and not two
Two surfaces render a post body: the post detail view
(`PostFeedService.get_post`) and the provenance panel (`ProvenanceService`).
Until issue #3965 the detail view did `_localize_inline_images(sanitize(...))`
and provenance did `sanitize(...)` alone — so the provenance panel hotlinked
the platform CDN for images FC had already downloaded, which is exactly what
#830 Phase 2 set out to stop.
That bug was available because the two halves were separately callable and
only one of them looked mandatory. `render_post_body` is therefore the whole
pipeline in a single call, and it is the only thing callers are meant to
reach for: sanitizing without localizing is not a supported operation, so it
should not be a reachable one. `localize_inline_images` stays public only
because a caller that already holds sanitized HTML needs it.
## Cost
`localize_inline_images` issues ZERO queries for a body with no inline
`<img>`, which is most of them — it returns before touching the session if
the body is empty, has no image tags, or has none carrying a parseable CDN
filehash. That early exit is why callers may run this per post in a loop
rather than needing a batched form; do not "optimize" it into one without a
measurement saying the loop is actually hot.
"""
from __future__ import annotations
from html import unescape
from sqlalchemy import select
from sqlalchemy.ext.asyncio import AsyncSession
from ..models import ImageRecord
from ..utils.html_sanitize import extract_img_srcs, rewrite_img_srcs, sanitize_post_html
from ..utils.paths import filehash_from_url
from .gallery_service import image_url
async def render_post_body(
session: AsyncSession, description: str | None, artist_id: int | None,
) -> str | None:
"""A post's body, ready to put in front of someone.
Sanitize, then repoint inline images at local copies. Use this rather than
calling either half on its own — see the module docstring.
"""
return await localize_inline_images(
session, sanitize_post_html(description), artist_id,
)
async def localize_inline_images(
session: AsyncSession, html: str | None, artist_id: int | None,
) -> str | None:
"""Rewrite a post body's inline `<img src=CDN>` to locally-served copies.
The join key is the CDN filehash the downloader persisted on each
ImageRecord (source_filehash): for every body image whose filehash maps
to a stored image of THIS artist, swap the src to /images/<path>. Images
we never captured (or pre-Phase-2 rows with no filehash) are left as-is —
they keep hotlinking, which is the prior behavior. Scoped to the post's
artist so one creator's body never resolves to another's file.
"""
if not html or artist_id is None:
return html
srcs = extract_img_srcs(html)
if not srcs:
return html
# filehash -> the raw (as-in-HTML) src strings carrying it. A body can
# repeat the same image; keep every raw form so each is substituted.
by_hash: dict[str, list[str]] = {}
for raw in srcs:
fh = filehash_from_url(unescape(raw))
if fh:
by_hash.setdefault(fh, []).append(raw)
if not by_hash:
return html
rows = (await session.execute(
select(ImageRecord.source_filehash, ImageRecord.path)
.where(
ImageRecord.artist_id == artist_id,
ImageRecord.source_filehash.in_(list(by_hash)),
)
)).all()
replace: dict[str, str] = {}
for fh, path in rows:
for raw in by_hash.get(fh, ()):
replace[raw] = image_url(path)
return rewrite_img_srcs(html, replace)
+101 -60
View File
@@ -11,8 +11,6 @@ attachments from PostAttachment) so the API layer can jsonify directly.
"""
from __future__ import annotations
from html import unescape
from sqlalchemy import and_, func, or_, select
from sqlalchemy.ext.asyncio import AsyncSession
@@ -26,23 +24,50 @@ from ..models import (
Source,
attachment_download_url,
)
from ..utils.html_sanitize import (
extract_img_srcs,
rewrite_img_srcs,
sanitize_post_html,
)
from ..utils.paths import filehash_from_url
from ..utils.text import html_to_plain, truncate_at_word
from .gallery_service import image_url, thumbnail_url
from .gallery_service import thumbnail_url
from .pagination import decode_cursor, encode_cursor
from .post_body import render_post_body
DESCRIPTION_LIMIT = 280
THUMBNAIL_LIMIT = 6
def _sort_key():
"""Postgres COALESCE expression used in ORDER BY and WHERE clauses."""
return func.coalesce(Post.post_date, Post.downloaded_at)
"""Postgres COALESCE expression used in ORDER BY and WHERE clauses.
`resurfaced_at` leads (milestone 388 E3). A synthetic post stays OPEN — a
creator who adds variants the next day extends the existing post — so such
a post has two dates, and which one orders the feed is a real decision:
* ordering by when the drop STARTED buries a group that grows a week later
under a week of other posts, so the operator never sees the new content —
which defeats keeping the group open at all;
* ordering by every growth lets a group that gains one image a day sit
permanently at the top, so chat out-competes authored posts for the front
page — the opposite of "post pacing stays front and centre".
So the feed orders by neither directly. `resurfaced_at` moves only when the
anti-thrash rule fires (discord_grouping.should_resurface: enough new
images AND enough time since the last move), which means a drip-feed
updates IN PLACE and a genuine second wave resurfaces exactly once.
It is NULL on every ordinary post, so this COALESCE cannot move anything
that is not a grouping. Used identically in ORDER BY and in the cursor's
WHERE, which is what keeps pagination stable across the change.
"""
return func.coalesce(Post.resurfaced_at, Post.post_date, Post.downloaded_at)
def _post_sort_value(post: Post):
"""The Python twin of `_sort_key()`, for building a cursor from a loaded row.
Kept next to it on purpose: these two are one expression in two languages,
and the failure when they disagree is not an error but a quiet one — rows
skipped or repeated at page boundaries, which reads as a backend bug
anywhere but here.
"""
return post.resurfaced_at or post.post_date or post.downloaded_at
class PostFeedService:
@@ -86,6 +111,13 @@ class PostFeedService:
.join(Artist, Post.artist_id == Artist.id)
.outerjoin(Source, Post.source_id == Source.id)
)
# Absorbed posts are the individual chat messages a synthetic post
# replaced (milestone 388 E2). They stay in the table — they are the
# images' true origin and the grouping has to be auditable — but the
# feed shows the post FC authored, not the dozen lines it was built
# from. `around` and `get_post` deliberately do NOT apply this: reaching
# a member by id is how you inspect a grouping.
stmt = stmt.where(Post.absorbed_by_post_id.is_(None))
if artist_id is not None:
stmt = stmt.where(Post.artist_id == artist_id)
if platform is not None:
@@ -127,15 +159,19 @@ class PostFeedService:
# Far edge in the travel direction: oldest row going older,
# newest row going newer (rows is descending for display).
edge_post = rows[-1][0] if direction == "older" else rows[0][0]
edge_key = edge_post.post_date or edge_post.downloaded_at
# Must match _sort_key() exactly, including resurfaced_at's
# precedence: a cursor built from a different expression than the
# ORDER BY silently skips or repeats rows at every page boundary.
edge_key = _post_sort_value(edge_post)
next_cursor = encode_cursor(edge_key, edge_post.id)
post_ids = [p.id for p, _, _ in rows]
thumbs_map = await self._thumbnails_for(post_ids)
atts_map = await self._attachments_for(post_ids)
links_map = await self._links_for(post_ids)
items = [
self._to_dict(post, artist, source, thumbs_map, atts_map)
self._to_dict(post, artist, source, thumbs_map, atts_map, links_map)
for post, artist, source in rows
]
return {"items": items, "next_cursor": next_cursor}
@@ -161,7 +197,7 @@ class PostFeedService:
if anchor is None:
return None
anchor_post, anchor_artist, anchor_source = anchor
anchor_key = anchor_post.post_date or anchor_post.downloaded_at
anchor_key = _post_sort_value(anchor_post)
anchor_cursor = encode_cursor(anchor_key, anchor_post.id)
older = await self.scroll(
@@ -176,6 +212,7 @@ class PostFeedService:
atts_map = await self._attachments_for([anchor_post.id])
anchor_item = self._to_dict(
anchor_post, anchor_artist, anchor_source, thumbs_map, atts_map,
await self._links_for([anchor_post.id]),
)
return {
"items": newer["items"] + [anchor_item] + older["items"],
@@ -199,58 +236,25 @@ class PostFeedService:
# the default arg.
thumbs_map = await self._thumbnails_for([post.id], limit=None)
atts_map = await self._attachments_for([post.id])
item = self._to_dict(post, artist, source, thumbs_map, atts_map)
item = self._to_dict(
post, artist, source, thumbs_map, atts_map,
await self._links_for([post.id]),
)
item["description_full"] = html_to_plain(post.description)
# Full (uncapped) translated description for the detail view (#143).
item["description_translated_full"] = post.description_translated
# Sanitized HTML body for faithful (semantic) rendering in the post view;
# detail-only (the feed list stays lightweight plain text). None when the
# post has no body. Inline `<img>` sources are remapped to locally-served
# copies (#830 Phase 2) so the body never hotlinks the public CDN.
item["description_html"] = await self._localize_inline_images(
sanitize_post_html(post.description), post.artist_id,
# Rendered body for faithful (semantic) display in the post view;
# detail-only, which is what keeps the feed list lightweight plain text
# (measured: note #3962). None when the post has no body. What
# "rendered" involves — sanitize, then repoint inline images at our own
# copies — belongs to `render_post_body`, which the provenance panel
# shares so the two surfaces cannot drift apart again (#3965).
item["description_html"] = await render_post_body(
self.session, post.description, post.artist_id,
)
item["external_links"] = await self._external_links_for(post.id)
return item
async def _localize_inline_images(
self, html: str | None, artist_id: int | None,
) -> str | None:
"""Rewrite a post body's inline `<img src=CDN>` to locally-served copies.
The join key is the CDN filehash the downloader persisted on each
ImageRecord (source_filehash): for every body image whose filehash maps
to a stored image of THIS artist, swap the src to /images/<path>. Images
we never captured (or pre-Phase-2 rows with no filehash) are left as-is —
they keep hotlinking, which is the prior behavior. Scoped to the post's
artist so one creator's body never resolves to another's file."""
if not html or artist_id is None:
return html
srcs = extract_img_srcs(html)
if not srcs:
return html
# filehash -> the raw (as-in-HTML) src strings carrying it. A body can
# repeat the same image; keep every raw form so each is substituted.
by_hash: dict[str, list[str]] = {}
for raw in srcs:
fh = filehash_from_url(unescape(raw))
if fh:
by_hash.setdefault(fh, []).append(raw)
if not by_hash:
return html
rows = (await self.session.execute(
select(ImageRecord.source_filehash, ImageRecord.path)
.where(
ImageRecord.artist_id == artist_id,
ImageRecord.source_filehash.in_(list(by_hash)),
)
)).all()
replace: dict[str, str] = {}
for fh, path in rows:
for raw in by_hash.get(fh, ()):
replace[raw] = image_url(path)
return rewrite_img_srcs(html, replace)
async def _external_links_for(self, post_id: int) -> list[dict]:
"""Off-platform file-host links recorded for a post (detail-only). Each
carries its host, full url, label, and download status so the post view
@@ -365,9 +369,25 @@ class PostFeedService:
})
return out
async def _links_for(self, post_ids: list[int]) -> dict[int, list[dict]]:
"""Accepted announcement links touching these posts (#388 E5).
Only ACCEPTED ones. A pending proposal is a question for the review
queue, not a claim to render beside the artwork — showing one here
would assert a link the operator has not agreed to, which is the exact
failure the confirm-only design exists to prevent.
"""
# Imported here rather than at module scope: post_association_service
# imports discord_grouping, which imports the models, and the feed
# service is imported by the API at startup. A local import keeps that
# chain out of the import graph for a purely optional read.
from .post_association_service import PostAssociationService
return await PostAssociationService(self.session).linked_for(post_ids)
def _to_dict(
self, post: Post, artist: Artist, source: Source | None,
thumbs_map: dict, atts_map: dict,
thumbs_map: dict, atts_map: dict, links_map: dict | None = None,
) -> dict:
plain_full = html_to_plain(post.description) if post.description else None
if plain_full is None:
@@ -400,6 +420,27 @@ class PostFeedService:
"translated_source_lang": post.translated_source_lang,
# Sticky per-post translation choice (auto/force/original, #155).
"translation_override": post.translation_override,
# Milestone 388 E2. Non-null means FC AUTHORED this post by grouping
# a creator's drop — the UI must say so wherever the post appears,
# and `synthesis` carries what it was built from so the operator can
# audit a grouping FC invented. Null for every real post; the two
# keys are always present so the frontend never branches on absence.
"synthesized_by": post.synthesized_by,
"synthesis": post.synthesis_details,
# #388 E3. A grouping stays open, so the card can say "updated N
# ago" — which is the whole signal that chat content is trickling
# in. NULL means it has not grown since it was created.
"last_grew_at": post.last_grew_at.isoformat() if post.last_grew_at else None,
# Accepted links only (#388 E5): [{role, post_id, id}], where role
# is "announces" (this post is the teaser) or "announced_by" (this
# post is the drop). Always a list so the UI never branches on
# absence.
"associations": (links_map or {}).get(post.id, []),
# Non-null on a chat message a synthetic post absorbed. The feed
# filters these out, but `around`/`get_post` still reach them, and
# the UI uses this to explain why a post it linked to is not in the
# stream.
"absorbed_by_post_id": post.absorbed_by_post_id,
"artist": {"id": artist.id, "name": artist.name, "slug": artist.slug},
"source": (
{"id": source.id, "platform": source.platform}
+20 -5
View File
@@ -18,17 +18,32 @@ from ..models import (
Source,
attachment_download_url,
)
from ..utils.html_sanitize import sanitize_post_html
from .post_body import render_post_body
def _post_dict(p: Post) -> dict:
async def _post_dict(session: AsyncSession, p: Post) -> dict:
"""One provenance entry's post.
NOTE: the key names here deliberately differ from
`PostFeedService._to_dict` (`url`/`title`/`date` vs
`post_url`/`post_title`/`post_date`, `attachment_count` vs `attachments`),
and `description_translated` is the FULL text here where the feed truncates
it to DESCRIPTION_LIMIT. That divergence is issue #3965's wider half and is
deliberately NOT addressed here — renaming is a breaking payload change for
ProvenancePanel with no second reason to spend it.
What IS fixed here is the body: this used to call `sanitize_post_html`
alone, so provenance bodies hotlinked the platform CDN for images already
on disk while the post detail view served local copies. `render_post_body`
is the whole pipeline, so the two surfaces cannot drift apart again.
"""
return {
"id": p.id,
"external_post_id": p.external_post_id,
"url": p.post_url,
"title": p.post_title,
"date": p.post_date.isoformat() if p.post_date else None,
"description_html": sanitize_post_html(p.description),
"description_html": await render_post_body(session, p.description, p.artist_id),
"attachment_count": p.attachment_count,
# Translation (#143): the English title/description shown by default when
# a translation exists; the UI toggles to the original. Source lang labels
@@ -131,7 +146,7 @@ class ProvenanceService:
"provenance_id": ip.id,
"captured_at": ip.captured_at.isoformat()
if ip.captured_at else None,
"post": _post_dict(post),
"post": await _post_dict(self.session, post),
"source": _source_dict(src) if src is not None else None,
"artist": _artist_dict(art),
}
@@ -154,7 +169,7 @@ class ProvenanceService:
return None
post, src, art = row
return {
"post": _post_dict(post),
"post": await _post_dict(self.session, post),
"source": _source_dict(src) if src is not None else None,
"artist": _artist_dict(art),
"attachments": await self._attachments_for_posts([post.id]),
+30 -1
View File
@@ -8,12 +8,13 @@ from __future__ import annotations
from datetime import UTC, datetime, timedelta
from sqlalchemy import select
from sqlalchemy import func, select
from sqlalchemy.dialects.postgresql import insert as pg_insert
from sqlalchemy.ext.asyncio import AsyncSession
from sqlalchemy.orm import selectinload
from ..models import AppSetting, Artist, ImportSettings, Source
from .db_helpers import failing_sources_clause, no_access_sources_clause
MIN_INTERVAL_SECONDS = 60
MAX_INTERVAL_SECONDS = 86400
@@ -219,10 +220,38 @@ async def scheduler_status(session: AsyncSession) -> dict:
cooldowns = await active_platform_cooldowns(session)
# Ingestion health for the front-door ribbon (#387 B3). Counted over ENABLED
# sources rather than the auto_check subset walked above: a source that is
# erroring or paywalled is worth surfacing whether or not a schedule happens
# to poll it. Two scalar COUNTs, not a second pass over `rows`.
#
# Both predicates are the shared ones, so the ribbon and the surfaces it
# links to cannot disagree about what they are counting.
failing_sources = (await session.execute(
select(func.count()).select_from(Source)
.where(Source.enabled.is_(True), failing_sources_clause())
)).scalar_one()
no_access_sources = (await session.execute(
select(func.count()).select_from(Source)
.where(Source.enabled.is_(True), no_access_sources_clause())
)).scalar_one()
# #387 B4: lets the front door tell "nothing configured yet" (a fresh
# install — show the on-ramp) apart from "configured, still fetching" (a
# first run in progress — show what's running). Telling someone to add a
# source when they already have three and are mid-backfill is worse than
# saying nothing. Deliberately NOT auto_sources, which counts only what is
# on a schedule: a source with auto_check off still means "configured".
total_sources = (await session.execute(
select(func.count()).select_from(Source).where(Source.enabled.is_(True))
)).scalar_one()
return {
"last_tick_at": last_tick_at,
"next_due_at": next_due_at.isoformat() if next_due_at else None,
"due_now": due_now,
"auto_sources": len(rows),
"failing_sources": failing_sources,
"no_access_sources": no_access_sources,
"total_sources": total_sources,
"platform_cooldowns": {p: dt.isoformat() for p, dt in cooldowns.items()},
}
+175
View File
@@ -0,0 +1,175 @@
"""The learned roster: which of FabledCurator's parts have checked in, and when.
Milestone 365. `celery inspect` answers "who is here"; this answers "who is
missing", which nothing in the application could do before — see
`models/service_seen.py` for why the identity is a queue set and not a
worker hostname.
## Who does the observing, and why it is the web process
Three candidates, and the choice matters more than the code:
* **A celery beat sweep.** Rejected. If the scheduler dies, the sweep stops,
every row goes stale, and the page reports that everything is down when one
thing is. An alarm that cannot distinguish "one part died" from "the
observer died" is worse than no alarm.
* **A background task in web.** Rejected on a detail of how this deploys:
hypercorn runs `--workers 4`, so a `before_serving` loop would be FOUR
concurrent inspect loops hammering the broker, forever, per container.
* **Refresh on demand, rate-limited by the data itself.** Taken. Whichever web
process happens to serve a health request refreshes the roster if it is
older than REFRESH_TTL, and otherwise reads what is already there.
The third has the property the other two lack: **the observer is the thing
serving the page.** If web is down you get a browser error rather than a
confidently green page, which is the honest failure. It also self-limits
without coordination — the TTL lives in the row everybody can see.
"""
from __future__ import annotations
import asyncio
import logging
from sqlalchemy import func, select
from sqlalchemy.dialects.postgresql import insert as pg_insert
from sqlalchemy.ext.asyncio import AsyncSession
from ..models import ServiceSeen
log = logging.getLogger(__name__)
# How stale the roster may be before a health request refreshes it. Comfortably
# under the staleness thresholds that decide a service is missing, so the
# verdict is never limited by how often anyone looked.
REFRESH_TTL_SECONDS = 20.0
# celery inspect is a broker round trip and this sits on a request path, so it
# gets a deadline (rule 156). A broker that has stopped answering must make the
# roster stale — which is a true statement about the system — not hang the one
# page that exists to explain it.
INSPECT_TIMEOUT_SECONDS = 2.0
# Queue set -> the name an operator recognises. Sorted-tuple keys, because the
# order celery reports them in is not guaranteed.
#
# A deployment that slices CELERY_QUEUES differently falls through to the raw
# queue list rather than being given a name this table invented for it: a
# wrong-but-confident label on a status page is worse than an ugly true one.
ROLE_NAMES: dict[tuple[str, ...], str] = {
("default", "download", "import", "thumbnail"): "Worker",
("maintenance", "scan"): "Scheduler",
("ml",): "ML worker",
}
def role_display_name(queues: tuple[str, ...]) -> str:
known = ROLE_NAMES.get(queues)
if known:
return known
return "Worker (" + ", ".join(queues) + ")"
def _inspect_celery_sync() -> dict[tuple[str, ...], dict]:
"""celery inspect, grouped by queue set rather than by worker.
Returns {queue_set: {"hostnames": [...], "active": int}}. Two replicas of
one role collapse into one entry on purpose — the question is whether the
role is being served, not how many containers exist.
"""
from ..celery_app import celery as celery_app
insp = celery_app.control.inspect(timeout=INSPECT_TIMEOUT_SECONDS)
active_queues = insp.active_queues() or {}
active_tasks = insp.active() or {}
grouped: dict[tuple[str, ...], dict] = {}
for hostname, queues in active_queues.items():
key = tuple(sorted({q["name"] for q in queues}))
entry = grouped.setdefault(key, {"hostnames": [], "active": 0})
entry["hostnames"].append(hostname)
entry["active"] += len(active_tasks.get(hostname, []))
for entry in grouped.values():
entry["hostnames"].sort()
return grouped
async def touch_service(
session: AsyncSession, *, key: str, kind: str, display_name: str, details: dict
) -> None:
"""Record that a part checked in just now.
Upsert rather than read-modify-write: several web processes and several
agents can be doing this at once, and the last writer is simply the most
recent sighting. `first_seen_at` is deliberately NOT updated — it is the
one field that answers "has this ever run", which the learned-roster design
depends on.
"""
stmt = pg_insert(ServiceSeen).values(
key=key, kind=kind, display_name=display_name, details=details,
)
stmt = stmt.on_conflict_do_update(
index_elements=[ServiceSeen.key],
set_={
"kind": stmt.excluded.kind,
"display_name": stmt.excluded.display_name,
"details": stmt.excluded.details,
"last_seen_at": func.now(),
},
)
await session.execute(stmt)
async def refresh_celery_roster(session: AsyncSession) -> None:
"""Inspect the broker and record what answered. Never raises.
A failure here means the roster does not advance, and the rows going stale
is then a TRUE report about a broker nobody can reach. Letting the
exception out would instead break the health endpoint, which is the one
thing that must keep answering when the stack is unwell.
"""
try:
grouped = await asyncio.wait_for(
asyncio.to_thread(_inspect_celery_sync),
timeout=INSPECT_TIMEOUT_SECONDS * 2,
)
except Exception:
log.warning("service roster: celery inspect failed; roster not refreshed", exc_info=True)
return
for queues, entry in grouped.items():
await touch_service(
session,
key="celery:" + ",".join(queues),
kind="celery",
display_name=role_display_name(queues),
details={
"queues": list(queues),
"hostnames": entry["hostnames"],
"replicas": len(entry["hostnames"]),
"active": entry["active"],
},
)
async def refresh_if_stale(session: AsyncSession) -> None:
"""Refresh the celery roster if nobody has for REFRESH_TTL_SECONDS.
Rate-limited by the data rather than by a lock: the gate is the newest
last_seen_at across the celery rows, which every web process can see. Two
processes racing through the gate costs one redundant inspect and writes
the same values twice, so the benign outcome needs no coordination to
prevent.
"""
newest = (
await session.execute(
select(func.max(ServiceSeen.last_seen_at)).where(ServiceSeen.kind == "celery")
)
).scalar_one_or_none()
if newest is not None:
age = (await session.execute(select(func.now()))).scalar_one() - newest
if age.total_seconds() < REFRESH_TTL_SECONDS:
return
await refresh_celery_roster(session)
+134 -3
View File
@@ -10,12 +10,16 @@ from sqlalchemy.ext.asyncio import AsyncSession
from ..models import (
Artist,
DownloadEvent,
ImageProvenance,
ImageRecord,
ImportSettings,
Post,
Source,
)
from .db_helpers import failing_sources_clause
from .gallery_dl import ErrorType
from .membership_roster import gated_reasons_for_sources
from .platforms import known_platform_keys
from .scheduler_service import compute_next_check_at
@@ -84,6 +88,17 @@ class SourceRecord:
# plan #704: cumulative posts processed across the walk's chunks — live
# progress for the badge.
backfill_posts: int
# Milestone #387 A3: posts the last walk skipped because the account can't
# view them. Lives on the EVENT (run_stats.tier_gated_count), not the
# source, so it is joined in by `list()` only — None everywhere else, which
# the UI renders as the bare no-access state with no fabricated number.
tier_gated_count: int | None = None
# Milestone #387 C5: WHY those posts are out of reach, when the learned
# roster can say — "lapsed" / "tier" / "free", or None for "no words beyond
# the count". Joined in by `list()` beside the count, and for the same
# reason: it annotates a gated state rather than producing one. Nothing on
# a fetch path may read it (see `membership_roster.gated_reason`).
gated_reason: str | None = None
def to_dict(self) -> dict:
return {
@@ -107,6 +122,8 @@ class SourceRecord:
"backfill_bypass_seen": self.backfill_bypass_seen,
"backfill_recapture": self.backfill_recapture,
"backfill_posts": self.backfill_posts,
"tier_gated_count": self.tier_gated_count,
"gated_reason": self.gated_reason,
}
@@ -114,6 +131,24 @@ class SourceRecord:
_EDITABLE = {"enabled", "url", "config_overrides", "check_interval_override", "platform"}
# `config_overrides` carries two unrelated things under one column: the
# operator's per-source download settings, and state FC writes for ITSELF. An
# operator edit replaces the first wholesale — removing a key has to be able to
# remove it — but it must never take the second with it.
#
# Two families, both matching data already on disk:
# `_*` the #693 backfill state machine (_backfill_state, _cursor,
# _cursor_stalls, _chunks, _posts, _bypass_seen, _recapture)
# `*_campaign_id` the resolved platform identity cache, written by
# `download_service._phase3_persist`
_CAMPAIGN_ID_SUFFIX = "_campaign_id"
def _is_app_managed(key: str) -> bool:
"""Is this a key FC maintains, rather than one the operator edits?"""
return key.startswith("_") or key.endswith(_CAMPAIGN_ID_SUFFIX)
# Plan #693: backfill safety cap. "Start backfill" (and a newly created
# enabled source) arms a run-until-done walk; this caps how many time-boxed
# chunks it may spend before pausing as "stalled", so a pathological walk that
@@ -156,11 +191,67 @@ class SourceService:
raise InvalidConfigError("config_overrides must be a JSON object")
return config
@staticmethod
def _merged_config(source: Source, incoming: dict | None) -> dict | None:
"""The operator's keys replace wholesale; FC's own keys survive.
Without this, one edit in the Subscriptions dialog silently discarded
the resolved campaign id AND the entire backfill position — the dialog
posts the whole object back (`SourceFormDialog`), so anything absent
from its JSON box was simply gone.
FC's keys are applied LAST so they win: the dialog round-trips whatever
it last read, and a client echoing a stale `_backfill_state` must not be
able to overwrite what the walk has since written.
"""
managed = {
k: v for k, v in (source.config_overrides or {}).items()
if _is_app_managed(k)
}
if incoming is None:
# An explicit null clears the operator's settings. It is not a
# request to forget where a backfill had got to.
return managed or None
operator = {k: v for k, v in incoming.items() if not _is_app_managed(k)}
return {**operator, **managed}
async def _load_settings(self) -> ImportSettings:
return await ImportSettings.load(self.session)
async def _tier_gated_counts(self, source_ids: list[int]) -> dict[int, int]:
"""Latest walk's tier-gated post count, per source, in ONE query.
Selects the `run_stats` sub-object rather than whole `metadata` blobs:
those carry truncated stdout/stderr up to 500KB each, and pulling one
per source to read a single integer would make the subscriptions list
pay for the Logs view. DISTINCT ON + ORDER BY takes the newest event per
source (Postgres-only, like the rest of this codebase).
Callers pass only the sources that actually need it — the count is
meaningless for a source that isn't tier-gated.
"""
if not source_ids:
return {}
rows = (await self.session.execute(
select(
DownloadEvent.source_id,
DownloadEvent.metadata_["run_stats"],
)
.where(DownloadEvent.source_id.in_(source_ids))
.distinct(DownloadEvent.source_id)
.order_by(DownloadEvent.source_id, DownloadEvent.started_at.desc())
)).all()
counts: dict[int, int] = {}
for source_id, run_stats in rows:
n = (run_stats or {}).get("tier_gated_count") or 0
if n:
counts[source_id] = int(n)
return counts
def _build_record(
self, source: Source, artist: Artist, settings: ImportSettings,
gated_counts: dict[int, int] | None = None,
gated_reasons: dict[int, str] | None = None,
) -> SourceRecord:
nxt = compute_next_check_at(source, artist, settings)
co = source.config_overrides or {}
@@ -185,6 +276,8 @@ class SourceService:
backfill_bypass_seen=bool(co.get("_backfill_bypass_seen")),
backfill_recapture=bool(co.get("_backfill_recapture")),
backfill_posts=int(co.get("_backfill_posts", 0)),
tier_gated_count=(gated_counts or {}).get(source.id),
gated_reason=(gated_reasons or {}).get(source.id),
)
async def _row_to_record(self, source: Source) -> SourceRecord:
@@ -210,14 +303,24 @@ class SourceService:
stmt = stmt.where(~Source.url.like("sidecar:%"))
if failing:
# Worst-first so the rollup card surfaces the loudest failures.
stmt = stmt.where(Source.consecutive_failures > 0).order_by(
# Shared clause: the front-door ribbon counts with the same one, so
# it can never report a number this list then contradicts.
stmt = stmt.where(failing_sources_clause()).order_by(
Source.consecutive_failures.desc(), Artist.name.asc(),
)
else:
stmt = stmt.order_by(Artist.name.asc(), Source.id.asc())
rows = (await self.session.execute(stmt)).all()
settings = await self._load_settings()
return [self._build_record(s, a, settings) for s, a in rows]
# Only tier-gated rows need either join — on a healthy library that is
# an empty list and both helpers short-circuit without a query.
gated = [s for s, _a in rows if s.error_type == ErrorType.TIER_LIMITED]
gated_counts = await self._tier_gated_counts([s.id for s in gated])
gated_reasons = await gated_reasons_for_sources(self.session, gated)
return [
self._build_record(s, a, settings, gated_counts, gated_reasons)
for s, a in rows
]
async def get(self, source_id: int) -> SourceRecord | None:
source = (await self.session.execute(
@@ -301,10 +404,38 @@ class SourceService:
if "url" in fields:
fields["url"] = self._validate_url(fields["url"])
if "config_overrides" in fields:
fields["config_overrides"] = self._validate_config(fields["config_overrides"])
fields["config_overrides"] = self._merged_config(
source, self._validate_config(fields["config_overrides"])
)
# Computed BEFORE the setattr loop, while `source.url` is still the old
# one. See the invalidation below.
url_changed = "url" in fields and fields["url"] != source.url
for key, value in fields.items():
setattr(source, key, value)
if url_changed:
# Repointing a source at a different creator makes a cached campaign
# id WRONG, not merely stale, and `patreon_resolver` consults that
# cache BEFORE attempting any lookup — so a kept id would resolve the
# old creator forever, and the membership join (#387 C4) would report
# a confident wrong match.
#
# This needs saying explicitly only because of the merge above: until
# then the wholesale overwrite wiped the id as an accident of the
# bug, which masked this. Preserving the id makes the invalidation
# this service's job.
#
# The backfill cursor is deliberately NOT cleared. It is opaque
# platform state the walk already validates with its own stall
# guard, and dropping it would restart a long backfill over a
# cosmetic URL edit (http->https, adding `/c/`).
co = dict(source.config_overrides or {})
for key in [k for k in co if k.endswith(_CAMPAIGN_ID_SUFFIX)]:
co.pop(key)
source.config_overrides = co
# Disabling a source clears its failure state (operator: disable the subs
# you're not paying for without them lingering as "failing"). Re-enabling
# then starts clean; the next real run re-derives health. Only on the
@@ -43,6 +43,7 @@ import requests
from ..utils.paths import filehash_from_url
from .native_ingest_common import (
_MAX_429_RETRIES,
Membership,
NativeAuthError,
NativeDriftError,
NativeIngestError,
@@ -297,6 +298,206 @@ def _extract_creator_name(html: str) -> str | None:
return name or None
# -- membership roster (#387 D1) ------------------------------------------
#
# Characterized from a live operator capture of the account's /subscriptions
# page, 2026-09-13 — Scribe note #3989. Read that note before changing any of
# this; each constant below is a finding from it, not a guess.
# The account page is fetched from `.adult`. The `.art` age wall never clears
# with the 18+ cookie for FC's requests (see _normalize_ss_host, issues #1259 /
# #1284). The capture itself came from `.art` only because a human had clicked
# through the gate in the browser.
_ROSTER_BASE = "https://subscribestar.adult"
_ROSTER_URL = f"{_ROSTER_BASE}/subscriptions"
# Two tables, and WHICH table a creator sits in is the only status the page
# gives — there is no per-row status word. Keyed on each card's
# `data-identifier`, the one vocabulary that names a state: the table class
# inside the cancelled card says `for-unsubscribed_users`, a different word for
# the same list (note #3989, CORRECTION 1). The identifier is stored verbatim as
# Membership.status and mapped in membership_roster.MEMBERSHIP_STATUS.
_ROSTER_ACTIVE = "active_subscriptions"
_ROSTER_CANCELLED = "cancelled_subscriptions"
_ROSTER_ROW_OPEN = '<td class="for-name">'
# Active rows nest a second `<tr class="for-actions">` INSIDE the row's own
# <tr> — a narrow-screen duplicate of the actions cell. Its <td>s are not
# columns, so every row is cut here before its cells are read.
_ROSTER_NESTED_ROW = '<tr class="for-actions"'
_ROSTER_HREF_RE = re.compile(r'<a href="/([^"/?#]+)"')
_ROSTER_USER_ID_RE = re.compile(r'data-user-id="([^"]*)"')
_ROSTER_NAME_RE = re.compile(r"<img [^>]*>([^<]*)</div>")
_ROSTER_HEAD_RE = re.compile(r'<th class="[^"]*"[^>]*>(.*?)</th>', re.DOTALL)
_ROSTER_CELL_RE = re.compile(r'<td class="[^"]*"[^>]*>(.*?)</td>', re.DOTALL)
_ROSTER_PAGE_LINK_RE = re.compile(r'href="[^"]*[?&]page=\d')
_TAG_RE = re.compile(r"<[^>]+>")
# Columns that hold identity or controls rather than facts about the
# subscription, so they stay out of `details`. Matched on the header's own text,
# lowercased — the page's words, not ours.
_ROSTER_SKIP_COLUMNS = frozenset({"profile", "updates", "actions"})
def _cell_text(fragment: str) -> str:
"""Visible text of a cell: tags dropped, entities decoded, whitespace folded.
Decoding matters here specifically: an active row with no Discord link
renders its cell as the entity `&mdash;`, not as an empty cell.
"""
return " ".join(unescape(_TAG_RE.sub(" ", fragment)).split())
def _roster_table(html: str, identifier: str) -> tuple[str, str] | None:
"""One roster card: (its table markup, whatever trails `</table>` inside it).
None when the card is absent. The trailing part is returned rather than
discarded because it is the pagination check: in the characterized page a
card closes the moment its table does.
"""
start = html.find(f'data-identifier="{identifier}"')
if start < 0:
return None
end = html.find("</table>", start)
if end < 0:
raise SubscribeStarDriftError(
f"SubscribeStar roster card {identifier!r} has no table"
)
close = html.find("</div>", end)
trailing = html[end + len("</table>"): close if close >= 0 else len(html)]
return html[start:end], trailing
def _roster_rows(table: str, identifier: str, base: str) -> list[Membership]:
labels = [_cell_text(h).lower() for h in _ROSTER_HEAD_RE.findall(table)]
body = table[table.find("<tbody>"):] if "<tbody>" in table else ""
starts = [m.start() for m in re.finditer(re.escape(_ROSTER_ROW_OPEN), body)]
rows = []
for n, start in enumerate(starts):
row = body[start: starts[n + 1] if n + 1 < len(starts) else len(body)]
row = row.split(_ROSTER_NESTED_ROW, 1)[0]
href = _ROSTER_HREF_RE.search(row)
if href is None:
raise SubscribeStarDriftError(
f"SubscribeStar roster row in {identifier!r} has no creator link"
)
# The creator's numeric id, NOT the slug, is the key (note #3989,
# CORRECTION 2). A slug re-keys when a creator renames; the old row then
# stops appearing, and a disappearance is exactly what reconciliation
# reads as a lapse. The id survives a rename.
user_id = _ROSTER_USER_ID_RE.search(row)
if user_id is None or not user_id.group(1).isdigit():
raise SubscribeStarDriftError(
f"SubscribeStar roster row in {identifier!r} has no numeric "
f"data-user-id — a membership that cannot be attributed to a "
f"creator is not usable"
)
name = _ROSTER_NAME_RE.search(row)
slug = unescape(href.group(1))
cells = _ROSTER_CELL_RE.findall(row)
if len(cells) != len(labels):
# Canary, not a refusal. Identity above does not depend on columns,
# so a shifted column must not fail the whole roster — but it would
# silently mislabel `details` (a price filed under "discord"), so
# say so in the worker log where it is diagnosable.
log.warning(
"SubscribeStar roster %r: %d cells against %d headers — column "
"details may be mislabelled; markup likely changed (note #3989)",
identifier, len(cells), len(labels),
)
rows.append(Membership(
campaign_id=user_id.group(1),
display_name=(_cell_text(name.group(1)) if name else "") or None,
url=f"{base}/{slug}",
vanity=slug,
status=identifier,
# No free-follow concept on this page (#3970 §2: False when a
# platform has none).
is_free_member=False,
# Tier names live behind a per-row modal, not inline. Fetching every
# modal would be N authenticated requests for a field nothing reads.
tier_names=[],
# Deliberately NOT parsed from the price cell: a bare `$` names no
# currency, and a page price is not proven to be the charge (#3970
# finding 4). None keeps "unknown" distinct from zero. The raw text
# is kept in `details`.
amount_cents=None,
currency=None,
details={
# Paired with the header text by POSITION: two columns share the
# `for-date` class, and the updates column's <td> does not carry
# its <th>'s class at all.
"columns": {
label: _cell_text(cell)
for label, cell in zip(labels, cells, strict=False)
if label not in _ROSTER_SKIP_COLUMNS
},
},
))
return rows
def parse_subscriptions_page(html: str, *, base: str = _ROSTER_BASE) -> list[Membership]:
"""Every membership on the account's /subscriptions page.
Refuses rather than guessing, because every conclusion downstream is drawn
from ABSENCE — a roster that comes back short reads as "you cancelled
those". So this raises when:
* the active card is missing — as SubscribeStarAuthError if the page is a
login or age wall (the fix is credentials), otherwise as drift;
* a row has no creator link or no numeric creator id;
* anything renders after a card's table, or the page carries a `page=` link.
Both cards are paginatable (`data-view="app#embed_pagination"`), and the
characterized account was too small to show what pagination looks like —
so possible pagination is treated as a roster FC cannot prove complete.
A missing cancelled card is NOT drift: an account that has never cancelled
plausibly has no such table. A creator present in both tables is reported
once, as active — a current subscription is the fact that matters.
"""
active = _roster_table(html, _ROSTER_ACTIVE)
if active is None:
if any(marker in html for marker in _LOGIN_MARKERS):
raise SubscribeStarAuthError(
"SubscribeStar served a login/age wall instead of the "
"subscriptions page (cookies expired or age cookie missing)"
)
raise SubscribeStarDriftError(
f"SubscribeStar subscriptions page has no {_ROSTER_ACTIVE!r} card "
f"{_describe_page(html)}"
)
roster_region = html[html.find(f'data-identifier="{_ROSTER_ACTIVE}"'):]
if _ROSTER_PAGE_LINK_RE.search(roster_region):
raise SubscribeStarDriftError(
"SubscribeStar subscriptions page carries a page= link — the roster "
"may be paginated, and FC cannot prove it is complete (note #3989)"
)
memberships: list[Membership] = []
seen: set[str] = set()
for identifier, found in (
(_ROSTER_ACTIVE, active),
(_ROSTER_CANCELLED, _roster_table(html, _ROSTER_CANCELLED)),
):
if found is None:
continue
table, trailing = found
if trailing.strip():
raise SubscribeStarDriftError(
f"SubscribeStar roster card {identifier!r} renders content after "
f"its table — possibly pagination, so the roster cannot be "
f"proven complete (note #3989)"
)
for membership in _roster_rows(table, identifier, base):
if membership.campaign_id in seen:
continue
seen.add(membership.campaign_id)
memberships.append(membership)
return memberships
class SubscribeStarClient:
"""Synchronous SubscribeStar HTML-scrape read client. Construct with a path
to a Netscape cookies.txt (the same file CredentialService.get_cookies_path
@@ -645,6 +846,23 @@ class SubscribeStarClient:
return None
return _extract_creator_name(html)
# -- membership roster (#387 D1) ----------------------------------------
def iter_memberships(self, user_id: str | None = None) -> Iterator[Membership]:
"""Yield every subscription the account holds (note #3989).
`user_id` exists for the seam's signature (note #3970) and is ignored:
the page is the logged-in account's own, so there is nothing to resolve.
The sweep only resolves an id for a client that exposes
`current_user_id`, which this one does not.
One request, and the whole page is parsed before anything is yielded, so
a drift error can never leave a caller holding part of a roster.
"""
self._session.headers["Referer"] = f"{_ROSTER_BASE}/"
resp = self._get(_ROSTER_URL)
yield from parse_subscriptions_page(resp.text or "", base=_ROSTER_BASE)
# -- verify ------------------------------------------------------------
def verify_auth(self, campaign_id: str) -> tuple[bool | None, str]:
+196
View File
@@ -1131,3 +1131,199 @@ def vacuum_analyze() -> dict:
done.append(table)
log.info("vacuum_analyze complete: %s", done)
return {"vacuumed": done}
@celery.task(
name="backend.app.tasks.maintenance.group_discord_drops",
soft_time_limit=1800, time_limit=2100,
)
def group_discord_drops() -> str:
"""Milestone 388 E2: group Discord message-posts into the drops FC authors.
Lives on the MAINTENANCE lane, not the ml lane, even though it reads SigLIP
vectors — it does no inference and imports no ML library, and the ml-worker
is an OPTIONAL container (B3). Routing it to 'ml' would silently disable
grouping on every stack that runs a GPU agent and drops that container,
which is the same trap gpu_queue.py was moved here to avoid.
Async body under its own loop, per the _async_session contract: the sweep
needs pgvector column reads and the shared services are async.
"""
import asyncio
from ..services.discord_grouping import sweep
from ._async_session import async_session_factory
async def _run() -> dict:
async_factory, engine = async_session_factory()
try:
async with async_factory() as session:
result = await sweep(session)
await session.commit()
return result
finally:
await engine.dispose()
res = asyncio.run(_run())
if not res["enabled"]:
return "disabled"
return (
f"sources={res['sources']} created={res['posts_created']} "
f"joined={res['images_joined']}"
)
@celery.task(
name="backend.app.tasks.maintenance.match_post_associations",
soft_time_limit=900, time_limit=1200,
)
def match_post_associations() -> str:
"""Milestone 388 E5: propose which Patreon post announced which Discord drop.
Proposes only — every pair lands in a review queue and nothing is linked
until the operator accepts. Maintenance lane for the same reason as the
grouper: no inference, no ML library, and it must not depend on the
optional ml-worker being present.
"""
import asyncio
from ..services.post_association_service import rescan
from ._async_session import async_session_factory
async def _run() -> dict:
async_factory, engine = async_session_factory()
try:
async with async_factory() as session:
result = await rescan(session)
await session.commit()
return result
finally:
await engine.dispose()
res = asyncio.run(_run())
if not res["enabled"]:
return "disabled"
return f"scanned={res['scanned']} proposed={res['proposed']}"
# The wall-clock budget for ONE platform's roster walk. Rule 156 distinguishes
# this from the per-REQUEST timeout the client already has: a paginated roster
# behind an endpoint that answers every page slowly-but-within-timeout would
# never trip that one, and would sit on a worker indefinitely. This is the wait
# that bounds the whole walk.
MEMBERSHIP_SYNC_BUDGET_SECONDS = 240.0
@celery.task(
name="backend.app.tasks.maintenance.sync_memberships",
soft_time_limit=900, time_limit=1200,
)
def sync_memberships() -> str:
"""Milestone 387 C3: walk each platform's membership roster into the DB.
Daily, because memberships change on a BILLING cycle rather than a download
cadence — polling an account-scoped endpoint more often would be both
pointless and less polite than the browser.
Rule 89's four, and where each actually lives:
* recovery — `touch_membership` is an upsert and never deletes, so
recovery is simply the next run; a dead run leaves the previous roster
intact rather than a half-written one (`sync_platform` completes the
fetch before writing anything).
* retention — per C1, rows age out and are never deleted on
disappearance, because disappearing IS the signal C4 reads.
* wall-clock timeout — MEMBERSHIP_SYNC_BUDGET_SECONDS per platform,
plus the task's own soft/hard limits.
* duration tracking — the TaskRun celery-signal plumbing, same as every
other sweep here.
Rule 164: this is the rule's own named exception — a genuinely external
feature that may call out. It never gates startup, and a failure leaves a
VISIBLE stale state (membership_sync) rather than an empty roster that
reads as "you subscribe to nothing".
"""
import asyncio
from ..services.artist_membership_service import rescan as membership_rescan
from ..services.credential_crypto import CredentialCrypto
from ..services.credential_service import CredentialService
from ..services.membership_roster import roster_user_id, sync_platform
from ..services.patreon_client import PatreonClient
from ..services.subscribestar_client import SubscribeStarClient
from ._async_session import async_session_factory
key_path = IMAGES_ROOT / "secrets" / "credential_key.b64"
# platform -> how to build a client from a cookies path. A platform is in
# the sweep only if it is here AND its client exposes `iter_memberships`
# AND a credential exists — three independent gates, each silent, so a
# platform is added with one line here and nothing else. SubscribeStar (D1)
# was the second; the only other change it needed was `roster_user_id`
# replacing an unconditional Patreon-only call below.
builders = {"patreon": PatreonClient, "subscribestar": SubscribeStarClient}
async def _run() -> dict:
async_factory, engine = async_session_factory()
results = []
try:
for platform, build in builders.items():
async with async_factory() as session:
cred = CredentialService(session, CredentialCrypto(key_path))
cookies = await cred.get_cookies_path(platform)
if cookies is None:
# No credential is not an error — the operator simply has
# not connected this platform. Recording a failure here
# would light up the UI for a feature they never enabled.
results.append({"platform": platform, "skipped": "no credential"})
continue
client = build(str(cookies))
if getattr(client, "iter_memberships", None) is None:
# The seam, probed not required (#387 C2). A client without
# it makes the feature invisible for that platform — no
# flag, no config row, no "unsupported" branch.
results.append({"platform": platform, "skipped": "no seam"})
continue
async def fetch(_client=client):
# The client is sync (`requests`); run it off the loop so a
# slow roster does not block the event loop, and bound the
# whole walk rather than only its individual requests.
def _walk():
return list(_client.iter_memberships(roster_user_id(_client)))
return await asyncio.wait_for(
asyncio.to_thread(_walk),
timeout=MEMBERSHIP_SYNC_BUDGET_SECONDS,
)
async with async_factory() as session:
results.append(
await sync_platform(session, platform=platform, fetch=fetch)
)
# #388 E4: offer the freshly-synced roster to the artists FC already
# tracks. Chained here rather than given its own beat entry because
# a suggestion can only be as good as the roster behind it — running
# it on any other cadence would just propose from staler data.
suggested = None
if any(r.get("ok") for r in results):
async with async_factory() as session:
suggested = (await membership_rescan(session))["proposed"]
await session.commit()
return {"results": results, "suggested": suggested}
finally:
await engine.dispose()
res = asyncio.run(_run())
parts = []
for r in res["results"]:
if "skipped" in r:
parts.append(f"{r['platform']}=skipped({r['skipped']})")
elif r.get("ok"):
parts.append(f"{r['platform']}={r['count']}")
else:
parts.append(f"{r['platform']}=FAILED({r['error']})")
if res.get("suggested") is not None:
parts.append(f"suggested={res['suggested']}")
return " ".join(parts) or "no platforms"
+15 -8
View File
@@ -198,14 +198,21 @@ per `docs/process.md`'s "add deps to the image when used by >1 project".
refresh from being undone.
- **`pull: true` on the scheduled path only** is the mechanism: a moved base
tag changes the `FROM` layer's cache key and everything above it rebuilds.
**It does not currently make the unmoved case free.** Measured on the first
real fire (run 4934, 2026-08-30): every content step reported `CACHED` and
the bases resolved to unchanged digests, yet all three `:latest` tags got a
NEW manifest digest, because buildkit mints a fresh image config per run and
republishes identical layers under it. So `:latest` is rewritten weekly
whether or not anything changed, and `:c-<sha>` is handed a new manifest to
diverge from on the same cadence — a digest change stops meaning anything.
Tracked as #3265; the likely fix is a deterministic `SOURCE_DATE_EPOCH`.
It did not always make the unmoved case free. Measured on the first real
fire (run 4934, 2026-08-30): every content step reported `CACHED` and the
bases resolved to unchanged digests, yet all three `:latest` tags got a NEW
manifest digest, because buildkit stamps a fresh image config per run and
republishes identical layers under it — so `:latest` was rewritten weekly
whether or not anything changed, and a digest change stopped meaning
anything (#3265).
- **`SOURCE_DATE_EPOCH` is what makes it free.** Set on each build step from
`artifacts.sh epoch <artifact>` — the unix timestamp of the same commit
`revision` and `version` name, so all three are views of one `newest()`
lookup and cannot drift into disagreeing. With the config's `created` field
and history timestamps pinned to the content rather than to the wall clock,
identical source produces an identical manifest digest and the push is a
registry no-op. That restores the property the whole scheme rests on: a
channel tag's digest changes when, and only when, its content does.
Separately not caught: a Debian package update inside the `apt-get install`
layer while the base tag stands still — a lag rather than a hole, since the
official python/cuda images rebuild with those updates baked in.
+46 -15
View File
@@ -1,12 +1,21 @@
# Base compose stack. Uses ${VAR:-default} interpolation throughout so the
# stack boots with zero config — sane dev defaults baked in. For production
# deployments, override the defaults via shell env vars or a .env file:
# Base compose stack, and the install path. Uses ${VAR:-default} throughout so
# the stack boots with zero config — but those defaults are DEV defaults, and
# two of them (DB_PASSWORD, SECRET_KEY) are published in this file. Copy
# .env.example to .env and set them before running this anywhere real.
#
# DB_PASSWORD=...real... SECRET_KEY=...real... docker compose up
# To run FabledCurator:
#
# The dev override (docker-compose.override.yml) is auto-merged when you
# run `docker compose up` from this directory and switches images to
# local builds + DEBUG logging.
# docker compose -f docker-compose.yml up -d
#
# The -f is load-bearing. Without it Compose auto-merges
# docker-compose.override.yml, which replaces every image: with a local
# build: and turns on DEBUG logging — the contributor path. Naming this file
# explicitly skips the override and pulls the published :latest images.
#
# FabledCurator has no authentication. Whatever can reach ${PORT} is an
# administrator, including over the stored Patreon/SubscribeStar session
# cookies. Do not publish this port beyond a network you trust — see
# "Before you expose it" in README.md.
# Rolling-deploy safety (Swarm / `docker stack deploy`): update one task at a
# time, START the new task before stopping the old (zero-downtime via the ingress
@@ -123,18 +132,40 @@ services:
CELERY_BROKER_URL: redis://redis:6379/0
CELERY_RESULT_BACKEND: redis://redis:6379/0
SECRET_KEY: ${SECRET_KEY:-dev_secret_key_not_for_production_change_me}
EXTENSION_API_KEY: ${EXTENSION_API_KEY:-}
LOG_LEVEL: ${LOG_LEVEL:-INFO}
# First boot only. FabledCurator refuses to start until the credential
# encryption key at /images/secrets/credential_key.b64 exists, and
# refuses to create one unless told to — auto-creating is
# indistinguishable from a restore that lost ./images/secrets, where it
# would mint a key that decrypts nothing and leave an instance that looks
# healthy while every paywalled download fails.
#
# Passed through EXPLICITLY because a variable in `.env` is only used for
# ${...} interpolation; it does not reach the container unless it is
# named here. Defaulted to empty so the refusal stands for everyone who
# has not opted in — the app tests for exactly "1".
#
# Set it in .env for one `up`, then remove it. See .env.example.
CURATOR_BOOTSTRAP_NEW_KEY: ${CURATOR_BOOTSTRAP_NEW_KEY:-}
volumes:
- ./images:/images
- ./import:/import
# FC-5 legacy migration: bind-mount the host's ImageRepo images dir
# under /import (FC's existing filesystem scan picks them up). Read-only
# is sufficient — FC copies into /images during the scan. The worker +
# scheduler services see the same /import via their own mounts below
# because of /import volume reuse. Edit the host path to match your
# install before running Settings → Maintenance → Legacy migration.
# - /var/lib/imagerepo/images:/import/imagerepo:ro
# /import is a staging area for scripting a one-off ingest of a library
# you already have on disk. Drop files in ./import, or bind-mount an
# existing directory under it as below, then trigger the scan:
#
# curl -X POST http://localhost:8080/api/import/trigger
#
# Read-only is sufficient — FC copies into /images during the scan. The
# worker + scheduler services mount the same /import so the scan can run
# on whichever lane picks it up.
#
# Deliberately has no UI. The manual-scan surface was retired 2026-07-02
# once imports arrived via subscriptions + the extension, and the call
# not to restore it stands (operator, 2026-09-02): folder ingestion
# brings complexity the product does not need. The endpoint stays as an
# unsupported escape hatch; the supported way in is Subscriptions.
# - /srv/media/my-library:/import/my-library:ro
depends_on:
postgres: { condition: service_healthy }
redis: { condition: service_healthy }
+17 -115
View File
@@ -1,32 +1,35 @@
/**
* Background script — message router + Discord token capture
* (webRequest) + Pixiv PKCE OAuth. Direct port of GS background.js;
* api.js client points at FC instead of GS.
* Background script — message router + Discord token capture (webRequest).
* Direct port of GS background.js; api.js client points at FC instead of GS.
*
* pixiv's PKCE OAuth flow lived here until FC retired pixiv (milestone #406).
* It was removed together with pixiv's host permissions rather than left
* behind: a webRequest listener on a host the manifest no longer grants is at
* best dead and at worst a startup failure for the whole background script.
*/
let discordToken = null;
let discordTokenCapturedAt = null;
let pixivRefreshToken = null;
let pixivTokenCapturedAt = null;
let pixivOAuthPending = null;
const PIXIV_CLIENT_ID = 'MOBrBDS8blbauoSck0ZfDbtuzpyT';
const PIXIV_CLIENT_SECRET = 'lsACyCD94FhDUtGTXi3QzcFE2uU1hqtDaKeqrdwj';
const PIXIV_OAUTH_URL = 'https://app-api.pixiv.net/web/v1/login';
const PIXIV_TOKEN_URL = 'https://oauth.secure.pixiv.net/auth/token';
const PIXIV_REDIRECT_URI = 'https://app-api.pixiv.net/web/v1/users/auth/pixiv/callback';
let initialized = false;
async function ensureInitialized() {
if (initialized) return;
await api.init();
await loadDiscordToken();
await loadPixivToken();
await forgetRetiredPixivToken();
initialized = true;
}
// A browser that authenticated pixiv before the retirement still holds a live
// OAuth refresh token in extension storage. Nothing reads it any more, and a
// credential for a service FC no longer uses is a liability with no benefit
// (the same reasoning as the server-side cleanup, issue #3980). Removing keys
// that are absent is a no-op, so this is safe on every startup.
async function forgetRetiredPixivToken() {
await browser.storage.local.remove(['pixivRefreshToken', 'pixivTokenCapturedAt']);
}
browser.runtime.onInstalled.addListener(() => ensureInitialized());
browser.runtime.onStartup.addListener(() => ensureInitialized());
ensureInitialized().catch(e => console.error('init failed:', e));
@@ -141,98 +144,6 @@ async function saveDiscordToken(token) {
await browser.storage.local.set({ discordToken: token, discordTokenCapturedAt });
}
// ---- Pixiv PKCE OAuth ----
async function loadPixivToken() {
const s = await browser.storage.local.get(['pixivRefreshToken', 'pixivTokenCapturedAt']);
pixivRefreshToken = s.pixivRefreshToken || null;
pixivTokenCapturedAt = s.pixivTokenCapturedAt || null;
}
async function savePixivToken(token) {
pixivRefreshToken = token;
pixivTokenCapturedAt = new Date().toISOString();
await browser.storage.local.set({ pixivRefreshToken: token, pixivTokenCapturedAt });
}
function generateCodeVerifier() {
const a = new Uint8Array(32);
crypto.getRandomValues(a);
return base64UrlEncode(a);
}
async function generateCodeChallenge(verifier) {
const data = new TextEncoder().encode(verifier);
const hash = await crypto.subtle.digest('SHA-256', data);
return base64UrlEncode(new Uint8Array(hash));
}
function base64UrlEncode(buf) {
return btoa(String.fromCharCode(...buf)).replace(/\+/g, '-').replace(/\//g, '_').replace(/=/g, '');
}
async function initiatePixivOAuth() {
const codeVerifier = generateCodeVerifier();
const codeChallenge = await generateCodeChallenge(codeVerifier);
const params = new URLSearchParams({
code_challenge: codeChallenge,
code_challenge_method: 'S256',
client: 'pixiv-android',
});
const tab = await browser.tabs.create({ url: `${PIXIV_OAUTH_URL}?${params}` });
return new Promise((resolve, reject) => {
pixivOAuthPending = { codeVerifier, tabId: tab.id, resolve, reject };
setTimeout(() => {
if (pixivOAuthPending && pixivOAuthPending.tabId === tab.id) {
pixivOAuthPending = null;
reject(new Error('Pixiv OAuth timed out (5 min)'));
}
}, 5 * 60 * 1000);
});
}
browser.webRequest.onBeforeRedirect.addListener(
async (details) => {
if (!pixivOAuthPending) return;
if (details.tabId !== pixivOAuthPending.tabId) return;
const url = new URL(details.redirectUrl);
const code = url.searchParams.get('code');
if (!code) return;
const verifier = pixivOAuthPending.codeVerifier;
const resolve = pixivOAuthPending.resolve;
const reject = pixivOAuthPending.reject;
pixivOAuthPending = null;
try {
const tokenResp = await fetch(PIXIV_TOKEN_URL, {
method: 'POST',
headers: { 'Content-Type': 'application/x-www-form-urlencoded' },
body: new URLSearchParams({
client_id: PIXIV_CLIENT_ID,
client_secret: PIXIV_CLIENT_SECRET,
code,
code_verifier: verifier,
grant_type: 'authorization_code',
include_policy: 'true',
redirect_uri: PIXIV_REDIRECT_URI,
}),
});
const body = await tokenResp.json();
if (!body.refresh_token) {
reject(new Error(`Pixiv token exchange failed: ${JSON.stringify(body)}`));
return;
}
await savePixivToken(body.refresh_token);
try { await browser.tabs.remove(details.tabId); } catch {}
resolve(body.refresh_token);
} catch (e) {
reject(e);
}
},
{ urls: ['https://app-api.pixiv.net/web/v1/users/auth/pixiv/callback*'] },
);
// Extract → verify → upload one cookie-auth platform. Returns a structured
// outcome so the two callers (EXPORT_COOKIES single, EXPORT_ALL_COOKIES) shape
// their own response + skip semantics. Verifies the captured cookies are
@@ -277,8 +188,6 @@ browser.runtime.onMessage.addListener(async (msg) => {
}
} else if (key === 'discord') {
status[key] = { hasToken: !!discordToken, capturedAt: discordTokenCapturedAt };
} else if (key === 'pixiv') {
status[key] = { hasToken: !!pixivRefreshToken, capturedAt: pixivTokenCapturedAt };
} else {
status[key] = {};
}
@@ -306,13 +215,6 @@ browser.runtime.onMessage.addListener(async (msg) => {
await api.uploadCredentials('discord', 'token', discordToken);
return { success: true };
}
if (key === 'pixiv') {
if (!pixivRefreshToken) {
await initiatePixivOAuth();
}
await api.uploadCredentials('pixiv', 'token', pixivRefreshToken);
return { success: true };
}
return { error: 'Unsupported platform.' };
} catch (e) {
return { error: e.message };
-9
View File
@@ -60,14 +60,6 @@ const PLATFORMS = {
urlPattern: /^https?:\/\/(www\.)?discord\.com/,
note: 'Open Discord in browser to capture token',
},
pixiv: {
name: 'Pixiv',
domains: ['.pixiv.net', 'www.pixiv.net', 'pixiv.net'],
authType: 'token',
color: '#0096FA',
urlPattern: /^https?:\/\/(www\.)?pixiv\.net/,
note: 'Click to authenticate via OAuth',
},
};
/**
@@ -88,7 +80,6 @@ const PLATFORM_ARTIST_PATTERNS = {
patreon: /^https?:\/\/(www\.)?patreon\.com\/(?:cw\/|c\/)?(?!(?:home|search|messages|notifications|library|settings|posts)(?:[\/?#]|$))[^/?#]+/i,
subscribestar: /^https?:\/\/(www\.)?subscribestar\.(com|adult)\/(?!feed$|messages$|library$)[^/?#]+\/?$/i,
hentaifoundry: /^https?:\/\/(www\.)?hentai-foundry\.com\/user\/[^/?#]+/i,
pixiv: /^https?:\/\/(www\.)?pixiv\.net\/(en\/)?users\/\d+/i,
};
function getPlatformFromUrl(url) {
+1 -5
View File
@@ -32,9 +32,6 @@
"*://*.subscribestar.adult/*",
"*://*.hentai-foundry.com/*",
"*://*.discord.com/*",
"*://*.pixiv.net/*",
"*://app-api.pixiv.net/*",
"*://oauth.secure.pixiv.net/*",
"*://*/*"
],
@@ -59,8 +56,7 @@
"*://*.patreon.com/*",
"*://*.subscribestar.com/*",
"*://*.subscribestar.adult/*",
"*://*.hentai-foundry.com/*",
"*://*.pixiv.net/*"
"*://*.hentai-foundry.com/*"
],
"js": ["lib/platforms.js", "content/content-script.js"],
"css": ["content/content-script.css"],
+1 -3
View File
@@ -108,7 +108,7 @@ function createPlatformCard(key, platform, status) {
card.className = 'platform-card';
card.dataset.platform = key;
const isTokenOnly = platform.authType === 'token' && !['discord', 'pixiv'].includes(key);
const isTokenOnly = platform.authType === 'token' && key !== 'discord';
const discordNeedsToken = key === 'discord' && !status.hasToken;
if (isTokenOnly || discordNeedsToken) card.classList.add('disabled');
@@ -141,7 +141,6 @@ function createPlatformCard(key, platform, status) {
function statusText(s, platform, key) {
if (key === 'discord') return s.hasToken ? 'Token captured — ready' : 'Open Discord to capture token';
if (key === 'pixiv') return s.hasToken ? 'Token captured — ready' : 'Click to authenticate via OAuth';
if (platform.authType === 'token') return 'Manual token entry required';
if (s.error) return 'Error checking cookies';
if (!s.hasCookies || !s.cookieCount) return 'No cookies — log in first';
@@ -149,7 +148,6 @@ function statusText(s, platform, key) {
}
function statusClass(s, platform, key) {
if (key === 'discord') return s.hasToken ? 'ready' : 'no-cookies';
if (key === 'pixiv') return s.hasToken ? 'ready' : 'no-cookies';
if (platform.authType === 'token') return 'no-cookies';
if (s.error) return 'error';
if (!s.hasCookies || !s.cookieCount) return 'no-cookies';
+13 -11
View File
@@ -18,7 +18,6 @@ describe('getPlatformFromUrl', () => {
expect(getPlatformFromUrl('https://subscribestar.adult/someone')).toBe('subscribestar')
expect(getPlatformFromUrl('https://www.hentai-foundry.com/user/someone')).toBe('hentaifoundry')
expect(getPlatformFromUrl('https://discord.com/channels/@me')).toBe('discord')
expect(getPlatformFromUrl('https://www.pixiv.net/en/users/123')).toBe('pixiv')
})
it('accepts http as well as https, with or without www', () => {
@@ -32,6 +31,15 @@ describe('getPlatformFromUrl', () => {
expect(getPlatformFromUrl('')).toBe(null)
})
it('returns null for pixiv, retired at milestone #406', () => {
// Retired on the operator's platform-focus decision (rule #171). Same guard
// as deviantart's below, and for the same reason: an absence nothing asserts
// is an absence a later edit can quietly undo.
expect(getPlatformFromUrl('https://www.pixiv.net/en/users/12345')).toBe(null)
expect(PLATFORMS.pixiv).toBeUndefined()
expect(PLATFORM_ARTIST_PATTERNS.pixiv).toBeUndefined()
})
it('returns null for deviantart, retired at #3069', () => {
// The 2026-07-05 product decision (FC downloaders = art-dedicated services
// only) left deviantart wired for seven weeks. Asserting the negative is
@@ -80,12 +88,6 @@ describe('isArtistPage', () => {
)
})
it('matches Pixiv numeric user pages, with or without the /en/ prefix', () => {
expect(isArtistPage('https://www.pixiv.net/users/12345', 'pixiv')).toBe(true)
expect(isArtistPage('https://www.pixiv.net/en/users/12345', 'pixiv')).toBe(true)
expect(isArtistPage('https://www.pixiv.net/en/artworks/999', 'pixiv')).toBe(false)
})
it('returns false for a platform with no artist pattern (discord)', () => {
expect(isArtistPage('https://discord.com/channels/@me', 'discord')).toBe(false)
})
@@ -123,8 +125,7 @@ describe('platform table integrity', () => {
const samples = {
patreon: 'https://www.patreon.com/cw/Atole',
subscribestar: 'https://subscribestar.adult/someone',
hentaifoundry: 'https://www.hentai-foundry.com/user/someone',
pixiv: 'https://www.pixiv.net/en/users/12345'
hentaifoundry: 'https://www.hentai-foundry.com/user/someone'
}
for (const [key, url] of Object.entries(samples)) {
expect(isArtistPage(url, key), `${key} artist pattern`).toBe(true)
@@ -179,8 +180,9 @@ describe('manifest.json agrees with the platform table', () => {
for (const h of manifest.host_permissions) {
if (h === '*://*/*') continue
const host = hostOf(h)
// pixiv's OAuth/API hosts are pixiv infrastructure, not creator pages,
// so they are matched by suffix rather than by the domains list.
// Suffix matching lets a platform's infrastructure subdomains belong to
// it without listing each one. (It was added for pixiv's OAuth hosts,
// which left with pixiv at milestone #406; the rule itself is general.)
const claimed = Object.values(PLATFORMS).some(
(p) => p.domains.includes(host) || p.domains.some((d) => host.endsWith(d))
)
+22 -3
View File
@@ -1,5 +1,24 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 32 32">
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 32 32" role="img" aria-label="FabledCurator">
<title>FabledCurator</title>
<!-- Hand-drawn, NOT traced from the source raster. This renders at 16px in a
browser tab and 22px in the nav (TopNav.vue), and a traced mark carries
hundreds of sub-pixel wobble nodes that read as fuzz at those sizes and
cannot be tidied afterwards.
Two elements only. The full logo's glove, cufflink, fleur-de-lis finials,
outer ring and sparkle rays are all deliberately ABSENT rather than drawn
and lost: measured at 16px, ring + frame + star merged into an
indistinct blob, and the whole logo was unreadable mush. The frame won
over a plain ring because it carries the meaning — a framed work is what
curation looks like — and it is the full logo's own centrepiece.
Colours are theme tokens (frontend/src/theme/fabled-tokens.js): obsidian
plate, accent gold. The plate is kept here (unlike logo.svg) so the tab
icon is self-contained against any browser chrome; on the nav it is
invisible because it matches --fc-chrome-rgb exactly. -->
<rect width="32" height="32" rx="6" fill="#14171A"/>
<path d="M8 9 L8 23 L24 23 L24 9 L20 9 L16 13 L12 9 Z"
fill="#A87338" stroke="#E8E4D8" stroke-width="1" stroke-linejoin="round"/>
<rect x="6.2" y="4.2" width="19.6" height="23.6" rx="1.4"
fill="none" stroke="#A87338" stroke-width="2.4"/>
<path d="M16 9.6 Q16.75 15.25 22.4 16 Q16.75 16.75 16 22.4 Q15.25 16.75 9.6 16 Q15.25 15.25 16 9.6 Z"
fill="#A87338"/>
</svg>

Before

Width:  |  Height:  |  Size: 263 B

After

Width:  |  Height:  |  Size: 1.4 KiB

+165
View File
@@ -0,0 +1,165 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 1254 1254" role="img" aria-label="FabledCurator">
<title>FabledCurator</title>
<!-- Traced from the source raster, then re-painted from the design tokens:
gold -> accent.curator #A87338, glove -> text.parchment #E8E4D8. The
original plate was a warm brown (#1B1105), NOT the app obsidian, so it
is dropped entirely — the mark sits on whatever surface hosts it and the
frame interior shows that surface through. Do not reintroduce a
background rect. Too detailed below ~48px: use favicon.svg there. -->
<g transform="translate(0,1254) scale(0.1,-0.1)" fill="#A87338" stroke="none">
<path d="M5900 12310 c-1934 -145 -3621 -1109 -4645 -2655 -1080 -1631 -1218
-3942 -340 -5694 978 -1952 3013 -3175 5285 -3176 389 0 474 5 760 41 2363
299 4374 1950 4995 4099 307 1062 261 2393 -120 3503 -693 2017 -2533 3520
-4690 3831 -347 50 -934 74 -1245 51z m745 -220 c2580 -167 4716 -2034 5170
-4520 289 -1587 -50 -3083 -969 -4280 -1592 -2071 -4404 -2832 -6832 -1847
-379 153 -528 230 -532 272 -7 83 -97 361 -202 625 -317 800 -808 1466 -1325
1796 -175 112 -163 111 -325 31 -151 -75 -153 -80 -55 -136 594 -337 1122
-1023 1484 -1923 101 -252 103 -250 -106 -108 -1075 724 -1897 1902 -2207
3160 -281 1138 -174 2503 279 3565 505 1183 1429 2155 2611 2748 734 368 1453
560 2324 620 114 8 541 6 685 -3z M6222 11867 c-54 -109 -138 -190 -252 -244
-44 -21 -80 -41 -80 -44 0 -4 37 -24 83 -46 129 -63 217 -155 265 -279 25 -63
38 -62 66 7 57 141 180 255 329 303 30 10 19 20 -56 54 -134 60 -227 154 -275
278 -27 68 -33 65 -80 -29z M5470 11590 c-2341 -327 -4149 -2143 -4370 -4389
-11 -117 -8 -125 46 -136 50 -9 60 4 68 93 120 1359 922 2698 2099 3503 671
459 1567 790 2288 846 88 7 102 16 97 65 -4 42 -34 45 -228 18z M6843 11599
c-36 -36 -8 -89 47 -89 105 0 456 -65 709 -130 776 -202 1571 -624 2098 -1112
62 -57 63 -58 122 -58 102 0 97 10 -74 158 -726 626 -1646 1053 -2605 1206
-219 35 -282 41 -297 25z M7855 10996 c-63 -64 -96 -114 -109 -170 -7 -27 -9
-28 -50 -23 -61 9 -109 -8 -145 -49 -48 -54 -42 -67 43 -108 185 -89 270 -345
181 -544 -28 -63 -28 -63 2 -99 39 -48 43 -49 86 -10 34 32 36 36 27 68 -61
210 106 482 311 506 27 3 54 12 60 19 34 42 -49 139 -127 149 -50 7 -50 7 -62
76 -17 99 -100 249 -137 249 -9 0 -44 -29 -80 -64z M7369 10610 c-60 -10 -136
-49 -234 -120 -215 -155 -399 -202 -543 -140 -52 23 -112 78 -112 105 0 7 -7
21 -16 29 -42 43 -114 -13 -114 -88 0 -97 102 -161 311 -197 198 -34 306 -24
490 47 46 18 46 18 92 -30 44 -46 88 -60 78 -25 -49 169 188 238 269 78 35
-68 25 -84 -32 -49 -45 27 -121 27 -148 0 -25 -25 -17 -46 24 -60 18 -6 48
-26 66 -45 36 -37 107 -65 166 -65 62 0 111 131 94 252 -29 203 -204 341 -391
308z M5710 10589 c-215 -36 -300 -314 -163 -533 17 -27 17 -29 -1 -40 -116
-68 -166 -195 -120 -304 57 -138 229 -168 289 -52 38 74 -4 170 -74 170 -12 0
-21 7 -21 15 0 20 52 20 90 0 131 -68 45 -361 -118 -405 -74 -19 -87 -61 -46
-142 67 -132 71 -285 20 -683 -20 -154 -43 -332 -52 -395 -8 -63 -17 -127 -20
-143 -5 -23 -1 -30 26 -45 18 -9 54 -34 80 -54 36 -27 51 -34 57 -25 10 13 11
22 68 457 35 267 88 663 140 1039 42 307 23 421 -85 513 -40 35 -51 58 -27 58
28 0 107 -61 129 -99 41 -72 49 -56 64 126 6 68 4 84 -14 120 -57 110 -16 243
74 243 34 0 30 -23 -6 -37 -76 -29 -82 -173 -8 -173 34 0 34 13 3 -225 -14
-104 -36 -278 -50 -385 -14 -107 -48 -357 -76 -555 -27 -198 -57 -412 -65
-475 -8 -63 -31 -232 -51 -375 -20 -143 -36 -282 -36 -310 -1 -53 58 -214 78
-215 6 0 18 58 28 133 10 72 38 281 62 462 25 182 61 445 80 585 19 140 55
408 80 595 25 187 55 408 66 490 11 83 24 154 29 159 5 5 109 -5 242 -24 556
-77 880 -123 1128 -161 146 -22 337 -51 425 -64 88 -13 322 -49 520 -79 198
-30 513 -78 700 -106 187 -28 350 -53 363 -56 31 -6 31 -3 -33 -565 -48 -428
-137 -1243 -180 -1659 -16 -145 -64 -594 -110 -1020 -14 -124 -36 -337 -50
-475 -14 -137 -28 -256 -30 -264 -6 -17 -45 -14 -298 24 -98 15 -184 24 -192
21 -22 -8 -61 -94 -49 -106 16 -16 632 -106 644 -94 9 9 41 267 75 614 17 170
33 325 100 955 24 231 58 553 75 715 31 294 142 1311 180 1636 34 298 33 317
-20 326 -58 10 -1143 178 -1443 224 -277 41 -274 41 -325 4 -53 -38 -53 -38
-105 17 -51 54 -44 52 -692 143 -510 72 -621 87 -789 107 -101 12 -150 22
-153 31 -9 23 4 25 118 13 62 -7 115 -8 119 -4 5 4 -34 58 -85 119 -180 214
-321 273 -565 233z M8170 10491 c-170 -56 -291 -270 -240 -425 24 -72 147 -96
240 -47 25 14 52 28 60 31 28 13 16 44 -25 68 -39 22 -52 23 -127 6 -59 -13
-10 69 63 106 130 67 273 -61 202 -181 -29 -49 22 -51 92 -3 47 31 43 32 122
-28 194 -147 622 -235 773 -158 95 49 107 170 20 215 -35 18 -45 14 -89 -37
-123 -141 -436 -34 -643 218 -163 200 -315 279 -448 235z M9741 10040 c-90
-24 -163 -63 -286 -155 -130 -97 -132 -93 62 -124 67 -11 141 -25 163 -31 51
-14 74 -2 91 45 17 46 39 46 39 -1 0 -39 -12 -67 -42 -101 -33 -36 -77 -313
-50 -313 5 0 19 11 32 24 39 42 170 75 170 43 0 -7 -19 -18 -43 -24 -126 -34
-187 -164 -216 -458 -6 -60 -20 -191 -31 -290 -44 -392 -101 -929 -140 -1310
-22 -220 -49 -479 -60 -575 -11 -96 -29 -265 -40 -375 -11 -110 -27 -260 -36
-333 -21 -174 0 -313 57 -386 23 -29 24 -46 4 -46 -9 0 -30 20 -47 45 -53 75
-64 52 -53 -118 11 -185 11 -221 -4 -254 -17 -42 -45 -45 -36 -4 15 67 -48 99
-235 120 -47 5 -104 13 -126 16 -53 9 -61 -4 -24 -40 32 -30 39 -55 17 -55 -8
0 -22 10 -33 21 -52 59 -288 135 -352 115 -21 -7 -147 -177 -147 -199 0 -15
13 -14 90 7 148 40 315 -9 315 -94 0 -33 26 -44 70 -31 44 13 80 52 80 87 0
40 24 27 80 -42 210 -258 551 -349 702 -188 95 101 116 316 45 455 -37 74 -44
99 -28 99 4 0 29 18 53 40 141 124 92 363 -84 411 -159 42 -259 -179 -118
-259 16 -9 30 -23 30 -29 0 -24 -49 -13 -83 17 -118 103 -48 355 120 431 89
40 102 73 59 150 -91 164 -95 351 -22 1044 26 248 58 564 72 704 36 360 77
503 176 616 54 62 49 93 -29 160 -89 76 -130 251 -80 342 21 40 89 77 123 69
33 -9 31 -36 -2 -36 -133 0 -118 -214 18 -251 187 -50 305 240 154 380 -33 31
-33 36 10 84 129 143 137 385 17 509 -89 91 -252 127 -402 88z M6187 10008
c-8 -14 -34 -172 -61 -383 -26 -198 -70 -522 -97 -720 -27 -198 -65 -477 -84
-620 -19 -143 -44 -325 -55 -405 -24 -176 -45 -353 -60 -520 -16 -173 -57
-544 -80 -728 -44 -338 -51 -425 -38 -438 7 -7 85 -24 173 -38 88 -14 372 -62
630 -105 536 -91 1299 -211 1336 -211 14 0 63 16 110 35 288 117 571 67 630
-112 13 -40 2 -37 248 -73 236 -35 226 -37 235 63 4 39 29 272 56 517 27 245
61 553 75 685 14 132 35 321 45 420 11 99 40 369 64 600 25 231 66 596 91 810
73 615 84 738 71 752 -6 6 -225 43 -486 82 -261 39 -608 92 -770 116 -952 146
-1345 203 -1770 260 -102 14 -199 27 -217 30 -26 5 -35 1 -46 -17z m598 -459
c116 -16 521 -75 900 -130 380 -54 850 -122 1045 -149 207 -30 360 -56 368
-63 10 -10 3 -92 -33 -412 -25 -220 -61 -537 -80 -705 -19 -168 -53 -465 -75
-660 -22 -195 -52 -456 -66 -580 -57 -525 -86 -746 -97 -757 -10 -10 -31 -10
-102 2 -90 15 -410 64 -1795 277 -412 63 -731 116 -739 123 -11 11 -7 58 23
296 20 156 67 529 106 829 38 300 90 700 115 890 25 190 66 505 91 700 48 377
50 383 98 375 17 -3 126 -19 241 -36z M7696 8744 c-121 -421 -482 -746 -868
-782 -60 -5 -81 -28 -38 -39 40 -10 160 -72 217 -112 265 -181 435 -496 468
-863 10 -102 24 -100 59 10 136 415 458 696 859 747 64 8 60 25 -17 58 -350
153 -580 495 -627 935 -14 130 -26 140 -53 46z M10390 9624 c-16 -42 -12 -105
8 -128 64 -71 251 -371 366 -584 279 -520 462 -1104 531 -1696 16 -142 23
-156 69 -156 63 0 66 16 37 242 -109 836 -404 1569 -903 2241 -86 115 -93 121
-108 81z M5657 6918 c-37 -73 -67 -146 -67 -162 0 -15 -20 -179 -45 -364 -37
-271 -43 -338 -33 -348 7 -7 153 -35 323 -64 171 -28 443 -75 605 -105 434
-78 1022 -175 1166 -191 60 -6 60 -6 110 36 27 24 50 49 52 56 4 16 8 16 -283
59 -680 102 -1828 300 -1839 317 -8 14 2 111 79 731 26 211 22 213 -68 35z
M1124 6843 c-33 -140 -181 -299 -326 -351 -63 -22 -61 -35 10 -61 126 -46 245
-167 298 -301 33 -84 45 -87 67 -16 40 125 154 250 283 308 89 40 91 49 17 77
-146 53 -255 176 -318 359 -8 21 -24 13 -31 -15z M11345 6824 c-8 -20 -15 -40
-15 -44 0 -21 -105 -162 -146 -195 -46 -38 -168 -105 -191 -105 -36 0 -3 -30
75 -67 132 -63 209 -146 268 -291 29 -71 39 -71 68 0 65 159 179 273 319 318
50 16 47 32 -12 53 -125 44 -246 162 -300 292 -35 82 -45 89 -66 39z M5443
6629 c-40 -46 -53 -68 -53 -93 0 -91 -103 -184 -173 -158 -47 18 -113 -55
-109 -120 2 -19 -4 -31 -19 -38 -30 -16 -64 -97 -94 -223 -30 -125 -23 -179
36 -250 105 -128 397 -130 570 -5 44 32 69 31 69 -3 0 -27 73 -101 91 -94 8 3
23 20 34 38 46 73 165 102 297 73 140 -31 251 -97 418 -251 162 -147 309 -196
461 -152 152 45 250 183 206 289 -34 82 -227 104 -274 31 -21 -32 25 -83 77
-83 40 0 39 -25 -1 -39 -139 -48 -317 64 -229 144 55 50 -44 81 -105 33 -23
-18 -30 -19 -48 -9 -129 81 -219 110 -511 162 -172 31 -207 32 -315 2 -22 -5
-23 -4 -16 26 3 17 5 31 3 31 -2 0 -46 9 -98 20 -142 30 -175 27 -217 -15 -19
-19 -39 -35 -44 -35 -19 0 -8 60 21 115 17 31 30 69 30 84 0 14 4 42 9 61 27
105 19 150 -20 104 -14 -17 -82 -54 -99 -54 -20 0 -10 27 16 41 70 37 123 151
148 319 20 134 15 138 -61 49z M1115 5844 c-19 -19 -18 -36 22 -269 67 -398
241 -929 406 -1239 49 -93 63 -104 105 -87 46 19 46 34 -3 131 -114 228 -240
575 -314 865 -46 182 -111 501 -111 543 0 62 -65 96 -105 56z M11312 5848 c-7
-7 -12 -26 -12 -43 0 -51 -48 -312 -86 -463 -293 -1183 -1186 -2343 -2304
-2992 -1397 -811 -3104 -938 -4572 -339 -97 40 -97 40 -148 14 -87 -44 -66
-59 249 -176 1105 -412 2330 -434 3476 -62 1535 498 2757 1657 3282 3113 115
322 240 900 203 945 -15 18 -71 20 -88 3z M7260 5633 c-76 -68 -38 -250 66
-313 18 -12 42 -19 51 -16 14 5 68 86 91 138 2 4 -15 18 -37 33 -59 37 -55 55
12 48 49 -5 57 -3 87 24 45 40 36 55 -53 86 -103 37 -176 37 -217 0z M8484
5146 c-81 -19 -146 -53 -223 -119 -68 -58 -115 -143 -461 -831 -106 -210 -141
-249 -325 -361 -267 -161 -503 -324 -564 -389 -53 -56 54 -28 549 145 554 194
522 173 700 469 198 331 438 660 628 862 72 78 72 78 -3 147 -86 79 -183 104
-301 77z M1620 3910 c0 -50 807 -1065 1163 -1465 56 -62 47 -20 -24 114 -246
463 -531 859 -786 1088 -180 163 -353 291 -353 263z M3579 2997 c-316 -89
-350 -543 -51 -680 162 -75 385 23 452 198 102 269 -137 557 -401 482z m168
-72 c194 -58 258 -317 117 -474 -174 -193 -475 -70 -474 194 1 200 173 335
357 280z M3596 2764 c-7 -61 -22 -96 -61 -143 -33 -39 -26 -51 31 -51 54 0 77
-10 125 -57 41 -39 44 -38 54 32 8 58 27 99 62 137 32 33 23 48 -28 48 -52 0
-97 20 -133 59 -40 41 -42 40 -50 -25z"/>
</g>
<g transform="translate(0,1254) scale(0.1,-0.1)" fill="#E8E4D8" stroke="none">
<path d="M5355 7946 c-40 -18 -64 -47 -110 -131 -159 -293 -388 -522 -783
-786 -172 -115 -170 -113 -244 -233 -34 -56 -172 -252 -308 -436 -398 -541
-402 -548 -645 -950 -208 -345 -310 -504 -368 -572 -80 -95 -241 -195 -372
-232 -138 -39 -138 -39 -177 -91 -35 -46 -47 -54 -155 -98 -229 -94 -393 -169
-393 -181 0 -7 16 -20 36 -30 240 -114 591 -436 850 -781 309 -409 618 -1048
770 -1588 14 -51 31 -95 37 -99 6 -3 49 26 96 66 170 145 370 281 630 431 137
79 156 94 159 118 4 32 -31 141 -99 310 -27 67 -49 137 -49 157 0 52 25 142
43 152 32 18 45 38 52 84 8 53 37 86 135 154 393 269 1017 301 1726 90 280
-84 322 -75 522 114 186 175 346 293 697 513 188 118 167 88 470 687 290 576
315 619 473 838 59 81 117 162 130 179 69 98 -43 209 -213 209 -172 0 -415
-121 -554 -277 -89 -98 -241 -341 -241 -384 0 -6 -38 -51 -84 -100 -47 -49
-113 -127 -147 -174 -50 -66 -75 -91 -119 -114 -72 -40 -300 -201 -314 -223
-9 -14 -4 -25 26 -55 44 -46 88 -111 88 -129 0 -26 -25 -14 -100 46 -97 78
-205 134 -317 164 -104 27 -117 34 -277 133 -112 70 -521 308 -851 496 -402
230 -627 313 -945 350 -97 11 -105 14 -108 35 -4 24 16 29 48 12 13 -7 33 -6
64 2 60 16 212 -3 344 -42 113 -34 126 -29 77 30 -34 40 -36 48 -31 89 3 26
10 96 15 156 13 172 62 341 126 435 28 41 63 96 77 121 17 29 70 82 142 143
355 299 509 573 510 906 0 301 -179 559 -339 486z m-1542 -4944 c201 -96 278
-319 183 -527 -125 -273 -529 -286 -663 -21 -169 334 154 703 480 548z M7135
5383 c-50 -44 -145 -90 -203 -99 -35 -5 -64 -22 -126 -70 -111 -88 -300 -177
-561 -265 -64 -21 -110 -42 -110 -50 0 -11 130 -93 248 -155 40 -21 153 -55
264 -81 17 -4 53 14 130 65 60 38 138 86 174 106 54 29 82 55 147 134 44 53
114 134 156 179 90 96 88 84 28 138 -27 24 -56 62 -66 84 -25 53 -34 55 -81
14z"/>
</g>
</svg>

After

Width:  |  Height:  |  Size: 12 KiB

+60 -7
View File
@@ -5,9 +5,13 @@
<img src="/favicon.svg" alt="" class="fc-brand__glyph" width="22" height="22" />
<span class="fc-brand__text">FabledCurator</span>
</RouterLink>
<span class="fc-health" :title="health.label">
<RouterLink
:to="{ name: 'settings', query: { tab: 'system' } }"
class="fc-health" :title="health.label"
:aria-label="`System health: ${health.label}`"
>
<v-icon size="x-small" :color="health.color">{{ health.icon }}</v-icon>
</span>
</RouterLink>
<PipelineStatusChip />
</div>
@@ -64,13 +68,15 @@
</template>
<script setup>
import { computed, onBeforeUnmount, onMounted, ref } from 'vue'
import { computed, onBeforeUnmount, onMounted, onUnmounted, ref } from 'vue'
import { useRoute } from 'vue-router'
import router, { FRONT_DOOR } from '../router.js'
import { useSystemStore } from '../stores/system.js'
import { useSystemHealthStore } from '../stores/systemHealth.js'
import PipelineStatusChip from './PipelineStatusChip.vue'
const system = useSystemStore()
const healthStore = useSystemHealthStore()
// Publish the nav's REAL height as --fc-nav-h so full-height workspaces
// (Explore/Subscriptions) and sticky sub-headers pin to it exactly instead of a
@@ -116,15 +122,55 @@ const settingsRoute = computed(() =>
navRoutes.value.find(r => r.name === 'settings') || null
)
// The dot beside the brand, and the only ambient signal that something in the
// stack has stopped (milestone 365).
//
// It used to read /api/health — a no-DB liveness check that proves the WEB
// container is serving and nothing else. Green there while the worker was dead
// is exactly what it looked like, and a green dot next to the product name is
// read as "everything is fine". It now reflects the whole-stack verdict.
//
// Deliberately re-using this element rather than adding a second indicator:
// there were already three partial surfaces (this, the pipeline chip, the
// Settings Activity tab) and a fourth would have made the question harder to
// answer, not easier. This is the one that already occupied the slot.
const health = computed(() => {
if (system.healthy === null) {
const overall = healthStore.overall
if (overall === null) {
return { icon: 'mdi-circle-outline', color: 'on-surface', label: 'checking…' }
}
if (system.healthy === true) {
return { icon: 'mdi-circle', color: 'success', label: 'healthy' }
if (overall === 'ok') {
return { icon: 'mdi-circle', color: 'success', label: 'All parts running' }
}
return { icon: 'mdi-alert-circle', color: 'error', label: 'unreachable' }
// Name what is wrong in the tooltip. "Something is unhealthy" sends someone
// hunting; "Scheduler has not checked in for 6 min" does not.
const worst = healthStore.problems[0]
const others = healthStore.problems.length - 1
const suffix = others > 0 ? ` (+${others} more)` : ''
if (overall === 'down') {
return {
icon: 'mdi-alert-circle', color: 'error',
label: (worst?.detail || 'A part has stopped') + suffix,
}
}
if (overall === 'stale') {
return {
icon: 'mdi-alert', color: 'warning',
label: (worst?.detail || 'A part is quiet') + suffix,
}
}
return { icon: 'mdi-help-circle-outline', color: 'on-surface', label: 'Health unknown' }
})
const HEALTH_POLL_MS = 15_000
let healthTimer = null
onMounted(() => {
healthStore.refresh()
healthTimer = setInterval(() => {
if (!document.hidden) healthStore.refresh()
}, HEALTH_POLL_MS)
})
onUnmounted(() => { if (healthTimer) clearInterval(healthTimer) })
</script>
<style scoped>
@@ -237,7 +283,14 @@ const health = computed(() => {
display: flex;
align-items: center;
flex-shrink: 0;
/* A RouterLink since milestone 365 — it is the path to the Settings System
tab, not just an indicator. Reset the anchor so turning a span into a
link changed nothing about how the nav reads. */
text-decoration: none;
color: inherit;
border-radius: 50%;
}
.fc-health:hover { background: rgb(var(--v-theme-on-surface) / 0.12); }
.fc-nav-right {
flex: 1 1 0;
min-width: 0;
@@ -0,0 +1,296 @@
<template>
<div class="fc-empty">
<!-- The brand mark's second home (#387 A-side logo work, operator-confirmed).
It earns its place HERE and not on a populated feed: this is the first
screen a fresh install shows anyone, and it gives the on-ramp something
to be composed around rather than a bare line of prose plus buttons.
Large on purpose — logo.svg stops reading below ~48px, which is why the
22px nav slot has its own mark. No wrapping card: the file deliberately
carries no plate and is meant to sit on the page surface. Vendored
locally, so this renders on an install with no network — which is
exactly the state this screen appears in. -->
<img src="/logo.svg" alt="" class="fc-empty__mark" width="150" height="150" />
<!-- FIRST RUN: sources exist, nothing has landed yet. Telling someone to
add a source when they already have three and are mid-backfill is
worse than saying nothing at all. -->
<template v-if="hasSources">
<p class="fc-empty__lead">Nothing has arrived yet.</p>
<p class="fc-empty__sub">
{{ sourceCountLabel }} configured. The first check can take a while —
deep history is fetched in chunks.
</p>
<div class="fc-empty__actions">
<v-btn
size="small" variant="tonal" color="accent"
prepend-icon="mdi-progress-download"
:to="{ path: '/subscriptions', query: { tab: 'downloads' } }"
>See what's running</v-btn>
</div>
</template>
<!-- FRESH INSTALL: the on-ramp, in the order the steps actually depend on
each other a source cannot fetch anything without a credential.
#387 C6 adds the third rung: once a credential exists, FC can just
LOOK UP what the operator already subscribes to instead of making
them retype it. Exactly one rung shows, chosen by `rung` below. -->
<template v-else>
<p class="fc-empty__lead">Nothing here yet.</p>
<!-- 1. No credential. Discovery is not yet possible, so it is not
offered a button that cannot work is worse than its absence. -->
<template v-if="rung === 'credential'">
<p class="fc-empty__sub">
FabledCurator follows the creators you subscribe to and files what
they post. Two steps to start.
</p>
<div class="fc-empty__actions">
<v-btn
size="small" variant="tonal" color="accent" prepend-icon="mdi-key-variant"
:to="{ path: '/subscriptions', query: { tab: 'settings' } }"
>Add a credential</v-btn>
<v-btn
size="small" variant="text" prepend-icon="mdi-plus"
:to="{ path: '/subscriptions' }"
>Add a source</v-btn>
</div>
<p class="fc-empty__hint">
A credential comes first a source can't fetch anything without your
logged-in session.
</p>
</template>
<!-- 2. Credential, roster never synced. The payoff rung. -->
<template v-else-if="rung === 'discover'">
<p class="fc-empty__sub">
You've connected {{ credentialLabel }}. FabledCurator can look up the
creators you already subscribe to, so you don't have to add them by
hand.
</p>
<div class="fc-empty__actions">
<v-btn
size="small" variant="tonal" color="accent"
prepend-icon="mdi-account-search-outline"
:loading="discovering"
@click="discover"
>Find what you already subscribe to</v-btn>
<v-btn
size="small" variant="text" prepend-icon="mdi-plus"
:to="{ path: '/subscriptions' }"
>Add a source</v-btn>
</div>
<p class="fc-empty__hint">
Nothing is added automatically — you'll get a list to pick from.
</p>
</template>
<!-- 3. Synced, and there is something to adopt. The count is the whole
message; picking happens in C4's card. -->
<template v-else-if="rung === 'adopt'">
<p class="fc-empty__sub">
{{ unmatchedCount }}
{{ unmatchedCount === 1 ? 'creator you subscribe to is' : 'creators you subscribe to are' }}
not being followed here yet.
</p>
<div class="fc-empty__actions">
<v-btn
size="small" variant="tonal" color="accent"
prepend-icon="mdi-account-plus-outline"
:to="{ path: '/subscriptions' }"
>Choose who to follow</v-btn>
</div>
<p class="fc-empty__hint">
Added one at a time, so a fresh install doesn't start a dozen
backfills at once.
</p>
</template>
<!-- 4. Discovery was tried and could not reach the platform. Rule #164:
say so plainly. Not a spinner, not a crash, not a retry button
that will fail the same way. The manual path still works. -->
<template v-else-if="rung === 'unavailable'">
<p class="fc-empty__sub">
FabledCurator couldn't reach {{ credentialLabel }} to look up your
subscriptions{{ discoveryError ? ` (${discoveryError})` : '' }}. You
can still add creators by hand.
</p>
<div class="fc-empty__actions">
<v-btn
size="small" variant="tonal" color="accent" prepend-icon="mdi-plus"
:to="{ path: '/subscriptions' }"
>Add a source</v-btn>
<v-btn
size="small" variant="text" prepend-icon="mdi-key-variant"
:to="{ path: '/subscriptions', query: { tab: 'settings' } }"
>Check the credential</v-btn>
</div>
</template>
<!-- 5. Synced and nothing to discover: the roster is fully tracked, or
the account subscribes to nothing. Either way the manual path is
genuinely the only thing left to offer. -->
<template v-else>
<p class="fc-empty__sub">
Everything you subscribe to on {{ credentialLabel }} is already
followed here. Add a creator by hand to start filling the feed.
</p>
<div class="fc-empty__actions">
<v-btn
size="small" variant="tonal" color="accent" prepend-icon="mdi-plus"
:to="{ path: '/subscriptions' }"
>Add a source</v-btn>
</div>
</template>
</template>
</div>
</template>
<script setup>
import { computed, onMounted, ref, watch } from 'vue'
import { storeToRefs } from 'pinia'
import { useCredentialsStore } from '../../stores/credentials.js'
import { useMembershipReconcileStore } from '../../stores/membershipReconcile.js'
import { useMembershipSyncStore } from '../../stores/membershipSync.js'
import { useSourcesStore } from '../../stores/sources.js'
const store = useSourcesStore()
const { scheduleStatus: status } = storeToRefs(store)
const credentials = useCredentialsStore()
const syncStore = useMembershipSyncStore()
const reconcile = useMembershipReconcileStore()
const { byPlatform } = storeToRefs(credentials)
const { platforms: syncPlatforms } = storeToRefs(syncStore)
const { platforms: reconcilePlatforms } = storeToRefs(reconcile)
// Absent status is treated as "fresh install", which is the safe way round:
// the on-ramp is useful to a first-run operator and merely redundant to an
// established one, whereas "see what's running" shown to someone with nothing
// configured is a dead end.
const hasSources = computed(() => (status.value?.total_sources || 0) > 0)
const sourceCountLabel = computed(() => {
const n = status.value?.total_sources || 0
return `${n} ${n === 1 ? 'source' : 'sources'}`
})
// --- #387 C6: which on-ramp rung this install is actually on ---------------
//
// ONE predicate, not four independent v-ifs. The rungs are mutually exclusive
// by construction — an install is at exactly one point in the loop — and four
// separate conditions would eventually let two of them render at once, which
// on the first screen a new installer ever sees is the worst place for it.
const hasCredential = computed(() => byPlatform.value.size > 0)
// Any platform that has ever completed a sweep. NULL last_success_at means
// NEVER, which the C3 endpoint is deliberately careful to preserve — rendering
// it as 0 is the conflation that whole surface exists to prevent.
const syncedPlatform = computed(
() => syncPlatforms.value.find((p) => p.last_success_at) || null,
)
// Errored WITHOUT ever having succeeded. A platform that synced once and
// failed since is not "unavailable" — it has a roster, just an ageing one, and
// C4's freshness gate is what handles that. This is specifically the
// never-worked case, which is what an install with no outbound network looks
// like (rule #164).
const discoveryError = computed(() => {
if (syncedPlatform.value) return null
const failed = syncPlatforms.value.find((p) => p.last_error_type)
return failed ? failed.last_error_type : null
})
const unmatchedCount = computed(() =>
reconcilePlatforms.value.reduce(
(n, p) => n + (p.subscribed_not_tracked?.length || 0), 0,
),
)
// Which platforms the operator has connected, for the copy. Named rather than
// counted: "you've connected patreon" is a sentence about their setup, where
// "1 credential" is a sentence about our data model.
const credentialLabel = computed(() => {
const names = [...byPlatform.value.keys()]
if (names.length === 0) return 'your platform'
if (names.length === 1) return names[0]
return `${names.slice(0, -1).join(', ')} and ${names[names.length - 1]}`
})
const rung = computed(() => {
// Absent data reads as "no credential", the same safe-direction default B4
// chose for hasSources: the first rung is useful to a genuine fresh install
// and merely redundant to anyone further along, whereas a discovery button
// shown to someone with no credential is a dead end.
if (!hasCredential.value) return 'credential'
if (discoveryError.value) return 'unavailable'
if (!syncedPlatform.value) return 'discover'
return unmatchedCount.value > 0 ? 'adopt' : 'manual'
})
const discovering = ref(false)
async function discover () {
// The sweep is queued, not inline (C3) — so this cannot report a result, and
// must not pretend to. syncNow() raises a toast saying it runs in the
// background; re-reading the state is what eventually moves the rung.
discovering.value = true
try {
await syncStore.syncNow()
await syncStore.load().catch(() => {})
} finally {
discovering.value = false
}
}
// Bucket 1's count is only meaningful once something has synced, so it is
// fetched then rather than on every empty render. `immediate` covers the case
// where the sync state arrived before this watcher was set up.
watch(syncedPlatform, (p) => {
if (p && reconcilePlatforms.value.length === 0) reconcile.load().catch(() => {})
}, { immediate: true })
onMounted(() => {
// The ribbon may already have loaded this on the front door; in Browse's
// Posts tab nothing has. Failure is swallowed — an empty feed must still
// explain itself when the status call is unavailable (rule #164). The same
// goes for every call below: this screen renders on an install that can
// reach nothing, and a rejected promise here would leave it blank.
if (!status.value) store.loadScheduleStatus().catch(() => {})
if (byPlatform.value.size === 0) credentials.loadAll().catch(() => {})
if (syncPlatforms.value.length === 0) syncStore.load().catch(() => {})
})
</script>
<style scoped>
.fc-empty {
display: flex;
flex-direction: column;
align-items: center;
text-align: center;
gap: 6px;
padding: 48px 16px 32px;
}
.fc-empty__mark {
/* Sits back a little: it frames the message rather than competing with it. */
opacity: 0.9;
margin-bottom: 8px;
}
.fc-empty__lead {
font-size: 1.05rem;
color: rgb(var(--v-theme-on-surface));
}
.fc-empty__sub, .fc-empty__hint {
color: rgb(var(--v-theme-on-surface-variant));
max-width: 34rem;
}
.fc-empty__hint { font-size: 0.8rem; margin-top: 4px; }
.fc-empty__actions {
display: flex;
flex-wrap: wrap;
justify-content: center;
gap: 8px;
margin-top: 10px;
}
</style>
@@ -0,0 +1,80 @@
<template>
<!-- Renders nothing at all until it has something true to say. A ribbon that
shows a skeleton or an error on the front door would make the app look
broken every cold load; the feed is the point, this is an aside. -->
<div v-if="status" class="fc-ribbon">
<span class="fc-ribbon__item">
<v-icon size="x-small">mdi-clock-outline</v-icon>
Checked {{ lastCheckedLabel }}
</span>
<RouterLink
v-if="failing" class="fc-ribbon__item fc-ribbon__item--err"
:to="{ path: '/subscriptions', query: { status: 'errors' } }"
>
<v-icon size="x-small">mdi-alert-circle-outline</v-icon>
{{ failing }} {{ failing === 1 ? 'source is' : 'sources are' }} failing
</RouterLink>
<!-- The reason this ribbon exists. Phase A made "we can't see this
creator's posts" a durable fact; without a line here it stays buried
three clicks into Subscriptions, which is exactly where nobody looks
until they already suspect something. -->
<RouterLink
v-if="noAccess" class="fc-ribbon__item fc-ribbon__item--gated"
:to="{ path: '/subscriptions', query: { status: 'no_access' } }"
>
<v-icon size="x-small">mdi-lock-outline</v-icon>
{{ noAccess }} you can't see
</RouterLink>
</div>
</template>
<script setup>
import { computed, onMounted } from 'vue'
import { storeToRefs } from 'pinia'
import { useSourcesStore } from '../../stores/sources.js'
import { formatRelative } from '../../utils/date.js'
const store = useSourcesStore()
const { scheduleStatus: status } = storeToRefs(store)
const failing = computed(() => status.value?.failing_sources || 0)
const noAccess = computed(() => status.value?.no_access_sources || 0)
const lastCheckedLabel = computed(() =>
formatRelative(status.value?.last_tick_at, { nullText: 'never' }),
)
onMounted(() => {
// Swallowed on purpose: this is an aside on the front door, and the feed
// must render whether or not the status call succeeds (rule #164 — a
// genuinely-external-ish fact degrades to absent, it never gates the page).
store.loadScheduleStatus().catch(() => {})
})
</script>
<style scoped>
.fc-ribbon {
display: flex;
align-items: center;
flex-wrap: wrap;
gap: 4px 14px;
padding: 2px 0 10px;
font-size: 0.78rem;
color: rgb(var(--v-theme-on-surface-variant));
}
.fc-ribbon__item {
display: inline-flex;
align-items: center;
gap: 4px;
color: inherit;
text-decoration: none;
}
/* Only the actionable items take a colour, so a healthy instance reads as one
quiet grey line rather than a status dashboard. */
.fc-ribbon__item--err { color: rgb(var(--v-theme-error)); }
.fc-ribbon__item--gated { color: rgb(var(--v-theme-info)); }
.fc-ribbon__item--err:hover,
.fc-ribbon__item--gated:hover { text-decoration: underline; }
</style>
+95 -4
View File
@@ -6,6 +6,19 @@
<v-chip size="x-small" variant="tonal">
{{ post.source?.platform ?? 'filesystem import' }}
</v-chip>
<!-- The honesty marker (#388 E2). Discord doesn't publish posts, so FC
groups a creator's drop and writes the post itself. This chip is the
only thing standing between "FC assembled this" and the card reading
as something the artist authored it must never be conditional on
anything but the flag, and it says what it was built FROM so the
claim is checkable rather than just disclosed. -->
<v-chip
v-if="synthesized" size="x-small" variant="outlined"
class="fc-post-card__synthetic" :title="synthesisTitle"
>
<v-icon icon="mdi-auto-fix" size="x-small" start />
grouped by FabledCurator
</v-chip>
<RouterLink
:to="{ name: 'artist', params: { slug: post.artist.slug } }"
class="fc-post-card__artist"
@@ -14,6 +27,12 @@
<span v-if="totalImages" class="fc-post-card__meta">
· {{ totalImages }} image{{ totalImages === 1 ? '' : 's' }}
</span>
<!-- Only on a grouping that has actually grown. An ordinary post can
never show this, and a grouping that has not grown says nothing
an "updated" label that is always present teaches you to ignore it. -->
<span v-if="grewAt" class="fc-post-card__meta fc-post-card__grew">
· updated {{ grewRelative }}
</span>
<v-spacer />
<PostSeriesMenu :post="post" />
<v-btn
@@ -59,6 +78,13 @@
<div class="fc-post-card__text">
<h3 v-if="displayTitle" class="fc-post-card__title">{{ displayTitle }}</h3>
<!-- A synthetic post has no title on purpose (inventing one is the one
place this feature could put words in a creator's mouth), so name
it for what it is rather than leaking `fc-drop:<message id>`. -->
<h3
v-else-if="synthesized"
class="fc-post-card__title fc-post-card__title--missing"
>{{ synthesisTitle }}</h3>
<h3 v-else class="fc-post-card__title fc-post-card__title--missing">
Post {{ post.external_post_id }}
</h3>
@@ -102,6 +128,25 @@
@click="toggleDesc"
>{{ descExpanded ? 'Show less' : 'Show more' }}</button>
<!-- #388 E5: the accepted announcement link, both directions. Only
ACCEPTED ones reach the payload, so anything rendered here is
something the operator agreed to — never a proposal. -->
<!-- A location with no name/path is "the current route, with these
query params" — so this works from Latest or Browse alike without
the card reaching for useRoute(), which it has no other reason to
know about. -->
<div v-if="associations.length" class="fc-post-card__assoc">
<RouterLink
v-for="a in associations" :key="a.id"
:to="{ query: { post_id: a.post_id } }"
class="fc-post-card__assoc-link"
>
<v-icon icon="mdi-link-variant" size="x-small" />
{{ a.role === 'announces'
? 'The full set is in Discord' : 'Announced on Patreon' }}
</RouterLink>
</div>
<!-- Off-platform file-host links (mega/gdrive/…) found in the body.
Shown once expanded; clickable now, auto-downloaded by the worker
slice. -->
@@ -163,6 +208,18 @@ const images = computed(() => props.post.thumbnails || [])
const totalImages = computed(() => images.value.length + (props.post.thumbnails_more || 0))
const plainTitle = computed(() => toPlainText(props.post.post_title))
// #388 E2. Non-null `synthesized_by` means FC authored this row by grouping a
// creator's drop; `synthesis` carries what it was built from. Read defensively
// — a post fetched before the field existed, or any surface that composes a
// post dict by hand, must degrade to "not synthetic" rather than throw.
const synthesized = computed(() => Boolean(props.post.synthesized_by))
const messageCount = computed(() => props.post.synthesis?.message_count ?? 0)
const synthesisTitle = computed(() => {
const n = messageCount.value
if (!n) return 'Grouped from Discord'
return `Grouped from ${n} Discord message${n === 1 ? '' : 's'}`
})
const hero = computed(() => images.value[0])
// The thumbnail strip spans the hero's full width (CSS grid, equal columns),
// rather than a fixed 3-cell cap. Show up to RAIL_MAX cells; when there are
@@ -189,16 +246,29 @@ const moreCount = computed(() => {
const railCols = computed(() => rail.value.length + (moreCount.value > 0 ? 1 : 0))
const sortDateIso = computed(() => props.post.post_date || props.post.downloaded_at)
// #388 E3. A grouping stays OPEN, so its own date and its latest activity are
// different facts. The card keeps showing when the drop STARTED — that is the
// post's identity — and reports growth separately, because "this post is from
// Tuesday but gained images this morning" is the whole signal that chat
// content is trickling in.
const grewAt = computed(() => (synthesized.value ? props.post.last_grew_at : null))
// #388 E5. Accepted links only — a pending proposal lives in the review queue,
// never beside the artwork. Defaults to [] so a post dict from before the
// feature (or composed by hand) renders without a link rather than throwing.
const associations = computed(() => props.post.associations || [])
const grewRelative = computed(() => (grewAt.value ? relativeFrom(grewAt.value) : ''))
const absoluteDate = computed(() => new Date(sortDateIso.value).toLocaleString())
const relativeDate = computed(() => {
const then = new Date(sortDateIso.value).getTime()
function relativeFrom (iso) {
const then = new Date(iso).getTime()
const diff = (Date.now() - then) / 1000
if (diff < 60) return `${Math.floor(diff)}s ago`
if (diff < 3600) return `${Math.floor(diff / 60)}m ago`
if (diff < 86400) return `${Math.floor(diff / 3600)}h ago`
if (diff < 86400 * 30) return `${Math.floor(diff / 86400)}d ago`
return new Date(sortDateIso.value).toLocaleDateString()
})
return new Date(iso).toLocaleDateString()
}
const relativeDate = computed(() => relativeFrom(sortDateIso.value))
// --- images → post-scoped modal ---------------------------------------
async function fullImageIds () {
@@ -337,6 +407,27 @@ function formatBytes (n) {
font-weight: 600;
}
.fc-post-card__artist:hover { color: rgb(var(--v-theme-accent)); }
/* Quiet, not decorative: the marker has to be legible on every card without
turning a synthetic post into the loudest thing in the feed. */
.fc-post-card__synthetic {
color: rgb(var(--v-theme-on-surface-variant));
}
/* Growth is news, so it gets the accent the rest of the meta line doesn't —
but it is still the meta line, not a badge competing with the artwork. */
.fc-post-card__grew { color: rgb(var(--v-theme-accent)); }
.fc-post-card__assoc { margin-top: 8px; }
.fc-post-card__assoc-link {
display: inline-flex;
align-items: center;
gap: 4px;
font-size: 0.8125rem;
color: rgb(var(--v-theme-accent));
text-decoration: none;
}
.fc-post-card__assoc-link:hover { text-decoration: underline; }
.fc-post-card__date,
.fc-post-card__meta { white-space: nowrap; }
@@ -19,7 +19,7 @@
<v-card-text>
<p class="fc-muted text-body-2">
Pushes session cookies from supported platforms
(patreon, subscribestar, hentaifoundry, discord, pixiv)
(patreon, subscribestar, hentaifoundry, discord)
into FabledCurator, and lets you add a creator as a source from
their page in one click.
</p>
@@ -0,0 +1,134 @@
<template>
<MaintenanceTile
icon="mdi-image-multiple"
title="Discord drop grouping"
blurb="Group a creator's variant drop into one post FC writes itself."
>
<div v-if="store.settings">
<div class="text-caption fc-muted mb-3">
Discord is a delivery channel, not a publisher one message isn't one
post. When this is on, FC groups a creator's variant drop (same piece,
different hair colour or outfit) into a single post it authors, with the
messages' text as the body. Grouped posts are always marked as FC's own,
and deleting one returns its messages to the feed unchanged.
</div>
<v-switch
v-model="local.discord_grouping_enabled" color="accent" hide-details
density="compact" label="Group Discord drops into posts"
@update:model-value="onSave"
/>
<v-row class="mt-2">
<v-col cols="12" sm="6">
<SettingNumberField
v-model="local.discord_group_max_distance"
label="Max visual distance" :min="0" :max="1" :step="0.01"
density="comfortable" max-width="none"
:disabled="!local.discord_grouping_enabled" @change="onSave"
/>
<div class="text-caption fc-muted mt-1">
0 is identical, 1 is unrelated. Lower groups less. Raising this is
the fix when a drop comes out scattered across several posts
but raise it slowly: too high merges pieces that only look alike.
</div>
</v-col>
<v-col cols="12" sm="6">
<SettingNumberField
v-model="local.discord_group_window_minutes"
label="Drop window (minutes)" :min="1" :step="5"
density="comfortable" max-width="none"
:disabled="!local.discord_grouping_enabled" @change="onSave"
/>
<div class="text-caption fc-muted mt-1">
The quiet gap that ends a drop, measured between consecutive
messages so variants trickling out over an evening stay one post.
This is the setting doing most of the work: without it, everything
an artist ever drew of one character would collapse into one post.
</div>
</v-col>
</v-row>
<!-- E3: a grouping stays open, and these decide for how long and how
loudly it announces that it grew. -->
<div class="text-caption fc-muted mt-4 mb-2">
A grouped post stays <strong>open</strong>: variants the creator adds
later join the existing post instead of starting a new one, and its
text grows with them.
</div>
<v-row>
<v-col cols="12" sm="4">
<SettingNumberField
v-model="local.discord_group_close_after_hours"
label="Stays open for (hours)" :min="1" :step="24"
density="comfortable" max-width="none"
:disabled="!local.discord_grouping_enabled" @change="onSave"
/>
<div class="text-caption fc-muted mt-1">
How long after its last addition a post still accepts new variants.
Not the drop window above that cuts one session into drops; this
decides how late a follow-up can still join. Too long and the same
character coming round again months later gets absorbed by mistake.
</div>
</v-col>
<v-col cols="12" sm="4">
<SettingNumberField
v-model="local.discord_group_resurface_min_images"
label="New images before it resurfaces" :min="1" :step="1"
density="comfortable" max-width="none"
:disabled="!local.discord_grouping_enabled" @change="onSave"
/>
<div class="text-caption fc-muted mt-1">
Growth smaller than this updates the post where it sits instead of
moving it back to the top of the feed.
</div>
</v-col>
<v-col cols="12" sm="4">
<SettingNumberField
v-model="local.discord_group_resurface_cooldown_hours"
label="Resurface at most every (hours)" :min="0" :step="1"
density="comfortable" max-width="none"
:disabled="!local.discord_grouping_enabled" @change="onSave"
/>
<div class="text-caption fc-muted mt-1">
Together with the count above, this is what stops a post that gains
an image a day from living permanently at the top of the feed.
</div>
</v-col>
</v-row>
</div>
<div v-else><v-skeleton-loader type="paragraph" /></div>
</MaintenanceTile>
</template>
<script setup>
// #388 E2. Every value here is operator-tunable because the quality bar is a
// judgement no test can settle: grouping too greedy merges distinct pieces,
// too shy leaves a drop scattered. The two failure modes are not symmetric —
// scattered is visible and fixable, a wrong merge is a post asserting that
// unrelated art belongs together — so the shipped default sits on the tight
// side and this card is how it gets loosened.
import { reactive, watch } from 'vue'
import { useSettingSave } from '../../composables/useSettingSave.js'
import { useMLStore } from '../../stores/ml.js'
import MaintenanceTile from '../common/MaintenanceTile.vue'
import SettingNumberField from '../common/SettingNumberField.vue'
const store = useMLStore()
const { save } = useSettingSave(store.patchSettings)
const local = reactive({})
watch(() => store.settings, (s) => { if (s) Object.assign(local, s) }, { immediate: true })
function onSave() {
save({
discord_grouping_enabled: Boolean(local.discord_grouping_enabled),
discord_group_max_distance: Number(local.discord_group_max_distance),
discord_group_window_minutes: Number(local.discord_group_window_minutes),
discord_group_close_after_hours: Number(local.discord_group_close_after_hours),
discord_group_resurface_min_images: Number(local.discord_group_resurface_min_images),
discord_group_resurface_cooldown_hours:
Number(local.discord_group_resurface_cooldown_hours),
})
}
</script>
@@ -14,6 +14,10 @@
<div class="fc-tile-stack">
<ImportFiltersForm />
<TranslationCard />
<DiscordGroupingCard />
<PostAssociationsCard />
<MembershipRosterCard />
<MembershipSuggestionsCard />
</div>
</section>
@@ -80,6 +84,10 @@ import DbMaintenanceCard from './DbMaintenanceCard.vue'
import VideoEmbeddingCard from './VideoEmbeddingCard.vue'
import CropProposersCard from './CropProposersCard.vue'
import HeadsCard from './HeadsCard.vue'
import DiscordGroupingCard from './DiscordGroupingCard.vue'
import MembershipRosterCard from './MembershipRosterCard.vue'
import MembershipSuggestionsCard from './MembershipSuggestionsCard.vue'
import PostAssociationsCard from './PostAssociationsCard.vue'
import GpuAgentCard from './GpuAgentCard.vue'
import AliasTable from './AliasTable.vue'
import BackupCard from './BackupCard.vue'
@@ -0,0 +1,97 @@
<template>
<MaintenanceTile
icon="mdi-account-heart-outline"
title="Subscription roster"
blurb="What your accounts actually subscribe to, as the platform reports it."
>
<div class="text-caption fc-muted mb-3">
FC reads the list of memberships from each connected platform so it can
tell which of your sources you still pay for. It only ever reads nothing
is added, removed or cancelled from here.
</div>
<div v-if="store.loading && !store.platforms.length" class="text-caption fc-muted">
Checking
</div>
<div
v-else-if="!store.platforms.length"
class="text-caption fc-muted"
>
No platform has been synced yet. Connect a credential in Subscriptions,
then use Sync now.
</div>
<div v-for="p in store.platforms" :key="p.platform" class="fc-roster">
<div class="fc-roster__head">
<strong>{{ p.platform }}</strong>
<!-- The three states this card exists to keep apart. "Never synced" is
words, never a number rendering it as 0 is the exact conflation
that would have C4 telling the operator to cancel things they pay
for. -->
<span v-if="!p.last_success_at" class="fc-roster__never">
never synced
</span>
<span v-else-if="!p.fresh" class="fc-roster__stale">
last synced {{ formatRelative(p.last_success_at) }} too old to rely on
</span>
<span v-else class="fc-roster__ok">
synced {{ formatRelative(p.last_success_at) }}
</span>
</div>
<div class="text-caption fc-muted">
<!-- Only stated when there is a successful sync behind it. A count
with no sync behind it is a guess wearing a number's clothes. -->
<template v-if="p.last_success_at">
{{ p.last_count }} membership{{ p.last_count === 1 ? '' : 's' }} found
</template>
<template v-else-if="p.last_attempt_at">
tried {{ formatRelative(p.last_attempt_at) }}, no successful sync yet
</template>
</div>
<div v-if="p.last_error_type" class="fc-roster__err text-caption">
{{ p.last_error_type }}: {{ p.last_error_message }}
</div>
</div>
<div class="d-flex align-center mt-4" style="gap:12px">
<v-spacer />
<v-btn size="small" variant="text" :loading="store.loading" @click="refresh">
Sync now
</v-btn>
</div>
</MaintenanceTile>
</template>
<script setup>
import { onMounted } from 'vue'
import { formatRelative } from '../../utils/date.js'
import { useMembershipSyncStore } from '../../stores/membershipSync.js'
import MaintenanceTile from '../common/MaintenanceTile.vue'
const store = useMembershipSyncStore()
async function refresh () {
await store.syncNow()
await store.load()
}
onMounted(() => { store.load() })
</script>
<style scoped>
.fc-roster {
padding: 8px 0;
border-top: 1px solid rgba(var(--v-theme-on-surface), 0.12);
}
.fc-roster__head { display: flex; align-items: baseline; gap: 8px; }
/* Never-synced and stale are the states that must not read as healthy, so
they take colour and the healthy one does not. */
.fc-roster__never { color: rgb(var(--v-theme-on-surface-variant)); font-style: italic; }
.fc-roster__stale { color: rgb(var(--v-theme-warning, var(--v-theme-accent))); }
.fc-roster__ok { color: rgb(var(--v-theme-on-surface-variant)); }
.fc-roster__err { color: rgb(var(--v-theme-error)); }
</style>
@@ -0,0 +1,74 @@
<template>
<MaintenanceTile
icon="mdi-account-multiple-plus-outline"
title="Same creator, another channel"
blurb="Creators you subscribe to who look like creators you already track."
>
<div class="text-caption fc-muted mb-3">
When a membership in your roster looks like a creator you already follow
elsewhere a Discord server, say FC offers to add the missing channel
to that same creator. Accepting <strong>adds a source</strong>; it never
merges two creators together.
</div>
<div class="d-flex align-center mb-2" style="gap:12px">
<strong class="text-body-2">
{{ store.suggestions.length }} waiting for review
</strong>
<v-spacer />
<v-btn size="small" variant="text" :loading="store.loading" @click="store.rescan">
Look again
</v-btn>
</div>
<div v-if="!store.suggestions.length" class="text-caption fc-muted">
Nothing proposed. That is the usual state a pair only appears when the
names line up, or when one of your creator's posts links to the other
channel.
</div>
<div v-for="s in store.suggestions" :key="s.id" class="fc-sugg">
<div class="fc-sugg__body">
<div>
<strong>{{ s.artist.name }}</strong>
<span class="fc-sugg__arrow">and your</span>
<strong>{{ s.membership.platform }}</strong>
membership
<em>{{ s.membership.display_name }}</em>
</div>
<!-- The per-signal breakdown, not just the total: "why did it suggest
this" is the question the operator actually has, and a lone
percentage cannot answer it. -->
<div class="text-caption fc-muted">
{{ Math.round(s.score * 100) }}% ·
name {{ Math.round((s.signals?.name ?? 0) * 100) }}% ·
links to it {{ (s.signals?.declared ?? 0) > 0 ? 'yes' : 'no' }}
</div>
</div>
<v-btn size="small" variant="tonal" @click="store.accept(s.id)">Add channel</v-btn>
<v-btn size="small" variant="text" @click="store.dismiss(s.id)">Dismiss</v-btn>
</div>
</MaintenanceTile>
</template>
<script setup>
import { onMounted } from 'vue'
import { useMembershipSuggestionsStore } from '../../stores/membershipSuggestions.js'
import MaintenanceTile from '../common/MaintenanceTile.vue'
const store = useMembershipSuggestionsStore()
onMounted(() => { store.load() })
</script>
<style scoped>
.fc-sugg {
display: flex;
align-items: center;
gap: 8px;
padding: 8px 0;
border-top: 1px solid rgba(var(--v-theme-on-surface), 0.12);
}
.fc-sugg__body { flex: 1; min-width: 0; }
.fc-sugg__arrow { color: rgb(var(--v-theme-on-surface-variant)); margin: 0 6px; }
</style>
@@ -0,0 +1,128 @@
<template>
<MaintenanceTile
icon="mdi-link-variant"
title="Announcement links"
blurb="Pair a Patreon teaser with the Discord drop it announced."
>
<div class="text-caption fc-muted mb-3">
Some creators post a cropped fragment on Patreon to say the real thing has
landed in their Discord. FC proposes those pairs; nothing is linked until
you accept one. A wrong link would tell you two different pieces are the
same, so this stays a suggestion.
</div>
<v-switch
v-model="enabled" color="accent" hide-details density="compact"
label="Look for announcement pairs"
@update:model-value="store.setEnabled"
/>
<v-row class="mt-2">
<v-col cols="12" sm="6">
<SettingNumberField
v-model="threshold" label="Confidence needed"
:min="0" :max="1" :step="0.05"
density="comfortable" max-width="none"
:disabled="!enabled" @change="store.setThreshold(Number(threshold))"
/>
<div class="text-caption fc-muted mt-1">
A pair always needs <strong>two</strong> reasons being close in time
and the post itself mentioning Discord. Neither is enough alone at any
setting at or above 0.55, which is what stops a busy posting day from
producing false pairs.
</div>
</v-col>
<v-col cols="12" sm="6">
<SettingNumberField
v-model="windowHours" label="How far apart (hours)"
:min="1" :step="1"
density="comfortable" max-width="none"
:disabled="!enabled" @change="store.setWindowHours(Number(windowHours))"
/>
<div class="text-caption fc-muted mt-1">
The announcement exists in order to point at the drop, so the two are
usually minutes to hours apart.
</div>
</v-col>
</v-row>
<div class="d-flex align-center mt-4 mb-2" style="gap:12px">
<strong class="text-body-2">
{{ store.proposals.length }} waiting for review
</strong>
<v-spacer />
<v-btn size="small" variant="text" :loading="store.loading" @click="store.rescan">
Scan now
</v-btn>
</div>
<div v-if="!store.proposals.length" class="text-caption fc-muted">
Nothing proposed. That is the expected state most of the time pairs only
appear when a post both lands near a drop and says it is about Discord.
</div>
<div
v-for="p in store.proposals" :key="p.id"
class="fc-assoc"
>
<div class="fc-assoc__body">
<RouterLink :to="{ name: 'latest', query: { post_id: p.announcement_post_id } }">
post {{ p.announcement_post_id }}
</RouterLink>
<span class="fc-assoc__arrow">announced</span>
<RouterLink :to="{ name: 'latest', query: { post_id: p.payload_post_id } }">
drop {{ p.payload_post_id }}
</RouterLink>
<!-- The per-signal breakdown, not just the total: "why did it suggest
this" is the question the operator actually has, and a lone score
cannot answer it. -->
<div class="text-caption fc-muted">
{{ Math.round(p.score * 100) }}% ·
timing {{ Math.round((p.signals?.proximity ?? 0) * 100) }}% ·
says so {{ Math.round((p.signals?.declared ?? 0) * 100) }}%
</div>
</div>
<v-btn size="small" variant="tonal" @click="store.accept(p.id)">Link</v-btn>
<v-btn size="small" variant="text" @click="store.dismiss(p.id)">Dismiss</v-btn>
</div>
</MaintenanceTile>
</template>
<script setup>
import { onMounted, ref, watch } from 'vue'
import { RouterLink } from 'vue-router'
import { usePostAssociationsStore } from '../../stores/postAssociations.js'
import MaintenanceTile from '../common/MaintenanceTile.vue'
import SettingNumberField from '../common/SettingNumberField.vue'
const store = usePostAssociationsStore()
const enabled = ref(true)
const threshold = ref(0.6)
const windowHours = ref(24)
watch(() => store.enabled, (v) => { enabled.value = v }, { immediate: true })
watch(() => store.threshold, (v) => { threshold.value = v }, { immediate: true })
watch(() => store.windowHours, (v) => { windowHours.value = v }, { immediate: true })
onMounted(async () => {
// Both swallow their own failures: a settings read that fails should not
// leave the queue unrendered, and vice versa.
await Promise.allSettled([store.loadSettings(), store.load()])
})
</script>
<style scoped>
.fc-assoc {
display: flex;
align-items: center;
gap: 8px;
padding: 8px 0;
border-top: 1px solid rgba(var(--v-theme-on-surface), 0.12);
}
.fc-assoc__body { flex: 1; min-width: 0; }
.fc-assoc__arrow {
color: rgb(var(--v-theme-on-surface-variant));
margin: 0 6px;
}
</style>
@@ -0,0 +1,128 @@
<template>
<!-- A Settings tab, not a page of its own (operator 2026-09-02): the first
cut hung this off the health dot alone, which is a target you have to
already suspect something to look for. Settings is where someone goes
to ask the instance about itself, so it lives beside Activity. -->
<div>
<div class="d-flex align-center mb-1">
<v-spacer />
<span class="fc-sys__checked">
{{ store.checkedAt ? `checked ${formatRelative(store.checkedAt)}` : 'checking…' }}
</span>
</div>
<p class="fc-sys__lede text-body-2 mb-5">
Every moving part of FabledCurator and whether it is still checking in.
Parts are learned as they appear, so anything that has run at least once
stays listed that is what lets a stopped one be noticed rather than
simply vanishing.
</p>
<v-alert
v-if="store.lastError" type="error" variant="tonal" density="compact" class="mb-4"
>
Could not reach FabledCurator: {{ store.lastError }}
</v-alert>
<v-card v-else variant="flat" class="fc-sys__card">
<div v-if="!store.parts.length" class="pa-6 text-center fc-sys__muted">
Still gathering this fills in on the first check.
</div>
<div
v-for="part in store.parts" :key="part.key"
class="fc-sys__row" :class="`fc-sys__row--${part.state}`"
>
<span class="fc-sys__dot" :class="`fc-sys__dot--${part.state}`" />
<div class="fc-sys__body">
<div class="fc-sys__name">
{{ part.name }}
<span class="fc-sys__kind">{{ kindLabel(part.kind) }}</span>
</div>
<!-- The sentence, not just a chip. At the moment someone is deciding
whether to go and open Portainer, "has not checked in for 6 min"
is the thing that answers them. -->
<div class="fc-sys__detail">{{ part.detail }}</div>
</div>
<div class="fc-sys__meta">
<div v-if="part.last_seen_at" :title="part.last_seen_at">
seen {{ formatRelative(part.last_seen_at) }}
</div>
<div v-if="part.latency_ms != null">{{ part.latency_ms }} ms</div>
<div v-if="part.queues?.length" class="fc-sys__queues">{{ part.queues.join(', ') }}</div>
</div>
</div>
</v-card>
<p v-if="store.thresholds" class="fc-sys__foot text-caption mt-4">
A part is called stale after
{{ Math.round(store.thresholds.stale_after_seconds / 60) }} min without a
check-in and treated as stopped after
{{ Math.round(store.thresholds.down_after_seconds / 60) }} min. The window
is deliberately wide: a rolling deploy briefly runs two of a service and
then neither, and an indicator that reddened on every update would stop
being read.
</p>
</div>
</template>
<script setup>
import { onMounted } from 'vue'
import { useSystemHealthStore } from '../../stores/systemHealth.js'
import { formatRelative } from '../../utils/date.js'
const store = useSystemHealthStore()
function kindLabel(kind) {
if (kind === 'celery') return 'background worker'
if (kind === 'agent') return 'GPU agent'
if (kind === 'datastore') return 'datastore'
return kind
}
// No timer of its own. TopNav already polls this same pinia store every 15s
// for the health dot, and it is mounted on every route this tab is reachable
// from — a second interval here would just double the request rate for a 5s
// freshness gain. v-window keeps a visited item MOUNTED (hidden, not
// destroyed), so a local timer would also have kept firing behind Maintenance.
// One refresh on open, so arriving at the tab doesn't wait out the nav's tick.
onMounted(() => { store.refresh() })
</script>
<style scoped>
.fc-sys__lede, .fc-sys__muted, .fc-sys__checked, .fc-sys__foot {
color: rgb(var(--v-theme-on-surface) / 0.66);
}
.fc-sys__checked { font-size: 0.78rem; }
.fc-sys__card { background: rgb(var(--v-theme-on-surface) / 0.04); }
.fc-sys__row {
display: flex; align-items: center; gap: 12px;
padding: 12px 16px;
border-bottom: 1px solid rgb(var(--v-theme-on-surface) / 0.08);
}
.fc-sys__row:last-child { border-bottom: 0; }
.fc-sys__dot { width: 9px; height: 9px; border-radius: 50%; flex: 0 0 auto; }
.fc-sys__dot--ok { background: rgb(var(--v-theme-success)); }
.fc-sys__dot--stale { background: rgb(var(--v-theme-warning)); }
.fc-sys__dot--down { background: rgb(var(--v-theme-error)); }
.fc-sys__dot--unknown { background: rgb(var(--v-theme-on-surface) / 0.35); }
.fc-sys__body { min-width: 0; flex: 1 1 auto; }
.fc-sys__name { font-weight: 600; }
.fc-sys__kind {
margin-left: 8px; font-weight: 400; font-size: 0.72rem; text-transform: uppercase;
letter-spacing: 0.04em; color: rgb(var(--v-theme-on-surface) / 0.5);
}
.fc-sys__detail { font-size: 0.82rem; color: rgb(var(--v-theme-on-surface) / 0.72); }
.fc-sys__meta {
text-align: right; font-size: 0.75rem; flex: 0 0 auto;
font-variant-numeric: tabular-nums; color: rgb(var(--v-theme-on-surface) / 0.6);
}
.fc-sys__queues { opacity: 0.75; }
</style>
@@ -32,6 +32,13 @@
v-if="e.live.errors" class="fc-active__count fc-active__count--err"
title="errors"
> {{ e.live.errors }}</span>
<!-- Tier-gated posts (#874): not an error the walk is working, the
content just isn't ours. Shown only when non-zero so a healthy
run stays uncluttered. -->
<span
v-if="e.live.gated" class="fc-active__count fc-active__count--gated"
title="posts skipped — no access at your tier"
>🔒 {{ e.live.gated }}</span>
<span class="fc-active__count fc-active__count--posts" title="posts scanned">
{{ e.live.posts }} posts</span>
</span>
@@ -141,6 +148,9 @@ function elapsed (startedIso) {
}
.fc-active__count { color: rgb(var(--v-theme-on-surface-variant)); }
.fc-active__count--err { color: rgb(var(--v-theme-error)); }
/* Matches FailingSourcesCard's severity map, which already colours
tier_limited as 'info' no-access is information, not a failure. */
.fc-active__count--gated { color: rgb(var(--v-theme-info)); }
.fc-active__count--posts { opacity: 0.7; }
@keyframes fc-active-pulse {
0%, 100% { opacity: 1; transform: scale(1); }
@@ -0,0 +1,144 @@
<template>
<!-- Renders nothing when there is nothing to say same posture as
NeedsAttentionCard. A reconciliation card that always shows would train
the operator to scroll past it. -->
<v-card v-if="anythingToSay" variant="tonal" class="mb-4 fc-recon">
<v-card-text>
<div v-for="p in interesting" :key="p.platform" class="fc-recon__platform">
<!-- Bucket 1: the adoption win. The only direction with an action,
because adding a source is the reversible half. -->
<template v-if="p.subscribed_not_tracked.length">
<div class="fc-recon__head">
<v-icon icon="mdi-account-plus-outline" size="small" class="me-2" />
<strong>
{{ p.subscribed_not_tracked.length }}
{{ p.platform }}
{{ p.subscribed_not_tracked.length === 1 ? 'creator' : 'creators' }}
you subscribe to but don't follow here
</strong>
</div>
<div v-for="m in p.subscribed_not_tracked" :key="m.id" class="fc-recon__row">
<div class="fc-recon__body">
<strong>{{ m.display_name }}</strong>
<span v-if="m.tier_names?.length" class="fc-recon__dim">
· {{ m.tier_names.join(', ') }}
</span>
<!-- A membership whose status this build has not been taught is
still offered, and says so rather than being presented as a
confirmed active subscription. -->
<span v-if="m.paid_access === null" class="fc-recon__dim">
· status not recognised ({{ m.status }})
</span>
</div>
<v-btn size="small" variant="tonal" @click="store.adopt(m.id)">
Add subscription
</v-btn>
</div>
</template>
<!-- Bucket 2: REPORT ONLY. No disable control here by design — the
source list on this same page is where that decision belongs. -->
<template v-if="p.tracked_not_subscribed.length">
<div class="fc-recon__head">
<v-icon icon="mdi-help-circle-outline" size="small" class="me-2" />
<strong>
{{ p.tracked_not_subscribed.length }}
{{ p.platform }}
{{ p.tracked_not_subscribed.length === 1 ? 'source' : 'sources' }}
your roster doesn't account for
</strong>
</div>
<div v-for="s in p.tracked_not_subscribed" :key="s.id" class="fc-recon__row">
<div class="fc-recon__body">
<strong>{{ s.artist.name }}</strong>
<div class="fc-recon__dim">{{ reasonFor(s) }}</div>
</div>
</div>
<div class="fc-recon__dim fc-recon__note">
Nothing has been changed. If you want one of these to stop checking,
disable it in the list below.
</div>
</template>
<!-- The roster could not be trusted, so bucket 2 was not computed. Said
in words: an empty list here must never read as "all clear". -->
<div v-if="!p.fresh" class="fc-recon__dim fc-recon__note">
<v-icon icon="mdi-clock-alert-outline" size="small" class="me-1" />
<template v-if="!p.last_success_at">
Your {{ p.platform }} roster has never synced, so FC can't tell which
sources you still subscribe to.
</template>
<template v-else>
Your {{ p.platform }} roster last synced
{{ formatRelative(p.last_success_at) }}, which is too old to draw
conclusions from.
</template>
Sync it from Settings Subscription roster.
</div>
</div>
</v-card-text>
</v-card>
</template>
<script setup>
import { computed, onMounted } from 'vue'
import { formatRelative } from '../../utils/date.js'
import { useMembershipReconcileStore } from '../../stores/membershipReconcile.js'
// #387 C4. Two directions, deliberately unequal: subscriptions FC doesn't
// follow get a one-click add, while sources the roster doesn't account for are
// REPORTED ONLY (operator decision, 2026-09-11). The asymmetry is the point —
// adding a source is trivially undone, and "you no longer subscribe to this" is
// computed from an absence that has three possible causes.
const store = useMembershipReconcileStore()
// An untrustworthy roster is only worth mentioning when there is something it
// would have reconciled — otherwise a failed sweep on a platform with no
// sources yet would put a warning on a page with nothing to warn about.
const interesting = computed(() => store.platforms.filter(
p => p.subscribed_not_tracked.length
|| p.tracked_not_subscribed.length
|| (!p.fresh && p.tracked_total)
))
const anythingToSay = computed(() => interesting.value.length > 0)
function reasonFor (s) {
if (s.basis === 'lapsed') {
return `Your membership reads "${s.membership?.status}" — you're not a paying supporter of this creator right now.`
}
if (s.basis === 'absent_exact') {
return "FC knows this creator's id on the platform, and it isn't in your roster."
}
// absent_handle — the weakest claim, and it says so. A renamed creator looks
// exactly like this, which is why it is not phrased as a conclusion.
return "No membership matched this source's address. It may simply have been renamed."
}
onMounted(() => { store.load() })
</script>
<style scoped>
.fc-recon__platform + .fc-recon__platform {
margin-top: 16px;
}
.fc-recon__head {
display: flex;
align-items: center;
margin: 8px 0 4px;
}
.fc-recon__row {
display: flex;
align-items: center;
gap: 8px;
padding: 6px 0;
border-top: 1px solid rgba(var(--v-theme-on-surface), 0.12);
}
.fc-recon__body { flex: 1; min-width: 0; }
/* Rule 96: dim text is the theme token, never bare opacity. */
.fc-recon__dim {
color: rgb(var(--v-theme-on-surface-variant));
font-size: 0.8125rem;
}
.fc-recon__note { margin-top: 8px; }
</style>
@@ -27,7 +27,10 @@ import { computed } from 'vue'
import { formatRelative } from '../../utils/date.js'
const props = defineProps({
// { last_tick_at, next_due_at, due_now, auto_sources } | null
// { last_tick_at, next_due_at, due_now, auto_sources,
// failing_sources, no_access_sources, platform_cooldowns } | null
// This bar renders the scheduling half; the ingestion counts are read by the
// front door's FeedStatusRibbon (#387 B3) off the same payload.
status: { type: Object, default: null },
})
@@ -77,7 +77,7 @@ const recapturing = computed(() => !!props.source.backfill_recapture)
// Recover / recapture are native-ingester features (ledger-bypass re-walk and
// post-text re-grab), available to every native platform — not just Patreon.
// Mirrors backend download_backends.NATIVE_INGESTER_PLATFORMS.
const NATIVE_PLATFORMS = ['patreon', 'subscribestar', 'pixiv']
const NATIVE_PLATFORMS = ['patreon', 'subscribestar']
const isNative = computed(() => NATIVE_PLATFORMS.includes(props.source.platform))
</script>
@@ -10,6 +10,10 @@
<div class="fc-health-tip">
<div>Last checked: {{ lastCheckedText }}</div>
<div v-if="nextCheckText">Next check: {{ nextCheckText }}</div>
<div v-if="noAccess" class="fc-health-tip__gated">{{ noAccessText }}</div>
<div v-if="noAccess && noAccessReason" class="fc-health-tip__why">
{{ noAccessReason }}
</div>
<div v-if="(source.consecutive_failures || 0) > 0">
Failures: {{ source.consecutive_failures }}
</div>
@@ -29,14 +33,42 @@ const props = defineProps({
warningThreshold: { type: Number, default: 5 },
})
const noAccess = computed(() => props.source.error_type === 'tier_limited')
const level = computed(() => {
if (!props.source.last_checked_at) return 'unchecked'
const f = props.source.consecutive_failures || 0
if (f === 0) return 'healthy'
// No-access outranks 'healthy' but is NOT a failure grade: the walk worked,
// the content simply isn't ours. Checked after failures so a source that is
// genuinely erroring still reads as erroring.
if (f === 0) return noAccess.value ? 'no-access' : 'healthy'
if (f < props.warningThreshold) return 'warning'
return 'critical'
})
// The count comes from the last walk's run_stats and is only joined in by the
// list endpoint, so it can legitimately be absent — say the state without it
// rather than printing a fabricated zero.
const noAccessText = computed(() => {
const n = props.source.tier_gated_count
return n
? `${n} post${n === 1 ? '' : 's'} you don't have access to`
: "Some posts are behind a tier you don't hold"
})
// #387 C5: the learned roster's explanation for the line above, when it has
// one. The backend sends null for every case where the roster is not evidence —
// campaign absent, roster stale, never swept, status not yet characterised — so
// there is deliberately NO fallback sentence here. A default would turn "we
// don't know why" into a reason, which is the one thing this step must not do.
const GATED_REASONS = {
lapsed: "You're not a patron any more — resubscribe, or disable this source.",
tier: "Your tier doesn't cover these posts — upgrade, or leave them be.",
free: "You follow this creator for free — these posts are for paying patrons.",
}
const noAccessReason = computed(() => GATED_REASONS[props.source.gated_reason] || null)
const ariaLabel = computed(() => `source health: ${level.value}`)
const lastCheckedText = computed(() => formatRelative(props.source.last_checked_at))
@@ -63,6 +95,9 @@ const truncatedError = computed(() => {
}
.fc-health-dot--unchecked { background-color: rgb(var(--v-theme-on-surface-variant)); opacity: 0.5; }
.fc-health-dot--healthy { background-color: rgb(var(--v-theme-success, 76 175 80)); }
/* Matches the 'info' severity FailingSourcesCard already assigns tier_limited —
deliberately not a warning/error hue: nothing is broken. */
.fc-health-dot--no-access { background-color: rgb(var(--v-theme-info, 33 150 243)); }
.fc-health-dot--warning { background-color: rgb(var(--v-theme-warning, 255 167 38)); }
.fc-health-dot--critical { background-color: rgb(var(--v-theme-error, 244 67 54)); }
@@ -70,6 +105,15 @@ const truncatedError = computed(() => {
font-size: 0.85rem;
line-height: 1.4;
}
.fc-health-tip__gated {
color: rgb(var(--v-theme-info, 33 150 243));
}
/* The reason is subordinate to the count it explains: same block, quieter, so
a tooltip that has one does not read as two separate findings. */
.fc-health-tip__why {
color: rgb(var(--v-theme-on-surface-variant));
max-width: 24rem;
}
.fc-health-tip__err {
margin-top: 0.25rem;
color: rgb(var(--v-theme-error, 244 67 54));
@@ -48,6 +48,20 @@
<span class="fc-source-row__err-text">{{ source.last_error }}</span>
</v-tooltip>
</v-chip>
<!-- No access (#387 A3). Sits directly after the failure chip and before
the backfill states: a source we can't see is the more useful thing
to say about it than which walk phase it's in, and unlike those it
doesn't resolve on its own. Info-coloured, never error — the walk
worked, the content just isn't ours. -->
<v-chip
v-else-if="source.error_type === 'tier_limited'"
size="x-small" color="info" variant="tonal" label
prepend-icon="mdi-lock-outline"
>{{ source.tier_gated_count ? `${source.tier_gated_count} gated` : 'No access' }}
<v-tooltip activator="parent" location="top" max-width="480">
<span>{{ noAccessTip }}</span>
</v-tooltip>
</v-chip>
<v-chip
v-else-if="source.backfill_state === 'running'"
size="x-small" color="info" variant="tonal" label
@@ -79,6 +93,8 @@
</template>
<script setup>
import { computed } from 'vue'
import SourceActions from './SourceActions.vue'
import SourceHealthDot from './SourceHealthDot.vue'
import { formatRelative } from '../../utils/date.js'
@@ -88,6 +104,19 @@ const props = defineProps({
checking: { type: Boolean, default: false },
warningThreshold: { type: Number, default: 5 },
})
// Says what to DO about it, not just what it is — the action here is the
// operator's subscription, not anything FC can retry. The count is only joined
// in by the list endpoint, so phrase it without one when it's absent rather
// than rendering a fabricated zero.
const noAccessTip = computed(() => {
const n = props.source.tier_gated_count
const what = n
? `The last check skipped ${n} post${n === 1 ? '' : 's'}`
: 'The last check skipped posts'
return `${what} this account can't view. Nothing is broken — your `
+ 'subscription tier does not grant access to them.'
})
const emit = defineEmits(['edit', 'remove', 'toggle', 'check', 'backfill', 'recover', 'recapture'])
function onToggleEnabled(value) {
@@ -7,6 +7,9 @@
nothing to say. -->
<NeedsAttentionCard />
<RecentArrivalsCard />
<!-- #387 C4: what you pay for vs. what FC follows. Renders nothing when
the two agree. -->
<MembershipReconcileCard />
<div class="fc-subs__bar">
<v-btn color="accent" prepend-icon="mdi-plus" @click="openAddSource(null)">
@@ -327,6 +330,7 @@ import { usePlatformsStore } from '../../stores/platforms.js'
import { useImportStore } from '../../stores/import.js'
import NeedsAttentionCard from './NeedsAttentionCard.vue'
import RecentArrivalsCard from './RecentArrivalsCard.vue'
import MembershipReconcileCard from './MembershipReconcileCard.vue'
import SourceRow from './SourceRow.vue'
import SourceCard from './SourceCard.vue'
import SourceHealthDot from './SourceHealthDot.vue'
@@ -348,6 +352,12 @@ const STATUS_OPTIONS = [
{ title: 'Enabled', value: 'enabled' },
{ title: 'Disabled', value: 'disabled' },
{ title: 'Has errors', value: 'errors' },
// #387 A3: no-access is deliberately its own filter and NOT folded into
// "Has errors" — nothing failed, and it is the only status here whose fix is
// the operator's subscription rather than anything FC can retry. Without a
// filter a gated source is invisible in a long list, since it correctly
// stays out of the failing rollup.
{ title: 'No access', value: 'no_access' },
{ title: 'Stale', value: 'stale' },
]
@@ -362,7 +372,19 @@ const platformsStore = usePlatformsStore()
const importStore = useImportStore()
const search = ref('')
const statusFilter = ref('all')
// URL-addressable (#387 B3) so the front-door status ribbon can link straight
// to "the sources this number is about" — a count that lands you on an
// unfiltered list makes the reader do the filtering the ribbon just did.
// Mirrors how artistFilter already reads from route.query below.
const statusFilter = computed({
get: () => route.query.status || 'all',
set: (v) => {
const q = { ...route.query }
if (!v || v === 'all') delete q.status
else q.status = v
router.replace({ query: q })
},
})
const needsAttention = ref(false)
const expanded = ref([])
const selected = ref([])
@@ -506,6 +528,7 @@ function groupMatchesStatus(g, status) {
if (status === 'enabled') return g.sources.some((s) => s.enabled)
if (status === 'disabled') return g.sources.every((s) => !s.enabled)
if (status === 'errors') return g.sources.some((s) => (s.consecutive_failures || 0) > 0)
if (status === 'no_access') return g.sources.some((s) => s.error_type === 'tier_limited')
if (status === 'stale') return g.sources.some((s) => !s.last_checked_at)
return true
}
+30 -1
View File
@@ -9,15 +9,37 @@ import SeriesView from './views/SeriesView.vue'
import SeriesManageView from './views/SeriesManageView.vue'
import SeriesReaderView from './views/SeriesReaderView.vue'
import SubscriptionsView from './views/SubscriptionsView.vue'
import PostsView from './views/PostsView.vue'
// The application's front door. `/` redirects here. Changing the front door
// is a one-line edit (e.g. '/gallery' or '/tags').
export const FRONT_DOOR = '/showcase'
//
// Moved from '/showcase' to '/latest' (milestone #387 B1). Showcase answers
// "show me something" — a random TABLESAMPLE, lean-back, and it can never tell
// you anything is wrong. The feed answers "what arrived?", which is a question
// the operator has every day, and it is the only view where a failing source
// shows up on its own as a creator who has gone quiet. The per-artist "new
// since last visit" badges (#597) were a workaround for this view not being
// the door. Showcase is demoted to a nav entry, not removed.
export const FRONT_DOOR = '/latest'
const routes = [
// Root is a redirect only — no meta.title so it stays out of the nav.
{ path: '/', redirect: FRONT_DOOR },
// The front door (#387 B1): the post feed as its own surface. Reuses
// PostsView unchanged — it is already a self-contained feed (own container,
// own store, infinite scroll, deep-link anchoring), and Browse only ever
// wrapped it in a tab strip. Mounting it directly IS the difference the
// promotion was after: a door you arrive at, not a hub you navigate out of.
//
// No stickyChrome: unlike Browse/Gallery/Settings this view has no sticky
// sub-header for the nav to butt against, so the nav keeps its normal fade.
// `props` turns on the ingestion-status ribbon (#387 B3). Only here: inside
// Browse's Posts tab the same view renders without it.
{ path: '/latest', name: 'latest', component: PostsView, props: { statusRibbon: true },
meta: { title: 'Latest', navOrder: 5 } },
// FC-2: image backbone
{ path: '/showcase', name: 'showcase', component: ShowcaseView, meta: { title: 'Showcase', navOrder: 10 } },
{ path: '/gallery', name: 'gallery', component: GalleryView, meta: { title: 'Gallery', navOrder: 20, stickyChrome: true } },
@@ -45,6 +67,13 @@ const routes = [
// Settings — config, pinned to the right of the nav (TopNav special-cases it).
{ path: '/settings', name: 'settings', component: SettingsView, meta: { title: 'Settings', stickyChrome: true } },
// System health is a Settings TAB, not a route of its own (operator
// 2026-09-02). It first shipped as /system reachable only from the health
// dot, which is a target you have to already suspect something to go
// looking for. Settings is where someone goes to ask the instance about
// itself. The path stays as a redirect so the dot's old link, and any
// bookmark from that build, still land somewhere real.
{ path: '/system', name: 'system', redirect: () => ({ name: 'settings', query: { tab: 'system' } }) },
// The old standalone paths now redirect into the Browse hub, preserving any
// deep-link query (e.g. /posts?post_id=N → /browse?tab=posts&post_id=N). The
@@ -0,0 +1,48 @@
import { defineStore } from 'pinia'
import { ref } from 'vue'
import { useApi } from '../composables/useApi.js'
import { useAsyncAction } from '../composables/useAsyncAction.js'
import { toast } from '../utils/toast.js'
// Backs the reconciliation card (#387 C4). Two directions with deliberately
// different weights: adopting a creator you already pay for is a per-row
// action, while "you follow this and your roster doesn't show it" only ever
// REPORTS — the operator disables from the Subscriptions row itself.
//
// `fresh` is the field that matters most here. A stale or never-run roster
// returns an empty tracked_not_subscribed list, and the card must say WHY it is
// empty rather than letting it read as "everything is fine".
export const useMembershipReconcileStore = defineStore('membershipReconcile', () => {
const api = useApi()
const platforms = ref([])
const { loading, error, run } = useAsyncAction({ errorAs: 'message' })
async function load () {
await run(async () => {
const body = await api.get('/api/sources/reconciliation')
platforms.value = body.platforms || []
})
}
async function adopt (membershipId) {
try {
const res = await api.post('/api/sources/reconciliation/adopt', {
membership_id: membershipId
})
toast({
text: res.already_tracked
? 'Already tracked — nothing to add'
: 'Now tracking this creator',
type: 'success'
})
// Reload rather than splice: adopting moves the row from one bucket to
// another, and the counts beside them have to move with it.
await load()
} catch (e) {
toast({ text: `Could not add: ${e?.body?.detail || e.message}`, type: 'error' })
}
}
return { platforms, loading, error, load, adopt }
})
@@ -0,0 +1,52 @@
import { defineStore } from 'pinia'
import { ref } from 'vue'
import { useApi } from '../composables/useApi.js'
import { useAsyncAction } from '../composables/useAsyncAction.js'
import { toast } from '../utils/toast.js'
// Backs the creator/membership review queue (#388 E4). Confirm-only: accepting
// ADDS A SOURCE to the artist that already has the other channel — it never
// merges two artists. Adding a source is trivially undone; a wrong merge
// silently mixes two creators' work with nothing left to separate them by.
export const useMembershipSuggestionsStore = defineStore('membershipSuggestions', () => {
const api = useApi()
const suggestions = ref([])
const { loading, error, run } = useAsyncAction({ errorAs: 'message' })
async function load () {
await run(async () => {
const body = await api.get('/api/sources/membership-suggestions')
suggestions.value = body.items || []
})
}
async function accept (id) {
try {
const res = await api.post(`/api/sources/membership-suggestions/${id}/accept`, {})
suggestions.value = suggestions.value.filter(s => s.id !== id)
toast({
text: res.already_linked ? 'Already linked' : 'Channel added to this creator',
type: 'success'
})
} catch (e) {
toast({ text: `Link failed: ${e.message}`, type: 'error' })
}
}
async function dismiss (id) {
try {
await api.post(`/api/sources/membership-suggestions/${id}/dismiss`, {})
suggestions.value = suggestions.value.filter(s => s.id !== id)
} catch (e) {
toast({ text: `Dismiss failed: ${e.message}`, type: 'error' })
}
}
async function rescan () {
await api.post('/api/sources/membership-suggestions/rescan', {})
await load()
}
return { suggestions, loading, error, load, accept, dismiss, rescan }
})
+38
View File
@@ -0,0 +1,38 @@
import { defineStore } from 'pinia'
import { ref } from 'vue'
import { useApi } from '../composables/useApi.js'
import { useAsyncAction } from '../composables/useAsyncAction.js'
import { toast } from '../utils/toast.js'
// Backs the subscription-roster card (#387 C3). The whole reason this surface
// exists: a roster that never synced, or failed to, must be DISTINGUISHABLE
// from an account that subscribes to nothing. All three states look like zero
// rows, and only one of them means "you are tracking things you do not pay
// for" — so the UI must never render "never synced" as a number.
export const useMembershipSyncStore = defineStore('membershipSync', () => {
const api = useApi()
const platforms = ref([])
const { loading, error, run } = useAsyncAction({ errorAs: 'message' })
async function load () {
await run(async () => {
const body = await api.get('/api/sources/membership-sync')
platforms.value = body.platforms || []
})
}
async function syncNow () {
try {
await api.post('/api/sources/membership-sync', {})
toast({
text: 'Roster sync queued — runs in the background; reload to see the result.',
type: 'info'
})
} catch (e) {
toast({ text: `Sync failed to queue: ${e.message}`, type: 'error' })
}
}
return { platforms, loading, error, load, syncNow }
})
+83
View File
@@ -0,0 +1,83 @@
import { defineStore } from 'pinia'
import { ref } from 'vue'
import { useApi } from '../composables/useApi.js'
import { useAsyncAction } from '../composables/useAsyncAction.js'
import { toast } from '../utils/toast.js'
// Backs the announcement review queue (#388 E5): "this Patreon post announced
// that Discord drop". Confirm-only, deliberately — a wrongly-asserted link
// tells the operator two different pieces are one, which is worse than no link
// at all, so accept is the ONLY thing that makes a link real. Mirrors
// seriesSuggestions (FC-6.3), which is the same shape for the same reason.
export const usePostAssociationsStore = defineStore('postAssociations', () => {
const api = useApi()
const proposals = ref([])
const enabled = ref(true)
const threshold = ref(0.6)
const windowHours = ref(24)
const { loading, error, run } = useAsyncAction({ errorAs: 'message' })
async function load () {
await run(async () => {
const body = await api.get('/api/posts/associations')
proposals.value = body.items || []
})
}
async function loadSettings () {
const s = await api.get('/api/settings/import')
enabled.value = s.discord_link_enabled
threshold.value = s.discord_link_threshold
windowHours.value = s.discord_link_window_hours
}
async function saveSettings (patch) {
await api.patch('/api/settings/import', { body: patch })
}
async function setEnabled (v) {
enabled.value = v
await saveSettings({ discord_link_enabled: v })
}
async function setThreshold (v) {
threshold.value = v
await saveSettings({ discord_link_threshold: v })
}
async function setWindowHours (v) {
windowHours.value = v
await saveSettings({ discord_link_window_hours: v })
}
async function accept (id) {
try {
await api.post(`/api/posts/associations/${id}/accept`, {})
proposals.value = proposals.value.filter(p => p.id !== id)
toast({ text: 'Linked', type: 'success' })
} catch (e) {
toast({ text: `Link failed: ${e.message}`, type: 'error' })
}
}
async function dismiss (id) {
try {
await api.post(`/api/posts/associations/${id}/dismiss`, {})
proposals.value = proposals.value.filter(p => p.id !== id)
} catch (e) {
toast({ text: `Dismiss failed: ${e.message}`, type: 'error' })
}
}
async function rescan () {
await api.post('/api/posts/associations/rescan', {})
await load()
}
return {
proposals, enabled, threshold, windowHours, loading, error,
load, loadSettings, setEnabled, setThreshold, setWindowHours,
accept, dismiss, rescan
}
})
+6 -1
View File
@@ -4,7 +4,12 @@ import { useApi } from '../composables/useApi.js'
export const useSystemStore = defineStore('system', () => {
const api = useApi()
const healthy = ref(null) // null=unknown, true=ok, false=down
// NOT what the nav dot reads any more (milestone 365): that is the
// whole-stack verdict in systemHealth.js. /api/health only proves the web
// container is serving, which is why a green dot here sat happily beside a
// dead worker. refreshHealth() is still called — it is also how build/version
// info arrives — so this stays as its by-product rather than its purpose.
const healthy = ref(null)
// What the instance says it is. Since milestone 318 stopped publishing
// version image tags, this is the only answer to "which build is this?" —
// there is no registry name left to check it against.
+49
View File
@@ -0,0 +1,49 @@
import { defineStore } from 'pinia'
import { computed, ref } from 'vue'
import { useApi } from '../composables/useApi.js'
// Whole-stack health: is every part of FabledCurator running (milestone 365)?
//
// Distinct from `system.js`, which polls /api/health — a no-DB liveness check
// that only proves the web container is serving. That endpoint answers "can I
// reach the API"; this one answers "is anything broken", which is the question
// a green dot beside the brand was already being read as answering.
//
// Also distinct from `systemActivity.js`, which is about what the pipeline is
// DOING — queue depths, running tasks, failures. Running and alive are
// different questions and they fail independently: a perfectly idle stack with
// a dead worker looks identical to a healthy one on the activity surfaces.
export const useSystemHealthStore = defineStore('systemHealth', () => {
const api = useApi()
const overall = ref(null) // null until the first answer: unknown ≠ ok
const parts = ref([])
const checkedAt = ref(null)
const thresholds = ref(null) // server-owned, so the UI keeps no second copy
const lastError = ref(null)
async function refresh() {
try {
const body = await api.get('/api/system/health')
overall.value = body.overall
parts.value = body.parts || []
checkedAt.value = body.checked_at
thresholds.value = body.thresholds || null
lastError.value = null
} catch (e) {
// The endpoint is built never to fail because a dependency failed, so a
// throw here means the API itself is unreachable — which is its own kind
// of unhealthy and must not be shown as "ok".
lastError.value = e.message
overall.value = 'unknown'
}
return overall.value
}
// The parts worth naming in a tooltip — everything that is not ok, worst
// first. The endpoint already sorts that way.
const problems = computed(() => parts.value.filter(p => p.state !== 'ok'))
return { overall, parts, checkedAt, thresholds, lastError, problems, refresh }
})
+30 -5
View File
@@ -3,9 +3,14 @@
<!-- In-context view: deep-linked to one post, with bidirectional infinite
scroll newer posts load above, older posts below. -->
<template v-if="postIdFilter != null">
<!-- Returns to whichever surface you deep-linked FROM. This used to be a
hard `{ name: 'posts' }`, which redirects into Browse fine while
this view only ever rendered inside Browse's tab, but it now also
serves the front door (#387 B1), where it would have yanked the
operator sideways into a different view. -->
<v-btn
variant="text" size="small" prepend-icon="mdi-arrow-left"
:to="{ name: 'posts' }" class="mb-2"
:to="allPostsTarget" class="mb-2"
>All posts</v-btn>
<v-alert v-if="store.error" type="error" variant="tonal" closable class="mb-3">
@@ -44,6 +49,8 @@
<!-- Normal feed -->
<template v-else>
<FeedStatusRibbon v-if="statusRibbon" />
<PostsFilterBar
:artist-id="artistFilter"
:platform="platformFilter"
@@ -59,11 +66,11 @@
</div>
<div v-else-if="store.items.length === 0 && store.done" class="fc-posts__empty">
<!-- A filtered miss is NOT the onboarding case: the operator has posts,
they just narrowed past them. Showing a fresh-install on-ramp here
would be telling someone with a full library to go and set it up. -->
<p v-if="hasActiveFilter">No posts match your search or filters.</p>
<p v-else>No posts yet. Subscribe to a source on the
<RouterLink to="/subscriptions">Subscriptions</RouterLink>
tab to start capturing posts.
</p>
<FeedEmptyState v-else />
</div>
<div v-else>
@@ -84,6 +91,17 @@ import { useRoute, useRouter } from 'vue-router'
import { usePostsStore } from '../stores/posts.js'
import PostsFilterBar from '../components/posts/PostsFilterBar.vue'
import PostCard from '../components/posts/PostCard.vue'
import FeedStatusRibbon from '../components/posts/FeedStatusRibbon.vue'
import FeedEmptyState from '../components/posts/FeedEmptyState.vue'
// The ingestion-status ribbon is a FRONT-DOOR concern, not a feed concern —
// inside Browse's Posts tab you are looking FOR something, and the Subscriptions
// hub is a click away. Passed as a route prop rather than sniffed from
// route.name so the view does not have to know what it is mounted as, and so
// the router file states the intent in one place.
defineProps({
statusRibbon: { type: Boolean, default: false },
})
const route = useRoute()
const router = useRouter()
@@ -103,6 +121,13 @@ const hasActiveFilter = computed(() =>
artistFilter.value != null || platformFilter.value != null || searchFilter.value != null
)
// Drop only `post_id` and stay where we are — keeps Browse's `tab=posts` (and
// any active artist/platform scope) intact instead of resetting the surface.
const allPostsTarget = computed(() => {
const { post_id: _drop, ...rest } = route.query
return { name: route.name, query: rest }
})
// --- normal feed (downward infinite scroll) ---
const sentinel = ref(null)
let observer = null
+16 -2
View File
@@ -14,6 +14,7 @@
style="position: sticky; top: var(--fc-nav-h, 64px); z-index: 4;"
>
<v-tab value="overview">Overview</v-tab>
<v-tab value="system">System</v-tab>
<v-tab value="activity">Activity</v-tab>
<v-tab value="cleanup">Cleanup</v-tab>
<v-tab value="maintenance">Maintenance</v-tab>
@@ -42,6 +43,13 @@
</v-alert>
</v-window-item>
<!-- Is every part of the stack still running (milestone 365). Sits
beside Activity deliberately: Activity answers "what is the queue
doing", this answers "is anything left to do it". -->
<v-window-item value="system">
<SystemHealthTab />
</v-window-item>
<v-window-item value="activity">
<SystemActivityTab @open-maintenance="tab = 'maintenance'" />
</v-window-item>
@@ -73,18 +81,24 @@
</template>
<script setup>
import { onMounted, onUnmounted, ref, watch } from 'vue'
import { onMounted, onUnmounted, watch } from 'vue'
import { useSystemStore } from '../stores/system.js'
import SystemStatsCards from '../components/settings/SystemStatsCards.vue'
import SystemActivitySummary from '../components/settings/SystemActivitySummary.vue'
import SystemActivityTab from '../components/settings/SystemActivityTab.vue'
import SystemHealthTab from '../components/settings/SystemHealthTab.vue'
import GpuActivityPanel from '../components/settings/GpuActivityPanel.vue'
import DownloadsActivityPanel from '../components/settings/DownloadsActivityPanel.vue'
import MaintenancePanel from '../components/settings/MaintenancePanel.vue'
import CleanupView from './CleanupView.vue'
import { useTabQuery } from '../composables/useTabQuery.js'
import { useMLStore } from '../stores/ml.js'
const tab = ref('overview')
// ?tab= sync (the same composable Browse/Subscriptions use) so a tab can be
// linked TO — the health dot beside the brand points at ?tab=system, and the
// old /system path redirects there.
const VALID_TABS = ['overview', 'system', 'activity', 'cleanup', 'maintenance']
const { tab } = useTabQuery(VALID_TABS, 'overview')
const system = useSystemStore()
const mlStore = useMLStore()
@@ -20,6 +20,35 @@ describe('ActiveDownloadsPanel', () => {
w.unmount() // clear the 1s elapsed-timer interval
})
// #387 A1: the live payload gained a `gated` count. Shown only when non-zero
// so a healthy run stays uncluttered — a "🔒 0" on every download would be
// noise, and noise is what stops the number being noticed when it matters.
it('ticks the tier-gated count mid-walk when there is one', () => {
const pinia = freshPinia()
useDownloadsStore().activeEvents = [{
id: 1, status: 'running',
started_at: new Date(Date.now() - 65000).toISOString(),
platform: 'patreon', artist_name: 'Alice',
live: { downloaded: 2, skipped: 0, errors: 0, posts: 12, gated: 9 },
}]
const w = mountComponent(ActiveDownloadsPanel, { pinia })
expect(w.text()).toContain('9')
w.unmount()
})
it('omits the gated count when nothing was gated', () => {
const pinia = freshPinia()
useDownloadsStore().activeEvents = [{
id: 1, status: 'running',
started_at: new Date(Date.now() - 65000).toISOString(),
platform: 'patreon', artist_name: 'Alice',
live: { downloaded: 2, skipped: 0, errors: 0, posts: 12, gated: 0 },
}]
const w = mountComponent(ActiveDownloadsPanel, { pinia })
expect(w.find('.fc-active__count--gated').exists()).toBe(false)
w.unmount()
})
it('renders nothing when there is no active work', () => {
const pinia = freshPinia()
useDownloadsStore().activeEvents = []
@@ -0,0 +1,216 @@
// @vitest-environment happy-dom
import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest'
import FeedEmptyState from '../../src/components/posts/FeedEmptyState.vue'
import { useCredentialsStore } from '../../src/stores/credentials.js'
import { useMembershipReconcileStore } from '../../src/stores/membershipReconcile.js'
import { useMembershipSyncStore } from '../../src/stores/membershipSync.js'
import { useSourcesStore } from '../../src/stores/sources.js'
import { mountWithStore } from '../support/mountComponent.js'
// #387 B4. This is the first screen a fresh install shows anyone, so the thing
// worth pinning is that it tells the two empties apart. Telling an operator to
// "add a source" when they already have three and are mid-backfill is worse
// than saying nothing — it reads as the app not knowing its own state.
// The second argument is C6's world. Omitting it means "no credential", which
// is both the fresh-install default and the rung every B4 case above asserts —
// so those keep passing unchanged rather than needing the new vocabulary.
const mountWith = (status, { credentials = [], sync = [], reconcile = [] } = {}) =>
mountWithStore(FeedEmptyState, () => {
useSourcesStore().scheduleStatus = status
useCredentialsStore().byPlatform = new Map(
credentials.map((name) => [name, { platform: name }]),
)
useMembershipSyncStore().platforms = sync
useMembershipReconcileStore().platforms = reconcile
})
const EMPTY = { total_sources: 0 }
const SYNCED = [{ platform: 'patreon', last_success_at: '2026-09-12T00:00:00Z' }]
const NEVER = [{ platform: 'patreon', last_success_at: null }]
// Every rung's distinguishing phrase, so a test can assert that exactly ONE is
// on screen. Each is chosen to sit on a SINGLE template line: a phrase spanning
// a line break would never match once the renderer keeps the newline, and a
// mutual-exclusion check whose phrases never match passes vacuously — coverage
// in appearance only (rule #167). Both sides are whitespace-normalised anyway,
// so indentation changes cannot quietly break it either.
const RUNG_PHRASES = {
credential: 'Add a credential',
discover: 'Find what you already subscribe to',
adopt: 'not being followed here yet',
unavailable: "couldn't reach",
manual: 'Everything you subscribe to on',
}
const flat = (s) => s.replace(/\s+/g, ' ').trim()
function rungsShown (w) {
const text = flat(w.text())
return Object.entries(RUNG_PHRASES)
.filter(([, phrase]) => text.includes(flat(phrase)))
.map(([name]) => name)
}
describe('FeedEmptyState', () => {
beforeEach(() => {
globalThis.fetch = vi.fn(async () => { throw new Error('offline') })
})
afterEach(() => vi.restoreAllMocks())
it('a fresh install gets the on-ramp, credential first', () => {
const w = mountWith({ total_sources: 0 })
expect(w.text()).toContain('Add a credential')
expect(w.text()).toContain('Add a source')
expect(w.text()).not.toContain("See what's running")
})
it('an install that is already fetching is told so, not told to set up', () => {
const w = mountWith({ total_sources: 3 })
expect(w.text()).toContain('3 sources')
expect(w.text()).toContain("See what's running")
expect(w.text()).not.toContain('Add a credential')
})
it('singularises a lone source', () => {
const w = mountWith({ total_sources: 1 })
expect(w.text()).toContain('1 source')
expect(w.text()).not.toContain('1 sources')
})
it('falls back to the on-ramp when the status call never answered', () => {
// Safe direction: the on-ramp is merely redundant to an established
// operator, whereas "see what's running" shown to someone with nothing
// configured is a dead end.
const w = mountWith(null)
expect(w.text()).toContain('Add a credential')
})
it('shows the brand mark, sized to be readable', () => {
// logo.svg stops reading below ~48px — that is why the 22px nav slot has a
// different mark. If this ever shrinks to a glyph, it is the wrong asset.
const img = mountWith({ total_sources: 0 }).find('img.fc-empty__mark')
expect(img.exists()).toBe(true)
expect(img.attributes('src')).toBe('/logo.svg')
expect(Number(img.attributes('width'))).toBeGreaterThanOrEqual(48)
// Decorative: the surrounding prose already carries the meaning.
expect(img.attributes('alt')).toBe('')
})
// --- #387 C6: the discovery rung -----------------------------------------
//
// The loop this closes: a new installer has already told Patreon which
// creators they follow, and making them retype that list is the friction the
// whole milestone is about. These pin that the rung shown matches what is
// actually POSSIBLE — offering discovery before a credential exists, or
// after it has proven unreachable, is a button that cannot work.
it('offers discovery once a credential exists and nothing has synced', () => {
const w = mountWith(EMPTY, { credentials: ['patreon'], sync: NEVER })
expect(w.text()).toContain('Find what you already subscribe to')
expect(w.text()).toContain('patreon')
// Nothing is added for them — the offer/never-auto-add line from C4.
expect(w.text()).toContain('Nothing is added automatically')
})
it('never offers discovery before a credential exists', () => {
// The button would 404 against an unconnected platform. Absence beats a
// dead control on the first screen anyone sees.
const w = mountWith(EMPTY, { sync: NEVER })
expect(w.text()).not.toContain('Find what you already subscribe to')
expect(w.text()).toContain('Add a credential')
})
it('a sweep that has never worked reads as unavailable, not as a retry', () => {
// Rule #164: an install with no outbound network must reach this screen and
// be told plainly. Not a spinner, not a crash, not a button that will fail
// the same way — and the manual path stays open.
const w = mountWith(EMPTY, {
credentials: ['patreon'],
sync: [{ platform: 'patreon', last_success_at: null, last_error_type: 'ConnectionError' }],
})
expect(w.text()).toContain("couldn't reach")
expect(w.text()).toContain('ConnectionError')
expect(w.text()).toContain('Add a source')
expect(w.text()).not.toContain('Find what you already subscribe to')
})
it('a roster that synced once and failed since is NOT unavailable', () => {
// It has a roster, just an ageing one — C4's freshness gate handles that.
// Collapsing the two would hide a usable roster behind an error banner.
const w = mountWith(EMPTY, {
credentials: ['patreon'],
sync: [{
platform: 'patreon',
last_success_at: '2026-09-12T00:00:00Z',
last_error_type: 'PatreonAuthError',
}],
reconcile: [{ platform: 'patreon', subscribed_not_tracked: [{ id: 1 }] }],
})
expect(w.text()).not.toContain("couldn't reach")
expect(w.text()).toContain('not being followed here yet')
})
it('counts what there is to adopt, and sends them to the picker', () => {
const w = mountWith(EMPTY, {
credentials: ['patreon'],
sync: SYNCED,
reconcile: [{ platform: 'patreon', subscribed_not_tracked: [{ id: 1 }, { id: 2 }] }],
})
expect(w.text()).toContain('2 creators you subscribe to are')
expect(w.text()).toContain('Choose who to follow')
})
it('singularises a lone unmatched creator', () => {
const w = mountWith(EMPTY, {
credentials: ['patreon'],
sync: SYNCED,
reconcile: [{ platform: 'patreon', subscribed_not_tracked: [{ id: 1 }] }],
})
expect(w.text()).toContain('1 creator you subscribe to is')
expect(w.text()).not.toContain('1 creators')
})
it('a fully-tracked roster stops offering discovery', () => {
// Nothing left to find, so the manual path is the honest last rung rather
// than a discovery button that would return an empty list.
const w = mountWith(EMPTY, {
credentials: ['patreon'],
sync: SYNCED,
reconcile: [{ platform: 'patreon', subscribed_not_tracked: [] }],
})
expect(w.text()).toContain('Everything you subscribe to on')
expect(w.text()).toContain('Add a source')
expect(w.text()).not.toContain('Find what you already subscribe to')
expect(w.text()).not.toContain('not being followed here yet')
})
it('shows exactly one rung in every state', () => {
// The component claims the rungs are mutually exclusive by construction.
// This is that claim, asserted — two on screen at once would be worst
// exactly here, on the first screen a new installer ever sees.
const states = [
[EMPTY, {}],
[EMPTY, { credentials: ['patreon'], sync: NEVER }],
[EMPTY, { credentials: ['patreon'], sync: [{ platform: 'patreon', last_success_at: null, last_error_type: 'ConnectionError' }] }],
[EMPTY, { credentials: ['patreon'], sync: SYNCED, reconcile: [{ platform: 'patreon', subscribed_not_tracked: [{ id: 1 }] }] }],
[EMPTY, { credentials: ['patreon'], sync: SYNCED, reconcile: [{ platform: 'patreon', subscribed_not_tracked: [] }] }],
]
for (const [status, world] of states) {
expect(rungsShown(mountWith(status, world))).toHaveLength(1)
}
})
it('an install that is already fetching never sees an on-ramp rung', () => {
// hasSources wins over everything C6 added — someone mid-backfill is not
// onboarding, whatever their roster says.
const w = mountWith({ total_sources: 3 }, {
credentials: ['patreon'],
sync: SYNCED,
reconcile: [{ platform: 'patreon', subscribed_not_tracked: [{ id: 1 }] }],
})
expect(w.text()).toContain("See what's running")
expect(rungsShown(w)).toHaveLength(0)
})
})
@@ -0,0 +1,74 @@
// @vitest-environment happy-dom
import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest'
import FeedStatusRibbon from '../../src/components/posts/FeedStatusRibbon.vue'
import { useSourcesStore } from '../../src/stores/sources.js'
import { mountWithStore } from '../support/mountComponent.js'
// #387 B3. This ribbon is the ONLY place phase A's work reaches someone who
// wasn't already looking for it, so what it does and does not say is the whole
// feature. These pin: it stays silent when it has nothing true to report, it
// keeps "failing" and "no access" as separate claims, and it never renders a
// zero — a "0 failing" on the front door is noise that trains you to ignore
// the line, which is exactly what would hide the real number later.
const mountWith = (status) => mountWithStore(FeedStatusRibbon, () => {
useSourcesStore().scheduleStatus = status
})
describe('FeedStatusRibbon', () => {
beforeEach(() => {
// The component fetches on mount and swallows failures by design; stub it
// so the assertions are about the seeded state, not a race with the call.
globalThis.fetch = vi.fn(async () => { throw new Error('offline') })
})
afterEach(() => vi.restoreAllMocks())
it('renders nothing before it has a status', () => {
const w = mountWith(null)
expect(w.find('.fc-ribbon').exists()).toBe(false)
})
it('a healthy instance is one quiet line, with no zeroes', () => {
const w = mountWith({
last_tick_at: new Date().toISOString(), failing_sources: 0, no_access_sources: 0,
})
expect(w.find('.fc-ribbon').exists()).toBe(true)
expect(w.text()).toContain('Checked')
// Assert on the claims, not on the digit — a just-now timestamp can render
// its own "0 minutes ago" and a substring check would catch that instead.
expect(w.text()).not.toContain('failing')
expect(w.text()).not.toContain("can't see")
expect(w.find('.fc-ribbon__item--err').exists()).toBe(false)
expect(w.find('.fc-ribbon__item--gated').exists()).toBe(false)
})
it('reports failing and no-access as two separate claims', () => {
const w = mountWith({
last_tick_at: new Date().toISOString(), failing_sources: 2, no_access_sources: 5,
})
expect(w.find('.fc-ribbon__item--err').text()).toContain('2 sources are failing')
expect(w.find('.fc-ribbon__item--gated').text()).toContain("5 you can't see")
})
it('no-access alone does not light the failing item', () => {
// The whole point of phase A: a paywalled creator is not a broken one.
const w = mountWith({
last_tick_at: new Date().toISOString(), failing_sources: 0, no_access_sources: 3,
})
expect(w.find('.fc-ribbon__item--err').exists()).toBe(false)
expect(w.find('.fc-ribbon__item--gated').exists()).toBe(true)
})
it('says "never" rather than a blank when nothing has ever run', () => {
const w = mountWith({ last_tick_at: null, failing_sources: 0, no_access_sources: 0 })
expect(w.text()).toContain('never')
})
it('singularises one failing source', () => {
const w = mountWith({
last_tick_at: new Date().toISOString(), failing_sources: 1, no_access_sources: 0,
})
expect(w.find('.fc-ribbon__item--err').text()).toContain('1 source is failing')
})
})
@@ -0,0 +1,77 @@
// @vitest-environment happy-dom
import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest'
import MembershipRosterCard from '../../src/components/settings/MembershipRosterCard.vue'
import { useMembershipSyncStore } from '../../src/stores/membershipSync.js'
import { mountWithStore } from '../support/mountComponent.js'
// #387 C3. The one thing this card must never do is render "never synced" as a
// number. Three situations produce zero memberships — the account subscribes to
// nothing, the sweep never ran, the sweep failed — and only the first means
// "you are tracking sources you do not pay for". Conflating them is how C4 ends
// up telling the operator to cancel things they are actively paying for.
const mountWith = (platforms) => mountWithStore(MembershipRosterCard, () => {
useMembershipSyncStore().platforms = platforms
})
describe('MembershipRosterCard', () => {
beforeEach(() => {
// The card loads on mount; stub it so assertions are about seeded state.
globalThis.fetch = vi.fn(async () => { throw new Error('offline') })
})
afterEach(() => vi.restoreAllMocks())
it('says "never synced" in words, and states no count', () => {
const w = mountWith([
{ platform: 'patreon', last_success_at: null, last_attempt_at: null,
last_count: null, fresh: false },
])
expect(w.text()).toContain('never synced')
// The count line must be absent entirely — not "0 memberships found".
expect(w.text()).not.toContain('memberships found')
})
it('does not claim a count when a sync has been tried but never succeeded', () => {
const w = mountWith([
{ platform: 'patreon', last_success_at: null,
last_attempt_at: new Date(Date.now() - 3600e3).toISOString(),
last_count: null, fresh: false, last_error_type: 'PatreonAuthError',
last_error_message: 'cookies expired' },
])
expect(w.text()).toContain('no successful sync yet')
expect(w.text()).not.toContain('memberships found')
// The error is surfaced, because "rotate your credential" is actionable.
expect(w.text()).toContain('PatreonAuthError')
})
it('marks a stale roster as too old to rely on', () => {
const w = mountWith([
{ platform: 'patreon',
last_success_at: new Date(Date.now() - 9 * 86400e3).toISOString(),
last_count: 12, fresh: false },
])
expect(w.text()).toContain('too old to rely on')
})
it('states the count only when a successful sync stands behind it', () => {
const w = mountWith([
{ platform: 'patreon',
last_success_at: new Date(Date.now() - 3600e3).toISOString(),
last_count: 12, fresh: true },
])
expect(w.text()).toContain('12 memberships found')
expect(w.text()).not.toContain('never synced')
expect(w.text()).not.toContain('too old')
})
it('a real zero is reported as zero, because that one IS an answer', () => {
const w = mountWith([
{ platform: 'patreon',
last_success_at: new Date(Date.now() - 3600e3).toISOString(),
last_count: 0, fresh: true },
])
expect(w.text()).toContain('0 memberships found')
expect(w.text()).not.toContain('never synced')
})
})
+118
View File
@@ -40,6 +40,124 @@ describe('PostCard', () => {
expect(full.text()).not.toContain('Show more')
})
// #388 E2. The honesty marker is the only thing standing between "FC
// assembled this" and the card reading as something the artist authored, so
// these pin BOTH directions: it appears when the flag is set, and — the one
// that actually matters — it is absent on every ordinary post.
describe('synthetic posts', () => {
const SYNTH = {
...BASE,
post_title: null,
synthesized_by: 'discord_drop',
synthesis: { message_count: 4, member_post_ids: [1, 2, 3, 4] },
}
it('says FC grouped it, and what from', () => {
const w = mountComponent(PostCard, { props: { post: SYNTH }, pinia: freshPinia() })
expect(w.find('.fc-post-card__synthetic').exists()).toBe(true)
expect(w.text()).toContain('grouped by FabledCurator')
expect(w.text()).toContain('Grouped from 4 Discord messages')
})
it('never marks an ordinary post', () => {
const w = mountComponent(PostCard, { props: { post: BASE }, pinia: freshPinia() })
expect(w.find('.fc-post-card__synthetic').exists()).toBe(false)
expect(w.text()).not.toContain('grouped by FabledCurator')
})
it('does not leak the internal drop key as a title', () => {
// post_title is NULL on purpose (inventing one would put words in a
// creator's mouth), so the untitled fallback must not print
// `fc-drop:<message id>` the way it does for a real untitled post.
const w = mountComponent(PostCard, {
props: { post: { ...SYNTH, external_post_id: 'fc-drop:99887766' } },
pinia: freshPinia(),
})
expect(w.text()).not.toContain('fc-drop:')
})
it('degrades to unmarked when the field is absent entirely', () => {
// A post dict composed before the field existed must read as "not
// synthetic" rather than throw on `synthesis.message_count`.
const { synthesized_by: _drop, synthesis: _also, ...legacy } = SYNTH
const w = mountComponent(PostCard, { props: { post: legacy }, pinia: freshPinia() })
expect(w.find('.fc-post-card__synthetic').exists()).toBe(false)
})
it('reports growth separately from the post date', () => {
// The drop's own date is its identity; growth is news about it. "From
// Tuesday, gained images this morning" is the trickling-in signal, and
// collapsing the two would erase it.
const w = mountComponent(PostCard, {
props: {
post: { ...SYNTH, last_grew_at: new Date(Date.now() - 3600e3).toISOString() },
},
pinia: freshPinia(),
})
expect(w.text()).toContain('updated 1h ago')
})
it('says nothing about growth on a grouping that has not grown', () => {
// An "updated" label that is always there teaches you to ignore it.
const w = mountComponent(PostCard, {
props: { post: { ...SYNTH, last_grew_at: null } },
pinia: freshPinia(),
})
expect(w.text()).not.toContain('updated')
})
it('never claims an ordinary post grew, even if the field leaks in', () => {
const w = mountComponent(PostCard, {
props: { post: { ...BASE, last_grew_at: new Date().toISOString() } },
pinia: freshPinia(),
})
expect(w.text()).not.toContain('updated')
})
it('shows an accepted announcement link from either end', () => {
const teaser = mountComponent(PostCard, {
props: {
post: {
...BASE,
associations: [{ id: 7, role: 'announces', post_id: 42 }],
},
},
pinia: freshPinia(),
})
expect(teaser.text()).toContain('The full set is in Discord')
const drop = mountComponent(PostCard, {
props: {
post: {
...SYNTH,
associations: [{ id: 7, role: 'announced_by', post_id: 1 }],
},
},
pinia: freshPinia(),
})
expect(drop.text()).toContain('Announced on Patreon')
})
it('shows no link when there is no accepted association', () => {
// Pending proposals never reach the payload, so an empty list here is
// the normal case and must render as silence, not an empty row.
const w = mountComponent(PostCard, {
props: { post: { ...BASE, associations: [] } },
pinia: freshPinia(),
})
expect(w.find('.fc-post-card__assoc').exists()).toBe(false)
})
it('singularises a one-message drop', () => {
const w = mountComponent(PostCard, {
props: { post: { ...SYNTH, synthesis: { message_count: 1 } } },
pinia: freshPinia(),
})
expect(w.text()).toContain('Grouped from 1 Discord message')
expect(w.text()).not.toContain('1 Discord messages')
})
})
const thumbs = (n) =>
Array.from({ length: n }, (_, i) => ({ image_id: 100 + i, thumbnail_url: `/t${i}` }))
@@ -0,0 +1,141 @@
// @vitest-environment happy-dom
import { describe, it, expect } from 'vitest'
import SourceHealthDot from '../../src/components/subscriptions/SourceHealthDot.vue'
import { VTooltipStub, mountComponent } from '../support/mountComponent.js'
// Milestone #387 A3. The dot is the only always-visible signal per source, so
// the grade it picks IS the claim FC makes about that subscription. These pin
// that no-access is graded as its own thing — not as healthy (which hides it)
// and not as a failure (which would send the operator hunting for a break that
// isn't there).
const checked = { last_checked_at: '2026-09-09T12:00:00+00:00' }
function dotClass (w) {
return w.find('.fc-health-dot').classes().join(' ')
}
describe('SourceHealthDot', () => {
it('grades a tier-gated source as no-access, not healthy', () => {
const w = mountComponent(SourceHealthDot, {
stubs: { VTooltip: VTooltipStub },
props: {
source: { ...checked, consecutive_failures: 0, error_type: 'tier_limited' },
},
})
expect(dotClass(w)).toContain('fc-health-dot--no-access')
expect(dotClass(w)).not.toContain('fc-health-dot--healthy')
})
it('shows the gated count when the list endpoint supplied one', () => {
const w = mountComponent(SourceHealthDot, {
stubs: { VTooltip: VTooltipStub },
props: {
source: {
...checked, consecutive_failures: 0,
error_type: 'tier_limited', tier_gated_count: 47,
},
},
})
expect(w.text()).toContain('47 posts')
})
it('states the condition without a number when no count was joined in', () => {
const w = mountComponent(SourceHealthDot, {
stubs: { VTooltip: VTooltipStub },
props: {
source: {
...checked, consecutive_failures: 0,
error_type: 'tier_limited', tier_gated_count: null,
},
},
})
// Absent is not zero: never render "0 posts you don't have access to".
expect(w.text()).not.toContain('0 post')
expect(w.text()).toContain("tier you don't hold")
})
// Milestone #387 C5. The count says WHAT; these say WHY — a much stronger
// claim, so the cases that must stay silent are pinned alongside the ones
// that speak.
function gated (extra) {
return mountComponent(SourceHealthDot, {
stubs: { VTooltip: VTooltipStub },
props: {
source: {
...checked, consecutive_failures: 0,
error_type: 'tier_limited', tier_gated_count: 47, ...extra,
},
},
})
}
it('a lapsed membership says the subscription ended', () => {
expect(gated({ gated_reason: 'lapsed' }).text()).toContain('not a patron any more')
})
it('an active membership says the tier is the limit, not the subscription', () => {
const text = gated({ gated_reason: 'tier' }).text()
expect(text).toContain("tier doesn't cover these posts")
expect(text).not.toContain('not a patron any more')
})
it('a free follow is not described as a lapsed subscription', () => {
const text = gated({ gated_reason: 'free' }).text()
expect(text).toContain('follow this creator for free')
expect(text).not.toContain('not a patron any more')
})
it('no reason means the count stands alone, with no invented explanation', () => {
// null covers all four not-evidence cases the backend collapses into it:
// campaign absent, roster stale, never swept, status uncharacterised.
const text = gated({ gated_reason: null }).text()
expect(text).toContain('47 posts')
expect(text).not.toContain('patron')
expect(text).not.toContain('tier doesn')
})
it('a reason the frontend has not been taught renders nothing', () => {
// The backend's status vocabulary grows per platform (D1). An unknown word
// must degrade to the bare count, never to `undefined` in the tooltip.
const text = gated({ gated_reason: 'some_future_word' }).text()
expect(text).toContain('47 posts')
expect(text).not.toContain('undefined')
})
it('a reason never appears on a source that is not gated', () => {
const w = mountComponent(SourceHealthDot, {
stubs: { VTooltip: VTooltipStub },
props: {
source: {
...checked, consecutive_failures: 0,
error_type: null, gated_reason: 'lapsed',
},
},
})
// The roster annotates a gated state; it never asserts one on its own.
expect(w.text()).not.toContain('not a patron any more')
})
it('a genuinely failing source still grades as a failure, gated or not', () => {
const w = mountComponent(SourceHealthDot, {
stubs: { VTooltip: VTooltipStub },
props: {
source: { ...checked, consecutive_failures: 9, error_type: 'tier_limited' },
warningThreshold: 5,
},
})
expect(dotClass(w)).toContain('fc-health-dot--critical')
expect(dotClass(w)).not.toContain('fc-health-dot--no-access')
})
it('an unchecked source is still unchecked', () => {
const w = mountComponent(SourceHealthDot, {
stubs: { VTooltip: VTooltipStub },
props: { source: { last_checked_at: null, consecutive_failures: 0 } },
})
expect(dotClass(w)).toContain('fc-health-dot--unchecked')
})
})
+31 -2
View File
@@ -2,8 +2,29 @@ import { describe, it, expect } from 'vitest'
import router, { FRONT_DOOR } from '../src/router.js'
describe('router', () => {
it('FRONT_DOOR defaults to /showcase', () => {
expect(FRONT_DOOR).toBe('/showcase')
it('FRONT_DOOR is the post feed', () => {
// Moved from /showcase in #387 B1. Asserted on the constant rather than on
// a navigation so the intent is pinned: the door is the feed, and moving it
// again should be a deliberate edit to a failing test, not a quiet change.
expect(FRONT_DOOR).toBe('/latest')
})
it('/latest is the feed, mounted outside the Browse tab strip', () => {
const r = router.resolve('/latest')
expect(r.name).toBe('latest')
// A nav entry in its own right (meta.title is what TopNav lists on), and
// first in the order.
expect(r.meta.title).toBe('Latest')
expect(r.meta.navOrder).toBe(5)
// Deliberately NOT stickyChrome — it has no sticky sub-header for the nav
// to butt against, unlike Browse/Gallery/Settings.
expect(r.meta.stickyChrome).toBeUndefined()
})
it('showcase is demoted, not removed — still reachable and still in the nav', () => {
const r = router.resolve('/showcase')
expect(r.name).toBe('showcase')
expect(r.meta.title).toBe('Showcase')
})
it('/ redirects to FRONT_DOOR', async () => {
@@ -40,6 +61,14 @@ describe('router', () => {
expect(router.currentRoute.value.query.post_id).toBe('7')
})
it('/system redirects into the Settings System tab', async () => {
// It shipped as a standalone page for one build; the health dot and any
// bookmark from it must still land on the surface, which is now a tab.
await router.push('/system')
expect(router.currentRoute.value.name).toBe('settings')
expect(router.currentRoute.value.query.tab).toBe('system')
})
it('series-read is an immersive route', () => {
const r = router.resolve('/series/5/read')
expect(r.name).toBe('series-read')
+25 -2
View File
@@ -13,12 +13,35 @@ export function freshPinia () {
return pinia
}
export function mountComponent (Component, { props = {}, pinia } = {}) {
// `stubs` is merged over the defaults. Needed whenever the component under
// test puts content in a NAMED slot of a Vuetify component: leaving those
// unresolved renders default-slot children only, so a named slot (v-tooltip's
// `#activator`, say) silently renders nothing and assertions find an empty
// wrapper rather than failing loudly.
export function mountComponent (Component, { props = {}, pinia, stubs = {} } = {}) {
return mount(Component, {
props,
global: {
plugins: pinia ? [pinia] : [],
stubs: { RouterLink: { template: '<a><slot /></a>' } },
stubs: { RouterLink: { template: '<a><slot /></a>' }, ...stubs },
},
})
}
// Renders both halves of a v-tooltip: the activator (the thing the operator
// actually sees) and the tip body. `props` is passed as an empty object so the
// activator's `v-bind="tipProps"` binds cleanly.
export const VTooltipStub = {
name: 'VTooltip',
template: '<div><slot name="activator" :props="{}" /><slot /></div>',
}
// Mount with a fresh pinia and the store already seeded, for components that
// read store state during render. Without it, the seeding has to be inlined
// between createPinia and mount in every spec — the same copy-paste that issue
// #3109 tracks for the backend row factories.
export function mountWithStore (Component, seed, opts = {}) {
const pinia = freshPinia()
seed()
return mountComponent(Component, { ...opts, pinia })
}
+22 -1
View File
@@ -90,7 +90,7 @@ DERIVER='scripts/artifacts.sh'
usage() {
echo "usage: artifacts.sh {paths|revision|version} {web|ml|agent|extension}" >&2
echo "usage: artifacts.sh {paths|revision|version|epoch} {web|ml|agent|extension}" >&2
exit 2
}
@@ -149,6 +149,26 @@ cmd_revision() {
echo "$(newest "$1")" | cut -d' ' -f2 | cut -c1-12
}
# The BUILD CLOCK: the same commit's unix timestamp, for SOURCE_DATE_EPOCH.
#
# buildkit stamps the image config's `created` field and every history entry
# with the wall clock of the build unless this is set, so two builds of
# identical source produce different config blobs and therefore different
# manifest digests. That is #3265: the weekly refresh republished all three
# `:latest` tags on 2026-08-30 with every content step CACHED and the bases
# resolved to unchanged digests — nothing was different, and the digest moved
# anyway. A digest that changes on a calendar cannot also mean "the content
# changed", which is the only thing anyone wants it for.
#
# It is the same commit `revision` and `version` name — deliberately, and this
# is the point of routing it through `newest()` rather than taking git's word
# separately. Three values derived from three lookups can disagree; three
# views of one lookup cannot. Note #3127 §2 is the record of what a second
# clock costs.
cmd_epoch() {
echo "$(newest "$1")" | cut -d' ' -f1
}
# The VERSION: `YYYY.MM.DD.HHMM`, zero-padded, UTC. One shape across the whole
# family (note #3127 §1, rule 148) — the number an instance reports about
# itself, and, with a `v` in front, the release tag naming the same build.
@@ -197,5 +217,6 @@ case "$1" in
paths) cmd_paths "$2" ;;
revision) cmd_revision "$2" ;;
version) cmd_version "$2" ;;
epoch) cmd_epoch "$2" ;;
*) usage ;;
esac

Some files were not shown because too many files have changed in this diff Show More