The point of the milestone rather than its tail. Two of the operator's artists
post a deliberately cropped fragment on Patreon to signal that the real thing
has landed in their Discord; this proposes those pairs.
Confirm-only, following the FC-6.3 series matcher. A wrongly-asserted
association tells the operator two different pieces are one, which is strictly
worse than no link: no link leaves them where they already were, a wrong one
actively misinforms and then propagates into whatever reads it. So the
matcher's job is a SHORT list worth reading, not a long list worth trusting.
**The threshold sits above every single signal weight, and that is the
design.** Proximity is 0.55, declaration 0.45, the cut 0.60 — so neither
signal can carry a pair alone. That makes "time proximity alone is never
sufficient" an arithmetic property rather than an aspiration: on a busy day an
artist posts several times, and a matcher that could pair on proximity alone
would turn every one of those days into false pairs until the review queue got
abandoned. A guard test asserts the relationship against WEIGHTS directly, so
it survives any refactor of the scorer, and says in its own failure message not
to fix it by lowering the assertion.
**Crop-to-source matching is HELD, on the plan's instruction** — real work with
real false-positive risk, worth building only once signals 1 and 2 are shown
insufficient against the operator's actual artists. Worth stating: a naive
whole-image SigLIP similarity is NOT that signal. A cropped teaser and its full
version are precisely the pair a whole-image comparison handles worst, so
adding one as a "bonus" would mostly add noise while looking like progress.
Two premises in the plan corrected in the building:
* **E4 is not actually a prerequisite.** A Patreon Source and a Discord Source
the operator has added under one Artist already share `Post.artist_id`, and
the synthetic grouping inherits it. E4 EXTENDS this to creators FC has to
learn the association for; it is not needed to represent one FC was told.
Same-artist is then a hard filter, not a scored signal — two different
creators posting minutes apart is a coincidence, not evidence.
* **`link_extract` cannot supply the declaration signal.** It exists, but
`SUPPORTED_HOSTS` is file hosts only and `host_for()` returns None for a
Discord URL, so no ExternalLink row is ever written for one. The signal
reads the post body directly instead.
And a bug my own test would have caught: `declared_signal` stripped the HTML
before looking for an invite, but `html_to_plain` discards attributes and
these creators put the invite in an anchor's `href` — so the strongest form of
the signal was being thrown away, leaving only whatever the link text said.
The invite now matches the raw body; the bare mention still matches stripped
text, so `\bdiscord\b` is tested against prose rather than against markup.
Dismissed rows are kept, not deleted: the row is what remembers the rejection,
and re-proposing a rejected pair on every scan is the one behaviour that makes
a review queue get ignored. Both FKs CASCADE, so E3's one-DELETE reversal
cannot leave a proposal pointing at a post that no longer exists.
Only ACCEPTED links reach the post payload. A pending proposal is a question
for the review queue, not a claim to render beside the artwork.
UI (rule 27) follows in the next commit.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9
Consolidate duplication accrued across the ML tagging + settings backend,
behavior-preserving (over-DRY guard applied — the three auto-apply sweep
BODIES stay separate; only their shared inner helpers are extracted).
- _sigmoid / _conflict_scores / _insert_presentation_review (heads.py): the
score→prob transform (6 inlined sites), the presentation conflict signal
(2 sites), and the ring-loud PresentationReview insert (2 sites, single-
sourced so the mode column can't drift on the shared composite PK).
- _applied_or_rejected (training_data.py): the per-tag "applied ∪ rejected"
skip-set, byte-identical at 3 sweep sites (heads.py x2, tasks/ml.py ccip).
- ccip sweep divergence fixes: import ccip._FIGURE_KINDS + training_data._l2norm
instead of local copies that silently drift when the canonical changes.
- MLSettings.load / .load_sync classmethods (mirror ImportSettings); route all
8 scalar_one singleton reads through them (the session.get None-path stays).
- GET serializers for MLSettings + ImportSettings are now table-driven off the
same _EDITABLE tuples PATCH writes, so a new field can't be silently absent
from GET (the split that historically dropped fields).
- AUTO_APPLY_THRESHOLD_MIN/MAX constant single-sources the [0.5,0.999] operating
range across the service clamp + the 5 API validators.
- test_ml_dry_helpers.py pins _applied_or_rejected + _sigmoid.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NsmJSQxnNxGgtM5Yz4GAqi
Extends WIP title-tagging to lower-precision cues (sketch/doodle/scribble) safely.
- wip_title.py: soft matcher (word-anchored; sketchbook/kadoodle don't trip it);
WIP_TITLE_SOFT_SOURCE + soft SQL prefilter; apply_wip_image_tags takes a source arg.
- training_data._AUTO_SOURCES += 'wip_title_soft' → the soft tier is PROVISIONAL and
never trains the wip head (a finished "sketch" can't pollute it). Only the hard
tier (wip_title) + manual train.
- ImportSettings.wip_soft_title_tagging_enabled (OFF by default, opt-in). Migration 0087.
- importer: hard tier wins, soft is the fallback (source wip_title_soft).
- backfill: refactored into a shared _backfill_wip_tier; hard always, soft when enabled.
- heads.soft_wip_conflict_audit + daily beat: score soft-tagged images against content
heads, flag ring-loud ones (PresentationReview mode=process) for the review strip —
the operator's "measure if they got falsely tagged" safety.
- api settings toggle; ImportFiltersForm soft toggle.
- tests: soft matcher pos/neg; soft source not a training positive; audit flags
ring-loud + spares quiet.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Auto-apply the `wip` system tag to posts whose TITLE explicitly declares
work-in-progress ("WIP" / "work in progress") — a deterministic, high-precision
complement to the image-based ML `wip` head. WIP images are excluded from the
Explore/gallery browse, so honouring the artist's own label keeps unfinished
pieces out of the main browse.
- services/wip_title.py: precision-first token-anchored matcher (swipe/wiped
never trip it) + sync apply helpers (source='wip_title', ON CONFLICT DO
NOTHING, chunked under the psycopg param ceiling).
- importer: live hook on FRESH import only (never on deep-scan/supersede), so a
manually-removed WIP tag is never re-applied by a routine re-scan.
- maintenance.backfill_wip_title_tags: operator-triggered back-catalogue sweep
(coarse SQL prefilter + regex confirm, keyset-paginated). Deliberately NOT a
beat — a periodic re-run would silently undo manual removals.
- ImportSettings.wip_title_tagging_enabled (default ON, migration 0085) gating
the live hook; GET/PATCH + POST /settings/wip-title/scan.
- Settings UI: toggle + "Scan existing posts" button.
- Tests: pure matcher unit tests + integration (apply idempotency, backfill
precision).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The gate at a fixed 0.80 couldn't catch the real pain: Interpreter (fresh ==
cached, verified by probe) confidently mis-detects short ASCII English like
"... WIP Part 1" as German at 0.86 — above the floor — so it was accepted and a
re-translate reproduced it. Confidence alone can't separate the 0.86 collision
(genuine German lands there too), and single-word mis-flags sit at a confident
1.0 no floor catches.
Two operator-approved levers:
- Acceptance floor is now a live Settings value (ImportSettings.
translation_min_confidence, default 0.90; surfaced in the Translation card), so
it's tunable without a redeploy. _accept takes the threshold as a parameter.
- Per-post sticky override (Post.translation_override: auto/force/original).
'force' stores a translation even below the floor (rescue a skipped
legit-foreign title); 'original' keeps the original and clears any stored
translation (kill a confident mis-flag no floor catches). The sweep honors it
on every run and _reset_translations skips 'original', so the choice survives a
Re-translate-all. POST /api/posts/<id>/translation-override applies it
immediately (translate now when the service is up, else queue for the sweep).
UI: PostTranslationControl on the posts-feed card.
Migration 0084 (both columns + a CHECK on the override). The feed + provenance
serializers expose translation_override.
With a stricter floor the rollback finally works: raise it -> Re-translate all ->
the 0.86 mis-flags are rejected and restored to the original; force /
keep-original handle the residual either way.
Tests: gate thresholds against the param (0.86 rejected at 0.90, explicit-floor
cases); sweep force/original + re-translate-skips-original; override endpoint
(validation, original clears, force queues when disabled, feed exposes it);
settings min_confidence default/save/validate.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CgZP9v2otxVJymiYsnVuMy
Throughput: translate_posts now runs every 8h (was daily) as the
steady-state cadence for newly-imported posts, and the Settings
"Translate now" button runs it in drain mode (run-until-done, no reset)
so one press clears the whole untranslated backlog instead of a single
300-post chunk. The interrupt/backoff re-enqueue now preserves the drain
flag so a bulk drain resumes cleanly after an Interpreter restart.
Misdetection groundwork: surface the detector's confidence from the
Interpreter client (it was in the detectedLanguage payload but discarded)
and add a read-only "Test translation" box — POST /settings/translation/
probe + TranslationCard UI — that shows detected language + confidence +
engine + result for pasted text, without saving. Lets the operator see
why a short/abbreviation-heavy English title gets mis-detected so the
detection guard (min-length + confidence floor) can be tuned from real
numbers. The guard itself follows once the mis-detected cases are probed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CgZP9v2otxVJymiYsnVuMy
translation/status now reports `active` (a translate/retranslate sweep is
running, from the TaskRun table) and `last_run` (the most recent finished run's
task + status). The Settings card polls live while a sweep runs, showing a
spinner + "Translating… N remaining" that ticks down, and flags a last run that
ended in error/timeout. No migration.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
retranslate_posts resets the 5 translation columns to NULL for a scoped set of posts (all, or WHERE artist_id IN ids) then reuses the untranslated sweep to re-run them, chasing the tail until drained (run-until-done). Interpreter cache keys on engine_version so a changed model re-translates, an unchanged one is cache-fast. Reset only happens when the service is configured+healthy so translations are never wiped when they can't be rebuilt. New POST /settings/translation/retranslate (artist_id | all=true). UI: per-artist 'Re-translate posts' on the Artist Management tab + 'Re-translate all' in the Settings Translation card, both with confirm dialogs. No migration (reuses m143 columns).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
New POST /api/settings/translation/test pings /v1/health for a GIVEN base URL (not
the saved one), so the operator can verify a URL before enabling it. TranslationCard
gains a Test-connection button that reports reachable/unreachable inline and
updates the status dot. Endpoint test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM
tasks/translation.py — translate_posts: picks untranslated posts (title OR
description non-empty), per-post [title, description] batch via the Interpreter
client, stores translations + detected lang + engine_version; passthrough /
already-target posts are marked handled with no stored translation. 503 or a
connection error interrupts (retry next cycle), 400 stops (fix config), per-post
commit keeps progress; wall-clock bounded. Wired into celery (maintenance_long
lane) + a daily beat. No-op unless enabled + base URL set + healthy. GET
/settings/translation/status + POST .../run for the Settings card. Task tests
(stubbed client, monkeypatched session).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM
Post gains post_title_translated / description_translated / translated_source_lang
/ translation_engine_version / translated_at — filled by the translate sweep so
viewing is instant. ImportSettings gains translation_enabled (OFF by default),
interpreter_base_url (EMPTY — no default host; the operator points it at their own
Interpreter proxy behind a reverse proxy) and translation_target_lang (en),
exposed + validated via /settings/import. Migration 0083. Settings defaults +
patch + validation test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM
Operator lever: disable a single file host (e.g. mega.nz when it's banning)
without touching the others. Five booleans on import_settings
(extdl_<host>_enabled, default true — works out of the box, rule #26); the
worker already reads them via getattr so no worker change. Migration 0050 +
model fields + settings GET/PATCH (uniform boolean validation) + a
'External file-host downloads' card in the subscriptions Settings tab.
Completes Phase 4. Refs FC #830.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Confirm-only "this post may continue this series" matcher.
- series_suggestion table (post_id, series_tag_id, score, signals jsonb, status
pending|added|dismissed, UNIQUE(post,series)); migration 0041 + two settings
knobs (series_suggest_enabled, series_suggest_threshold).
- series_match_service: weighted additive score (title-stem / same-artist /
page-continuity / shared-distinctive-tags), no single signal gating. The title
"pattern" is derived on the fly from the post titles already in a series, so it
sharpens as more are confirmed (no persisted state to drift). Candidates are
bounded to the post's artist. match_post upserts pending suggestions (UNIQUE +
on-conflict, respecting prior added/dismissed decisions).
- accept reuses add_post_as_chapter then marks 'added'; dismiss marks 'dismissed'.
- rescan_series_suggestions_task: settings-gated, time-boxed + self-resuming from
a post-id cursor (maintenance_long lane), like normalize_tags_task.
- API: GET /series/suggestions, POST .../<id>/accept|dismiss, POST .../rescan.
- Settings: enabled + threshold exposed via /settings/import.
- Tests: pure scoring helpers + matcher/accept/dismiss/rescan lifecycle + UNIQUE
dedup.
Frontend (Suggestions tab + settings card) lands next.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The `select(ImportSettings).where(id == 1)).scalar_one()` singleton load was
repeated 15× across services, API, and 5 task modules. Added async load() +
sync load_sync() classmethods on the model and migrated all 15 full-row sites
(callers already imported ImportSettings, so no new imports; dropped download's
now-orphaned select import). Left maintenance.py's deliberate column-select
(import_scan_path only) as-is.
Rest of the service layer was already adequately DRY — the Record/to_dict
pattern is only 2 instances and the savepoint find-or-create recovery is
correctly per-entity, so neither was forced into a shared abstraction.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Settings GET returns the single-row ImportSettings DB record; PATCH does
selective field updates (only known editable fields are applied; unknown
keys silently ignored). System stats returns image/tag/storage counts,
task status histogram, and the active batch summary in one shot for the
Settings overview tab.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>