Files
FabledCurator/backend/app/tasks/download.py
T
bvandeusenandClaude Opus 5 2b093958a4
CI and images / extension-version (push) Successful in 3s
CI and images / lint (push) Successful in 3s
CI and images / frontend-build (push) Successful in 19s
CI and images / backend-lint-and-test (push) Successful in 40s
CI and images / integration (push) Failing after 2m17s
CI and images / sign-extension (push) Skipped
CI and images / build-web (push) Skipped
CI and images / smoke-web (push) Skipped
CI and images / promote (push) Skipped
CI and images / build-agent (push) Skipped
feat: a tick keeps looking back 30 days, so an EDITED post is reached (4386)
Operator, 2026-09-23, on a Floppystack post: "this post has been updated as
he implements hot fixes — any chance we have a way to scan for or see updated
posts so we can update ours to match and pull the new attachments and
pictures etc."

The download half already worked: extract_media reads the media list off the
LIVE feed response every walk, so a newly attached hotfix build is a ledger
key we have never seen. Only REACHING the post was missing — a tick stopped
after 20 contiguous already-have-it items, and a post edited three days after
publication sits well below twenty. Not a bug in the early-out; a count
cannot express "recent".

The early-out now needs BOTH conditions: the run of seen items AND a post
published before the horizon. Strictly a widening — window 0 is exactly the
old behaviour, and no window can make a tick stop EARLIER than it used to, so
a source paused for months still walks its whole unseen backlog. The horizon
is a floor on how far to look, never a ceiling.

Inside the window the post-record gate is bypassed too (write_post_record
revisit=True): the body is re-read from the feed response already in hand, so
a revisit costs zero requests, and a body that comes back empty writes
NOTHING rather than blanking one a detail-fetch had filled. Revisits are kept
out of the #862 body-drift canary's sample for the same reason — an empty
revisit is healthy, and counting it would walk the alarm toward firing on
good ticks.

The run summary names what changed ("3 post(s) updated (5 new file(s))") with
a line per post; the ask was to SEE updated posts, not only to end up with
their bytes.

download_revisit_days is a settings row, not a constant (rule 25) — how long
a creator keeps editing is a property of the creator. Default 30, 0 turns it
off. Migration 0108.

Also corrects two stale docstrings: both clients described post_meta as
feeding an Ingester.preview that no longer calls it. It had no consumer at
all until this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-23 19:21:06 -04:00

245 lines
11 KiB
Python

"""download_source Celery task — runs DownloadService for one source."""
import asyncio
import logging
from datetime import UTC, datetime
from pathlib import Path
from celery.exceptions import SoftTimeLimitExceeded
from sqlalchemy import select
from sqlalchemy.exc import DBAPIError, OperationalError
from sqlalchemy.orm import Session as SyncSession
from ..celery_app import celery
from ..models import DownloadEvent, ImportSettings, Source
from ..services.credential_crypto import CredentialCrypto
from ..services.credential_service import CredentialService
from ..services.download_service import DownloadService
from ..services.gallery_dl import GalleryDLService
from ..services.importer import Importer
from ..services.thumbnailer import Thumbnailer
from ._async_session import async_session_factory
from .import_file import _sync_session_factory
log = logging.getLogger(__name__)
IMAGES_ROOT = Path("/images")
_KEY_PATH = IMAGES_ROOT / "secrets" / "credential_key.b64"
# Celery time budget for one download_source run. The ceiling that
# governs *clean* teardown is the SOFT limit: it raises a catchable
# SoftTimeLimitExceeded in-process, whereas the HARD limit SIGKILLs the
# worker (no chance to finalize). Both gallery-dl subprocess budgets
# (gallery_dl.py: _DEFAULT_GDL_TIMEOUT_SECONDS=870 tick,
# BACKFILL_CHUNK_SECONDS=600 per backfill chunk, plan #693) MUST sit below the soft limit
# so subprocess.run raises its own TimeoutExpired first — that path
# captures partial stdout/stderr and finalizes the DownloadEvent. soft is
# max-subprocess (1170) + ~180s phase-3 persist headroom; hard is soft +
# 150s SIGKILL backstop. Audit 2026-06-03 (Anduo #39912): the old
# soft=900 sat BELOW the 1170 backfill budget, so SoftTimeLimitExceeded
# preempted TimeoutExpired and the event stranded empty. The recovery
# sweep's DOWNLOAD_STALL_THRESHOLD_MINUTES (30 min) still trails the new
# 25-min hard kill by 5 min, so it stays a true backstop. Invariant
# guarded by test_timeout_ladder_keeps_subprocess_budgets_under_soft_limit.
DOWNLOAD_SOFT_TIME_LIMIT = 1350
DOWNLOAD_HARD_TIME_LIMIT = 1500
# Per-platform serialization (plan: concurrency cap). When a serialized
# platform (Patreon) is already walking, defer this run by re-enqueuing it a
# little later rather than holding a worker slot or bowling into the same rate
# limit. Bounded so a wedged platform eventually runs anyway. The lock TTL sits
# just past the hard kill so a SIGKILL'd worker's lock auto-expires; a backfill
# chunk only holds it ~10 min, so the bounded wait stays well under the 30-min
# DownloadEvent recovery sweep.
_PLATFORM_LOCK_TTL = DOWNLOAD_HARD_TIME_LIMIT + 120
_SERIALIZE_COUNTDOWN = 20
_MAX_SERIALIZE_WAITS = 45 # ~15 min ceiling, then run uncapped as a safety valve
def _peek_platform(source_id: int) -> str | None:
SyncFactory = _sync_session_factory()
with SyncFactory() as session:
return session.execute(
select(Source.platform).where(Source.id == source_id)
).scalar_one_or_none()
def _finalize_soft_limited(session: SyncSession, source_id: int) -> None:
"""Defense in depth for the soft-time-limit kill path.
A SoftTimeLimitExceeded unwinds download_source before phase 3 can
finalize the DownloadEvent, leaving it 'running' until the recovery
sweep stamps a context-free "stranded" error 30 min later — AND
leaving backfill_runs_remaining undecremented so the source re-runs
and re-strands every tick (Anduo #39912, 2026-06-03). Flip the
in-flight event to error with a real reason, mirror phase 3's
source-health write, and decrement any backfill budget so a
chronically-slow source self-heals back to tick mode.
The caller owns the commit. All mutations are gated on actually
finding a running event, so a benign late soft-limit (phase 3 already
committed) is a no-op.
"""
now = datetime.now(UTC)
ev = session.execute(
select(DownloadEvent)
.where(DownloadEvent.source_id == source_id)
.where(DownloadEvent.status == "running")
.order_by(DownloadEvent.id.desc())
.limit(1)
).scalar_one_or_none()
if ev is None:
return
ev.status = "error"
ev.finished_at = now
ev.error = (
f"killed by Celery soft time limit ({DOWNLOAD_SOFT_TIME_LIMIT}s) "
"before the download finished — the run exceeded its time budget. "
"Progress is checkpointed per page, so the next tick resumes near "
"the cut; the backfill budget was decremented so it walks less. If "
"this recurs, a single post is heavy enough to overrun the chunk "
"time-box on its own."
)
ev.metadata_ = {
**(ev.metadata_ or {}),
"error_type": "timeout",
"soft_time_limited": True,
}
src = session.get(Source, source_id)
if src is not None:
src.consecutive_failures = (src.consecutive_failures or 0) + 1
src.last_error = "soft time limit exceeded"
src.error_type = "timeout"
src.last_checked_at = now
if (src.backfill_runs_remaining or 0) > 0:
src.backfill_runs_remaining = max(0, src.backfill_runs_remaining - 1)
@celery.task(
name="backend.app.tasks.download.download_source",
bind=True,
acks_late=True,
autoretry_for=(OperationalError, DBAPIError, OSError),
retry_backoff=10,
retry_backoff_max=120,
retry_jitter=True,
max_retries=3,
soft_time_limit=DOWNLOAD_SOFT_TIME_LIMIT,
time_limit=DOWNLOAD_HARD_TIME_LIMIT,
)
def download_source(self, source_id: int, _serialize_waits: int = 0) -> int:
"""Returns the DownloadEvent.id."""
# Per-platform concurrency cap: only one Patreon walk runs at a time.
from ..services.platform_lock import platform_lock
platform = _peek_platform(source_id)
lock = (
platform_lock(platform, ttl_seconds=_PLATFORM_LOCK_TTL)
if platform is not None
else None
)
if lock is not None:
try:
got_lock = bool(lock.acquire(blocking=False))
except Exception: # noqa: BLE001 — broker hiccup → degrade to uncapped
log.warning("platform lock acquire failed for %s; running uncapped", platform)
lock = None
got_lock = True
if lock is not None and not got_lock:
if _serialize_waits < _MAX_SERIALIZE_WAITS:
# Another walk on this platform holds the lock — re-enqueue
# ourselves shortly (the pending DownloadEvent stays; no new
# event, no log spam) rather than running concurrently into
# the rate limit.
download_source.apply_async(
(source_id,),
{"_serialize_waits": _serialize_waits + 1},
countdown=_SERIALIZE_COUNTDOWN,
)
log.info(
"download_source(%s) deferred — %s already walking (wait %d/%d)",
source_id, platform, _serialize_waits + 1, _MAX_SERIALIZE_WAITS,
)
return -1
log.warning(
"download_source(%s) running WITHOUT the %s lock — max serialize "
"waits hit, proceeding uncapped",
source_id, platform,
)
lock = None
async def _run():
async_factory, async_engine = async_session_factory()
SyncFactory = _sync_session_factory()
try:
with SyncFactory() as sync_session:
settings = ImportSettings.load_sync(sync_session)
rate_limit = settings.download_rate_limit_seconds
validate_files = settings.download_validate_files
# How far back a tick keeps looking for EDITED posts. Read here
# with the other downloader knobs, off the row this block is
# already holding open.
revisit_days = settings.download_revisit_days
gdl = GalleryDLService(
images_root=IMAGES_ROOT,
rate_limit=rate_limit,
validate_files=validate_files,
)
crypto = CredentialCrypto(_KEY_PATH)
async with async_factory() as async_session:
cred_service = CredentialService(async_session, crypto)
with SyncFactory() as sync_session:
sync_settings = ImportSettings.load_sync(sync_session)
importer = Importer(
session=sync_session,
images_root=IMAGES_ROOT,
import_root=IMAGES_ROOT,
thumbnailer=Thumbnailer(images_root=IMAGES_ROOT),
settings=sync_settings,
)
svc = DownloadService(
async_session=async_session,
sync_session=sync_session,
gdl=gdl,
importer=importer,
cred_service=cred_service,
# The native Patreon ingester opens its own short-lived
# sync sessions for the seen-ledger (never held across
# the walk). Same factory the importer's sync session
# comes from — a different DB connection per checkout.
sync_session_factory=SyncFactory,
revisit_days=revisit_days,
)
return await svc.download_source(source_id)
finally:
await async_engine.dispose()
try:
return asyncio.run(_run())
except SoftTimeLimitExceeded:
# phase 3 never ran — salvage the in-flight event so the operator
# sees a real reason instead of the recovery sweep's generic
# "stranded" 30 min later (Anduo #39912). Best-effort: a failure
# here must not mask the timeout. Re-raise so Celery + the
# task_run signal handler still record the kill.
try:
SyncFactory = _sync_session_factory()
with SyncFactory() as session:
_finalize_soft_limited(session, source_id)
session.commit()
except Exception: # noqa: BLE001 — cleanup must not swallow the kill
log.exception("soft-limit finalize failed for source %s", source_id)
raise
finally:
# Release the per-platform lock so the next walk can proceed. The TTL is
# a backstop for a SIGKILL; here we free it the instant the run ends. A
# late release after TTL expiry raises (lock already gone) — ignore it.
if lock is not None:
try:
lock.release()
except Exception: # noqa: BLE001 — TTL may have already freed it
pass