The walk's time budget covered the walk alone, but phase 3 runs in the same
Celery task under the same 1350s soft limit. TamadaHeijun's recapture walked
for about two minutes and handed phase 3 431 orphan imports and ~3000
relinks. Phase 3 ran for 20 minutes and died at the soft limit (event
90808). This predates the worker consolidation; the limits are unchanged
since June.
The walk now also stops when elapsed time plus phase 3's estimated cost
(2.5s per import, 0.25s per relink, measured on the live instance) passes
CHUNK_TOTAL_SECONDS (1200). Work handed to phase 3 counts as progress, so
such a stop is a PARTIAL chunk boundary and the next chunk resumes the page.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
Two concurrent Patreon walks could trip the server rate limit even with each
source pacing its own requests. The platform-cooldown handled the aftermath of
a 429; this adds the preventive half — a per-platform Redis lock so only one
Patreon walk runs at a time. Different platforms still run concurrently up to
the worker concurrency; only a second walk on the SAME serialized platform
waits.
download_source acquires fc:download_lock:<platform> (non-blocking) before the
run. On contention it re-enqueues itself with a short countdown (the pending
event stays — no new event, no log spam), bounded to ~15 min then runs uncapped
as a safety valve. The lock TTL sits just past the hard kill so a SIGKILL'd
worker auto-releases; a backfill chunk only holds it ~10 min, well under the
30-min DownloadEvent recovery sweep. A broker hiccup degrades to uncapped
(prior behaviour) rather than stalling downloads. SERIALIZED_PLATFORMS={patreon};
gallery-dl platforms are left uncapped (self-pacing subprocesses).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Plan #693. Large-catalog backfill (Anduo) no longer sprints to the timeout
wall and dies as an error each run. Builds on the cursor checkpoint (#689).
- Time-boxed chunks: BACKFILL_TIMEOUT_SECONDS(1170)→BACKFILL_CHUNK_SECONDS(600),
far under the 1350 soft limit. Hitting it = normal chunk boundary (the
TimeoutExpired path already captures partial output + the cursor), not a
near-wall death.
- Run-until-done state machine driven by config_overrides[_backfill_state]
(running/complete/stalled). A running backfill auto-continues in chunks
across ticks until gallery-dl exits cleanly (rc=0 = reached the bottom →
'complete'); a safety-cap (BACKFILL_MAX_CHUNKS=200) + the #689 stall-guard
pause a pathological walk as 'stalled'. Replaces the N-runs counter
(backfill_runs_remaining repurposed as the cap countdown).
- Progress, not error: a chunk that timed out but advanced (cursor moved
and/or files written) is reclassified TIMEOUT→PARTIAL (status 'ok').
- Retry storm tamed: gallery-dl retries 3→2, downloader timeout 120→60s, so
one stuck CDN file fails in ~1-2 min not ~10 (Anduo #40838).
- API: POST /sources/{id}/backfill now takes {action: start|stop}; service
start_backfill/stop_backfill; new enabled sources auto-arm run-until-done;
source dict exposes backfill_state + backfill_chunks.
Frontend (Start/Stop control + state badge) lands in the next push.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Backfill downloads stranded with empty logs + a generic "stranded by
recovery sweep" error. Root cause: the backfill gallery-dl subprocess
timeout (1170s) exceeded download_source's Celery soft_time_limit (900s),
so SoftTimeLimitExceeded preempted subprocess.TimeoutExpired. The
TimeoutExpired path (which captures partial stdout/stderr and finalizes
the event) never ran, the event was left 'running', and phase 3 never
decremented backfill_runs_remaining — so the source re-ran and
re-stranded every tick (Anduo #39912).
Two layers:
1. Raise download_source limits (soft 900→1350, hard 1200→1500) so both
subprocess budgets (870 tick / 1170 backfill) sit below the soft
limit with phase-3 persist headroom. Promote to module constants and
guard the invariant with a test.
2. Catch SoftTimeLimitExceeded in download_source and finalize the
in-flight event with a real reason, mirror phase-3 source-health, and
decrement backfill so a chronically-slow source self-heals to tick
mode. The existing celery_signals handler only covered TaskRun, not
DownloadEvent — that was the gap.
Updates stale 900/1200 references in gallery_dl.py + maintenance.py.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>