CI and images / lint (push) Failing after 4s
CI and images / extension-version (push) Successful in 4s
CI and images / extension-test (push) Successful in 18s
CI and images / frontend-build (push) Successful in 21s
CI and images / backend-lint-and-test (push) Successful in 32s
CI and images / integration (push) Successful in 2m22s
CI and images / sign-extension (push) Skipped
CI and images / build-web (push) Skipped
CI and images / smoke-web (push) Skipped
CI and images / promote (push) Skipped
CI and images / build-agent (push) Skipped
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
58 lines
2.3 KiB
Python
58 lines
2.3 KiB
Python
"""Closing the download runs a worker restart orphaned (#4433).
|
|
|
|
Its own module, importing nothing but the model, because its caller is the
|
|
worker boot hook in `celery_signals` — which every download task imports via
|
|
`celery_app`. Living in `tasks.maintenance` put the whole maintenance import
|
|
graph, the membership roster included, on the fetch path, which
|
|
`test_gated_reason` forbids.
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
from datetime import UTC, datetime
|
|
|
|
from sqlalchemy import literal, update
|
|
from sqlalchemy.dialects.postgresql import JSONB
|
|
|
|
from ..models import DownloadEvent
|
|
|
|
DOWNLOAD_INTERRUPTED_MESSAGE = (
|
|
"interrupted by a worker restart — the next check picks it up where it left off"
|
|
)
|
|
|
|
|
|
def interrupt_orphaned_download_events(session, *, booted_at: datetime) -> int:
|
|
"""Close the download events a restart orphaned, without blaming the source.
|
|
|
|
Called when the download lane comes up (#4433). Anything still
|
|
pending/running from before this boot belongs to the previous process:
|
|
a walk that outlived the 90s stop grace was SIGKILLed, and a queued or
|
|
serialize-deferred task is held unacked until Redis redelivers it about an
|
|
hour later. Left alone, the 30-min stall sweep would error each one and
|
|
bump `consecutive_failures`, backing the source off as if the platform had
|
|
failed.
|
|
|
|
Instead they end as `skipped` (terminal, not a failure) and the source is
|
|
not touched: `last_checked_at` keeps its old value, so the next tick finds
|
|
it due and the walk resumes from its checkpoint. A redelivered message
|
|
that arrives later finds no pending event and opens a fresh one.
|
|
|
|
An event promoted to running after the boot has `started_at` reset to its
|
|
real start (download_service), so it is never caught here. Does NOT commit.
|
|
"""
|
|
now = datetime.now(UTC)
|
|
result = session.execute(
|
|
update(DownloadEvent)
|
|
.where(DownloadEvent.status.in_(["pending", "running"]))
|
|
.where(DownloadEvent.started_at < booted_at)
|
|
.values(
|
|
status="skipped",
|
|
finished_at=now,
|
|
error=DOWNLOAD_INTERRUPTED_MESSAGE,
|
|
metadata_=DownloadEvent.metadata_.op("||")(
|
|
literal({"error_type": "interrupted"}, JSONB)
|
|
),
|
|
)
|
|
.returning(DownloadEvent.id)
|
|
)
|
|
return len(result.all())
|