fix: backfill_phash runs on the long maintenance lane, not the scheduler's quick one (4411)
CI and images / lint (push) Successful in 2s
CI and images / extension-version (push) Successful in 3s
CI and images / frontend-build (push) Successful in 21s
CI and images / backend-lint-and-test (push) Successful in 33s
CI and images / integration (push) Successful in 2m21s
CI and images / sign-extension (push) Successful in 4s
CI and images / build-agent (push) Successful in 6s
CI and images / build-web (push) Successful in 1m47s
CI and images / smoke-web (push) Successful in 55s
CI and images / promote (push) Successful in 1s
CI and images / lint (push) Successful in 2s
CI and images / extension-version (push) Successful in 3s
CI and images / frontend-build (push) Successful in 21s
CI and images / backend-lint-and-test (push) Successful in 33s
CI and images / integration (push) Successful in 2m21s
CI and images / sign-extension (push) Successful in 4s
CI and images / build-agent (push) Successful in 6s
CI and images / build-web (push) Successful in 1m47s
CI and images / smoke-web (push) Successful in 55s
CI and images / promote (push) Successful in 1s
A whole-library rehash with a 35-minute limit was matched by the maintenance.* glob. It held a scheduler process for its whole run, and the minute ticks queued behind it. An exact-name route now sends it to maintenance_long. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
This commit is contained in:
@@ -62,6 +62,13 @@ def make_celery() -> Celery:
|
|||||||
# can never starve the quick self-healing sweeps (operator-flagged
|
# can never starve the quick self-healing sweeps (operator-flagged
|
||||||
# 2026-06-07: a 2h audit blocked vacuum/backup/normalize for hours).
|
# 2026-06-07: a 2h audit blocked vacuum/backup/normalize for hours).
|
||||||
"backend.app.tasks.maintenance.*": {"queue": "maintenance"},
|
"backend.app.tasks.maintenance.*": {"queue": "maintenance"},
|
||||||
|
# The one long job in maintenance.py: a whole-library phash
|
||||||
|
# recompute (35 min hard limit; the library was cleared for
|
||||||
|
# re-hashing by migration 0098). On the quick lane it held a
|
||||||
|
# scheduler process for its whole run, and the minute ticks queued
|
||||||
|
# up behind it (2026-09-24: 7 waiting, "all workers busy for 18
|
||||||
|
# minutes"). An exact name wins over the glob above.
|
||||||
|
"backend.app.tasks.maintenance.backfill_phash": {"queue": "maintenance_long"},
|
||||||
"backend.app.tasks.backup.*": {"queue": "maintenance_long"},
|
"backend.app.tasks.backup.*": {"queue": "maintenance_long"},
|
||||||
"backend.app.tasks.admin.*": {"queue": "maintenance_long"},
|
"backend.app.tasks.admin.*": {"queue": "maintenance_long"},
|
||||||
"backend.app.tasks.library_audit.*": {"queue": "maintenance_long"},
|
"backend.app.tasks.library_audit.*": {"queue": "maintenance_long"},
|
||||||
|
|||||||
@@ -36,3 +36,17 @@ def test_queue_for_mirrors_external_to_download():
|
|||||||
celery.conf.task_routes["backend.app.tasks.external.*"]["queue"]
|
celery.conf.task_routes["backend.app.tasks.external.*"]["queue"]
|
||||||
== "download"
|
== "download"
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_backfill_phash_runs_on_the_long_lane():
|
||||||
|
"""It lives in maintenance.py, so the quick-lane glob matches it too —
|
||||||
|
the router must pick the exact name. A 35-minute rehash on the scheduler
|
||||||
|
lane blocked the minute ticks behind it (2026-09-24)."""
|
||||||
|
route = celery.amqp.router.route(
|
||||||
|
{}, "backend.app.tasks.maintenance.backfill_phash",
|
||||||
|
)
|
||||||
|
assert route["queue"].name == "maintenance_long"
|
||||||
|
quick = celery.amqp.router.route(
|
||||||
|
{}, "backend.app.tasks.maintenance.recover_stalled_task_runs",
|
||||||
|
)
|
||||||
|
assert quick["queue"].name == "maintenance"
|
||||||
|
|||||||
Reference in New Issue
Block a user