5aa8e3d81b4b43440c6a1f797d2c36a25a86e99d
9
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5aa8e3d81b |
fix: a stopped source is not a failing one, and cannot be deep-scanned (4279)
CI / lint (push) Successful in 2s
CI / extension-version (push) Successful in 2s
Build images / sign-extension (push) Successful in 3s
Build images / build-agent (push) Successful in 6s
CI / frontend-build (push) Successful in 20s
CI / backend-lint-and-test (push) Successful in 33s
Build images / build-web (push) Successful in 1m3s
Build images / smoke-web (push) Skipped
Build images / build-ml (push) Successful in 2m12s
Build images / promote (push) Skipped
CI / integration (push) Successful in 2m16s
Ebi77 sat in the "1 source is failing" banner for six days with no action available, reading `stranded by recovery sweep (no terminal status after time_limit)`. Four things lined up: 1. The membership sweep did its job — saw `former_patron`, disabled the source, cleared its failure state. Clean at 02:50. 2. Twenty minutes later a deep scan was armed on it. `/backfill` had a credential pre-flight but NO `enabled` guard, while `/check` has carried one all along. The two trigger endpoints disagreed, and the ungated one is the one that arms the long walk. 3. Without a membership the walk cannot finish, never reaches a terminal status, and the recovery sweep strands it with consecutive_failures = 1. 4. Nothing could clear that. A disabled source is never scheduled, so no successful run resets the count; `SourceService.update` clears only on an explicit disable and it was already disabled; and the banner's Retry routes to `/check`, which refuses a disabled source. The card offered a button structurally incapable of acting on the only source it was showing. `failing_sources_clause()` now means "enabled AND erroring". That also settles a disagreement its two callers already had: the scheduler's count paired it with `enabled.is_(True)` and `SourceService.list(failing=True)` did not, so one counted Ebi77 and the other did not — exactly the drift the note above that function warns about, which is why the test belongs IN the predicate rather than beside it. The scheduler's now-duplicate clause is dropped so one place decides. `/backfill` gains the guard for start/recover/recapture. `stop` stays open on a disabled source, or arming becomes a one-way door. Migration 0101 clears failure state on sources that are already disabled — the predicate fixes what the surfaces report, not what the rows carry, and the rows are why the operator had no way out (lesson #4202). It matches what `update` already does on an explicit disable, so rows disabled by any other path come into line. Enabled sources are untouched: a real failure on a live source must keep showing, which the second new test pins. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR |
||
|
|
6b19012bb6 |
feat: the empty front door is the install's first screen (milestone 387 step B4)
CI / lint (push) Successful in 5s
CI / extension-version (push) Successful in 6s
CI / frontend-build (push) Successful in 26s
CI / backend-lint-and-test (push) Successful in 36s
CI / integration (push) Successful in 2m3s
Build images / sign-extension (push) Successful in 5s
Build images / build-agent (push) Successful in 9s
Build images / build-ml (push) Successful in 9s
Build images / build-web (push) Successful in 6s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
The front door is now a feed, and a blank feed implies things should be here in a way a blank masonry does not. On a fresh install this is the first screen anyone sees — including someone who is not the operator, which is what milestone 328 is making possible. Tells the two empties apart, which is the point. "No sources yet" gets the on-ramp; "sources configured, nothing landed yet" gets told that the first check takes a while and pointed at Downloads. Telling someone to add a source when they already have three and are mid-backfill reads as the app not knowing its own state. Needs total_sources on schedule-status to distinguish them — deliberately not auto_sources, which counts only what is on a schedule, so a source with auto_check off would have read as "nothing configured". Both exact-shape assertions updated in THIS change rather than after CI caught them, which is the lesson from B3's red push. An absent status falls back to the on-ramp on purpose: it is merely redundant to an established operator, whereas "see what's running" shown to someone with nothing configured is a dead end. A filtered miss is deliberately NOT the onboarding case — the operator has posts, they just narrowed past them. Showing a fresh-install on-ramp there would tell someone with a full library to go set it up. This is where logo.svg lands, as the operator asked. It earns its place on a first-run screen and not on a populated feed, and gives the on-ramp something to compose around instead of prose plus two buttons. Large: the mark stops reading below ~48px, which is why the 22px nav slot has a different one. Pinned by test so a later tidy-up cannot quietly shrink it to a glyph. Also extracts mountWithStore into the shared test support module. Writing the second spec created exactly the copy-paste that open issue 3109 tracks for the backend row factories, so it is consolidated now rather than at copy three, and recorded as snippet 3829 with the two traps it does NOT solve — named slots rendering nothing, and components that fetch on mount. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9 |
||
|
|
a708f5e9db |
feat: the front door says whether ingestion is working (milestone 387 step B3)
CI / extension-version (push) Successful in 4s
CI / lint (push) Successful in 4s
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 9s
CI / frontend-build (push) Successful in 27s
CI / backend-lint-and-test (push) Successful in 34s
Build images / build-web (push) Successful in 1m17s
Build images / smoke-web (push) Skipped
CI / integration (push) Failing after 2m8s
Build images / build-ml (push) Successful in 2m18s
Build images / promote (push) Skipped
The step phase A was building toward. A1 made the gated count true, A2 made it a durable state, A3 made it visible in Subscriptions — but Subscriptions is where you go once you already suspect something. This is the line that reaches someone who wasn't looking. A thin grey strip above the feed, front door only: last check, sources failing, sources you can't see. Only the actionable items take a colour, and nothing renders at zero — a permanent "0 failing" trains you to skip the line, which would hide the real number when it appears. Two predicates, defined once. The ribbon counts and the surfaces it links to have to agree on what "failing" and "no access" MEAN, or the ribbon says 3 and the card shows 4. They live in db_helpers, which exists for exactly this reason (its docstring: divergent copies are how the race bugs crept in). Not in source_service, because scheduler_service needs them too and source_service already imports scheduler_service — the other direction is a cycle. Counting deliberately spans all ENABLED sources rather than the auto_check subset scheduler_status already walks: a source erroring on a manual-only artist is still erroring. Disabled sources count for nothing, which is what makes issue 1285 the real escape hatch for a sub you stopped paying for. Extends the existing schedule-status endpoint rather than adding a parallel aggregate — the store already fetches it. Two scalar COUNTs. The status filter is now URL-addressable, which it had to be for the ribbon's links to land anywhere: a count that drops you on an unfiltered list makes the reader redo the filtering the ribbon just did. Mirrors how artistFilter already reads from route.query. Front-door-only via a route prop, not a route.name check, so the view doesn't need to know what it's mounted as and the router states the intent in one place. Inside Browse's Posts tab you're looking FOR something and the hub is one click away. The fetch is swallowed on mount by design (rule 164): this is an aside, and the feed must render whether or not the status call succeeds. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LNXXULQDjVZmbuNa2G9mD9 |
||
|
|
a5101494b6 |
feat(downloads): bulk retry respects cooldown; single-source RETRY overrides
Today's platform-cooldown commit (
|
||
|
|
77f7a23410 |
feat(scheduler): order due sources by last_checked_at — most overdue first
select_due_sources returned rows in undefined order (Postgres-determined, typically PK). At tick rates that outpace download-queue throughput, a freshly-rerun source could keep getting re-queued ahead of one that's still waiting for its first attempt this cycle. Operator-flagged 2026-05-30: > if there are 8 hours before a source is due again and 40 full time > downloads can happen in that period that means that there's a chance > the first one to fire gets back into the download queue before item 41 > has a chance to get downloaded. Added `ORDER BY last_checked_at ASC NULLS FIRST, id` to the due-source SELECT. Never-checked sources go first, then longest-since-checked, then ties broken by id. Combined with Celery's FIFO `download` queue, the oldest-overdue source in each tick now reaches a worker before any fresher one. Test pins the ordering: a NULL-last_checked source, a 4-hour-overdue source, and a 2-min-overdue source come back in that exact order from select_due_sources. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
61ce1ce13c |
feat(scheduler): platform-wide cooldown on RATE_LIMITED — burst prevention
The scan tick fired download_source.delay() for every due source without
grouping by platform; with multiple download workers, N due Patreon
sources could all hit Patreon's API in parallel and rate-limit each
other. Per-source consecutive_failures backoff REACTS to that (slows the
offender across cycles) but didn't PREVENT the first-tick burst.
When DownloadService._update_source_health sees a source error
classified as ErrorType.RATE_LIMITED, it now stamps an AppSetting row
`platform_cooldown:<platform>` with the cooldown expiry (now + 15 min,
PLATFORM_RATE_LIMIT_COOLDOWN_SECONDS). select_due_sources queries every
platform_cooldown:* key at the start of each tick and excludes every
source whose platform is in active cooldown. scheduler_status surfaces
active cooldowns as platform_cooldowns: {platform: expires_iso} so the
TopNav pipeline chip / activity summary can display them.
INSERT...ON CONFLICT DO UPDATE for the upsert so two workers racing
RATE_LIMITED responses on the same platform don't let one's
IntegrityError roll back the other's event-finalize transaction
(stranding the event for the recovery sweep). Atomic at the SQL level.
Tests cover: select_due_sources skips a platform in cooldown; other
platforms unaffected during single-platform cooldown; expired cooldown
rows don't filter; set_platform_cooldown is upsert-safe under repeated
calls.
Operator-flagged 2026-05-30 ("running multiple workers I don't know how
we'd keep the downloader from hitting a rate limit on a source").
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
171c486939 |
refactor(dry-B2): ImportSettings.load()/load_sync() classmethods for the singleton row
The `select(ImportSettings).where(id == 1)).scalar_one()` singleton load was repeated 15× across services, API, and 5 task modules. Added async load() + sync load_sync() classmethods on the model and migrated all 15 full-row sites (callers already imported ImportSettings, so no new imports; dropped download's now-orphaned select import). Left maintenance.py's deliberate column-select (import_scan_path only) as-is. Rest of the service layer was already adequately DRY — the Record/to_dict pattern is only 2 instances and the savepoint find-or-create recovery is correctly per-entity, so neither was forced into a shared abstraction. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
2358cedf3e |
feat(dashboards): scheduler health strip, failing-source rollup, 24h activity sparkline, credential staleness nudge
D1 scheduler visibility: AppSetting last-tick stamp on every Beat tick + GET /api/sources/schedule-status (last_tick_at/next_due_at/due_now/auto_sources) + SchedulerStatusBar on the Subscriptions tab (re-polled every 30s). D2 failing-source rollup: ?failing=true on the sources list + FailingSourcesCard on Downloads with per-source and bulk "retry" (re-runs the feed via /check). D3 activity sparkline: GET /api/downloads/activity hourly buckets + CSS bar chart by the stat chips (failures stacked in error color); refreshes on live poll. D4 credential staleness: surface last_verified age + "re-verify recommended" warning past 30d; also fixes the dead last_verified_at field-name mismatch so the verification row renders at all. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
0493fdc466 |
fc3d: scheduler_service — pure logic for due-source selection + backoff
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |