95d2ae1agent bandwidth governor: one aggregate MB/s budget (default 8; live dial + net readout; ffmpeg governed via /proc + SIGSTOP) so the agent can't saturate the desktop's network — the root cause behind the serving slowness measurements.
sample_frames_from_url now returns (frames, reason) — reason carries the
SPECIFIC cause on failure (ffmpeg's stderr tail, e.g. "moov atom not found",
or the timeout) instead of only logging it agent-side. The worker folds it
into the failure it reports, so curator's GpuJob.error reads e.g.
no frames sampled from video — ffmpeg exit 183: moov atom not found ...
instead of the bare "(unprocessable)". The errored-jobs list becomes
self-describing: after a retry sweep, surviving errors name their real
defect without needing the agent log. Return-value plumbing (not shared
state) so concurrent downloaders stay isolated. Agent VERSION → 2026-07-02.2.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
GpuAgentCard was hardcoded :open=true, HeadsCard opened whenever any head
existed, TagEvalCard whenever a persisted run existed — so a fresh Settings
load greeted the operator with several tiles already expanded. All three now
force-open only while their task is actually running (the #877 resurface
behavior on the busy-driven tiles is untouched).
MaintenanceTile additionally persists MANUAL expand/collapse per tile in
localStorage, so the section reloads the way the operator left it; a forced
open while a task runs stays transient and is never saved as a preference.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM
The hourly ccip backfill's skip-list lacked 'error' (and the daily
siglip/embed variants re-gated failures on their missing results), so every
permanently-bad file got a fresh doomed job each run — ~24 duplicate error
rows/day per file, the perpetual 'unprocessable' flood. An errored job is now
a TOMBSTONE: no backfill re-enqueues it; retry is deliberate-only via
/retry_errors (an errored back-catalogue needs one button press after a
model swap).
One shared set of dedupe DELETEs (services/ml/gpu_jobs.error_dedupe_statements)
runs before every backfill and inside /retry_errors: error rows made moot by a
later pending/leased/done row go first, then older duplicates (newest reason
survives) — so the error count reads as distinct failing files and a retry
can't fan one file out into duplicate pending jobs. /retry_errors now returns
{requeued, pruned} and the toast shows both.
Poison-loop guards (release and lease-expiry burn no attempts, so a job that
stalls its transfer or crashes the agent every time cycled forever —
operator-observed jobs 99044/125288/131594/143131):
- agent: 3 in-session transient bounces (fetch or submit) → fail with the real
reason instead of another release; strikes never count while stopping, and
clear on submit success. Agent build 2026-07-02.3.
- server: the 60s orphan sweep (statements shared between the beat task and
GpuJobService so they can't drift) converts expired leases with >=5 lease
grants and pending jobs with >=10 to 'error', preserving the last stored
failure reason. Backstops old agent builds.
Tests: tombstone rule across all three backfill variants, moot-row pruning,
poison conversions, and the extended /retry_errors dedupe contract.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM
One shared TokenBucket (default 8 MB/s; BANDWIDTH_LIMIT_MB_S, 0 = unlimited;
live MB/s dial + net readout in the control UI) is charged by every still
download (streamed chunk reads) and every ffmpeg video stream (metered from
outside via /proc/<pid>/io and SIGSTOP/SIGCONTed into budget).
Why: D1 re-measurement 2026-07-02 — the idle link moves ~38 MB/s, but 8
unthrottled downloaders bufferbloated it to ~1-1.5 MB/s PER STREAM (operator's
browser included). Capping the aggregate keeps the desktop usable and still
beats the collapsed sweep throughput it replaces. Agent build 2026-07-02.4.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM
- ml-backfill-daily: the CPU tag_and_embed backfill raced the GPU agent's
daily embed backfill for the same NULL-embedding images at ~100x the cost
(B1 audit verdict, milestone #124). The backfill TASK stays — the manual
/api/ml/backfill button remains the deliberate CPU fallback pending B3.
- purge-legacy: one-time IR-migration cleanup, dry-run verified 0 targets on
the live library before removal (A2 audit, milestone #123). Fully retired
per rule 22: tile, store action, route, service fn, tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM
An errored GPU job's stored reason is a suspicion; the file probe is the
verdict. A 15-min beat sweep (triage_gpu_errors) runs verify_integrity's own
probe (sha256 + decode) on each errored image ONCE and writes both verdicts:
ImageRecord.integrity_status and the new GpuJob.triage_status ('defect' |
'file_ok', migration 0072). Every classification logs at WARNING so it
surfaces in Logs/System Activity.
- 'defect' rows are excluded from /retry_errors (re-running a known-bad file
burns agent time re-minting the tombstone); response now reports
defects_kept and the GpuAgentCard toast says so.
- GET /api/gpu/errors: triage view — reason buckets (classify_reason),
probe verdicts, per-job detail. POST /errors/triage runs the sweep now.
- POST /api/gpu/errors/<id>/recover: reuses the Layer-2 refetch pattern —
delete the defective copy + record (full cascade takes the tombstones too)
and re-poll its subscription Source so a fresh copy re-imports and re-enters
the pipeline; 'no_source' when nothing pollable resolves.
- New 'Failed processing' card (GpuTriageCard) in Maintenance: verdict counts,
reason summary, probe-now, defect list with thumbnails + per-image Recover.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM
The head-vs-centroid eval (#1130) existed to prove the 'frozen embedding +
trained head' spine; the operator accepted the tagging system and dropped the
harness. Removed per rule 22: TagEvalCard + store, /api/tag_eval blueprint,
tag_eval_run ml task, recover-stalled-tag-eval-runs sweep + beat entry,
TagEvalRun model + table (migration 0073), and its tests.
The eval's data loaders + metric helpers were NOT eval-specific — the nightly
heads trainer runs on them — so they moved verbatim to
services/ml/training_data.py (heads.py import updated; behavior unchanged).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Eight CI-green commits from dev (runs 1854–1856, 1858):
3d62017collapse-by-default: maintenance tiles start closed; manual open/close persists per tile; in-flight tasks still force-open.1b1d373error-tombstone semantics: backfills skip errored; duplicate/moot tombstones self-prune; poison-loop guards agent- and server-side.09e2772agent build 2026-07-02.3 (poison guard).31c416bbeat-comment hygiene.95d2ae1agent bandwidth governor: one aggregate MB/s budget (default 8; live dial + net readout; ffmpeg governed via /proc + SIGSTOP) so the agent can't saturate the desktop's network — the root cause behind the serving slowness measurements.1f27189retireml-backfill-dailybeat + the spent purge-legacy action (operator-approved; dry-run verified 0 targets).a7abcc4failure triage → integrity → recovery (#125): probe sweep writes file verdicts (migration 0072); defects excluded from retry;/api/gpu/errors+ recover endpoint; "Failed processing" card.eaea430retire the tag-eval harness (operator-approved; migration 0073); its shared helpers live on intraining_data.pyfor the heads trainer.Deploy notes: migrations 0072 (add column) + 0073 (drop table) run on web boot; agent image re-pull → build 2026-07-02.4.
🤖 Generated with Claude Code
https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM
The hourly ccip backfill's skip-list lacked 'error' (and the daily siglip/embed variants re-gated failures on their missing results), so every permanently-bad file got a fresh doomed job each run — ~24 duplicate error rows/day per file, the perpetual 'unprocessable' flood. An errored job is now a TOMBSTONE: no backfill re-enqueues it; retry is deliberate-only via /retry_errors (an errored back-catalogue needs one button press after a model swap). One shared set of dedupe DELETEs (services/ml/gpu_jobs.error_dedupe_statements) runs before every backfill and inside /retry_errors: error rows made moot by a later pending/leased/done row go first, then older duplicates (newest reason survives) — so the error count reads as distinct failing files and a retry can't fan one file out into duplicate pending jobs. /retry_errors now returns {requeued, pruned} and the toast shows both. Poison-loop guards (release and lease-expiry burn no attempts, so a job that stalls its transfer or crashes the agent every time cycled forever — operator-observed jobs 99044/125288/131594/143131): - agent: 3 in-session transient bounces (fetch or submit) → fail with the real reason instead of another release; strikes never count while stopping, and clear on submit success. Agent build 2026-07-02.3. - server: the 60s orphan sweep (statements shared between the beat task and GpuJobService so they can't drift) converts expired leases with >=5 lease grants and pending jobs with >=10 to 'error', preserving the last stored failure reason. Backstops old agent builds. Tests: tombstone rule across all three backfill variants, moot-row pruning, poison conversions, and the extended /retry_errors dedupe contract. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnMAn errored GPU job's stored reason is a suspicion; the file probe is the verdict. A 15-min beat sweep (triage_gpu_errors) runs verify_integrity's own probe (sha256 + decode) on each errored image ONCE and writes both verdicts: ImageRecord.integrity_status and the new GpuJob.triage_status ('defect' | 'file_ok', migration 0072). Every classification logs at WARNING so it surfaces in Logs/System Activity. - 'defect' rows are excluded from /retry_errors (re-running a known-bad file burns agent time re-minting the tombstone); response now reports defects_kept and the GpuAgentCard toast says so. - GET /api/gpu/errors: triage view — reason buckets (classify_reason), probe verdicts, per-job detail. POST /errors/triage runs the sweep now. - POST /api/gpu/errors/<id>/recover: reuses the Layer-2 refetch pattern — delete the defective copy + record (full cascade takes the tombstones too) and re-poll its subscription Source so a fresh copy re-imports and re-enters the pipeline; 'no_source' when nothing pollable resolves. - New 'Failed processing' card (GpuTriageCard) in Maintenance: verdict counts, reason summary, probe-now, defect list with thumbnails + per-image Recover. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CDgx8bQS5YrGRK76v8HUnM