Worker count flopping every cycle (GPU spikes). Replaces the throughput hill-climb (which probed +1 / reverted −1 every tick on noisy windows) with a GPU-utilization-band controller that HOLDS while util is in a healthy band, growing only on clear spare capacity and shrinking under saturation / memory pressure. EWMA-smoothed util + spaced decisions → steady load.
GPU util/VRAM bars only updating on manual refresh. Adds a dedicated /gpu endpoint (local nvidia-smi, no curator round-trip) polled every 1.5s, and drops the curator queue-status timeout 15s→5s so /status stays snappy.
Agent-only. Fixes two operator-reported issues:
- **Worker count flopping every cycle (GPU spikes).** Replaces the throughput hill-climb (which probed +1 / reverted −1 every tick on noisy windows) with a GPU-utilization-band controller that HOLDS while util is in a healthy band, growing only on clear spare capacity and shrinking under saturation / memory pressure. EWMA-smoothed util + spaced decisions → steady load.
- **GPU util/VRAM bars only updating on manual refresh.** Adds a dedicated `/gpu` endpoint (local nvidia-smi, no curator round-trip) polled every 1.5s, and drops the curator queue-status timeout 15s→5s so `/status` stays snappy.
Ships in the agent image.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Two operator-reported issues with the GPU agent:
1. Worker count flopped almost every cycle, spiking the GPU. The hill-climb
probed +1, judged it over a too-short noisy throughput window, saw no clear
gain and reverted -1 — every tick. Replace it with a GPU-utilization-band
controller: HOLD while smoothed util sits in a healthy band, grow only on
clear spare capacity (util below the low mark + VRAM headroom), shrink under
saturation or memory pressure. Util is EWMA-smoothed and decisions are spaced
(DECIDE_EVERY samples), so a noisy nvidia-smi reading can't move the pool.
Load stays consistent instead of probe/reverting.
2. GPU util/VRAM bars only updated on manual refresh. They rode the /status
poll, which blocks on the curator queue call (slow when curator is busy), so
the meters froze between refreshes. Give them a dedicated /gpu endpoint
(local nvidia-smi only, no curator round-trip) polled every 1.5s, and drop
the curator queue-status timeout 15s -> 5s so /status itself stays snappy.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ttrj5P7upUTueSfoJcxEqa
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Agent-only. Fixes two operator-reported issues:
/gpuendpoint (local nvidia-smi, no curator round-trip) polled every 1.5s, and drops the curator queue-status timeout 15s→5s so/statusstays snappy.Ships in the agent image.
🤖 Generated with Claude Code