torch and onnxruntime both fall back to the CPU without raising, so the agent that ran CPU-bound for weeks after a driver update leased and checked in like a healthy one. - The agent sends its startup accel report on every lease and heartbeat. - The server keeps a bounded copy on the roster row. A running agent with a runtime off the GPU becomes `degraded`, with a sentence naming the runtime and the reason. - The top nav shows it amber. - The agent page carries a banner, and its pill reads "CPU only". Also: the bandwidth field gets the page's − / + stepper, and both number fields drop the browser's spin arrows. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
FabledCurator GPU agent
A desktop-GPU worker that embeds characters (CCIP) + figure crops for FabledCurator. It talks to FC only over HTTP — it leases jobs, fetches image pixels, runs the models on your GPU, and posts results back. Your FC database and Redis stay private; the agent never touches them.
You run it when you want a burst and stop it to reclaim the card.
0. Host prerequisite — NVIDIA Container Toolkit
Docker needs the toolkit to hand the GPU to a container (else: "could not select device driver nvidia with capabilities gpu"). On Arch/CachyOS:
sudo pacman -S nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
# verify:
docker run --rm --gpus all nvidia/cuda:13.0.3-base-ubuntu24.04 nvidia-smi
# the header's CUDA version must be 13.0 or later (driver 580+)
After a driver update: regenerate the CDI spec
If the agent's first log lines say accel: torch is NOT on the GPU or report
cudaGetDeviceCount: unknown error (999) while nvidia-smi still works, the
toolkit's saved device list (/etc/cdi/nvidia.yaml) is out of date. The
nvidia-uvm device number changes between driver versions, and a spec
generated before the update hands the container a device node that no longer
exists (2026-09-24: host 511,0, container 235,0). Compare
ls -l /dev/nvidia-uvm on the host with the same inside the container, then:
sudo nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml
# if your toolkit ships it, this keeps it current on every driver update:
sudo systemctl enable --now nvidia-cdi-refresh.path
1. Get a token
In FC: Settings → Tagging → GPU agent → Generate token (or Rotate). Copy it.
2. Pull (CI publishes it alongside the web image)
docker pull git.fabledsword.com/bvandeusen/fabledcurator-agent:latest
Local build for development instead:
docker build -t fc-gpu-agent agent/
3. Run (on the machine with the GPU)
docker run --rm --gpus all -p 8770:8770 \
-e FC_URL=http://curator.traefik.internal \
-e FC_TOKEN=<paste-the-token> \
-v fc-agent-models:/models \
git.fabledsword.com/bvandeusen/fabledcurator-agent:latest
Then open http://localhost:8770 — the control page. Click Start to begin
draining the queue; Pause/Stop to yield the GPU. The -v fc-agent-models
volume caches the downloaded ONNX models so restarts are fast.
Kick off a backfill from FC (GPU agent card → Queue character embedding), then watch the queue counts on the control page (or FC's card) drain.
Config (env)
| var | default | meaning |
|---|---|---|
FC_URL |
http://localhost:8000 |
FC base URL |
FC_TOKEN |
— | the bearer token (required) |
AGENT_ID |
desktop-agent |
identifies this agent's leases |
BATCH_SIZE |
4 |
jobs leased per round (still processed one at a time) |
CCIP_MODEL |
imgutils default | CCIP model name |
DETECTOR_LEVEL |
m |
person-detector size: n < s < m < x |
POLL_IDLE_SECONDS |
10 |
wait between empty leases |
⚠️ Verify on first run
This part can't be CI-tested (no GPU/models in CI), so confirm against your
installed dghs-imgutils (pip show dghs-imgutils) — see fc_agent/models.py:
imgutils.detect.detect_person(image, level=...)returns[((x0,y0,x1,y1), label, score), ...].imgutils.metrics.ccip_extract_feature(image, model=...)returns a vector (768-d for caformer). If you want the F1-0.94 variant, setCCIP_MODEL=ccip-caformer_b36-24(verify the exact string in imgutils).
If FC's matcher under/over-fires, tune the cosine threshold in
backend/app/services/ml/ccip.py (DEFAULT_SIM_THRESHOLD) and use
GET /api/ccip/overview + /api/ccip/images/<id> to spot-check.
CPU fallback
Swap onnxruntime-gpu → onnxruntime in requirements.txt and drop --gpus all
to grind it slowly on the server instead. Same agent, no card.