Files
FabledCurator/agent
bvandeusenandClaude Opus 5.5 cc53d8db7b
CI and images / lint (push) Successful in 3s
CI and images / extension-version (push) Successful in 3s
CI and images / frontend-build (push) Successful in 27s
CI and images / backend-lint-and-test (push) Successful in 34s
CI and images / integration (push) Successful in 2m22s
CI and images / sign-extension (push) Successful in 4s
CI and images / build-web (push) Successful in 2m46s
CI and images / smoke-web (push) Successful in 50s
CI and images / build-agent (push) Successful in 6m35s
CI and images / promote (push) Successful in 2s
feat: a GPU agent on the CPU shows as degraded, in the System view and on its own page (4410)
torch and onnxruntime both fall back to the CPU without raising, so the agent
that ran CPU-bound for weeks after a driver update leased and checked in like
a healthy one.

- The agent sends its startup accel report on every lease and heartbeat.
- The server keeps a bounded copy on the roster row. A running agent with a
  runtime off the GPU becomes `degraded`, with a sentence naming the runtime
  and the reason.
- The top nav shows it amber.
- The agent page carries a banner, and its pill reads "CPU only".

Also: the bandwidth field gets the page's − / + stepper, and both number
fields drop the browser's spin arrows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-24 19:09:07 -04:00
..

FabledCurator GPU agent

A desktop-GPU worker that embeds characters (CCIP) + figure crops for FabledCurator. It talks to FC only over HTTP — it leases jobs, fetches image pixels, runs the models on your GPU, and posts results back. Your FC database and Redis stay private; the agent never touches them.

You run it when you want a burst and stop it to reclaim the card.

0. Host prerequisite — NVIDIA Container Toolkit

Docker needs the toolkit to hand the GPU to a container (else: "could not select device driver nvidia with capabilities gpu"). On Arch/CachyOS:

sudo pacman -S nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
# verify:
docker run --rm --gpus all nvidia/cuda:13.0.3-base-ubuntu24.04 nvidia-smi
# the header's CUDA version must be 13.0 or later (driver 580+)

After a driver update: regenerate the CDI spec

If the agent's first log lines say accel: torch is NOT on the GPU or report cudaGetDeviceCount: unknown error (999) while nvidia-smi still works, the toolkit's saved device list (/etc/cdi/nvidia.yaml) is out of date. The nvidia-uvm device number changes between driver versions, and a spec generated before the update hands the container a device node that no longer exists (2026-09-24: host 511,0, container 235,0). Compare ls -l /dev/nvidia-uvm on the host with the same inside the container, then:

sudo nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml
# if your toolkit ships it, this keeps it current on every driver update:
sudo systemctl enable --now nvidia-cdi-refresh.path

1. Get a token

In FC: Settings → Tagging → GPU agent → Generate token (or Rotate). Copy it.

2. Pull (CI publishes it alongside the web image)

docker pull git.fabledsword.com/bvandeusen/fabledcurator-agent:latest

Local build for development instead: docker build -t fc-gpu-agent agent/

3. Run (on the machine with the GPU)

docker run --rm --gpus all -p 8770:8770 \
  -e FC_URL=http://curator.traefik.internal \
  -e FC_TOKEN=<paste-the-token> \
  -v fc-agent-models:/models \
  git.fabledsword.com/bvandeusen/fabledcurator-agent:latest

Then open http://localhost:8770 — the control page. Click Start to begin draining the queue; Pause/Stop to yield the GPU. The -v fc-agent-models volume caches the downloaded ONNX models so restarts are fast.

Kick off a backfill from FC (GPU agent card → Queue character embedding), then watch the queue counts on the control page (or FC's card) drain.

Config (env)

var default meaning
FC_URL http://localhost:8000 FC base URL
FC_TOKEN — the bearer token (required)
AGENT_ID desktop-agent identifies this agent's leases
BATCH_SIZE 4 jobs leased per round (still processed one at a time)
CCIP_MODEL imgutils default CCIP model name
DETECTOR_LEVEL m person-detector size: n < s < m < x
POLL_IDLE_SECONDS 10 wait between empty leases

⚠️ Verify on first run

This part can't be CI-tested (no GPU/models in CI), so confirm against your installed dghs-imgutils (pip show dghs-imgutils) — see fc_agent/models.py:

  • imgutils.detect.detect_person(image, level=...) returns [((x0,y0,x1,y1), label, score), ...].
  • imgutils.metrics.ccip_extract_feature(image, model=...) returns a vector (768-d for caformer). If you want the F1-0.94 variant, set CCIP_MODEL=ccip-caformer_b36-24 (verify the exact string in imgutils).

If FC's matcher under/over-fires, tune the cosine threshold in backend/app/services/ml/ccip.py (DEFAULT_SIM_THRESHOLD) and use GET /api/ccip/overview + /api/ccip/images/<id> to spot-check.

CPU fallback

Swap onnxruntime-gpu → onnxruntime in requirements.txt and drop --gpus all to grind it slowly on the server instead. Same agent, no card.