Files
FabledCurator/agent
bvandeusenandClaude Opus 5 693759f2bb
CI and images / lint (push) Successful in 3s
CI and images / extension-version (push) Successful in 2s
CI and images / frontend-build (push) Successful in 21s
CI and images / backend-lint-and-test (push) Successful in 31s
CI and images / integration (push) Failing after 2m10s
CI and images / sign-extension (push) Skipped
CI and images / build-web (push) Skipped
CI and images / smoke-web (push) Skipped
CI and images / promote (push) Skipped
CI and images / build-agent (push) Skipped
fix: an idle GPU agent could not check in, so the roster called it stopped
Operator, 2026-09-23: *"I'm running the gpu agent on my device and it
currently reads as 'offline' but it's running and has checked in recently."*

It had checked in — twelve minutes ago. Two cadences that never agreed:

    idle lease poll ceiling   900s   agent/fc_agent/worker.py (sleep mode)
    heartbeat while idle      never  gated on holding leases
    roster "stopped" after    300s   api/system_health.py

The roster records an agent check-in on `lease` and `heartbeat`. The heartbeat
loop was gated on `if ids:`, so an agent holding no leases sent nothing at
all — leaving the lease poll as the only check-in, and sleep mode backs that
off exponentially to a 900s ceiling. 900 against 300: an IDLE agent was
structurally guaranteed to read as stopped. Nothing was broken; nothing was
misconfigured; the two halves simply disagreed.

Not a recent regression. Sleep mode landed 2026-07-02; the roster adopted the
lease as its check-in on 2026-09-02 — *"A lease IS the check-in … Recorded on
the call that was already happening"* — without noticing that the call it was
piggybacking on had been deliberately slowed ten weeks earlier.

The heartbeat now sends whether or not it holds leases. An empty one extends
nothing (`id.in_([])` matches no rows) and costs one small POST every 45s —
against the 6/min lease poll sleep mode exists to avoid, that is not a cadence
worth protecting, and it is what makes "is the agent alive" answerable at all.

Still gated on `self._running`: a worker that has been stopped is not checking
in for work, and reporting it as present would be a different lie.

Two things I could NOT determine from the code, both needing the live table:
whether a stale `agent:agent` row exists from an older build that omitted
`agent_id` (the server defaults it), and whether changing `AGENT_ID` has ever
stranded an abandoned row — nothing prunes `service_seen`, so either would sit
there reading "stopped" forever.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
2026-09-23 14:52:36 -04:00
..

FabledCurator GPU agent

A desktop-GPU worker that embeds characters (CCIP) + figure crops for FabledCurator. It talks to FC only over HTTP — it leases jobs, fetches image pixels, runs the models on your GPU, and posts results back. Your FC database and Redis stay private; the agent never touches them.

You run it when you want a burst and stop it to reclaim the card.

0. Host prerequisite — NVIDIA Container Toolkit

Docker needs the toolkit to hand the GPU to a container (else: "could not select device driver nvidia with capabilities gpu"). On Arch/CachyOS:

sudo pacman -S nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
# verify:
docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi

1. Get a token

In FC: Settings → Tagging → GPU agent → Generate token (or Rotate). Copy it.

2. Pull (CI publishes it alongside the web image)

docker pull git.fabledsword.com/bvandeusen/fabledcurator-agent:latest

Local build for development instead: docker build -t fc-gpu-agent agent/

3. Run (on the machine with the GPU)

docker run --rm --gpus all -p 8770:8770 \
  -e FC_URL=http://curator.traefik.internal \
  -e FC_TOKEN=<paste-the-token> \
  -v fc-agent-models:/models \
  git.fabledsword.com/bvandeusen/fabledcurator-agent:latest

Then open http://localhost:8770 — the control page. Click Start to begin draining the queue; Pause/Stop to yield the GPU. The -v fc-agent-models volume caches the downloaded ONNX models so restarts are fast.

Kick off a backfill from FC (GPU agent card → Queue character embedding), then watch the queue counts on the control page (or FC's card) drain.

Config (env)

var default meaning
FC_URL http://localhost:8000 FC base URL
FC_TOKEN the bearer token (required)
AGENT_ID desktop-agent identifies this agent's leases
BATCH_SIZE 4 jobs leased per round (still processed one at a time)
CCIP_MODEL imgutils default CCIP model name
DETECTOR_LEVEL m person-detector size: n < s < m < x
POLL_IDLE_SECONDS 10 wait between empty leases

⚠️ Verify on first run

This part can't be CI-tested (no GPU/models in CI), so confirm against your installed dghs-imgutils (pip show dghs-imgutils) — see fc_agent/models.py:

  • imgutils.detect.detect_person(image, level=...) returns [((x0,y0,x1,y1), label, score), ...].
  • imgutils.metrics.ccip_extract_feature(image, model=...) returns a vector (768-d for caformer). If you want the F1-0.94 variant, set CCIP_MODEL=ccip-caformer_b36-24 (verify the exact string in imgutils).

If FC's matcher under/over-fires, tune the cosine threshold in backend/app/services/ml/ccip.py (DEFAULT_SIM_THRESHOLD) and use GET /api/ccip/overview + /api/ccip/images/<id> to spot-check.

CPU fallback

Swap onnxruntime-gpuonnxruntime in requirements.txt and drop --gpus all to grind it slowly on the server instead. Same agent, no card.