Compare commits
10
Commits
d01f33dea6
...
dev
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
ad8392b790 | ||
|
|
5084ba666b | ||
|
|
fe4e0f2b71 | ||
|
|
dc8af8b1a7 | ||
|
|
131237143b | ||
|
|
59d27ef76e | ||
|
|
f630e50e75 | ||
|
|
86abaf0b94 | ||
|
|
4815040d74 | ||
|
|
81b7b6f308 |
@@ -34,6 +34,34 @@ DB_PASSWORD=
|
|||||||
SECRET_KEY=
|
SECRET_KEY=
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# FIRST BOOT ONLY — then delete this line
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
# FabledCurator encrypts your stored platform credentials with a Fernet key it
|
||||||
|
# keeps at /images/secrets/credential_key.b64 — inside the ./images bind mount,
|
||||||
|
# so it outlives the container. On a brand-new install that file does not exist
|
||||||
|
# yet, and the app REFUSES TO START rather than quietly create one:
|
||||||
|
#
|
||||||
|
# MissingCredentialKey: Fernet key file not found at
|
||||||
|
# /images/secrets/credential_key.b64
|
||||||
|
#
|
||||||
|
# That refusal is deliberate. Auto-creating a key is indistinguishable from the
|
||||||
|
# disaster case — a restore that brought the database back but lost
|
||||||
|
# ./images/secrets — and there it would mint a key that cannot decrypt anything,
|
||||||
|
# leaving an instance that looks healthy while every paywalled download fails.
|
||||||
|
# So the choice is yours to make explicitly, once.
|
||||||
|
#
|
||||||
|
# Set this for your first `up`, watch the container come up, then DELETE THE
|
||||||
|
# LINE. Leaving it set disarms the protection permanently, on an instance that
|
||||||
|
# by then has credentials worth protecting.
|
||||||
|
#
|
||||||
|
# BACK UP ./images/secrets/ ALONGSIDE YOUR DATABASE. The key is the only thing
|
||||||
|
# that can read your stored credentials; a database restored without it needs
|
||||||
|
# every credential re-entered by hand.
|
||||||
|
CURATOR_BOOTSTRAP_NEW_KEY=1
|
||||||
|
|
||||||
|
|
||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
# Optional — defaults are fine
|
# Optional — defaults are fine
|
||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
|
|||||||
+177
-233
@@ -452,6 +452,18 @@ jobs:
|
|||||||
# to. Same source of truth; no double-store.
|
# to. Same source of truth; no double-store.
|
||||||
|
|
||||||
build-web:
|
build-web:
|
||||||
|
# Consumed by smoke-web's job-level `if:`. It cannot read `env` — the env
|
||||||
|
# context is available to STEP `if:` and step bodies, never to a job's own
|
||||||
|
# condition, and an unresolvable context there is empty rather than an
|
||||||
|
# error. `smoke-web` skipped silently on run 5290 for exactly that reason.
|
||||||
|
#
|
||||||
|
# Keying off the reuse step's own output is better than re-deriving the
|
||||||
|
# trigger anyway: it is the same single decision the build, the XPI
|
||||||
|
# download and the promote all take (build.yml's "one decision drives
|
||||||
|
# everything downstream"), and it says the thing smoke-web actually needs
|
||||||
|
# to know — a candidate was published — rather than restating why.
|
||||||
|
outputs:
|
||||||
|
candidate: ${{ steps.reuse.outputs.promote }}
|
||||||
# A plain `needs` — no `always()`. That expression existed to let a
|
# A plain `needs` — no `always()`. That expression existed to let a
|
||||||
# SKIPPED sign-extension through on a tag push while still blocking a
|
# SKIPPED sign-extension through on a tag push while still blocking a
|
||||||
# FAILED one. With no tag trigger, sign-extension always runs, so the
|
# FAILED one. With no tag trigger, sign-extension always runs, so the
|
||||||
@@ -722,10 +734,16 @@ jobs:
|
|||||||
# already allows for :buildcache, not the per-build tag family that
|
# already allows for :buildcache, not the per-build tag family that
|
||||||
# milestone 318 withdrew.
|
# milestone 318 withdrew.
|
||||||
#
|
#
|
||||||
# Both values are decided HERE, beside `hit`, for the reason the
|
# Decided HERE, beside `hit`, for the reason the force/schedule
|
||||||
# force/schedule branch below gives: one step decides what this job
|
# branch below gives: one step decides what this job does. A
|
||||||
# does. A promote condition derived independently could disagree with
|
# condition derived independently could disagree with the tag the
|
||||||
# the tag the build actually wrote.
|
# build actually wrote.
|
||||||
|
#
|
||||||
|
# build-web additionally exposes this as `outputs.candidate`, which is
|
||||||
|
# what gates the `promote` job — a job's `if:` cannot read `env`, and
|
||||||
|
# one flag is enough because all three derive it from the same
|
||||||
|
# IS_REFRESH. ml and agent do not re-emit it; a second copy nothing
|
||||||
|
# reads is the kind of thing that later reads as load-bearing.
|
||||||
if [ "${IS_REFRESH:-}" = "true" ]; then
|
if [ "${IS_REFRESH:-}" = "true" ]; then
|
||||||
echo "build_ref=$IMAGE:refresh-candidate" >> "$GITHUB_OUTPUT"
|
echo "build_ref=$IMAGE:refresh-candidate" >> "$GITHUB_OUTPUT"
|
||||||
echo "promote=true" >> "$GITHUB_OUTPUT"
|
echo "promote=true" >> "$GITHUB_OUTPUT"
|
||||||
@@ -947,77 +965,6 @@ jobs:
|
|||||||
FC_CHANNEL=${{ steps.tag.outputs.channel }}
|
FC_CHANNEL=${{ steps.tag.outputs.channel }}
|
||||||
FC_VERSION=${{ steps.reuse.outputs.version }}
|
FC_VERSION=${{ steps.reuse.outputs.version }}
|
||||||
|
|
||||||
# Point the channel tag at the candidate the refresh just built.
|
|
||||||
#
|
|
||||||
# Unconditional TODAY, so this milestone never leaves the refresh in a
|
|
||||||
# state where it builds and publishes nothing. Step 4 wraps it in the
|
|
||||||
# smoke suite's verdict; until then the scheduled path behaves exactly
|
|
||||||
# as it did, just via two operations instead of one.
|
|
||||||
#
|
|
||||||
# NOT `imagetools create`. That wraps its source in an INDEX, and an
|
|
||||||
# indexed channel tag is the one thing this pipeline cannot survive:
|
|
||||||
# `.Image.Config.Labels` does not resolve through an index, so the
|
|
||||||
# fc.revision the reuse check reads off the channel tag would come back
|
|
||||||
# empty, every subsequent push would miss and rebuild, and nothing would
|
|
||||||
# go red. That is #3183, observed on run 4751 — reuse worked exactly once
|
|
||||||
# and the only symptom was the bill. The repoint step below excludes its
|
|
||||||
# own source tag for precisely this reason; a promote that re-introduced
|
|
||||||
# the wrap through a different door would undo that care.
|
|
||||||
#
|
|
||||||
# A manifest PUT is what "make this tag name that image" means at the
|
|
||||||
# registry level: the same bytes under the same media type, so the digest
|
|
||||||
# is identical, the media type is preserved, and no layer moves.
|
|
||||||
- name: Promote the refresh candidate to the channel
|
|
||||||
if: steps.reuse.outputs.promote == 'true'
|
|
||||||
env:
|
|
||||||
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator
|
|
||||||
CHANNEL_REF: ${{ steps.reuse.outputs.channel_ref }}
|
|
||||||
TOKEN: ${{ secrets.RELEASE_TOKEN }}
|
|
||||||
ACTOR: ${{ github.actor }}
|
|
||||||
run: |
|
|
||||||
set -eu
|
|
||||||
REPO=${IMAGE#git.fabledsword.com/}
|
|
||||||
TAG=${CHANNEL_REF##*:}
|
|
||||||
|
|
||||||
# Registry auth is its own token exchange — the `docker login` above
|
|
||||||
# authenticates the docker client, not curl. Deadline on every call
|
|
||||||
# (rule 156): a registry that stops answering must fail this step,
|
|
||||||
# not hang the weekly refresh until the job times out.
|
|
||||||
BEARER=$(curl -fsS --max-time 30 -u "$ACTOR:$TOKEN" \
|
|
||||||
"https://git.fabledsword.com/v2/token?scope=repository:$REPO:pull,push&service=git.fabledsword.com" \
|
|
||||||
| python3 -c 'import sys,json; print(json.load(sys.stdin)["token"])')
|
|
||||||
|
|
||||||
# Ask for the image manifest media types ONLY. Offering the index
|
|
||||||
# types too would let the registry hand back an index if one ever
|
|
||||||
# existed at this tag, and we would faithfully copy the thing we are
|
|
||||||
# trying not to create.
|
|
||||||
ACCEPT='application/vnd.oci.image.manifest.v1+json, application/vnd.docker.distribution.manifest.v2+json'
|
|
||||||
CT=$(curl -fsS --max-time 60 -o manifest.json -D headers.txt \
|
|
||||||
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
|
|
||||||
"https://git.fabledsword.com/v2/$REPO/manifests/refresh-candidate" \
|
|
||||||
&& tr -d '\r' < headers.txt | awk -F': ' '/^[Cc]ontent-[Tt]ype:/{print $2}')
|
|
||||||
test -n "$CT"
|
|
||||||
SRC_DIGEST=$(tr -d '\r' < headers.txt | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
|
|
||||||
echo "promote: candidate is $SRC_DIGEST ($CT)"
|
|
||||||
|
|
||||||
curl -fsS --max-time 120 -X PUT \
|
|
||||||
-H "Authorization: Bearer $BEARER" -H "Content-Type: $CT" \
|
|
||||||
--data-binary @manifest.json \
|
|
||||||
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG"
|
|
||||||
|
|
||||||
# Read it back. A PUT that returned 2xx but landed something else is
|
|
||||||
# exactly the silent-and-plausible failure this pipeline keeps
|
|
||||||
# producing, and the check costs one request.
|
|
||||||
NOW=$(curl -fsS --max-time 30 -o /dev/null -D - \
|
|
||||||
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
|
|
||||||
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG" \
|
|
||||||
| tr -d '\r' | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
|
|
||||||
if [ "$NOW" != "$SRC_DIGEST" ]; then
|
|
||||||
echo "promote: $IMAGE:$TAG is $NOW, expected $SRC_DIGEST" >&2
|
|
||||||
exit 1
|
|
||||||
fi
|
|
||||||
echo "promote: $IMAGE:$TAG now names $NOW"
|
|
||||||
|
|
||||||
# Every tag but the channel's own is written HERE, registry-side,
|
# Every tag but the channel's own is written HERE, registry-side,
|
||||||
# whether or not a build ran. Each -t becomes another reference to the
|
# whether or not a build ran. Each -t becomes another reference to the
|
||||||
# SAME manifest the channel tag holds, so :c-<sha> is byte-identical to
|
# SAME manifest the channel tag holds, so :c-<sha> is byte-identical to
|
||||||
@@ -1123,8 +1070,8 @@ jobs:
|
|||||||
# time anyone saw this job run would also be the first time it could stop a
|
# time anyone saw this job run would also be the first time it could stop a
|
||||||
# publish.
|
# publish.
|
||||||
smoke-web:
|
smoke-web:
|
||||||
if: env.IS_REFRESH == 'true'
|
|
||||||
needs: [build-web]
|
needs: [build-web]
|
||||||
|
if: needs.build-web.outputs.candidate == 'true'
|
||||||
runs-on: python-ci
|
runs-on: python-ci
|
||||||
container:
|
container:
|
||||||
image: git.fabledsword.com/bvandeusen/ci-python:3.14
|
image: git.fabledsword.com/bvandeusen/ci-python:3.14
|
||||||
@@ -1201,6 +1148,14 @@ jobs:
|
|||||||
ENVOPTS="$ENVOPTS -e DB_PORT=5432 -e DB_NAME=$DB_NAME -e SECRET_KEY=$SECRET_KEY"
|
ENVOPTS="$ENVOPTS -e DB_PORT=5432 -e DB_NAME=$DB_NAME -e SECRET_KEY=$SECRET_KEY"
|
||||||
ENVOPTS="$ENVOPTS -e CELERY_BROKER_URL=redis://$RD_IP:6379/0"
|
ENVOPTS="$ENVOPTS -e CELERY_BROKER_URL=redis://$RD_IP:6379/0"
|
||||||
ENVOPTS="$ENVOPTS -e CELERY_RESULT_BACKEND=redis://$RD_IP:6379/0"
|
ENVOPTS="$ENVOPTS -e CELERY_RESULT_BACKEND=redis://$RD_IP:6379/0"
|
||||||
|
# A throwaway CI instance IS first-time setup, which is the one case
|
||||||
|
# credential_crypto allows a key to be minted in. Without it the web
|
||||||
|
# role refuses to boot — deliberately, since silently generating a
|
||||||
|
# key on a restored-DB-but-lost-secrets deployment would leave every
|
||||||
|
# Credential row undecryptable (the 2026-06-02 audit). Discovered by
|
||||||
|
# this job on its first real run; see #3422 for the fact that no
|
||||||
|
# user-facing file mentions this variable at all.
|
||||||
|
ENVOPTS="$ENVOPTS -e CURATOR_BOOTSTRAP_NEW_KEY=1"
|
||||||
|
|
||||||
# 1. The schema builds from empty, using the image's OWN libpq and
|
# 1. The schema builds from empty, using the image's OWN libpq and
|
||||||
# psycopg. This is the same call entrypoint.sh makes before it
|
# psycopg. This is the same call entrypoint.sh makes before it
|
||||||
@@ -1235,18 +1190,141 @@ jobs:
|
|||||||
healthy=1
|
healthy=1
|
||||||
break
|
break
|
||||||
fi
|
fi
|
||||||
|
# A container that has EXITED will never answer, so stop asking.
|
||||||
|
# Without this the loop spent 3m35s polling a dead container on
|
||||||
|
# this job's first run, and — because docker recycles the IP — got
|
||||||
|
# a confusing mix of connection-refused and 5s timeouts from
|
||||||
|
# whatever took the address next. The trap's log dump had the real
|
||||||
|
# answer the whole time; this just stops burying it.
|
||||||
|
if [ "$(docker inspect -f '{{.State.Running}}' "$CID" 2>/dev/null)" != "true" ]; then
|
||||||
|
echo "smoke: FAILED — the web container exited during boot." >&2
|
||||||
|
echo "smoke: its log follows; entrypoint runs alembic BEFORE" >&2
|
||||||
|
echo "smoke: serving, so a startup exception lands here." >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
sleep 2
|
sleep 2
|
||||||
done
|
done
|
||||||
if [ -z "$healthy" ]; then
|
if [ -z "$healthy" ]; then
|
||||||
echo "smoke: FAILED — web did not answer /api/health within 120s." >&2
|
# 60 iterations of (up to 5s connect + 2s sleep) — up to ~7min, not
|
||||||
echo "smoke: entrypoint runs alembic BEFORE serving, and step 1" >&2
|
# the 120s an earlier version of this message claimed.
|
||||||
echo "smoke: passed, so look at hypercorn and the python base." >&2
|
echo "smoke: FAILED — web is running but never answered" >&2
|
||||||
|
echo "smoke: /api/health. It is up, so look at hypercorn and the" >&2
|
||||||
|
echo "smoke: python base rather than at startup." >&2
|
||||||
exit 1
|
exit 1
|
||||||
fi
|
fi
|
||||||
curl -fsS --max-time 5 "http://$WEB_IP:8080/api/health"
|
curl -fsS --max-time 5 "http://$WEB_IP:8080/api/health"
|
||||||
echo
|
echo
|
||||||
|
|
||||||
echo "smoke: all checks passed against $CANDIDATE"
|
echo "smoke: all checks passed against $CANDIDATE"
|
||||||
|
|
||||||
|
# Move the channel tags — the whole point of the gate.
|
||||||
|
#
|
||||||
|
# Lives in its own job because the verdict it depends on cannot exist until
|
||||||
|
# after build-web has finished, and the promote used to run INSIDE build-web.
|
||||||
|
#
|
||||||
|
# `needs` on smoke-web is the gate. A failed smoke skips this job, so a
|
||||||
|
# refresh that broke something leaves :latest naming the build that works —
|
||||||
|
# "the refresh failed" and "production is broken" must not be the same event.
|
||||||
|
# A SKIPPED smoke also skips this job, which is the behaviour that matters
|
||||||
|
# most: on run 5290 the gate silently skipped itself, and a design where only
|
||||||
|
# a FAILED gate blocks would have published unverified images while reporting
|
||||||
|
# success. Not running is not the same as passing.
|
||||||
|
#
|
||||||
|
# All three images promote TOGETHER, or none do. They are one stack: build.yml
|
||||||
|
# already refuses to publish a :dev web image beside a stale :dev ml, because
|
||||||
|
# the mismatch only shows up as a runtime failure. A refresh that published ml
|
||||||
|
# and withheld web would be that same trap, arrived at through the gate.
|
||||||
|
#
|
||||||
|
# The gate covers the web image only (milestone 362 step 3 scoped it there),
|
||||||
|
# so ml and agent are being held to web's verdict rather than their own. That
|
||||||
|
# is deliberate and it is the conservative direction — they ship together, so
|
||||||
|
# the weakest evidence should govern all three — but it is not the same as
|
||||||
|
# having smoked them, and it should not be read as if it were.
|
||||||
|
promote:
|
||||||
|
needs: [build-web, build-ml, build-agent, smoke-web]
|
||||||
|
# Only a refresh publishes through a candidate; a push writes its channel
|
||||||
|
# tag directly from the build. Reads the same reuse-step decision the build
|
||||||
|
# took, via a job output — a job's `if:` cannot see the `env` context.
|
||||||
|
if: needs.build-web.outputs.candidate == 'true'
|
||||||
|
runs-on: python-ci
|
||||||
|
container:
|
||||||
|
image: git.fabledsword.com/bvandeusen/ci-python:3.14
|
||||||
|
steps:
|
||||||
|
- name: Point the channel tags at the smoked candidates
|
||||||
|
env:
|
||||||
|
TOKEN: ${{ secrets.RELEASE_TOKEN }}
|
||||||
|
ACTOR: ${{ github.actor }}
|
||||||
|
run: |
|
||||||
|
set -eu
|
||||||
|
# `latest` is not a guess: a refresh always builds `main` (BUILD_REF),
|
||||||
|
# and the "must have checked out main" guard in every build job fails
|
||||||
|
# the run if that did not hold. So the channel is main's.
|
||||||
|
TAG=latest
|
||||||
|
FAILED=""
|
||||||
|
|
||||||
|
for NAME in fabledcurator fabledcurator-ml fabledcurator-agent; do
|
||||||
|
REPO="bvandeusen/$NAME"
|
||||||
|
echo "promote: $REPO"
|
||||||
|
|
||||||
|
# Registry auth is its own token exchange — `docker login`
|
||||||
|
# authenticates the docker client, not curl. Deadline on every call
|
||||||
|
# (rule 156): a registry that stops answering must fail this step,
|
||||||
|
# not hang the weekly refresh until the job times out.
|
||||||
|
BEARER=$(curl -fsS --max-time 30 -u "$ACTOR:$TOKEN" \
|
||||||
|
"https://git.fabledsword.com/v2/token?scope=repository:$REPO:pull,push&service=git.fabledsword.com" \
|
||||||
|
| python3 -c 'import sys,json; print(json.load(sys.stdin)["token"])')
|
||||||
|
|
||||||
|
# Ask for the IMAGE manifest media types only. Offering the index
|
||||||
|
# types too would let the registry hand back an index if one ever
|
||||||
|
# existed at this tag, and we would faithfully copy the thing this
|
||||||
|
# whole approach exists to avoid creating.
|
||||||
|
ACCEPT='application/vnd.oci.image.manifest.v1+json, application/vnd.docker.distribution.manifest.v2+json'
|
||||||
|
CT=$(curl -fsS --max-time 60 -o manifest.json -D headers.txt \
|
||||||
|
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
|
||||||
|
"https://git.fabledsword.com/v2/$REPO/manifests/refresh-candidate" \
|
||||||
|
&& tr -d '\r' < headers.txt | awk -F': ' '/^[Cc]ontent-[Tt]ype:/{print $2}')
|
||||||
|
test -n "$CT"
|
||||||
|
SRC=$(tr -d '\r' < headers.txt | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
|
||||||
|
echo "promote: candidate $SRC ($CT)"
|
||||||
|
|
||||||
|
# NOT `imagetools create`. That wraps its source in an INDEX, and
|
||||||
|
# `.Image.Config.Labels` does not resolve through one — the
|
||||||
|
# fc.revision the reuse check reads off the channel tag would come
|
||||||
|
# back empty, every later push would miss and rebuild, and nothing
|
||||||
|
# would go red (#3183, run 4751). A manifest PUT is what "make this
|
||||||
|
# tag name that image" means at the registry: same bytes, same media
|
||||||
|
# type, same digest, no layer transfer.
|
||||||
|
curl -fsS --max-time 120 -X PUT \
|
||||||
|
-H "Authorization: Bearer $BEARER" -H "Content-Type: $CT" \
|
||||||
|
--data-binary @manifest.json \
|
||||||
|
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG"
|
||||||
|
|
||||||
|
# Read it back. A PUT that returned 2xx but landed something else is
|
||||||
|
# exactly the silent-and-plausible failure this pipeline keeps
|
||||||
|
# producing, and the check costs one request.
|
||||||
|
NOW=$(curl -fsS --max-time 30 -o /dev/null -D - \
|
||||||
|
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
|
||||||
|
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG" \
|
||||||
|
| tr -d '\r' | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
|
||||||
|
if [ "$NOW" != "$SRC" ]; then
|
||||||
|
echo "promote: FAILED — $NAME:$TAG is $NOW, expected $SRC" >&2
|
||||||
|
FAILED="$FAILED $NAME"
|
||||||
|
continue
|
||||||
|
fi
|
||||||
|
echo "promote: $NAME:$TAG now names $NOW"
|
||||||
|
done
|
||||||
|
|
||||||
|
if [ -n "$FAILED" ]; then
|
||||||
|
echo "" >&2
|
||||||
|
echo "promote: FAILED for:$FAILED" >&2
|
||||||
|
echo "promote: the channel tags are now INCONSISTENT — some images" >&2
|
||||||
|
echo "promote: moved and some did not. Re-run this refresh; the" >&2
|
||||||
|
echo "promote: candidates are still published and the promote is" >&2
|
||||||
|
echo "promote: idempotent." >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
echo "promote: all three channel tags moved"
|
||||||
|
|
||||||
build-ml:
|
build-ml:
|
||||||
runs-on: python-ci
|
runs-on: python-ci
|
||||||
container:
|
container:
|
||||||
@@ -1452,16 +1530,20 @@ jobs:
|
|||||||
# already allows for :buildcache, not the per-build tag family that
|
# already allows for :buildcache, not the per-build tag family that
|
||||||
# milestone 318 withdrew.
|
# milestone 318 withdrew.
|
||||||
#
|
#
|
||||||
# Both values are decided HERE, beside `hit`, for the reason the
|
# Decided HERE, beside `hit`, for the reason the force/schedule
|
||||||
# force/schedule branch below gives: one step decides what this job
|
# branch below gives: one step decides what this job does. A
|
||||||
# does. A promote condition derived independently could disagree with
|
# condition derived independently could disagree with the tag the
|
||||||
# the tag the build actually wrote.
|
# build actually wrote.
|
||||||
|
#
|
||||||
|
# build-web additionally exposes this as `outputs.candidate`, which is
|
||||||
|
# what gates the `promote` job — a job's `if:` cannot read `env`, and
|
||||||
|
# one flag is enough because all three derive it from the same
|
||||||
|
# IS_REFRESH. ml and agent do not re-emit it; a second copy nothing
|
||||||
|
# reads is the kind of thing that later reads as load-bearing.
|
||||||
if [ "${IS_REFRESH:-}" = "true" ]; then
|
if [ "${IS_REFRESH:-}" = "true" ]; then
|
||||||
echo "build_ref=$IMAGE:refresh-candidate" >> "$GITHUB_OUTPUT"
|
echo "build_ref=$IMAGE:refresh-candidate" >> "$GITHUB_OUTPUT"
|
||||||
echo "promote=true" >> "$GITHUB_OUTPUT"
|
|
||||||
else
|
else
|
||||||
echo "build_ref=$IMAGE:$T" >> "$GITHUB_OUTPUT"
|
echo "build_ref=$IMAGE:$T" >> "$GITHUB_OUTPUT"
|
||||||
echo "promote=false" >> "$GITHUB_OUTPUT"
|
|
||||||
fi
|
fi
|
||||||
|
|
||||||
# Compare VALUES, never exit codes. Measured on buildx v0.36.1
|
# Compare VALUES, never exit codes. Measured on buildx v0.36.1
|
||||||
@@ -1595,77 +1677,6 @@ jobs:
|
|||||||
cache-from: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator-ml:buildcache
|
cache-from: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator-ml:buildcache
|
||||||
cache-to: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator-ml:buildcache,mode=max
|
cache-to: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator-ml:buildcache,mode=max
|
||||||
|
|
||||||
# Point the channel tag at the candidate the refresh just built.
|
|
||||||
#
|
|
||||||
# Unconditional TODAY, so this milestone never leaves the refresh in a
|
|
||||||
# state where it builds and publishes nothing. Step 4 wraps it in the
|
|
||||||
# smoke suite's verdict; until then the scheduled path behaves exactly
|
|
||||||
# as it did, just via two operations instead of one.
|
|
||||||
#
|
|
||||||
# NOT `imagetools create`. That wraps its source in an INDEX, and an
|
|
||||||
# indexed channel tag is the one thing this pipeline cannot survive:
|
|
||||||
# `.Image.Config.Labels` does not resolve through an index, so the
|
|
||||||
# fc.revision the reuse check reads off the channel tag would come back
|
|
||||||
# empty, every subsequent push would miss and rebuild, and nothing would
|
|
||||||
# go red. That is #3183, observed on run 4751 — reuse worked exactly once
|
|
||||||
# and the only symptom was the bill. The repoint step below excludes its
|
|
||||||
# own source tag for precisely this reason; a promote that re-introduced
|
|
||||||
# the wrap through a different door would undo that care.
|
|
||||||
#
|
|
||||||
# A manifest PUT is what "make this tag name that image" means at the
|
|
||||||
# registry level: the same bytes under the same media type, so the digest
|
|
||||||
# is identical, the media type is preserved, and no layer moves.
|
|
||||||
- name: Promote the refresh candidate to the channel
|
|
||||||
if: steps.reuse.outputs.promote == 'true'
|
|
||||||
env:
|
|
||||||
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator-ml
|
|
||||||
CHANNEL_REF: ${{ steps.reuse.outputs.channel_ref }}
|
|
||||||
TOKEN: ${{ secrets.RELEASE_TOKEN }}
|
|
||||||
ACTOR: ${{ github.actor }}
|
|
||||||
run: |
|
|
||||||
set -eu
|
|
||||||
REPO=${IMAGE#git.fabledsword.com/}
|
|
||||||
TAG=${CHANNEL_REF##*:}
|
|
||||||
|
|
||||||
# Registry auth is its own token exchange — the `docker login` above
|
|
||||||
# authenticates the docker client, not curl. Deadline on every call
|
|
||||||
# (rule 156): a registry that stops answering must fail this step,
|
|
||||||
# not hang the weekly refresh until the job times out.
|
|
||||||
BEARER=$(curl -fsS --max-time 30 -u "$ACTOR:$TOKEN" \
|
|
||||||
"https://git.fabledsword.com/v2/token?scope=repository:$REPO:pull,push&service=git.fabledsword.com" \
|
|
||||||
| python3 -c 'import sys,json; print(json.load(sys.stdin)["token"])')
|
|
||||||
|
|
||||||
# Ask for the image manifest media types ONLY. Offering the index
|
|
||||||
# types too would let the registry hand back an index if one ever
|
|
||||||
# existed at this tag, and we would faithfully copy the thing we are
|
|
||||||
# trying not to create.
|
|
||||||
ACCEPT='application/vnd.oci.image.manifest.v1+json, application/vnd.docker.distribution.manifest.v2+json'
|
|
||||||
CT=$(curl -fsS --max-time 60 -o manifest.json -D headers.txt \
|
|
||||||
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
|
|
||||||
"https://git.fabledsword.com/v2/$REPO/manifests/refresh-candidate" \
|
|
||||||
&& tr -d '\r' < headers.txt | awk -F': ' '/^[Cc]ontent-[Tt]ype:/{print $2}')
|
|
||||||
test -n "$CT"
|
|
||||||
SRC_DIGEST=$(tr -d '\r' < headers.txt | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
|
|
||||||
echo "promote: candidate is $SRC_DIGEST ($CT)"
|
|
||||||
|
|
||||||
curl -fsS --max-time 120 -X PUT \
|
|
||||||
-H "Authorization: Bearer $BEARER" -H "Content-Type: $CT" \
|
|
||||||
--data-binary @manifest.json \
|
|
||||||
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG"
|
|
||||||
|
|
||||||
# Read it back. A PUT that returned 2xx but landed something else is
|
|
||||||
# exactly the silent-and-plausible failure this pipeline keeps
|
|
||||||
# producing, and the check costs one request.
|
|
||||||
NOW=$(curl -fsS --max-time 30 -o /dev/null -D - \
|
|
||||||
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
|
|
||||||
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG" \
|
|
||||||
| tr -d '\r' | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
|
|
||||||
if [ "$NOW" != "$SRC_DIGEST" ]; then
|
|
||||||
echo "promote: $IMAGE:$TAG is $NOW, expected $SRC_DIGEST" >&2
|
|
||||||
exit 1
|
|
||||||
fi
|
|
||||||
echo "promote: $IMAGE:$TAG now names $NOW"
|
|
||||||
|
|
||||||
# Every tag but the channel's own is written HERE, registry-side,
|
# Every tag but the channel's own is written HERE, registry-side,
|
||||||
# whether or not a build ran. Each -t becomes another reference to the
|
# whether or not a build ran. Each -t becomes another reference to the
|
||||||
# SAME manifest the channel tag holds, so :c-<sha> is byte-identical to
|
# SAME manifest the channel tag holds, so :c-<sha> is byte-identical to
|
||||||
@@ -1947,16 +1958,20 @@ jobs:
|
|||||||
# already allows for :buildcache, not the per-build tag family that
|
# already allows for :buildcache, not the per-build tag family that
|
||||||
# milestone 318 withdrew.
|
# milestone 318 withdrew.
|
||||||
#
|
#
|
||||||
# Both values are decided HERE, beside `hit`, for the reason the
|
# Decided HERE, beside `hit`, for the reason the force/schedule
|
||||||
# force/schedule branch below gives: one step decides what this job
|
# branch below gives: one step decides what this job does. A
|
||||||
# does. A promote condition derived independently could disagree with
|
# condition derived independently could disagree with the tag the
|
||||||
# the tag the build actually wrote.
|
# build actually wrote.
|
||||||
|
#
|
||||||
|
# build-web additionally exposes this as `outputs.candidate`, which is
|
||||||
|
# what gates the `promote` job — a job's `if:` cannot read `env`, and
|
||||||
|
# one flag is enough because all three derive it from the same
|
||||||
|
# IS_REFRESH. ml and agent do not re-emit it; a second copy nothing
|
||||||
|
# reads is the kind of thing that later reads as load-bearing.
|
||||||
if [ "${IS_REFRESH:-}" = "true" ]; then
|
if [ "${IS_REFRESH:-}" = "true" ]; then
|
||||||
echo "build_ref=$IMAGE:refresh-candidate" >> "$GITHUB_OUTPUT"
|
echo "build_ref=$IMAGE:refresh-candidate" >> "$GITHUB_OUTPUT"
|
||||||
echo "promote=true" >> "$GITHUB_OUTPUT"
|
|
||||||
else
|
else
|
||||||
echo "build_ref=$IMAGE:$T" >> "$GITHUB_OUTPUT"
|
echo "build_ref=$IMAGE:$T" >> "$GITHUB_OUTPUT"
|
||||||
echo "promote=false" >> "$GITHUB_OUTPUT"
|
|
||||||
fi
|
fi
|
||||||
|
|
||||||
# Compare VALUES, never exit codes. Measured on buildx v0.36.1
|
# Compare VALUES, never exit codes. Measured on buildx v0.36.1
|
||||||
@@ -2090,77 +2105,6 @@ jobs:
|
|||||||
cache-from: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator-agent:buildcache
|
cache-from: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator-agent:buildcache
|
||||||
cache-to: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator-agent:buildcache,mode=max
|
cache-to: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator-agent:buildcache,mode=max
|
||||||
|
|
||||||
# Point the channel tag at the candidate the refresh just built.
|
|
||||||
#
|
|
||||||
# Unconditional TODAY, so this milestone never leaves the refresh in a
|
|
||||||
# state where it builds and publishes nothing. Step 4 wraps it in the
|
|
||||||
# smoke suite's verdict; until then the scheduled path behaves exactly
|
|
||||||
# as it did, just via two operations instead of one.
|
|
||||||
#
|
|
||||||
# NOT `imagetools create`. That wraps its source in an INDEX, and an
|
|
||||||
# indexed channel tag is the one thing this pipeline cannot survive:
|
|
||||||
# `.Image.Config.Labels` does not resolve through an index, so the
|
|
||||||
# fc.revision the reuse check reads off the channel tag would come back
|
|
||||||
# empty, every subsequent push would miss and rebuild, and nothing would
|
|
||||||
# go red. That is #3183, observed on run 4751 — reuse worked exactly once
|
|
||||||
# and the only symptom was the bill. The repoint step below excludes its
|
|
||||||
# own source tag for precisely this reason; a promote that re-introduced
|
|
||||||
# the wrap through a different door would undo that care.
|
|
||||||
#
|
|
||||||
# A manifest PUT is what "make this tag name that image" means at the
|
|
||||||
# registry level: the same bytes under the same media type, so the digest
|
|
||||||
# is identical, the media type is preserved, and no layer moves.
|
|
||||||
- name: Promote the refresh candidate to the channel
|
|
||||||
if: steps.reuse.outputs.promote == 'true'
|
|
||||||
env:
|
|
||||||
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator-agent
|
|
||||||
CHANNEL_REF: ${{ steps.reuse.outputs.channel_ref }}
|
|
||||||
TOKEN: ${{ secrets.RELEASE_TOKEN }}
|
|
||||||
ACTOR: ${{ github.actor }}
|
|
||||||
run: |
|
|
||||||
set -eu
|
|
||||||
REPO=${IMAGE#git.fabledsword.com/}
|
|
||||||
TAG=${CHANNEL_REF##*:}
|
|
||||||
|
|
||||||
# Registry auth is its own token exchange — the `docker login` above
|
|
||||||
# authenticates the docker client, not curl. Deadline on every call
|
|
||||||
# (rule 156): a registry that stops answering must fail this step,
|
|
||||||
# not hang the weekly refresh until the job times out.
|
|
||||||
BEARER=$(curl -fsS --max-time 30 -u "$ACTOR:$TOKEN" \
|
|
||||||
"https://git.fabledsword.com/v2/token?scope=repository:$REPO:pull,push&service=git.fabledsword.com" \
|
|
||||||
| python3 -c 'import sys,json; print(json.load(sys.stdin)["token"])')
|
|
||||||
|
|
||||||
# Ask for the image manifest media types ONLY. Offering the index
|
|
||||||
# types too would let the registry hand back an index if one ever
|
|
||||||
# existed at this tag, and we would faithfully copy the thing we are
|
|
||||||
# trying not to create.
|
|
||||||
ACCEPT='application/vnd.oci.image.manifest.v1+json, application/vnd.docker.distribution.manifest.v2+json'
|
|
||||||
CT=$(curl -fsS --max-time 60 -o manifest.json -D headers.txt \
|
|
||||||
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
|
|
||||||
"https://git.fabledsword.com/v2/$REPO/manifests/refresh-candidate" \
|
|
||||||
&& tr -d '\r' < headers.txt | awk -F': ' '/^[Cc]ontent-[Tt]ype:/{print $2}')
|
|
||||||
test -n "$CT"
|
|
||||||
SRC_DIGEST=$(tr -d '\r' < headers.txt | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
|
|
||||||
echo "promote: candidate is $SRC_DIGEST ($CT)"
|
|
||||||
|
|
||||||
curl -fsS --max-time 120 -X PUT \
|
|
||||||
-H "Authorization: Bearer $BEARER" -H "Content-Type: $CT" \
|
|
||||||
--data-binary @manifest.json \
|
|
||||||
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG"
|
|
||||||
|
|
||||||
# Read it back. A PUT that returned 2xx but landed something else is
|
|
||||||
# exactly the silent-and-plausible failure this pipeline keeps
|
|
||||||
# producing, and the check costs one request.
|
|
||||||
NOW=$(curl -fsS --max-time 30 -o /dev/null -D - \
|
|
||||||
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
|
|
||||||
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG" \
|
|
||||||
| tr -d '\r' | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
|
|
||||||
if [ "$NOW" != "$SRC_DIGEST" ]; then
|
|
||||||
echo "promote: $IMAGE:$TAG is $NOW, expected $SRC_DIGEST" >&2
|
|
||||||
exit 1
|
|
||||||
fi
|
|
||||||
echo "promote: $IMAGE:$TAG now names $NOW"
|
|
||||||
|
|
||||||
# Every tag but the channel's own is written HERE, registry-side,
|
# Every tag but the channel's own is written HERE, registry-side,
|
||||||
# whether or not a build ran. Each -t becomes another reference to the
|
# whether or not a build ran. Each -t becomes another reference to the
|
||||||
# SAME manifest the channel tag holds, so :c-<sha> is byte-identical to
|
# SAME manifest the channel tag holds, so :c-<sha> is byte-identical to
|
||||||
|
|||||||
@@ -90,7 +90,31 @@ If you forget it, the symptom is a long build instead of a quick pull.
|
|||||||
The database schema is created automatically on first start — the web container
|
The database schema is created automatically on first start — the web container
|
||||||
runs its migrations before serving. Nothing to initialise by hand.
|
runs its migrations before serving. Nothing to initialise by hand.
|
||||||
|
|
||||||
A few things are worth knowing about the first few minutes:
|
**One thing does need a deliberate act, and the app will not start without it.**
|
||||||
|
FabledCurator encrypts your stored platform credentials with a key it keeps at
|
||||||
|
`./images/secrets/credential_key.b64`. On a brand-new install that file does not
|
||||||
|
exist, and rather than quietly creating one the app stops:
|
||||||
|
|
||||||
|
```
|
||||||
|
MissingCredentialKey: Fernet key file not found at /images/secrets/credential_key.b64
|
||||||
|
```
|
||||||
|
|
||||||
|
Set `CURATOR_BOOTSTRAP_NEW_KEY=1` in your `.env` for the first `up`, then delete
|
||||||
|
the line once the container is running. `.env.example` ships it with that
|
||||||
|
instruction attached.
|
||||||
|
|
||||||
|
The refusal is deliberate, and worth understanding rather than working around:
|
||||||
|
auto-creating a key is indistinguishable from the disaster case — a restore that
|
||||||
|
brought the database back but lost `./images/secrets` — where it would mint a key
|
||||||
|
that cannot decrypt anything, leaving an instance that looks healthy while every
|
||||||
|
paywalled download fails. Making you say so once, on an empty install, is the
|
||||||
|
price of that not happening silently later.
|
||||||
|
|
||||||
|
**Which means: back up `./images/secrets/` alongside your database.** It is the
|
||||||
|
only thing that can read your stored credentials. A database restored without it
|
||||||
|
needs every credential entered again by hand.
|
||||||
|
|
||||||
|
A few other things are worth knowing about the first few minutes:
|
||||||
|
|
||||||
- **The ML worker downloads its model weights on first boot**, several GB from
|
- **The ML worker downloads its model weights on first boot**, several GB from
|
||||||
HuggingFace into `./models`. Until that finishes, tagging is queued rather
|
HuggingFace into `./models`. Until that finishes, tagging is queued rather
|
||||||
|
|||||||
@@ -0,0 +1,64 @@
|
|||||||
|
"""service_seen — the learned roster that makes a stopped part observable.
|
||||||
|
|
||||||
|
Milestone 365. Nothing in FabledCurator knew what was SUPPOSED to be running:
|
||||||
|
`celery inspect` reports the workers that answer, so a dead worker was a
|
||||||
|
shorter list rather than a red light, and the only surface that could tell an
|
||||||
|
operator otherwise was Portainer. This table is the memory that turns an
|
||||||
|
absence into something the app can see.
|
||||||
|
|
||||||
|
Keyed on the queue set for a celery role and on agent_id for the GPU agent —
|
||||||
|
NOT on the celery worker name, which here is `celery@<container id>` and is
|
||||||
|
minted fresh on every deploy. See the model docstring for why that choice is
|
||||||
|
the whole design.
|
||||||
|
|
||||||
|
## First migration on the collapsed baseline
|
||||||
|
|
||||||
|
0089 is the single generated baseline that replaced revisions 0001..0089
|
||||||
|
(milestone 328). This is the first revision written on top of it, so it is
|
||||||
|
also the first evidence that the chain steps forward from the collapse rather
|
||||||
|
than merely reproducing the schema — which nothing had demonstrated yet.
|
||||||
|
|
||||||
|
An existing install is at 0089 because it ran the real 0089; a fresh one is at
|
||||||
|
0089 because it ran the baseline. Both arrive here identically, which was the
|
||||||
|
property the collapse was designed around.
|
||||||
|
|
||||||
|
Revision ID: 0090
|
||||||
|
Revises: 0089
|
||||||
|
Create Date: 2026-09-02
|
||||||
|
|
||||||
|
"""
|
||||||
|
from typing import Sequence, Union
|
||||||
|
|
||||||
|
import sqlalchemy as sa
|
||||||
|
from alembic import op
|
||||||
|
|
||||||
|
revision: str = "0090"
|
||||||
|
down_revision: Union[str, None] = "0089"
|
||||||
|
branch_labels: Union[str, Sequence[str], None] = None
|
||||||
|
depends_on: Union[str, Sequence[str], None] = None
|
||||||
|
|
||||||
|
|
||||||
|
def upgrade() -> None:
|
||||||
|
op.create_table(
|
||||||
|
"service_seen",
|
||||||
|
sa.Column("key", sa.String(length=128), nullable=False),
|
||||||
|
sa.Column("kind", sa.String(length=16), nullable=False),
|
||||||
|
sa.Column("display_name", sa.String(length=64), nullable=False),
|
||||||
|
sa.Column(
|
||||||
|
"first_seen_at", sa.DateTime(timezone=True),
|
||||||
|
server_default=sa.text("now()"), nullable=False,
|
||||||
|
),
|
||||||
|
sa.Column(
|
||||||
|
"last_seen_at", sa.DateTime(timezone=True),
|
||||||
|
server_default=sa.text("now()"), nullable=False,
|
||||||
|
),
|
||||||
|
sa.Column("details", sa.JSON(), nullable=False),
|
||||||
|
sa.PrimaryKeyConstraint("key", name=op.f("pk_service_seen")),
|
||||||
|
)
|
||||||
|
# No secondary indexes, deliberately: one row per moving part means every
|
||||||
|
# read is a handful of rows and an index would be write cost buying
|
||||||
|
# nothing (#3301 removed seven of exactly that shape).
|
||||||
|
|
||||||
|
|
||||||
|
def downgrade() -> None:
|
||||||
|
op.drop_table("service_seen")
|
||||||
@@ -38,6 +38,7 @@ def all_blueprints() -> list[Blueprint]:
|
|||||||
from .suggestions import suggestions_bp
|
from .suggestions import suggestions_bp
|
||||||
from .system_activity import system_activity_bp
|
from .system_activity import system_activity_bp
|
||||||
from .system_backup import system_backup_bp
|
from .system_backup import system_backup_bp
|
||||||
|
from .system_health import system_health_bp
|
||||||
from .tags import tags_bp
|
from .tags import tags_bp
|
||||||
from .thumbnails import thumbnails_bp
|
from .thumbnails import thumbnails_bp
|
||||||
return [
|
return [
|
||||||
@@ -51,6 +52,7 @@ def all_blueprints() -> list[Blueprint]:
|
|||||||
showcase_bp,
|
showcase_bp,
|
||||||
settings_bp,
|
settings_bp,
|
||||||
system_activity_bp,
|
system_activity_bp,
|
||||||
|
system_health_bp,
|
||||||
system_backup_bp,
|
system_backup_bp,
|
||||||
admin_bp,
|
admin_bp,
|
||||||
cleanup_bp,
|
cleanup_bp,
|
||||||
|
|||||||
@@ -21,6 +21,7 @@ from ..services.gallery_service import image_url
|
|||||||
from ..services.ml.gpu_jobs import GpuJobService, error_dedupe_statements
|
from ..services.ml.gpu_jobs import GpuJobService, error_dedupe_statements
|
||||||
from ..services.ml.gpu_triage import classify_reason, recover_defective_image
|
from ..services.ml.gpu_triage import classify_reason, recover_defective_image
|
||||||
from ..services.ml.regions import RegionService
|
from ..services.ml.regions import RegionService
|
||||||
|
from ..services.service_roster import touch_service
|
||||||
|
|
||||||
gpu_bp = Blueprint("gpu", __name__, url_prefix="/api/gpu")
|
gpu_bp = Blueprint("gpu", __name__, url_prefix="/api/gpu")
|
||||||
|
|
||||||
@@ -256,6 +257,18 @@ async def lease():
|
|||||||
if not await _agent_authed(session):
|
if not await _agent_authed(session):
|
||||||
return jsonify({"error": "unauthorized"}), 401
|
return jsonify({"error": "unauthorized"}), 401
|
||||||
jobs = await GpuJobService(session).lease(agent_id, batch_size=batch)
|
jobs = await GpuJobService(session).lease(agent_id, batch_size=batch)
|
||||||
|
# The agent cannot be polled — it is HTTP-only and pulls from here, so
|
||||||
|
# web never dials it. A lease IS the check-in, and until milestone 365
|
||||||
|
# it was thrown away: an agent sitting idle with nothing to lease left
|
||||||
|
# no trace at all and was indistinguishable from one switched off a
|
||||||
|
# week ago. Recorded on the call that was already happening.
|
||||||
|
await touch_service(
|
||||||
|
session,
|
||||||
|
key=f"agent:{agent_id}",
|
||||||
|
kind="agent",
|
||||||
|
display_name="GPU agent" if agent_id == "agent" else f"GPU agent ({agent_id})",
|
||||||
|
details={"agent_id": agent_id, "last_call": "lease", "leased": len(jobs)},
|
||||||
|
)
|
||||||
ml = await MLSettings.load(session)
|
ml = await MLSettings.load(session)
|
||||||
# image rows for url/mime in one shot
|
# image rows for url/mime in one shot
|
||||||
ids = [j.image_record_id for j in jobs]
|
ids = [j.image_record_id for j in jobs]
|
||||||
@@ -329,6 +342,13 @@ async def heartbeat():
|
|||||||
if not await _agent_authed(session):
|
if not await _agent_authed(session):
|
||||||
return jsonify({"error": "unauthorized"}), 401
|
return jsonify({"error": "unauthorized"}), 401
|
||||||
n = await GpuJobService(session).heartbeat(agent_id, job_ids)
|
n = await GpuJobService(session).heartbeat(agent_id, job_ids)
|
||||||
|
await touch_service(
|
||||||
|
session,
|
||||||
|
key=f"agent:{agent_id}",
|
||||||
|
kind="agent",
|
||||||
|
display_name="GPU agent" if agent_id == "agent" else f"GPU agent ({agent_id})",
|
||||||
|
details={"agent_id": agent_id, "last_call": "heartbeat", "extended": n},
|
||||||
|
)
|
||||||
await session.commit()
|
await session.commit()
|
||||||
return jsonify({"extended": n})
|
return jsonify({"extended": n})
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,192 @@
|
|||||||
|
"""Is every part of FabledCurator running? One verdict, one endpoint.
|
||||||
|
|
||||||
|
Milestone 365. The nav indicator and the System page both read this and
|
||||||
|
nothing else — composing a verdict is this module's job, not the UI's.
|
||||||
|
|
||||||
|
## Two kinds of part, answered two different ways
|
||||||
|
|
||||||
|
**Learned** — celery roles and the GPU agent, from `service_seen`. The
|
||||||
|
question is "how long since it checked in", and these are the parts that can
|
||||||
|
be ABSENT, which is the whole point: `celery inspect` alone reports presence,
|
||||||
|
so a dead worker is a shorter list rather than a red light.
|
||||||
|
|
||||||
|
**Probed live** — Postgres and Redis. Always expected, never learned, and a
|
||||||
|
last-seen for them would be actively misleading: that Redis answered thirty
|
||||||
|
seconds ago says nothing about now.
|
||||||
|
|
||||||
|
## This endpoint must never fail because something it checks has failed
|
||||||
|
|
||||||
|
The inversion is easy to write by accident and it destroys the feature exactly
|
||||||
|
when it is needed — a 500 when Redis is down, instead of `redis: down`. Every
|
||||||
|
probe is wrapped, every wait has a deadline (rule 156), and the roster refresh
|
||||||
|
swallows its own errors. The worst case is a part reported `unknown`, which is
|
||||||
|
a true statement.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import asyncio
|
||||||
|
import logging
|
||||||
|
import time
|
||||||
|
from datetime import UTC, datetime
|
||||||
|
|
||||||
|
from quart import Blueprint, jsonify
|
||||||
|
from sqlalchemy import select, text
|
||||||
|
|
||||||
|
from ..config import get_config
|
||||||
|
from ..extensions import get_session
|
||||||
|
from ..models import ServiceSeen
|
||||||
|
from ..services.service_roster import refresh_if_stale
|
||||||
|
|
||||||
|
log = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
system_health_bp = Blueprint("system_health", __name__, url_prefix="/api/system")
|
||||||
|
|
||||||
|
# How long a learned part may go quiet before it is doubted, then disbelieved.
|
||||||
|
#
|
||||||
|
# These are deliberately generous, and the reason is a deploy rather than a
|
||||||
|
# worker: `docker compose up -d` rolls start-first, so a role is briefly served
|
||||||
|
# by two containers and then by neither while the old one drains. Thresholds
|
||||||
|
# tight enough to catch a crash in seconds would paint the page red every time
|
||||||
|
# the stack is updated, and an alarm that cries wolf on every deploy is one
|
||||||
|
# nobody reads. Tune down only after watching a real deploy pass through.
|
||||||
|
STALE_AFTER_SECONDS = 90
|
||||||
|
DOWN_AFTER_SECONDS = 300
|
||||||
|
|
||||||
|
# Probes cross a process boundary, so they carry deadlines. A hung Postgres
|
||||||
|
# must make this endpoint say "postgres: down", not hang alongside it.
|
||||||
|
PROBE_TIMEOUT_SECONDS = 2.0
|
||||||
|
|
||||||
|
_OK, _STALE, _DOWN, _UNKNOWN = "ok", "stale", "down", "unknown"
|
||||||
|
|
||||||
|
# Worst-first, so an overall verdict is just the max.
|
||||||
|
_SEVERITY = {_OK: 0, _UNKNOWN: 1, _STALE: 2, _DOWN: 3}
|
||||||
|
|
||||||
|
|
||||||
|
def _age_state(age_seconds: float) -> str:
|
||||||
|
if age_seconds >= DOWN_AFTER_SECONDS:
|
||||||
|
return _DOWN
|
||||||
|
if age_seconds >= STALE_AFTER_SECONDS:
|
||||||
|
return _STALE
|
||||||
|
return _OK
|
||||||
|
|
||||||
|
|
||||||
|
def _describe_learned(name: str, state: str, age: float, details: dict) -> str:
|
||||||
|
"""Say what the state MEANS. A red chip tells an operator less than a
|
||||||
|
sentence does at the moment they are deciding whether to go and look."""
|
||||||
|
if state == _OK:
|
||||||
|
replicas = details.get("replicas")
|
||||||
|
if replicas and replicas > 1:
|
||||||
|
return f"{name} is running ({replicas} replicas)"
|
||||||
|
return f"{name} is running"
|
||||||
|
mins = int(age // 60)
|
||||||
|
ago = f"{mins} min" if mins else f"{int(age)}s"
|
||||||
|
if state == _STALE:
|
||||||
|
return f"{name} has not checked in for {ago}"
|
||||||
|
return f"{name} has not checked in for {ago} — treat it as stopped"
|
||||||
|
|
||||||
|
|
||||||
|
async def _probe_postgres(session) -> dict:
|
||||||
|
started = time.monotonic()
|
||||||
|
try:
|
||||||
|
await asyncio.wait_for(
|
||||||
|
session.execute(text("SELECT 1")), timeout=PROBE_TIMEOUT_SECONDS
|
||||||
|
)
|
||||||
|
except Exception as exc: # noqa: BLE001 — a probe reports, it never raises
|
||||||
|
return {
|
||||||
|
"key": "postgres", "kind": "datastore", "name": "PostgreSQL",
|
||||||
|
"state": _DOWN, "detail": f"not answering: {type(exc).__name__}",
|
||||||
|
}
|
||||||
|
return {
|
||||||
|
"key": "postgres", "kind": "datastore", "name": "PostgreSQL", "state": _OK,
|
||||||
|
"detail": "answering", "latency_ms": round((time.monotonic() - started) * 1000, 1),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _ping_redis_sync() -> None:
|
||||||
|
import redis # local import; mirrors system_activity's pattern
|
||||||
|
|
||||||
|
client = redis.Redis.from_url(
|
||||||
|
get_config().celery_broker_url,
|
||||||
|
socket_connect_timeout=PROBE_TIMEOUT_SECONDS,
|
||||||
|
socket_timeout=PROBE_TIMEOUT_SECONDS,
|
||||||
|
)
|
||||||
|
client.ping()
|
||||||
|
|
||||||
|
|
||||||
|
async def _probe_redis() -> dict:
|
||||||
|
started = time.monotonic()
|
||||||
|
try:
|
||||||
|
await asyncio.wait_for(
|
||||||
|
asyncio.to_thread(_ping_redis_sync), timeout=PROBE_TIMEOUT_SECONDS * 2
|
||||||
|
)
|
||||||
|
except Exception as exc: # noqa: BLE001
|
||||||
|
return {
|
||||||
|
"key": "redis", "kind": "datastore", "name": "Redis",
|
||||||
|
"state": _DOWN,
|
||||||
|
"detail": f"not answering: {type(exc).__name__} — queues and workers "
|
||||||
|
f"cannot be reached either",
|
||||||
|
}
|
||||||
|
return {
|
||||||
|
"key": "redis", "kind": "datastore", "name": "Redis", "state": _OK,
|
||||||
|
"detail": "answering", "latency_ms": round((time.monotonic() - started) * 1000, 1),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
@system_health_bp.route("/health", methods=["GET"])
|
||||||
|
async def system_health():
|
||||||
|
"""Every part, its state, and one overall verdict.
|
||||||
|
|
||||||
|
Response: {overall, parts: [{key, kind, name, state, detail, last_seen_at,
|
||||||
|
…}], checked_at}
|
||||||
|
"""
|
||||||
|
parts: list[dict] = []
|
||||||
|
now = datetime.now(UTC)
|
||||||
|
|
||||||
|
async with get_session() as session:
|
||||||
|
# Postgres first, and if it is unreachable nothing else can be read —
|
||||||
|
# say so rather than failing, because "the database is down" is the
|
||||||
|
# single most useful thing this endpoint can ever report.
|
||||||
|
pg = await _probe_postgres(session)
|
||||||
|
parts.append(pg)
|
||||||
|
|
||||||
|
if pg["state"] == _OK:
|
||||||
|
# Rate-limited inside; see service_roster on why the web process
|
||||||
|
# is the right observer.
|
||||||
|
try:
|
||||||
|
await refresh_if_stale(session)
|
||||||
|
await session.commit()
|
||||||
|
except Exception: # noqa: BLE001
|
||||||
|
log.warning("system health: roster refresh failed", exc_info=True)
|
||||||
|
|
||||||
|
rows = (
|
||||||
|
await session.execute(select(ServiceSeen).order_by(ServiceSeen.display_name))
|
||||||
|
).scalars().all()
|
||||||
|
for row in rows:
|
||||||
|
age = (now - row.last_seen_at).total_seconds()
|
||||||
|
state = _age_state(age)
|
||||||
|
parts.append({
|
||||||
|
"key": row.key,
|
||||||
|
"kind": row.kind,
|
||||||
|
"name": row.display_name,
|
||||||
|
"state": state,
|
||||||
|
"detail": _describe_learned(row.display_name, state, age, row.details or {}),
|
||||||
|
"last_seen_at": row.last_seen_at.isoformat(),
|
||||||
|
"first_seen_at": row.first_seen_at.isoformat(),
|
||||||
|
**{k: v for k, v in (row.details or {}).items() if k != "agent_id"},
|
||||||
|
})
|
||||||
|
|
||||||
|
parts.append(await _probe_redis())
|
||||||
|
|
||||||
|
overall = max((p["state"] for p in parts), key=lambda s: _SEVERITY[s], default=_UNKNOWN)
|
||||||
|
return jsonify({
|
||||||
|
"overall": overall,
|
||||||
|
"parts": sorted(parts, key=lambda p: (-_SEVERITY[p["state"]], p["name"])),
|
||||||
|
"checked_at": now.isoformat(),
|
||||||
|
# So the UI can explain a `stale` without hard-coding the same numbers
|
||||||
|
# in a second place.
|
||||||
|
"thresholds": {
|
||||||
|
"stale_after_seconds": STALE_AFTER_SECONDS,
|
||||||
|
"down_after_seconds": DOWN_AFTER_SECONDS,
|
||||||
|
},
|
||||||
|
})
|
||||||
@@ -32,6 +32,7 @@ from .presentation_review import PresentationReview
|
|||||||
from .series_chapter import SeriesChapter
|
from .series_chapter import SeriesChapter
|
||||||
from .series_page import SeriesPage
|
from .series_page import SeriesPage
|
||||||
from .series_suggestion import SeriesSuggestion
|
from .series_suggestion import SeriesSuggestion
|
||||||
|
from .service_seen import ServiceSeen
|
||||||
from .source import Source
|
from .source import Source
|
||||||
from .subscribestar_failed_media import SubscribeStarFailedMedia
|
from .subscribestar_failed_media import SubscribeStarFailedMedia
|
||||||
from .subscribestar_seen_media import SubscribeStarSeenMedia
|
from .subscribestar_seen_media import SubscribeStarSeenMedia
|
||||||
@@ -63,6 +64,7 @@ __all__ = [
|
|||||||
"SeriesChapter",
|
"SeriesChapter",
|
||||||
"SeriesPage",
|
"SeriesPage",
|
||||||
"SeriesSuggestion",
|
"SeriesSuggestion",
|
||||||
|
"ServiceSeen",
|
||||||
"ImageRecord",
|
"ImageRecord",
|
||||||
"ImageProvenance",
|
"ImageProvenance",
|
||||||
"ImageRegion",
|
"ImageRegion",
|
||||||
|
|||||||
@@ -0,0 +1,88 @@
|
|||||||
|
"""service_seen — the learned roster of FabledCurator's own moving parts.
|
||||||
|
|
||||||
|
Nothing else in this application knows what is SUPPOSED to be running.
|
||||||
|
`celery inspect` reports the workers that answer, so a stopped worker is a
|
||||||
|
shorter list rather than a red light, and Postgres and Redis have no
|
||||||
|
representation at all. That is why the only place an operator could see a
|
||||||
|
dead service was Portainer, which knows the intended set (milestone 365).
|
||||||
|
|
||||||
|
This table is the memory that makes an absence observable: every part that
|
||||||
|
has ever checked in, and when it last did. A row that stops advancing is a
|
||||||
|
part that stopped.
|
||||||
|
|
||||||
|
## Why the key is not the hostname
|
||||||
|
|
||||||
|
`_read_workers_sync()` returns celery's worker names, which here are
|
||||||
|
`celery@<container id>`. Those are minted fresh on every deploy. Keyed on
|
||||||
|
them, this table would record a death and a birth every time the stack is
|
||||||
|
updated — and a status page that goes red on every deploy is a status page
|
||||||
|
nobody reads, which is worse than not having one.
|
||||||
|
|
||||||
|
So a celery role is keyed on its **queue set**, which is assigned per role in
|
||||||
|
docker-compose.yml (`CELERY_QUEUES`) and survives container replacement:
|
||||||
|
|
||||||
|
default,import,thumbnail,download -> worker
|
||||||
|
maintenance,scan -> scheduler (celery worker --beat)
|
||||||
|
ml -> ml-worker
|
||||||
|
|
||||||
|
Two replicas of one role share a queue set and are therefore ONE row — which
|
||||||
|
is right, because the question being answered is "is that role being served",
|
||||||
|
not "how many containers exist". The replica count and their hostnames go in
|
||||||
|
`details`, where they can change without the identity changing.
|
||||||
|
|
||||||
|
The GPU agent is keyed on its `agent_id`, the identity its lease protocol
|
||||||
|
already uses (`api/gpu.py`).
|
||||||
|
|
||||||
|
## What is NOT in here
|
||||||
|
|
||||||
|
Postgres and Redis. They are always expected and never learned, and a
|
||||||
|
last-seen for them would be actively misleading — that one answered thirty
|
||||||
|
seconds ago says nothing about now. They are probed live at request time.
|
||||||
|
|
||||||
|
## kind
|
||||||
|
|
||||||
|
Plain `String`, not a Postgres ENUM and not CHECK-gated, matching
|
||||||
|
`gpu_job.status` and `backup_run.status`. The value set here is expected to
|
||||||
|
grow as parts are added, and a constraint swap per new kind (rule 36) would
|
||||||
|
be cost with no invariant behind it.
|
||||||
|
|
||||||
|
celery — a worker role, keyed on its queue set
|
||||||
|
agent — a GPU agent, keyed on its agent_id
|
||||||
|
"""
|
||||||
|
|
||||||
|
from datetime import datetime
|
||||||
|
|
||||||
|
from sqlalchemy import JSON, DateTime, String, func
|
||||||
|
from sqlalchemy.orm import Mapped, mapped_column
|
||||||
|
|
||||||
|
from .base import Base
|
||||||
|
|
||||||
|
|
||||||
|
class ServiceSeen(Base):
|
||||||
|
__tablename__ = "service_seen"
|
||||||
|
|
||||||
|
# No indexes beyond the primary key, deliberately. This table holds one row
|
||||||
|
# per moving part — a handful, forever — so every query against it is a
|
||||||
|
# full read of a few rows and an index would be write cost buying nothing
|
||||||
|
# (the lesson of #3301, which removed seven redundant ones).
|
||||||
|
key: Mapped[str] = mapped_column(String(128), primary_key=True)
|
||||||
|
kind: Mapped[str] = mapped_column(String(16), nullable=False)
|
||||||
|
|
||||||
|
# What to call it in the UI. Derived from the queue set where it is
|
||||||
|
# recognised, and falling back to the raw queue list where it is not — a
|
||||||
|
# deployment that slices its queues differently should still show something
|
||||||
|
# true rather than a name this code invented for it.
|
||||||
|
display_name: Mapped[str] = mapped_column(String(64), nullable=False)
|
||||||
|
|
||||||
|
first_seen_at: Mapped[datetime] = mapped_column(
|
||||||
|
DateTime(timezone=True), nullable=False, server_default=func.now(),
|
||||||
|
)
|
||||||
|
last_seen_at: Mapped[datetime] = mapped_column(
|
||||||
|
DateTime(timezone=True), nullable=False, server_default=func.now(),
|
||||||
|
)
|
||||||
|
|
||||||
|
# The parts that change without changing identity: replica hostnames,
|
||||||
|
# active task counts, the queues actually being served. Kept as a blob
|
||||||
|
# because it is displayed and never queried — giving it columns would
|
||||||
|
# invite filtering on it, which is what the activity endpoints are for.
|
||||||
|
details: Mapped[dict] = mapped_column(JSON, nullable=False, default=dict)
|
||||||
@@ -0,0 +1,175 @@
|
|||||||
|
"""The learned roster: which of FabledCurator's parts have checked in, and when.
|
||||||
|
|
||||||
|
Milestone 365. `celery inspect` answers "who is here"; this answers "who is
|
||||||
|
missing", which nothing in the application could do before — see
|
||||||
|
`models/service_seen.py` for why the identity is a queue set and not a
|
||||||
|
worker hostname.
|
||||||
|
|
||||||
|
## Who does the observing, and why it is the web process
|
||||||
|
|
||||||
|
Three candidates, and the choice matters more than the code:
|
||||||
|
|
||||||
|
* **A celery beat sweep.** Rejected. If the scheduler dies, the sweep stops,
|
||||||
|
every row goes stale, and the page reports that everything is down when one
|
||||||
|
thing is. An alarm that cannot distinguish "one part died" from "the
|
||||||
|
observer died" is worse than no alarm.
|
||||||
|
* **A background task in web.** Rejected on a detail of how this deploys:
|
||||||
|
hypercorn runs `--workers 4`, so a `before_serving` loop would be FOUR
|
||||||
|
concurrent inspect loops hammering the broker, forever, per container.
|
||||||
|
* **Refresh on demand, rate-limited by the data itself.** Taken. Whichever web
|
||||||
|
process happens to serve a health request refreshes the roster if it is
|
||||||
|
older than REFRESH_TTL, and otherwise reads what is already there.
|
||||||
|
|
||||||
|
The third has the property the other two lack: **the observer is the thing
|
||||||
|
serving the page.** If web is down you get a browser error rather than a
|
||||||
|
confidently green page, which is the honest failure. It also self-limits
|
||||||
|
without coordination — the TTL lives in the row everybody can see.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import asyncio
|
||||||
|
import logging
|
||||||
|
|
||||||
|
from sqlalchemy import func, select
|
||||||
|
from sqlalchemy.dialects.postgresql import insert as pg_insert
|
||||||
|
from sqlalchemy.ext.asyncio import AsyncSession
|
||||||
|
|
||||||
|
from ..models import ServiceSeen
|
||||||
|
|
||||||
|
log = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
# How stale the roster may be before a health request refreshes it. Comfortably
|
||||||
|
# under the staleness thresholds that decide a service is missing, so the
|
||||||
|
# verdict is never limited by how often anyone looked.
|
||||||
|
REFRESH_TTL_SECONDS = 20.0
|
||||||
|
|
||||||
|
# celery inspect is a broker round trip and this sits on a request path, so it
|
||||||
|
# gets a deadline (rule 156). A broker that has stopped answering must make the
|
||||||
|
# roster stale — which is a true statement about the system — not hang the one
|
||||||
|
# page that exists to explain it.
|
||||||
|
INSPECT_TIMEOUT_SECONDS = 2.0
|
||||||
|
|
||||||
|
# Queue set -> the name an operator recognises. Sorted-tuple keys, because the
|
||||||
|
# order celery reports them in is not guaranteed.
|
||||||
|
#
|
||||||
|
# A deployment that slices CELERY_QUEUES differently falls through to the raw
|
||||||
|
# queue list rather than being given a name this table invented for it: a
|
||||||
|
# wrong-but-confident label on a status page is worse than an ugly true one.
|
||||||
|
ROLE_NAMES: dict[tuple[str, ...], str] = {
|
||||||
|
("default", "download", "import", "thumbnail"): "Worker",
|
||||||
|
("maintenance", "scan"): "Scheduler",
|
||||||
|
("ml",): "ML worker",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def role_display_name(queues: tuple[str, ...]) -> str:
|
||||||
|
known = ROLE_NAMES.get(queues)
|
||||||
|
if known:
|
||||||
|
return known
|
||||||
|
return "Worker (" + ", ".join(queues) + ")"
|
||||||
|
|
||||||
|
|
||||||
|
def _inspect_celery_sync() -> dict[tuple[str, ...], dict]:
|
||||||
|
"""celery inspect, grouped by queue set rather than by worker.
|
||||||
|
|
||||||
|
Returns {queue_set: {"hostnames": [...], "active": int}}. Two replicas of
|
||||||
|
one role collapse into one entry on purpose — the question is whether the
|
||||||
|
role is being served, not how many containers exist.
|
||||||
|
"""
|
||||||
|
from ..celery_app import celery as celery_app
|
||||||
|
|
||||||
|
insp = celery_app.control.inspect(timeout=INSPECT_TIMEOUT_SECONDS)
|
||||||
|
active_queues = insp.active_queues() or {}
|
||||||
|
active_tasks = insp.active() or {}
|
||||||
|
|
||||||
|
grouped: dict[tuple[str, ...], dict] = {}
|
||||||
|
for hostname, queues in active_queues.items():
|
||||||
|
key = tuple(sorted({q["name"] for q in queues}))
|
||||||
|
entry = grouped.setdefault(key, {"hostnames": [], "active": 0})
|
||||||
|
entry["hostnames"].append(hostname)
|
||||||
|
entry["active"] += len(active_tasks.get(hostname, []))
|
||||||
|
for entry in grouped.values():
|
||||||
|
entry["hostnames"].sort()
|
||||||
|
return grouped
|
||||||
|
|
||||||
|
|
||||||
|
async def touch_service(
|
||||||
|
session: AsyncSession, *, key: str, kind: str, display_name: str, details: dict
|
||||||
|
) -> None:
|
||||||
|
"""Record that a part checked in just now.
|
||||||
|
|
||||||
|
Upsert rather than read-modify-write: several web processes and several
|
||||||
|
agents can be doing this at once, and the last writer is simply the most
|
||||||
|
recent sighting. `first_seen_at` is deliberately NOT updated — it is the
|
||||||
|
one field that answers "has this ever run", which the learned-roster design
|
||||||
|
depends on.
|
||||||
|
"""
|
||||||
|
stmt = pg_insert(ServiceSeen).values(
|
||||||
|
key=key, kind=kind, display_name=display_name, details=details,
|
||||||
|
)
|
||||||
|
stmt = stmt.on_conflict_do_update(
|
||||||
|
index_elements=[ServiceSeen.key],
|
||||||
|
set_={
|
||||||
|
"kind": stmt.excluded.kind,
|
||||||
|
"display_name": stmt.excluded.display_name,
|
||||||
|
"details": stmt.excluded.details,
|
||||||
|
"last_seen_at": func.now(),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
await session.execute(stmt)
|
||||||
|
|
||||||
|
|
||||||
|
async def refresh_celery_roster(session: AsyncSession) -> None:
|
||||||
|
"""Inspect the broker and record what answered. Never raises.
|
||||||
|
|
||||||
|
A failure here means the roster does not advance, and the rows going stale
|
||||||
|
is then a TRUE report about a broker nobody can reach. Letting the
|
||||||
|
exception out would instead break the health endpoint, which is the one
|
||||||
|
thing that must keep answering when the stack is unwell.
|
||||||
|
"""
|
||||||
|
try:
|
||||||
|
grouped = await asyncio.wait_for(
|
||||||
|
asyncio.to_thread(_inspect_celery_sync),
|
||||||
|
timeout=INSPECT_TIMEOUT_SECONDS * 2,
|
||||||
|
)
|
||||||
|
except Exception:
|
||||||
|
log.warning("service roster: celery inspect failed; roster not refreshed", exc_info=True)
|
||||||
|
return
|
||||||
|
|
||||||
|
for queues, entry in grouped.items():
|
||||||
|
await touch_service(
|
||||||
|
session,
|
||||||
|
key="celery:" + ",".join(queues),
|
||||||
|
kind="celery",
|
||||||
|
display_name=role_display_name(queues),
|
||||||
|
details={
|
||||||
|
"queues": list(queues),
|
||||||
|
"hostnames": entry["hostnames"],
|
||||||
|
"replicas": len(entry["hostnames"]),
|
||||||
|
"active": entry["active"],
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
async def refresh_if_stale(session: AsyncSession) -> None:
|
||||||
|
"""Refresh the celery roster if nobody has for REFRESH_TTL_SECONDS.
|
||||||
|
|
||||||
|
Rate-limited by the data rather than by a lock: the gate is the newest
|
||||||
|
last_seen_at across the celery rows, which every web process can see. Two
|
||||||
|
processes racing through the gate costs one redundant inspect and writes
|
||||||
|
the same values twice, so the benign outcome needs no coordination to
|
||||||
|
prevent.
|
||||||
|
"""
|
||||||
|
newest = (
|
||||||
|
await session.execute(
|
||||||
|
select(func.max(ServiceSeen.last_seen_at)).where(ServiceSeen.kind == "celery")
|
||||||
|
)
|
||||||
|
).scalar_one_or_none()
|
||||||
|
|
||||||
|
if newest is not None:
|
||||||
|
age = (await session.execute(select(func.now()))).scalar_one() - newest
|
||||||
|
if age.total_seconds() < REFRESH_TTL_SECONDS:
|
||||||
|
return
|
||||||
|
|
||||||
|
await refresh_celery_roster(session)
|
||||||
@@ -133,6 +133,20 @@ services:
|
|||||||
CELERY_RESULT_BACKEND: redis://redis:6379/0
|
CELERY_RESULT_BACKEND: redis://redis:6379/0
|
||||||
SECRET_KEY: ${SECRET_KEY:-dev_secret_key_not_for_production_change_me}
|
SECRET_KEY: ${SECRET_KEY:-dev_secret_key_not_for_production_change_me}
|
||||||
LOG_LEVEL: ${LOG_LEVEL:-INFO}
|
LOG_LEVEL: ${LOG_LEVEL:-INFO}
|
||||||
|
# First boot only. FabledCurator refuses to start until the credential
|
||||||
|
# encryption key at /images/secrets/credential_key.b64 exists, and
|
||||||
|
# refuses to create one unless told to — auto-creating is
|
||||||
|
# indistinguishable from a restore that lost ./images/secrets, where it
|
||||||
|
# would mint a key that decrypts nothing and leave an instance that looks
|
||||||
|
# healthy while every paywalled download fails.
|
||||||
|
#
|
||||||
|
# Passed through EXPLICITLY because a variable in `.env` is only used for
|
||||||
|
# ${...} interpolation; it does not reach the container unless it is
|
||||||
|
# named here. Defaulted to empty so the refusal stands for everyone who
|
||||||
|
# has not opted in — the app tests for exactly "1".
|
||||||
|
#
|
||||||
|
# Set it in .env for one `up`, then remove it. See .env.example.
|
||||||
|
CURATOR_BOOTSTRAP_NEW_KEY: ${CURATOR_BOOTSTRAP_NEW_KEY:-}
|
||||||
volumes:
|
volumes:
|
||||||
- ./images:/images
|
- ./images:/images
|
||||||
- ./import:/import
|
- ./import:/import
|
||||||
|
|||||||
@@ -5,9 +5,13 @@
|
|||||||
<img src="/favicon.svg" alt="" class="fc-brand__glyph" width="22" height="22" />
|
<img src="/favicon.svg" alt="" class="fc-brand__glyph" width="22" height="22" />
|
||||||
<span class="fc-brand__text">FabledCurator</span>
|
<span class="fc-brand__text">FabledCurator</span>
|
||||||
</RouterLink>
|
</RouterLink>
|
||||||
<span class="fc-health" :title="health.label">
|
<RouterLink
|
||||||
|
:to="{ name: 'settings', query: { tab: 'system' } }"
|
||||||
|
class="fc-health" :title="health.label"
|
||||||
|
:aria-label="`System health: ${health.label}`"
|
||||||
|
>
|
||||||
<v-icon size="x-small" :color="health.color">{{ health.icon }}</v-icon>
|
<v-icon size="x-small" :color="health.color">{{ health.icon }}</v-icon>
|
||||||
</span>
|
</RouterLink>
|
||||||
<PipelineStatusChip />
|
<PipelineStatusChip />
|
||||||
</div>
|
</div>
|
||||||
|
|
||||||
@@ -64,13 +68,15 @@
|
|||||||
</template>
|
</template>
|
||||||
|
|
||||||
<script setup>
|
<script setup>
|
||||||
import { computed, onBeforeUnmount, onMounted, ref } from 'vue'
|
import { computed, onBeforeUnmount, onMounted, onUnmounted, ref } from 'vue'
|
||||||
import { useRoute } from 'vue-router'
|
import { useRoute } from 'vue-router'
|
||||||
import router, { FRONT_DOOR } from '../router.js'
|
import router, { FRONT_DOOR } from '../router.js'
|
||||||
import { useSystemStore } from '../stores/system.js'
|
import { useSystemStore } from '../stores/system.js'
|
||||||
|
import { useSystemHealthStore } from '../stores/systemHealth.js'
|
||||||
import PipelineStatusChip from './PipelineStatusChip.vue'
|
import PipelineStatusChip from './PipelineStatusChip.vue'
|
||||||
|
|
||||||
const system = useSystemStore()
|
const system = useSystemStore()
|
||||||
|
const healthStore = useSystemHealthStore()
|
||||||
|
|
||||||
// Publish the nav's REAL height as --fc-nav-h so full-height workspaces
|
// Publish the nav's REAL height as --fc-nav-h so full-height workspaces
|
||||||
// (Explore/Subscriptions) and sticky sub-headers pin to it exactly instead of a
|
// (Explore/Subscriptions) and sticky sub-headers pin to it exactly instead of a
|
||||||
@@ -116,15 +122,55 @@ const settingsRoute = computed(() =>
|
|||||||
navRoutes.value.find(r => r.name === 'settings') || null
|
navRoutes.value.find(r => r.name === 'settings') || null
|
||||||
)
|
)
|
||||||
|
|
||||||
|
// The dot beside the brand, and the only ambient signal that something in the
|
||||||
|
// stack has stopped (milestone 365).
|
||||||
|
//
|
||||||
|
// It used to read /api/health — a no-DB liveness check that proves the WEB
|
||||||
|
// container is serving and nothing else. Green there while the worker was dead
|
||||||
|
// is exactly what it looked like, and a green dot next to the product name is
|
||||||
|
// read as "everything is fine". It now reflects the whole-stack verdict.
|
||||||
|
//
|
||||||
|
// Deliberately re-using this element rather than adding a second indicator:
|
||||||
|
// there were already three partial surfaces (this, the pipeline chip, the
|
||||||
|
// Settings Activity tab) and a fourth would have made the question harder to
|
||||||
|
// answer, not easier. This is the one that already occupied the slot.
|
||||||
const health = computed(() => {
|
const health = computed(() => {
|
||||||
if (system.healthy === null) {
|
const overall = healthStore.overall
|
||||||
|
if (overall === null) {
|
||||||
return { icon: 'mdi-circle-outline', color: 'on-surface', label: 'checking…' }
|
return { icon: 'mdi-circle-outline', color: 'on-surface', label: 'checking…' }
|
||||||
}
|
}
|
||||||
if (system.healthy === true) {
|
if (overall === 'ok') {
|
||||||
return { icon: 'mdi-circle', color: 'success', label: 'healthy' }
|
return { icon: 'mdi-circle', color: 'success', label: 'All parts running' }
|
||||||
}
|
}
|
||||||
return { icon: 'mdi-alert-circle', color: 'error', label: 'unreachable' }
|
// Name what is wrong in the tooltip. "Something is unhealthy" sends someone
|
||||||
|
// hunting; "Scheduler has not checked in for 6 min" does not.
|
||||||
|
const worst = healthStore.problems[0]
|
||||||
|
const others = healthStore.problems.length - 1
|
||||||
|
const suffix = others > 0 ? ` (+${others} more)` : ''
|
||||||
|
if (overall === 'down') {
|
||||||
|
return {
|
||||||
|
icon: 'mdi-alert-circle', color: 'error',
|
||||||
|
label: (worst?.detail || 'A part has stopped') + suffix,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if (overall === 'stale') {
|
||||||
|
return {
|
||||||
|
icon: 'mdi-alert', color: 'warning',
|
||||||
|
label: (worst?.detail || 'A part is quiet') + suffix,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return { icon: 'mdi-help-circle-outline', color: 'on-surface', label: 'Health unknown' }
|
||||||
})
|
})
|
||||||
|
|
||||||
|
const HEALTH_POLL_MS = 15_000
|
||||||
|
let healthTimer = null
|
||||||
|
onMounted(() => {
|
||||||
|
healthStore.refresh()
|
||||||
|
healthTimer = setInterval(() => {
|
||||||
|
if (!document.hidden) healthStore.refresh()
|
||||||
|
}, HEALTH_POLL_MS)
|
||||||
|
})
|
||||||
|
onUnmounted(() => { if (healthTimer) clearInterval(healthTimer) })
|
||||||
</script>
|
</script>
|
||||||
|
|
||||||
<style scoped>
|
<style scoped>
|
||||||
@@ -237,7 +283,14 @@ const health = computed(() => {
|
|||||||
display: flex;
|
display: flex;
|
||||||
align-items: center;
|
align-items: center;
|
||||||
flex-shrink: 0;
|
flex-shrink: 0;
|
||||||
|
/* A RouterLink since milestone 365 — it is the path to the Settings System
|
||||||
|
tab, not just an indicator. Reset the anchor so turning a span into a
|
||||||
|
link changed nothing about how the nav reads. */
|
||||||
|
text-decoration: none;
|
||||||
|
color: inherit;
|
||||||
|
border-radius: 50%;
|
||||||
}
|
}
|
||||||
|
.fc-health:hover { background: rgb(var(--v-theme-on-surface) / 0.12); }
|
||||||
.fc-nav-right {
|
.fc-nav-right {
|
||||||
flex: 1 1 0;
|
flex: 1 1 0;
|
||||||
min-width: 0;
|
min-width: 0;
|
||||||
|
|||||||
@@ -0,0 +1,128 @@
|
|||||||
|
<template>
|
||||||
|
<!-- A Settings tab, not a page of its own (operator 2026-09-02): the first
|
||||||
|
cut hung this off the health dot alone, which is a target you have to
|
||||||
|
already suspect something to look for. Settings is where someone goes
|
||||||
|
to ask the instance about itself, so it lives beside Activity. -->
|
||||||
|
<div>
|
||||||
|
<div class="d-flex align-center mb-1">
|
||||||
|
<v-spacer />
|
||||||
|
<span class="fc-sys__checked">
|
||||||
|
{{ store.checkedAt ? `checked ${formatRelative(store.checkedAt)}` : 'checking…' }}
|
||||||
|
</span>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<p class="fc-sys__lede text-body-2 mb-5">
|
||||||
|
Every moving part of FabledCurator and whether it is still checking in.
|
||||||
|
Parts are learned as they appear, so anything that has run at least once
|
||||||
|
stays listed — that is what lets a stopped one be noticed rather than
|
||||||
|
simply vanishing.
|
||||||
|
</p>
|
||||||
|
|
||||||
|
<v-alert
|
||||||
|
v-if="store.lastError" type="error" variant="tonal" density="compact" class="mb-4"
|
||||||
|
>
|
||||||
|
Could not reach FabledCurator: {{ store.lastError }}
|
||||||
|
</v-alert>
|
||||||
|
|
||||||
|
<v-card v-else variant="flat" class="fc-sys__card">
|
||||||
|
<div v-if="!store.parts.length" class="pa-6 text-center fc-sys__muted">
|
||||||
|
Still gathering — this fills in on the first check.
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div
|
||||||
|
v-for="part in store.parts" :key="part.key"
|
||||||
|
class="fc-sys__row" :class="`fc-sys__row--${part.state}`"
|
||||||
|
>
|
||||||
|
<span class="fc-sys__dot" :class="`fc-sys__dot--${part.state}`" />
|
||||||
|
|
||||||
|
<div class="fc-sys__body">
|
||||||
|
<div class="fc-sys__name">
|
||||||
|
{{ part.name }}
|
||||||
|
<span class="fc-sys__kind">{{ kindLabel(part.kind) }}</span>
|
||||||
|
</div>
|
||||||
|
<!-- The sentence, not just a chip. At the moment someone is deciding
|
||||||
|
whether to go and open Portainer, "has not checked in for 6 min"
|
||||||
|
is the thing that answers them. -->
|
||||||
|
<div class="fc-sys__detail">{{ part.detail }}</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="fc-sys__meta">
|
||||||
|
<div v-if="part.last_seen_at" :title="part.last_seen_at">
|
||||||
|
seen {{ formatRelative(part.last_seen_at) }}
|
||||||
|
</div>
|
||||||
|
<div v-if="part.latency_ms != null">{{ part.latency_ms }} ms</div>
|
||||||
|
<div v-if="part.queues?.length" class="fc-sys__queues">{{ part.queues.join(', ') }}</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</v-card>
|
||||||
|
|
||||||
|
<p v-if="store.thresholds" class="fc-sys__foot text-caption mt-4">
|
||||||
|
A part is called stale after
|
||||||
|
{{ Math.round(store.thresholds.stale_after_seconds / 60) }} min without a
|
||||||
|
check-in and treated as stopped after
|
||||||
|
{{ Math.round(store.thresholds.down_after_seconds / 60) }} min. The window
|
||||||
|
is deliberately wide: a rolling deploy briefly runs two of a service and
|
||||||
|
then neither, and an indicator that reddened on every update would stop
|
||||||
|
being read.
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
|
</template>
|
||||||
|
|
||||||
|
<script setup>
|
||||||
|
import { onMounted } from 'vue'
|
||||||
|
|
||||||
|
import { useSystemHealthStore } from '../../stores/systemHealth.js'
|
||||||
|
import { formatRelative } from '../../utils/date.js'
|
||||||
|
|
||||||
|
const store = useSystemHealthStore()
|
||||||
|
|
||||||
|
function kindLabel(kind) {
|
||||||
|
if (kind === 'celery') return 'background worker'
|
||||||
|
if (kind === 'agent') return 'GPU agent'
|
||||||
|
if (kind === 'datastore') return 'datastore'
|
||||||
|
return kind
|
||||||
|
}
|
||||||
|
|
||||||
|
// No timer of its own. TopNav already polls this same pinia store every 15s
|
||||||
|
// for the health dot, and it is mounted on every route this tab is reachable
|
||||||
|
// from — a second interval here would just double the request rate for a 5s
|
||||||
|
// freshness gain. v-window keeps a visited item MOUNTED (hidden, not
|
||||||
|
// destroyed), so a local timer would also have kept firing behind Maintenance.
|
||||||
|
// One refresh on open, so arriving at the tab doesn't wait out the nav's tick.
|
||||||
|
onMounted(() => { store.refresh() })
|
||||||
|
</script>
|
||||||
|
|
||||||
|
<style scoped>
|
||||||
|
.fc-sys__lede, .fc-sys__muted, .fc-sys__checked, .fc-sys__foot {
|
||||||
|
color: rgb(var(--v-theme-on-surface) / 0.66);
|
||||||
|
}
|
||||||
|
.fc-sys__checked { font-size: 0.78rem; }
|
||||||
|
.fc-sys__card { background: rgb(var(--v-theme-on-surface) / 0.04); }
|
||||||
|
|
||||||
|
.fc-sys__row {
|
||||||
|
display: flex; align-items: center; gap: 12px;
|
||||||
|
padding: 12px 16px;
|
||||||
|
border-bottom: 1px solid rgb(var(--v-theme-on-surface) / 0.08);
|
||||||
|
}
|
||||||
|
.fc-sys__row:last-child { border-bottom: 0; }
|
||||||
|
|
||||||
|
.fc-sys__dot { width: 9px; height: 9px; border-radius: 50%; flex: 0 0 auto; }
|
||||||
|
.fc-sys__dot--ok { background: rgb(var(--v-theme-success)); }
|
||||||
|
.fc-sys__dot--stale { background: rgb(var(--v-theme-warning)); }
|
||||||
|
.fc-sys__dot--down { background: rgb(var(--v-theme-error)); }
|
||||||
|
.fc-sys__dot--unknown { background: rgb(var(--v-theme-on-surface) / 0.35); }
|
||||||
|
|
||||||
|
.fc-sys__body { min-width: 0; flex: 1 1 auto; }
|
||||||
|
.fc-sys__name { font-weight: 600; }
|
||||||
|
.fc-sys__kind {
|
||||||
|
margin-left: 8px; font-weight: 400; font-size: 0.72rem; text-transform: uppercase;
|
||||||
|
letter-spacing: 0.04em; color: rgb(var(--v-theme-on-surface) / 0.5);
|
||||||
|
}
|
||||||
|
.fc-sys__detail { font-size: 0.82rem; color: rgb(var(--v-theme-on-surface) / 0.72); }
|
||||||
|
|
||||||
|
.fc-sys__meta {
|
||||||
|
text-align: right; font-size: 0.75rem; flex: 0 0 auto;
|
||||||
|
font-variant-numeric: tabular-nums; color: rgb(var(--v-theme-on-surface) / 0.6);
|
||||||
|
}
|
||||||
|
.fc-sys__queues { opacity: 0.75; }
|
||||||
|
</style>
|
||||||
@@ -45,6 +45,13 @@ const routes = [
|
|||||||
|
|
||||||
// Settings — config, pinned to the right of the nav (TopNav special-cases it).
|
// Settings — config, pinned to the right of the nav (TopNav special-cases it).
|
||||||
{ path: '/settings', name: 'settings', component: SettingsView, meta: { title: 'Settings', stickyChrome: true } },
|
{ path: '/settings', name: 'settings', component: SettingsView, meta: { title: 'Settings', stickyChrome: true } },
|
||||||
|
// System health is a Settings TAB, not a route of its own (operator
|
||||||
|
// 2026-09-02). It first shipped as /system reachable only from the health
|
||||||
|
// dot, which is a target you have to already suspect something to go
|
||||||
|
// looking for. Settings is where someone goes to ask the instance about
|
||||||
|
// itself. The path stays as a redirect so the dot's old link, and any
|
||||||
|
// bookmark from that build, still land somewhere real.
|
||||||
|
{ path: '/system', name: 'system', redirect: () => ({ name: 'settings', query: { tab: 'system' } }) },
|
||||||
|
|
||||||
// The old standalone paths now redirect into the Browse hub, preserving any
|
// The old standalone paths now redirect into the Browse hub, preserving any
|
||||||
// deep-link query (e.g. /posts?post_id=N → /browse?tab=posts&post_id=N). The
|
// deep-link query (e.g. /posts?post_id=N → /browse?tab=posts&post_id=N). The
|
||||||
|
|||||||
@@ -4,7 +4,12 @@ import { useApi } from '../composables/useApi.js'
|
|||||||
|
|
||||||
export const useSystemStore = defineStore('system', () => {
|
export const useSystemStore = defineStore('system', () => {
|
||||||
const api = useApi()
|
const api = useApi()
|
||||||
const healthy = ref(null) // null=unknown, true=ok, false=down
|
// NOT what the nav dot reads any more (milestone 365): that is the
|
||||||
|
// whole-stack verdict in systemHealth.js. /api/health only proves the web
|
||||||
|
// container is serving, which is why a green dot here sat happily beside a
|
||||||
|
// dead worker. refreshHealth() is still called — it is also how build/version
|
||||||
|
// info arrives — so this stays as its by-product rather than its purpose.
|
||||||
|
const healthy = ref(null)
|
||||||
// What the instance says it is. Since milestone 318 stopped publishing
|
// What the instance says it is. Since milestone 318 stopped publishing
|
||||||
// version image tags, this is the only answer to "which build is this?" —
|
// version image tags, this is the only answer to "which build is this?" —
|
||||||
// there is no registry name left to check it against.
|
// there is no registry name left to check it against.
|
||||||
|
|||||||
@@ -0,0 +1,49 @@
|
|||||||
|
import { defineStore } from 'pinia'
|
||||||
|
import { computed, ref } from 'vue'
|
||||||
|
|
||||||
|
import { useApi } from '../composables/useApi.js'
|
||||||
|
|
||||||
|
// Whole-stack health: is every part of FabledCurator running (milestone 365)?
|
||||||
|
//
|
||||||
|
// Distinct from `system.js`, which polls /api/health — a no-DB liveness check
|
||||||
|
// that only proves the web container is serving. That endpoint answers "can I
|
||||||
|
// reach the API"; this one answers "is anything broken", which is the question
|
||||||
|
// a green dot beside the brand was already being read as answering.
|
||||||
|
//
|
||||||
|
// Also distinct from `systemActivity.js`, which is about what the pipeline is
|
||||||
|
// DOING — queue depths, running tasks, failures. Running and alive are
|
||||||
|
// different questions and they fail independently: a perfectly idle stack with
|
||||||
|
// a dead worker looks identical to a healthy one on the activity surfaces.
|
||||||
|
export const useSystemHealthStore = defineStore('systemHealth', () => {
|
||||||
|
const api = useApi()
|
||||||
|
|
||||||
|
const overall = ref(null) // null until the first answer: unknown ≠ ok
|
||||||
|
const parts = ref([])
|
||||||
|
const checkedAt = ref(null)
|
||||||
|
const thresholds = ref(null) // server-owned, so the UI keeps no second copy
|
||||||
|
const lastError = ref(null)
|
||||||
|
|
||||||
|
async function refresh() {
|
||||||
|
try {
|
||||||
|
const body = await api.get('/api/system/health')
|
||||||
|
overall.value = body.overall
|
||||||
|
parts.value = body.parts || []
|
||||||
|
checkedAt.value = body.checked_at
|
||||||
|
thresholds.value = body.thresholds || null
|
||||||
|
lastError.value = null
|
||||||
|
} catch (e) {
|
||||||
|
// The endpoint is built never to fail because a dependency failed, so a
|
||||||
|
// throw here means the API itself is unreachable — which is its own kind
|
||||||
|
// of unhealthy and must not be shown as "ok".
|
||||||
|
lastError.value = e.message
|
||||||
|
overall.value = 'unknown'
|
||||||
|
}
|
||||||
|
return overall.value
|
||||||
|
}
|
||||||
|
|
||||||
|
// The parts worth naming in a tooltip — everything that is not ok, worst
|
||||||
|
// first. The endpoint already sorts that way.
|
||||||
|
const problems = computed(() => parts.value.filter(p => p.state !== 'ok'))
|
||||||
|
|
||||||
|
return { overall, parts, checkedAt, thresholds, lastError, problems, refresh }
|
||||||
|
})
|
||||||
@@ -14,6 +14,7 @@
|
|||||||
style="position: sticky; top: var(--fc-nav-h, 64px); z-index: 4;"
|
style="position: sticky; top: var(--fc-nav-h, 64px); z-index: 4;"
|
||||||
>
|
>
|
||||||
<v-tab value="overview">Overview</v-tab>
|
<v-tab value="overview">Overview</v-tab>
|
||||||
|
<v-tab value="system">System</v-tab>
|
||||||
<v-tab value="activity">Activity</v-tab>
|
<v-tab value="activity">Activity</v-tab>
|
||||||
<v-tab value="cleanup">Cleanup</v-tab>
|
<v-tab value="cleanup">Cleanup</v-tab>
|
||||||
<v-tab value="maintenance">Maintenance</v-tab>
|
<v-tab value="maintenance">Maintenance</v-tab>
|
||||||
@@ -42,6 +43,13 @@
|
|||||||
</v-alert>
|
</v-alert>
|
||||||
</v-window-item>
|
</v-window-item>
|
||||||
|
|
||||||
|
<!-- Is every part of the stack still running (milestone 365). Sits
|
||||||
|
beside Activity deliberately: Activity answers "what is the queue
|
||||||
|
doing", this answers "is anything left to do it". -->
|
||||||
|
<v-window-item value="system">
|
||||||
|
<SystemHealthTab />
|
||||||
|
</v-window-item>
|
||||||
|
|
||||||
<v-window-item value="activity">
|
<v-window-item value="activity">
|
||||||
<SystemActivityTab @open-maintenance="tab = 'maintenance'" />
|
<SystemActivityTab @open-maintenance="tab = 'maintenance'" />
|
||||||
</v-window-item>
|
</v-window-item>
|
||||||
@@ -73,18 +81,24 @@
|
|||||||
</template>
|
</template>
|
||||||
|
|
||||||
<script setup>
|
<script setup>
|
||||||
import { onMounted, onUnmounted, ref, watch } from 'vue'
|
import { onMounted, onUnmounted, watch } from 'vue'
|
||||||
import { useSystemStore } from '../stores/system.js'
|
import { useSystemStore } from '../stores/system.js'
|
||||||
import SystemStatsCards from '../components/settings/SystemStatsCards.vue'
|
import SystemStatsCards from '../components/settings/SystemStatsCards.vue'
|
||||||
import SystemActivitySummary from '../components/settings/SystemActivitySummary.vue'
|
import SystemActivitySummary from '../components/settings/SystemActivitySummary.vue'
|
||||||
import SystemActivityTab from '../components/settings/SystemActivityTab.vue'
|
import SystemActivityTab from '../components/settings/SystemActivityTab.vue'
|
||||||
|
import SystemHealthTab from '../components/settings/SystemHealthTab.vue'
|
||||||
import GpuActivityPanel from '../components/settings/GpuActivityPanel.vue'
|
import GpuActivityPanel from '../components/settings/GpuActivityPanel.vue'
|
||||||
import DownloadsActivityPanel from '../components/settings/DownloadsActivityPanel.vue'
|
import DownloadsActivityPanel from '../components/settings/DownloadsActivityPanel.vue'
|
||||||
import MaintenancePanel from '../components/settings/MaintenancePanel.vue'
|
import MaintenancePanel from '../components/settings/MaintenancePanel.vue'
|
||||||
import CleanupView from './CleanupView.vue'
|
import CleanupView from './CleanupView.vue'
|
||||||
|
import { useTabQuery } from '../composables/useTabQuery.js'
|
||||||
import { useMLStore } from '../stores/ml.js'
|
import { useMLStore } from '../stores/ml.js'
|
||||||
|
|
||||||
const tab = ref('overview')
|
// ?tab= sync (the same composable Browse/Subscriptions use) so a tab can be
|
||||||
|
// linked TO — the health dot beside the brand points at ?tab=system, and the
|
||||||
|
// old /system path redirects there.
|
||||||
|
const VALID_TABS = ['overview', 'system', 'activity', 'cleanup', 'maintenance']
|
||||||
|
const { tab } = useTabQuery(VALID_TABS, 'overview')
|
||||||
const system = useSystemStore()
|
const system = useSystemStore()
|
||||||
const mlStore = useMLStore()
|
const mlStore = useMLStore()
|
||||||
|
|
||||||
|
|||||||
@@ -40,6 +40,14 @@ describe('router', () => {
|
|||||||
expect(router.currentRoute.value.query.post_id).toBe('7')
|
expect(router.currentRoute.value.query.post_id).toBe('7')
|
||||||
})
|
})
|
||||||
|
|
||||||
|
it('/system redirects into the Settings System tab', async () => {
|
||||||
|
// It shipped as a standalone page for one build; the health dot and any
|
||||||
|
// bookmark from it must still land on the surface, which is now a tab.
|
||||||
|
await router.push('/system')
|
||||||
|
expect(router.currentRoute.value.name).toBe('settings')
|
||||||
|
expect(router.currentRoute.value.query.tab).toBe('system')
|
||||||
|
})
|
||||||
|
|
||||||
it('series-read is an immersive route', () => {
|
it('series-read is an immersive route', () => {
|
||||||
const r = router.resolve('/series/5/read')
|
const r = router.resolve('/series/5/read')
|
||||||
expect(r.name).toBe('series-read')
|
expect(r.name).toBe('series-read')
|
||||||
|
|||||||
Reference in New Issue
Block a user