Compare commits

...
15 Commits
Author SHA1 Message Date
bvandeusen a8fdbd86bc Merge pull request 'FabledCurator can now tell you one of its own parts has stopped' (#250) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 4s
Build images / build-agent (push) Successful in 8s
Build images / build-ml (push) Successful in 13s
Build images / build-web (push) Successful in 13s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / frontend-build (push) Successful in 20s
CI / backend-lint-and-test (push) Successful in 33s
CI / integration (push) Successful in 1m53s
2026-09-02 18:01:38 -04:00
bvandeusenandClaude Opus 5 5084ba666b feat: the dot beside the brand now means the whole stack (milestone 365 step 4)
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Successful in 5s
CI / extension-version (push) Successful in 5s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 9s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 31s
Build images / build-web (push) Successful in 55s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / integration (push) Successful in 1m43s
The ask was a surface AND a path. The path is the part that was missing —
everything that could answer "is it running" lived inside Settings, which you
only open once you already suspect something.

**Re-used the indicator that already existed rather than adding a fourth.**
There were three partial surfaces: TopNav's health dot, PipelineStatusChip's
pulse, and the Settings Activity tab. None answered "is every part alive", and
a fourth would have made the question harder to answer, not easier.

TopNav's dot read /api/health — a no-DB liveness check proving only that the
WEB container is serving. Green there while a worker was dead is exactly what
it looked like, and a green dot beside the product name gets read as
"everything is fine". It now reflects the whole-stack verdict, and it is a
link: the place someone already looks when they suspect something is now also
the way to the detail.

The tooltip names the actual problem. "Scheduler has not checked in for 6 min"
sends someone somewhere; "something is unhealthy" sends them hunting.

/system is deliberately NOT in the nav row — TopNav builds that from routes
with a meta.title, and a sixth top-level tab for a page visited twice a year
costs more attention than it returns. It is reached from the dot.

The page lists every learned part with its state as a sentence rather than a
chip, and prints the staleness thresholds it was judged by, taken from the
endpoint so the UI keeps no second copy of them. PipelineStatusChip still
hand-rolls its own 3-minute scheduler window; that is now a duplicate of a
threshold the server owns, and worth collapsing once this has been watched
working.

The stores stay separate on purpose: system.js is "can I reach the API",
systemActivity.js is "what is the pipeline doing", systemHealth.js is "is
anything broken". Running and alive fail independently — an idle stack with a
dead worker looks identical to a healthy one on every activity surface, which
is the whole reason this milestone exists.

Not yet verified against a real stopped service. Rule 12 keeps a local stack
out of it, and frontend CI has no Vue type-check or visual regression, so this
needs an operator look rather than a green lane.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 17:20:55 -04:00
bvandeusenandClaude Opus 5 fe4e0f2b71 feat: /api/system/health — one verdict for the whole stack (milestone 365 step 3)
CI / lint (push) Successful in 5s
Build images / sign-extension (push) Successful in 5s
Build images / build-agent (push) Successful in 8s
CI / extension-version (push) Successful in 4s
CI / frontend-build (push) Successful in 20s
CI / backend-lint-and-test (push) Successful in 35s
Build images / build-web (push) Successful in 1m1s
Build images / smoke-web (push) Skipped
CI / integration (push) Successful in 1m51s
Build images / build-ml (push) Successful in 1m59s
Build images / promote (push) Skipped
The single endpoint the nav indicator and the System page will both read.
Composing a verdict is this module's job, not the UI's.

Two kinds of part, answered differently. LEARNED — celery roles and the GPU
agent, out of service_seen, where the question is "how long since it checked
in" and the answer can be "it has not". PROBED — Postgres and Redis, always
expected, never learned, because a last-seen for them would be actively
misleading: that Redis answered thirty seconds ago says nothing about now.

**The endpoint must never fail because something it checks has failed.** That
inversion is easy to write by accident and it destroys the feature exactly
when it is needed — a 500 when Redis is down instead of `redis: down`. Every
probe is wrapped, every wait carries a deadline (rule 156), and the roster
refresh swallows its own errors. The worst case is a part reported `unknown`,
which is a true statement about the system.

Postgres is probed first and gates the rest, because if it is unreachable
nothing else can be read — and "the database is down" is the most useful
single thing this can ever say.

The staleness thresholds are the design risk, not the code, and they are
deliberately generous: 90s to doubt, 300s to disbelieve. The constraint is a
deploy rather than a crash — `docker compose up -d` rolls start-first, so a
role is briefly served by two containers and then by neither while the old one
drains. Thresholds tight enough to catch a crash in seconds would paint the
page red on every update, and an alarm that cries wolf on every deploy is one
nobody reads. Tune down only after watching a real deploy pass through. The
numbers ship in the response so the UI can explain a `stale` without keeping a
second copy of them.

States are described in sentences rather than left as chips: "Scheduler has
not checked in for 6 min — treat it as stopped" is what someone needs at the
moment they are deciding whether to go and open Portainer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 17:17:22 -04:00
bvandeusenandClaude Opus 5 dc8af8b1a7 feat: a learned roster, so a stopped part is observable (milestone 365 steps 1-2)
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 30s
Build images / build-web (push) Successful in 55s
Build images / smoke-web (push) Skipped
Build images / build-ml (push) Successful in 1m45s
Build images / promote (push) Skipped
CI / integration (push) Successful in 1m49s
Nothing in FabledCurator knew what was SUPPOSED to be running. `celery
inspect` reports the workers that ANSWER, so a dead worker was a shorter list
rather than a red light, and grep for any notion of expected services returned
nothing. That is why Portainer was the only place an operator could see it:
Portainer knows the intended set.

`service_seen` is the memory that makes an absence observable — every part
that has checked in, and when it last did.

**Keyed on the queue set, not the worker hostname.** Celery's worker names
here are `celery@<container id>`, minted fresh on every deploy. Keyed on those,
this table would record a death and a birth every time the stack updates — and
a status page that goes red on every deploy is a status page nobody reads,
which is worse than not having one. CELERY_QUEUES is assigned per role in
compose and survives container replacement, so it is the stable identity. Two
replicas of a role are therefore ONE row, which is right: the question is
whether the role is served, not how many containers exist.

The GPU agent is keyed on agent_id, the identity its lease protocol already
uses. gpu.py received it on both lease and heartbeat and threw it away — an
idle agent with nothing to lease left no trace and was indistinguishable from
one switched off a week ago. Now recorded on the calls that were already
happening.

**Who observes, corrected from the plan.** The plan said "record from the
existing inspect path", which would only run when someone opened the Activity
tab. Two other candidates and why they lost:

- A beat sweep. If the scheduler dies the sweep stops, every row goes stale,
  and the page says everything is down when one thing is. An alarm that cannot
  distinguish "a part died" from "the observer died" is worse than none.
- A background task in web. hypercorn runs --workers 4, so that is four
  concurrent inspect loops per container, forever.

Taken instead: refresh on demand, rate-limited by the newest last_seen_at that
every process can already see. The observer is then the thing serving the page
— if web is down you get a browser error, not a confidently green page — and
it self-limits with no coordination, since a race costs one redundant inspect
that writes identical values.

Migration 0090 is the first written on the collapsed baseline (milestone 328),
so it is also the first evidence the chain steps FORWARD from 0089 rather than
merely reproducing the schema. No secondary indexes: one row per moving part
means every read is a handful of rows, and #3301 is the record of what
speculative indexes cost.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 17:15:46 -04:00
bvandeusen 0421fd3109 Merge pull request 'The weekly refresh now publishes only what a gate has proven — and a new install can start' (#249) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 4s
Build images / build-ml (push) Successful in 8s
Build images / build-agent (push) Successful in 9s
Build images / build-web (push) Successful in 7s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / frontend-build (push) Successful in 22s
extension / lint (push) Successful in 19s
CI / backend-lint-and-test (push) Successful in 31s
CI / integration (push) Successful in 1m49s
2026-09-02 16:31:55 -04:00
bvandeusenandClaude Opus 5 131237143b Revert "test: force the smoke gate to fail, to watch it block a publish"
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 6s
Build images / build-ml (push) Successful in 8s
Build images / build-agent (push) Successful in 9s
CI / frontend-build (push) Successful in 22s
extension / lint (push) Successful in 21s
Build images / build-web (push) Successful in 6s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / backend-lint-and-test (push) Successful in 34s
CI / integration (push) Successful in 2m0s
extension / lint (pull_request) Successful in 20s
The gate held. Run 5320, dispatched with the forced failure in place:

  build-web    success   (candidate published)
  build-ml     success
  build-agent  success
  smoke-web    FAILED
  promote      skipped
  run          failure

And the three channel tags did not move:

  fabledcurator        33d3d8332f74 -> 33d3d8332f74
  fabledcurator-ml     e94a5435cb45 -> e94a5435cb45
  fabledcurator-agent  bae27d34d811 -> bae27d34d811

So a refresh that breaks something now leaves :latest naming the build that
works, which is the property milestone 362 exists to establish. The rejected
candidate is still published under :refresh-candidate, so whoever reads the
red job on Monday can pull the exact image that failed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 16:23:49 -04:00
bvandeusenandClaude Opus 5 59d27ef76e test: force the smoke gate to fail, to watch it block a publish
CI / lint (push) Successful in 4s
Build images / sign-extension (push) Successful in 4s
CI / extension-version (push) Successful in 3s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 8s
CI / frontend-build (push) Successful in 19s
extension / lint (push) Successful in 18s
Build images / build-web (push) Successful in 7s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / backend-lint-and-test (push) Successful in 31s
CI / integration (push) Successful in 1m57s
TEMPORARY, reverted in the next commit. Milestone 362's verification section
requires the gate to be seen rejecting a build — a gate nobody has watched
reject anything is a gate nobody knows is wired up. Every real check passes,
so the rejection has to be forced.

Under test is the job dependency, not the assertions: a failed smoke-web must
skip the promote job, and the three :latest tags must still name the digests
they named before the run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 16:22:01 -04:00
bvandeusenandClaude Opus 5 f630e50e75 ci: the refresh publishes only what the gate passed (#3265 milestone step 4)
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 5s
Build images / build-ml (push) Successful in 8s
Build images / build-agent (push) Successful in 9s
Build images / build-web (push) Successful in 6s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
extension / lint (push) Successful in 19s
CI / frontend-build (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 34s
CI / integration (push) Successful in 1m53s
The gate reported a verdict nothing consulted. Now it decides.

The promote moved out of the three build jobs into its own `promote` job,
because the verdict cannot exist until build-web has finished and the promote
used to run inside it. `needs: [build-web, build-ml, build-agent, smoke-web]`
is the whole mechanism: a failed smoke skips the promote, so a refresh that
broke something leaves :latest naming the build that works. "The refresh
failed" and "production is broken" must not be the same event.

A SKIPPED smoke also skips it, and that is the case that matters most. On run
5290 the gate silently skipped itself — job-level `if:` cannot read the env
context — and a design where only a FAILED gate blocks would have published
unverified images while reporting success. Not running is not the same as
passing, and today produced two separate bugs of exactly that shape (#3414,
and the smoke-web skip).

All three images now promote together or not at all. They are one stack:
build.yml already refuses to publish a :dev web image beside a stale :dev ml
because the mismatch only surfaces as a runtime failure, and a refresh that
published ml while withholding web would be that same trap reached through the
gate. Stated plainly in the job comment: the gate covers web only, so ml and
agent are held to web's verdict rather than their own. That is the
conservative direction, not equivalent evidence, and should not be read as if
it were.

Three near-identical promote steps collapsed into one loop. A partial failure
now says which images moved and that the state is inconsistent, rather than
leaving that to be inferred — the promote is idempotent and the candidates are
still published, so the instruction is simply to re-run.

Also removed the now-dead `promote` output from the ml and agent reuse steps.
Only build-web's is read (as outputs.candidate); two more copies nothing
consults is the kind of thing that reads as load-bearing a year later.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 16:21:27 -04:00
bvandeusenandClaude Opus 5 86abaf0b94 docs: a new install could not start, and nothing told anyone why (#3422)
CI / lint (push) Successful in 4s
Build images / sign-extension (push) Successful in 4s
CI / extension-version (push) Successful in 3s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 8s
Build images / build-web (push) Successful in 6s
Build images / smoke-web (push) Skipped
CI / frontend-build (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 31s
CI / integration (push) Successful in 1m46s
The install path milestone 328 wrote produces a web container that exits on
boot. entrypoint.sh runs alembic, then app construction raises:

  MissingCredentialKey: Fernet key file not found at
  /images/secrets/credential_key.b64. For first-time setup, set
  CURATOR_BOOTSTRAP_NEW_KEY=1.

That variable appeared in no README, no .env.example and no compose file —
only in backend/. So a stranger following the documented steps got an app
that does not start and an error with no context. Found by the milestone-362
smoke gate on its first real run (#3422).

The product behaviour stays exactly as it is. credential_crypto refuses to
mint a key because the 2026-06-02 audit found a partial restore — database
back, ./images/secrets lost — silently generating a fresh one and producing a
healthy-looking instance where every authenticated download failed AUTH_ERROR.
Failing fast is right; not saying so is the bug.

So: .env.example carries the variable in its own FIRST BOOT ONLY section with
the reasoning and an instruction to delete the line afterwards, and README's
First run leads with it, because "the app will not start" belongs before "the
ML worker downloads weights". Both say to back up ./images/secrets/ alongside
the database, which is the part that costs real data if it is learned late.

**compose had to change too, and this is the part that would have shipped a
second broken instruction.** A variable in `.env` is only used for ${...}
interpolation — it does not reach the container unless the service names it.
Telling people to set it in .env, without that, would have documented a step
that does nothing. Added to the shared app_env anchor, defaulted to empty so
the refusal still stands for everyone who has not opted in.

Not taken: auto-bootstrapping when the credential table is empty, which would
remove the manual step entirely and keep the audit's protection for restores.
That is the better product and it is a code change with a predicate that has
to be exactly right; this is the smallest correct fix, and #3422 stays open
for the other one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 16:19:28 -04:00
bvandeusenandClaude Opus 5 4815040d74 ci: the smoke gate found a real one on its first run — and had two bugs of its own
Build images / sign-extension (push) Successful in 6s
CI / lint (push) Successful in 6s
CI / extension-version (push) Successful in 4s
Build images / build-ml (push) Successful in 11s
Build images / build-agent (push) Successful in 13s
extension / lint (push) Successful in 24s
CI / frontend-build (push) Successful in 28s
CI / backend-lint-and-test (push) Successful in 35s
Build images / build-web (push) Successful in 8s
Build images / smoke-web (push) Skipped
CI / integration (push) Successful in 2m5s
Run 5296 was `smoke-web`'s first genuine execution. Checks 1 and 2 passed:
alembic built the schema from empty inside the image, all five apt binaries
resolved, and the application's own Thumbnailer produced JPEG, PNG-with-alpha,
WebP and an ffmpeg video frame against the image's libraries. Check 3 failed,
and the trap's log dump said exactly why:

  MissingCredentialKey: Fernet key file not found at
  /images/secrets/credential_key.b64. For first-time setup, set
  CURATOR_BOOTSTRAP_NEW_KEY=1.

That is the product being right. credential_crypto refuses to mint a key
unless someone opts in, because the 2026-06-02 audit found a partial restore
(DB back, /images/secrets/ lost) silently generating a fresh one and leaving a
working-looking system where every authenticated download failed AUTH_ERROR.

It is also a first-run blocker for milestone 328, filed as #3422: the variable
appears in no README, no .env.example and no compose file, so the install path
that milestone just finished writing produces a container that exits on boot.
Not fixed here — the fix trades safety against friction and is the operator's
call.

Two defects in the gate itself, both surfaced by the same run:

- A throwaway CI instance IS first-time setup, so it now passes
  CURATOR_BOOTSTRAP_NEW_KEY=1. The check was asserting a condition no fresh
  container can satisfy.

- The health loop polled a dead container for 3m35s. Docker had already
  recycled its IP, so the replies were a baffling mix of connection-refused
  and 5s timeouts from whatever took the address next. It now checks
  `.State.Running` each iteration and fails immediately with the container's
  log. The trap had the real answer the whole time; this stops burying it
  under four minutes of noise.

Also corrected a message claiming a 120s budget: 60 iterations of up to 5s
connect plus 2s sleep is nearer seven minutes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 16:00:18 -04:00
bvandeusenandClaude Opus 5 81b7b6f308 ci: smoke-web never ran — a job's if: cannot read the env context
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 4s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 8s
CI / frontend-build (push) Successful in 22s
extension / lint (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 32s
Build images / build-web (push) Successful in 8s
Build images / smoke-web (push) Skipped
CI / integration (push) Successful in 1m51s
Run 5290 dispatched a refresh. Everything worked: the guard fired, the build
published the candidate, the promote pointed :latest at it. And `smoke-web`
reported conclusion "skipped", with no steps and no log.

Its condition was `if: env.IS_REFRESH == 'true'`. The env context is available
to STEP conditions and step bodies but never to a job's own `if:`, and an
unresolvable context there evaluates to empty rather than erroring. So the
gate skipped itself, silently, on the one run that existed to exercise it.

Second silent-skip of this family today, after #3414. Same shape both times:
something evaluated false, nothing failed, and the run reported success. It is
worth naming the pattern — on this pipeline, "green" and "ran" are different
claims, and the steps' own conclusions are the only place the difference shows.

Fixed by keying off a job output rather than re-deriving the trigger:
build-web now exposes the reuse step's `promote` decision as `outputs.candidate`
and smoke-web consumes it. That is better than duplicating the expression:
it is the same single decision the build, the XPI download and the promote all
take already — build.yml's own "one decision drives everything downstream" —
and it asserts the thing smoke-web actually depends on, that a candidate was
published, rather than restating the reason one would be.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 15:53:40 -04:00
bvandeusen d01f33dea6 Merge pull request 'A gate for the weekly refresh, and the lever that makes it testable' (#248) from dev into main
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 3s
CI / extension-version (push) Successful in 2s
Build images / build-ml (push) Successful in 6s
Build images / build-agent (push) Successful in 7s
Build images / build-web (push) Successful in 6s
Build images / smoke-web (push) Skipped
CI / frontend-build (push) Successful in 19s
extension / lint (push) Successful in 17s
CI / backend-lint-and-test (push) Successful in 31s
CI / integration (push) Successful in 1m42s
2026-09-02 15:51:50 -04:00
bvandeusenandClaude Opus 5 bfa9fd678b ci: smoke the refreshed image against real Postgres and Redis
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 3s
CI / extension-version (push) Successful in 4s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 9s
Build images / build-web (push) Successful in 7s
Build images / smoke-web (push) Skipped
extension / lint (push) Successful in 20s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 32s
CI / integration (push) Successful in 1m48s
extension / lint (pull_request) Successful in 20s
Milestone 362 step 3. This is the gate the weekly base refresh never had.

`ci.yml` cannot be that gate, and the reason matters more than the fix. Its
lanes run on ci-python:3.14 and install requirements.txt — a base refresh
changes neither, so all five stay green through a bump that breaks the product.
What a refresh re-resolves is the Dockerfile's apt layer:

    ffmpeg unar libpq5 postgresql-client zstd megatools
    libjpeg62-turbo libwebp7 libpng16-16 ca-certificates

Unpinned, every build, and nothing else in this repo looks at it. That line is
the dependency creep; it is also precisely what the test suite structurally
cannot observe, since the suite never runs inside the image and the image
carries no tests and no pytest.

So `smoke-web` runs the CANDIDATE IMAGE against real service containers:

  1. `alembic upgrade head` on an empty database — the image's own libpq and
     psycopg, and the same call entrypoint.sh makes before it serves anything,
     so a failure here is a failure to boot.
  2. The apt binaries, then the application's own `Thumbnailer` — JPEG, PNG
     with alpha, WebP, and a video frame through ffmpeg. `Thumbnailer` needs no
     database and no app context, so the check exercises real product code
     rather than a proxy for it. `ffmpeg -version` exiting 0 would pass while a
     codec removal broke every thumbnail in the library.
  3. The web role boots and answers /api/health.

Every failure names the package it implicates. This fires on a Sunday,
unattended, about a change nobody made deliberately — "assertion failed" a week
later teaches nobody anything.

The script is piped over stdin rather than bind-mounted: the workspace is a
docker volume belonging to the job's own container, so a host bind of $PWD does
not resolve for a sibling. Container logs are dumped only on failure, and the
trap re-exits with the real status rather than the status of `docker rm`.

Deliberately NOT gating the promote yet — that is step 4. Landing the gate and
the thing it gates together would mean the first time anyone saw this job run
would also be the first time it could stop a publish.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 15:45:10 -04:00
bvandeusenandClaude Opus 5 24a2b70a5a ci: a boolean input never equals the string 'true'
Build images / sign-extension (push) Successful in 4s
Build images / build-ml (push) Successful in 6s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 23s
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 2s
CI / backend-lint-and-test (push) Successful in 35s
Build images / build-web (push) Successful in 8s
CI / integration (push) Successful in 1m48s
The refresh lever did not work, and the way it did not work is the point.

Run 5270 dispatched with refresh=true. Its log:

  expression '(github.event_name == 'schedule'
               || github.event.inputs.refresh == 'true') && 'true' || 'false''
    evaluated to '%!t(string=false)'
  trigger: event=workflow_dispatch IS_REFRESH='false' BUILD_REF='refs/heads/dev'
  trigger: raw inputs refresh='true' force_build='false'

The input arrived as true and the comparison still said false. `type: boolean`
delivers a real boolean, and GitHub expression semantics cast operands to
numbers when their types differ — so `true == 'true'` compares 1 against NaN.
My comment on the previous commit asserted the opposite, that Forgejo delivers
inputs as strings, and asserted it without checking.

The run went GREEN with every step skipped, because a refresh that evaluates
false is indistinguishable from an ordinary push. A lever that silently does
nothing is worse than no lever: it would have been trusted.

Normalised through format(), which is representation-independent — a boolean
true and a string 'true' both render 'true'. That is also why force_build was
never bitten: it passes its raw value into an env var and compares in the
shell, where everything is a string already. format() buys the same thing at
expression level, which is where a step `if:` needs the answer.

The diagnostic from the previous commit stays. It is what turned this from a
guess into a measurement, and it is the only thing that would catch the same
class of failure next time.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 14:59:48 -04:00
bvandeusenandClaude Opus 5 2c88ad3efb ci: report the raw and normalised trigger values
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
CI / extension-version (push) Successful in 2s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 7s
CI / backend-lint-and-test (push) Successful in 31s
Build images / build-web (push) Successful in 9s
extension / lint (push) Successful in 16s
CI / frontend-build (push) Successful in 22s
CI / integration (push) Successful in 2m39s
The refresh dispatch on run 5265 went green with every step skipped: the
main-only guard did not fire, checkout took dev, and the reuse step read
IS_REFRESH as false. So both workflow-level expressions evaluated false while
the identical accessor works for force_build, which compares its value in the
shell rather than in an expression.

That is a guess until it is measured, and the failure is silent by
construction — a refresh that evaluates false behaves exactly like an ordinary
push and reports success. This prints the raw input beside the normalised
value in the step that already exists to say what a run derived, so the two
disagreeing is visible rather than inferred.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2026-09-02 14:58:20 -04:00
17 changed files with 1386 additions and 243 deletions
+28
View File
@@ -34,6 +34,34 @@ DB_PASSWORD=
SECRET_KEY=
# ---------------------------------------------------------------------------
# FIRST BOOT ONLY — then delete this line
# ---------------------------------------------------------------------------
# FabledCurator encrypts your stored platform credentials with a Fernet key it
# keeps at /images/secrets/credential_key.b64 — inside the ./images bind mount,
# so it outlives the container. On a brand-new install that file does not exist
# yet, and the app REFUSES TO START rather than quietly create one:
#
# MissingCredentialKey: Fernet key file not found at
# /images/secrets/credential_key.b64
#
# That refusal is deliberate. Auto-creating a key is indistinguishable from the
# disaster case — a restore that brought the database back but lost
# ./images/secrets — and there it would mint a key that cannot decrypt anything,
# leaving an instance that looks healthy while every paywalled download fails.
# So the choice is yours to make explicitly, once.
#
# Set this for your first `up`, watch the container come up, then DELETE THE
# LINE. Leaving it set disarms the protection permanently, on an instance that
# by then has credentials worth protecting.
#
# BACK UP ./images/secrets/ ALONGSIDE YOUR DATABASE. The key is the only thing
# that can read your stored credentials; a database restored without it needs
# every credential re-entered by hand.
CURATOR_BOOTSTRAP_NEW_KEY=1
# ---------------------------------------------------------------------------
# Optional — defaults are fine
# ---------------------------------------------------------------------------
+371 -234
View File
@@ -89,12 +89,30 @@ on:
# makes the milestone-362 gate verifiable at all: a gate has to be watched
# rejecting something before anyone can believe it is wired up.
#
# Note this is a STRING comparison, not a boolean. Forgejo delivers
# workflow_dispatch inputs as strings, so `inputs.refresh` is 'true'/'false'
# and `&&` on it would treat the string 'false' as truthy.
# The input is normalised through `format()` before it is compared, and that
# is not defensive styling — the direct comparison is WRONG and fails silently.
#
# `type: boolean` delivers a real boolean, and GitHub expression semantics cast
# operands to numbers when their types differ: `true == 'true'` compares 1
# against NaN and is FALSE. Measured on run 5270, whose own log says it —
#
# expression '(github.event_name == 'schedule'
# || github.event.inputs.refresh == 'true') && 'true' || 'false''
# evaluated to '%!t(string=false)'
# trigger: raw inputs refresh='true'
#
# — the input arrived as `true` and the expression still said false. The run
# then went green with every step skipped, because a refresh that evaluates
# false behaves exactly like an ordinary push. That is the whole hazard: the
# failure has no symptom.
#
# `force_build` never hit this because it never compares in an expression. It
# passes the raw value into an env var and tests it in the shell, where
# everything is already a string. `format('{0}', x)` buys the same thing here,
# where a step-level `if:` needs the answer before any shell runs.
env:
IS_REFRESH: ${{ (github.event_name == 'schedule' || github.event.inputs.refresh == 'true') && 'true' || 'false' }}
BUILD_REF: ${{ (github.event_name == 'schedule' || github.event.inputs.refresh == 'true') && 'main' || github.ref }}
IS_REFRESH: ${{ (github.event_name == 'schedule' || format('{0}', github.event.inputs.refresh) == 'true') && 'true' || 'false' }}
BUILD_REF: ${{ (github.event_name == 'schedule' || format('{0}', github.event.inputs.refresh) == 'true') && 'main' || github.ref }}
# Requires repo secret RELEASE_TOKEN — a Forgejo PAT with scopes:
# - write:package, read:package (for docker push to git.fabledsword.com)
@@ -434,6 +452,18 @@ jobs:
# to. Same source of truth; no double-store.
build-web:
# Consumed by smoke-web's job-level `if:`. It cannot read `env` — the env
# context is available to STEP `if:` and step bodies, never to a job's own
# condition, and an unresolvable context there is empty rather than an
# error. `smoke-web` skipped silently on run 5290 for exactly that reason.
#
# Keying off the reuse step's own output is better than re-deriving the
# trigger anyway: it is the same single decision the build, the XPI
# download and the promote all take (build.yml's "one decision drives
# everything downstream"), and it says the thing smoke-web actually needs
# to know — a candidate was published — rather than restating why.
outputs:
candidate: ${{ steps.reuse.outputs.promote }}
# A plain `needs` — no `always()`. That expression existed to let a
# SKIPPED sign-extension through on a tag push while still blocking a
# FAILED one. With no tag trigger, sign-extension always runs, so the
@@ -491,8 +521,18 @@ jobs:
# the PREVIOUS XPI while the freshly signed one is orphaned (#3156).
# * dev and main derive the same values for the same source.
- name: Report the derived artifact version
env:
# Diagnostic for the trigger normalisation. `refresh` is reported RAW
# as well as normalised, because the two disagreeing is the whole
# failure mode: a dispatch input whose type does not compare the way
# the expression assumes evaluates to false silently, and the only
# symptom is a refresh that quietly behaves like an ordinary push.
RAW_REFRESH: ${{ github.event.inputs.refresh }}
RAW_FORCE: ${{ github.event.inputs.force_build }}
run: |
set -u
echo "trigger: event=$GITHUB_EVENT_NAME IS_REFRESH='${IS_REFRESH:-<unset>}' BUILD_REF='${BUILD_REF:-<unset>}'"
echo "trigger: raw inputs refresh='${RAW_REFRESH:-<unset>}' force_build='${RAW_FORCE:-<unset>}'"
A=web
V=$(sh scripts/artifacts.sh version "$A" 2>&1 || echo UNAVAILABLE)
R=$(sh scripts/artifacts.sh revision "$A" 2>&1 || echo UNAVAILABLE)
@@ -694,10 +734,16 @@ jobs:
# already allows for :buildcache, not the per-build tag family that
# milestone 318 withdrew.
#
# Both values are decided HERE, beside `hit`, for the reason the
# force/schedule branch below gives: one step decides what this job
# does. A promote condition derived independently could disagree with
# the tag the build actually wrote.
# Decided HERE, beside `hit`, for the reason the force/schedule
# branch below gives: one step decides what this job does. A
# condition derived independently could disagree with the tag the
# build actually wrote.
#
# build-web additionally exposes this as `outputs.candidate`, which is
# what gates the `promote` job — a job's `if:` cannot read `env`, and
# one flag is enough because all three derive it from the same
# IS_REFRESH. ml and agent do not re-emit it; a second copy nothing
# reads is the kind of thing that later reads as load-bearing.
if [ "${IS_REFRESH:-}" = "true" ]; then
echo "build_ref=$IMAGE:refresh-candidate" >> "$GITHUB_OUTPUT"
echo "promote=true" >> "$GITHUB_OUTPUT"
@@ -919,77 +965,6 @@ jobs:
FC_CHANNEL=${{ steps.tag.outputs.channel }}
FC_VERSION=${{ steps.reuse.outputs.version }}
# Point the channel tag at the candidate the refresh just built.
#
# Unconditional TODAY, so this milestone never leaves the refresh in a
# state where it builds and publishes nothing. Step 4 wraps it in the
# smoke suite's verdict; until then the scheduled path behaves exactly
# as it did, just via two operations instead of one.
#
# NOT `imagetools create`. That wraps its source in an INDEX, and an
# indexed channel tag is the one thing this pipeline cannot survive:
# `.Image.Config.Labels` does not resolve through an index, so the
# fc.revision the reuse check reads off the channel tag would come back
# empty, every subsequent push would miss and rebuild, and nothing would
# go red. That is #3183, observed on run 4751 — reuse worked exactly once
# and the only symptom was the bill. The repoint step below excludes its
# own source tag for precisely this reason; a promote that re-introduced
# the wrap through a different door would undo that care.
#
# A manifest PUT is what "make this tag name that image" means at the
# registry level: the same bytes under the same media type, so the digest
# is identical, the media type is preserved, and no layer moves.
- name: Promote the refresh candidate to the channel
if: steps.reuse.outputs.promote == 'true'
env:
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator
CHANNEL_REF: ${{ steps.reuse.outputs.channel_ref }}
TOKEN: ${{ secrets.RELEASE_TOKEN }}
ACTOR: ${{ github.actor }}
run: |
set -eu
REPO=${IMAGE#git.fabledsword.com/}
TAG=${CHANNEL_REF##*:}
# Registry auth is its own token exchange — the `docker login` above
# authenticates the docker client, not curl. Deadline on every call
# (rule 156): a registry that stops answering must fail this step,
# not hang the weekly refresh until the job times out.
BEARER=$(curl -fsS --max-time 30 -u "$ACTOR:$TOKEN" \
"https://git.fabledsword.com/v2/token?scope=repository:$REPO:pull,push&service=git.fabledsword.com" \
| python3 -c 'import sys,json; print(json.load(sys.stdin)["token"])')
# Ask for the image manifest media types ONLY. Offering the index
# types too would let the registry hand back an index if one ever
# existed at this tag, and we would faithfully copy the thing we are
# trying not to create.
ACCEPT='application/vnd.oci.image.manifest.v1+json, application/vnd.docker.distribution.manifest.v2+json'
CT=$(curl -fsS --max-time 60 -o manifest.json -D headers.txt \
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
"https://git.fabledsword.com/v2/$REPO/manifests/refresh-candidate" \
&& tr -d '\r' < headers.txt | awk -F': ' '/^[Cc]ontent-[Tt]ype:/{print $2}')
test -n "$CT"
SRC_DIGEST=$(tr -d '\r' < headers.txt | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
echo "promote: candidate is $SRC_DIGEST ($CT)"
curl -fsS --max-time 120 -X PUT \
-H "Authorization: Bearer $BEARER" -H "Content-Type: $CT" \
--data-binary @manifest.json \
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG"
# Read it back. A PUT that returned 2xx but landed something else is
# exactly the silent-and-plausible failure this pipeline keeps
# producing, and the check costs one request.
NOW=$(curl -fsS --max-time 30 -o /dev/null -D - \
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG" \
| tr -d '\r' | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
if [ "$NOW" != "$SRC_DIGEST" ]; then
echo "promote: $IMAGE:$TAG is $NOW, expected $SRC_DIGEST" >&2
exit 1
fi
echo "promote: $IMAGE:$TAG now names $NOW"
# Every tag but the channel's own is written HERE, registry-side,
# whether or not a build ran. Each -t becomes another reference to the
# SAME manifest the channel tag holds, so :c-<sha> is byte-identical to
@@ -1074,6 +1049,282 @@ jobs:
docker buildx imagetools create $ARGS "$SOURCE"
echo "repointed from $SOURCE:$ARGS"
# Does the image a refresh just built still work?
#
# This is the gate the base refresh never had. `ci.yml` cannot be it: its
# lanes run on ci-python:3.14 and install requirements.txt, and a base bump
# changes neither — all five stay green through a refresh that breaks the
# product. What a refresh re-resolves is the Dockerfile's apt layer (ffmpeg,
# unar, libpq5, postgresql-client, zstd, megatools, libjpeg62-turbo,
# libwebp7, libpng16-16), unpinned, every build.
#
# So this runs the CANDIDATE IMAGE, against real Postgres and Redis. Not the
# source tree, and not a static inspection: `ffmpeg -version` exiting 0 would
# pass while a codec removal broke every thumbnail in the library.
#
# Refresh-only. On a push the bytes came from a commit, and a commit is what
# ci.yml already tests.
#
# Reports a verdict; it does not yet gate the promote (milestone 362 step 4).
# Landing the gate and the thing it gates in one change would mean the first
# time anyone saw this job run would also be the first time it could stop a
# publish.
smoke-web:
needs: [build-web]
if: needs.build-web.outputs.candidate == 'true'
runs-on: python-ci
container:
image: git.fabledsword.com/bvandeusen/ci-python:3.14
env:
DB_USER: fabledcurator
DB_PASSWORD: ci_smoke
DB_PORT: "5432"
DB_NAME: fabledcurator_smoke
SECRET_KEY: ci_smoke_placeholder
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator
services:
postgres:
image: pgvector/pgvector:pg16
env:
POSTGRES_USER: fabledcurator
POSTGRES_PASSWORD: ci_smoke
POSTGRES_DB: fabledcurator_smoke
options: >-
--health-cmd "pg_isready -U fabledcurator"
--health-interval 10s
--health-timeout 5s
--health-retries 10
redis:
image: redis:7-alpine
options: >-
--health-cmd "redis-cli ping"
--health-interval 10s
--health-timeout 5s
--health-retries 10
steps:
- uses: actions/checkout@v4
with:
# The same ref the image was built from, so the smoke script matches
# the code inside the candidate.
ref: ${{ env.BUILD_REF }}
- name: Smoke the candidate image
env:
TOKEN: ${{ secrets.RELEASE_TOKEN }}
ACTOR: ${{ github.actor }}
run: |
set -eux
# Service discovery mirrors ci.yml's integration lane: these jobs run
# in a container against a mounted docker socket, so the services are
# SIBLINGS reachable by IP, not by hostname.
PG=$(docker ps --filter "name=smoke" --filter "ancestor=pgvector/pgvector:pg16" -q | head -n1)
RD=$(docker ps --filter "name=smoke" --filter "ancestor=redis:7-alpine" -q | head -n1)
test -n "$PG" && test -n "$RD"
PG_IP=$(docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' "$PG")
RD_IP=$(docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' "$RD")
test -n "$PG_IP" && test -n "$RD_IP"
# Socket probe in python, not bash's /dev/tcp — these steps run under
# `sh -e`, where that path does not exist. Same fix and reasoning as
# ci.yml's integration job; see the comment there.
pg_ready=""
for i in $(seq 1 60); do
if python -c "import socket,sys; s=socket.socket(); s.settimeout(2); sys.exit(0 if s.connect_ex(('$PG_IP', 5432)) == 0 else 1)"; then
pg_ready=1
break
fi
sleep 2
done
if [ -z "$pg_ready" ]; then
echo "postgres at $PG_IP:5432 did not accept a connection within 120s"
exit 1
fi
echo "$TOKEN" | docker login git.fabledsword.com -u "$ACTOR" --password-stdin
CANDIDATE="$IMAGE:refresh-candidate"
docker pull "$CANDIDATE"
ENVOPTS="-e DB_USER=$DB_USER -e DB_PASSWORD=$DB_PASSWORD -e DB_HOST=$PG_IP"
ENVOPTS="$ENVOPTS -e DB_PORT=5432 -e DB_NAME=$DB_NAME -e SECRET_KEY=$SECRET_KEY"
ENVOPTS="$ENVOPTS -e CELERY_BROKER_URL=redis://$RD_IP:6379/0"
ENVOPTS="$ENVOPTS -e CELERY_RESULT_BACKEND=redis://$RD_IP:6379/0"
# A throwaway CI instance IS first-time setup, which is the one case
# credential_crypto allows a key to be minted in. Without it the web
# role refuses to boot — deliberately, since silently generating a
# key on a restored-DB-but-lost-secrets deployment would leave every
# Credential row undecryptable (the 2026-06-02 audit). Discovered by
# this job on its first real run; see #3422 for the fact that no
# user-facing file mentions this variable at all.
ENVOPTS="$ENVOPTS -e CURATOR_BOOTSTRAP_NEW_KEY=1"
# 1. The schema builds from empty, using the image's OWN libpq and
# psycopg. This is the same call entrypoint.sh makes before it
# serves anything, so a failure here is a failure to boot.
echo "smoke: alembic upgrade head"
docker run --rm $ENVOPTS "$CANDIDATE" alembic upgrade head
# 2. The apt layer's binaries and the app's own thumbnail path, run
# inside the image. Piped over stdin rather than bind-mounted: the
# workspace is a docker VOLUME belonging to this job's container,
# so a host bind of $PWD would not resolve for a sibling.
echo "smoke: image-internal checks"
docker run --rm -i $ENVOPTS "$CANDIDATE" shell -c 'python3 -' < scripts/smoke_image.py
# 3. It actually serves. `docker run -d` then poll the container's own
# IP — no port publishing, because the job container reaches
# siblings directly and a published port would collide with
# whatever else the runner is hosting.
echo "smoke: web boots and answers /api/health"
CID=$(docker run -d $ENVOPTS "$CANDIDATE" web)
# Clean up the container however this ends, and dump its log ONLY
# on failure — a boot that never answers must fail with the reason
# visible rather than as a bare timeout (rule 156), while a green run
# has nothing to say. `exit $rc` preserves the real status, which a
# trap that ends on a successful `docker rm` would otherwise mask.
trap 'rc=$?; [ $rc -eq 0 ] || docker logs "$CID" 2>&1 | tail -40; docker rm -f "$CID" >/dev/null 2>&1 || true; exit $rc' EXIT
WEB_IP=$(docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' "$CID")
test -n "$WEB_IP"
healthy=""
for i in $(seq 1 60); do
if curl -fsS --max-time 5 "http://$WEB_IP:8080/api/health" >/dev/null 2>&1; then
healthy=1
break
fi
# A container that has EXITED will never answer, so stop asking.
# Without this the loop spent 3m35s polling a dead container on
# this job's first run, and — because docker recycles the IP — got
# a confusing mix of connection-refused and 5s timeouts from
# whatever took the address next. The trap's log dump had the real
# answer the whole time; this just stops burying it.
if [ "$(docker inspect -f '{{.State.Running}}' "$CID" 2>/dev/null)" != "true" ]; then
echo "smoke: FAILED — the web container exited during boot." >&2
echo "smoke: its log follows; entrypoint runs alembic BEFORE" >&2
echo "smoke: serving, so a startup exception lands here." >&2
exit 1
fi
sleep 2
done
if [ -z "$healthy" ]; then
# 60 iterations of (up to 5s connect + 2s sleep) — up to ~7min, not
# the 120s an earlier version of this message claimed.
echo "smoke: FAILED — web is running but never answered" >&2
echo "smoke: /api/health. It is up, so look at hypercorn and the" >&2
echo "smoke: python base rather than at startup." >&2
exit 1
fi
curl -fsS --max-time 5 "http://$WEB_IP:8080/api/health"
echo
echo "smoke: all checks passed against $CANDIDATE"
# Move the channel tags — the whole point of the gate.
#
# Lives in its own job because the verdict it depends on cannot exist until
# after build-web has finished, and the promote used to run INSIDE build-web.
#
# `needs` on smoke-web is the gate. A failed smoke skips this job, so a
# refresh that broke something leaves :latest naming the build that works —
# "the refresh failed" and "production is broken" must not be the same event.
# A SKIPPED smoke also skips this job, which is the behaviour that matters
# most: on run 5290 the gate silently skipped itself, and a design where only
# a FAILED gate blocks would have published unverified images while reporting
# success. Not running is not the same as passing.
#
# All three images promote TOGETHER, or none do. They are one stack: build.yml
# already refuses to publish a :dev web image beside a stale :dev ml, because
# the mismatch only shows up as a runtime failure. A refresh that published ml
# and withheld web would be that same trap, arrived at through the gate.
#
# The gate covers the web image only (milestone 362 step 3 scoped it there),
# so ml and agent are being held to web's verdict rather than their own. That
# is deliberate and it is the conservative direction — they ship together, so
# the weakest evidence should govern all three — but it is not the same as
# having smoked them, and it should not be read as if it were.
promote:
needs: [build-web, build-ml, build-agent, smoke-web]
# Only a refresh publishes through a candidate; a push writes its channel
# tag directly from the build. Reads the same reuse-step decision the build
# took, via a job output — a job's `if:` cannot see the `env` context.
if: needs.build-web.outputs.candidate == 'true'
runs-on: python-ci
container:
image: git.fabledsword.com/bvandeusen/ci-python:3.14
steps:
- name: Point the channel tags at the smoked candidates
env:
TOKEN: ${{ secrets.RELEASE_TOKEN }}
ACTOR: ${{ github.actor }}
run: |
set -eu
# `latest` is not a guess: a refresh always builds `main` (BUILD_REF),
# and the "must have checked out main" guard in every build job fails
# the run if that did not hold. So the channel is main's.
TAG=latest
FAILED=""
for NAME in fabledcurator fabledcurator-ml fabledcurator-agent; do
REPO="bvandeusen/$NAME"
echo "promote: $REPO"
# Registry auth is its own token exchange — `docker login`
# authenticates the docker client, not curl. Deadline on every call
# (rule 156): a registry that stops answering must fail this step,
# not hang the weekly refresh until the job times out.
BEARER=$(curl -fsS --max-time 30 -u "$ACTOR:$TOKEN" \
"https://git.fabledsword.com/v2/token?scope=repository:$REPO:pull,push&service=git.fabledsword.com" \
| python3 -c 'import sys,json; print(json.load(sys.stdin)["token"])')
# Ask for the IMAGE manifest media types only. Offering the index
# types too would let the registry hand back an index if one ever
# existed at this tag, and we would faithfully copy the thing this
# whole approach exists to avoid creating.
ACCEPT='application/vnd.oci.image.manifest.v1+json, application/vnd.docker.distribution.manifest.v2+json'
CT=$(curl -fsS --max-time 60 -o manifest.json -D headers.txt \
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
"https://git.fabledsword.com/v2/$REPO/manifests/refresh-candidate" \
&& tr -d '\r' < headers.txt | awk -F': ' '/^[Cc]ontent-[Tt]ype:/{print $2}')
test -n "$CT"
SRC=$(tr -d '\r' < headers.txt | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
echo "promote: candidate $SRC ($CT)"
# NOT `imagetools create`. That wraps its source in an INDEX, and
# `.Image.Config.Labels` does not resolve through one — the
# fc.revision the reuse check reads off the channel tag would come
# back empty, every later push would miss and rebuild, and nothing
# would go red (#3183, run 4751). A manifest PUT is what "make this
# tag name that image" means at the registry: same bytes, same media
# type, same digest, no layer transfer.
curl -fsS --max-time 120 -X PUT \
-H "Authorization: Bearer $BEARER" -H "Content-Type: $CT" \
--data-binary @manifest.json \
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG"
# Read it back. A PUT that returned 2xx but landed something else is
# exactly the silent-and-plausible failure this pipeline keeps
# producing, and the check costs one request.
NOW=$(curl -fsS --max-time 30 -o /dev/null -D - \
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG" \
| tr -d '\r' | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
if [ "$NOW" != "$SRC" ]; then
echo "promote: FAILED — $NAME:$TAG is $NOW, expected $SRC" >&2
FAILED="$FAILED $NAME"
continue
fi
echo "promote: $NAME:$TAG now names $NOW"
done
if [ -n "$FAILED" ]; then
echo "" >&2
echo "promote: FAILED for:$FAILED" >&2
echo "promote: the channel tags are now INCONSISTENT — some images" >&2
echo "promote: moved and some did not. Re-run this refresh; the" >&2
echo "promote: candidates are still published and the promote is" >&2
echo "promote: idempotent." >&2
exit 1
fi
echo "promote: all three channel tags moved"
build-ml:
runs-on: python-ci
container:
@@ -1127,8 +1378,18 @@ jobs:
# the PREVIOUS XPI while the freshly signed one is orphaned (#3156).
# * dev and main derive the same values for the same source.
- name: Report the derived artifact version
env:
# Diagnostic for the trigger normalisation. `refresh` is reported RAW
# as well as normalised, because the two disagreeing is the whole
# failure mode: a dispatch input whose type does not compare the way
# the expression assumes evaluates to false silently, and the only
# symptom is a refresh that quietly behaves like an ordinary push.
RAW_REFRESH: ${{ github.event.inputs.refresh }}
RAW_FORCE: ${{ github.event.inputs.force_build }}
run: |
set -u
echo "trigger: event=$GITHUB_EVENT_NAME IS_REFRESH='${IS_REFRESH:-<unset>}' BUILD_REF='${BUILD_REF:-<unset>}'"
echo "trigger: raw inputs refresh='${RAW_REFRESH:-<unset>}' force_build='${RAW_FORCE:-<unset>}'"
A=ml
V=$(sh scripts/artifacts.sh version "$A" 2>&1 || echo UNAVAILABLE)
R=$(sh scripts/artifacts.sh revision "$A" 2>&1 || echo UNAVAILABLE)
@@ -1269,16 +1530,20 @@ jobs:
# already allows for :buildcache, not the per-build tag family that
# milestone 318 withdrew.
#
# Both values are decided HERE, beside `hit`, for the reason the
# force/schedule branch below gives: one step decides what this job
# does. A promote condition derived independently could disagree with
# the tag the build actually wrote.
# Decided HERE, beside `hit`, for the reason the force/schedule
# branch below gives: one step decides what this job does. A
# condition derived independently could disagree with the tag the
# build actually wrote.
#
# build-web additionally exposes this as `outputs.candidate`, which is
# what gates the `promote` job — a job's `if:` cannot read `env`, and
# one flag is enough because all three derive it from the same
# IS_REFRESH. ml and agent do not re-emit it; a second copy nothing
# reads is the kind of thing that later reads as load-bearing.
if [ "${IS_REFRESH:-}" = "true" ]; then
echo "build_ref=$IMAGE:refresh-candidate" >> "$GITHUB_OUTPUT"
echo "promote=true" >> "$GITHUB_OUTPUT"
else
echo "build_ref=$IMAGE:$T" >> "$GITHUB_OUTPUT"
echo "promote=false" >> "$GITHUB_OUTPUT"
fi
# Compare VALUES, never exit codes. Measured on buildx v0.36.1
@@ -1412,77 +1677,6 @@ jobs:
cache-from: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator-ml:buildcache
cache-to: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator-ml:buildcache,mode=max
# Point the channel tag at the candidate the refresh just built.
#
# Unconditional TODAY, so this milestone never leaves the refresh in a
# state where it builds and publishes nothing. Step 4 wraps it in the
# smoke suite's verdict; until then the scheduled path behaves exactly
# as it did, just via two operations instead of one.
#
# NOT `imagetools create`. That wraps its source in an INDEX, and an
# indexed channel tag is the one thing this pipeline cannot survive:
# `.Image.Config.Labels` does not resolve through an index, so the
# fc.revision the reuse check reads off the channel tag would come back
# empty, every subsequent push would miss and rebuild, and nothing would
# go red. That is #3183, observed on run 4751 — reuse worked exactly once
# and the only symptom was the bill. The repoint step below excludes its
# own source tag for precisely this reason; a promote that re-introduced
# the wrap through a different door would undo that care.
#
# A manifest PUT is what "make this tag name that image" means at the
# registry level: the same bytes under the same media type, so the digest
# is identical, the media type is preserved, and no layer moves.
- name: Promote the refresh candidate to the channel
if: steps.reuse.outputs.promote == 'true'
env:
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator-ml
CHANNEL_REF: ${{ steps.reuse.outputs.channel_ref }}
TOKEN: ${{ secrets.RELEASE_TOKEN }}
ACTOR: ${{ github.actor }}
run: |
set -eu
REPO=${IMAGE#git.fabledsword.com/}
TAG=${CHANNEL_REF##*:}
# Registry auth is its own token exchange — the `docker login` above
# authenticates the docker client, not curl. Deadline on every call
# (rule 156): a registry that stops answering must fail this step,
# not hang the weekly refresh until the job times out.
BEARER=$(curl -fsS --max-time 30 -u "$ACTOR:$TOKEN" \
"https://git.fabledsword.com/v2/token?scope=repository:$REPO:pull,push&service=git.fabledsword.com" \
| python3 -c 'import sys,json; print(json.load(sys.stdin)["token"])')
# Ask for the image manifest media types ONLY. Offering the index
# types too would let the registry hand back an index if one ever
# existed at this tag, and we would faithfully copy the thing we are
# trying not to create.
ACCEPT='application/vnd.oci.image.manifest.v1+json, application/vnd.docker.distribution.manifest.v2+json'
CT=$(curl -fsS --max-time 60 -o manifest.json -D headers.txt \
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
"https://git.fabledsword.com/v2/$REPO/manifests/refresh-candidate" \
&& tr -d '\r' < headers.txt | awk -F': ' '/^[Cc]ontent-[Tt]ype:/{print $2}')
test -n "$CT"
SRC_DIGEST=$(tr -d '\r' < headers.txt | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
echo "promote: candidate is $SRC_DIGEST ($CT)"
curl -fsS --max-time 120 -X PUT \
-H "Authorization: Bearer $BEARER" -H "Content-Type: $CT" \
--data-binary @manifest.json \
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG"
# Read it back. A PUT that returned 2xx but landed something else is
# exactly the silent-and-plausible failure this pipeline keeps
# producing, and the check costs one request.
NOW=$(curl -fsS --max-time 30 -o /dev/null -D - \
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG" \
| tr -d '\r' | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
if [ "$NOW" != "$SRC_DIGEST" ]; then
echo "promote: $IMAGE:$TAG is $NOW, expected $SRC_DIGEST" >&2
exit 1
fi
echo "promote: $IMAGE:$TAG now names $NOW"
# Every tag but the channel's own is written HERE, registry-side,
# whether or not a build ran. Each -t becomes another reference to the
# SAME manifest the channel tag holds, so :c-<sha> is byte-identical to
@@ -1617,8 +1811,18 @@ jobs:
# the PREVIOUS XPI while the freshly signed one is orphaned (#3156).
# * dev and main derive the same values for the same source.
- name: Report the derived artifact version
env:
# Diagnostic for the trigger normalisation. `refresh` is reported RAW
# as well as normalised, because the two disagreeing is the whole
# failure mode: a dispatch input whose type does not compare the way
# the expression assumes evaluates to false silently, and the only
# symptom is a refresh that quietly behaves like an ordinary push.
RAW_REFRESH: ${{ github.event.inputs.refresh }}
RAW_FORCE: ${{ github.event.inputs.force_build }}
run: |
set -u
echo "trigger: event=$GITHUB_EVENT_NAME IS_REFRESH='${IS_REFRESH:-<unset>}' BUILD_REF='${BUILD_REF:-<unset>}'"
echo "trigger: raw inputs refresh='${RAW_REFRESH:-<unset>}' force_build='${RAW_FORCE:-<unset>}'"
A=agent
V=$(sh scripts/artifacts.sh version "$A" 2>&1 || echo UNAVAILABLE)
R=$(sh scripts/artifacts.sh revision "$A" 2>&1 || echo UNAVAILABLE)
@@ -1754,16 +1958,20 @@ jobs:
# already allows for :buildcache, not the per-build tag family that
# milestone 318 withdrew.
#
# Both values are decided HERE, beside `hit`, for the reason the
# force/schedule branch below gives: one step decides what this job
# does. A promote condition derived independently could disagree with
# the tag the build actually wrote.
# Decided HERE, beside `hit`, for the reason the force/schedule
# branch below gives: one step decides what this job does. A
# condition derived independently could disagree with the tag the
# build actually wrote.
#
# build-web additionally exposes this as `outputs.candidate`, which is
# what gates the `promote` job — a job's `if:` cannot read `env`, and
# one flag is enough because all three derive it from the same
# IS_REFRESH. ml and agent do not re-emit it; a second copy nothing
# reads is the kind of thing that later reads as load-bearing.
if [ "${IS_REFRESH:-}" = "true" ]; then
echo "build_ref=$IMAGE:refresh-candidate" >> "$GITHUB_OUTPUT"
echo "promote=true" >> "$GITHUB_OUTPUT"
else
echo "build_ref=$IMAGE:$T" >> "$GITHUB_OUTPUT"
echo "promote=false" >> "$GITHUB_OUTPUT"
fi
# Compare VALUES, never exit codes. Measured on buildx v0.36.1
@@ -1897,77 +2105,6 @@ jobs:
cache-from: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator-agent:buildcache
cache-to: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator-agent:buildcache,mode=max
# Point the channel tag at the candidate the refresh just built.
#
# Unconditional TODAY, so this milestone never leaves the refresh in a
# state where it builds and publishes nothing. Step 4 wraps it in the
# smoke suite's verdict; until then the scheduled path behaves exactly
# as it did, just via two operations instead of one.
#
# NOT `imagetools create`. That wraps its source in an INDEX, and an
# indexed channel tag is the one thing this pipeline cannot survive:
# `.Image.Config.Labels` does not resolve through an index, so the
# fc.revision the reuse check reads off the channel tag would come back
# empty, every subsequent push would miss and rebuild, and nothing would
# go red. That is #3183, observed on run 4751 — reuse worked exactly once
# and the only symptom was the bill. The repoint step below excludes its
# own source tag for precisely this reason; a promote that re-introduced
# the wrap through a different door would undo that care.
#
# A manifest PUT is what "make this tag name that image" means at the
# registry level: the same bytes under the same media type, so the digest
# is identical, the media type is preserved, and no layer moves.
- name: Promote the refresh candidate to the channel
if: steps.reuse.outputs.promote == 'true'
env:
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator-agent
CHANNEL_REF: ${{ steps.reuse.outputs.channel_ref }}
TOKEN: ${{ secrets.RELEASE_TOKEN }}
ACTOR: ${{ github.actor }}
run: |
set -eu
REPO=${IMAGE#git.fabledsword.com/}
TAG=${CHANNEL_REF##*:}
# Registry auth is its own token exchange — the `docker login` above
# authenticates the docker client, not curl. Deadline on every call
# (rule 156): a registry that stops answering must fail this step,
# not hang the weekly refresh until the job times out.
BEARER=$(curl -fsS --max-time 30 -u "$ACTOR:$TOKEN" \
"https://git.fabledsword.com/v2/token?scope=repository:$REPO:pull,push&service=git.fabledsword.com" \
| python3 -c 'import sys,json; print(json.load(sys.stdin)["token"])')
# Ask for the image manifest media types ONLY. Offering the index
# types too would let the registry hand back an index if one ever
# existed at this tag, and we would faithfully copy the thing we are
# trying not to create.
ACCEPT='application/vnd.oci.image.manifest.v1+json, application/vnd.docker.distribution.manifest.v2+json'
CT=$(curl -fsS --max-time 60 -o manifest.json -D headers.txt \
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
"https://git.fabledsword.com/v2/$REPO/manifests/refresh-candidate" \
&& tr -d '\r' < headers.txt | awk -F': ' '/^[Cc]ontent-[Tt]ype:/{print $2}')
test -n "$CT"
SRC_DIGEST=$(tr -d '\r' < headers.txt | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
echo "promote: candidate is $SRC_DIGEST ($CT)"
curl -fsS --max-time 120 -X PUT \
-H "Authorization: Bearer $BEARER" -H "Content-Type: $CT" \
--data-binary @manifest.json \
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG"
# Read it back. A PUT that returned 2xx but landed something else is
# exactly the silent-and-plausible failure this pipeline keeps
# producing, and the check costs one request.
NOW=$(curl -fsS --max-time 30 -o /dev/null -D - \
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG" \
| tr -d '\r' | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
if [ "$NOW" != "$SRC_DIGEST" ]; then
echo "promote: $IMAGE:$TAG is $NOW, expected $SRC_DIGEST" >&2
exit 1
fi
echo "promote: $IMAGE:$TAG now names $NOW"
# Every tag but the channel's own is written HERE, registry-side,
# whether or not a build ran. Each -t becomes another reference to the
# SAME manifest the channel tag holds, so :c-<sha> is byte-identical to
+25 -1
View File
@@ -90,7 +90,31 @@ If you forget it, the symptom is a long build instead of a quick pull.
The database schema is created automatically on first start — the web container
runs its migrations before serving. Nothing to initialise by hand.
A few things are worth knowing about the first few minutes:
**One thing does need a deliberate act, and the app will not start without it.**
FabledCurator encrypts your stored platform credentials with a key it keeps at
`./images/secrets/credential_key.b64`. On a brand-new install that file does not
exist, and rather than quietly creating one the app stops:
```
MissingCredentialKey: Fernet key file not found at /images/secrets/credential_key.b64
```
Set `CURATOR_BOOTSTRAP_NEW_KEY=1` in your `.env` for the first `up`, then delete
the line once the container is running. `.env.example` ships it with that
instruction attached.
The refusal is deliberate, and worth understanding rather than working around:
auto-creating a key is indistinguishable from the disaster case — a restore that
brought the database back but lost `./images/secrets` — where it would mint a key
that cannot decrypt anything, leaving an instance that looks healthy while every
paywalled download fails. Making you say so once, on an empty install, is the
price of that not happening silently later.
**Which means: back up `./images/secrets/` alongside your database.** It is the
only thing that can read your stored credentials. A database restored without it
needs every credential entered again by hand.
A few other things are worth knowing about the first few minutes:
- **The ML worker downloads its model weights on first boot**, several GB from
HuggingFace into `./models`. Until that finishes, tagging is queued rather
+64
View File
@@ -0,0 +1,64 @@
"""service_seen — the learned roster that makes a stopped part observable.
Milestone 365. Nothing in FabledCurator knew what was SUPPOSED to be running:
`celery inspect` reports the workers that answer, so a dead worker was a
shorter list rather than a red light, and the only surface that could tell an
operator otherwise was Portainer. This table is the memory that turns an
absence into something the app can see.
Keyed on the queue set for a celery role and on agent_id for the GPU agent —
NOT on the celery worker name, which here is `celery@<container id>` and is
minted fresh on every deploy. See the model docstring for why that choice is
the whole design.
## First migration on the collapsed baseline
0089 is the single generated baseline that replaced revisions 0001..0089
(milestone 328). This is the first revision written on top of it, so it is
also the first evidence that the chain steps forward from the collapse rather
than merely reproducing the schema — which nothing had demonstrated yet.
An existing install is at 0089 because it ran the real 0089; a fresh one is at
0089 because it ran the baseline. Both arrive here identically, which was the
property the collapse was designed around.
Revision ID: 0090
Revises: 0089
Create Date: 2026-09-02
"""
from typing import Sequence, Union
import sqlalchemy as sa
from alembic import op
revision: str = "0090"
down_revision: Union[str, None] = "0089"
branch_labels: Union[str, Sequence[str], None] = None
depends_on: Union[str, Sequence[str], None] = None
def upgrade() -> None:
op.create_table(
"service_seen",
sa.Column("key", sa.String(length=128), nullable=False),
sa.Column("kind", sa.String(length=16), nullable=False),
sa.Column("display_name", sa.String(length=64), nullable=False),
sa.Column(
"first_seen_at", sa.DateTime(timezone=True),
server_default=sa.text("now()"), nullable=False,
),
sa.Column(
"last_seen_at", sa.DateTime(timezone=True),
server_default=sa.text("now()"), nullable=False,
),
sa.Column("details", sa.JSON(), nullable=False),
sa.PrimaryKeyConstraint("key", name=op.f("pk_service_seen")),
)
# No secondary indexes, deliberately: one row per moving part means every
# read is a handful of rows and an index would be write cost buying
# nothing (#3301 removed seven of exactly that shape).
def downgrade() -> None:
op.drop_table("service_seen")
+2
View File
@@ -38,6 +38,7 @@ def all_blueprints() -> list[Blueprint]:
from .suggestions import suggestions_bp
from .system_activity import system_activity_bp
from .system_backup import system_backup_bp
from .system_health import system_health_bp
from .tags import tags_bp
from .thumbnails import thumbnails_bp
return [
@@ -51,6 +52,7 @@ def all_blueprints() -> list[Blueprint]:
showcase_bp,
settings_bp,
system_activity_bp,
system_health_bp,
system_backup_bp,
admin_bp,
cleanup_bp,
+20
View File
@@ -21,6 +21,7 @@ from ..services.gallery_service import image_url
from ..services.ml.gpu_jobs import GpuJobService, error_dedupe_statements
from ..services.ml.gpu_triage import classify_reason, recover_defective_image
from ..services.ml.regions import RegionService
from ..services.service_roster import touch_service
gpu_bp = Blueprint("gpu", __name__, url_prefix="/api/gpu")
@@ -256,6 +257,18 @@ async def lease():
if not await _agent_authed(session):
return jsonify({"error": "unauthorized"}), 401
jobs = await GpuJobService(session).lease(agent_id, batch_size=batch)
# The agent cannot be polled — it is HTTP-only and pulls from here, so
# web never dials it. A lease IS the check-in, and until milestone 365
# it was thrown away: an agent sitting idle with nothing to lease left
# no trace at all and was indistinguishable from one switched off a
# week ago. Recorded on the call that was already happening.
await touch_service(
session,
key=f"agent:{agent_id}",
kind="agent",
display_name="GPU agent" if agent_id == "agent" else f"GPU agent ({agent_id})",
details={"agent_id": agent_id, "last_call": "lease", "leased": len(jobs)},
)
ml = await MLSettings.load(session)
# image rows for url/mime in one shot
ids = [j.image_record_id for j in jobs]
@@ -329,6 +342,13 @@ async def heartbeat():
if not await _agent_authed(session):
return jsonify({"error": "unauthorized"}), 401
n = await GpuJobService(session).heartbeat(agent_id, job_ids)
await touch_service(
session,
key=f"agent:{agent_id}",
kind="agent",
display_name="GPU agent" if agent_id == "agent" else f"GPU agent ({agent_id})",
details={"agent_id": agent_id, "last_call": "heartbeat", "extended": n},
)
await session.commit()
return jsonify({"extended": n})
+192
View File
@@ -0,0 +1,192 @@
"""Is every part of FabledCurator running? One verdict, one endpoint.
Milestone 365. The nav indicator and the System page both read this and
nothing else — composing a verdict is this module's job, not the UI's.
## Two kinds of part, answered two different ways
**Learned** — celery roles and the GPU agent, from `service_seen`. The
question is "how long since it checked in", and these are the parts that can
be ABSENT, which is the whole point: `celery inspect` alone reports presence,
so a dead worker is a shorter list rather than a red light.
**Probed live** — Postgres and Redis. Always expected, never learned, and a
last-seen for them would be actively misleading: that Redis answered thirty
seconds ago says nothing about now.
## This endpoint must never fail because something it checks has failed
The inversion is easy to write by accident and it destroys the feature exactly
when it is needed — a 500 when Redis is down, instead of `redis: down`. Every
probe is wrapped, every wait has a deadline (rule 156), and the roster refresh
swallows its own errors. The worst case is a part reported `unknown`, which is
a true statement.
"""
from __future__ import annotations
import asyncio
import logging
import time
from datetime import UTC, datetime
from quart import Blueprint, jsonify
from sqlalchemy import select, text
from ..config import get_config
from ..extensions import get_session
from ..models import ServiceSeen
from ..services.service_roster import refresh_if_stale
log = logging.getLogger(__name__)
system_health_bp = Blueprint("system_health", __name__, url_prefix="/api/system")
# How long a learned part may go quiet before it is doubted, then disbelieved.
#
# These are deliberately generous, and the reason is a deploy rather than a
# worker: `docker compose up -d` rolls start-first, so a role is briefly served
# by two containers and then by neither while the old one drains. Thresholds
# tight enough to catch a crash in seconds would paint the page red every time
# the stack is updated, and an alarm that cries wolf on every deploy is one
# nobody reads. Tune down only after watching a real deploy pass through.
STALE_AFTER_SECONDS = 90
DOWN_AFTER_SECONDS = 300
# Probes cross a process boundary, so they carry deadlines. A hung Postgres
# must make this endpoint say "postgres: down", not hang alongside it.
PROBE_TIMEOUT_SECONDS = 2.0
_OK, _STALE, _DOWN, _UNKNOWN = "ok", "stale", "down", "unknown"
# Worst-first, so an overall verdict is just the max.
_SEVERITY = {_OK: 0, _UNKNOWN: 1, _STALE: 2, _DOWN: 3}
def _age_state(age_seconds: float) -> str:
if age_seconds >= DOWN_AFTER_SECONDS:
return _DOWN
if age_seconds >= STALE_AFTER_SECONDS:
return _STALE
return _OK
def _describe_learned(name: str, state: str, age: float, details: dict) -> str:
"""Say what the state MEANS. A red chip tells an operator less than a
sentence does at the moment they are deciding whether to go and look."""
if state == _OK:
replicas = details.get("replicas")
if replicas and replicas > 1:
return f"{name} is running ({replicas} replicas)"
return f"{name} is running"
mins = int(age // 60)
ago = f"{mins} min" if mins else f"{int(age)}s"
if state == _STALE:
return f"{name} has not checked in for {ago}"
return f"{name} has not checked in for {ago} — treat it as stopped"
async def _probe_postgres(session) -> dict:
started = time.monotonic()
try:
await asyncio.wait_for(
session.execute(text("SELECT 1")), timeout=PROBE_TIMEOUT_SECONDS
)
except Exception as exc: # noqa: BLE001 — a probe reports, it never raises
return {
"key": "postgres", "kind": "datastore", "name": "PostgreSQL",
"state": _DOWN, "detail": f"not answering: {type(exc).__name__}",
}
return {
"key": "postgres", "kind": "datastore", "name": "PostgreSQL", "state": _OK,
"detail": "answering", "latency_ms": round((time.monotonic() - started) * 1000, 1),
}
def _ping_redis_sync() -> None:
import redis # local import; mirrors system_activity's pattern
client = redis.Redis.from_url(
get_config().celery_broker_url,
socket_connect_timeout=PROBE_TIMEOUT_SECONDS,
socket_timeout=PROBE_TIMEOUT_SECONDS,
)
client.ping()
async def _probe_redis() -> dict:
started = time.monotonic()
try:
await asyncio.wait_for(
asyncio.to_thread(_ping_redis_sync), timeout=PROBE_TIMEOUT_SECONDS * 2
)
except Exception as exc: # noqa: BLE001
return {
"key": "redis", "kind": "datastore", "name": "Redis",
"state": _DOWN,
"detail": f"not answering: {type(exc).__name__} — queues and workers "
f"cannot be reached either",
}
return {
"key": "redis", "kind": "datastore", "name": "Redis", "state": _OK,
"detail": "answering", "latency_ms": round((time.monotonic() - started) * 1000, 1),
}
@system_health_bp.route("/health", methods=["GET"])
async def system_health():
"""Every part, its state, and one overall verdict.
Response: {overall, parts: [{key, kind, name, state, detail, last_seen_at,
…}], checked_at}
"""
parts: list[dict] = []
now = datetime.now(UTC)
async with get_session() as session:
# Postgres first, and if it is unreachable nothing else can be read —
# say so rather than failing, because "the database is down" is the
# single most useful thing this endpoint can ever report.
pg = await _probe_postgres(session)
parts.append(pg)
if pg["state"] == _OK:
# Rate-limited inside; see service_roster on why the web process
# is the right observer.
try:
await refresh_if_stale(session)
await session.commit()
except Exception: # noqa: BLE001
log.warning("system health: roster refresh failed", exc_info=True)
rows = (
await session.execute(select(ServiceSeen).order_by(ServiceSeen.display_name))
).scalars().all()
for row in rows:
age = (now - row.last_seen_at).total_seconds()
state = _age_state(age)
parts.append({
"key": row.key,
"kind": row.kind,
"name": row.display_name,
"state": state,
"detail": _describe_learned(row.display_name, state, age, row.details or {}),
"last_seen_at": row.last_seen_at.isoformat(),
"first_seen_at": row.first_seen_at.isoformat(),
**{k: v for k, v in (row.details or {}).items() if k != "agent_id"},
})
parts.append(await _probe_redis())
overall = max((p["state"] for p in parts), key=lambda s: _SEVERITY[s], default=_UNKNOWN)
return jsonify({
"overall": overall,
"parts": sorted(parts, key=lambda p: (-_SEVERITY[p["state"]], p["name"])),
"checked_at": now.isoformat(),
# So the UI can explain a `stale` without hard-coding the same numbers
# in a second place.
"thresholds": {
"stale_after_seconds": STALE_AFTER_SECONDS,
"down_after_seconds": DOWN_AFTER_SECONDS,
},
})
+2
View File
@@ -32,6 +32,7 @@ from .presentation_review import PresentationReview
from .series_chapter import SeriesChapter
from .series_page import SeriesPage
from .series_suggestion import SeriesSuggestion
from .service_seen import ServiceSeen
from .source import Source
from .subscribestar_failed_media import SubscribeStarFailedMedia
from .subscribestar_seen_media import SubscribeStarSeenMedia
@@ -63,6 +64,7 @@ __all__ = [
"SeriesChapter",
"SeriesPage",
"SeriesSuggestion",
"ServiceSeen",
"ImageRecord",
"ImageProvenance",
"ImageRegion",
+88
View File
@@ -0,0 +1,88 @@
"""service_seen — the learned roster of FabledCurator's own moving parts.
Nothing else in this application knows what is SUPPOSED to be running.
`celery inspect` reports the workers that answer, so a stopped worker is a
shorter list rather than a red light, and Postgres and Redis have no
representation at all. That is why the only place an operator could see a
dead service was Portainer, which knows the intended set (milestone 365).
This table is the memory that makes an absence observable: every part that
has ever checked in, and when it last did. A row that stops advancing is a
part that stopped.
## Why the key is not the hostname
`_read_workers_sync()` returns celery's worker names, which here are
`celery@<container id>`. Those are minted fresh on every deploy. Keyed on
them, this table would record a death and a birth every time the stack is
updated — and a status page that goes red on every deploy is a status page
nobody reads, which is worse than not having one.
So a celery role is keyed on its **queue set**, which is assigned per role in
docker-compose.yml (`CELERY_QUEUES`) and survives container replacement:
default,import,thumbnail,download -> worker
maintenance,scan -> scheduler (celery worker --beat)
ml -> ml-worker
Two replicas of one role share a queue set and are therefore ONE row — which
is right, because the question being answered is "is that role being served",
not "how many containers exist". The replica count and their hostnames go in
`details`, where they can change without the identity changing.
The GPU agent is keyed on its `agent_id`, the identity its lease protocol
already uses (`api/gpu.py`).
## What is NOT in here
Postgres and Redis. They are always expected and never learned, and a
last-seen for them would be actively misleading — that one answered thirty
seconds ago says nothing about now. They are probed live at request time.
## kind
Plain `String`, not a Postgres ENUM and not CHECK-gated, matching
`gpu_job.status` and `backup_run.status`. The value set here is expected to
grow as parts are added, and a constraint swap per new kind (rule 36) would
be cost with no invariant behind it.
celery — a worker role, keyed on its queue set
agent — a GPU agent, keyed on its agent_id
"""
from datetime import datetime
from sqlalchemy import JSON, DateTime, String, func
from sqlalchemy.orm import Mapped, mapped_column
from .base import Base
class ServiceSeen(Base):
__tablename__ = "service_seen"
# No indexes beyond the primary key, deliberately. This table holds one row
# per moving part — a handful, forever — so every query against it is a
# full read of a few rows and an index would be write cost buying nothing
# (the lesson of #3301, which removed seven redundant ones).
key: Mapped[str] = mapped_column(String(128), primary_key=True)
kind: Mapped[str] = mapped_column(String(16), nullable=False)
# What to call it in the UI. Derived from the queue set where it is
# recognised, and falling back to the raw queue list where it is not — a
# deployment that slices its queues differently should still show something
# true rather than a name this code invented for it.
display_name: Mapped[str] = mapped_column(String(64), nullable=False)
first_seen_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now(),
)
last_seen_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), nullable=False, server_default=func.now(),
)
# The parts that change without changing identity: replica hostnames,
# active task counts, the queues actually being served. Kept as a blob
# because it is displayed and never queried — giving it columns would
# invite filtering on it, which is what the activity endpoints are for.
details: Mapped[dict] = mapped_column(JSON, nullable=False, default=dict)
+175
View File
@@ -0,0 +1,175 @@
"""The learned roster: which of FabledCurator's parts have checked in, and when.
Milestone 365. `celery inspect` answers "who is here"; this answers "who is
missing", which nothing in the application could do before — see
`models/service_seen.py` for why the identity is a queue set and not a
worker hostname.
## Who does the observing, and why it is the web process
Three candidates, and the choice matters more than the code:
* **A celery beat sweep.** Rejected. If the scheduler dies, the sweep stops,
every row goes stale, and the page reports that everything is down when one
thing is. An alarm that cannot distinguish "one part died" from "the
observer died" is worse than no alarm.
* **A background task in web.** Rejected on a detail of how this deploys:
hypercorn runs `--workers 4`, so a `before_serving` loop would be FOUR
concurrent inspect loops hammering the broker, forever, per container.
* **Refresh on demand, rate-limited by the data itself.** Taken. Whichever web
process happens to serve a health request refreshes the roster if it is
older than REFRESH_TTL, and otherwise reads what is already there.
The third has the property the other two lack: **the observer is the thing
serving the page.** If web is down you get a browser error rather than a
confidently green page, which is the honest failure. It also self-limits
without coordination — the TTL lives in the row everybody can see.
"""
from __future__ import annotations
import asyncio
import logging
from sqlalchemy import func, select
from sqlalchemy.dialects.postgresql import insert as pg_insert
from sqlalchemy.ext.asyncio import AsyncSession
from ..models import ServiceSeen
log = logging.getLogger(__name__)
# How stale the roster may be before a health request refreshes it. Comfortably
# under the staleness thresholds that decide a service is missing, so the
# verdict is never limited by how often anyone looked.
REFRESH_TTL_SECONDS = 20.0
# celery inspect is a broker round trip and this sits on a request path, so it
# gets a deadline (rule 156). A broker that has stopped answering must make the
# roster stale — which is a true statement about the system — not hang the one
# page that exists to explain it.
INSPECT_TIMEOUT_SECONDS = 2.0
# Queue set -> the name an operator recognises. Sorted-tuple keys, because the
# order celery reports them in is not guaranteed.
#
# A deployment that slices CELERY_QUEUES differently falls through to the raw
# queue list rather than being given a name this table invented for it: a
# wrong-but-confident label on a status page is worse than an ugly true one.
ROLE_NAMES: dict[tuple[str, ...], str] = {
("default", "download", "import", "thumbnail"): "Worker",
("maintenance", "scan"): "Scheduler",
("ml",): "ML worker",
}
def role_display_name(queues: tuple[str, ...]) -> str:
known = ROLE_NAMES.get(queues)
if known:
return known
return "Worker (" + ", ".join(queues) + ")"
def _inspect_celery_sync() -> dict[tuple[str, ...], dict]:
"""celery inspect, grouped by queue set rather than by worker.
Returns {queue_set: {"hostnames": [...], "active": int}}. Two replicas of
one role collapse into one entry on purpose — the question is whether the
role is being served, not how many containers exist.
"""
from ..celery_app import celery as celery_app
insp = celery_app.control.inspect(timeout=INSPECT_TIMEOUT_SECONDS)
active_queues = insp.active_queues() or {}
active_tasks = insp.active() or {}
grouped: dict[tuple[str, ...], dict] = {}
for hostname, queues in active_queues.items():
key = tuple(sorted({q["name"] for q in queues}))
entry = grouped.setdefault(key, {"hostnames": [], "active": 0})
entry["hostnames"].append(hostname)
entry["active"] += len(active_tasks.get(hostname, []))
for entry in grouped.values():
entry["hostnames"].sort()
return grouped
async def touch_service(
session: AsyncSession, *, key: str, kind: str, display_name: str, details: dict
) -> None:
"""Record that a part checked in just now.
Upsert rather than read-modify-write: several web processes and several
agents can be doing this at once, and the last writer is simply the most
recent sighting. `first_seen_at` is deliberately NOT updated — it is the
one field that answers "has this ever run", which the learned-roster design
depends on.
"""
stmt = pg_insert(ServiceSeen).values(
key=key, kind=kind, display_name=display_name, details=details,
)
stmt = stmt.on_conflict_do_update(
index_elements=[ServiceSeen.key],
set_={
"kind": stmt.excluded.kind,
"display_name": stmt.excluded.display_name,
"details": stmt.excluded.details,
"last_seen_at": func.now(),
},
)
await session.execute(stmt)
async def refresh_celery_roster(session: AsyncSession) -> None:
"""Inspect the broker and record what answered. Never raises.
A failure here means the roster does not advance, and the rows going stale
is then a TRUE report about a broker nobody can reach. Letting the
exception out would instead break the health endpoint, which is the one
thing that must keep answering when the stack is unwell.
"""
try:
grouped = await asyncio.wait_for(
asyncio.to_thread(_inspect_celery_sync),
timeout=INSPECT_TIMEOUT_SECONDS * 2,
)
except Exception:
log.warning("service roster: celery inspect failed; roster not refreshed", exc_info=True)
return
for queues, entry in grouped.items():
await touch_service(
session,
key="celery:" + ",".join(queues),
kind="celery",
display_name=role_display_name(queues),
details={
"queues": list(queues),
"hostnames": entry["hostnames"],
"replicas": len(entry["hostnames"]),
"active": entry["active"],
},
)
async def refresh_if_stale(session: AsyncSession) -> None:
"""Refresh the celery roster if nobody has for REFRESH_TTL_SECONDS.
Rate-limited by the data rather than by a lock: the gate is the newest
last_seen_at across the celery rows, which every web process can see. Two
processes racing through the gate costs one redundant inspect and writes
the same values twice, so the benign outcome needs no coordination to
prevent.
"""
newest = (
await session.execute(
select(func.max(ServiceSeen.last_seen_at)).where(ServiceSeen.kind == "celery")
)
).scalar_one_or_none()
if newest is not None:
age = (await session.execute(select(func.now()))).scalar_one() - newest
if age.total_seconds() < REFRESH_TTL_SECONDS:
return
await refresh_celery_roster(session)
+14
View File
@@ -133,6 +133,20 @@ services:
CELERY_RESULT_BACKEND: redis://redis:6379/0
SECRET_KEY: ${SECRET_KEY:-dev_secret_key_not_for_production_change_me}
LOG_LEVEL: ${LOG_LEVEL:-INFO}
# First boot only. FabledCurator refuses to start until the credential
# encryption key at /images/secrets/credential_key.b64 exists, and
# refuses to create one unless told to — auto-creating is
# indistinguishable from a restore that lost ./images/secrets, where it
# would mint a key that decrypts nothing and leave an instance that looks
# healthy while every paywalled download fails.
#
# Passed through EXPLICITLY because a variable in `.env` is only used for
# ${...} interpolation; it does not reach the container unless it is
# named here. Defaulted to empty so the refusal stands for everyone who
# has not opted in — the app tests for exactly "1".
#
# Set it in .env for one `up`, then remove it. See .env.example.
CURATOR_BOOTSTRAP_NEW_KEY: ${CURATOR_BOOTSTRAP_NEW_KEY:-}
volumes:
- ./images:/images
- ./import:/import
+59 -7
View File
@@ -5,9 +5,12 @@
<img src="/favicon.svg" alt="" class="fc-brand__glyph" width="22" height="22" />
<span class="fc-brand__text">FabledCurator</span>
</RouterLink>
<span class="fc-health" :title="health.label">
<RouterLink
:to="{ name: 'system' }" class="fc-health" :title="health.label"
:aria-label="`System health: ${health.label}`"
>
<v-icon size="x-small" :color="health.color">{{ health.icon }}</v-icon>
</span>
</RouterLink>
<PipelineStatusChip />
</div>
@@ -64,13 +67,15 @@
</template>
<script setup>
import { computed, onBeforeUnmount, onMounted, ref } from 'vue'
import { computed, onBeforeUnmount, onMounted, onUnmounted, ref } from 'vue'
import { useRoute } from 'vue-router'
import router, { FRONT_DOOR } from '../router.js'
import { useSystemStore } from '../stores/system.js'
import { useSystemHealthStore } from '../stores/systemHealth.js'
import PipelineStatusChip from './PipelineStatusChip.vue'
const system = useSystemStore()
const healthStore = useSystemHealthStore()
// Publish the nav's REAL height as --fc-nav-h so full-height workspaces
// (Explore/Subscriptions) and sticky sub-headers pin to it exactly instead of a
@@ -116,15 +121,55 @@ const settingsRoute = computed(() =>
navRoutes.value.find(r => r.name === 'settings') || null
)
// The dot beside the brand, and the only ambient signal that something in the
// stack has stopped (milestone 365).
//
// It used to read /api/health — a no-DB liveness check that proves the WEB
// container is serving and nothing else. Green there while the worker was dead
// is exactly what it looked like, and a green dot next to the product name is
// read as "everything is fine". It now reflects the whole-stack verdict.
//
// Deliberately re-using this element rather than adding a second indicator:
// there were already three partial surfaces (this, the pipeline chip, the
// Settings Activity tab) and a fourth would have made the question harder to
// answer, not easier. This is the one that already occupied the slot.
const health = computed(() => {
if (system.healthy === null) {
const overall = healthStore.overall
if (overall === null) {
return { icon: 'mdi-circle-outline', color: 'on-surface', label: 'checking…' }
}
if (system.healthy === true) {
return { icon: 'mdi-circle', color: 'success', label: 'healthy' }
if (overall === 'ok') {
return { icon: 'mdi-circle', color: 'success', label: 'All parts running' }
}
return { icon: 'mdi-alert-circle', color: 'error', label: 'unreachable' }
// Name what is wrong in the tooltip. "Something is unhealthy" sends someone
// hunting; "Scheduler has not checked in for 6 min" does not.
const worst = healthStore.problems[0]
const others = healthStore.problems.length - 1
const suffix = others > 0 ? ` (+${others} more)` : ''
if (overall === 'down') {
return {
icon: 'mdi-alert-circle', color: 'error',
label: (worst?.detail || 'A part has stopped') + suffix,
}
}
if (overall === 'stale') {
return {
icon: 'mdi-alert', color: 'warning',
label: (worst?.detail || 'A part is quiet') + suffix,
}
}
return { icon: 'mdi-help-circle-outline', color: 'on-surface', label: 'Health unknown' }
})
const HEALTH_POLL_MS = 15_000
let healthTimer = null
onMounted(() => {
healthStore.refresh()
healthTimer = setInterval(() => {
if (!document.hidden) healthStore.refresh()
}, HEALTH_POLL_MS)
})
onUnmounted(() => { if (healthTimer) clearInterval(healthTimer) })
</script>
<style scoped>
@@ -237,7 +282,14 @@ const health = computed(() => {
display: flex;
align-items: center;
flex-shrink: 0;
/* A RouterLink since milestone 365 — it is the path to /system, not just an
indicator. Reset the anchor so turning a span into a link changed nothing
about how the nav reads. */
text-decoration: none;
color: inherit;
border-radius: 50%;
}
.fc-health:hover { background: rgb(var(--v-theme-on-surface) / 0.12); }
.fc-nav-right {
flex: 1 1 0;
min-width: 0;
+7
View File
@@ -1,5 +1,6 @@
import { createRouter, createWebHistory, createMemoryHistory } from 'vue-router'
import SettingsView from './views/SettingsView.vue'
import SystemView from './views/SystemView.vue'
import GalleryView from './views/GalleryView.vue'
import ShowcaseView from './views/ShowcaseView.vue'
import ExploreView from './views/ExploreView.vue'
@@ -45,6 +46,12 @@ const routes = [
// Settings — config, pinned to the right of the nav (TopNav special-cases it).
{ path: '/settings', name: 'settings', component: SettingsView, meta: { title: 'Settings', stickyChrome: true } },
// Deliberately NO meta.title: TopNav builds its nav row from routes that
// have one, and this is reached from the health indicator beside the
// brand — the place someone already looks when they suspect something is
// wrong. A sixth top-level tab for a page you visit twice a year would
// cost more attention than it returns.
{ path: '/system', name: 'system', component: SystemView },
// The old standalone paths now redirect into the Browse hub, preserving any
// deep-link query (e.g. /posts?post_id=N → /browse?tab=posts&post_id=N). The
+6 -1
View File
@@ -4,7 +4,12 @@ import { useApi } from '../composables/useApi.js'
export const useSystemStore = defineStore('system', () => {
const api = useApi()
const healthy = ref(null) // null=unknown, true=ok, false=down
// NOT what the nav dot reads any more (milestone 365): that is the
// whole-stack verdict in systemHealth.js. /api/health only proves the web
// container is serving, which is why a green dot here sat happily beside a
// dead worker. refreshHealth() is still called — it is also how build/version
// info arrives — so this stays as its by-product rather than its purpose.
const healthy = ref(null)
// What the instance says it is. Since milestone 318 stopped publishing
// version image tags, this is the only answer to "which build is this?" —
// there is no registry name left to check it against.
+49
View File
@@ -0,0 +1,49 @@
import { defineStore } from 'pinia'
import { computed, ref } from 'vue'
import { useApi } from '../composables/useApi.js'
// Whole-stack health: is every part of FabledCurator running (milestone 365)?
//
// Distinct from `system.js`, which polls /api/health — a no-DB liveness check
// that only proves the web container is serving. That endpoint answers "can I
// reach the API"; this one answers "is anything broken", which is the question
// a green dot beside the brand was already being read as answering.
//
// Also distinct from `systemActivity.js`, which is about what the pipeline is
// DOING — queue depths, running tasks, failures. Running and alive are
// different questions and they fail independently: a perfectly idle stack with
// a dead worker looks identical to a healthy one on the activity surfaces.
export const useSystemHealthStore = defineStore('systemHealth', () => {
const api = useApi()
const overall = ref(null) // null until the first answer: unknown ≠ ok
const parts = ref([])
const checkedAt = ref(null)
const thresholds = ref(null) // server-owned, so the UI keeps no second copy
const lastError = ref(null)
async function refresh() {
try {
const body = await api.get('/api/system/health')
overall.value = body.overall
parts.value = body.parts || []
checkedAt.value = body.checked_at
thresholds.value = body.thresholds || null
lastError.value = null
} catch (e) {
// The endpoint is built never to fail because a dependency failed, so a
// throw here means the API itself is unreachable — which is its own kind
// of unhealthy and must not be shown as "ok".
lastError.value = e.message
overall.value = 'unknown'
}
return overall.value
}
// The parts worth naming in a tooltip — everything that is not ok, worst
// first. The endpoint already sorts that way.
const problems = computed(() => parts.value.filter(p => p.state !== 'ok'))
return { overall, parts, checkedAt, thresholds, lastError, problems, refresh }
})
+129
View File
@@ -0,0 +1,129 @@
<template>
<v-container class="py-6" style="max-width: 900px">
<div class="d-flex align-center mb-1">
<h1 class="text-h5">System</h1>
<v-spacer />
<span class="fc-sys__checked">
{{ store.checkedAt ? `checked ${formatRelative(store.checkedAt)}` : 'checking…' }}
</span>
</div>
<p class="fc-sys__lede text-body-2 mb-5">
Every moving part of FabledCurator and whether it is still checking in.
Parts are learned as they appear, so anything that has run at least once
stays listed that is what lets a stopped one be noticed rather than
simply vanishing.
</p>
<v-alert
v-if="store.lastError" type="error" variant="tonal" density="compact" class="mb-4"
>
Could not reach FabledCurator: {{ store.lastError }}
</v-alert>
<v-card v-else variant="flat" class="fc-sys__card">
<div v-if="!store.parts.length" class="pa-6 text-center fc-sys__muted">
Still gathering this fills in on the first check.
</div>
<div
v-for="part in store.parts" :key="part.key"
class="fc-sys__row" :class="`fc-sys__row--${part.state}`"
>
<span class="fc-sys__dot" :class="`fc-sys__dot--${part.state}`" />
<div class="fc-sys__body">
<div class="fc-sys__name">
{{ part.name }}
<span class="fc-sys__kind">{{ kindLabel(part.kind) }}</span>
</div>
<!-- The sentence, not just a chip. At the moment someone is deciding
whether to go and open Portainer, "has not checked in for 6 min"
is the thing that answers them. -->
<div class="fc-sys__detail">{{ part.detail }}</div>
</div>
<div class="fc-sys__meta">
<div v-if="part.last_seen_at" :title="part.last_seen_at">
seen {{ formatRelative(part.last_seen_at) }}
</div>
<div v-if="part.latency_ms != null">{{ part.latency_ms }} ms</div>
<div v-if="part.queues?.length" class="fc-sys__queues">{{ part.queues.join(', ') }}</div>
</div>
</div>
</v-card>
<p v-if="store.thresholds" class="fc-sys__foot text-caption mt-4">
A part is called stale after
{{ Math.round(store.thresholds.stale_after_seconds / 60) }} min without a
check-in and treated as stopped after
{{ Math.round(store.thresholds.down_after_seconds / 60) }} min. The window
is deliberately wide: a rolling deploy briefly runs two of a service and
then neither, and an indicator that reddened on every update would stop
being read.
</p>
</v-container>
</template>
<script setup>
import { onMounted, onUnmounted } from 'vue'
import { useSystemHealthStore } from '../stores/systemHealth.js'
import { formatRelative } from '../utils/date.js'
const store = useSystemHealthStore()
// Slower than the pipeline chip's 8s: liveness changes on the scale of
// container restarts, not task starts, and this page is open while someone
// watches it.
const POLL_MS = 10_000
let timer = null
function kindLabel(kind) {
if (kind === 'celery') return 'background worker'
if (kind === 'agent') return 'GPU agent'
if (kind === 'datastore') return 'datastore'
return kind
}
onMounted(() => {
store.refresh()
timer = setInterval(() => { if (!document.hidden) store.refresh() }, POLL_MS)
})
onUnmounted(() => { if (timer) clearInterval(timer) })
</script>
<style scoped>
.fc-sys__lede, .fc-sys__muted, .fc-sys__checked, .fc-sys__foot {
color: rgb(var(--v-theme-on-surface) / 0.66);
}
.fc-sys__checked { font-size: 0.78rem; }
.fc-sys__card { background: rgb(var(--v-theme-on-surface) / 0.04); }
.fc-sys__row {
display: flex; align-items: center; gap: 12px;
padding: 12px 16px;
border-bottom: 1px solid rgb(var(--v-theme-on-surface) / 0.08);
}
.fc-sys__row:last-child { border-bottom: 0; }
.fc-sys__dot { width: 9px; height: 9px; border-radius: 50%; flex: 0 0 auto; }
.fc-sys__dot--ok { background: rgb(var(--v-theme-success)); }
.fc-sys__dot--stale { background: rgb(var(--v-theme-warning)); }
.fc-sys__dot--down { background: rgb(var(--v-theme-error)); }
.fc-sys__dot--unknown { background: rgb(var(--v-theme-on-surface) / 0.35); }
.fc-sys__body { min-width: 0; flex: 1 1 auto; }
.fc-sys__name { font-weight: 600; }
.fc-sys__kind {
margin-left: 8px; font-weight: 400; font-size: 0.72rem; text-transform: uppercase;
letter-spacing: 0.04em; color: rgb(var(--v-theme-on-surface) / 0.5);
}
.fc-sys__detail { font-size: 0.82rem; color: rgb(var(--v-theme-on-surface) / 0.72); }
.fc-sys__meta {
text-align: right; font-size: 0.75rem; flex: 0 0 auto;
font-variant-numeric: tabular-nums; color: rgb(var(--v-theme-on-surface) / 0.6);
}
.fc-sys__queues { opacity: 0.75; }
</style>
+155
View File
@@ -0,0 +1,155 @@
"""Prove a freshly built image can still do the things its OS packages provide.
Run INSIDE the image, not against the source tree. That distinction is the
entire reason this file exists.
`ci.yml`'s lanes run on `ci-python:3.14` and install `requirements.txt`. A base
refresh changes neither, so all five lanes stay green through a base bump that
breaks the product. What a refresh actually re-resolves is this, from the
Dockerfile:
RUN apt-get update && apt-get install -y --no-install-recommends \
ffmpeg unar libpq5 postgresql-client zstd megatools \
libjpeg62-turbo libwebp7 libpng16-16 ca-certificates
Unpinned, every build. Nothing else in this repo looks at it.
So the checks below run the APPLICATION'S OWN code — `Thumbnailer`, which needs
no database and no app context — against whatever Pillow and ffmpeg have
become. `ffmpeg -version` exiting 0 would pass while a codec removal or an
soname bump broke every thumbnail in the library; producing a thumbnail would
not.
Every failure names the package it implicates. This fires on a Sunday,
unattended, about a change nobody made deliberately — "assertion failed" a week
later teaches nobody anything.
Usage: docker run --rm -i <image> shell -c 'python3 -' < scripts/smoke_image.py
"""
from __future__ import annotations
import shutil
import subprocess
import tempfile
from pathlib import Path
try:
from PIL import Image
from backend.app.services.thumbnailer import Thumbnailer
except Exception as exc: # noqa: BLE001 — a smoke test reports, it never raises
print(f"smoke: FAILED — could not import the thumbnail path at all: {exc}")
print(" Implicates Pillow or its shared libraries (libjpeg62-turbo,")
print(" libpng16-16, libwebp7), or the python base image itself.")
raise SystemExit(1) from exc
# Binary → what stops working without it. Listed individually because
# `--no-install-recommends` means any one of them can vanish on its own when a
# dependency chain higher up changes.
REQUIRED_BINARIES = {
"ffmpeg": "video thumbnails and transcoding (Dockerfile: ffmpeg)",
"unar": "archive import — cbz/zip/rar members (Dockerfile: unar)",
"pg_dump": "database backup (Dockerfile: postgresql-client)",
"zstd": "backup compression, pg_dump | tar --zstd (Dockerfile: zstd)",
"megatools": "mega.nz public-link downloads, #830 (Dockerfile: megatools)",
}
def check_jpeg(thumbs: Thumbnailer, src: Path) -> None:
path = src / "flat.jpg"
Image.new("RGB", (900, 400), (30, 90, 160)).save(path, "JPEG")
result = thumbs.generate_image_thumbnail(path, "a" * 64)
assert result.mime == "image/jpeg", f"mime was {result.mime}"
assert result.path.stat().st_size > 0, "no bytes written"
# Re-open it. A file that writes but cannot be read back is the shape a
# half-broken codec produces, and size alone would not catch it.
with Image.open(result.path) as im:
im.load()
def check_png_alpha(thumbs: Thumbnailer, src: Path) -> None:
path = src / "alpha.png"
Image.new("RGBA", (400, 900), (200, 40, 40, 128)).save(path, "PNG")
result = thumbs.generate_image_thumbnail(path, "b" * 64)
assert result.mime == "image/png", f"mime was {result.mime}"
with Image.open(result.path) as im:
im.load()
assert im.mode in ("RGBA", "LA", "P"), f"alpha lost, mode={im.mode}"
def check_webp(thumbs: Thumbnailer, src: Path) -> None:
path = src / "sample.webp"
Image.new("RGB", (500, 500), (10, 140, 70)).save(path, "WEBP")
result = thumbs.generate_image_thumbnail(path, "c" * 64)
assert result.path.stat().st_size > 0, "no bytes written"
def check_video(thumbs: Thumbnailer, src: Path) -> None:
# Synthesised rather than committed as a fixture: a checked-in video is a
# binary blob nobody can review, and lavfi ships with every ffmpeg build.
#
# 3 seconds, not 2. The seek lands at max(1.0, duration * 0.05) = 1.0s, and
# a clip barely longer than its own seek is how #1231 produced zero frames.
# This check exists to exercise ffmpeg, not to re-litigate that edge.
clip = src / "clip.mp4"
subprocess.run(
["ffmpeg", "-nostdin", "-f", "lavfi", "-i", "testsrc=size=640x360:rate=10",
"-t", "3", "-pix_fmt", "yuv420p", "-y", str(clip)],
check=True, capture_output=True, timeout=120,
)
result = thumbs.generate_video_thumbnail(clip, "d" * 64, duration_seconds=3.0)
assert result.path.stat().st_size > 0, "no bytes written"
with Image.open(result.path) as im:
im.load()
CHECKS = (
("JPEG thumbnail", "libjpeg62-turbo / Pillow", check_jpeg),
("PNG thumbnail (alpha)", "libpng16-16 / Pillow", check_png_alpha),
("WebP decode", "libwebp7 / Pillow", check_webp),
("video thumbnail", "ffmpeg", check_video),
)
def main() -> int:
failures: list[str] = []
print("smoke: binaries the apt layer provides")
for binary, purpose in REQUIRED_BINARIES.items():
if shutil.which(binary) is None:
print(f" FAIL {binary}: not on PATH")
failures.append(f"{binary}{purpose}")
else:
print(f" ok {binary}")
print("smoke: the application's own thumbnail path, against this image's libraries")
with tempfile.TemporaryDirectory() as tmp:
root = Path(tmp)
src = root / "src"
src.mkdir()
thumbs = Thumbnailer(root)
for name, implicates, fn in CHECKS:
try:
fn(thumbs, src)
print(f" ok {name}")
except Exception as exc: # noqa: BLE001 — report every check, then fail once
print(f" FAIL {name}: {exc}")
failures.append(f"{name}{implicates}")
if failures:
print(f"\nsmoke: FAILED — {len(failures)} check(s)")
for failure in failures:
print(f" - {failure}")
print("\nThis image was built against freshly resolved base layers. The")
print("named packages are where to look: compare this build's apt versions")
print("against the previous :latest before assuming the app changed.")
return 1
print("\nsmoke: all checks passed")
return 0
if __name__ == "__main__":
raise SystemExit(main())