Files
FabledCurator/ci-requirements.md
T
bvandeusen 63e0a423d7
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
CI / extension-version (push) Successful in 3s
Build images / build-agent (push) Successful in 7s
Build images / build-ml (push) Successful in 8s
Build images / build-web (push) Successful in 6s
extension / lint (push) Successful in 20s
CI / frontend-build (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 31s
CI / integration (push) Successful in 3m50s
ci: a weekly base-image refresh on the channel tags (milestone 326 step 4)
Skip-if-exists is keyed on our own source, so an artifact whose source
stops moving stops picking up base-image updates. `agent/` last changed
2026-07-17; every push since has correctly declined to rebuild it, which
also means it will serve that day's nvidia/cuda layers indefinitely.

A `schedule:` trigger, Sunday 06:00 UTC, away from CI-runner's Monday
security sweep so the two are never diagnosing each other.

#3154's blocking open question is dissolved rather than answered. It was
written when the identity was a `r-<revision>` TAG, and asked how the
next ordinary push could avoid repointing :latest back off the refresh.
Milestone 318 replaced that tag with a LABEL, and #3183 made the repoint
step exclude its source tag so the label stays readable. Excluding the
source is what also keeps a refresh from being undone: on the next main
push the reuse check hits, :latest is not rewritten, and the new :c-<sha>
is written FROM the refreshed :latest. To be verified by digest, not by
this argument.

Four decisions, each commented where it lives:

* It builds `main`, not the branch that triggered it. Forgejo registers a
  cron from the default branch — `dev` here — so a scheduled run arrives
  with github.ref on dev, and a refresh of :dev would be refreshing the
  one channel that is rebuilt constantly anyway. The ref is decided once
  in a top-level `env: BUILD_REF` that all four checkouts take. Deriving
  it per job would let the halves disagree: sign-extension would derive
  dev's extension version while build-web bundled main's, and the release
  download would 404 on a version that exists perfectly well.

* It publishes only the channel tag. :c-<sha> for main's HEAD already
  names the bytes that commit built; re-pushing it over refreshed layers
  would break the one tag rule 145 makes immutable, and it is the
  rollback unit — so the breakage would surface on the day somebody
  needed it. The repoint step needs no schedule case: the tag list is the
  channel tag alone, SOURCE is the only entry, it is excluded as always,
  and the step correctly does nothing.

* It bypasses reuse by construction, since it rebuilds the same source
  and fc.revision always matches. Checked in the reuse step beside
  force_build, so one decision still drives both the build and the
  repoint.

* `pull: true`, on the scheduled path only, is the actual mechanism. A
  moved base tag changes the FROM layer's cache key and everything above
  it rebuilds; an unmoved one is satisfied by the registry cache and the
  refresh is a ~13s no-op that republishes nothing. That no-op is the
  point — :latest should change when there is something new in it, not
  every Sunday. The known lag, left deliberately: an apt package update
  while the base tag stands still is not caught, and closing it needs
  no-cache: true, which buys weekly churn for it.
2026-08-29 22:53:48 -04:00

15 KiB

CI Requirements — FabledCurator

Spec: https://git.fabledsword.com/bvandeusen/CI-runner/src/branch/main/docs/process.md

Runtime image

git.fabledsword.com/bvandeusen/ci-python:3.14

Image deps used

  • python 3.14
  • ruff (analyzer for backend/, tests/, alembic/, agent/, scripts/)
  • node (frontend job: npm install + vitest + vite build)
  • docker CLI + buildx (.forgejo/workflows/build.yml: build-web, build-ml, build-agent — Fabled-Git registry push, and imagetools inspect/create for the reuse path)

Secondary runtime image

node:24-bookworm-slim — .forgejo/workflows/extension.yml only.

.forgejo/workflows/release.yml runs on ci-python:3.14 like everything else and installs nothing: it needs git and stdlib python, and builds no image.

The extension lane is the one job that does NOT run on ci-python:3.14: it needs a current Node for web-ext and vitest and nothing Python at all. Kept on the upstream slim image rather than adding a Node toolchain to ci-python, per docs/process.md's "add deps to the image when used by >1 project".

Per-job tool installs

  • pip install -r requirements.txt pytest pytest-asyncio — in backend-lint-and-test and integration jobs
  • npm install --no-audit --no-fund — in frontend-build job
  • npm install --no-audit --no-fund — in extension.yml's lint job (web-ext + vitest)
  • unzip — in extension.yml's "Verify XPI contents" step, installed via apt only when absent (node:24-bookworm-slim may or may not carry it). Debian package, ~2s. Not worth baking into a shared image for a single consumer, per docs/process.md's ">1 project" rule.

Notes

  • Integration wall time ~3 min, dominated by pgvector container start + the pip install step (~30-45s on cold cache) + alembic + 300+ integration tests.
  • The pip install in two jobs is intentional and per docs/process.md's "add deps to image when used by >1 project" rule: FC alone is one Python project, so the deps live in requirements.txt and install per-job. Reconsider when a second Fabled-family Python backend lands.
  • Integration uses Fabled-Git Actions services: + socket-discovered bridge IPs because act_runner (swarm-runner v0.6+) puts services on the default bridge with no embedded DNS. The pattern is documented in the rulebook's fabled-git.md "CI philosophy" section and FC's ci.yml is the canonical example.
  • No package-lock.json is tracked yet (FC's feedback_no_local_runs memory bans npm install locally). Using npm install rather than npm ci until a lockfile lands.
  • No imagemagick / pandoc per-job installs needed.
  • extension/'s vitest specs load lib/*.js by evaluating the real file as a classic script (test/helpers/loadLib.js) rather than adding module.exports shims to production code — the libs ship as background.scripts, not ES modules, so the specs exercise exactly the bytes packaged into the XPI.
  • extension/scripts/packaging.sh is the single definition of what ships inside the XPI. Three consumers read from it rather than keeping their own copy: web-ext's --ignore-files (extension/package.json), the git log pathspec inside the script's own version derivation, and scripts/artifacts.sh, which appends the extension's set to web's because the web image bundles the signed XPI. Hand-kept copies of that one fact is what allowed issue #2397, so extension/test/version.spec.js asserts no workflow has reintroduced a literal :(exclude)extension/….
  • Packaged and version-relevant are two different sets (#3156). scripts/ is excluded from the XPI and is NOT excluded from the version derivation, because packaging.sh decides the version string stamped into the packaged manifest.json. The membership test is "can changing this file change the published bytes?", not "is this file copied in?" — which is why the script keeps two lists rather than one.
  • The shipped extension version is derived, not committed. It is the commit TIME of the newest packaged-extension change, rendered YYYY.M.D.HHMM UTC (family rules 148/149 — never a commit count, which orders by branch rather than by recency). build.yml's sign-extension computes it and stamps it into extension/manifest.json + package.json in the working tree before signing; the stamp is never committed. The version in the repo is wholly inert — since milestone 318 step 8 there is no hand-set MAJOR.MINOR either.
  • The extension is the one artifact that does not zero-pad, and that is not a drift (#3138). Mozilla's grammar for AMO is ^(0|[1-9][0-9]{0,8})([.](0|[1-9][0-9]{0,8})){0,3}$ — a segment is the single digit 0 or starts 1-9, and there are at most four. 2026.08.29.0201 is rejected; 2026.8.29.201 is the same value one character narrower per segment, and rule 148 defines comparison as numeric per segment, so nothing is reordered. ci.yml's extension-version lane asserts the derived string against that exact regex, plus a YYYY.M.D.HHMM shape check that would catch a regression to the pre-318 1.0.<minutes> — which AMO would accept and which orders below everything already signed. Checking here is the whole point: AMO 409s on re-signing, so a version it rejects is burned and cannot be reused. scripts/artifacts.sh version extension delegates to packaging.sh so the two cannot answer differently.
  • Every job that derives anything checks out with fetch-depth: 0 — all four build.yml jobs, ci.yml's extension-version and backend-lint-and-test (for tests/test_artifact_paths.py and test_artifact_identity.py), and release.yml, which additionally walks the tag graph. A depth-1 clone sees one commit and derives a wrong, too-low value rather than failing, so the full-history checkout is load-bearing rather than incidental.
  • scripts/artifacts.sh is the same shape one level up: one definition per artifact of what it is built from, and the two values derived from it. revision (12 hex of the newest commit touching that set) and version (YYYY.MM.DD.HHMM UTC, rule 148). Four artifacts, four independent answers, so a push touching only agent/ leaves web and ml alone. tests/test_artifact_paths.py reads each Dockerfile and asserts every COPY source is covered, so adding a COPY without updating the script fails CI.
  • A file that DECIDES an artifact's identity belongs in its set even though it is copied into nothingpackaging.sh for the extension and web (#3156), and artifacts.sh itself for web (#3202), which decides the FC_VERSION baked into that image. Only web needs the second entry: every artifact stamps a revision, but a revision has a backstop (a changed derivation stops matching the published label and forces a rebuild) and a version has none, because nothing compares it to anything. tests/test_artifact_paths.py's DERIVERS table is the guard.
  • Builds are skipped when the content is already published. Each image carries its revision as an fc.revision LABEL, and build.yml reads that label back off the moving channel tag (imagetools inspect --format). Equal to the derived revision means the bytes are already published, so the job repoints the remaining tags at the existing manifest instead of rebuilding. Two things this depends on: an inspect that errors for ANY reason reads as a MISS so no needed build is ever skipped, and the repoint must EXCLUDE the source tag — imagetools create wraps its source in a manifest index, and config labels do not resolve through an index, so writing the channel tag from itself destroys the label the next run reads (#3183).
  • The build pushes exactly ONE tag — the channel's — and every other tag is written registry-side afterwards (#3190). buildx on this runner pushes the first tag to the registry and then re-pushes the rest through the docker driver, out of a local image store that a registry-direct build never fills; it fails intermittently with tag does not exist. On dev that only reddens a job, but on main it silently skips :c-<sha> while :latest publishes fine — a missing rollback tag has no consumer that fails, so nothing but the red job would notice until somebody needs to roll back. imagetools create has no local store to be absent from, and it is the code the reuse path already ran, so both paths now share one proven route. The cost: :c-<sha> is an index rather than a plain image, so fc.revision does not resolve through it — nothing reads it there, and the index names the same manifest.
  • The image builds run on a docker-container buildx builder, and provenance/sbom are explicitly OFF (milestone 326 step 1). The builder is what makes a registry layer cache possible at all — the default docker driver cannot export one (#3114) — and it is #3190's leading suspect, since it is the driver that resolves image metadata against a local store a registry-direct push never fills. The attestation flags are load-bearing, not tidiness: on the container driver build-push-action@v5 defaults provenance to true when pushing, an attestation manifest makes the pushed tag a manifest INDEX, and config labels do not resolve through an index — so leaving them on would make every push read fc.revision=<none>, miss, and rebuild forever with every lane green. Same failure as #3183, different door. These jobs run inside a container against a mounted docker socket, so the buildkit container is a sibling rather than a child.
  • All three images import and export a registry layer cache (<image>:buildcache, mode=max). This is not an optimisation bolted onto the driver change — it is the other half of it. The docker-container driver gets a fresh buildkit instance per job and therefore has no local layer store at all, where the old docker driver at least reused whatever the runner's dockerd happened to hold. Measured on run 4896, the first builds after the driver moved: web 3m44s (was 2m23s), ml 3m49s (was 3m20s), agent 11m12s (was 9m26s) — every one slower. A :buildcache tag is read by every build that runs, is one moving ref per image, holds cache blobs rather than a shippable artifact, and is overwritten in place, so it is not a return of the per-version tags milestone 318 withdrew (#3114).
  • build.yml accepts a workflow_dispatch with force_build, which bypasses the reuse check for all three images. It exists because skip-if-exists makes its own build path untestable: agent/ has not changed since 2026-07-17, so the agent build has not run in six weeks and cannot be exercised on demand — and #3190 lives on exactly that path. Editing build.yml does not force a build either, deliberately: the workflow is not shipped bytes and is in no artifact's path set. The flag is read through github.event.inputs into an env var rather than interpolated into a run block, and it is checked inside the reuse step so that one decision drives both the build and the repoint.
  • A weekly schedule rebuilds all three images against fresh base layers (Sunday 06:00 UTC, milestone 326 step 4, #3154). Skip-if-exists is keyed on OUR source, so an artifact whose source stops moving stops picking up base updates — agent/ has not changed since 2026-07-17 and would otherwise serve that day's nvidia/cuda layers forever. Four things make it work:
    • It builds main, not the branch that triggered it. Forgejo registers a cron from the DEFAULT branch (dev here), so a scheduled run arrives with github.ref on dev. The ref is decided once in a top-level env: BUILD_REF that every checkout in the file takes, rather than per job — otherwise sign-extension would derive dev's extension version while build-web bundled main's, and the release download would 404 on a version that exists perfectly well.
    • It publishes only :latest. :c-<sha> for main's HEAD already names the bytes that commit built; re-pushing it over refreshed layers would break the one tag rule 145 makes immutable, and it is the rollback unit. The repoint step needs no schedule case for this — the tag list is the channel tag alone, so SOURCE is the only entry, it is excluded as always, and the step correctly does nothing.
    • :latest and :c-<sha> therefore diverge between a refresh and the next main push, by design. They re-converge on that push: it hits reuse (a refresh does not move fc.revision, because it does not touch the source), and the repoint writes the NEW :c-<sha> from the refreshed :latest. The push path needed no change for this, because the repoint already excluded the source tag — the same rule that keeps the label readable also keeps a refresh from being undone.
    • pull: true on the scheduled path only is the actual mechanism. If a base tag moved, the FROM layer's cache key changes and everything above it rebuilds; if it did not, the registry cache satisfies the whole graph and the refresh is a ~13s no-op that republishes nothing. That no-op is the point — :latest should change when there is something new in it, not every Sunday. The known lag: a Debian package update inside the apt-get install layer while the base tag stands still is not caught. Closing it needs no-cache: true, which buys weekly churn for it; the official python/cuda images rebuild with those updates baked in, so this is a lag rather than a hole.
  • FC_CHANNEL and FC_VERSION are build args, not runtime settings. build.yml passes them to the web image only — the ml and agent images have nothing to report them to. /api/health returns both, the foot of Settings renders them, and /api/extension/manifest reports the channel beside the extension version so an install can be traced to a channel. With no version image tags, that self-report is the ONLY answer to "which build is this?" — which is why a missing version renders unknown rather than a blank: an empty footer reads as "no version", a different and false claim. Both are declared LAST in the Dockerfile on purpose: an ARG invalidates every layer below it, and these are the values that differ between the dev and main builds of identical source, so placing them earlier would stop the two channels ever sharing a cached pip install. Empty by default — a local build then reports nothing rather than claiming a channel it is not on.
  • The channel is never folded into the version. A -dev suffix makes the extension's per-segment parseInt comparator read that segment as 0, so every dev build compares equal to every other — issue #2993 exactly (rule 149). frontend/test/systemBuild.spec.js pins the rendered version to the bare number.
  • Callers MUST set -f before substituting the script's output. Without it the shell expands test/** against the working tree and silently narrows the pattern to whatever files exist at that moment — a failure that looks like nothing until dev files start appearing in the XPI. test/version.spec.js asserts every --ignore-files consumer sets it, and that no consumer has quietly reinstated a hardcoded list.