Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / build-ml (push) Successful in 6s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 20s
extension / lint (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 33s
Build images / build-web (push) Successful in 56s
CI / integration (push) Successful in 1m50s
Milestone 362 step 1, closing #3265's root cause. The weekly base refresh rewrote all three `:latest` tags on 2026-08-30 with nothing changed in any of them. Not a cache miss — run 4934's log shows every content step CACHED and both bases resolved to unchanged pinned digests. buildkit stamps the image config with the wall clock of the build, so identical layers get republished under a new config blob and therefore a new manifest digest. The cost is not storage, it is meaning: `:latest` moved on a calendar, so a digest change stopped being evidence that anything was different. That is the one thing a digest is any use for, and it is load-bearing here — the reuse check, the `:c-<sha>` rollback story and any future redeploy signal all rest on it. SOURCE_DATE_EPOCH normalises `created` and the history timestamps, so the same source produces the same config bytes and the same digest, and pushing it is a registry no-op. The value is routed through artifacts.sh's existing `newest()` rather than taken from git separately. `revision`, `version` and now `epoch` are three fields of ONE lookup, so they cannot drift into naming different commits — a divergence that would stamp an image reproducibly against one commit while it reported being another, with both values looking perfectly well-formed. Note #3127 §2 is the record of what a second clock costs; this adds a view, not a clock. Also corrected: the build step comment and ci-requirements.md both described the churn as current behaviour with the fix as a "likely" future. They now describe what the file does. Tests pin the property the fix depends on, not the fix: epoch is the same commit version names, in both renderings including the extension's unpadded one, and it does not move between two calls on one checkout. A future refactor that gave epoch its own `git log` would pass every other test in that file. Not yet verified end to end — proving it needs two consecutive refreshes to land on the same digest, which is the next thing, and is the step #3265 exists because nobody did last time. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
1712 lines
94 KiB
YAML
1712 lines
94 KiB
YAML
name: Build images
|
|
|
|
on:
|
|
push:
|
|
# `:dev` builds were dropped 2026-05-26 to save a docker build per dev
|
|
# push, on the reasoning that "operator tests from `:latest` after
|
|
# merge-to-main". Restored 2026-08-27: that is testing by shipping, and
|
|
# family rules 146/147 now name it directly — `main` IS production, and a
|
|
# channel that can only be refreshed by shipping is not a channel. The
|
|
# pressure to merge in order to try something does not come from
|
|
# carelessness; it comes from `:dev` being unable to carry the build.
|
|
#
|
|
# All three images build on dev, deliberately: a `:dev` web image paired
|
|
# with a stale `:dev` ml or agent is a worse trap than no dev channel at
|
|
# all, since the mismatch only shows up as a runtime failure.
|
|
branches: [main, dev]
|
|
#
|
|
# NO tag trigger (milestone 318 step 2). A `v*` tag names a commit `main`
|
|
# already built and published; rebuilding it produces the same source under
|
|
# the same names and RE-PUSHES `:c-<sha>`, which rule 145 forbids even when
|
|
# the bytes match — image configs carry timestamps, so "same source" does
|
|
# not mean "same manifest". The release build was publishing nothing new
|
|
# and violating an immutability rule to do it.
|
|
#
|
|
# Releases still happen (rule 148, on explicit request per rule 2). They
|
|
# produce a changelog, not an image.
|
|
|
|
# The escape hatch for the one thing skip-if-exists makes untestable: a
|
|
# build that WOULD be skipped. `agent/` has not changed since 2026-07-17, so
|
|
# every push since has correctly declined to build it — which also means the
|
|
# agent build path has not run in six weeks and cannot be exercised on
|
|
# demand. #3190 lives on exactly that path.
|
|
#
|
|
# Editing build.yml does not force one either, and that is deliberate: the
|
|
# workflow is not shipped bytes, so it is in no artifact's path set. Putting
|
|
# it in one would re-version every artifact for a comment change.
|
|
#
|
|
# ONE input, not one per artifact. Forcing all three is cheap once the
|
|
# registry cache is warm (#3114), and three booleans is an interface nobody
|
|
# remembers the meaning of.
|
|
workflow_dispatch:
|
|
inputs:
|
|
force_build:
|
|
description: 'Rebuild every image even if the published revision matches'
|
|
type: boolean
|
|
default: false
|
|
|
|
# The base-image refresh (milestone 326 step 4, #3154).
|
|
#
|
|
# Skip-if-exists is keyed on OUR source, so an artifact whose source stops
|
|
# moving stops picking up base-image updates. `agent/` last changed
|
|
# 2026-07-17; every push since has correctly declined to rebuild it, which
|
|
# also means it will serve that day's `nvidia/cuda` layers forever. Nothing
|
|
# is wrong until it has been unchanged for months, which is precisely why
|
|
# this is a calendar trigger and not a condition on the push path.
|
|
#
|
|
# Weekly, Sunday 06:00 UTC. Away from CI-runner's Monday security sweep so
|
|
# the two are never diagnosing each other, and on the quietest day so a
|
|
# surprise rebuild is not competing with a push.
|
|
schedule:
|
|
- cron: '0 6 * * 0'
|
|
|
|
# Which branch a run BUILDS, as opposed to which one triggered it.
|
|
#
|
|
# They are the same thing on every trigger but `schedule`. Forgejo registers a
|
|
# cron from the DEFAULT branch — `dev` here — so a scheduled run arrives with
|
|
# `github.ref` pointing at dev, and a refresh that rebuilt `:dev` would be
|
|
# refreshing the one channel that gets rebuilt constantly anyway. Production is
|
|
# `main` (rule 147), and `:latest` is the tag that goes stale.
|
|
#
|
|
# So the ref is decided once, here, and every checkout in the file takes it.
|
|
# Deriving it per job invites the two halves to disagree: sign-extension would
|
|
# derive dev's extension version while build-web bundled main's, and the
|
|
# release download would 404 on a version that exists perfectly well.
|
|
env:
|
|
BUILD_REF: ${{ github.event_name == 'schedule' && 'main' || github.ref }}
|
|
|
|
# Requires repo secret RELEASE_TOKEN — a Forgejo PAT with scopes:
|
|
# - write:package, read:package (for docker push to git.fabledsword.com)
|
|
# - write:release (for ext-<version> release asset cache)
|
|
# - write:issue (for future issue-management automation)
|
|
# The injected GITHUB_TOKEN cannot be used — it lacks write:package.
|
|
|
|
jobs:
|
|
# Sign-or-fetch-from-cache: signs the extension via AMO if no ext-<version>
|
|
# Forgejo release exists yet, otherwise downloads the cached signed XPI.
|
|
# Result is uploaded as an Actions artifact for build-web to consume.
|
|
#
|
|
# Why this lives in build.yml (not a separate workflow): the image a push
|
|
# publishes MUST carry the XPI. A separate sign workflow racing build.yml
|
|
# leaves that image without one for ~5min (until the commit-back triggers
|
|
# another build). Inline ordering eliminates the race.
|
|
# Cache strategy: Forgejo Release Assets — picked 2026-05-25 over Generic
|
|
# Packages (cleaner API surface) and commit-back-to-side-branch (no extra
|
|
# branch to manage). AMO blocks re-signing the same version (returns 409),
|
|
# so signing is intentionally one-shot per version.
|
|
#
|
|
# BOTH branches sign (milestone 271 step 6, 2026-08-27). Not two signatures:
|
|
# the version is the commit TIME of the newest packaged-extension change, so
|
|
# dev and main derive the SAME number for the same extension source. A dev
|
|
# push that changes the extension signs it; the merge to main then finds the
|
|
# ext-<version> release already there, hits the cache, and bundles the
|
|
# byte-identical XPI into `:latest` with no second AMO call. One signature
|
|
# per extension CHANGE, shared by both channels — that is what makes two
|
|
# channels affordable, and it is why step 4 (derived version) had to land
|
|
# first. Ungating this while the version was still the hand-set 1.0.11 would
|
|
# have hit the existing ext-1.0.11 cache and bundled MAIN's stale XPI into
|
|
# `:dev` — a dev channel confidently serving old code.
|
|
#
|
|
# Unconditional since milestone 318 step 2: main and dev are now the only
|
|
# triggers, so the branch gate that used to exclude tag pushes matched
|
|
# everything. A condition that is always true reads as if some path avoids
|
|
# it, which is worse than no condition.
|
|
sign-extension:
|
|
runs-on: python-ci
|
|
container:
|
|
image: git.fabledsword.com/bvandeusen/ci-python:3.14
|
|
steps:
|
|
- uses: actions/checkout@v4
|
|
with:
|
|
# Not the triggering ref — see the `env:` block at the top. On a
|
|
# scheduled refresh this is `main`; on everything else it is the ref
|
|
# that fired, so this is a no-op on every ordinary path.
|
|
ref: ${{ env.BUILD_REF }}
|
|
# Full history is load-bearing, not a convenience: the version this
|
|
# job signs is derived from the commit TIME of the newest packaged
|
|
# extension change. A depth-1 clone sees one commit and derives a
|
|
# wrong, too-low value rather than failing (ci-requirements.md).
|
|
fetch-depth: 0
|
|
|
|
# BUILD_REF is what makes a scheduled run build `main` rather than the
|
|
# branch its cron fired from — and it is read through the `env` context
|
|
# inside `with:`, which this runner is NOT known to evaluate. If it does
|
|
# not, checkout silently falls back to the triggering ref and the weekly
|
|
# refresh publishes DEV's source to `:latest`, which is production.
|
|
# Every lane would stay green; the first sign of it would be production
|
|
# running code that was never merged.
|
|
#
|
|
# So assert the checkout instead of trusting the expression. A red
|
|
# weekly job is a fine outcome. Shipping dev to production is not.
|
|
#
|
|
# `if:` reads the `github` context, which the runner demonstrably does
|
|
# evaluate — this file already gates steps on it — so the guard cannot
|
|
# be disabled by the same uncertainty it exists to cover.
|
|
- name: Guard — a scheduled run must have checked out main
|
|
if: github.event_name == 'schedule'
|
|
run: |
|
|
set -eu
|
|
BRANCH=$(git rev-parse --abbrev-ref HEAD)
|
|
echo "schedule: HEAD is $BRANCH ($(git rev-parse --short HEAD))"
|
|
if [ "$BRANCH" != "main" ]; then
|
|
echo "schedule: expected main, got '$BRANCH'." >&2
|
|
echo "schedule: BUILD_REF was not honoured by the runner." >&2
|
|
echo "schedule: refusing to publish a channel tag from it." >&2
|
|
exit 1
|
|
fi
|
|
|
|
# The version is DERIVED, not read from the repo (milestone 271 step 4,
|
|
# cut over 2026-08-27). `packaging.sh version` returns `YYYY.M.D.HHMM`
|
|
# UTC — the commit TIME of the newest change to a PACKAGED extension
|
|
# file, per family rule 148/149. Never a commit count, which orders by
|
|
# branch rather than by recency.
|
|
#
|
|
# Unpadded, and only here: AMO's grammar rejects a leading zero, so the
|
|
# extension renders rule 148's numbers without the family's padding
|
|
# (milestone 318 step 8). Same value, one character narrower per segment;
|
|
# ci.yml's extension-version lane checks the string against Mozilla's
|
|
# published regex before this job ever calls AMO.
|
|
#
|
|
# The committed "version" in manifest.json / package.json decides NOTHING
|
|
# — not even a MAJOR.MINOR prefix, which step 8 removed. The stamp step
|
|
# below overwrites it in the working tree before web-ext ever reads it,
|
|
# and it is deliberately NOT committed back: the commit carrying the bump
|
|
# would itself be a change to the extension and would move the version
|
|
# again. The repo holds the source; the build derives the label.
|
|
- name: Derive extension version
|
|
id: extver
|
|
run: |
|
|
set -eu
|
|
VERSION=$(sh extension/scripts/packaging.sh version)
|
|
echo "version=$VERSION" >> "$GITHUB_OUTPUT"
|
|
echo "Derived extension version: $VERSION"
|
|
|
|
# Firefox refuses a downgrade and AMO never releases a burned version,
|
|
# so a version that moves BACKWARDS is unrecoverable: it strands every
|
|
# install that already took the higher one. Two ways it could happen —
|
|
# a checkout without full history (derives too low), or a rewritten
|
|
# history that drops the newest packaged commit.
|
|
#
|
|
# The test is `derived < highest already signed`, strictly. Equality is
|
|
# the ORDINARY case, not a fault: an unchanged extension derives the same
|
|
# version it did last build, which is exactly what lets the ext-<version>
|
|
# cache hit and holds AMO to one call per extension CHANGE. Only moving
|
|
# backwards is a failure, so this runs on every path — cache hit
|
|
# included — rather than only before a sign.
|
|
# --- shadow mode (milestone 313, step 2) -----------------------------
|
|
# Informational ONLY. Nothing reads this and it must never fail the
|
|
# build — no `set -e`, and every derivation falls back to UNAVAILABLE.
|
|
#
|
|
# What to watch across pushes, because this is what step 3 will trust:
|
|
# * a push touching only agent/ moves the agent and leaves web and ml
|
|
# STILL. If web moves, its path set is too wide.
|
|
# * a push touching only docs moves nothing.
|
|
# * a push touching the extension moves the extension AND web, since
|
|
# web bakes in the XPI. If web does not move, its set is too narrow:
|
|
# the reuse check hits, and the channel serves a web image bundling
|
|
# the PREVIOUS XPI while the freshly signed one is orphaned (#3156).
|
|
# * dev and main derive the same values for the same source.
|
|
- name: Shadow — derived artifact version (informational)
|
|
run: |
|
|
set -u
|
|
A=extension
|
|
V=$(sh scripts/artifacts.sh version "$A" 2>&1 || echo UNAVAILABLE)
|
|
R=$(sh scripts/artifacts.sh revision "$A" 2>&1 || echo UNAVAILABLE)
|
|
echo "derived: artifact=$A version=$V revision=$R sha=$GITHUB_SHA"
|
|
|
|
- name: Guard — the derived version must never go backwards
|
|
env:
|
|
TOKEN: ${{ secrets.RELEASE_TOKEN }}
|
|
DERIVED: ${{ steps.extver.outputs.version }}
|
|
run: |
|
|
python3 - <<'PY'
|
|
import json, os, sys, urllib.request
|
|
|
|
API = ("https://git.fabledsword.com/api/v1/repos/"
|
|
"bvandeusen/FabledCurator/releases")
|
|
headers = {"Authorization": "token " + os.environ["TOKEN"]}
|
|
|
|
# Paginated rather than first-page-only: ext-* releases share this
|
|
# list with the v* release tags, so one page would start missing them
|
|
# as those accumulate. The bound FAILS rather than silently scanning
|
|
# part of the list and calling the highest it saw the highest there is.
|
|
tags = []
|
|
for page in range(1, 21):
|
|
req = urllib.request.Request(
|
|
f"{API}?limit=50&page={page}", headers=headers)
|
|
with urllib.request.urlopen(req, timeout=30) as resp:
|
|
batch = json.load(resp)
|
|
if not batch:
|
|
break
|
|
tags += [r.get("tag_name", "") for r in batch]
|
|
else:
|
|
sys.exit("guard: >1000 releases — pagination bound reached")
|
|
|
|
def parse(v):
|
|
try:
|
|
return tuple(int(part) for part in v.split("."))
|
|
except ValueError:
|
|
return None
|
|
|
|
derived_s = os.environ["DERIVED"]
|
|
derived = parse(derived_s)
|
|
if derived is None:
|
|
sys.exit(f"guard: derived version {derived_s!r} is not numeric")
|
|
|
|
signed = sorted(
|
|
(v, t) for t in tags if t.startswith("ext-")
|
|
for v in [parse(t[4:])] if v
|
|
)
|
|
if not signed:
|
|
print("guard: no ext-* release yet — nothing to go backwards from")
|
|
raise SystemExit(0)
|
|
|
|
hi, hi_tag = signed[-1]
|
|
print(f"guard: derived={derived_s} highest already signed={hi_tag}")
|
|
if derived < hi:
|
|
sys.exit(
|
|
f"REFUSING TO SIGN: derived {derived_s} is OLDER than the "
|
|
f"already-signed {hi_tag}. Firefox would reject it as a "
|
|
f"downgrade, and AMO will not release the burned version. "
|
|
f"First thing to check: did this job check out with "
|
|
f"fetch-depth: 0?"
|
|
)
|
|
print("guard: ok")
|
|
PY
|
|
|
|
- name: Check Forgejo release-asset cache
|
|
id: cache
|
|
env:
|
|
TOKEN: ${{ secrets.RELEASE_TOKEN }}
|
|
run: |
|
|
set -eu
|
|
VERSION=${{ steps.extver.outputs.version }}
|
|
STATUS=$(curl -s -o release.json -w "%{http_code}" \
|
|
-H "Authorization: token $TOKEN" \
|
|
"https://git.fabledsword.com/api/v1/repos/bvandeusen/FabledCurator/releases/tags/ext-$VERSION" || echo 000)
|
|
echo "Tag lookup HTTP status: $STATUS"
|
|
# JSON parsing via python (ci-python:3.14 has stdlib json; jq is
|
|
# not in the image and adding it per ci-requirements.md is not
|
|
# warranted for a single consumer — operator-flagged 2026-05-26
|
|
# after a sign job failed with `jq: not found`).
|
|
if [ "$STATUS" = "200" ]; then
|
|
ASSET_ID=$(python3 -c "import json; r=json.load(open('release.json')); xpis=[a for a in r.get('assets', []) if a.get('name','').endswith('.xpi')]; print(xpis[0]['id'] if xpis else '')")
|
|
if [ -n "$ASSET_ID" ]; then
|
|
echo "cached=true" >> "$GITHUB_OUTPUT"
|
|
echo "asset_id=$ASSET_ID" >> "$GITHUB_OUTPUT"
|
|
echo "Cached XPI exists at ext-$VERSION (asset id $ASSET_ID); skipping AMO sign"
|
|
else
|
|
echo "cached=false" >> "$GITHUB_OUTPUT"
|
|
echo "Release ext-$VERSION exists but has no .xpi asset; will re-sign + re-upload"
|
|
fi
|
|
else
|
|
echo "cached=false" >> "$GITHUB_OUTPUT"
|
|
echo "No release named ext-$VERSION; will sign via AMO and upload"
|
|
fi
|
|
|
|
# No "download cached XPI in sign-extension" step: build-web
|
|
# fetches directly from the Forgejo ext-<version> release asset
|
|
# (removed 2026-05-26 alongside the actions/upload-artifact
|
|
# removal — sign-extension's job is just to ensure the cache
|
|
# exists on Forgejo; the build-web side reads it independently).
|
|
|
|
# web-ext signs whatever manifest.json says, so the derived value has to
|
|
# reach the tree before signing. package.json is written too: the two are
|
|
# required to agree (ci.yml's guard), and a local `npm run build` reads
|
|
# it. Working tree only — never committed, per the note on the derive
|
|
# step.
|
|
- name: Stamp the derived version into manifest.json + package.json
|
|
env:
|
|
DERIVED: ${{ steps.extver.outputs.version }}
|
|
run: |
|
|
python3 - <<'PY'
|
|
import json, os
|
|
|
|
version = os.environ["DERIVED"]
|
|
for path in ("extension/manifest.json", "extension/package.json"):
|
|
with open(path) as fh:
|
|
doc = json.load(fh)
|
|
doc["version"] = version
|
|
with open(path, "w") as fh:
|
|
json.dump(doc, fh, indent=2)
|
|
fh.write("\n")
|
|
print(f"{path}: version -> {version}")
|
|
PY
|
|
|
|
- name: Sign via AMO (cache miss)
|
|
if: steps.cache.outputs.cached != 'true'
|
|
run: |
|
|
cd extension && npm install --no-save --no-audit --no-fund && npm run sign
|
|
env:
|
|
WEB_EXT_API_KEY: ${{ secrets.MOZILLA_AMO_JWT_KEY }}
|
|
WEB_EXT_API_SECRET: ${{ secrets.MOZILLA_AMO_JWT_SECRET }}
|
|
|
|
- name: Upload signed XPI to ext-<version> release (cache miss)
|
|
if: steps.cache.outputs.cached != 'true'
|
|
env:
|
|
TOKEN: ${{ secrets.RELEASE_TOKEN }}
|
|
run: |
|
|
set -eux
|
|
VERSION=${{ steps.extver.outputs.version }}
|
|
# AMO renames signed XPIs with its internal addon-id-safe-string;
|
|
# canonicalize to fabledcurator-<version>.xpi so the FC server's
|
|
# whitelist (backend/app/frontend.py expects 'fabledcurator-*.xpi')
|
|
# keeps working.
|
|
SIGNED=$(ls extension/web-ext-artifacts/*.xpi | head -1)
|
|
XPI="extension/web-ext-artifacts/fabledcurator-$VERSION.xpi"
|
|
cp "$SIGNED" "$XPI"
|
|
# Find-or-create the ext-<version> release. Track whether WE
|
|
# created it so an upload failure below can roll back (don't
|
|
# leave an empty release tombstone that the next run's
|
|
# cache-check mistakes for a partial-failure state).
|
|
#
|
|
# target_commitish is the signing commit, not a branch name: since
|
|
# step 6 either branch can create this release, and hard-coding
|
|
# `main` would tag a dev-signed XPI against a main commit that may
|
|
# not even contain the extension source it was built from.
|
|
STATUS=$(curl -s -o release.json -w "%{http_code}" \
|
|
-H "Authorization: token $TOKEN" \
|
|
"https://git.fabledsword.com/api/v1/repos/bvandeusen/FabledCurator/releases/tags/ext-$VERSION" || echo 000)
|
|
if [ "$STATUS" = "200" ]; then
|
|
CREATED_BY_US=false
|
|
else
|
|
curl -s -X POST -H "Authorization: token $TOKEN" -H "Content-Type: application/json" \
|
|
-d "{\"tag_name\":\"ext-$VERSION\",\"name\":\"Extension $VERSION (signed XPI cache)\",\"body\":\"Internal cache for the signed XPI consumed by build.yml's build-web job. Not a user-facing FC release.\",\"target_commitish\":\"$GITHUB_SHA\"}" \
|
|
-o release.json \
|
|
"https://git.fabledsword.com/api/v1/repos/bvandeusen/FabledCurator/releases"
|
|
CREATED_BY_US=true
|
|
fi
|
|
RELEASE_ID=$(python3 -c "import json; print(json.load(open('release.json'))['id'])")
|
|
test -n "$RELEASE_ID"
|
|
# Rollback-on-failure: if the asset upload fails AND we just
|
|
# created the release in this run, delete it. Prevents an empty
|
|
# ext-<version> release from poisoning the next workflow run
|
|
# (operator-flagged 2026-05-26 — without rollback the next run
|
|
# saw 'release exists, no asset → cache miss → sign' which AMO
|
|
# then rejected with 409 'Version already exists').
|
|
rollback_if_we_created() {
|
|
if [ "$CREATED_BY_US" = "true" ]; then
|
|
echo "Rolling back: deleting just-created release $RELEASE_ID"
|
|
curl -s -X DELETE -H "Authorization: token $TOKEN" \
|
|
"https://git.fabledsword.com/api/v1/repos/bvandeusen/FabledCurator/releases/$RELEASE_ID" || true
|
|
curl -s -X DELETE -H "Authorization: token $TOKEN" \
|
|
"https://git.fabledsword.com/api/v1/repos/bvandeusen/FabledCurator/tags/ext-$VERSION" || true
|
|
fi
|
|
}
|
|
trap 'rollback_if_we_created' EXIT
|
|
HTTP_CODE=$(curl -s -X POST -H "Authorization: token $TOKEN" \
|
|
-F "attachment=@$XPI" \
|
|
-o /dev/null -w "%{http_code}" \
|
|
"https://git.fabledsword.com/api/v1/repos/bvandeusen/FabledCurator/releases/$RELEASE_ID/assets?name=fabledcurator-$VERSION.xpi")
|
|
if [ "$HTTP_CODE" != "201" ] && [ "$HTTP_CODE" != "200" ]; then
|
|
echo "Asset upload failed with HTTP $HTTP_CODE"
|
|
exit 1
|
|
fi
|
|
# Upload succeeded — clear the rollback trap.
|
|
trap - EXIT
|
|
echo "Uploaded fabledcurator-$VERSION.xpi to ext-$VERSION release"
|
|
|
|
# No actions/upload-artifact step: Forgejo Actions (and our
|
|
# act_runner) doesn't support upload-artifact@v4+ (GHES limitation
|
|
# surfaced 2026-05-26). Instead build-web reads the signed XPI
|
|
# straight from the ext-<version> Forgejo release we just uploaded
|
|
# to. Same source of truth; no double-store.
|
|
|
|
build-web:
|
|
# A plain `needs` — no `always()`. That expression existed to let a
|
|
# SKIPPED sign-extension through on a tag push while still blocking a
|
|
# FAILED one. With no tag trigger, sign-extension always runs, so the
|
|
# default behaviour is exactly what we want: a failed sign skips build-web
|
|
# rather than shipping an image without its XPI.
|
|
needs: [sign-extension]
|
|
runs-on: python-ci
|
|
container:
|
|
image: git.fabledsword.com/bvandeusen/ci-python:3.14
|
|
steps:
|
|
- uses: actions/checkout@v4
|
|
with:
|
|
# Not the triggering ref — see the `env:` block at the top. On a
|
|
# scheduled refresh this is `main`; on everything else it is the ref
|
|
# that fired, so this is a no-op on every ordinary path.
|
|
ref: ${{ env.BUILD_REF }}
|
|
# Full history: this job RE-DERIVES the extension version rather than
|
|
# being handed it, and a depth-1 clone derives a wrong, too-low value
|
|
# rather than failing — which would 404 the download of a release
|
|
# that exists perfectly well under its real name.
|
|
fetch-depth: 0
|
|
|
|
# See sign-extension's copy for why this guard exists.
|
|
- name: Guard — a scheduled run must have checked out main
|
|
if: github.event_name == 'schedule'
|
|
run: |
|
|
set -eu
|
|
BRANCH=$(git rev-parse --abbrev-ref HEAD)
|
|
echo "schedule: HEAD is $BRANCH ($(git rev-parse --short HEAD))"
|
|
if [ "$BRANCH" != "main" ]; then
|
|
echo "schedule: expected main, got '$BRANCH'." >&2
|
|
echo "schedule: BUILD_REF was not honoured by the runner." >&2
|
|
echo "schedule: refusing to publish a channel tag from it." >&2
|
|
exit 1
|
|
fi
|
|
|
|
# --- derived values, one line (milestone 313) ------------------------
|
|
# These stopped being shadow output at step 3. `revision` decides
|
|
# whether the build below runs at all and `version` is what the image
|
|
# reports about itself; the load-bearing steps each print only the one
|
|
# they use, so this is the only place the pair appears together. When a
|
|
# build is skipped, this is the line that says what the commit derived.
|
|
#
|
|
# Still diagnostic, so it still must not fail the build — no `set -e`,
|
|
# and every derivation falls back to UNAVAILABLE. A broken echo must
|
|
# never be the reason an image does not ship.
|
|
#
|
|
# What it should say:
|
|
# * a push touching only agent/ moves the agent and leaves web and ml
|
|
# STILL. If web moves, its path set is too wide.
|
|
# * a push touching only docs moves nothing.
|
|
# * a push touching the extension moves the extension AND web, since
|
|
# web bakes in the XPI. If web does not move, its set is too narrow:
|
|
# the reuse check hits, and the channel serves a web image bundling
|
|
# the PREVIOUS XPI while the freshly signed one is orphaned (#3156).
|
|
# * dev and main derive the same values for the same source.
|
|
- name: Report the derived artifact version
|
|
run: |
|
|
set -u
|
|
A=web
|
|
V=$(sh scripts/artifacts.sh version "$A" 2>&1 || echo UNAVAILABLE)
|
|
R=$(sh scripts/artifacts.sh revision "$A" 2>&1 || echo UNAVAILABLE)
|
|
echo "derived: artifact=$A version=$V revision=$R sha=$GITHUB_SHA"
|
|
|
|
- name: Determine tag
|
|
id: tag
|
|
run: |
|
|
# Two trigger shapes, and between them they publish three tags:
|
|
# main → :latest (production, moving — rule 147: main IS production)
|
|
# :c-<sha> (immutable, the rollback unit — rule 145)
|
|
# dev → :dev (the rolling test channel — rule 146)
|
|
#
|
|
# That is the whole list. No :<version>, and no :main — rule 145,
|
|
# narrowed 2026-08-28 once it was verified that nothing pins:
|
|
# "a third name for the same thing is upkeep for a model we do not
|
|
# run." The date tag published between milestone 313 step 3 and
|
|
# milestone 318 was exactly that; :main was a second moving name for
|
|
# whatever :latest already pointed at.
|
|
#
|
|
# `dev` gets no :c-<sha> deliberately. On a channel whose entire
|
|
# contract is that it moves, a per-push immutable tag is a rollback
|
|
# target nobody has ever pulled, accumulating forever. The accepted
|
|
# cost: on dev there is no rollback but the previous :dev, which is
|
|
# gone — recovery is revert-on-git plus a CI cycle.
|
|
#
|
|
# Reinstating :<version> is a real decision, not a default. It earns
|
|
# its place when something genuinely pins: a second instance held on
|
|
# a known-good build, or a deliberately frozen window. Tag at the
|
|
# moment you decide to freeze; no back-catalogue is needed.
|
|
#
|
|
# POSIX-safe substring (the runner shell is dash/BusyBox sh, not
|
|
# bash — `${var:0:7}` errors with "Bad substitution"; cut works
|
|
# everywhere). Operator-flagged 2026-06-01 after the first :c-<sha>
|
|
# main-push build failed at this step.
|
|
SHORT_SHA=$(printf '%s' "$GITHUB_SHA" | cut -c1-7)
|
|
|
|
# A scheduled refresh publishes the CHANNEL and nothing else
|
|
# (#3154). :c-<sha> for main's HEAD already exists and names the
|
|
# bytes that commit actually built; re-pushing it over refreshed
|
|
# base layers would break the one tag rule 145 makes immutable —
|
|
# and it is the rollback unit, so the breakage would surface on the
|
|
# day somebody needed it.
|
|
#
|
|
# The accepted consequence: between a refresh and the next main
|
|
# push, :latest and :c-<sha> point at different manifests. That is
|
|
# the design, not drift. They RE-CONVERGE on that push — it hits
|
|
# reuse (a refresh does not move fc.revision, because it does not
|
|
# touch the source), and the repoint step then writes the new
|
|
# :c-<sha> from the refreshed :latest. So the rollback unit ends up
|
|
# naming the bytes production is actually running, which is the
|
|
# property that matters.
|
|
#
|
|
# Checked BEFORE the ref test, not after: a scheduled run's
|
|
# GITHUB_REF is the default branch (dev), so the main test would
|
|
# never fire on it.
|
|
if [ "${GITHUB_EVENT_NAME:-}" = "schedule" ]; then
|
|
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator:latest" >> "$GITHUB_OUTPUT"
|
|
echo "channel=main" >> "$GITHUB_OUTPUT"
|
|
elif [ "${GITHUB_REF##*/}" = "main" ]; then
|
|
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator:latest,git.fabledsword.com/bvandeusen/fabledcurator:c-${SHORT_SHA}" >> "$GITHUB_OUTPUT"
|
|
echo "channel=main" >> "$GITHUB_OUTPUT"
|
|
else
|
|
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator:dev" >> "$GITHUB_OUTPUT"
|
|
echo "channel=dev" >> "$GITHUB_OUTPUT"
|
|
fi
|
|
|
|
# A shell step, not docker/login-action@v3, because the action's shared
|
|
# cache races itself (#3118). act_runner caches a remote action under one
|
|
# /root/.cache/act/<hash> per runner, and build-web, build-ml and
|
|
# build-agent all start in the same second and all want this same action.
|
|
# One job re-clones the directory — which empties and repopulates it —
|
|
# while another is walking it to copy into its container, and the walker
|
|
# lstat()s a file that has just vanished. It failed twice on 2026-08-27,
|
|
# naming a DIFFERENT missing file each time (`eslint.config.mjs`, then
|
|
# `jest.config.ts`), which is what rules out a corrupt cache and points at
|
|
# a race. The loser dies with MODULE_NOT_FOUND on dist/index.js before the
|
|
# action runs at all, so the secret is never even reached.
|
|
#
|
|
# Nothing is lost by dropping it: logging in is one command, the docker
|
|
# CLI is already in the CI image (ci-requirements.md), and the same
|
|
# reasoning as family rule 5 applies — a marketplace action buys nothing
|
|
# when the tool is baked into the image the workflow already selected.
|
|
#
|
|
# Password on stdin, never as an argument: an argument lands in the
|
|
# process table and draws docker's own deprecation warning.
|
|
- name: Login to Forgejo registry
|
|
env:
|
|
TOKEN: ${{ secrets.RELEASE_TOKEN }}
|
|
ACTOR: ${{ github.actor }}
|
|
run: echo "$TOKEN" | docker login git.fabledsword.com -u "$ACTOR" --password-stdin
|
|
|
|
# A REAL buildx builder, not the default `docker` driver (#3114, #3190).
|
|
#
|
|
# The default driver builds through the local dockerd. It cannot export a
|
|
# registry cache at all — which is why the agent rebuilds a ~6.3 GB CUDA
|
|
# + torch image from scratch whenever the runner's local cache is cold,
|
|
# measured at 9m26s against 7s warm. It is also #3190's leading suspect:
|
|
# after a registry-direct push it resolves image metadata against a local
|
|
# store the push never filled, and reports `No such image` on an image
|
|
# that published perfectly well three seconds earlier.
|
|
#
|
|
# These jobs run INSIDE a container against a mounted docker socket, so
|
|
# the buildkit container this starts is a SIBLING of the job container,
|
|
# not a child. That works over the socket mount; it had never been tried
|
|
# here before milestone 326 step 1.
|
|
- name: Set up buildx
|
|
uses: docker/setup-buildx-action@v3
|
|
|
|
# --- reuse-if-published (milestone 313, step 4) ----------------------
|
|
# Does the image the channel tag already points at carry THIS commit's
|
|
# revision? If so the bytes this job would produce are already published
|
|
# and the build is pure waste: the remaining tags get repointed at that
|
|
# existing manifest instead, registry-side, in seconds.
|
|
#
|
|
# Keyed on an `fc.revision` LABEL rather than on a tag of its own
|
|
# (milestone 318 step 3). A tag would be a name minted per build that one
|
|
# thing reads — what rule 145 narrowed against — and would be prunable
|
|
# under the registry's keep_pattern (#3157), silently expiring the cache.
|
|
# A label rides inside a tag that has to exist anyway.
|
|
#
|
|
# An image with no such label reads as a miss and rebuilds. That is the
|
|
# migration, not a fault: labels cannot be backfilled, since the reuse
|
|
# path copies a manifest and config labels are not manifest annotations.
|
|
# Each artifact pays one rebuild, once.
|
|
#
|
|
# This is what stops a push that touched only `agent/` from rebuilding
|
|
# web and ml, and a merge to main from rebuilding what dev already built.
|
|
#
|
|
# The failure direction is deliberate. An inspect that errors for ANY
|
|
# reason — network, auth, a registry hiccup — reads as a miss and the
|
|
# build runs. Only a genuine 200 skips one, so there is no path here
|
|
# that skips a build that was actually needed; the worst case is paying
|
|
# for a build we could have avoided.
|
|
#
|
|
# BASE-IMAGE FRESHNESS: an artifact whose source stops moving stops
|
|
# picking up base-image updates. Milestone 318 removed the argument this
|
|
# used to need rather than answering it — with no version tags there is
|
|
# no immutable name a refresh could contradict, and rule 145 already
|
|
# allows a rebuild with different contents to republish a MOVING tag.
|
|
# So a refresh is just a build. A scheduled channel-only one is tracked
|
|
# separately (#3154); it does not belong in the push path.
|
|
- name: Is this content already published?
|
|
id: reuse
|
|
env:
|
|
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator
|
|
CHANNEL: ${{ steps.tag.outputs.channel }}
|
|
# Empty on a push; the string "true" only from a workflow_dispatch
|
|
# that asked for it. `github.event.inputs` rather than the `inputs`
|
|
# context — release.yml already uses that form, and it is the one
|
|
# this runner is known to evaluate. Read through env rather than
|
|
# interpolated into the run block, same rule as release.yml's TAG.
|
|
FORCE: ${{ github.event.inputs.force_build }}
|
|
# A scheduled refresh has to bypass reuse by construction: it
|
|
# rebuilds the SAME source, so fc.revision always matches and the
|
|
# check would skip every refresh there has ever been.
|
|
EVENT: ${{ github.event_name }}
|
|
run: |
|
|
set -eu
|
|
DERIVED=$(sh scripts/artifacts.sh revision web)
|
|
echo "revision=$DERIVED" >> "$GITHUB_OUTPUT"
|
|
# Baked into the web image as FC_VERSION and reported by /api/health.
|
|
# A pure function of the revision — same commit, same string — so it
|
|
# adds no variability the reuse check would have to account for.
|
|
echo "version=$(sh scripts/artifacts.sh version web)" >> "$GITHUB_OUTPUT"
|
|
|
|
# The build clock, pinned to the same commit (#3265). Without it
|
|
# buildkit stamps the image config with the wall clock of the build,
|
|
# so identical layers republish under a new config blob and the
|
|
# channel tag gets a new manifest digest for no reason. Derived from
|
|
# `newest()` like revision and version, so all three name one commit
|
|
# and cannot drift apart.
|
|
echo "epoch=$(sh scripts/artifacts.sh epoch web)" >> "$GITHUB_OUTPUT"
|
|
|
|
# The moving tag for this channel. Which tag we ask IS the channel —
|
|
# that is why the revision needs no -main/-dev qualifier any more.
|
|
if [ "$CHANNEL" = "main" ]; then T=latest; else T=dev; fi
|
|
echo "channel_ref=$IMAGE:$T" >> "$GITHUB_OUTPUT"
|
|
|
|
# Compare VALUES, never exit codes. Measured on buildx v0.36.1
|
|
# (run 4732): a missing key returns an empty string and exits 0, so
|
|
# branching on the exit code would read "no label yet" as success.
|
|
# An unreachable tag also lands here as empty via the `|| echo`.
|
|
# Empty never equals a 12-char revision, so every uncertain case
|
|
# falls through to a build — the safe direction, with no special
|
|
# casing for it.
|
|
#
|
|
# Read the SPECIFIC key. The map also carries whatever the base image
|
|
# set, and `org.opencontainers.image.version` sits right beside ours
|
|
# looking like a plausible answer (it reads 24.04 on the agent).
|
|
PUBLISHED=$(docker buildx imagetools inspect "$IMAGE:$T" \
|
|
--format '{{ index .Image.Config.Labels "fc.revision" }}' \
|
|
2>/dev/null || echo "")
|
|
echo "reuse: $IMAGE:$T carries fc.revision=${PUBLISHED:-<none>}; derived=$DERIVED"
|
|
if [ -z "$PUBLISHED" ] && docker buildx imagetools inspect "$IMAGE:$T" >/dev/null 2>&1; then
|
|
# The tag resolves but carries no readable label. Expected exactly
|
|
# once per artifact, during the migration onto labels. If it recurs
|
|
# every push, something is rewriting the channel tag as a manifest
|
|
# index — see the repoint step's note.
|
|
echo "reuse: NOTE $IMAGE:$T exists but has no readable fc.revision."
|
|
echo "reuse: NOTE Fine once, while migrating. Every push means the"
|
|
echo "reuse: NOTE tag is being index-wrapped and reuse is dead."
|
|
fi
|
|
|
|
# FORCE is checked here rather than in the build step's `if:`, so
|
|
# that one decision drives everything downstream. The repoint step
|
|
# keys off `hit` too, and a force that bypassed only the build would
|
|
# leave the two disagreeing about what just happened.
|
|
if [ "${FORCE:-false}" = "true" ]; then
|
|
echo "hit=false" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: force_build set — building regardless"
|
|
elif [ "${EVENT:-}" = "schedule" ]; then
|
|
echo "hit=false" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: scheduled base refresh — building regardless"
|
|
elif [ -n "$PUBLISHED" ] && [ "$PUBLISHED" = "$DERIVED" ]; then
|
|
echo "hit=true" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: already published — skipping the build"
|
|
else
|
|
echo "hit=false" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: not published — building"
|
|
fi
|
|
|
|
- name: Download signed XPI from Forgejo release asset
|
|
# dev and main each bundle the XPI their own sign-extension just
|
|
# published — the point of the channel work (milestone 271 step 6): the
|
|
# dev image carries the extension being developed, rather than
|
|
# requiring a merge to try it.
|
|
#
|
|
# The 10-minute polling loop that used to live here is gone with the
|
|
# tag trigger (milestone 318 step 2). It existed for one shape only: a
|
|
# release cut fired the tag build and the main build together, the tag
|
|
# build skipped sign-extension and raced straight here, and it lost
|
|
# every time (operator-flagged 2026-05-27 after v26.05.27.0). Polling
|
|
# was the fix for a build that should not have been running.
|
|
#
|
|
# sign-extension is a `needs` dependency and it succeeded, so the
|
|
# release exists. A single fetch is correct, and a 404 now means a real
|
|
# disagreement about the derived version rather than a race — which is
|
|
# exactly what should fail loudly instead of being slept through.
|
|
#
|
|
# Still gated on the reuse miss: a published image already contains its
|
|
# XPI, so this would fetch a file nothing then reads.
|
|
if: steps.reuse.outputs.hit != 'true'
|
|
env:
|
|
TOKEN: ${{ secrets.RELEASE_TOKEN }}
|
|
run: |
|
|
set -eux
|
|
# Re-derived, not read from the repo: sign-extension published
|
|
# ext-<derived>, and the committed version has been inert since
|
|
# milestone 271 step 4. Both jobs run `packaging.sh version` over the
|
|
# same commit, so they agree by construction — and if they ever
|
|
# didn't, this download 404s and the build fails loudly instead of
|
|
# shipping a stale XPI.
|
|
VERSION=$(sh extension/scripts/packaging.sh version)
|
|
# One fetch, no retry. sign-extension ran to success in this same
|
|
# workflow and published ext-$VERSION; both jobs derive $VERSION from
|
|
# the same commit, so they agree by construction. A 404 here means
|
|
# they did NOT agree, and sleeping on that would only delay the
|
|
# report.
|
|
STATUS=$(curl -s -o release.json -w "%{http_code}" \
|
|
-H "Authorization: token $TOKEN" \
|
|
"https://git.fabledsword.com/api/v1/repos/bvandeusen/FabledCurator/releases/tags/ext-$VERSION" || echo 000)
|
|
if [ "$STATUS" != "200" ]; then
|
|
echo "ERROR: ext-$VERSION release not found (HTTP $STATUS)."
|
|
echo "sign-extension succeeded in this run, so it published some"
|
|
echo "other version — the two jobs derived different values for one"
|
|
echo "commit. Check that both checked out with fetch-depth: 0."
|
|
exit 1
|
|
fi
|
|
# Extract the .xpi asset's browser_download_url (Forgejo's
|
|
# /releases/assets/<id> endpoint returns ASSET METADATA, not
|
|
# the binary blob — operator-flagged 2026-05-26: my prior
|
|
# code curl'd the metadata endpoint without -f and wrote the
|
|
# resulting 404-page-not-found text into fabledcurator-*.xpi,
|
|
# which Firefox then rejected as "corrupt").
|
|
# browser_download_url is the canonical binary endpoint and
|
|
# is also publicly accessible (no token needed) but we pass
|
|
# the token anyway for symmetry with private-repo support.
|
|
DOWNLOAD_URL=$(python3 -c "import json; r=json.load(open('release.json')); xpis=[a for a in r.get('assets', []) if a.get('name','').endswith('.xpi')]; print(xpis[0]['browser_download_url'])")
|
|
test -n "$DOWNLOAD_URL"
|
|
echo "Downloading XPI from: $DOWNLOAD_URL"
|
|
mkdir -p frontend/public/extension
|
|
DEST="frontend/public/extension/fabledcurator-$VERSION.xpi"
|
|
# -f = fail on HTTP error (prevents silent corruption like the
|
|
# 2026-05-26 incident); -L = follow redirects.
|
|
curl -sfL -H "Authorization: token $TOKEN" -o "$DEST" "$DOWNLOAD_URL"
|
|
# Sanity check: the binary should start with the ZIP magic (PK\x03\x04).
|
|
# If it's anything else, the next docker build will ship a corrupt XPI.
|
|
MAGIC=$(head -c 2 "$DEST" | od -An -c | tr -d ' \n')
|
|
if [ "$MAGIC" != "PK" ]; then
|
|
echo "ERROR: downloaded XPI does not start with ZIP magic 'PK' (got '$MAGIC')"
|
|
echo "File contents preview:"
|
|
head -c 200 "$DEST"
|
|
exit 1
|
|
fi
|
|
cp "$DEST" "frontend/public/extension/fabledcurator-latest.xpi"
|
|
ls -la frontend/public/extension/
|
|
|
|
- name: Build and push web image
|
|
if: steps.reuse.outputs.hit != 'true'
|
|
# Read by buildx out of the ENVIRONMENT, not passed as a build-arg —
|
|
# it normalises the image config's `created` field and the history
|
|
# timestamps rather than being consumed by the Dockerfile. See #3265
|
|
# and the reuse step's `epoch` output.
|
|
env:
|
|
SOURCE_DATE_EPOCH: ${{ steps.reuse.outputs.epoch }}
|
|
uses: docker/build-push-action@v5
|
|
with:
|
|
context: .
|
|
file: Dockerfile
|
|
push: true
|
|
# Re-resolve the FROM references against the registry instead of
|
|
# trusting whatever digest the cache was built against. This is the
|
|
# whole mechanism of the scheduled refresh (#3154): if the base tag
|
|
# moved, the FROM layer's cache key changes, every layer above it
|
|
# invalidates, and the image genuinely rebuilds.
|
|
#
|
|
# MEASURED on the first real fire, run 4934 (#3265): when the base
|
|
# did NOT move, the build was ~13s with every content step CACHED —
|
|
# and the channel tag STILL got a new manifest digest, because
|
|
# buildkit stamps a fresh image config per run and republishes the
|
|
# identical layers under it. All three images moved that way on
|
|
# 2026-08-30 with nothing whatsoever changed in them.
|
|
#
|
|
# SOURCE_DATE_EPOCH (below) is the fix: pinned to the commit the
|
|
# content came from, the config is byte-identical across runs, so
|
|
# the manifest digest is too and the push is a registry no-op. A
|
|
# digest change means the content changed again, which is the only
|
|
# thing a digest is any use for.
|
|
#
|
|
# What `pull` does NOT catch either: a Debian package update inside
|
|
# the `apt-get install` layer while the base tag itself stands
|
|
# still. The official python/cuda images rebuild with those updates
|
|
# baked in, so this is a lag rather than a hole; closing it needs
|
|
# `no-cache: true`, which is a much larger version of the same
|
|
# churn #3265 is about.
|
|
#
|
|
# Only on the schedule. An ordinary push wants the cached base.
|
|
pull: ${{ github.event_name == 'schedule' }}
|
|
# ONE tag, the channel's. Every other tag is written by the step
|
|
# below, registry-side. buildx here pushes the first tag to the
|
|
# registry and then re-pushes the rest through the DOCKER driver,
|
|
# out of a local image store a registry-direct build never filled —
|
|
# #3190, which cost `main` its :c-<sha> on 2026-08-29 while :latest
|
|
# published perfectly well.
|
|
tags: ${{ steps.reuse.outputs.channel_ref }}
|
|
# The reuse key. Read back off the channel tag on the next push to
|
|
# decide whether that push needs to build at all, so this is not
|
|
# decoration — an unstamped image is one that will always rebuild.
|
|
labels: |
|
|
fc.revision=${{ steps.reuse.outputs.revision }}
|
|
# LOAD-BEARING, not a preference. On the default docker driver these
|
|
# were no-ops; on the docker-container driver above,
|
|
# build-push-action@v5 defaults provenance to TRUE when pushing.
|
|
# Provenance attaches an attestation manifest, which makes the pushed
|
|
# tag a manifest INDEX — and `.Image.Config.Labels` does not resolve
|
|
# through an index.
|
|
#
|
|
# The label directly above IS the reuse key. Wrap the channel tag in
|
|
# an index and the next push reads fc.revision=<none>, misses, and
|
|
# rebuilds. Then so does the one after that, forever. Nothing fails,
|
|
# nothing goes red, and the only symptom is the bill. That is #3183
|
|
# arriving through a different door, and note #3127 §4 records the
|
|
# same shape for `platforms:`.
|
|
provenance: false
|
|
sbom: false
|
|
# The ONLY cache this driver can have. `docker-container` gets a
|
|
# FRESH buildkit instance per job, so unlike the default docker
|
|
# driver it has no local layer store to fall back on — measured on
|
|
# run 4896, the first builds after the driver change: web 3m44s
|
|
# (was 2m23s), ml 3m49s (was 3m20s), agent 11m12s (was 9m26s). The
|
|
# driver change ALONE is a regression; this is the other half of it.
|
|
#
|
|
# mode=max so intermediate stages cache too. web's frontend-builder
|
|
# stage and the agent's two ~150s pip layers are the whole cost, and
|
|
# they are exactly what a min-mode cache would drop.
|
|
#
|
|
# A `:buildcache` tag is NOT the withdrawn tag scheme coming back.
|
|
# Rule 145 narrowed against names NOTHING reads; this one is read by
|
|
# every build that runs, is one moving ref per image rather than one
|
|
# per build, holds cache blobs rather than a shippable artifact, and
|
|
# is overwritten in place rather than accumulating. It is closer to
|
|
# :dev than to the :2026.8.28 tags milestone 318 deleted. (#3114.)
|
|
cache-from: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator:buildcache
|
|
cache-to: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator:buildcache,mode=max
|
|
# Only the web image carries these: it is the one with a UI and an
|
|
# HTTP surface to report them on. The ml and agent images have
|
|
# nothing to tell.
|
|
build-args: |
|
|
FC_CHANNEL=${{ steps.tag.outputs.channel }}
|
|
FC_VERSION=${{ steps.reuse.outputs.version }}
|
|
|
|
# Every tag but the channel's own is written HERE, registry-side,
|
|
# whether or not a build ran. Each -t becomes another reference to the
|
|
# SAME manifest the channel tag holds, so :c-<sha> is byte-identical to
|
|
# what is published rather than a lookalike rebuild.
|
|
#
|
|
# Owning the build path too is #3190's fix, not a tidy-up:
|
|
#
|
|
# #27 pushing …/fabledcurator:latest DONE 15.8s
|
|
# #28 pushing …/fabledcurator:c-0e15c44 with docker
|
|
# #28 ERROR: tag does not exist: …:c-0e15c44
|
|
#
|
|
# Intermittent — build-ml made the identical two-tag push seconds later
|
|
# and succeeded — and worse than it looks. `:latest` had already
|
|
# published, so production was correct while the immutable rollback tag
|
|
# rule 145 requires of every main push simply did not exist. Nothing but
|
|
# the red job would ever have noticed: a missing :c-<sha> has no
|
|
# consumer that fails, so it surfaces when somebody needs to roll back.
|
|
#
|
|
# `imagetools create` is a registry-side manifest copy — no layer
|
|
# transfer, no local daemon, nothing that can be absent. The reuse case
|
|
# has always gone this way, so this puts the build case on the code that
|
|
# was already proven rather than on a second path.
|
|
#
|
|
# Running on every path also keeps family rule 146 true: a rolling
|
|
# channel refreshes itself, so skipping a build must never leave :dev or
|
|
# :latest pointing at something older than the commit just pushed.
|
|
#
|
|
# The cost, accepted knowingly: `imagetools create` wraps its source in
|
|
# an index, so :c-<sha> becomes an index and fc.revision does not
|
|
# resolve through it. Nothing reads that label off :c-<sha> — the reuse
|
|
# check only ever inspects the CHANNEL tag — and the index names the
|
|
# same manifest, so a pull is byte-identical. The reuse path already
|
|
# produced :c-<sha> this way; this only makes it uniform.
|
|
- name: Write the remaining tags from the published image
|
|
env:
|
|
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator
|
|
SOURCE: ${{ steps.reuse.outputs.channel_ref }}
|
|
TAGS: ${{ steps.tag.outputs.tags }}
|
|
run: |
|
|
set -euf
|
|
# The source tag is EXCLUDED from the targets, and that is load-
|
|
# bearing rather than an optimisation.
|
|
#
|
|
# `imagetools create` wraps the source manifest in an INDEX. Point it
|
|
# at the channel tag with that same tag as a target and the tag stops
|
|
# being a plain image — after which `.Image.Config.Labels` no longer
|
|
# resolves through it and the fc.revision label reads as absent. The
|
|
# next push then misses and rebuilds, so reuse worked exactly once
|
|
# and every subsequent push paid full price. Observed on run 4751:
|
|
# ml:dev reported fc.revision=<none> one push after run 4749 had read
|
|
# a7e626a67a79 off it. Nothing failed; the savings just evaporated.
|
|
#
|
|
# Excluding the source means the channel tag is only ever written
|
|
# by a real build, so it stays a plain image and stays readable.
|
|
# On dev that leaves nothing to do either way: the build pushed :dev
|
|
# itself, or the hit established it was already right. On main it
|
|
# leaves :c-<sha>, which rule 145 requires of every main push whether
|
|
# or not a build ran.
|
|
#
|
|
# steps.tag emits ONE comma-separated list; imagetools wants a -t per
|
|
# ref. (That list used to feed docker/build-push-action directly —
|
|
# which is exactly what #3190 made unsafe.)
|
|
ARGS=""
|
|
IFS=,
|
|
for t in $TAGS; do
|
|
[ "$t" = "$SOURCE" ] && continue
|
|
ARGS="$ARGS -t $t"
|
|
done
|
|
unset IFS
|
|
#
|
|
# This is also the whole of the scheduled refresh's tag handling
|
|
# (#3154): a refresh's tag list is the channel tag alone, so SOURCE
|
|
# is the only entry, it gets excluded, and this step correctly does
|
|
# nothing. No `if:` on the step and no schedule special-case —
|
|
# excluding the source was already the right rule.
|
|
if [ -z "$ARGS" ]; then
|
|
echo "repoint: $SOURCE is the only tag for this channel and"
|
|
echo "repoint: already holds this revision — nothing to write."
|
|
exit 0
|
|
fi
|
|
# shellcheck disable=SC2086
|
|
docker buildx imagetools create $ARGS "$SOURCE"
|
|
echo "repointed from $SOURCE:$ARGS"
|
|
|
|
build-ml:
|
|
runs-on: python-ci
|
|
container:
|
|
image: git.fabledsword.com/bvandeusen/ci-python:3.14
|
|
steps:
|
|
- uses: actions/checkout@v4
|
|
with:
|
|
# Not the triggering ref — see the `env:` block at the top. On a
|
|
# scheduled refresh this is `main`; on everything else it is the ref
|
|
# that fired, so this is a no-op on every ordinary path.
|
|
ref: ${{ env.BUILD_REF }}
|
|
# Full history: this job derives its artifact's version from the
|
|
# commit its shipped files last changed in (milestone 313). A
|
|
# depth-1 clone cannot see that commit — it either derives a wrong,
|
|
# too-low value or finds nothing at all, and neither is a failure
|
|
# the build would otherwise notice.
|
|
fetch-depth: 0
|
|
|
|
# See sign-extension's copy for why this guard exists.
|
|
- name: Guard — a scheduled run must have checked out main
|
|
if: github.event_name == 'schedule'
|
|
run: |
|
|
set -eu
|
|
BRANCH=$(git rev-parse --abbrev-ref HEAD)
|
|
echo "schedule: HEAD is $BRANCH ($(git rev-parse --short HEAD))"
|
|
if [ "$BRANCH" != "main" ]; then
|
|
echo "schedule: expected main, got '$BRANCH'." >&2
|
|
echo "schedule: BUILD_REF was not honoured by the runner." >&2
|
|
echo "schedule: refusing to publish a channel tag from it." >&2
|
|
exit 1
|
|
fi
|
|
|
|
# --- derived values, one line (milestone 313) ------------------------
|
|
# These stopped being shadow output at step 3. `revision` decides
|
|
# whether the build below runs at all and `version` is what the image
|
|
# reports about itself; the load-bearing steps each print only the one
|
|
# they use, so this is the only place the pair appears together. When a
|
|
# build is skipped, this is the line that says what the commit derived.
|
|
#
|
|
# Still diagnostic, so it still must not fail the build — no `set -e`,
|
|
# and every derivation falls back to UNAVAILABLE. A broken echo must
|
|
# never be the reason an image does not ship.
|
|
#
|
|
# What it should say:
|
|
# * a push touching only agent/ moves the agent and leaves web and ml
|
|
# STILL. If web moves, its path set is too wide.
|
|
# * a push touching only docs moves nothing.
|
|
# * a push touching the extension moves the extension AND web, since
|
|
# web bakes in the XPI. If web does not move, its set is too narrow:
|
|
# the reuse check hits, and the channel serves a web image bundling
|
|
# the PREVIOUS XPI while the freshly signed one is orphaned (#3156).
|
|
# * dev and main derive the same values for the same source.
|
|
- name: Report the derived artifact version
|
|
run: |
|
|
set -u
|
|
A=ml
|
|
V=$(sh scripts/artifacts.sh version "$A" 2>&1 || echo UNAVAILABLE)
|
|
R=$(sh scripts/artifacts.sh revision "$A" 2>&1 || echo UNAVAILABLE)
|
|
echo "derived: artifact=$A version=$V revision=$R sha=$GITHUB_SHA"
|
|
|
|
- name: Determine tag
|
|
id: tag
|
|
run: |
|
|
# Mirrors build-web's tag list; see the comment there.
|
|
# POSIX-safe substring (the runner shell is dash/BusyBox sh, not
|
|
# bash — `${var:0:7}` errors with "Bad substitution"; cut works
|
|
# everywhere). Operator-flagged 2026-06-01 after first :c-<sha>
|
|
# main-push build failed at this step.
|
|
SHORT_SHA=$(printf '%s' "$GITHUB_SHA" | cut -c1-7)
|
|
# Mirrors build-web's tag list and its schedule handling; see
|
|
# the comments there.
|
|
if [ "${GITHUB_EVENT_NAME:-}" = "schedule" ]; then
|
|
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator-ml:latest" >> "$GITHUB_OUTPUT"
|
|
echo "channel=main" >> "$GITHUB_OUTPUT"
|
|
elif [ "${GITHUB_REF##*/}" = "main" ]; then
|
|
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator-ml:latest,git.fabledsword.com/bvandeusen/fabledcurator-ml:c-${SHORT_SHA}" >> "$GITHUB_OUTPUT"
|
|
echo "channel=main" >> "$GITHUB_OUTPUT"
|
|
else
|
|
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator-ml:dev" >> "$GITHUB_OUTPUT"
|
|
echo "channel=dev" >> "$GITHUB_OUTPUT"
|
|
fi
|
|
|
|
# Shell step rather than docker/login-action — see build-web's note on
|
|
# the shared action-cache race (#3118).
|
|
- name: Login to Forgejo registry
|
|
env:
|
|
TOKEN: ${{ secrets.RELEASE_TOKEN }}
|
|
ACTOR: ${{ github.actor }}
|
|
run: echo "$TOKEN" | docker login git.fabledsword.com -u "$ACTOR" --password-stdin
|
|
|
|
# A REAL buildx builder, not the default `docker` driver (#3114, #3190).
|
|
#
|
|
# The default driver builds through the local dockerd. It cannot export a
|
|
# registry cache at all — which is why the agent rebuilds a ~6.3 GB CUDA
|
|
# + torch image from scratch whenever the runner's local cache is cold,
|
|
# measured at 9m26s against 7s warm. It is also #3190's leading suspect:
|
|
# after a registry-direct push it resolves image metadata against a local
|
|
# store the push never filled, and reports `No such image` on an image
|
|
# that published perfectly well three seconds earlier.
|
|
#
|
|
# These jobs run INSIDE a container against a mounted docker socket, so
|
|
# the buildkit container this starts is a SIBLING of the job container,
|
|
# not a child. That works over the socket mount; it had never been tried
|
|
# here before milestone 326 step 1.
|
|
- name: Set up buildx
|
|
uses: docker/setup-buildx-action@v3
|
|
|
|
# --- reuse-if-published (milestone 313, step 4) ----------------------
|
|
# Does the image the channel tag already points at carry THIS commit's
|
|
# revision? If so the bytes this job would produce are already published
|
|
# and the build is pure waste: the remaining tags get repointed at that
|
|
# existing manifest instead, registry-side, in seconds.
|
|
#
|
|
# Keyed on an `fc.revision` LABEL rather than on a tag of its own
|
|
# (milestone 318 step 3). A tag would be a name minted per build that one
|
|
# thing reads — what rule 145 narrowed against — and would be prunable
|
|
# under the registry's keep_pattern (#3157), silently expiring the cache.
|
|
# A label rides inside a tag that has to exist anyway.
|
|
#
|
|
# An image with no such label reads as a miss and rebuilds. That is the
|
|
# migration, not a fault: labels cannot be backfilled, since the reuse
|
|
# path copies a manifest and config labels are not manifest annotations.
|
|
# Each artifact pays one rebuild, once.
|
|
#
|
|
# This is what stops a push that touched only `agent/` from rebuilding
|
|
# web and ml, and a merge to main from rebuilding what dev already built.
|
|
#
|
|
# The failure direction is deliberate. An inspect that errors for ANY
|
|
# reason — network, auth, a registry hiccup — reads as a miss and the
|
|
# build runs. Only a genuine 200 skips one, so there is no path here
|
|
# that skips a build that was actually needed; the worst case is paying
|
|
# for a build we could have avoided.
|
|
#
|
|
# BASE-IMAGE FRESHNESS: an artifact whose source stops moving stops
|
|
# picking up base-image updates. Milestone 318 removed the argument this
|
|
# used to need rather than answering it — with no version tags there is
|
|
# no immutable name a refresh could contradict, and rule 145 already
|
|
# allows a rebuild with different contents to republish a MOVING tag.
|
|
# So a refresh is just a build. A scheduled channel-only one is tracked
|
|
# separately (#3154); it does not belong in the push path.
|
|
- name: Is this content already published?
|
|
id: reuse
|
|
env:
|
|
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator-ml
|
|
CHANNEL: ${{ steps.tag.outputs.channel }}
|
|
# Empty on a push; the string "true" only from a workflow_dispatch
|
|
# that asked for it. `github.event.inputs` rather than the `inputs`
|
|
# context — release.yml already uses that form, and it is the one
|
|
# this runner is known to evaluate. Read through env rather than
|
|
# interpolated into the run block, same rule as release.yml's TAG.
|
|
FORCE: ${{ github.event.inputs.force_build }}
|
|
# A scheduled refresh has to bypass reuse by construction: it
|
|
# rebuilds the SAME source, so fc.revision always matches and the
|
|
# check would skip every refresh there has ever been.
|
|
EVENT: ${{ github.event_name }}
|
|
run: |
|
|
set -eu
|
|
DERIVED=$(sh scripts/artifacts.sh revision ml)
|
|
echo "revision=$DERIVED" >> "$GITHUB_OUTPUT"
|
|
|
|
# The build clock, pinned to the same commit (#3265). Without it
|
|
# buildkit stamps the image config with the wall clock of the build,
|
|
# so identical layers republish under a new config blob and the
|
|
# channel tag gets a new manifest digest for no reason. Derived from
|
|
# `newest()` like revision and version, so all three name one commit
|
|
# and cannot drift apart.
|
|
echo "epoch=$(sh scripts/artifacts.sh epoch ml)" >> "$GITHUB_OUTPUT"
|
|
|
|
# The moving tag for this channel. Which tag we ask IS the channel —
|
|
# that is why the revision needs no -main/-dev qualifier any more.
|
|
if [ "$CHANNEL" = "main" ]; then T=latest; else T=dev; fi
|
|
echo "channel_ref=$IMAGE:$T" >> "$GITHUB_OUTPUT"
|
|
|
|
# Compare VALUES, never exit codes. Measured on buildx v0.36.1
|
|
# (run 4732): a missing key returns an empty string and exits 0, so
|
|
# branching on the exit code would read "no label yet" as success.
|
|
# An unreachable tag also lands here as empty via the `|| echo`.
|
|
# Empty never equals a 12-char revision, so every uncertain case
|
|
# falls through to a build — the safe direction, with no special
|
|
# casing for it.
|
|
#
|
|
# Read the SPECIFIC key. The map also carries whatever the base image
|
|
# set, and `org.opencontainers.image.version` sits right beside ours
|
|
# looking like a plausible answer (it reads 24.04 on the agent).
|
|
PUBLISHED=$(docker buildx imagetools inspect "$IMAGE:$T" \
|
|
--format '{{ index .Image.Config.Labels "fc.revision" }}' \
|
|
2>/dev/null || echo "")
|
|
echo "reuse: $IMAGE:$T carries fc.revision=${PUBLISHED:-<none>}; derived=$DERIVED"
|
|
if [ -z "$PUBLISHED" ] && docker buildx imagetools inspect "$IMAGE:$T" >/dev/null 2>&1; then
|
|
# The tag resolves but carries no readable label. Expected exactly
|
|
# once per artifact, during the migration onto labels. If it recurs
|
|
# every push, something is rewriting the channel tag as a manifest
|
|
# index — see the repoint step's note.
|
|
echo "reuse: NOTE $IMAGE:$T exists but has no readable fc.revision."
|
|
echo "reuse: NOTE Fine once, while migrating. Every push means the"
|
|
echo "reuse: NOTE tag is being index-wrapped and reuse is dead."
|
|
fi
|
|
|
|
# FORCE is checked here rather than in the build step's `if:`, so
|
|
# that one decision drives everything downstream. The repoint step
|
|
# keys off `hit` too, and a force that bypassed only the build would
|
|
# leave the two disagreeing about what just happened.
|
|
if [ "${FORCE:-false}" = "true" ]; then
|
|
echo "hit=false" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: force_build set — building regardless"
|
|
elif [ "${EVENT:-}" = "schedule" ]; then
|
|
echo "hit=false" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: scheduled base refresh — building regardless"
|
|
elif [ -n "$PUBLISHED" ] && [ "$PUBLISHED" = "$DERIVED" ]; then
|
|
echo "hit=true" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: already published — skipping the build"
|
|
else
|
|
echo "hit=false" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: not published — building"
|
|
fi
|
|
|
|
- name: Build and push ml image
|
|
if: steps.reuse.outputs.hit != 'true'
|
|
# Read by buildx out of the ENVIRONMENT, not passed as a build-arg —
|
|
# it normalises the image config's `created` field and the history
|
|
# timestamps rather than being consumed by the Dockerfile. See #3265
|
|
# and the reuse step's `epoch` output.
|
|
env:
|
|
SOURCE_DATE_EPOCH: ${{ steps.reuse.outputs.epoch }}
|
|
uses: docker/build-push-action@v5
|
|
with:
|
|
context: .
|
|
file: Dockerfile.ml
|
|
push: true
|
|
# Re-resolve the FROM references against the registry instead of
|
|
# trusting whatever digest the cache was built against. This is the
|
|
# whole mechanism of the scheduled refresh (#3154): if the base tag
|
|
# moved, the FROM layer's cache key changes, every layer above it
|
|
# invalidates, and the image genuinely rebuilds.
|
|
#
|
|
# MEASURED on the first real fire, run 4934 (#3265): when the base
|
|
# did NOT move, the build was ~13s with every content step CACHED —
|
|
# and the channel tag STILL got a new manifest digest, because
|
|
# buildkit stamps a fresh image config per run and republishes the
|
|
# identical layers under it. All three images moved that way on
|
|
# 2026-08-30 with nothing whatsoever changed in them.
|
|
#
|
|
# SOURCE_DATE_EPOCH (below) is the fix: pinned to the commit the
|
|
# content came from, the config is byte-identical across runs, so
|
|
# the manifest digest is too and the push is a registry no-op. A
|
|
# digest change means the content changed again, which is the only
|
|
# thing a digest is any use for.
|
|
#
|
|
# What `pull` does NOT catch either: a Debian package update inside
|
|
# the `apt-get install` layer while the base tag itself stands
|
|
# still. The official python/cuda images rebuild with those updates
|
|
# baked in, so this is a lag rather than a hole; closing it needs
|
|
# `no-cache: true`, which is a much larger version of the same
|
|
# churn #3265 is about.
|
|
#
|
|
# Only on the schedule. An ordinary push wants the cached base.
|
|
pull: ${{ github.event_name == 'schedule' }}
|
|
# ONE tag, the channel's. Every other tag is written by the step
|
|
# below, registry-side. buildx here pushes the first tag to the
|
|
# registry and then re-pushes the rest through the DOCKER driver,
|
|
# out of a local image store a registry-direct build never filled —
|
|
# #3190, which cost `main` its :c-<sha> on 2026-08-29 while :latest
|
|
# published perfectly well.
|
|
tags: ${{ steps.reuse.outputs.channel_ref }}
|
|
# The reuse key. Read back off the channel tag on the next push to
|
|
# decide whether that push needs to build at all, so this is not
|
|
# decoration — an unstamped image is one that will always rebuild.
|
|
labels: |
|
|
fc.revision=${{ steps.reuse.outputs.revision }}
|
|
# LOAD-BEARING, not a preference. On the default docker driver these
|
|
# were no-ops; on the docker-container driver above,
|
|
# build-push-action@v5 defaults provenance to TRUE when pushing.
|
|
# Provenance attaches an attestation manifest, which makes the pushed
|
|
# tag a manifest INDEX — and `.Image.Config.Labels` does not resolve
|
|
# through an index.
|
|
#
|
|
# The label directly above IS the reuse key. Wrap the channel tag in
|
|
# an index and the next push reads fc.revision=<none>, misses, and
|
|
# rebuilds. Then so does the one after that, forever. Nothing fails,
|
|
# nothing goes red, and the only symptom is the bill. That is #3183
|
|
# arriving through a different door, and note #3127 §4 records the
|
|
# same shape for `platforms:`.
|
|
provenance: false
|
|
sbom: false
|
|
# The ONLY cache this driver can have. `docker-container` gets a
|
|
# FRESH buildkit instance per job, so unlike the default docker
|
|
# driver it has no local layer store to fall back on — measured on
|
|
# run 4896, the first builds after the driver change: web 3m44s
|
|
# (was 2m23s), ml 3m49s (was 3m20s), agent 11m12s (was 9m26s). The
|
|
# driver change ALONE is a regression; this is the other half of it.
|
|
#
|
|
# mode=max so intermediate stages cache too. web's frontend-builder
|
|
# stage and the agent's two ~150s pip layers are the whole cost, and
|
|
# they are exactly what a min-mode cache would drop.
|
|
#
|
|
# A `:buildcache` tag is NOT the withdrawn tag scheme coming back.
|
|
# Rule 145 narrowed against names NOTHING reads; this one is read by
|
|
# every build that runs, is one moving ref per image rather than one
|
|
# per build, holds cache blobs rather than a shippable artifact, and
|
|
# is overwritten in place rather than accumulating. It is closer to
|
|
# :dev than to the :2026.8.28 tags milestone 318 deleted. (#3114.)
|
|
cache-from: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator-ml:buildcache
|
|
cache-to: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator-ml:buildcache,mode=max
|
|
|
|
# Every tag but the channel's own is written HERE, registry-side,
|
|
# whether or not a build ran. Each -t becomes another reference to the
|
|
# SAME manifest the channel tag holds, so :c-<sha> is byte-identical to
|
|
# what is published rather than a lookalike rebuild.
|
|
#
|
|
# Owning the build path too is #3190's fix, not a tidy-up:
|
|
#
|
|
# #27 pushing …/fabledcurator:latest DONE 15.8s
|
|
# #28 pushing …/fabledcurator:c-0e15c44 with docker
|
|
# #28 ERROR: tag does not exist: …:c-0e15c44
|
|
#
|
|
# Intermittent — build-ml made the identical two-tag push seconds later
|
|
# and succeeded — and worse than it looks. `:latest` had already
|
|
# published, so production was correct while the immutable rollback tag
|
|
# rule 145 requires of every main push simply did not exist. Nothing but
|
|
# the red job would ever have noticed: a missing :c-<sha> has no
|
|
# consumer that fails, so it surfaces when somebody needs to roll back.
|
|
#
|
|
# `imagetools create` is a registry-side manifest copy — no layer
|
|
# transfer, no local daemon, nothing that can be absent. The reuse case
|
|
# has always gone this way, so this puts the build case on the code that
|
|
# was already proven rather than on a second path.
|
|
#
|
|
# Running on every path also keeps family rule 146 true: a rolling
|
|
# channel refreshes itself, so skipping a build must never leave :dev or
|
|
# :latest pointing at something older than the commit just pushed.
|
|
#
|
|
# The cost, accepted knowingly: `imagetools create` wraps its source in
|
|
# an index, so :c-<sha> becomes an index and fc.revision does not
|
|
# resolve through it. Nothing reads that label off :c-<sha> — the reuse
|
|
# check only ever inspects the CHANNEL tag — and the index names the
|
|
# same manifest, so a pull is byte-identical. The reuse path already
|
|
# produced :c-<sha> this way; this only makes it uniform.
|
|
- name: Write the remaining tags from the published image
|
|
env:
|
|
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator-ml
|
|
SOURCE: ${{ steps.reuse.outputs.channel_ref }}
|
|
TAGS: ${{ steps.tag.outputs.tags }}
|
|
run: |
|
|
set -euf
|
|
# The source tag is EXCLUDED from the targets, and that is load-
|
|
# bearing rather than an optimisation.
|
|
#
|
|
# `imagetools create` wraps the source manifest in an INDEX. Point it
|
|
# at the channel tag with that same tag as a target and the tag stops
|
|
# being a plain image — after which `.Image.Config.Labels` no longer
|
|
# resolves through it and the fc.revision label reads as absent. The
|
|
# next push then misses and rebuilds, so reuse worked exactly once
|
|
# and every subsequent push paid full price. Observed on run 4751:
|
|
# ml:dev reported fc.revision=<none> one push after run 4749 had read
|
|
# a7e626a67a79 off it. Nothing failed; the savings just evaporated.
|
|
#
|
|
# Excluding the source means the channel tag is only ever written
|
|
# by a real build, so it stays a plain image and stays readable.
|
|
# On dev that leaves nothing to do either way: the build pushed :dev
|
|
# itself, or the hit established it was already right. On main it
|
|
# leaves :c-<sha>, which rule 145 requires of every main push whether
|
|
# or not a build ran.
|
|
#
|
|
# steps.tag emits ONE comma-separated list; imagetools wants a -t per
|
|
# ref. (That list used to feed docker/build-push-action directly —
|
|
# which is exactly what #3190 made unsafe.)
|
|
ARGS=""
|
|
IFS=,
|
|
for t in $TAGS; do
|
|
[ "$t" = "$SOURCE" ] && continue
|
|
ARGS="$ARGS -t $t"
|
|
done
|
|
unset IFS
|
|
if [ -z "$ARGS" ]; then
|
|
echo "repoint: $SOURCE is the only tag for this channel and"
|
|
echo "repoint: already holds this revision — nothing to write."
|
|
exit 0
|
|
fi
|
|
# shellcheck disable=SC2086
|
|
docker buildx imagetools create $ARGS "$SOURCE"
|
|
echo "repointed from $SOURCE:$ARGS"
|
|
|
|
# The desktop GPU agent (#114) — published so the operator pulls + runs it on
|
|
# the GPU machine instead of building locally. Independent of web/ml (its own
|
|
# CUDA + onnxruntime-gpu image, context = agent/). Same tag cadence.
|
|
build-agent:
|
|
runs-on: python-ci
|
|
container:
|
|
image: git.fabledsword.com/bvandeusen/ci-python:3.14
|
|
steps:
|
|
- uses: actions/checkout@v4
|
|
with:
|
|
# Not the triggering ref — see the `env:` block at the top. On a
|
|
# scheduled refresh this is `main`; on everything else it is the ref
|
|
# that fired, so this is a no-op on every ordinary path.
|
|
ref: ${{ env.BUILD_REF }}
|
|
# Full history: this job derives its artifact's version from the
|
|
# commit its shipped files last changed in (milestone 313). A
|
|
# depth-1 clone cannot see that commit — it either derives a wrong,
|
|
# too-low value or finds nothing at all, and neither is a failure
|
|
# the build would otherwise notice.
|
|
fetch-depth: 0
|
|
|
|
# See sign-extension's copy for why this guard exists.
|
|
- name: Guard — a scheduled run must have checked out main
|
|
if: github.event_name == 'schedule'
|
|
run: |
|
|
set -eu
|
|
BRANCH=$(git rev-parse --abbrev-ref HEAD)
|
|
echo "schedule: HEAD is $BRANCH ($(git rev-parse --short HEAD))"
|
|
if [ "$BRANCH" != "main" ]; then
|
|
echo "schedule: expected main, got '$BRANCH'." >&2
|
|
echo "schedule: BUILD_REF was not honoured by the runner." >&2
|
|
echo "schedule: refusing to publish a channel tag from it." >&2
|
|
exit 1
|
|
fi
|
|
|
|
# --- derived values, one line (milestone 313) ------------------------
|
|
# These stopped being shadow output at step 3. `revision` decides
|
|
# whether the build below runs at all and `version` is what the image
|
|
# reports about itself; the load-bearing steps each print only the one
|
|
# they use, so this is the only place the pair appears together. When a
|
|
# build is skipped, this is the line that says what the commit derived.
|
|
#
|
|
# Still diagnostic, so it still must not fail the build — no `set -e`,
|
|
# and every derivation falls back to UNAVAILABLE. A broken echo must
|
|
# never be the reason an image does not ship.
|
|
#
|
|
# What it should say:
|
|
# * a push touching only agent/ moves the agent and leaves web and ml
|
|
# STILL. If web moves, its path set is too wide.
|
|
# * a push touching only docs moves nothing.
|
|
# * a push touching the extension moves the extension AND web, since
|
|
# web bakes in the XPI. If web does not move, its set is too narrow:
|
|
# the reuse check hits, and the channel serves a web image bundling
|
|
# the PREVIOUS XPI while the freshly signed one is orphaned (#3156).
|
|
# * dev and main derive the same values for the same source.
|
|
- name: Report the derived artifact version
|
|
run: |
|
|
set -u
|
|
A=agent
|
|
V=$(sh scripts/artifacts.sh version "$A" 2>&1 || echo UNAVAILABLE)
|
|
R=$(sh scripts/artifacts.sh revision "$A" 2>&1 || echo UNAVAILABLE)
|
|
echo "derived: artifact=$A version=$V revision=$R sha=$GITHUB_SHA"
|
|
|
|
- name: Determine tag
|
|
id: tag
|
|
run: |
|
|
SHORT_SHA=$(printf '%s' "$GITHUB_SHA" | cut -c1-7)
|
|
# Mirrors build-web's tag list and its schedule handling; see
|
|
# the comments there.
|
|
if [ "${GITHUB_EVENT_NAME:-}" = "schedule" ]; then
|
|
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator-agent:latest" >> "$GITHUB_OUTPUT"
|
|
echo "channel=main" >> "$GITHUB_OUTPUT"
|
|
elif [ "${GITHUB_REF##*/}" = "main" ]; then
|
|
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator-agent:latest,git.fabledsword.com/bvandeusen/fabledcurator-agent:c-${SHORT_SHA}" >> "$GITHUB_OUTPUT"
|
|
echo "channel=main" >> "$GITHUB_OUTPUT"
|
|
else
|
|
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator-agent:dev" >> "$GITHUB_OUTPUT"
|
|
echo "channel=dev" >> "$GITHUB_OUTPUT"
|
|
fi
|
|
|
|
# Shell step rather than docker/login-action — see build-web's note on
|
|
# the shared action-cache race (#3118).
|
|
- name: Login to Forgejo registry
|
|
env:
|
|
TOKEN: ${{ secrets.RELEASE_TOKEN }}
|
|
ACTOR: ${{ github.actor }}
|
|
run: echo "$TOKEN" | docker login git.fabledsword.com -u "$ACTOR" --password-stdin
|
|
|
|
# A REAL buildx builder, not the default `docker` driver (#3114, #3190).
|
|
#
|
|
# The default driver builds through the local dockerd. It cannot export a
|
|
# registry cache at all — which is why the agent rebuilds a ~6.3 GB CUDA
|
|
# + torch image from scratch whenever the runner's local cache is cold,
|
|
# measured at 9m26s against 7s warm. It is also #3190's leading suspect:
|
|
# after a registry-direct push it resolves image metadata against a local
|
|
# store the push never filled, and reports `No such image` on an image
|
|
# that published perfectly well three seconds earlier.
|
|
#
|
|
# These jobs run INSIDE a container against a mounted docker socket, so
|
|
# the buildkit container this starts is a SIBLING of the job container,
|
|
# not a child. That works over the socket mount; it had never been tried
|
|
# here before milestone 326 step 1.
|
|
- name: Set up buildx
|
|
uses: docker/setup-buildx-action@v3
|
|
|
|
# --- reuse-if-published (milestone 313, step 4) ----------------------
|
|
# Does the image the channel tag already points at carry THIS commit's
|
|
# revision? If so the bytes this job would produce are already published
|
|
# and the build is pure waste: the remaining tags get repointed at that
|
|
# existing manifest instead, registry-side, in seconds.
|
|
#
|
|
# Keyed on an `fc.revision` LABEL rather than on a tag of its own
|
|
# (milestone 318 step 3). A tag would be a name minted per build that one
|
|
# thing reads — what rule 145 narrowed against — and would be prunable
|
|
# under the registry's keep_pattern (#3157), silently expiring the cache.
|
|
# A label rides inside a tag that has to exist anyway.
|
|
#
|
|
# An image with no such label reads as a miss and rebuilds. That is the
|
|
# migration, not a fault: labels cannot be backfilled, since the reuse
|
|
# path copies a manifest and config labels are not manifest annotations.
|
|
# Each artifact pays one rebuild, once.
|
|
#
|
|
# This is what stops a push that touched only `agent/` from rebuilding
|
|
# web and ml, and a merge to main from rebuilding what dev already built.
|
|
#
|
|
# The failure direction is deliberate. An inspect that errors for ANY
|
|
# reason — network, auth, a registry hiccup — reads as a miss and the
|
|
# build runs. Only a genuine 200 skips one, so there is no path here
|
|
# that skips a build that was actually needed; the worst case is paying
|
|
# for a build we could have avoided.
|
|
#
|
|
# BASE-IMAGE FRESHNESS: an artifact whose source stops moving stops
|
|
# picking up base-image updates. Milestone 318 removed the argument this
|
|
# used to need rather than answering it — with no version tags there is
|
|
# no immutable name a refresh could contradict, and rule 145 already
|
|
# allows a rebuild with different contents to republish a MOVING tag.
|
|
# So a refresh is just a build. A scheduled channel-only one is tracked
|
|
# separately (#3154); it does not belong in the push path.
|
|
- name: Is this content already published?
|
|
id: reuse
|
|
env:
|
|
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator-agent
|
|
CHANNEL: ${{ steps.tag.outputs.channel }}
|
|
# Empty on a push; the string "true" only from a workflow_dispatch
|
|
# that asked for it. `github.event.inputs` rather than the `inputs`
|
|
# context — release.yml already uses that form, and it is the one
|
|
# this runner is known to evaluate. Read through env rather than
|
|
# interpolated into the run block, same rule as release.yml's TAG.
|
|
FORCE: ${{ github.event.inputs.force_build }}
|
|
# A scheduled refresh has to bypass reuse by construction: it
|
|
# rebuilds the SAME source, so fc.revision always matches and the
|
|
# check would skip every refresh there has ever been.
|
|
EVENT: ${{ github.event_name }}
|
|
run: |
|
|
set -eu
|
|
DERIVED=$(sh scripts/artifacts.sh revision agent)
|
|
echo "revision=$DERIVED" >> "$GITHUB_OUTPUT"
|
|
|
|
# The build clock, pinned to the same commit (#3265). Without it
|
|
# buildkit stamps the image config with the wall clock of the build,
|
|
# so identical layers republish under a new config blob and the
|
|
# channel tag gets a new manifest digest for no reason. Derived from
|
|
# `newest()` like revision and version, so all three name one commit
|
|
# and cannot drift apart.
|
|
echo "epoch=$(sh scripts/artifacts.sh epoch agent)" >> "$GITHUB_OUTPUT"
|
|
|
|
# The moving tag for this channel. Which tag we ask IS the channel —
|
|
# that is why the revision needs no -main/-dev qualifier any more.
|
|
if [ "$CHANNEL" = "main" ]; then T=latest; else T=dev; fi
|
|
echo "channel_ref=$IMAGE:$T" >> "$GITHUB_OUTPUT"
|
|
|
|
# Compare VALUES, never exit codes. Measured on buildx v0.36.1
|
|
# (run 4732): a missing key returns an empty string and exits 0, so
|
|
# branching on the exit code would read "no label yet" as success.
|
|
# An unreachable tag also lands here as empty via the `|| echo`.
|
|
# Empty never equals a 12-char revision, so every uncertain case
|
|
# falls through to a build — the safe direction, with no special
|
|
# casing for it.
|
|
#
|
|
# Read the SPECIFIC key. The map also carries whatever the base image
|
|
# set, and `org.opencontainers.image.version` sits right beside ours
|
|
# looking like a plausible answer (it reads 24.04 on the agent).
|
|
PUBLISHED=$(docker buildx imagetools inspect "$IMAGE:$T" \
|
|
--format '{{ index .Image.Config.Labels "fc.revision" }}' \
|
|
2>/dev/null || echo "")
|
|
echo "reuse: $IMAGE:$T carries fc.revision=${PUBLISHED:-<none>}; derived=$DERIVED"
|
|
if [ -z "$PUBLISHED" ] && docker buildx imagetools inspect "$IMAGE:$T" >/dev/null 2>&1; then
|
|
# The tag resolves but carries no readable label. Expected exactly
|
|
# once per artifact, during the migration onto labels. If it recurs
|
|
# every push, something is rewriting the channel tag as a manifest
|
|
# index — see the repoint step's note.
|
|
echo "reuse: NOTE $IMAGE:$T exists but has no readable fc.revision."
|
|
echo "reuse: NOTE Fine once, while migrating. Every push means the"
|
|
echo "reuse: NOTE tag is being index-wrapped and reuse is dead."
|
|
fi
|
|
|
|
# FORCE is checked here rather than in the build step's `if:`, so
|
|
# that one decision drives everything downstream. The repoint step
|
|
# keys off `hit` too, and a force that bypassed only the build would
|
|
# leave the two disagreeing about what just happened.
|
|
if [ "${FORCE:-false}" = "true" ]; then
|
|
echo "hit=false" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: force_build set — building regardless"
|
|
elif [ "${EVENT:-}" = "schedule" ]; then
|
|
echo "hit=false" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: scheduled base refresh — building regardless"
|
|
elif [ -n "$PUBLISHED" ] && [ "$PUBLISHED" = "$DERIVED" ]; then
|
|
echo "hit=true" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: already published — skipping the build"
|
|
else
|
|
echo "hit=false" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: not published — building"
|
|
fi
|
|
|
|
- name: Build and push agent image
|
|
if: steps.reuse.outputs.hit != 'true'
|
|
# Read by buildx out of the ENVIRONMENT, not passed as a build-arg —
|
|
# it normalises the image config's `created` field and the history
|
|
# timestamps rather than being consumed by the Dockerfile. See #3265
|
|
# and the reuse step's `epoch` output.
|
|
env:
|
|
SOURCE_DATE_EPOCH: ${{ steps.reuse.outputs.epoch }}
|
|
uses: docker/build-push-action@v5
|
|
with:
|
|
context: agent
|
|
file: agent/Dockerfile
|
|
push: true
|
|
# Re-resolve the FROM references against the registry instead of
|
|
# trusting whatever digest the cache was built against. This is the
|
|
# whole mechanism of the scheduled refresh (#3154): if the base tag
|
|
# moved, the FROM layer's cache key changes, every layer above it
|
|
# invalidates, and the image genuinely rebuilds.
|
|
#
|
|
# MEASURED on the first real fire, run 4934 (#3265): when the base
|
|
# did NOT move, the build was ~13s with every content step CACHED —
|
|
# and the channel tag STILL got a new manifest digest, because
|
|
# buildkit stamps a fresh image config per run and republishes the
|
|
# identical layers under it. All three images moved that way on
|
|
# 2026-08-30 with nothing whatsoever changed in them.
|
|
#
|
|
# SOURCE_DATE_EPOCH (below) is the fix: pinned to the commit the
|
|
# content came from, the config is byte-identical across runs, so
|
|
# the manifest digest is too and the push is a registry no-op. A
|
|
# digest change means the content changed again, which is the only
|
|
# thing a digest is any use for.
|
|
#
|
|
# What `pull` does NOT catch either: a Debian package update inside
|
|
# the `apt-get install` layer while the base tag itself stands
|
|
# still. The official python/cuda images rebuild with those updates
|
|
# baked in, so this is a lag rather than a hole; closing it needs
|
|
# `no-cache: true`, which is a much larger version of the same
|
|
# churn #3265 is about.
|
|
#
|
|
# Only on the schedule. An ordinary push wants the cached base.
|
|
pull: ${{ github.event_name == 'schedule' }}
|
|
# ONE tag, the channel's. Every other tag is written by the step
|
|
# below, registry-side. buildx here pushes the first tag to the
|
|
# registry and then re-pushes the rest through the DOCKER driver,
|
|
# out of a local image store a registry-direct build never filled —
|
|
# #3190, which cost `main` its :c-<sha> on 2026-08-29 while :latest
|
|
# published perfectly well.
|
|
tags: ${{ steps.reuse.outputs.channel_ref }}
|
|
# The reuse key. Read back off the channel tag on the next push to
|
|
# decide whether that push needs to build at all, so this is not
|
|
# decoration — an unstamped image is one that will always rebuild.
|
|
labels: |
|
|
fc.revision=${{ steps.reuse.outputs.revision }}
|
|
# LOAD-BEARING, not a preference. On the default docker driver these
|
|
# were no-ops; on the docker-container driver above,
|
|
# build-push-action@v5 defaults provenance to TRUE when pushing.
|
|
# Provenance attaches an attestation manifest, which makes the pushed
|
|
# tag a manifest INDEX — and `.Image.Config.Labels` does not resolve
|
|
# through an index.
|
|
#
|
|
# The label directly above IS the reuse key. Wrap the channel tag in
|
|
# an index and the next push reads fc.revision=<none>, misses, and
|
|
# rebuilds. Then so does the one after that, forever. Nothing fails,
|
|
# nothing goes red, and the only symptom is the bill. That is #3183
|
|
# arriving through a different door, and note #3127 §4 records the
|
|
# same shape for `platforms:`.
|
|
provenance: false
|
|
sbom: false
|
|
# The ONLY cache this driver can have. `docker-container` gets a
|
|
# FRESH buildkit instance per job, so unlike the default docker
|
|
# driver it has no local layer store to fall back on — measured on
|
|
# run 4896, the first builds after the driver change: web 3m44s
|
|
# (was 2m23s), ml 3m49s (was 3m20s), agent 11m12s (was 9m26s). The
|
|
# driver change ALONE is a regression; this is the other half of it.
|
|
#
|
|
# mode=max so intermediate stages cache too. web's frontend-builder
|
|
# stage and the agent's two ~150s pip layers are the whole cost, and
|
|
# they are exactly what a min-mode cache would drop.
|
|
#
|
|
# A `:buildcache` tag is NOT the withdrawn tag scheme coming back.
|
|
# Rule 145 narrowed against names NOTHING reads; this one is read by
|
|
# every build that runs, is one moving ref per image rather than one
|
|
# per build, holds cache blobs rather than a shippable artifact, and
|
|
# is overwritten in place rather than accumulating. It is closer to
|
|
# :dev than to the :2026.8.28 tags milestone 318 deleted. (#3114.)
|
|
cache-from: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator-agent:buildcache
|
|
cache-to: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator-agent:buildcache,mode=max
|
|
|
|
# Every tag but the channel's own is written HERE, registry-side,
|
|
# whether or not a build ran. Each -t becomes another reference to the
|
|
# SAME manifest the channel tag holds, so :c-<sha> is byte-identical to
|
|
# what is published rather than a lookalike rebuild.
|
|
#
|
|
# Owning the build path too is #3190's fix, not a tidy-up:
|
|
#
|
|
# #27 pushing …/fabledcurator:latest DONE 15.8s
|
|
# #28 pushing …/fabledcurator:c-0e15c44 with docker
|
|
# #28 ERROR: tag does not exist: …:c-0e15c44
|
|
#
|
|
# Intermittent — build-ml made the identical two-tag push seconds later
|
|
# and succeeded — and worse than it looks. `:latest` had already
|
|
# published, so production was correct while the immutable rollback tag
|
|
# rule 145 requires of every main push simply did not exist. Nothing but
|
|
# the red job would ever have noticed: a missing :c-<sha> has no
|
|
# consumer that fails, so it surfaces when somebody needs to roll back.
|
|
#
|
|
# `imagetools create` is a registry-side manifest copy — no layer
|
|
# transfer, no local daemon, nothing that can be absent. The reuse case
|
|
# has always gone this way, so this puts the build case on the code that
|
|
# was already proven rather than on a second path.
|
|
#
|
|
# Running on every path also keeps family rule 146 true: a rolling
|
|
# channel refreshes itself, so skipping a build must never leave :dev or
|
|
# :latest pointing at something older than the commit just pushed.
|
|
#
|
|
# The cost, accepted knowingly: `imagetools create` wraps its source in
|
|
# an index, so :c-<sha> becomes an index and fc.revision does not
|
|
# resolve through it. Nothing reads that label off :c-<sha> — the reuse
|
|
# check only ever inspects the CHANNEL tag — and the index names the
|
|
# same manifest, so a pull is byte-identical. The reuse path already
|
|
# produced :c-<sha> this way; this only makes it uniform.
|
|
- name: Write the remaining tags from the published image
|
|
env:
|
|
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator-agent
|
|
SOURCE: ${{ steps.reuse.outputs.channel_ref }}
|
|
TAGS: ${{ steps.tag.outputs.tags }}
|
|
run: |
|
|
set -euf
|
|
# The source tag is EXCLUDED from the targets, and that is load-
|
|
# bearing rather than an optimisation.
|
|
#
|
|
# `imagetools create` wraps the source manifest in an INDEX. Point it
|
|
# at the channel tag with that same tag as a target and the tag stops
|
|
# being a plain image — after which `.Image.Config.Labels` no longer
|
|
# resolves through it and the fc.revision label reads as absent. The
|
|
# next push then misses and rebuilds, so reuse worked exactly once
|
|
# and every subsequent push paid full price. Observed on run 4751:
|
|
# ml:dev reported fc.revision=<none> one push after run 4749 had read
|
|
# a7e626a67a79 off it. Nothing failed; the savings just evaporated.
|
|
#
|
|
# Excluding the source means the channel tag is only ever written
|
|
# by a real build, so it stays a plain image and stays readable.
|
|
# On dev that leaves nothing to do either way: the build pushed :dev
|
|
# itself, or the hit established it was already right. On main it
|
|
# leaves :c-<sha>, which rule 145 requires of every main push whether
|
|
# or not a build ran.
|
|
#
|
|
# steps.tag emits ONE comma-separated list; imagetools wants a -t per
|
|
# ref. (That list used to feed docker/build-push-action directly —
|
|
# which is exactly what #3190 made unsafe.)
|
|
ARGS=""
|
|
IFS=,
|
|
for t in $TAGS; do
|
|
[ "$t" = "$SOURCE" ] && continue
|
|
ARGS="$ARGS -t $t"
|
|
done
|
|
unset IFS
|
|
if [ -z "$ARGS" ]; then
|
|
echo "repoint: $SOURCE is the only tag for this channel and"
|
|
echo "repoint: already holds this revision — nothing to write."
|
|
exit 0
|
|
fi
|
|
# shellcheck disable=SC2086
|
|
docker buildx imagetools create $ARGS "$SOURCE"
|
|
echo "repointed from $SOURCE:$ARGS"
|