CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / build-agent (push) Successful in 7s
Build images / build-ml (push) Successful in 7s
Build images / build-web (push) Successful in 8s
CI / frontend-build (push) Successful in 20s
extension / lint (push) Successful in 18s
CI / backend-lint-and-test (push) Successful in 37s
CI / integration (push) Successful in 1m45s
extension / lint (pull_request) Successful in 24s
Milestone 362, enabling step 2's verification and everything after it. The weekly refresh was testable once a week. That is not a cadence anything can be developed against, and milestone 362's whole point is a gate — which has to be watched rejecting something before anyone can believe it is wired up. So `refresh` joins `force_build` as a dispatch input, on the same reasoning that added that one (#3252: confirm #3190 was gone rather than wait for it to recur). Adding it meant confronting that "is this a refresh?" was asked in five places and spelled five ways: `github.event_name == 'schedule'` in an `if:`, `$GITHUB_EVENT_NAME` in one shell, an `EVENT:` env passed into another, and a bare expression on `pull:`. Five spellings of one fact is how half of them come to disagree once somebody adds a sixth trigger — which is precisely what this commit is. So it is derived once at the top, next to BUILD_REF, which already exists for exactly this reason on exactly this question. String comparison, not boolean: Forgejo delivers dispatch inputs as strings, so `inputs.refresh` is 'true'/'false' and `&&` on it would read the string 'false' as truthy. **A constraint this makes visible, which pre-dates it.** A refresh checks out `main` (BUILD_REF) while running the workflow definition from the branch that triggered it — the cron registers from the default branch. So dev's workflow builds main's source, and dev's workflow cannot depend on anything main's tree does not have yet. It does now: the reuse step calls `artifacts.sh epoch`, which lands on main with this batch. Until then a refresh dispatch fails loudly at that call, which is the right failure — the alternative is tolerating a missing epoch and silently rebuilding #3265 into every refresh. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
2048 lines
113 KiB
YAML
2048 lines
113 KiB
YAML
name: Build images
|
|
|
|
on:
|
|
push:
|
|
# `:dev` builds were dropped 2026-05-26 to save a docker build per dev
|
|
# push, on the reasoning that "operator tests from `:latest` after
|
|
# merge-to-main". Restored 2026-08-27: that is testing by shipping, and
|
|
# family rules 146/147 now name it directly — `main` IS production, and a
|
|
# channel that can only be refreshed by shipping is not a channel. The
|
|
# pressure to merge in order to try something does not come from
|
|
# carelessness; it comes from `:dev` being unable to carry the build.
|
|
#
|
|
# All three images build on dev, deliberately: a `:dev` web image paired
|
|
# with a stale `:dev` ml or agent is a worse trap than no dev channel at
|
|
# all, since the mismatch only shows up as a runtime failure.
|
|
branches: [main, dev]
|
|
#
|
|
# NO tag trigger (milestone 318 step 2). A `v*` tag names a commit `main`
|
|
# already built and published; rebuilding it produces the same source under
|
|
# the same names and RE-PUSHES `:c-<sha>`, which rule 145 forbids even when
|
|
# the bytes match — image configs carry timestamps, so "same source" does
|
|
# not mean "same manifest". The release build was publishing nothing new
|
|
# and violating an immutability rule to do it.
|
|
#
|
|
# Releases still happen (rule 148, on explicit request per rule 2). They
|
|
# produce a changelog, not an image.
|
|
|
|
# The escape hatch for the one thing skip-if-exists makes untestable: a
|
|
# build that WOULD be skipped. `agent/` has not changed since 2026-07-17, so
|
|
# every push since has correctly declined to build it — which also means the
|
|
# agent build path has not run in six weeks and cannot be exercised on
|
|
# demand. #3190 lives on exactly that path.
|
|
#
|
|
# Editing build.yml does not force one either, and that is deliberate: the
|
|
# workflow is not shipped bytes, so it is in no artifact's path set. Putting
|
|
# it in one would re-version every artifact for a comment change.
|
|
#
|
|
# ONE input, not one per artifact. Forcing all three is cheap once the
|
|
# registry cache is warm (#3114), and three booleans is an interface nobody
|
|
# remembers the meaning of.
|
|
workflow_dispatch:
|
|
inputs:
|
|
force_build:
|
|
description: 'Rebuild every image even if the published revision matches'
|
|
type: boolean
|
|
default: false
|
|
refresh:
|
|
description: 'Behave as the weekly base refresh: build main against fresh bases, publish through the candidate tag'
|
|
type: boolean
|
|
default: false
|
|
|
|
# The base-image refresh (milestone 326 step 4, #3154).
|
|
#
|
|
# Skip-if-exists is keyed on OUR source, so an artifact whose source stops
|
|
# moving stops picking up base-image updates. `agent/` last changed
|
|
# 2026-07-17; every push since has correctly declined to rebuild it, which
|
|
# also means it will serve that day's `nvidia/cuda` layers forever. Nothing
|
|
# is wrong until it has been unchanged for months, which is precisely why
|
|
# this is a calendar trigger and not a condition on the push path.
|
|
#
|
|
# Weekly, Sunday 06:00 UTC. Away from CI-runner's Monday security sweep so
|
|
# the two are never diagnosing each other, and on the quietest day so a
|
|
# surprise rebuild is not competing with a push.
|
|
schedule:
|
|
- cron: '0 6 * * 0'
|
|
|
|
# Which branch a run BUILDS, as opposed to which one triggered it.
|
|
#
|
|
# They are the same thing on every trigger but `schedule`. Forgejo registers a
|
|
# cron from the DEFAULT branch — `dev` here — so a scheduled run arrives with
|
|
# `github.ref` pointing at dev, and a refresh that rebuilt `:dev` would be
|
|
# refreshing the one channel that gets rebuilt constantly anyway. Production is
|
|
# `main` (rule 147), and `:latest` is the tag that goes stale.
|
|
#
|
|
# So the ref is decided once, here, and every checkout in the file takes it.
|
|
# Deriving it per job invites the two halves to disagree: sign-extension would
|
|
# derive dev's extension version while build-web bundled main's, and the
|
|
# release download would 404 on a version that exists perfectly well.
|
|
# IS THIS A BASE REFRESH? Asked in five places and previously spelled five
|
|
# ways — `github.event_name == 'schedule'` in an `if:`, `$GITHUB_EVENT_NAME` in
|
|
# one shell, an `EVENT:` env passed into another, and a bare expression on
|
|
# `pull:`. Five spellings of one fact is how half of them come to disagree
|
|
# after somebody adds a sixth trigger.
|
|
#
|
|
# The `refresh` dispatch input is here so this path can be EXERCISED. A weekly
|
|
# cron is otherwise testable once a week, which is not a cadence anything can
|
|
# be developed against — the same reason `force_build` exists (#3252, added to
|
|
# confirm #3190 was gone rather than wait for it to recur). It is also what
|
|
# makes the milestone-362 gate verifiable at all: a gate has to be watched
|
|
# rejecting something before anyone can believe it is wired up.
|
|
#
|
|
# Note this is a STRING comparison, not a boolean. Forgejo delivers
|
|
# workflow_dispatch inputs as strings, so `inputs.refresh` is 'true'/'false'
|
|
# and `&&` on it would treat the string 'false' as truthy.
|
|
env:
|
|
IS_REFRESH: ${{ (github.event_name == 'schedule' || github.event.inputs.refresh == 'true') && 'true' || 'false' }}
|
|
BUILD_REF: ${{ (github.event_name == 'schedule' || github.event.inputs.refresh == 'true') && 'main' || github.ref }}
|
|
|
|
# Requires repo secret RELEASE_TOKEN — a Forgejo PAT with scopes:
|
|
# - write:package, read:package (for docker push to git.fabledsword.com)
|
|
# - write:release (for ext-<version> release asset cache)
|
|
# - write:issue (for future issue-management automation)
|
|
# The injected GITHUB_TOKEN cannot be used — it lacks write:package.
|
|
|
|
jobs:
|
|
# Sign-or-fetch-from-cache: signs the extension via AMO if no ext-<version>
|
|
# Forgejo release exists yet, otherwise downloads the cached signed XPI.
|
|
# Result is uploaded as an Actions artifact for build-web to consume.
|
|
#
|
|
# Why this lives in build.yml (not a separate workflow): the image a push
|
|
# publishes MUST carry the XPI. A separate sign workflow racing build.yml
|
|
# leaves that image without one for ~5min (until the commit-back triggers
|
|
# another build). Inline ordering eliminates the race.
|
|
# Cache strategy: Forgejo Release Assets — picked 2026-05-25 over Generic
|
|
# Packages (cleaner API surface) and commit-back-to-side-branch (no extra
|
|
# branch to manage). AMO blocks re-signing the same version (returns 409),
|
|
# so signing is intentionally one-shot per version.
|
|
#
|
|
# BOTH branches sign (milestone 271 step 6, 2026-08-27). Not two signatures:
|
|
# the version is the commit TIME of the newest packaged-extension change, so
|
|
# dev and main derive the SAME number for the same extension source. A dev
|
|
# push that changes the extension signs it; the merge to main then finds the
|
|
# ext-<version> release already there, hits the cache, and bundles the
|
|
# byte-identical XPI into `:latest` with no second AMO call. One signature
|
|
# per extension CHANGE, shared by both channels — that is what makes two
|
|
# channels affordable, and it is why step 4 (derived version) had to land
|
|
# first. Ungating this while the version was still the hand-set 1.0.11 would
|
|
# have hit the existing ext-1.0.11 cache and bundled MAIN's stale XPI into
|
|
# `:dev` — a dev channel confidently serving old code.
|
|
#
|
|
# Unconditional since milestone 318 step 2: main and dev are now the only
|
|
# triggers, so the branch gate that used to exclude tag pushes matched
|
|
# everything. A condition that is always true reads as if some path avoids
|
|
# it, which is worse than no condition.
|
|
sign-extension:
|
|
runs-on: python-ci
|
|
container:
|
|
image: git.fabledsword.com/bvandeusen/ci-python:3.14
|
|
steps:
|
|
- uses: actions/checkout@v4
|
|
with:
|
|
# Not the triggering ref — see the `env:` block at the top. On a
|
|
# scheduled refresh this is `main`; on everything else it is the ref
|
|
# that fired, so this is a no-op on every ordinary path.
|
|
ref: ${{ env.BUILD_REF }}
|
|
# Full history is load-bearing, not a convenience: the version this
|
|
# job signs is derived from the commit TIME of the newest packaged
|
|
# extension change. A depth-1 clone sees one commit and derives a
|
|
# wrong, too-low value rather than failing (ci-requirements.md).
|
|
fetch-depth: 0
|
|
|
|
# BUILD_REF is what makes a scheduled run build `main` rather than the
|
|
# branch its cron fired from — and it is read through the `env` context
|
|
# inside `with:`, which this runner is NOT known to evaluate. If it does
|
|
# not, checkout silently falls back to the triggering ref and the weekly
|
|
# refresh publishes DEV's source to `:latest`, which is production.
|
|
# Every lane would stay green; the first sign of it would be production
|
|
# running code that was never merged.
|
|
#
|
|
# So assert the checkout instead of trusting the expression. A red
|
|
# weekly job is a fine outcome. Shipping dev to production is not.
|
|
#
|
|
# `if:` reads the `github` context, which the runner demonstrably does
|
|
# evaluate — this file already gates steps on it — so the guard cannot
|
|
# be disabled by the same uncertainty it exists to cover.
|
|
- name: Guard — a scheduled run must have checked out main
|
|
if: env.IS_REFRESH == 'true'
|
|
run: |
|
|
set -eu
|
|
BRANCH=$(git rev-parse --abbrev-ref HEAD)
|
|
echo "schedule: HEAD is $BRANCH ($(git rev-parse --short HEAD))"
|
|
if [ "$BRANCH" != "main" ]; then
|
|
echo "schedule: expected main, got '$BRANCH'." >&2
|
|
echo "schedule: BUILD_REF was not honoured by the runner." >&2
|
|
echo "schedule: refusing to publish a channel tag from it." >&2
|
|
exit 1
|
|
fi
|
|
|
|
# The version is DERIVED, not read from the repo (milestone 271 step 4,
|
|
# cut over 2026-08-27). `packaging.sh version` returns `YYYY.M.D.HHMM`
|
|
# UTC — the commit TIME of the newest change to a PACKAGED extension
|
|
# file, per family rule 148/149. Never a commit count, which orders by
|
|
# branch rather than by recency.
|
|
#
|
|
# Unpadded, and only here: AMO's grammar rejects a leading zero, so the
|
|
# extension renders rule 148's numbers without the family's padding
|
|
# (milestone 318 step 8). Same value, one character narrower per segment;
|
|
# ci.yml's extension-version lane checks the string against Mozilla's
|
|
# published regex before this job ever calls AMO.
|
|
#
|
|
# The committed "version" in manifest.json / package.json decides NOTHING
|
|
# — not even a MAJOR.MINOR prefix, which step 8 removed. The stamp step
|
|
# below overwrites it in the working tree before web-ext ever reads it,
|
|
# and it is deliberately NOT committed back: the commit carrying the bump
|
|
# would itself be a change to the extension and would move the version
|
|
# again. The repo holds the source; the build derives the label.
|
|
- name: Derive extension version
|
|
id: extver
|
|
run: |
|
|
set -eu
|
|
VERSION=$(sh extension/scripts/packaging.sh version)
|
|
echo "version=$VERSION" >> "$GITHUB_OUTPUT"
|
|
echo "Derived extension version: $VERSION"
|
|
|
|
# Firefox refuses a downgrade and AMO never releases a burned version,
|
|
# so a version that moves BACKWARDS is unrecoverable: it strands every
|
|
# install that already took the higher one. Two ways it could happen —
|
|
# a checkout without full history (derives too low), or a rewritten
|
|
# history that drops the newest packaged commit.
|
|
#
|
|
# The test is `derived < highest already signed`, strictly. Equality is
|
|
# the ORDINARY case, not a fault: an unchanged extension derives the same
|
|
# version it did last build, which is exactly what lets the ext-<version>
|
|
# cache hit and holds AMO to one call per extension CHANGE. Only moving
|
|
# backwards is a failure, so this runs on every path — cache hit
|
|
# included — rather than only before a sign.
|
|
# --- shadow mode (milestone 313, step 2) -----------------------------
|
|
# Informational ONLY. Nothing reads this and it must never fail the
|
|
# build — no `set -e`, and every derivation falls back to UNAVAILABLE.
|
|
#
|
|
# What to watch across pushes, because this is what step 3 will trust:
|
|
# * a push touching only agent/ moves the agent and leaves web and ml
|
|
# STILL. If web moves, its path set is too wide.
|
|
# * a push touching only docs moves nothing.
|
|
# * a push touching the extension moves the extension AND web, since
|
|
# web bakes in the XPI. If web does not move, its set is too narrow:
|
|
# the reuse check hits, and the channel serves a web image bundling
|
|
# the PREVIOUS XPI while the freshly signed one is orphaned (#3156).
|
|
# * dev and main derive the same values for the same source.
|
|
- name: Shadow — derived artifact version (informational)
|
|
run: |
|
|
set -u
|
|
A=extension
|
|
V=$(sh scripts/artifacts.sh version "$A" 2>&1 || echo UNAVAILABLE)
|
|
R=$(sh scripts/artifacts.sh revision "$A" 2>&1 || echo UNAVAILABLE)
|
|
echo "derived: artifact=$A version=$V revision=$R sha=$GITHUB_SHA"
|
|
|
|
- name: Guard — the derived version must never go backwards
|
|
env:
|
|
TOKEN: ${{ secrets.RELEASE_TOKEN }}
|
|
DERIVED: ${{ steps.extver.outputs.version }}
|
|
run: |
|
|
python3 - <<'PY'
|
|
import json, os, sys, urllib.request
|
|
|
|
API = ("https://git.fabledsword.com/api/v1/repos/"
|
|
"bvandeusen/FabledCurator/releases")
|
|
headers = {"Authorization": "token " + os.environ["TOKEN"]}
|
|
|
|
# Paginated rather than first-page-only: ext-* releases share this
|
|
# list with the v* release tags, so one page would start missing them
|
|
# as those accumulate. The bound FAILS rather than silently scanning
|
|
# part of the list and calling the highest it saw the highest there is.
|
|
tags = []
|
|
for page in range(1, 21):
|
|
req = urllib.request.Request(
|
|
f"{API}?limit=50&page={page}", headers=headers)
|
|
with urllib.request.urlopen(req, timeout=30) as resp:
|
|
batch = json.load(resp)
|
|
if not batch:
|
|
break
|
|
tags += [r.get("tag_name", "") for r in batch]
|
|
else:
|
|
sys.exit("guard: >1000 releases — pagination bound reached")
|
|
|
|
def parse(v):
|
|
try:
|
|
return tuple(int(part) for part in v.split("."))
|
|
except ValueError:
|
|
return None
|
|
|
|
derived_s = os.environ["DERIVED"]
|
|
derived = parse(derived_s)
|
|
if derived is None:
|
|
sys.exit(f"guard: derived version {derived_s!r} is not numeric")
|
|
|
|
signed = sorted(
|
|
(v, t) for t in tags if t.startswith("ext-")
|
|
for v in [parse(t[4:])] if v
|
|
)
|
|
if not signed:
|
|
print("guard: no ext-* release yet — nothing to go backwards from")
|
|
raise SystemExit(0)
|
|
|
|
hi, hi_tag = signed[-1]
|
|
print(f"guard: derived={derived_s} highest already signed={hi_tag}")
|
|
if derived < hi:
|
|
sys.exit(
|
|
f"REFUSING TO SIGN: derived {derived_s} is OLDER than the "
|
|
f"already-signed {hi_tag}. Firefox would reject it as a "
|
|
f"downgrade, and AMO will not release the burned version. "
|
|
f"First thing to check: did this job check out with "
|
|
f"fetch-depth: 0?"
|
|
)
|
|
print("guard: ok")
|
|
PY
|
|
|
|
- name: Check Forgejo release-asset cache
|
|
id: cache
|
|
env:
|
|
TOKEN: ${{ secrets.RELEASE_TOKEN }}
|
|
run: |
|
|
set -eu
|
|
VERSION=${{ steps.extver.outputs.version }}
|
|
STATUS=$(curl -s -o release.json -w "%{http_code}" \
|
|
-H "Authorization: token $TOKEN" \
|
|
"https://git.fabledsword.com/api/v1/repos/bvandeusen/FabledCurator/releases/tags/ext-$VERSION" || echo 000)
|
|
echo "Tag lookup HTTP status: $STATUS"
|
|
# JSON parsing via python (ci-python:3.14 has stdlib json; jq is
|
|
# not in the image and adding it per ci-requirements.md is not
|
|
# warranted for a single consumer — operator-flagged 2026-05-26
|
|
# after a sign job failed with `jq: not found`).
|
|
if [ "$STATUS" = "200" ]; then
|
|
ASSET_ID=$(python3 -c "import json; r=json.load(open('release.json')); xpis=[a for a in r.get('assets', []) if a.get('name','').endswith('.xpi')]; print(xpis[0]['id'] if xpis else '')")
|
|
if [ -n "$ASSET_ID" ]; then
|
|
echo "cached=true" >> "$GITHUB_OUTPUT"
|
|
echo "asset_id=$ASSET_ID" >> "$GITHUB_OUTPUT"
|
|
echo "Cached XPI exists at ext-$VERSION (asset id $ASSET_ID); skipping AMO sign"
|
|
else
|
|
echo "cached=false" >> "$GITHUB_OUTPUT"
|
|
echo "Release ext-$VERSION exists but has no .xpi asset; will re-sign + re-upload"
|
|
fi
|
|
else
|
|
echo "cached=false" >> "$GITHUB_OUTPUT"
|
|
echo "No release named ext-$VERSION; will sign via AMO and upload"
|
|
fi
|
|
|
|
# No "download cached XPI in sign-extension" step: build-web
|
|
# fetches directly from the Forgejo ext-<version> release asset
|
|
# (removed 2026-05-26 alongside the actions/upload-artifact
|
|
# removal — sign-extension's job is just to ensure the cache
|
|
# exists on Forgejo; the build-web side reads it independently).
|
|
|
|
# web-ext signs whatever manifest.json says, so the derived value has to
|
|
# reach the tree before signing. package.json is written too: the two are
|
|
# required to agree (ci.yml's guard), and a local `npm run build` reads
|
|
# it. Working tree only — never committed, per the note on the derive
|
|
# step.
|
|
- name: Stamp the derived version into manifest.json + package.json
|
|
env:
|
|
DERIVED: ${{ steps.extver.outputs.version }}
|
|
run: |
|
|
python3 - <<'PY'
|
|
import json, os
|
|
|
|
version = os.environ["DERIVED"]
|
|
for path in ("extension/manifest.json", "extension/package.json"):
|
|
with open(path) as fh:
|
|
doc = json.load(fh)
|
|
doc["version"] = version
|
|
with open(path, "w") as fh:
|
|
json.dump(doc, fh, indent=2)
|
|
fh.write("\n")
|
|
print(f"{path}: version -> {version}")
|
|
PY
|
|
|
|
- name: Sign via AMO (cache miss)
|
|
if: steps.cache.outputs.cached != 'true'
|
|
run: |
|
|
cd extension && npm install --no-save --no-audit --no-fund && npm run sign
|
|
env:
|
|
WEB_EXT_API_KEY: ${{ secrets.MOZILLA_AMO_JWT_KEY }}
|
|
WEB_EXT_API_SECRET: ${{ secrets.MOZILLA_AMO_JWT_SECRET }}
|
|
|
|
- name: Upload signed XPI to ext-<version> release (cache miss)
|
|
if: steps.cache.outputs.cached != 'true'
|
|
env:
|
|
TOKEN: ${{ secrets.RELEASE_TOKEN }}
|
|
run: |
|
|
set -eux
|
|
VERSION=${{ steps.extver.outputs.version }}
|
|
# AMO renames signed XPIs with its internal addon-id-safe-string;
|
|
# canonicalize to fabledcurator-<version>.xpi so the FC server's
|
|
# whitelist (backend/app/frontend.py expects 'fabledcurator-*.xpi')
|
|
# keeps working.
|
|
SIGNED=$(ls extension/web-ext-artifacts/*.xpi | head -1)
|
|
XPI="extension/web-ext-artifacts/fabledcurator-$VERSION.xpi"
|
|
cp "$SIGNED" "$XPI"
|
|
# Find-or-create the ext-<version> release. Track whether WE
|
|
# created it so an upload failure below can roll back (don't
|
|
# leave an empty release tombstone that the next run's
|
|
# cache-check mistakes for a partial-failure state).
|
|
#
|
|
# target_commitish is the signing commit, not a branch name: since
|
|
# step 6 either branch can create this release, and hard-coding
|
|
# `main` would tag a dev-signed XPI against a main commit that may
|
|
# not even contain the extension source it was built from.
|
|
STATUS=$(curl -s -o release.json -w "%{http_code}" \
|
|
-H "Authorization: token $TOKEN" \
|
|
"https://git.fabledsword.com/api/v1/repos/bvandeusen/FabledCurator/releases/tags/ext-$VERSION" || echo 000)
|
|
if [ "$STATUS" = "200" ]; then
|
|
CREATED_BY_US=false
|
|
else
|
|
curl -s -X POST -H "Authorization: token $TOKEN" -H "Content-Type: application/json" \
|
|
-d "{\"tag_name\":\"ext-$VERSION\",\"name\":\"Extension $VERSION (signed XPI cache)\",\"body\":\"Internal cache for the signed XPI consumed by build.yml's build-web job. Not a user-facing FC release.\",\"target_commitish\":\"$GITHUB_SHA\"}" \
|
|
-o release.json \
|
|
"https://git.fabledsword.com/api/v1/repos/bvandeusen/FabledCurator/releases"
|
|
CREATED_BY_US=true
|
|
fi
|
|
RELEASE_ID=$(python3 -c "import json; print(json.load(open('release.json'))['id'])")
|
|
test -n "$RELEASE_ID"
|
|
# Rollback-on-failure: if the asset upload fails AND we just
|
|
# created the release in this run, delete it. Prevents an empty
|
|
# ext-<version> release from poisoning the next workflow run
|
|
# (operator-flagged 2026-05-26 — without rollback the next run
|
|
# saw 'release exists, no asset → cache miss → sign' which AMO
|
|
# then rejected with 409 'Version already exists').
|
|
rollback_if_we_created() {
|
|
if [ "$CREATED_BY_US" = "true" ]; then
|
|
echo "Rolling back: deleting just-created release $RELEASE_ID"
|
|
curl -s -X DELETE -H "Authorization: token $TOKEN" \
|
|
"https://git.fabledsword.com/api/v1/repos/bvandeusen/FabledCurator/releases/$RELEASE_ID" || true
|
|
curl -s -X DELETE -H "Authorization: token $TOKEN" \
|
|
"https://git.fabledsword.com/api/v1/repos/bvandeusen/FabledCurator/tags/ext-$VERSION" || true
|
|
fi
|
|
}
|
|
trap 'rollback_if_we_created' EXIT
|
|
HTTP_CODE=$(curl -s -X POST -H "Authorization: token $TOKEN" \
|
|
-F "attachment=@$XPI" \
|
|
-o /dev/null -w "%{http_code}" \
|
|
"https://git.fabledsword.com/api/v1/repos/bvandeusen/FabledCurator/releases/$RELEASE_ID/assets?name=fabledcurator-$VERSION.xpi")
|
|
if [ "$HTTP_CODE" != "201" ] && [ "$HTTP_CODE" != "200" ]; then
|
|
echo "Asset upload failed with HTTP $HTTP_CODE"
|
|
exit 1
|
|
fi
|
|
# Upload succeeded — clear the rollback trap.
|
|
trap - EXIT
|
|
echo "Uploaded fabledcurator-$VERSION.xpi to ext-$VERSION release"
|
|
|
|
# No actions/upload-artifact step: Forgejo Actions (and our
|
|
# act_runner) doesn't support upload-artifact@v4+ (GHES limitation
|
|
# surfaced 2026-05-26). Instead build-web reads the signed XPI
|
|
# straight from the ext-<version> Forgejo release we just uploaded
|
|
# to. Same source of truth; no double-store.
|
|
|
|
build-web:
|
|
# A plain `needs` — no `always()`. That expression existed to let a
|
|
# SKIPPED sign-extension through on a tag push while still blocking a
|
|
# FAILED one. With no tag trigger, sign-extension always runs, so the
|
|
# default behaviour is exactly what we want: a failed sign skips build-web
|
|
# rather than shipping an image without its XPI.
|
|
needs: [sign-extension]
|
|
runs-on: python-ci
|
|
container:
|
|
image: git.fabledsword.com/bvandeusen/ci-python:3.14
|
|
steps:
|
|
- uses: actions/checkout@v4
|
|
with:
|
|
# Not the triggering ref — see the `env:` block at the top. On a
|
|
# scheduled refresh this is `main`; on everything else it is the ref
|
|
# that fired, so this is a no-op on every ordinary path.
|
|
ref: ${{ env.BUILD_REF }}
|
|
# Full history: this job RE-DERIVES the extension version rather than
|
|
# being handed it, and a depth-1 clone derives a wrong, too-low value
|
|
# rather than failing — which would 404 the download of a release
|
|
# that exists perfectly well under its real name.
|
|
fetch-depth: 0
|
|
|
|
# See sign-extension's copy for why this guard exists.
|
|
- name: Guard — a scheduled run must have checked out main
|
|
if: env.IS_REFRESH == 'true'
|
|
run: |
|
|
set -eu
|
|
BRANCH=$(git rev-parse --abbrev-ref HEAD)
|
|
echo "schedule: HEAD is $BRANCH ($(git rev-parse --short HEAD))"
|
|
if [ "$BRANCH" != "main" ]; then
|
|
echo "schedule: expected main, got '$BRANCH'." >&2
|
|
echo "schedule: BUILD_REF was not honoured by the runner." >&2
|
|
echo "schedule: refusing to publish a channel tag from it." >&2
|
|
exit 1
|
|
fi
|
|
|
|
# --- derived values, one line (milestone 313) ------------------------
|
|
# These stopped being shadow output at step 3. `revision` decides
|
|
# whether the build below runs at all and `version` is what the image
|
|
# reports about itself; the load-bearing steps each print only the one
|
|
# they use, so this is the only place the pair appears together. When a
|
|
# build is skipped, this is the line that says what the commit derived.
|
|
#
|
|
# Still diagnostic, so it still must not fail the build — no `set -e`,
|
|
# and every derivation falls back to UNAVAILABLE. A broken echo must
|
|
# never be the reason an image does not ship.
|
|
#
|
|
# What it should say:
|
|
# * a push touching only agent/ moves the agent and leaves web and ml
|
|
# STILL. If web moves, its path set is too wide.
|
|
# * a push touching only docs moves nothing.
|
|
# * a push touching the extension moves the extension AND web, since
|
|
# web bakes in the XPI. If web does not move, its set is too narrow:
|
|
# the reuse check hits, and the channel serves a web image bundling
|
|
# the PREVIOUS XPI while the freshly signed one is orphaned (#3156).
|
|
# * dev and main derive the same values for the same source.
|
|
- name: Report the derived artifact version
|
|
run: |
|
|
set -u
|
|
A=web
|
|
V=$(sh scripts/artifacts.sh version "$A" 2>&1 || echo UNAVAILABLE)
|
|
R=$(sh scripts/artifacts.sh revision "$A" 2>&1 || echo UNAVAILABLE)
|
|
echo "derived: artifact=$A version=$V revision=$R sha=$GITHUB_SHA"
|
|
|
|
- name: Determine tag
|
|
id: tag
|
|
run: |
|
|
# Two trigger shapes, and between them they publish three tags:
|
|
# main → :latest (production, moving — rule 147: main IS production)
|
|
# :c-<sha> (immutable, the rollback unit — rule 145)
|
|
# dev → :dev (the rolling test channel — rule 146)
|
|
#
|
|
# That is the whole list. No :<version>, and no :main — rule 145,
|
|
# narrowed 2026-08-28 once it was verified that nothing pins:
|
|
# "a third name for the same thing is upkeep for a model we do not
|
|
# run." The date tag published between milestone 313 step 3 and
|
|
# milestone 318 was exactly that; :main was a second moving name for
|
|
# whatever :latest already pointed at.
|
|
#
|
|
# `dev` gets no :c-<sha> deliberately. On a channel whose entire
|
|
# contract is that it moves, a per-push immutable tag is a rollback
|
|
# target nobody has ever pulled, accumulating forever. The accepted
|
|
# cost: on dev there is no rollback but the previous :dev, which is
|
|
# gone — recovery is revert-on-git plus a CI cycle.
|
|
#
|
|
# Reinstating :<version> is a real decision, not a default. It earns
|
|
# its place when something genuinely pins: a second instance held on
|
|
# a known-good build, or a deliberately frozen window. Tag at the
|
|
# moment you decide to freeze; no back-catalogue is needed.
|
|
#
|
|
# POSIX-safe substring (the runner shell is dash/BusyBox sh, not
|
|
# bash — `${var:0:7}` errors with "Bad substitution"; cut works
|
|
# everywhere). Operator-flagged 2026-06-01 after the first :c-<sha>
|
|
# main-push build failed at this step.
|
|
SHORT_SHA=$(printf '%s' "$GITHUB_SHA" | cut -c1-7)
|
|
|
|
# A scheduled refresh publishes the CHANNEL and nothing else
|
|
# (#3154). :c-<sha> for main's HEAD already exists and names the
|
|
# bytes that commit actually built; re-pushing it over refreshed
|
|
# base layers would break the one tag rule 145 makes immutable —
|
|
# and it is the rollback unit, so the breakage would surface on the
|
|
# day somebody needed it.
|
|
#
|
|
# The accepted consequence: between a refresh and the next main
|
|
# push, :latest and :c-<sha> point at different manifests. That is
|
|
# the design, not drift. They RE-CONVERGE on that push — it hits
|
|
# reuse (a refresh does not move fc.revision, because it does not
|
|
# touch the source), and the repoint step then writes the new
|
|
# :c-<sha> from the refreshed :latest. So the rollback unit ends up
|
|
# naming the bytes production is actually running, which is the
|
|
# property that matters.
|
|
#
|
|
# Checked BEFORE the ref test, not after: a scheduled run's
|
|
# GITHUB_REF is the default branch (dev), so the main test would
|
|
# never fire on it.
|
|
if [ "${IS_REFRESH:-}" = "true" ]; then
|
|
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator:latest" >> "$GITHUB_OUTPUT"
|
|
echo "channel=main" >> "$GITHUB_OUTPUT"
|
|
elif [ "${GITHUB_REF##*/}" = "main" ]; then
|
|
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator:latest,git.fabledsword.com/bvandeusen/fabledcurator:c-${SHORT_SHA}" >> "$GITHUB_OUTPUT"
|
|
echo "channel=main" >> "$GITHUB_OUTPUT"
|
|
else
|
|
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator:dev" >> "$GITHUB_OUTPUT"
|
|
echo "channel=dev" >> "$GITHUB_OUTPUT"
|
|
fi
|
|
|
|
# A shell step, not docker/login-action@v3, because the action's shared
|
|
# cache races itself (#3118). act_runner caches a remote action under one
|
|
# /root/.cache/act/<hash> per runner, and build-web, build-ml and
|
|
# build-agent all start in the same second and all want this same action.
|
|
# One job re-clones the directory — which empties and repopulates it —
|
|
# while another is walking it to copy into its container, and the walker
|
|
# lstat()s a file that has just vanished. It failed twice on 2026-08-27,
|
|
# naming a DIFFERENT missing file each time (`eslint.config.mjs`, then
|
|
# `jest.config.ts`), which is what rules out a corrupt cache and points at
|
|
# a race. The loser dies with MODULE_NOT_FOUND on dist/index.js before the
|
|
# action runs at all, so the secret is never even reached.
|
|
#
|
|
# Nothing is lost by dropping it: logging in is one command, the docker
|
|
# CLI is already in the CI image (ci-requirements.md), and the same
|
|
# reasoning as family rule 5 applies — a marketplace action buys nothing
|
|
# when the tool is baked into the image the workflow already selected.
|
|
#
|
|
# Password on stdin, never as an argument: an argument lands in the
|
|
# process table and draws docker's own deprecation warning.
|
|
- name: Login to Forgejo registry
|
|
env:
|
|
TOKEN: ${{ secrets.RELEASE_TOKEN }}
|
|
ACTOR: ${{ github.actor }}
|
|
run: echo "$TOKEN" | docker login git.fabledsword.com -u "$ACTOR" --password-stdin
|
|
|
|
# A REAL buildx builder, not the default `docker` driver (#3114, #3190).
|
|
#
|
|
# The default driver builds through the local dockerd. It cannot export a
|
|
# registry cache at all — which is why the agent rebuilds a ~6.3 GB CUDA
|
|
# + torch image from scratch whenever the runner's local cache is cold,
|
|
# measured at 9m26s against 7s warm. It is also #3190's leading suspect:
|
|
# after a registry-direct push it resolves image metadata against a local
|
|
# store the push never filled, and reports `No such image` on an image
|
|
# that published perfectly well three seconds earlier.
|
|
#
|
|
# These jobs run INSIDE a container against a mounted docker socket, so
|
|
# the buildkit container this starts is a SIBLING of the job container,
|
|
# not a child. That works over the socket mount; it had never been tried
|
|
# here before milestone 326 step 1.
|
|
- name: Set up buildx
|
|
uses: docker/setup-buildx-action@v3
|
|
|
|
# --- reuse-if-published (milestone 313, step 4) ----------------------
|
|
# Does the image the channel tag already points at carry THIS commit's
|
|
# revision? If so the bytes this job would produce are already published
|
|
# and the build is pure waste: the remaining tags get repointed at that
|
|
# existing manifest instead, registry-side, in seconds.
|
|
#
|
|
# Keyed on an `fc.revision` LABEL rather than on a tag of its own
|
|
# (milestone 318 step 3). A tag would be a name minted per build that one
|
|
# thing reads — what rule 145 narrowed against — and would be prunable
|
|
# under the registry's keep_pattern (#3157), silently expiring the cache.
|
|
# A label rides inside a tag that has to exist anyway.
|
|
#
|
|
# An image with no such label reads as a miss and rebuilds. That is the
|
|
# migration, not a fault: labels cannot be backfilled, since the reuse
|
|
# path copies a manifest and config labels are not manifest annotations.
|
|
# Each artifact pays one rebuild, once.
|
|
#
|
|
# This is what stops a push that touched only `agent/` from rebuilding
|
|
# web and ml, and a merge to main from rebuilding what dev already built.
|
|
#
|
|
# The failure direction is deliberate. An inspect that errors for ANY
|
|
# reason — network, auth, a registry hiccup — reads as a miss and the
|
|
# build runs. Only a genuine 200 skips one, so there is no path here
|
|
# that skips a build that was actually needed; the worst case is paying
|
|
# for a build we could have avoided.
|
|
#
|
|
# BASE-IMAGE FRESHNESS: an artifact whose source stops moving stops
|
|
# picking up base-image updates. Milestone 318 removed the argument this
|
|
# used to need rather than answering it — with no version tags there is
|
|
# no immutable name a refresh could contradict, and rule 145 already
|
|
# allows a rebuild with different contents to republish a MOVING tag.
|
|
# So a refresh is just a build. A scheduled channel-only one is tracked
|
|
# separately (#3154); it does not belong in the push path.
|
|
- name: Is this content already published?
|
|
id: reuse
|
|
env:
|
|
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator
|
|
CHANNEL: ${{ steps.tag.outputs.channel }}
|
|
# Empty on a push; the string "true" only from a workflow_dispatch
|
|
# that asked for it. `github.event.inputs` rather than the `inputs`
|
|
# context — release.yml already uses that form, and it is the one
|
|
# this runner is known to evaluate. Read through env rather than
|
|
# interpolated into the run block, same rule as release.yml's TAG.
|
|
FORCE: ${{ github.event.inputs.force_build }}
|
|
# A scheduled refresh has to bypass reuse by construction: it
|
|
# rebuilds the SAME source, so fc.revision always matches and the
|
|
# check would skip every refresh there has ever been.
|
|
run: |
|
|
set -eu
|
|
DERIVED=$(sh scripts/artifacts.sh revision web)
|
|
echo "revision=$DERIVED" >> "$GITHUB_OUTPUT"
|
|
# Baked into the web image as FC_VERSION and reported by /api/health.
|
|
# A pure function of the revision — same commit, same string — so it
|
|
# adds no variability the reuse check would have to account for.
|
|
echo "version=$(sh scripts/artifacts.sh version web)" >> "$GITHUB_OUTPUT"
|
|
|
|
# The build clock, pinned to the same commit (#3265). Without it
|
|
# buildkit stamps the image config with the wall clock of the build,
|
|
# so identical layers republish under a new config blob and the
|
|
# channel tag gets a new manifest digest for no reason. Derived from
|
|
# `newest()` like revision and version, so all three name one commit
|
|
# and cannot drift apart.
|
|
echo "epoch=$(sh scripts/artifacts.sh epoch web)" >> "$GITHUB_OUTPUT"
|
|
|
|
# The moving tag for this channel. Which tag we ask IS the channel —
|
|
# that is why the revision needs no -main/-dev qualifier any more.
|
|
if [ "$CHANNEL" = "main" ]; then T=latest; else T=dev; fi
|
|
echo "channel_ref=$IMAGE:$T" >> "$GITHUB_OUTPUT"
|
|
|
|
# WHERE THE BUILD PUBLISHES, which is not always the channel — and
|
|
# whether the channel then has to be written separately.
|
|
#
|
|
# On a push the build writes the channel tag directly: the bytes came
|
|
# from a commit, and a commit is the thing CI tests. Nothing to hold
|
|
# it behind.
|
|
#
|
|
# On the scheduled refresh it writes a CANDIDATE tag instead. A
|
|
# refresh rebuilds against freshly resolved base images, and the web
|
|
# image's runtime is a line of UNPINNED Debian packages (ffmpeg,
|
|
# libjpeg62-turbo, libpq5, megatools…) re-resolved on every build.
|
|
# Nothing in ci.yml can see that: its lanes run on ci-python:3.14 and
|
|
# install requirements.txt, and a base bump changes neither. So
|
|
# refreshed bytes have to be proven before :latest names them, and
|
|
# proving needs a moment between "built" and "published" to occupy.
|
|
# This is that moment; :latest goes on naming the build that works
|
|
# until something says otherwise.
|
|
#
|
|
# `:refresh-candidate` is one moving ref per image, overwritten in
|
|
# place, holding a build nobody is told to pull — the shape rule 145
|
|
# already allows for :buildcache, not the per-build tag family that
|
|
# milestone 318 withdrew.
|
|
#
|
|
# Both values are decided HERE, beside `hit`, for the reason the
|
|
# force/schedule branch below gives: one step decides what this job
|
|
# does. A promote condition derived independently could disagree with
|
|
# the tag the build actually wrote.
|
|
if [ "${IS_REFRESH:-}" = "true" ]; then
|
|
echo "build_ref=$IMAGE:refresh-candidate" >> "$GITHUB_OUTPUT"
|
|
echo "promote=true" >> "$GITHUB_OUTPUT"
|
|
else
|
|
echo "build_ref=$IMAGE:$T" >> "$GITHUB_OUTPUT"
|
|
echo "promote=false" >> "$GITHUB_OUTPUT"
|
|
fi
|
|
|
|
# Compare VALUES, never exit codes. Measured on buildx v0.36.1
|
|
# (run 4732): a missing key returns an empty string and exits 0, so
|
|
# branching on the exit code would read "no label yet" as success.
|
|
# An unreachable tag also lands here as empty via the `|| echo`.
|
|
# Empty never equals a 12-char revision, so every uncertain case
|
|
# falls through to a build — the safe direction, with no special
|
|
# casing for it.
|
|
#
|
|
# Read the SPECIFIC key. The map also carries whatever the base image
|
|
# set, and `org.opencontainers.image.version` sits right beside ours
|
|
# looking like a plausible answer (it reads 24.04 on the agent).
|
|
PUBLISHED=$(docker buildx imagetools inspect "$IMAGE:$T" \
|
|
--format '{{ index .Image.Config.Labels "fc.revision" }}' \
|
|
2>/dev/null || echo "")
|
|
echo "reuse: $IMAGE:$T carries fc.revision=${PUBLISHED:-<none>}; derived=$DERIVED"
|
|
if [ -z "$PUBLISHED" ] && docker buildx imagetools inspect "$IMAGE:$T" >/dev/null 2>&1; then
|
|
# The tag resolves but carries no readable label. Expected exactly
|
|
# once per artifact, during the migration onto labels. If it recurs
|
|
# every push, something is rewriting the channel tag as a manifest
|
|
# index — see the repoint step's note.
|
|
echo "reuse: NOTE $IMAGE:$T exists but has no readable fc.revision."
|
|
echo "reuse: NOTE Fine once, while migrating. Every push means the"
|
|
echo "reuse: NOTE tag is being index-wrapped and reuse is dead."
|
|
fi
|
|
|
|
# FORCE is checked here rather than in the build step's `if:`, so
|
|
# that one decision drives everything downstream. The repoint step
|
|
# keys off `hit` too, and a force that bypassed only the build would
|
|
# leave the two disagreeing about what just happened.
|
|
if [ "${FORCE:-false}" = "true" ]; then
|
|
echo "hit=false" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: force_build set — building regardless"
|
|
elif [ "${IS_REFRESH:-}" = "true" ]; then
|
|
echo "hit=false" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: scheduled base refresh — building regardless"
|
|
elif [ -n "$PUBLISHED" ] && [ "$PUBLISHED" = "$DERIVED" ]; then
|
|
echo "hit=true" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: already published — skipping the build"
|
|
else
|
|
echo "hit=false" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: not published — building"
|
|
fi
|
|
|
|
- name: Download signed XPI from Forgejo release asset
|
|
# dev and main each bundle the XPI their own sign-extension just
|
|
# published — the point of the channel work (milestone 271 step 6): the
|
|
# dev image carries the extension being developed, rather than
|
|
# requiring a merge to try it.
|
|
#
|
|
# The 10-minute polling loop that used to live here is gone with the
|
|
# tag trigger (milestone 318 step 2). It existed for one shape only: a
|
|
# release cut fired the tag build and the main build together, the tag
|
|
# build skipped sign-extension and raced straight here, and it lost
|
|
# every time (operator-flagged 2026-05-27 after v26.05.27.0). Polling
|
|
# was the fix for a build that should not have been running.
|
|
#
|
|
# sign-extension is a `needs` dependency and it succeeded, so the
|
|
# release exists. A single fetch is correct, and a 404 now means a real
|
|
# disagreement about the derived version rather than a race — which is
|
|
# exactly what should fail loudly instead of being slept through.
|
|
#
|
|
# Still gated on the reuse miss: a published image already contains its
|
|
# XPI, so this would fetch a file nothing then reads.
|
|
if: steps.reuse.outputs.hit != 'true'
|
|
env:
|
|
TOKEN: ${{ secrets.RELEASE_TOKEN }}
|
|
run: |
|
|
set -eux
|
|
# Re-derived, not read from the repo: sign-extension published
|
|
# ext-<derived>, and the committed version has been inert since
|
|
# milestone 271 step 4. Both jobs run `packaging.sh version` over the
|
|
# same commit, so they agree by construction — and if they ever
|
|
# didn't, this download 404s and the build fails loudly instead of
|
|
# shipping a stale XPI.
|
|
VERSION=$(sh extension/scripts/packaging.sh version)
|
|
# One fetch, no retry. sign-extension ran to success in this same
|
|
# workflow and published ext-$VERSION; both jobs derive $VERSION from
|
|
# the same commit, so they agree by construction. A 404 here means
|
|
# they did NOT agree, and sleeping on that would only delay the
|
|
# report.
|
|
STATUS=$(curl -s -o release.json -w "%{http_code}" \
|
|
-H "Authorization: token $TOKEN" \
|
|
"https://git.fabledsword.com/api/v1/repos/bvandeusen/FabledCurator/releases/tags/ext-$VERSION" || echo 000)
|
|
if [ "$STATUS" != "200" ]; then
|
|
echo "ERROR: ext-$VERSION release not found (HTTP $STATUS)."
|
|
echo "sign-extension succeeded in this run, so it published some"
|
|
echo "other version — the two jobs derived different values for one"
|
|
echo "commit. Check that both checked out with fetch-depth: 0."
|
|
exit 1
|
|
fi
|
|
# Extract the .xpi asset's browser_download_url (Forgejo's
|
|
# /releases/assets/<id> endpoint returns ASSET METADATA, not
|
|
# the binary blob — operator-flagged 2026-05-26: my prior
|
|
# code curl'd the metadata endpoint without -f and wrote the
|
|
# resulting 404-page-not-found text into fabledcurator-*.xpi,
|
|
# which Firefox then rejected as "corrupt").
|
|
# browser_download_url is the canonical binary endpoint and
|
|
# is also publicly accessible (no token needed) but we pass
|
|
# the token anyway for symmetry with private-repo support.
|
|
DOWNLOAD_URL=$(python3 -c "import json; r=json.load(open('release.json')); xpis=[a for a in r.get('assets', []) if a.get('name','').endswith('.xpi')]; print(xpis[0]['browser_download_url'])")
|
|
test -n "$DOWNLOAD_URL"
|
|
echo "Downloading XPI from: $DOWNLOAD_URL"
|
|
mkdir -p frontend/public/extension
|
|
DEST="frontend/public/extension/fabledcurator-$VERSION.xpi"
|
|
# -f = fail on HTTP error (prevents silent corruption like the
|
|
# 2026-05-26 incident); -L = follow redirects.
|
|
curl -sfL -H "Authorization: token $TOKEN" -o "$DEST" "$DOWNLOAD_URL"
|
|
# Sanity check: the binary should start with the ZIP magic (PK\x03\x04).
|
|
# If it's anything else, the next docker build will ship a corrupt XPI.
|
|
MAGIC=$(head -c 2 "$DEST" | od -An -c | tr -d ' \n')
|
|
if [ "$MAGIC" != "PK" ]; then
|
|
echo "ERROR: downloaded XPI does not start with ZIP magic 'PK' (got '$MAGIC')"
|
|
echo "File contents preview:"
|
|
head -c 200 "$DEST"
|
|
exit 1
|
|
fi
|
|
cp "$DEST" "frontend/public/extension/fabledcurator-latest.xpi"
|
|
ls -la frontend/public/extension/
|
|
|
|
- name: Build and push web image
|
|
if: steps.reuse.outputs.hit != 'true'
|
|
# Read by buildx out of the ENVIRONMENT, not passed as a build-arg —
|
|
# it normalises the image config's `created` field and the history
|
|
# timestamps rather than being consumed by the Dockerfile. See #3265
|
|
# and the reuse step's `epoch` output.
|
|
env:
|
|
SOURCE_DATE_EPOCH: ${{ steps.reuse.outputs.epoch }}
|
|
uses: docker/build-push-action@v5
|
|
with:
|
|
context: .
|
|
file: Dockerfile
|
|
push: true
|
|
# Re-resolve the FROM references against the registry instead of
|
|
# trusting whatever digest the cache was built against. This is the
|
|
# whole mechanism of the scheduled refresh (#3154): if the base tag
|
|
# moved, the FROM layer's cache key changes, every layer above it
|
|
# invalidates, and the image genuinely rebuilds.
|
|
#
|
|
# MEASURED on the first real fire, run 4934 (#3265): when the base
|
|
# did NOT move, the build was ~13s with every content step CACHED —
|
|
# and the channel tag STILL got a new manifest digest, because
|
|
# buildkit stamps a fresh image config per run and republishes the
|
|
# identical layers under it. All three images moved that way on
|
|
# 2026-08-30 with nothing whatsoever changed in them.
|
|
#
|
|
# SOURCE_DATE_EPOCH (below) is the fix: pinned to the commit the
|
|
# content came from, the config is byte-identical across runs, so
|
|
# the manifest digest is too and the push is a registry no-op. A
|
|
# digest change means the content changed again, which is the only
|
|
# thing a digest is any use for.
|
|
#
|
|
# What `pull` does NOT catch either: a Debian package update inside
|
|
# the `apt-get install` layer while the base tag itself stands
|
|
# still. The official python/cuda images rebuild with those updates
|
|
# baked in, so this is a lag rather than a hole; closing it needs
|
|
# `no-cache: true`, which is a much larger version of the same
|
|
# churn #3265 is about.
|
|
#
|
|
# Only on the schedule. An ordinary push wants the cached base.
|
|
pull: ${{ env.IS_REFRESH == 'true' }}
|
|
# ONE tag, the channel's. Every other tag is written by the step
|
|
# below, registry-side. buildx here pushes the first tag to the
|
|
# registry and then re-pushes the rest through the DOCKER driver,
|
|
# out of a local image store a registry-direct build never filled —
|
|
# #3190, which cost `main` its :c-<sha> on 2026-08-29 while :latest
|
|
# published perfectly well.
|
|
tags: ${{ steps.reuse.outputs.build_ref }}
|
|
# The reuse key. Read back off the channel tag on the next push to
|
|
# decide whether that push needs to build at all, so this is not
|
|
# decoration — an unstamped image is one that will always rebuild.
|
|
labels: |
|
|
fc.revision=${{ steps.reuse.outputs.revision }}
|
|
# LOAD-BEARING, not a preference. On the default docker driver these
|
|
# were no-ops; on the docker-container driver above,
|
|
# build-push-action@v5 defaults provenance to TRUE when pushing.
|
|
# Provenance attaches an attestation manifest, which makes the pushed
|
|
# tag a manifest INDEX — and `.Image.Config.Labels` does not resolve
|
|
# through an index.
|
|
#
|
|
# The label directly above IS the reuse key. Wrap the channel tag in
|
|
# an index and the next push reads fc.revision=<none>, misses, and
|
|
# rebuilds. Then so does the one after that, forever. Nothing fails,
|
|
# nothing goes red, and the only symptom is the bill. That is #3183
|
|
# arriving through a different door, and note #3127 §4 records the
|
|
# same shape for `platforms:`.
|
|
provenance: false
|
|
sbom: false
|
|
# The ONLY cache this driver can have. `docker-container` gets a
|
|
# FRESH buildkit instance per job, so unlike the default docker
|
|
# driver it has no local layer store to fall back on — measured on
|
|
# run 4896, the first builds after the driver change: web 3m44s
|
|
# (was 2m23s), ml 3m49s (was 3m20s), agent 11m12s (was 9m26s). The
|
|
# driver change ALONE is a regression; this is the other half of it.
|
|
#
|
|
# mode=max so intermediate stages cache too. web's frontend-builder
|
|
# stage and the agent's two ~150s pip layers are the whole cost, and
|
|
# they are exactly what a min-mode cache would drop.
|
|
#
|
|
# A `:buildcache` tag is NOT the withdrawn tag scheme coming back.
|
|
# Rule 145 narrowed against names NOTHING reads; this one is read by
|
|
# every build that runs, is one moving ref per image rather than one
|
|
# per build, holds cache blobs rather than a shippable artifact, and
|
|
# is overwritten in place rather than accumulating. It is closer to
|
|
# :dev than to the :2026.8.28 tags milestone 318 deleted. (#3114.)
|
|
cache-from: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator:buildcache
|
|
cache-to: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator:buildcache,mode=max
|
|
# Only the web image carries these: it is the one with a UI and an
|
|
# HTTP surface to report them on. The ml and agent images have
|
|
# nothing to tell.
|
|
build-args: |
|
|
FC_CHANNEL=${{ steps.tag.outputs.channel }}
|
|
FC_VERSION=${{ steps.reuse.outputs.version }}
|
|
|
|
# Point the channel tag at the candidate the refresh just built.
|
|
#
|
|
# Unconditional TODAY, so this milestone never leaves the refresh in a
|
|
# state where it builds and publishes nothing. Step 4 wraps it in the
|
|
# smoke suite's verdict; until then the scheduled path behaves exactly
|
|
# as it did, just via two operations instead of one.
|
|
#
|
|
# NOT `imagetools create`. That wraps its source in an INDEX, and an
|
|
# indexed channel tag is the one thing this pipeline cannot survive:
|
|
# `.Image.Config.Labels` does not resolve through an index, so the
|
|
# fc.revision the reuse check reads off the channel tag would come back
|
|
# empty, every subsequent push would miss and rebuild, and nothing would
|
|
# go red. That is #3183, observed on run 4751 — reuse worked exactly once
|
|
# and the only symptom was the bill. The repoint step below excludes its
|
|
# own source tag for precisely this reason; a promote that re-introduced
|
|
# the wrap through a different door would undo that care.
|
|
#
|
|
# A manifest PUT is what "make this tag name that image" means at the
|
|
# registry level: the same bytes under the same media type, so the digest
|
|
# is identical, the media type is preserved, and no layer moves.
|
|
- name: Promote the refresh candidate to the channel
|
|
if: steps.reuse.outputs.promote == 'true'
|
|
env:
|
|
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator
|
|
CHANNEL_REF: ${{ steps.reuse.outputs.channel_ref }}
|
|
TOKEN: ${{ secrets.RELEASE_TOKEN }}
|
|
ACTOR: ${{ github.actor }}
|
|
run: |
|
|
set -eu
|
|
REPO=${IMAGE#git.fabledsword.com/}
|
|
TAG=${CHANNEL_REF##*:}
|
|
|
|
# Registry auth is its own token exchange — the `docker login` above
|
|
# authenticates the docker client, not curl. Deadline on every call
|
|
# (rule 156): a registry that stops answering must fail this step,
|
|
# not hang the weekly refresh until the job times out.
|
|
BEARER=$(curl -fsS --max-time 30 -u "$ACTOR:$TOKEN" \
|
|
"https://git.fabledsword.com/v2/token?scope=repository:$REPO:pull,push&service=git.fabledsword.com" \
|
|
| python3 -c 'import sys,json; print(json.load(sys.stdin)["token"])')
|
|
|
|
# Ask for the image manifest media types ONLY. Offering the index
|
|
# types too would let the registry hand back an index if one ever
|
|
# existed at this tag, and we would faithfully copy the thing we are
|
|
# trying not to create.
|
|
ACCEPT='application/vnd.oci.image.manifest.v1+json, application/vnd.docker.distribution.manifest.v2+json'
|
|
CT=$(curl -fsS --max-time 60 -o manifest.json -D headers.txt \
|
|
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
|
|
"https://git.fabledsword.com/v2/$REPO/manifests/refresh-candidate" \
|
|
&& tr -d '\r' < headers.txt | awk -F': ' '/^[Cc]ontent-[Tt]ype:/{print $2}')
|
|
test -n "$CT"
|
|
SRC_DIGEST=$(tr -d '\r' < headers.txt | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
|
|
echo "promote: candidate is $SRC_DIGEST ($CT)"
|
|
|
|
curl -fsS --max-time 120 -X PUT \
|
|
-H "Authorization: Bearer $BEARER" -H "Content-Type: $CT" \
|
|
--data-binary @manifest.json \
|
|
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG"
|
|
|
|
# Read it back. A PUT that returned 2xx but landed something else is
|
|
# exactly the silent-and-plausible failure this pipeline keeps
|
|
# producing, and the check costs one request.
|
|
NOW=$(curl -fsS --max-time 30 -o /dev/null -D - \
|
|
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
|
|
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG" \
|
|
| tr -d '\r' | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
|
|
if [ "$NOW" != "$SRC_DIGEST" ]; then
|
|
echo "promote: $IMAGE:$TAG is $NOW, expected $SRC_DIGEST" >&2
|
|
exit 1
|
|
fi
|
|
echo "promote: $IMAGE:$TAG now names $NOW"
|
|
|
|
# Every tag but the channel's own is written HERE, registry-side,
|
|
# whether or not a build ran. Each -t becomes another reference to the
|
|
# SAME manifest the channel tag holds, so :c-<sha> is byte-identical to
|
|
# what is published rather than a lookalike rebuild.
|
|
#
|
|
# Owning the build path too is #3190's fix, not a tidy-up:
|
|
#
|
|
# #27 pushing …/fabledcurator:latest DONE 15.8s
|
|
# #28 pushing …/fabledcurator:c-0e15c44 with docker
|
|
# #28 ERROR: tag does not exist: …:c-0e15c44
|
|
#
|
|
# Intermittent — build-ml made the identical two-tag push seconds later
|
|
# and succeeded — and worse than it looks. `:latest` had already
|
|
# published, so production was correct while the immutable rollback tag
|
|
# rule 145 requires of every main push simply did not exist. Nothing but
|
|
# the red job would ever have noticed: a missing :c-<sha> has no
|
|
# consumer that fails, so it surfaces when somebody needs to roll back.
|
|
#
|
|
# `imagetools create` is a registry-side manifest copy — no layer
|
|
# transfer, no local daemon, nothing that can be absent. The reuse case
|
|
# has always gone this way, so this puts the build case on the code that
|
|
# was already proven rather than on a second path.
|
|
#
|
|
# Running on every path also keeps family rule 146 true: a rolling
|
|
# channel refreshes itself, so skipping a build must never leave :dev or
|
|
# :latest pointing at something older than the commit just pushed.
|
|
#
|
|
# The cost, accepted knowingly: `imagetools create` wraps its source in
|
|
# an index, so :c-<sha> becomes an index and fc.revision does not
|
|
# resolve through it. Nothing reads that label off :c-<sha> — the reuse
|
|
# check only ever inspects the CHANNEL tag — and the index names the
|
|
# same manifest, so a pull is byte-identical. The reuse path already
|
|
# produced :c-<sha> this way; this only makes it uniform.
|
|
- name: Write the remaining tags from the published image
|
|
env:
|
|
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator
|
|
SOURCE: ${{ steps.reuse.outputs.channel_ref }}
|
|
TAGS: ${{ steps.tag.outputs.tags }}
|
|
run: |
|
|
set -euf
|
|
# The source tag is EXCLUDED from the targets, and that is load-
|
|
# bearing rather than an optimisation.
|
|
#
|
|
# `imagetools create` wraps the source manifest in an INDEX. Point it
|
|
# at the channel tag with that same tag as a target and the tag stops
|
|
# being a plain image — after which `.Image.Config.Labels` no longer
|
|
# resolves through it and the fc.revision label reads as absent. The
|
|
# next push then misses and rebuilds, so reuse worked exactly once
|
|
# and every subsequent push paid full price. Observed on run 4751:
|
|
# ml:dev reported fc.revision=<none> one push after run 4749 had read
|
|
# a7e626a67a79 off it. Nothing failed; the savings just evaporated.
|
|
#
|
|
# Excluding the source means the channel tag is only ever written
|
|
# by a real build, so it stays a plain image and stays readable.
|
|
# On dev that leaves nothing to do either way: the build pushed :dev
|
|
# itself, or the hit established it was already right. On main it
|
|
# leaves :c-<sha>, which rule 145 requires of every main push whether
|
|
# or not a build ran.
|
|
#
|
|
# steps.tag emits ONE comma-separated list; imagetools wants a -t per
|
|
# ref. (That list used to feed docker/build-push-action directly —
|
|
# which is exactly what #3190 made unsafe.)
|
|
ARGS=""
|
|
IFS=,
|
|
for t in $TAGS; do
|
|
[ "$t" = "$SOURCE" ] && continue
|
|
ARGS="$ARGS -t $t"
|
|
done
|
|
unset IFS
|
|
#
|
|
# This is also the whole of the scheduled refresh's tag handling
|
|
# (#3154): a refresh's tag list is the channel tag alone, so SOURCE
|
|
# is the only entry, it gets excluded, and this step correctly does
|
|
# nothing. No `if:` on the step and no schedule special-case —
|
|
# excluding the source was already the right rule.
|
|
if [ -z "$ARGS" ]; then
|
|
echo "repoint: $SOURCE is the only tag for this channel and"
|
|
echo "repoint: already holds this revision — nothing to write."
|
|
exit 0
|
|
fi
|
|
# shellcheck disable=SC2086
|
|
docker buildx imagetools create $ARGS "$SOURCE"
|
|
echo "repointed from $SOURCE:$ARGS"
|
|
|
|
build-ml:
|
|
runs-on: python-ci
|
|
container:
|
|
image: git.fabledsword.com/bvandeusen/ci-python:3.14
|
|
steps:
|
|
- uses: actions/checkout@v4
|
|
with:
|
|
# Not the triggering ref — see the `env:` block at the top. On a
|
|
# scheduled refresh this is `main`; on everything else it is the ref
|
|
# that fired, so this is a no-op on every ordinary path.
|
|
ref: ${{ env.BUILD_REF }}
|
|
# Full history: this job derives its artifact's version from the
|
|
# commit its shipped files last changed in (milestone 313). A
|
|
# depth-1 clone cannot see that commit — it either derives a wrong,
|
|
# too-low value or finds nothing at all, and neither is a failure
|
|
# the build would otherwise notice.
|
|
fetch-depth: 0
|
|
|
|
# See sign-extension's copy for why this guard exists.
|
|
- name: Guard — a scheduled run must have checked out main
|
|
if: env.IS_REFRESH == 'true'
|
|
run: |
|
|
set -eu
|
|
BRANCH=$(git rev-parse --abbrev-ref HEAD)
|
|
echo "schedule: HEAD is $BRANCH ($(git rev-parse --short HEAD))"
|
|
if [ "$BRANCH" != "main" ]; then
|
|
echo "schedule: expected main, got '$BRANCH'." >&2
|
|
echo "schedule: BUILD_REF was not honoured by the runner." >&2
|
|
echo "schedule: refusing to publish a channel tag from it." >&2
|
|
exit 1
|
|
fi
|
|
|
|
# --- derived values, one line (milestone 313) ------------------------
|
|
# These stopped being shadow output at step 3. `revision` decides
|
|
# whether the build below runs at all and `version` is what the image
|
|
# reports about itself; the load-bearing steps each print only the one
|
|
# they use, so this is the only place the pair appears together. When a
|
|
# build is skipped, this is the line that says what the commit derived.
|
|
#
|
|
# Still diagnostic, so it still must not fail the build — no `set -e`,
|
|
# and every derivation falls back to UNAVAILABLE. A broken echo must
|
|
# never be the reason an image does not ship.
|
|
#
|
|
# What it should say:
|
|
# * a push touching only agent/ moves the agent and leaves web and ml
|
|
# STILL. If web moves, its path set is too wide.
|
|
# * a push touching only docs moves nothing.
|
|
# * a push touching the extension moves the extension AND web, since
|
|
# web bakes in the XPI. If web does not move, its set is too narrow:
|
|
# the reuse check hits, and the channel serves a web image bundling
|
|
# the PREVIOUS XPI while the freshly signed one is orphaned (#3156).
|
|
# * dev and main derive the same values for the same source.
|
|
- name: Report the derived artifact version
|
|
run: |
|
|
set -u
|
|
A=ml
|
|
V=$(sh scripts/artifacts.sh version "$A" 2>&1 || echo UNAVAILABLE)
|
|
R=$(sh scripts/artifacts.sh revision "$A" 2>&1 || echo UNAVAILABLE)
|
|
echo "derived: artifact=$A version=$V revision=$R sha=$GITHUB_SHA"
|
|
|
|
- name: Determine tag
|
|
id: tag
|
|
run: |
|
|
# Mirrors build-web's tag list; see the comment there.
|
|
# POSIX-safe substring (the runner shell is dash/BusyBox sh, not
|
|
# bash — `${var:0:7}` errors with "Bad substitution"; cut works
|
|
# everywhere). Operator-flagged 2026-06-01 after first :c-<sha>
|
|
# main-push build failed at this step.
|
|
SHORT_SHA=$(printf '%s' "$GITHUB_SHA" | cut -c1-7)
|
|
# Mirrors build-web's tag list and its schedule handling; see
|
|
# the comments there.
|
|
if [ "${IS_REFRESH:-}" = "true" ]; then
|
|
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator-ml:latest" >> "$GITHUB_OUTPUT"
|
|
echo "channel=main" >> "$GITHUB_OUTPUT"
|
|
elif [ "${GITHUB_REF##*/}" = "main" ]; then
|
|
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator-ml:latest,git.fabledsword.com/bvandeusen/fabledcurator-ml:c-${SHORT_SHA}" >> "$GITHUB_OUTPUT"
|
|
echo "channel=main" >> "$GITHUB_OUTPUT"
|
|
else
|
|
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator-ml:dev" >> "$GITHUB_OUTPUT"
|
|
echo "channel=dev" >> "$GITHUB_OUTPUT"
|
|
fi
|
|
|
|
# Shell step rather than docker/login-action — see build-web's note on
|
|
# the shared action-cache race (#3118).
|
|
- name: Login to Forgejo registry
|
|
env:
|
|
TOKEN: ${{ secrets.RELEASE_TOKEN }}
|
|
ACTOR: ${{ github.actor }}
|
|
run: echo "$TOKEN" | docker login git.fabledsword.com -u "$ACTOR" --password-stdin
|
|
|
|
# A REAL buildx builder, not the default `docker` driver (#3114, #3190).
|
|
#
|
|
# The default driver builds through the local dockerd. It cannot export a
|
|
# registry cache at all — which is why the agent rebuilds a ~6.3 GB CUDA
|
|
# + torch image from scratch whenever the runner's local cache is cold,
|
|
# measured at 9m26s against 7s warm. It is also #3190's leading suspect:
|
|
# after a registry-direct push it resolves image metadata against a local
|
|
# store the push never filled, and reports `No such image` on an image
|
|
# that published perfectly well three seconds earlier.
|
|
#
|
|
# These jobs run INSIDE a container against a mounted docker socket, so
|
|
# the buildkit container this starts is a SIBLING of the job container,
|
|
# not a child. That works over the socket mount; it had never been tried
|
|
# here before milestone 326 step 1.
|
|
- name: Set up buildx
|
|
uses: docker/setup-buildx-action@v3
|
|
|
|
# --- reuse-if-published (milestone 313, step 4) ----------------------
|
|
# Does the image the channel tag already points at carry THIS commit's
|
|
# revision? If so the bytes this job would produce are already published
|
|
# and the build is pure waste: the remaining tags get repointed at that
|
|
# existing manifest instead, registry-side, in seconds.
|
|
#
|
|
# Keyed on an `fc.revision` LABEL rather than on a tag of its own
|
|
# (milestone 318 step 3). A tag would be a name minted per build that one
|
|
# thing reads — what rule 145 narrowed against — and would be prunable
|
|
# under the registry's keep_pattern (#3157), silently expiring the cache.
|
|
# A label rides inside a tag that has to exist anyway.
|
|
#
|
|
# An image with no such label reads as a miss and rebuilds. That is the
|
|
# migration, not a fault: labels cannot be backfilled, since the reuse
|
|
# path copies a manifest and config labels are not manifest annotations.
|
|
# Each artifact pays one rebuild, once.
|
|
#
|
|
# This is what stops a push that touched only `agent/` from rebuilding
|
|
# web and ml, and a merge to main from rebuilding what dev already built.
|
|
#
|
|
# The failure direction is deliberate. An inspect that errors for ANY
|
|
# reason — network, auth, a registry hiccup — reads as a miss and the
|
|
# build runs. Only a genuine 200 skips one, so there is no path here
|
|
# that skips a build that was actually needed; the worst case is paying
|
|
# for a build we could have avoided.
|
|
#
|
|
# BASE-IMAGE FRESHNESS: an artifact whose source stops moving stops
|
|
# picking up base-image updates. Milestone 318 removed the argument this
|
|
# used to need rather than answering it — with no version tags there is
|
|
# no immutable name a refresh could contradict, and rule 145 already
|
|
# allows a rebuild with different contents to republish a MOVING tag.
|
|
# So a refresh is just a build. A scheduled channel-only one is tracked
|
|
# separately (#3154); it does not belong in the push path.
|
|
- name: Is this content already published?
|
|
id: reuse
|
|
env:
|
|
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator-ml
|
|
CHANNEL: ${{ steps.tag.outputs.channel }}
|
|
# Empty on a push; the string "true" only from a workflow_dispatch
|
|
# that asked for it. `github.event.inputs` rather than the `inputs`
|
|
# context — release.yml already uses that form, and it is the one
|
|
# this runner is known to evaluate. Read through env rather than
|
|
# interpolated into the run block, same rule as release.yml's TAG.
|
|
FORCE: ${{ github.event.inputs.force_build }}
|
|
# A scheduled refresh has to bypass reuse by construction: it
|
|
# rebuilds the SAME source, so fc.revision always matches and the
|
|
# check would skip every refresh there has ever been.
|
|
run: |
|
|
set -eu
|
|
DERIVED=$(sh scripts/artifacts.sh revision ml)
|
|
echo "revision=$DERIVED" >> "$GITHUB_OUTPUT"
|
|
|
|
# The build clock, pinned to the same commit (#3265). Without it
|
|
# buildkit stamps the image config with the wall clock of the build,
|
|
# so identical layers republish under a new config blob and the
|
|
# channel tag gets a new manifest digest for no reason. Derived from
|
|
# `newest()` like revision and version, so all three name one commit
|
|
# and cannot drift apart.
|
|
echo "epoch=$(sh scripts/artifacts.sh epoch ml)" >> "$GITHUB_OUTPUT"
|
|
|
|
# The moving tag for this channel. Which tag we ask IS the channel —
|
|
# that is why the revision needs no -main/-dev qualifier any more.
|
|
if [ "$CHANNEL" = "main" ]; then T=latest; else T=dev; fi
|
|
echo "channel_ref=$IMAGE:$T" >> "$GITHUB_OUTPUT"
|
|
|
|
# WHERE THE BUILD PUBLISHES, which is not always the channel — and
|
|
# whether the channel then has to be written separately.
|
|
#
|
|
# On a push the build writes the channel tag directly: the bytes came
|
|
# from a commit, and a commit is the thing CI tests. Nothing to hold
|
|
# it behind.
|
|
#
|
|
# On the scheduled refresh it writes a CANDIDATE tag instead. A
|
|
# refresh rebuilds against freshly resolved base images, and the web
|
|
# image's runtime is a line of UNPINNED Debian packages (ffmpeg,
|
|
# libjpeg62-turbo, libpq5, megatools…) re-resolved on every build.
|
|
# Nothing in ci.yml can see that: its lanes run on ci-python:3.14 and
|
|
# install requirements.txt, and a base bump changes neither. So
|
|
# refreshed bytes have to be proven before :latest names them, and
|
|
# proving needs a moment between "built" and "published" to occupy.
|
|
# This is that moment; :latest goes on naming the build that works
|
|
# until something says otherwise.
|
|
#
|
|
# `:refresh-candidate` is one moving ref per image, overwritten in
|
|
# place, holding a build nobody is told to pull — the shape rule 145
|
|
# already allows for :buildcache, not the per-build tag family that
|
|
# milestone 318 withdrew.
|
|
#
|
|
# Both values are decided HERE, beside `hit`, for the reason the
|
|
# force/schedule branch below gives: one step decides what this job
|
|
# does. A promote condition derived independently could disagree with
|
|
# the tag the build actually wrote.
|
|
if [ "${IS_REFRESH:-}" = "true" ]; then
|
|
echo "build_ref=$IMAGE:refresh-candidate" >> "$GITHUB_OUTPUT"
|
|
echo "promote=true" >> "$GITHUB_OUTPUT"
|
|
else
|
|
echo "build_ref=$IMAGE:$T" >> "$GITHUB_OUTPUT"
|
|
echo "promote=false" >> "$GITHUB_OUTPUT"
|
|
fi
|
|
|
|
# Compare VALUES, never exit codes. Measured on buildx v0.36.1
|
|
# (run 4732): a missing key returns an empty string and exits 0, so
|
|
# branching on the exit code would read "no label yet" as success.
|
|
# An unreachable tag also lands here as empty via the `|| echo`.
|
|
# Empty never equals a 12-char revision, so every uncertain case
|
|
# falls through to a build — the safe direction, with no special
|
|
# casing for it.
|
|
#
|
|
# Read the SPECIFIC key. The map also carries whatever the base image
|
|
# set, and `org.opencontainers.image.version` sits right beside ours
|
|
# looking like a plausible answer (it reads 24.04 on the agent).
|
|
PUBLISHED=$(docker buildx imagetools inspect "$IMAGE:$T" \
|
|
--format '{{ index .Image.Config.Labels "fc.revision" }}' \
|
|
2>/dev/null || echo "")
|
|
echo "reuse: $IMAGE:$T carries fc.revision=${PUBLISHED:-<none>}; derived=$DERIVED"
|
|
if [ -z "$PUBLISHED" ] && docker buildx imagetools inspect "$IMAGE:$T" >/dev/null 2>&1; then
|
|
# The tag resolves but carries no readable label. Expected exactly
|
|
# once per artifact, during the migration onto labels. If it recurs
|
|
# every push, something is rewriting the channel tag as a manifest
|
|
# index — see the repoint step's note.
|
|
echo "reuse: NOTE $IMAGE:$T exists but has no readable fc.revision."
|
|
echo "reuse: NOTE Fine once, while migrating. Every push means the"
|
|
echo "reuse: NOTE tag is being index-wrapped and reuse is dead."
|
|
fi
|
|
|
|
# FORCE is checked here rather than in the build step's `if:`, so
|
|
# that one decision drives everything downstream. The repoint step
|
|
# keys off `hit` too, and a force that bypassed only the build would
|
|
# leave the two disagreeing about what just happened.
|
|
if [ "${FORCE:-false}" = "true" ]; then
|
|
echo "hit=false" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: force_build set — building regardless"
|
|
elif [ "${IS_REFRESH:-}" = "true" ]; then
|
|
echo "hit=false" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: scheduled base refresh — building regardless"
|
|
elif [ -n "$PUBLISHED" ] && [ "$PUBLISHED" = "$DERIVED" ]; then
|
|
echo "hit=true" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: already published — skipping the build"
|
|
else
|
|
echo "hit=false" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: not published — building"
|
|
fi
|
|
|
|
- name: Build and push ml image
|
|
if: steps.reuse.outputs.hit != 'true'
|
|
# Read by buildx out of the ENVIRONMENT, not passed as a build-arg —
|
|
# it normalises the image config's `created` field and the history
|
|
# timestamps rather than being consumed by the Dockerfile. See #3265
|
|
# and the reuse step's `epoch` output.
|
|
env:
|
|
SOURCE_DATE_EPOCH: ${{ steps.reuse.outputs.epoch }}
|
|
uses: docker/build-push-action@v5
|
|
with:
|
|
context: .
|
|
file: Dockerfile.ml
|
|
push: true
|
|
# Re-resolve the FROM references against the registry instead of
|
|
# trusting whatever digest the cache was built against. This is the
|
|
# whole mechanism of the scheduled refresh (#3154): if the base tag
|
|
# moved, the FROM layer's cache key changes, every layer above it
|
|
# invalidates, and the image genuinely rebuilds.
|
|
#
|
|
# MEASURED on the first real fire, run 4934 (#3265): when the base
|
|
# did NOT move, the build was ~13s with every content step CACHED —
|
|
# and the channel tag STILL got a new manifest digest, because
|
|
# buildkit stamps a fresh image config per run and republishes the
|
|
# identical layers under it. All three images moved that way on
|
|
# 2026-08-30 with nothing whatsoever changed in them.
|
|
#
|
|
# SOURCE_DATE_EPOCH (below) is the fix: pinned to the commit the
|
|
# content came from, the config is byte-identical across runs, so
|
|
# the manifest digest is too and the push is a registry no-op. A
|
|
# digest change means the content changed again, which is the only
|
|
# thing a digest is any use for.
|
|
#
|
|
# What `pull` does NOT catch either: a Debian package update inside
|
|
# the `apt-get install` layer while the base tag itself stands
|
|
# still. The official python/cuda images rebuild with those updates
|
|
# baked in, so this is a lag rather than a hole; closing it needs
|
|
# `no-cache: true`, which is a much larger version of the same
|
|
# churn #3265 is about.
|
|
#
|
|
# Only on the schedule. An ordinary push wants the cached base.
|
|
pull: ${{ env.IS_REFRESH == 'true' }}
|
|
# ONE tag, the channel's. Every other tag is written by the step
|
|
# below, registry-side. buildx here pushes the first tag to the
|
|
# registry and then re-pushes the rest through the DOCKER driver,
|
|
# out of a local image store a registry-direct build never filled —
|
|
# #3190, which cost `main` its :c-<sha> on 2026-08-29 while :latest
|
|
# published perfectly well.
|
|
tags: ${{ steps.reuse.outputs.build_ref }}
|
|
# The reuse key. Read back off the channel tag on the next push to
|
|
# decide whether that push needs to build at all, so this is not
|
|
# decoration — an unstamped image is one that will always rebuild.
|
|
labels: |
|
|
fc.revision=${{ steps.reuse.outputs.revision }}
|
|
# LOAD-BEARING, not a preference. On the default docker driver these
|
|
# were no-ops; on the docker-container driver above,
|
|
# build-push-action@v5 defaults provenance to TRUE when pushing.
|
|
# Provenance attaches an attestation manifest, which makes the pushed
|
|
# tag a manifest INDEX — and `.Image.Config.Labels` does not resolve
|
|
# through an index.
|
|
#
|
|
# The label directly above IS the reuse key. Wrap the channel tag in
|
|
# an index and the next push reads fc.revision=<none>, misses, and
|
|
# rebuilds. Then so does the one after that, forever. Nothing fails,
|
|
# nothing goes red, and the only symptom is the bill. That is #3183
|
|
# arriving through a different door, and note #3127 §4 records the
|
|
# same shape for `platforms:`.
|
|
provenance: false
|
|
sbom: false
|
|
# The ONLY cache this driver can have. `docker-container` gets a
|
|
# FRESH buildkit instance per job, so unlike the default docker
|
|
# driver it has no local layer store to fall back on — measured on
|
|
# run 4896, the first builds after the driver change: web 3m44s
|
|
# (was 2m23s), ml 3m49s (was 3m20s), agent 11m12s (was 9m26s). The
|
|
# driver change ALONE is a regression; this is the other half of it.
|
|
#
|
|
# mode=max so intermediate stages cache too. web's frontend-builder
|
|
# stage and the agent's two ~150s pip layers are the whole cost, and
|
|
# they are exactly what a min-mode cache would drop.
|
|
#
|
|
# A `:buildcache` tag is NOT the withdrawn tag scheme coming back.
|
|
# Rule 145 narrowed against names NOTHING reads; this one is read by
|
|
# every build that runs, is one moving ref per image rather than one
|
|
# per build, holds cache blobs rather than a shippable artifact, and
|
|
# is overwritten in place rather than accumulating. It is closer to
|
|
# :dev than to the :2026.8.28 tags milestone 318 deleted. (#3114.)
|
|
cache-from: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator-ml:buildcache
|
|
cache-to: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator-ml:buildcache,mode=max
|
|
|
|
# Point the channel tag at the candidate the refresh just built.
|
|
#
|
|
# Unconditional TODAY, so this milestone never leaves the refresh in a
|
|
# state where it builds and publishes nothing. Step 4 wraps it in the
|
|
# smoke suite's verdict; until then the scheduled path behaves exactly
|
|
# as it did, just via two operations instead of one.
|
|
#
|
|
# NOT `imagetools create`. That wraps its source in an INDEX, and an
|
|
# indexed channel tag is the one thing this pipeline cannot survive:
|
|
# `.Image.Config.Labels` does not resolve through an index, so the
|
|
# fc.revision the reuse check reads off the channel tag would come back
|
|
# empty, every subsequent push would miss and rebuild, and nothing would
|
|
# go red. That is #3183, observed on run 4751 — reuse worked exactly once
|
|
# and the only symptom was the bill. The repoint step below excludes its
|
|
# own source tag for precisely this reason; a promote that re-introduced
|
|
# the wrap through a different door would undo that care.
|
|
#
|
|
# A manifest PUT is what "make this tag name that image" means at the
|
|
# registry level: the same bytes under the same media type, so the digest
|
|
# is identical, the media type is preserved, and no layer moves.
|
|
- name: Promote the refresh candidate to the channel
|
|
if: steps.reuse.outputs.promote == 'true'
|
|
env:
|
|
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator-ml
|
|
CHANNEL_REF: ${{ steps.reuse.outputs.channel_ref }}
|
|
TOKEN: ${{ secrets.RELEASE_TOKEN }}
|
|
ACTOR: ${{ github.actor }}
|
|
run: |
|
|
set -eu
|
|
REPO=${IMAGE#git.fabledsword.com/}
|
|
TAG=${CHANNEL_REF##*:}
|
|
|
|
# Registry auth is its own token exchange — the `docker login` above
|
|
# authenticates the docker client, not curl. Deadline on every call
|
|
# (rule 156): a registry that stops answering must fail this step,
|
|
# not hang the weekly refresh until the job times out.
|
|
BEARER=$(curl -fsS --max-time 30 -u "$ACTOR:$TOKEN" \
|
|
"https://git.fabledsword.com/v2/token?scope=repository:$REPO:pull,push&service=git.fabledsword.com" \
|
|
| python3 -c 'import sys,json; print(json.load(sys.stdin)["token"])')
|
|
|
|
# Ask for the image manifest media types ONLY. Offering the index
|
|
# types too would let the registry hand back an index if one ever
|
|
# existed at this tag, and we would faithfully copy the thing we are
|
|
# trying not to create.
|
|
ACCEPT='application/vnd.oci.image.manifest.v1+json, application/vnd.docker.distribution.manifest.v2+json'
|
|
CT=$(curl -fsS --max-time 60 -o manifest.json -D headers.txt \
|
|
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
|
|
"https://git.fabledsword.com/v2/$REPO/manifests/refresh-candidate" \
|
|
&& tr -d '\r' < headers.txt | awk -F': ' '/^[Cc]ontent-[Tt]ype:/{print $2}')
|
|
test -n "$CT"
|
|
SRC_DIGEST=$(tr -d '\r' < headers.txt | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
|
|
echo "promote: candidate is $SRC_DIGEST ($CT)"
|
|
|
|
curl -fsS --max-time 120 -X PUT \
|
|
-H "Authorization: Bearer $BEARER" -H "Content-Type: $CT" \
|
|
--data-binary @manifest.json \
|
|
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG"
|
|
|
|
# Read it back. A PUT that returned 2xx but landed something else is
|
|
# exactly the silent-and-plausible failure this pipeline keeps
|
|
# producing, and the check costs one request.
|
|
NOW=$(curl -fsS --max-time 30 -o /dev/null -D - \
|
|
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
|
|
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG" \
|
|
| tr -d '\r' | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
|
|
if [ "$NOW" != "$SRC_DIGEST" ]; then
|
|
echo "promote: $IMAGE:$TAG is $NOW, expected $SRC_DIGEST" >&2
|
|
exit 1
|
|
fi
|
|
echo "promote: $IMAGE:$TAG now names $NOW"
|
|
|
|
# Every tag but the channel's own is written HERE, registry-side,
|
|
# whether or not a build ran. Each -t becomes another reference to the
|
|
# SAME manifest the channel tag holds, so :c-<sha> is byte-identical to
|
|
# what is published rather than a lookalike rebuild.
|
|
#
|
|
# Owning the build path too is #3190's fix, not a tidy-up:
|
|
#
|
|
# #27 pushing …/fabledcurator:latest DONE 15.8s
|
|
# #28 pushing …/fabledcurator:c-0e15c44 with docker
|
|
# #28 ERROR: tag does not exist: …:c-0e15c44
|
|
#
|
|
# Intermittent — build-ml made the identical two-tag push seconds later
|
|
# and succeeded — and worse than it looks. `:latest` had already
|
|
# published, so production was correct while the immutable rollback tag
|
|
# rule 145 requires of every main push simply did not exist. Nothing but
|
|
# the red job would ever have noticed: a missing :c-<sha> has no
|
|
# consumer that fails, so it surfaces when somebody needs to roll back.
|
|
#
|
|
# `imagetools create` is a registry-side manifest copy — no layer
|
|
# transfer, no local daemon, nothing that can be absent. The reuse case
|
|
# has always gone this way, so this puts the build case on the code that
|
|
# was already proven rather than on a second path.
|
|
#
|
|
# Running on every path also keeps family rule 146 true: a rolling
|
|
# channel refreshes itself, so skipping a build must never leave :dev or
|
|
# :latest pointing at something older than the commit just pushed.
|
|
#
|
|
# The cost, accepted knowingly: `imagetools create` wraps its source in
|
|
# an index, so :c-<sha> becomes an index and fc.revision does not
|
|
# resolve through it. Nothing reads that label off :c-<sha> — the reuse
|
|
# check only ever inspects the CHANNEL tag — and the index names the
|
|
# same manifest, so a pull is byte-identical. The reuse path already
|
|
# produced :c-<sha> this way; this only makes it uniform.
|
|
- name: Write the remaining tags from the published image
|
|
env:
|
|
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator-ml
|
|
SOURCE: ${{ steps.reuse.outputs.channel_ref }}
|
|
TAGS: ${{ steps.tag.outputs.tags }}
|
|
run: |
|
|
set -euf
|
|
# The source tag is EXCLUDED from the targets, and that is load-
|
|
# bearing rather than an optimisation.
|
|
#
|
|
# `imagetools create` wraps the source manifest in an INDEX. Point it
|
|
# at the channel tag with that same tag as a target and the tag stops
|
|
# being a plain image — after which `.Image.Config.Labels` no longer
|
|
# resolves through it and the fc.revision label reads as absent. The
|
|
# next push then misses and rebuilds, so reuse worked exactly once
|
|
# and every subsequent push paid full price. Observed on run 4751:
|
|
# ml:dev reported fc.revision=<none> one push after run 4749 had read
|
|
# a7e626a67a79 off it. Nothing failed; the savings just evaporated.
|
|
#
|
|
# Excluding the source means the channel tag is only ever written
|
|
# by a real build, so it stays a plain image and stays readable.
|
|
# On dev that leaves nothing to do either way: the build pushed :dev
|
|
# itself, or the hit established it was already right. On main it
|
|
# leaves :c-<sha>, which rule 145 requires of every main push whether
|
|
# or not a build ran.
|
|
#
|
|
# steps.tag emits ONE comma-separated list; imagetools wants a -t per
|
|
# ref. (That list used to feed docker/build-push-action directly —
|
|
# which is exactly what #3190 made unsafe.)
|
|
ARGS=""
|
|
IFS=,
|
|
for t in $TAGS; do
|
|
[ "$t" = "$SOURCE" ] && continue
|
|
ARGS="$ARGS -t $t"
|
|
done
|
|
unset IFS
|
|
if [ -z "$ARGS" ]; then
|
|
echo "repoint: $SOURCE is the only tag for this channel and"
|
|
echo "repoint: already holds this revision — nothing to write."
|
|
exit 0
|
|
fi
|
|
# shellcheck disable=SC2086
|
|
docker buildx imagetools create $ARGS "$SOURCE"
|
|
echo "repointed from $SOURCE:$ARGS"
|
|
|
|
# The desktop GPU agent (#114) — published so the operator pulls + runs it on
|
|
# the GPU machine instead of building locally. Independent of web/ml (its own
|
|
# CUDA + onnxruntime-gpu image, context = agent/). Same tag cadence.
|
|
build-agent:
|
|
runs-on: python-ci
|
|
container:
|
|
image: git.fabledsword.com/bvandeusen/ci-python:3.14
|
|
steps:
|
|
- uses: actions/checkout@v4
|
|
with:
|
|
# Not the triggering ref — see the `env:` block at the top. On a
|
|
# scheduled refresh this is `main`; on everything else it is the ref
|
|
# that fired, so this is a no-op on every ordinary path.
|
|
ref: ${{ env.BUILD_REF }}
|
|
# Full history: this job derives its artifact's version from the
|
|
# commit its shipped files last changed in (milestone 313). A
|
|
# depth-1 clone cannot see that commit — it either derives a wrong,
|
|
# too-low value or finds nothing at all, and neither is a failure
|
|
# the build would otherwise notice.
|
|
fetch-depth: 0
|
|
|
|
# See sign-extension's copy for why this guard exists.
|
|
- name: Guard — a scheduled run must have checked out main
|
|
if: env.IS_REFRESH == 'true'
|
|
run: |
|
|
set -eu
|
|
BRANCH=$(git rev-parse --abbrev-ref HEAD)
|
|
echo "schedule: HEAD is $BRANCH ($(git rev-parse --short HEAD))"
|
|
if [ "$BRANCH" != "main" ]; then
|
|
echo "schedule: expected main, got '$BRANCH'." >&2
|
|
echo "schedule: BUILD_REF was not honoured by the runner." >&2
|
|
echo "schedule: refusing to publish a channel tag from it." >&2
|
|
exit 1
|
|
fi
|
|
|
|
# --- derived values, one line (milestone 313) ------------------------
|
|
# These stopped being shadow output at step 3. `revision` decides
|
|
# whether the build below runs at all and `version` is what the image
|
|
# reports about itself; the load-bearing steps each print only the one
|
|
# they use, so this is the only place the pair appears together. When a
|
|
# build is skipped, this is the line that says what the commit derived.
|
|
#
|
|
# Still diagnostic, so it still must not fail the build — no `set -e`,
|
|
# and every derivation falls back to UNAVAILABLE. A broken echo must
|
|
# never be the reason an image does not ship.
|
|
#
|
|
# What it should say:
|
|
# * a push touching only agent/ moves the agent and leaves web and ml
|
|
# STILL. If web moves, its path set is too wide.
|
|
# * a push touching only docs moves nothing.
|
|
# * a push touching the extension moves the extension AND web, since
|
|
# web bakes in the XPI. If web does not move, its set is too narrow:
|
|
# the reuse check hits, and the channel serves a web image bundling
|
|
# the PREVIOUS XPI while the freshly signed one is orphaned (#3156).
|
|
# * dev and main derive the same values for the same source.
|
|
- name: Report the derived artifact version
|
|
run: |
|
|
set -u
|
|
A=agent
|
|
V=$(sh scripts/artifacts.sh version "$A" 2>&1 || echo UNAVAILABLE)
|
|
R=$(sh scripts/artifacts.sh revision "$A" 2>&1 || echo UNAVAILABLE)
|
|
echo "derived: artifact=$A version=$V revision=$R sha=$GITHUB_SHA"
|
|
|
|
- name: Determine tag
|
|
id: tag
|
|
run: |
|
|
SHORT_SHA=$(printf '%s' "$GITHUB_SHA" | cut -c1-7)
|
|
# Mirrors build-web's tag list and its schedule handling; see
|
|
# the comments there.
|
|
if [ "${IS_REFRESH:-}" = "true" ]; then
|
|
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator-agent:latest" >> "$GITHUB_OUTPUT"
|
|
echo "channel=main" >> "$GITHUB_OUTPUT"
|
|
elif [ "${GITHUB_REF##*/}" = "main" ]; then
|
|
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator-agent:latest,git.fabledsword.com/bvandeusen/fabledcurator-agent:c-${SHORT_SHA}" >> "$GITHUB_OUTPUT"
|
|
echo "channel=main" >> "$GITHUB_OUTPUT"
|
|
else
|
|
echo "tags=git.fabledsword.com/bvandeusen/fabledcurator-agent:dev" >> "$GITHUB_OUTPUT"
|
|
echo "channel=dev" >> "$GITHUB_OUTPUT"
|
|
fi
|
|
|
|
# Shell step rather than docker/login-action — see build-web's note on
|
|
# the shared action-cache race (#3118).
|
|
- name: Login to Forgejo registry
|
|
env:
|
|
TOKEN: ${{ secrets.RELEASE_TOKEN }}
|
|
ACTOR: ${{ github.actor }}
|
|
run: echo "$TOKEN" | docker login git.fabledsword.com -u "$ACTOR" --password-stdin
|
|
|
|
# A REAL buildx builder, not the default `docker` driver (#3114, #3190).
|
|
#
|
|
# The default driver builds through the local dockerd. It cannot export a
|
|
# registry cache at all — which is why the agent rebuilds a ~6.3 GB CUDA
|
|
# + torch image from scratch whenever the runner's local cache is cold,
|
|
# measured at 9m26s against 7s warm. It is also #3190's leading suspect:
|
|
# after a registry-direct push it resolves image metadata against a local
|
|
# store the push never filled, and reports `No such image` on an image
|
|
# that published perfectly well three seconds earlier.
|
|
#
|
|
# These jobs run INSIDE a container against a mounted docker socket, so
|
|
# the buildkit container this starts is a SIBLING of the job container,
|
|
# not a child. That works over the socket mount; it had never been tried
|
|
# here before milestone 326 step 1.
|
|
- name: Set up buildx
|
|
uses: docker/setup-buildx-action@v3
|
|
|
|
# --- reuse-if-published (milestone 313, step 4) ----------------------
|
|
# Does the image the channel tag already points at carry THIS commit's
|
|
# revision? If so the bytes this job would produce are already published
|
|
# and the build is pure waste: the remaining tags get repointed at that
|
|
# existing manifest instead, registry-side, in seconds.
|
|
#
|
|
# Keyed on an `fc.revision` LABEL rather than on a tag of its own
|
|
# (milestone 318 step 3). A tag would be a name minted per build that one
|
|
# thing reads — what rule 145 narrowed against — and would be prunable
|
|
# under the registry's keep_pattern (#3157), silently expiring the cache.
|
|
# A label rides inside a tag that has to exist anyway.
|
|
#
|
|
# An image with no such label reads as a miss and rebuilds. That is the
|
|
# migration, not a fault: labels cannot be backfilled, since the reuse
|
|
# path copies a manifest and config labels are not manifest annotations.
|
|
# Each artifact pays one rebuild, once.
|
|
#
|
|
# This is what stops a push that touched only `agent/` from rebuilding
|
|
# web and ml, and a merge to main from rebuilding what dev already built.
|
|
#
|
|
# The failure direction is deliberate. An inspect that errors for ANY
|
|
# reason — network, auth, a registry hiccup — reads as a miss and the
|
|
# build runs. Only a genuine 200 skips one, so there is no path here
|
|
# that skips a build that was actually needed; the worst case is paying
|
|
# for a build we could have avoided.
|
|
#
|
|
# BASE-IMAGE FRESHNESS: an artifact whose source stops moving stops
|
|
# picking up base-image updates. Milestone 318 removed the argument this
|
|
# used to need rather than answering it — with no version tags there is
|
|
# no immutable name a refresh could contradict, and rule 145 already
|
|
# allows a rebuild with different contents to republish a MOVING tag.
|
|
# So a refresh is just a build. A scheduled channel-only one is tracked
|
|
# separately (#3154); it does not belong in the push path.
|
|
- name: Is this content already published?
|
|
id: reuse
|
|
env:
|
|
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator-agent
|
|
CHANNEL: ${{ steps.tag.outputs.channel }}
|
|
# Empty on a push; the string "true" only from a workflow_dispatch
|
|
# that asked for it. `github.event.inputs` rather than the `inputs`
|
|
# context — release.yml already uses that form, and it is the one
|
|
# this runner is known to evaluate. Read through env rather than
|
|
# interpolated into the run block, same rule as release.yml's TAG.
|
|
FORCE: ${{ github.event.inputs.force_build }}
|
|
# A scheduled refresh has to bypass reuse by construction: it
|
|
# rebuilds the SAME source, so fc.revision always matches and the
|
|
# check would skip every refresh there has ever been.
|
|
run: |
|
|
set -eu
|
|
DERIVED=$(sh scripts/artifacts.sh revision agent)
|
|
echo "revision=$DERIVED" >> "$GITHUB_OUTPUT"
|
|
|
|
# The build clock, pinned to the same commit (#3265). Without it
|
|
# buildkit stamps the image config with the wall clock of the build,
|
|
# so identical layers republish under a new config blob and the
|
|
# channel tag gets a new manifest digest for no reason. Derived from
|
|
# `newest()` like revision and version, so all three name one commit
|
|
# and cannot drift apart.
|
|
echo "epoch=$(sh scripts/artifacts.sh epoch agent)" >> "$GITHUB_OUTPUT"
|
|
|
|
# The moving tag for this channel. Which tag we ask IS the channel —
|
|
# that is why the revision needs no -main/-dev qualifier any more.
|
|
if [ "$CHANNEL" = "main" ]; then T=latest; else T=dev; fi
|
|
echo "channel_ref=$IMAGE:$T" >> "$GITHUB_OUTPUT"
|
|
|
|
# WHERE THE BUILD PUBLISHES, which is not always the channel — and
|
|
# whether the channel then has to be written separately.
|
|
#
|
|
# On a push the build writes the channel tag directly: the bytes came
|
|
# from a commit, and a commit is the thing CI tests. Nothing to hold
|
|
# it behind.
|
|
#
|
|
# On the scheduled refresh it writes a CANDIDATE tag instead. A
|
|
# refresh rebuilds against freshly resolved base images, and the web
|
|
# image's runtime is a line of UNPINNED Debian packages (ffmpeg,
|
|
# libjpeg62-turbo, libpq5, megatools…) re-resolved on every build.
|
|
# Nothing in ci.yml can see that: its lanes run on ci-python:3.14 and
|
|
# install requirements.txt, and a base bump changes neither. So
|
|
# refreshed bytes have to be proven before :latest names them, and
|
|
# proving needs a moment between "built" and "published" to occupy.
|
|
# This is that moment; :latest goes on naming the build that works
|
|
# until something says otherwise.
|
|
#
|
|
# `:refresh-candidate` is one moving ref per image, overwritten in
|
|
# place, holding a build nobody is told to pull — the shape rule 145
|
|
# already allows for :buildcache, not the per-build tag family that
|
|
# milestone 318 withdrew.
|
|
#
|
|
# Both values are decided HERE, beside `hit`, for the reason the
|
|
# force/schedule branch below gives: one step decides what this job
|
|
# does. A promote condition derived independently could disagree with
|
|
# the tag the build actually wrote.
|
|
if [ "${IS_REFRESH:-}" = "true" ]; then
|
|
echo "build_ref=$IMAGE:refresh-candidate" >> "$GITHUB_OUTPUT"
|
|
echo "promote=true" >> "$GITHUB_OUTPUT"
|
|
else
|
|
echo "build_ref=$IMAGE:$T" >> "$GITHUB_OUTPUT"
|
|
echo "promote=false" >> "$GITHUB_OUTPUT"
|
|
fi
|
|
|
|
# Compare VALUES, never exit codes. Measured on buildx v0.36.1
|
|
# (run 4732): a missing key returns an empty string and exits 0, so
|
|
# branching on the exit code would read "no label yet" as success.
|
|
# An unreachable tag also lands here as empty via the `|| echo`.
|
|
# Empty never equals a 12-char revision, so every uncertain case
|
|
# falls through to a build — the safe direction, with no special
|
|
# casing for it.
|
|
#
|
|
# Read the SPECIFIC key. The map also carries whatever the base image
|
|
# set, and `org.opencontainers.image.version` sits right beside ours
|
|
# looking like a plausible answer (it reads 24.04 on the agent).
|
|
PUBLISHED=$(docker buildx imagetools inspect "$IMAGE:$T" \
|
|
--format '{{ index .Image.Config.Labels "fc.revision" }}' \
|
|
2>/dev/null || echo "")
|
|
echo "reuse: $IMAGE:$T carries fc.revision=${PUBLISHED:-<none>}; derived=$DERIVED"
|
|
if [ -z "$PUBLISHED" ] && docker buildx imagetools inspect "$IMAGE:$T" >/dev/null 2>&1; then
|
|
# The tag resolves but carries no readable label. Expected exactly
|
|
# once per artifact, during the migration onto labels. If it recurs
|
|
# every push, something is rewriting the channel tag as a manifest
|
|
# index — see the repoint step's note.
|
|
echo "reuse: NOTE $IMAGE:$T exists but has no readable fc.revision."
|
|
echo "reuse: NOTE Fine once, while migrating. Every push means the"
|
|
echo "reuse: NOTE tag is being index-wrapped and reuse is dead."
|
|
fi
|
|
|
|
# FORCE is checked here rather than in the build step's `if:`, so
|
|
# that one decision drives everything downstream. The repoint step
|
|
# keys off `hit` too, and a force that bypassed only the build would
|
|
# leave the two disagreeing about what just happened.
|
|
if [ "${FORCE:-false}" = "true" ]; then
|
|
echo "hit=false" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: force_build set — building regardless"
|
|
elif [ "${IS_REFRESH:-}" = "true" ]; then
|
|
echo "hit=false" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: scheduled base refresh — building regardless"
|
|
elif [ -n "$PUBLISHED" ] && [ "$PUBLISHED" = "$DERIVED" ]; then
|
|
echo "hit=true" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: already published — skipping the build"
|
|
else
|
|
echo "hit=false" >> "$GITHUB_OUTPUT"
|
|
echo "reuse: not published — building"
|
|
fi
|
|
|
|
- name: Build and push agent image
|
|
if: steps.reuse.outputs.hit != 'true'
|
|
# Read by buildx out of the ENVIRONMENT, not passed as a build-arg —
|
|
# it normalises the image config's `created` field and the history
|
|
# timestamps rather than being consumed by the Dockerfile. See #3265
|
|
# and the reuse step's `epoch` output.
|
|
env:
|
|
SOURCE_DATE_EPOCH: ${{ steps.reuse.outputs.epoch }}
|
|
uses: docker/build-push-action@v5
|
|
with:
|
|
context: agent
|
|
file: agent/Dockerfile
|
|
push: true
|
|
# Re-resolve the FROM references against the registry instead of
|
|
# trusting whatever digest the cache was built against. This is the
|
|
# whole mechanism of the scheduled refresh (#3154): if the base tag
|
|
# moved, the FROM layer's cache key changes, every layer above it
|
|
# invalidates, and the image genuinely rebuilds.
|
|
#
|
|
# MEASURED on the first real fire, run 4934 (#3265): when the base
|
|
# did NOT move, the build was ~13s with every content step CACHED —
|
|
# and the channel tag STILL got a new manifest digest, because
|
|
# buildkit stamps a fresh image config per run and republishes the
|
|
# identical layers under it. All three images moved that way on
|
|
# 2026-08-30 with nothing whatsoever changed in them.
|
|
#
|
|
# SOURCE_DATE_EPOCH (below) is the fix: pinned to the commit the
|
|
# content came from, the config is byte-identical across runs, so
|
|
# the manifest digest is too and the push is a registry no-op. A
|
|
# digest change means the content changed again, which is the only
|
|
# thing a digest is any use for.
|
|
#
|
|
# What `pull` does NOT catch either: a Debian package update inside
|
|
# the `apt-get install` layer while the base tag itself stands
|
|
# still. The official python/cuda images rebuild with those updates
|
|
# baked in, so this is a lag rather than a hole; closing it needs
|
|
# `no-cache: true`, which is a much larger version of the same
|
|
# churn #3265 is about.
|
|
#
|
|
# Only on the schedule. An ordinary push wants the cached base.
|
|
pull: ${{ env.IS_REFRESH == 'true' }}
|
|
# ONE tag, the channel's. Every other tag is written by the step
|
|
# below, registry-side. buildx here pushes the first tag to the
|
|
# registry and then re-pushes the rest through the DOCKER driver,
|
|
# out of a local image store a registry-direct build never filled —
|
|
# #3190, which cost `main` its :c-<sha> on 2026-08-29 while :latest
|
|
# published perfectly well.
|
|
tags: ${{ steps.reuse.outputs.build_ref }}
|
|
# The reuse key. Read back off the channel tag on the next push to
|
|
# decide whether that push needs to build at all, so this is not
|
|
# decoration — an unstamped image is one that will always rebuild.
|
|
labels: |
|
|
fc.revision=${{ steps.reuse.outputs.revision }}
|
|
# LOAD-BEARING, not a preference. On the default docker driver these
|
|
# were no-ops; on the docker-container driver above,
|
|
# build-push-action@v5 defaults provenance to TRUE when pushing.
|
|
# Provenance attaches an attestation manifest, which makes the pushed
|
|
# tag a manifest INDEX — and `.Image.Config.Labels` does not resolve
|
|
# through an index.
|
|
#
|
|
# The label directly above IS the reuse key. Wrap the channel tag in
|
|
# an index and the next push reads fc.revision=<none>, misses, and
|
|
# rebuilds. Then so does the one after that, forever. Nothing fails,
|
|
# nothing goes red, and the only symptom is the bill. That is #3183
|
|
# arriving through a different door, and note #3127 §4 records the
|
|
# same shape for `platforms:`.
|
|
provenance: false
|
|
sbom: false
|
|
# The ONLY cache this driver can have. `docker-container` gets a
|
|
# FRESH buildkit instance per job, so unlike the default docker
|
|
# driver it has no local layer store to fall back on — measured on
|
|
# run 4896, the first builds after the driver change: web 3m44s
|
|
# (was 2m23s), ml 3m49s (was 3m20s), agent 11m12s (was 9m26s). The
|
|
# driver change ALONE is a regression; this is the other half of it.
|
|
#
|
|
# mode=max so intermediate stages cache too. web's frontend-builder
|
|
# stage and the agent's two ~150s pip layers are the whole cost, and
|
|
# they are exactly what a min-mode cache would drop.
|
|
#
|
|
# A `:buildcache` tag is NOT the withdrawn tag scheme coming back.
|
|
# Rule 145 narrowed against names NOTHING reads; this one is read by
|
|
# every build that runs, is one moving ref per image rather than one
|
|
# per build, holds cache blobs rather than a shippable artifact, and
|
|
# is overwritten in place rather than accumulating. It is closer to
|
|
# :dev than to the :2026.8.28 tags milestone 318 deleted. (#3114.)
|
|
cache-from: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator-agent:buildcache
|
|
cache-to: type=registry,ref=git.fabledsword.com/bvandeusen/fabledcurator-agent:buildcache,mode=max
|
|
|
|
# Point the channel tag at the candidate the refresh just built.
|
|
#
|
|
# Unconditional TODAY, so this milestone never leaves the refresh in a
|
|
# state where it builds and publishes nothing. Step 4 wraps it in the
|
|
# smoke suite's verdict; until then the scheduled path behaves exactly
|
|
# as it did, just via two operations instead of one.
|
|
#
|
|
# NOT `imagetools create`. That wraps its source in an INDEX, and an
|
|
# indexed channel tag is the one thing this pipeline cannot survive:
|
|
# `.Image.Config.Labels` does not resolve through an index, so the
|
|
# fc.revision the reuse check reads off the channel tag would come back
|
|
# empty, every subsequent push would miss and rebuild, and nothing would
|
|
# go red. That is #3183, observed on run 4751 — reuse worked exactly once
|
|
# and the only symptom was the bill. The repoint step below excludes its
|
|
# own source tag for precisely this reason; a promote that re-introduced
|
|
# the wrap through a different door would undo that care.
|
|
#
|
|
# A manifest PUT is what "make this tag name that image" means at the
|
|
# registry level: the same bytes under the same media type, so the digest
|
|
# is identical, the media type is preserved, and no layer moves.
|
|
- name: Promote the refresh candidate to the channel
|
|
if: steps.reuse.outputs.promote == 'true'
|
|
env:
|
|
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator-agent
|
|
CHANNEL_REF: ${{ steps.reuse.outputs.channel_ref }}
|
|
TOKEN: ${{ secrets.RELEASE_TOKEN }}
|
|
ACTOR: ${{ github.actor }}
|
|
run: |
|
|
set -eu
|
|
REPO=${IMAGE#git.fabledsword.com/}
|
|
TAG=${CHANNEL_REF##*:}
|
|
|
|
# Registry auth is its own token exchange — the `docker login` above
|
|
# authenticates the docker client, not curl. Deadline on every call
|
|
# (rule 156): a registry that stops answering must fail this step,
|
|
# not hang the weekly refresh until the job times out.
|
|
BEARER=$(curl -fsS --max-time 30 -u "$ACTOR:$TOKEN" \
|
|
"https://git.fabledsword.com/v2/token?scope=repository:$REPO:pull,push&service=git.fabledsword.com" \
|
|
| python3 -c 'import sys,json; print(json.load(sys.stdin)["token"])')
|
|
|
|
# Ask for the image manifest media types ONLY. Offering the index
|
|
# types too would let the registry hand back an index if one ever
|
|
# existed at this tag, and we would faithfully copy the thing we are
|
|
# trying not to create.
|
|
ACCEPT='application/vnd.oci.image.manifest.v1+json, application/vnd.docker.distribution.manifest.v2+json'
|
|
CT=$(curl -fsS --max-time 60 -o manifest.json -D headers.txt \
|
|
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
|
|
"https://git.fabledsword.com/v2/$REPO/manifests/refresh-candidate" \
|
|
&& tr -d '\r' < headers.txt | awk -F': ' '/^[Cc]ontent-[Tt]ype:/{print $2}')
|
|
test -n "$CT"
|
|
SRC_DIGEST=$(tr -d '\r' < headers.txt | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
|
|
echo "promote: candidate is $SRC_DIGEST ($CT)"
|
|
|
|
curl -fsS --max-time 120 -X PUT \
|
|
-H "Authorization: Bearer $BEARER" -H "Content-Type: $CT" \
|
|
--data-binary @manifest.json \
|
|
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG"
|
|
|
|
# Read it back. A PUT that returned 2xx but landed something else is
|
|
# exactly the silent-and-plausible failure this pipeline keeps
|
|
# producing, and the check costs one request.
|
|
NOW=$(curl -fsS --max-time 30 -o /dev/null -D - \
|
|
-H "Authorization: Bearer $BEARER" -H "Accept: $ACCEPT" \
|
|
"https://git.fabledsword.com/v2/$REPO/manifests/$TAG" \
|
|
| tr -d '\r' | awk -F': ' '/^[Dd]ocker-[Cc]ontent-[Dd]igest:/{print $2}')
|
|
if [ "$NOW" != "$SRC_DIGEST" ]; then
|
|
echo "promote: $IMAGE:$TAG is $NOW, expected $SRC_DIGEST" >&2
|
|
exit 1
|
|
fi
|
|
echo "promote: $IMAGE:$TAG now names $NOW"
|
|
|
|
# Every tag but the channel's own is written HERE, registry-side,
|
|
# whether or not a build ran. Each -t becomes another reference to the
|
|
# SAME manifest the channel tag holds, so :c-<sha> is byte-identical to
|
|
# what is published rather than a lookalike rebuild.
|
|
#
|
|
# Owning the build path too is #3190's fix, not a tidy-up:
|
|
#
|
|
# #27 pushing …/fabledcurator:latest DONE 15.8s
|
|
# #28 pushing …/fabledcurator:c-0e15c44 with docker
|
|
# #28 ERROR: tag does not exist: …:c-0e15c44
|
|
#
|
|
# Intermittent — build-ml made the identical two-tag push seconds later
|
|
# and succeeded — and worse than it looks. `:latest` had already
|
|
# published, so production was correct while the immutable rollback tag
|
|
# rule 145 requires of every main push simply did not exist. Nothing but
|
|
# the red job would ever have noticed: a missing :c-<sha> has no
|
|
# consumer that fails, so it surfaces when somebody needs to roll back.
|
|
#
|
|
# `imagetools create` is a registry-side manifest copy — no layer
|
|
# transfer, no local daemon, nothing that can be absent. The reuse case
|
|
# has always gone this way, so this puts the build case on the code that
|
|
# was already proven rather than on a second path.
|
|
#
|
|
# Running on every path also keeps family rule 146 true: a rolling
|
|
# channel refreshes itself, so skipping a build must never leave :dev or
|
|
# :latest pointing at something older than the commit just pushed.
|
|
#
|
|
# The cost, accepted knowingly: `imagetools create` wraps its source in
|
|
# an index, so :c-<sha> becomes an index and fc.revision does not
|
|
# resolve through it. Nothing reads that label off :c-<sha> — the reuse
|
|
# check only ever inspects the CHANNEL tag — and the index names the
|
|
# same manifest, so a pull is byte-identical. The reuse path already
|
|
# produced :c-<sha> this way; this only makes it uniform.
|
|
- name: Write the remaining tags from the published image
|
|
env:
|
|
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator-agent
|
|
SOURCE: ${{ steps.reuse.outputs.channel_ref }}
|
|
TAGS: ${{ steps.tag.outputs.tags }}
|
|
run: |
|
|
set -euf
|
|
# The source tag is EXCLUDED from the targets, and that is load-
|
|
# bearing rather than an optimisation.
|
|
#
|
|
# `imagetools create` wraps the source manifest in an INDEX. Point it
|
|
# at the channel tag with that same tag as a target and the tag stops
|
|
# being a plain image — after which `.Image.Config.Labels` no longer
|
|
# resolves through it and the fc.revision label reads as absent. The
|
|
# next push then misses and rebuilds, so reuse worked exactly once
|
|
# and every subsequent push paid full price. Observed on run 4751:
|
|
# ml:dev reported fc.revision=<none> one push after run 4749 had read
|
|
# a7e626a67a79 off it. Nothing failed; the savings just evaporated.
|
|
#
|
|
# Excluding the source means the channel tag is only ever written
|
|
# by a real build, so it stays a plain image and stays readable.
|
|
# On dev that leaves nothing to do either way: the build pushed :dev
|
|
# itself, or the hit established it was already right. On main it
|
|
# leaves :c-<sha>, which rule 145 requires of every main push whether
|
|
# or not a build ran.
|
|
#
|
|
# steps.tag emits ONE comma-separated list; imagetools wants a -t per
|
|
# ref. (That list used to feed docker/build-push-action directly —
|
|
# which is exactly what #3190 made unsafe.)
|
|
ARGS=""
|
|
IFS=,
|
|
for t in $TAGS; do
|
|
[ "$t" = "$SOURCE" ] && continue
|
|
ARGS="$ARGS -t $t"
|
|
done
|
|
unset IFS
|
|
if [ -z "$ARGS" ]; then
|
|
echo "repoint: $SOURCE is the only tag for this channel and"
|
|
echo "repoint: already holds this revision — nothing to write."
|
|
exit 0
|
|
fi
|
|
# shellcheck disable=SC2086
|
|
docker buildx imagetools create $ARGS "$SOURCE"
|
|
echo "repointed from $SOURCE:$ARGS"
|