ci: smoke the refreshed image against real Postgres and Redis
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 3s
CI / extension-version (push) Successful in 4s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 9s
Build images / build-web (push) Successful in 7s
Build images / smoke-web (push) Skipped
extension / lint (push) Successful in 20s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 32s
CI / integration (push) Successful in 1m48s
extension / lint (pull_request) Successful in 20s

Milestone 362 step 3. This is the gate the weekly base refresh never had.

`ci.yml` cannot be that gate, and the reason matters more than the fix. Its
lanes run on ci-python:3.14 and install requirements.txt — a base refresh
changes neither, so all five stay green through a bump that breaks the product.
What a refresh re-resolves is the Dockerfile's apt layer:

    ffmpeg unar libpq5 postgresql-client zstd megatools
    libjpeg62-turbo libwebp7 libpng16-16 ca-certificates

Unpinned, every build, and nothing else in this repo looks at it. That line is
the dependency creep; it is also precisely what the test suite structurally
cannot observe, since the suite never runs inside the image and the image
carries no tests and no pytest.

So `smoke-web` runs the CANDIDATE IMAGE against real service containers:

  1. `alembic upgrade head` on an empty database — the image's own libpq and
     psycopg, and the same call entrypoint.sh makes before it serves anything,
     so a failure here is a failure to boot.
  2. The apt binaries, then the application's own `Thumbnailer` — JPEG, PNG
     with alpha, WebP, and a video frame through ffmpeg. `Thumbnailer` needs no
     database and no app context, so the check exercises real product code
     rather than a proxy for it. `ffmpeg -version` exiting 0 would pass while a
     codec removal broke every thumbnail in the library.
  3. The web role boots and answers /api/health.

Every failure names the package it implicates. This fires on a Sunday,
unattended, about a change nobody made deliberately — "assertion failed" a week
later teaches nobody anything.

The script is piped over stdin rather than bind-mounted: the workspace is a
docker volume belonging to the job's own container, so a host bind of $PWD does
not resolve for a sibling. Container logs are dumped only on failure, and the
trap re-exits with the real status rather than the status of `docker rm`.

Deliberately NOT gating the promote yet — that is step 4. Landing the gate and
the thing it gates together would mean the first time anyone saw this job run
would also be the first time it could stop a publish.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
This commit is contained in:
2026-09-02 15:45:10 -04:00
co-authored by Claude Opus 5
parent 24a2b70a5a
commit bfa9fd678b
2 changed files with 300 additions and 0 deletions
+145
View File
@@ -1102,6 +1102,151 @@ jobs:
docker buildx imagetools create $ARGS "$SOURCE"
echo "repointed from $SOURCE:$ARGS"
# Does the image a refresh just built still work?
#
# This is the gate the base refresh never had. `ci.yml` cannot be it: its
# lanes run on ci-python:3.14 and install requirements.txt, and a base bump
# changes neither — all five stay green through a refresh that breaks the
# product. What a refresh re-resolves is the Dockerfile's apt layer (ffmpeg,
# unar, libpq5, postgresql-client, zstd, megatools, libjpeg62-turbo,
# libwebp7, libpng16-16), unpinned, every build.
#
# So this runs the CANDIDATE IMAGE, against real Postgres and Redis. Not the
# source tree, and not a static inspection: `ffmpeg -version` exiting 0 would
# pass while a codec removal broke every thumbnail in the library.
#
# Refresh-only. On a push the bytes came from a commit, and a commit is what
# ci.yml already tests.
#
# Reports a verdict; it does not yet gate the promote (milestone 362 step 4).
# Landing the gate and the thing it gates in one change would mean the first
# time anyone saw this job run would also be the first time it could stop a
# publish.
smoke-web:
if: env.IS_REFRESH == 'true'
needs: [build-web]
runs-on: python-ci
container:
image: git.fabledsword.com/bvandeusen/ci-python:3.14
env:
DB_USER: fabledcurator
DB_PASSWORD: ci_smoke
DB_PORT: "5432"
DB_NAME: fabledcurator_smoke
SECRET_KEY: ci_smoke_placeholder
IMAGE: git.fabledsword.com/bvandeusen/fabledcurator
services:
postgres:
image: pgvector/pgvector:pg16
env:
POSTGRES_USER: fabledcurator
POSTGRES_PASSWORD: ci_smoke
POSTGRES_DB: fabledcurator_smoke
options: >-
--health-cmd "pg_isready -U fabledcurator"
--health-interval 10s
--health-timeout 5s
--health-retries 10
redis:
image: redis:7-alpine
options: >-
--health-cmd "redis-cli ping"
--health-interval 10s
--health-timeout 5s
--health-retries 10
steps:
- uses: actions/checkout@v4
with:
# The same ref the image was built from, so the smoke script matches
# the code inside the candidate.
ref: ${{ env.BUILD_REF }}
- name: Smoke the candidate image
env:
TOKEN: ${{ secrets.RELEASE_TOKEN }}
ACTOR: ${{ github.actor }}
run: |
set -eux
# Service discovery mirrors ci.yml's integration lane: these jobs run
# in a container against a mounted docker socket, so the services are
# SIBLINGS reachable by IP, not by hostname.
PG=$(docker ps --filter "name=smoke" --filter "ancestor=pgvector/pgvector:pg16" -q | head -n1)
RD=$(docker ps --filter "name=smoke" --filter "ancestor=redis:7-alpine" -q | head -n1)
test -n "$PG" && test -n "$RD"
PG_IP=$(docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' "$PG")
RD_IP=$(docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' "$RD")
test -n "$PG_IP" && test -n "$RD_IP"
# Socket probe in python, not bash's /dev/tcp — these steps run under
# `sh -e`, where that path does not exist. Same fix and reasoning as
# ci.yml's integration job; see the comment there.
pg_ready=""
for i in $(seq 1 60); do
if python -c "import socket,sys; s=socket.socket(); s.settimeout(2); sys.exit(0 if s.connect_ex(('$PG_IP', 5432)) == 0 else 1)"; then
pg_ready=1
break
fi
sleep 2
done
if [ -z "$pg_ready" ]; then
echo "postgres at $PG_IP:5432 did not accept a connection within 120s"
exit 1
fi
echo "$TOKEN" | docker login git.fabledsword.com -u "$ACTOR" --password-stdin
CANDIDATE="$IMAGE:refresh-candidate"
docker pull "$CANDIDATE"
ENVOPTS="-e DB_USER=$DB_USER -e DB_PASSWORD=$DB_PASSWORD -e DB_HOST=$PG_IP"
ENVOPTS="$ENVOPTS -e DB_PORT=5432 -e DB_NAME=$DB_NAME -e SECRET_KEY=$SECRET_KEY"
ENVOPTS="$ENVOPTS -e CELERY_BROKER_URL=redis://$RD_IP:6379/0"
ENVOPTS="$ENVOPTS -e CELERY_RESULT_BACKEND=redis://$RD_IP:6379/0"
# 1. The schema builds from empty, using the image's OWN libpq and
# psycopg. This is the same call entrypoint.sh makes before it
# serves anything, so a failure here is a failure to boot.
echo "smoke: alembic upgrade head"
docker run --rm $ENVOPTS "$CANDIDATE" alembic upgrade head
# 2. The apt layer's binaries and the app's own thumbnail path, run
# inside the image. Piped over stdin rather than bind-mounted: the
# workspace is a docker VOLUME belonging to this job's container,
# so a host bind of $PWD would not resolve for a sibling.
echo "smoke: image-internal checks"
docker run --rm -i $ENVOPTS "$CANDIDATE" shell -c 'python3 -' < scripts/smoke_image.py
# 3. It actually serves. `docker run -d` then poll the container's own
# IP — no port publishing, because the job container reaches
# siblings directly and a published port would collide with
# whatever else the runner is hosting.
echo "smoke: web boots and answers /api/health"
CID=$(docker run -d $ENVOPTS "$CANDIDATE" web)
# Clean up the container however this ends, and dump its log ONLY
# on failure — a boot that never answers must fail with the reason
# visible rather than as a bare timeout (rule 156), while a green run
# has nothing to say. `exit $rc` preserves the real status, which a
# trap that ends on a successful `docker rm` would otherwise mask.
trap 'rc=$?; [ $rc -eq 0 ] || docker logs "$CID" 2>&1 | tail -40; docker rm -f "$CID" >/dev/null 2>&1 || true; exit $rc' EXIT
WEB_IP=$(docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' "$CID")
test -n "$WEB_IP"
healthy=""
for i in $(seq 1 60); do
if curl -fsS --max-time 5 "http://$WEB_IP:8080/api/health" >/dev/null 2>&1; then
healthy=1
break
fi
sleep 2
done
if [ -z "$healthy" ]; then
echo "smoke: FAILED — web did not answer /api/health within 120s." >&2
echo "smoke: entrypoint runs alembic BEFORE serving, and step 1" >&2
echo "smoke: passed, so look at hypercorn and the python base." >&2
exit 1
fi
curl -fsS --max-time 5 "http://$WEB_IP:8080/api/health"
echo
echo "smoke: all checks passed against $CANDIDATE"
build-ml:
runs-on: python-ci
container: