main
1415
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a8fdbd86bc |
Merge pull request 'FabledCurator can now tell you one of its own parts has stopped' (#250) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 4s
Build images / build-agent (push) Successful in 8s
Build images / build-ml (push) Successful in 13s
Build images / build-web (push) Successful in 13s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / frontend-build (push) Successful in 20s
CI / backend-lint-and-test (push) Successful in 33s
CI / integration (push) Successful in 1m53s
|
||
|
|
5084ba666b |
feat: the dot beside the brand now means the whole stack (milestone 365 step 4)
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Successful in 5s
CI / extension-version (push) Successful in 5s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 9s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 31s
Build images / build-web (push) Successful in 55s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / integration (push) Successful in 1m43s
The ask was a surface AND a path. The path is the part that was missing — everything that could answer "is it running" lived inside Settings, which you only open once you already suspect something. **Re-used the indicator that already existed rather than adding a fourth.** There were three partial surfaces: TopNav's health dot, PipelineStatusChip's pulse, and the Settings Activity tab. None answered "is every part alive", and a fourth would have made the question harder to answer, not easier. TopNav's dot read /api/health — a no-DB liveness check proving only that the WEB container is serving. Green there while a worker was dead is exactly what it looked like, and a green dot beside the product name gets read as "everything is fine". It now reflects the whole-stack verdict, and it is a link: the place someone already looks when they suspect something is now also the way to the detail. The tooltip names the actual problem. "Scheduler has not checked in for 6 min" sends someone somewhere; "something is unhealthy" sends them hunting. /system is deliberately NOT in the nav row — TopNav builds that from routes with a meta.title, and a sixth top-level tab for a page visited twice a year costs more attention than it returns. It is reached from the dot. The page lists every learned part with its state as a sentence rather than a chip, and prints the staleness thresholds it was judged by, taken from the endpoint so the UI keeps no second copy of them. PipelineStatusChip still hand-rolls its own 3-minute scheduler window; that is now a duplicate of a threshold the server owns, and worth collapsing once this has been watched working. The stores stay separate on purpose: system.js is "can I reach the API", systemActivity.js is "what is the pipeline doing", systemHealth.js is "is anything broken". Running and alive fail independently — an idle stack with a dead worker looks identical to a healthy one on every activity surface, which is the whole reason this milestone exists. Not yet verified against a real stopped service. Rule 12 keeps a local stack out of it, and frontend CI has no Vue type-check or visual regression, so this needs an operator look rather than a green lane. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA |
||
|
|
fe4e0f2b71 |
feat: /api/system/health — one verdict for the whole stack (milestone 365 step 3)
CI / lint (push) Successful in 5s
Build images / sign-extension (push) Successful in 5s
Build images / build-agent (push) Successful in 8s
CI / extension-version (push) Successful in 4s
CI / frontend-build (push) Successful in 20s
CI / backend-lint-and-test (push) Successful in 35s
Build images / build-web (push) Successful in 1m1s
Build images / smoke-web (push) Skipped
CI / integration (push) Successful in 1m51s
Build images / build-ml (push) Successful in 1m59s
Build images / promote (push) Skipped
The single endpoint the nav indicator and the System page will both read. Composing a verdict is this module's job, not the UI's. Two kinds of part, answered differently. LEARNED — celery roles and the GPU agent, out of service_seen, where the question is "how long since it checked in" and the answer can be "it has not". PROBED — Postgres and Redis, always expected, never learned, because a last-seen for them would be actively misleading: that Redis answered thirty seconds ago says nothing about now. **The endpoint must never fail because something it checks has failed.** That inversion is easy to write by accident and it destroys the feature exactly when it is needed — a 500 when Redis is down instead of `redis: down`. Every probe is wrapped, every wait carries a deadline (rule 156), and the roster refresh swallows its own errors. The worst case is a part reported `unknown`, which is a true statement about the system. Postgres is probed first and gates the rest, because if it is unreachable nothing else can be read — and "the database is down" is the most useful single thing this can ever say. The staleness thresholds are the design risk, not the code, and they are deliberately generous: 90s to doubt, 300s to disbelieve. The constraint is a deploy rather than a crash — `docker compose up -d` rolls start-first, so a role is briefly served by two containers and then by neither while the old one drains. Thresholds tight enough to catch a crash in seconds would paint the page red on every update, and an alarm that cries wolf on every deploy is one nobody reads. Tune down only after watching a real deploy pass through. The numbers ship in the response so the UI can explain a `stale` without keeping a second copy of them. States are described in sentences rather than left as chips: "Scheduler has not checked in for 6 min — treat it as stopped" is what someone needs at the moment they are deciding whether to go and open Portainer. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA |
||
|
|
dc8af8b1a7 |
feat: a learned roster, so a stopped part is observable (milestone 365 steps 1-2)
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 30s
Build images / build-web (push) Successful in 55s
Build images / smoke-web (push) Skipped
Build images / build-ml (push) Successful in 1m45s
Build images / promote (push) Skipped
CI / integration (push) Successful in 1m49s
Nothing in FabledCurator knew what was SUPPOSED to be running. `celery inspect` reports the workers that ANSWER, so a dead worker was a shorter list rather than a red light, and grep for any notion of expected services returned nothing. That is why Portainer was the only place an operator could see it: Portainer knows the intended set. `service_seen` is the memory that makes an absence observable — every part that has checked in, and when it last did. **Keyed on the queue set, not the worker hostname.** Celery's worker names here are `celery@<container id>`, minted fresh on every deploy. Keyed on those, this table would record a death and a birth every time the stack updates — and a status page that goes red on every deploy is a status page nobody reads, which is worse than not having one. CELERY_QUEUES is assigned per role in compose and survives container replacement, so it is the stable identity. Two replicas of a role are therefore ONE row, which is right: the question is whether the role is served, not how many containers exist. The GPU agent is keyed on agent_id, the identity its lease protocol already uses. gpu.py received it on both lease and heartbeat and threw it away — an idle agent with nothing to lease left no trace and was indistinguishable from one switched off a week ago. Now recorded on the calls that were already happening. **Who observes, corrected from the plan.** The plan said "record from the existing inspect path", which would only run when someone opened the Activity tab. Two other candidates and why they lost: - A beat sweep. If the scheduler dies the sweep stops, every row goes stale, and the page says everything is down when one thing is. An alarm that cannot distinguish "a part died" from "the observer died" is worse than none. - A background task in web. hypercorn runs --workers 4, so that is four concurrent inspect loops per container, forever. Taken instead: refresh on demand, rate-limited by the newest last_seen_at that every process can already see. The observer is then the thing serving the page — if web is down you get a browser error, not a confidently green page — and it self-limits with no coordination, since a race costs one redundant inspect that writes identical values. Migration 0090 is the first written on the collapsed baseline (milestone 328), so it is also the first evidence the chain steps FORWARD from 0089 rather than merely reproducing the schema. No secondary indexes: one row per moving part means every read is a handful of rows, and #3301 is the record of what speculative indexes cost. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA |
||
|
|
0421fd3109 |
Merge pull request 'The weekly refresh now publishes only what a gate has proven — and a new install can start' (#249) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 4s
Build images / build-ml (push) Successful in 8s
Build images / build-agent (push) Successful in 9s
Build images / build-web (push) Successful in 7s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / frontend-build (push) Successful in 22s
extension / lint (push) Successful in 19s
CI / backend-lint-and-test (push) Successful in 31s
CI / integration (push) Successful in 1m49s
|
||
|
|
131237143b |
Revert "test: force the smoke gate to fail, to watch it block a publish"
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 6s
Build images / build-ml (push) Successful in 8s
Build images / build-agent (push) Successful in 9s
CI / frontend-build (push) Successful in 22s
extension / lint (push) Successful in 21s
Build images / build-web (push) Successful in 6s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / backend-lint-and-test (push) Successful in 34s
CI / integration (push) Successful in 2m0s
extension / lint (pull_request) Successful in 20s
The gate held. Run 5320, dispatched with the forced failure in place: build-web success (candidate published) build-ml success build-agent success smoke-web FAILED promote skipped run failure And the three channel tags did not move: fabledcurator 33d3d8332f74 -> 33d3d8332f74 fabledcurator-ml e94a5435cb45 -> e94a5435cb45 fabledcurator-agent bae27d34d811 -> bae27d34d811 So a refresh that breaks something now leaves :latest naming the build that works, which is the property milestone 362 exists to establish. The rejected candidate is still published under :refresh-candidate, so whoever reads the red job on Monday can pull the exact image that failed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA |
||
|
|
59d27ef76e |
test: force the smoke gate to fail, to watch it block a publish
CI / lint (push) Successful in 4s
Build images / sign-extension (push) Successful in 4s
CI / extension-version (push) Successful in 3s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 8s
CI / frontend-build (push) Successful in 19s
extension / lint (push) Successful in 18s
Build images / build-web (push) Successful in 7s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / backend-lint-and-test (push) Successful in 31s
CI / integration (push) Successful in 1m57s
TEMPORARY, reverted in the next commit. Milestone 362's verification section requires the gate to be seen rejecting a build — a gate nobody has watched reject anything is a gate nobody knows is wired up. Every real check passes, so the rejection has to be forced. Under test is the job dependency, not the assertions: a failed smoke-web must skip the promote job, and the three :latest tags must still name the digests they named before the run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA |
||
|
|
f630e50e75 |
ci: the refresh publishes only what the gate passed (#3265 milestone step 4)
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 5s
Build images / build-ml (push) Successful in 8s
Build images / build-agent (push) Successful in 9s
Build images / build-web (push) Successful in 6s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
extension / lint (push) Successful in 19s
CI / frontend-build (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 34s
CI / integration (push) Successful in 1m53s
The gate reported a verdict nothing consulted. Now it decides. The promote moved out of the three build jobs into its own `promote` job, because the verdict cannot exist until build-web has finished and the promote used to run inside it. `needs: [build-web, build-ml, build-agent, smoke-web]` is the whole mechanism: a failed smoke skips the promote, so a refresh that broke something leaves :latest naming the build that works. "The refresh failed" and "production is broken" must not be the same event. A SKIPPED smoke also skips it, and that is the case that matters most. On run 5290 the gate silently skipped itself — job-level `if:` cannot read the env context — and a design where only a FAILED gate blocks would have published unverified images while reporting success. Not running is not the same as passing, and today produced two separate bugs of exactly that shape (#3414, and the smoke-web skip). All three images now promote together or not at all. They are one stack: build.yml already refuses to publish a :dev web image beside a stale :dev ml because the mismatch only surfaces as a runtime failure, and a refresh that published ml while withholding web would be that same trap reached through the gate. Stated plainly in the job comment: the gate covers web only, so ml and agent are held to web's verdict rather than their own. That is the conservative direction, not equivalent evidence, and should not be read as if it were. Three near-identical promote steps collapsed into one loop. A partial failure now says which images moved and that the state is inconsistent, rather than leaving that to be inferred — the promote is idempotent and the candidates are still published, so the instruction is simply to re-run. Also removed the now-dead `promote` output from the ml and agent reuse steps. Only build-web's is read (as outputs.candidate); two more copies nothing consults is the kind of thing that reads as load-bearing a year later. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA |
||
|
|
86abaf0b94 |
docs: a new install could not start, and nothing told anyone why (#3422)
CI / lint (push) Successful in 4s
Build images / sign-extension (push) Successful in 4s
CI / extension-version (push) Successful in 3s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 8s
Build images / build-web (push) Successful in 6s
Build images / smoke-web (push) Skipped
CI / frontend-build (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 31s
CI / integration (push) Successful in 1m46s
The install path milestone 328 wrote produces a web container that exits on boot. entrypoint.sh runs alembic, then app construction raises: MissingCredentialKey: Fernet key file not found at /images/secrets/credential_key.b64. For first-time setup, set CURATOR_BOOTSTRAP_NEW_KEY=1. That variable appeared in no README, no .env.example and no compose file — only in backend/. So a stranger following the documented steps got an app that does not start and an error with no context. Found by the milestone-362 smoke gate on its first real run (#3422). The product behaviour stays exactly as it is. credential_crypto refuses to mint a key because the 2026-06-02 audit found a partial restore — database back, ./images/secrets lost — silently generating a fresh one and producing a healthy-looking instance where every authenticated download failed AUTH_ERROR. Failing fast is right; not saying so is the bug. So: .env.example carries the variable in its own FIRST BOOT ONLY section with the reasoning and an instruction to delete the line afterwards, and README's First run leads with it, because "the app will not start" belongs before "the ML worker downloads weights". Both say to back up ./images/secrets/ alongside the database, which is the part that costs real data if it is learned late. **compose had to change too, and this is the part that would have shipped a second broken instruction.** A variable in `.env` is only used for ${...} interpolation — it does not reach the container unless the service names it. Telling people to set it in .env, without that, would have documented a step that does nothing. Added to the shared app_env anchor, defaulted to empty so the refusal still stands for everyone who has not opted in. Not taken: auto-bootstrapping when the credential table is empty, which would remove the manual step entirely and keep the audit's protection for restores. That is the better product and it is a code change with a predicate that has to be exactly right; this is the smallest correct fix, and #3422 stays open for the other one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA |
||
|
|
4815040d74 |
ci: the smoke gate found a real one on its first run — and had two bugs of its own
Build images / sign-extension (push) Successful in 6s
CI / lint (push) Successful in 6s
CI / extension-version (push) Successful in 4s
Build images / build-ml (push) Successful in 11s
Build images / build-agent (push) Successful in 13s
extension / lint (push) Successful in 24s
CI / frontend-build (push) Successful in 28s
CI / backend-lint-and-test (push) Successful in 35s
Build images / build-web (push) Successful in 8s
Build images / smoke-web (push) Skipped
CI / integration (push) Successful in 2m5s
Run 5296 was `smoke-web`'s first genuine execution. Checks 1 and 2 passed: alembic built the schema from empty inside the image, all five apt binaries resolved, and the application's own Thumbnailer produced JPEG, PNG-with-alpha, WebP and an ffmpeg video frame against the image's libraries. Check 3 failed, and the trap's log dump said exactly why: MissingCredentialKey: Fernet key file not found at /images/secrets/credential_key.b64. For first-time setup, set CURATOR_BOOTSTRAP_NEW_KEY=1. That is the product being right. credential_crypto refuses to mint a key unless someone opts in, because the 2026-06-02 audit found a partial restore (DB back, /images/secrets/ lost) silently generating a fresh one and leaving a working-looking system where every authenticated download failed AUTH_ERROR. It is also a first-run blocker for milestone 328, filed as #3422: the variable appears in no README, no .env.example and no compose file, so the install path that milestone just finished writing produces a container that exits on boot. Not fixed here — the fix trades safety against friction and is the operator's call. Two defects in the gate itself, both surfaced by the same run: - A throwaway CI instance IS first-time setup, so it now passes CURATOR_BOOTSTRAP_NEW_KEY=1. The check was asserting a condition no fresh container can satisfy. - The health loop polled a dead container for 3m35s. Docker had already recycled its IP, so the replies were a baffling mix of connection-refused and 5s timeouts from whatever took the address next. It now checks `.State.Running` each iteration and fails immediately with the container's log. The trap had the real answer the whole time; this stops burying it under four minutes of noise. Also corrected a message claiming a 120s budget: 60 iterations of up to 5s connect plus 2s sleep is nearer seven minutes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA |
||
|
|
81b7b6f308 |
ci: smoke-web never ran — a job's if: cannot read the env context
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 4s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 8s
CI / frontend-build (push) Successful in 22s
extension / lint (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 32s
Build images / build-web (push) Successful in 8s
Build images / smoke-web (push) Skipped
CI / integration (push) Successful in 1m51s
Run 5290 dispatched a refresh. Everything worked: the guard fired, the build published the candidate, the promote pointed :latest at it. And `smoke-web` reported conclusion "skipped", with no steps and no log. Its condition was `if: env.IS_REFRESH == 'true'`. The env context is available to STEP conditions and step bodies but never to a job's own `if:`, and an unresolvable context there evaluates to empty rather than erroring. So the gate skipped itself, silently, on the one run that existed to exercise it. Second silent-skip of this family today, after #3414. Same shape both times: something evaluated false, nothing failed, and the run reported success. It is worth naming the pattern — on this pipeline, "green" and "ran" are different claims, and the steps' own conclusions are the only place the difference shows. Fixed by keying off a job output rather than re-deriving the trigger: build-web now exposes the reuse step's `promote` decision as `outputs.candidate` and smoke-web consumes it. That is better than duplicating the expression: it is the same single decision the build, the XPI download and the promote all take already — build.yml's own "one decision drives everything downstream" — and it asserts the thing smoke-web actually depends on, that a candidate was published, rather than restating the reason one would be. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA |
||
|
|
d01f33dea6 |
Merge pull request 'A gate for the weekly refresh, and the lever that makes it testable' (#248) from dev into main
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 3s
CI / extension-version (push) Successful in 2s
Build images / build-ml (push) Successful in 6s
Build images / build-agent (push) Successful in 7s
Build images / build-web (push) Successful in 6s
Build images / smoke-web (push) Skipped
CI / frontend-build (push) Successful in 19s
extension / lint (push) Successful in 17s
CI / backend-lint-and-test (push) Successful in 31s
CI / integration (push) Successful in 1m42s
|
||
|
|
bfa9fd678b |
ci: smoke the refreshed image against real Postgres and Redis
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 3s
CI / extension-version (push) Successful in 4s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 9s
Build images / build-web (push) Successful in 7s
Build images / smoke-web (push) Skipped
extension / lint (push) Successful in 20s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 32s
CI / integration (push) Successful in 1m48s
extension / lint (pull_request) Successful in 20s
Milestone 362 step 3. This is the gate the weekly base refresh never had.
`ci.yml` cannot be that gate, and the reason matters more than the fix. Its
lanes run on ci-python:3.14 and install requirements.txt — a base refresh
changes neither, so all five stay green through a bump that breaks the product.
What a refresh re-resolves is the Dockerfile's apt layer:
ffmpeg unar libpq5 postgresql-client zstd megatools
libjpeg62-turbo libwebp7 libpng16-16 ca-certificates
Unpinned, every build, and nothing else in this repo looks at it. That line is
the dependency creep; it is also precisely what the test suite structurally
cannot observe, since the suite never runs inside the image and the image
carries no tests and no pytest.
So `smoke-web` runs the CANDIDATE IMAGE against real service containers:
1. `alembic upgrade head` on an empty database — the image's own libpq and
psycopg, and the same call entrypoint.sh makes before it serves anything,
so a failure here is a failure to boot.
2. The apt binaries, then the application's own `Thumbnailer` — JPEG, PNG
with alpha, WebP, and a video frame through ffmpeg. `Thumbnailer` needs no
database and no app context, so the check exercises real product code
rather than a proxy for it. `ffmpeg -version` exiting 0 would pass while a
codec removal broke every thumbnail in the library.
3. The web role boots and answers /api/health.
Every failure names the package it implicates. This fires on a Sunday,
unattended, about a change nobody made deliberately — "assertion failed" a week
later teaches nobody anything.
The script is piped over stdin rather than bind-mounted: the workspace is a
docker volume belonging to the job's own container, so a host bind of $PWD does
not resolve for a sibling. Container logs are dumped only on failure, and the
trap re-exits with the real status rather than the status of `docker rm`.
Deliberately NOT gating the promote yet — that is step 4. Landing the gate and
the thing it gates together would mean the first time anyone saw this job run
would also be the first time it could stop a publish.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
|
||
|
|
24a2b70a5a |
ci: a boolean input never equals the string 'true'
Build images / sign-extension (push) Successful in 4s
Build images / build-ml (push) Successful in 6s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 23s
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 2s
CI / backend-lint-and-test (push) Successful in 35s
Build images / build-web (push) Successful in 8s
CI / integration (push) Successful in 1m48s
The refresh lever did not work, and the way it did not work is the point.
Run 5270 dispatched with refresh=true. Its log:
expression '(github.event_name == 'schedule'
|| github.event.inputs.refresh == 'true') && 'true' || 'false''
evaluated to '%!t(string=false)'
trigger: event=workflow_dispatch IS_REFRESH='false' BUILD_REF='refs/heads/dev'
trigger: raw inputs refresh='true' force_build='false'
The input arrived as true and the comparison still said false. `type: boolean`
delivers a real boolean, and GitHub expression semantics cast operands to
numbers when their types differ — so `true == 'true'` compares 1 against NaN.
My comment on the previous commit asserted the opposite, that Forgejo delivers
inputs as strings, and asserted it without checking.
The run went GREEN with every step skipped, because a refresh that evaluates
false is indistinguishable from an ordinary push. A lever that silently does
nothing is worse than no lever: it would have been trusted.
Normalised through format(), which is representation-independent — a boolean
true and a string 'true' both render 'true'. That is also why force_build was
never bitten: it passes its raw value into an env var and compares in the
shell, where everything is a string already. format() buys the same thing at
expression level, which is where a step `if:` needs the answer.
The diagnostic from the previous commit stays. It is what turned this from a
guess into a measurement, and it is the only thing that would catch the same
class of failure next time.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA
|
||
|
|
2c88ad3efb |
ci: report the raw and normalised trigger values
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
CI / extension-version (push) Successful in 2s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 7s
CI / backend-lint-and-test (push) Successful in 31s
Build images / build-web (push) Successful in 9s
extension / lint (push) Successful in 16s
CI / frontend-build (push) Successful in 22s
CI / integration (push) Successful in 2m39s
The refresh dispatch on run 5265 went green with every step skipped: the main-only guard did not fire, checkout took dev, and the reuse step read IS_REFRESH as false. So both workflow-level expressions evaluated false while the identical accessor works for force_build, which compares its value in the shell rather than in an expression. That is a guess until it is measured, and the failure is silent by construction — a refresh that evaluates false behaves exactly like an ordinary push and reports success. This prints the raw input beside the normalised value in the step that already exists to say what a run derived, so the two disagreeing is visible rather than inferred. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA |
||
|
|
adab33694d |
Merge pull request 'Make a digest mean something again, and give the refresh somewhere to stand' (#247) from dev into main
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
CI / extension-version (push) Successful in 3s
Build images / build-agent (push) Successful in 7s
Build images / build-ml (push) Successful in 8s
Build images / build-web (push) Successful in 14s
CI / frontend-build (push) Successful in 21s
extension / lint (push) Successful in 18s
CI / backend-lint-and-test (push) Successful in 32s
CI / integration (push) Successful in 2m2s
|
||
|
|
bfc4f9cec9 |
ci: one fact for "is this a base refresh", and a lever to trigger one
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / build-agent (push) Successful in 7s
Build images / build-ml (push) Successful in 7s
Build images / build-web (push) Successful in 8s
CI / frontend-build (push) Successful in 20s
extension / lint (push) Successful in 18s
CI / backend-lint-and-test (push) Successful in 37s
CI / integration (push) Successful in 1m45s
extension / lint (pull_request) Successful in 24s
Milestone 362, enabling step 2's verification and everything after it. The weekly refresh was testable once a week. That is not a cadence anything can be developed against, and milestone 362's whole point is a gate — which has to be watched rejecting something before anyone can believe it is wired up. So `refresh` joins `force_build` as a dispatch input, on the same reasoning that added that one (#3252: confirm #3190 was gone rather than wait for it to recur). Adding it meant confronting that "is this a refresh?" was asked in five places and spelled five ways: `github.event_name == 'schedule'` in an `if:`, `$GITHUB_EVENT_NAME` in one shell, an `EVENT:` env passed into another, and a bare expression on `pull:`. Five spellings of one fact is how half of them come to disagree once somebody adds a sixth trigger — which is precisely what this commit is. So it is derived once at the top, next to BUILD_REF, which already exists for exactly this reason on exactly this question. String comparison, not boolean: Forgejo delivers dispatch inputs as strings, so `inputs.refresh` is 'true'/'false' and `&&` on it would read the string 'false' as truthy. **A constraint this makes visible, which pre-dates it.** A refresh checks out `main` (BUILD_REF) while running the workflow definition from the branch that triggered it — the cron registers from the default branch. So dev's workflow builds main's source, and dev's workflow cannot depend on anything main's tree does not have yet. It does now: the reuse step calls `artifacts.sh epoch`, which lands on main with this batch. Until then a refresh dispatch fails loudly at that call, which is the right failure — the alternative is tolerating a missing epoch and silently rebuilding #3265 into every refresh. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA |
||
|
|
b590d25f8f |
ci: the scheduled refresh builds a candidate, then names the channel
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 2s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 8s
Build images / build-web (push) Successful in 6s
CI / frontend-build (push) Successful in 24s
extension / lint (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 32s
CI / integration (push) Successful in 1m50s
Milestone 362 step 2. Structural: it creates a moment between "built" and "published" for step 3's gate to occupy. No behaviour change. A refresh rebuilds against freshly resolved base images, and the web image's runtime is a line of UNPINNED Debian packages — ffmpeg, libjpeg62-turbo, libpq5, megatools — re-resolved on every build. Nothing in ci.yml can see that: its lanes run on ci-python:3.14 and install requirements.txt, and a base bump changes neither. So refreshed bytes need proving before :latest names them, and proving needs somewhere to stand. On a push nothing changes: build_ref IS channel_ref, promote is false, and the build writes the channel tag directly the way it always has. On the schedule the build writes :refresh-candidate — one moving ref per image, overwritten in place, holding a build nobody is told to pull. That is the shape rule 145 already allows for :buildcache, not the per-build tag family 318 withdrew. Both values are decided in the reuse step beside `hit`, because that step already owns "what does this job do" (build.yml's own rule, at the force branch). A promote condition derived somewhere else could disagree with the tag the build actually wrote. **The promote is a manifest PUT, not `imagetools create`.** That distinction is the whole risk in this change. `imagetools create` wraps its source in an index, and an indexed channel tag is the one thing this pipeline cannot survive: `.Image.Config.Labels` does not resolve through an index, so the fc.revision the reuse check reads back would come up empty, every later push would miss and rebuild, and nothing would go red. That is #3183, observed on run 4751 — reuse worked exactly once and the only symptom was the bill. The repoint step already excludes its own source tag for this reason; a promote that re-introduced the wrap through another door would undo that care. A manifest PUT is what "make this tag name that image" means at the registry: same bytes, same media type, identical digest, no layer transfer. It reads the result back and fails if the tag does not name what was just written — a PUT that 2xx'd and landed something else is exactly the silent-and-plausible failure this pipeline keeps producing. Every call carries a deadline (rule 156); a registry that stops answering must fail the step, not hang the weekly refresh until the job times out. Promote is UNCONDITIONAL today, deliberately. Gating it before the gate exists would leave the refresh building something and publishing nothing for as long as this milestone takes. Step 4 wraps it in the smoke suite's verdict. Not yet verified on the refresh path — that needs a scheduled run, and the lever to trigger one on demand is the next commit. This one is verified by the push path being untouched. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA |
||
|
|
635138b0d1 |
ci: pin the build clock to the commit, so an unchanged refresh publishes nothing
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / build-ml (push) Successful in 6s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 20s
extension / lint (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 33s
Build images / build-web (push) Successful in 56s
CI / integration (push) Successful in 1m50s
Milestone 362 step 1, closing #3265's root cause. The weekly base refresh rewrote all three `:latest` tags on 2026-08-30 with nothing changed in any of them. Not a cache miss — run 4934's log shows every content step CACHED and both bases resolved to unchanged pinned digests. buildkit stamps the image config with the wall clock of the build, so identical layers get republished under a new config blob and therefore a new manifest digest. The cost is not storage, it is meaning: `:latest` moved on a calendar, so a digest change stopped being evidence that anything was different. That is the one thing a digest is any use for, and it is load-bearing here — the reuse check, the `:c-<sha>` rollback story and any future redeploy signal all rest on it. SOURCE_DATE_EPOCH normalises `created` and the history timestamps, so the same source produces the same config bytes and the same digest, and pushing it is a registry no-op. The value is routed through artifacts.sh's existing `newest()` rather than taken from git separately. `revision`, `version` and now `epoch` are three fields of ONE lookup, so they cannot drift into naming different commits — a divergence that would stamp an image reproducibly against one commit while it reported being another, with both values looking perfectly well-formed. Note #3127 §2 is the record of what a second clock costs; this adds a view, not a clock. Also corrected: the build step comment and ci-requirements.md both described the churn as current behaviour with the fix as a "likely" future. They now describe what the file does. Tests pin the property the fix depends on, not the fix: epoch is the same commit version names, in both renderings including the extension's unpadded one, and it does not move between two calls on one checkout. A future refactor that gave epoch its own `git log` would pass every other test in that file. Not yet verified end to end — proving it needs two consecutive refreshes to land on the same digest, which is the next thing, and is the step #3265 exists because nobody did last time. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA |
||
|
|
c0370069e0 |
release: the first release describes the product; it has nothing to diff against
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 5s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 8s
Build images / build-web (push) Successful in 7s
CI / frontend-build (push) Successful in 22s
CI / backend-lint-and-test (push) Successful in 36s
CI / integration (push) Successful in 1m42s
Step 7 needs a release that reads as "what is FabledCurator and how do I run
it". What the script would actually have published is "changes since
v26.06.04.0" over 533 commits, truncated to 200 — a release page whose first
screen is the internal build-out that milestone 328 exists to stop shipping,
addressed to a reader who has never seen this project.
Two causes, fixed separately.
**A pre-convention tag is history, not a predecessor.** The 28 `v26.*` tags
were kept when their releases were deleted, so `--match v*` walks ancestry
straight back to one of them. Reachable is not comparable: nobody has run
v26.06.04.0 and its release page no longer exists to compare against. The
match is now `v[0-9][0-9][0-9][0-9].*` — rule 148's shape, which is exactly
the set of tags naming a release a reader could have been running.
**With that narrowed, the first rule-148 tag reaches no predecessor**, and the
old fallback — diff against the whole history — is worse than the problem it
replaced. A release with no predecessor now renders the product overview and
no commit list at all.
The overview is READ OUT OF README.md between `<!-- overview:start -->` and
`<!-- overview:end -->`, not written into the script, for the same reason the
changelog is derived: two hand-maintained descriptions of one product drift
and nothing ever catches it. The release page and the repo front page are one
source. Missing markers are reported as a note and publish anyway, on
cross_checks()'s reasoning — the release is still the useful object.
Every later release goes back to being a changelog, which is what note #3127
§5 says a release is for. MAX_COMMITS still guards the case it now guards:
two real releases far enough apart that the list stops being readable.
Also corrected while marking up the README: "Importing — ingests an existing
library from disk" was still advertising the folder-import feature that
|
||
|
|
3e4d39b111 |
Merge pull request 'The install path a stranger reads, and a CI probe that never probed' (#246) from dev into main
Build images / sign-extension (push) Successful in 3s
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / build-agent (push) Successful in 8s
Build images / build-ml (push) Successful in 14s
extension / lint (push) Successful in 17s
Build images / build-web (push) Successful in 15s
CI / frontend-build (push) Successful in 20s
CI / backend-lint-and-test (push) Successful in 37s
CI / integration (push) Successful in 1m55s
|
||
|
|
3590c478f5 |
docs+ci: folder import stays retired, and fix a readiness probe that never probed
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
CI / extension-version (push) Successful in 3s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 13s
Build images / build-web (push) Successful in 7s
CI / frontend-build (push) Successful in 22s
extension / lint (push) Successful in 29s
CI / backend-lint-and-test (push) Successful in 49s
CI / integration (push) Successful in 1m57s
extension / lint (pull_request) Successful in 22s
Two unrelated things, both found while closing out milestone 328. **Folder import (#3367).** The operator's call, this session: the import-from-file surface was abandoned on purpose and is not coming back — "it has its own complexities that we didn't need." The README and the compose comment both described the missing button as a rough edge with a tracking issue, which promised a fix that is not coming. Both now say the retirement is the decision, name Subscriptions as the supported way to fill a new install, and describe /api/import/trigger as an unsupported escape hatch for anyone who wants to script one. **The CI readiness probe.** ci.yml's integration job and baseline.yml both waited for Postgres with `(echo > /dev/tcp/$PG_IP/5432)`. Those steps run under `sh -e` — act's default shell — where /dev/tcp is not a magic path but a filename that does not exist. The probe could therefore never succeed: run 18035, a GREEN run, spends 05:20:53 → 05:22:53 in that loop and exits it by exhaustion, not by connecting. Every integration run has been paying a flat 120s for a check that established nothing, and proceeding regardless. Replaced with a socket connect in python (present in the image, no package needed), and exhausting the budget is now a named failure instead of a silent fall-through — rule 156. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TTjbZZ6JirCMSaJzQV1RhA |
||
|
|
8a4af589f1 |
docs: write the install path for someone who is not the operator (#3271)
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 9s
Build images / build-ml (push) Successful in 29s
Build images / build-web (push) Successful in 23s
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 4s
CI / frontend-build (push) Successful in 20s
CI / backend-lint-and-test (push) Successful in 30s
CI / integration (push) Successful in 3m42s
FC has no login — no User model, no session auth, nothing. That was a deliberate call for a single-operator tool and it stays (operator, this session), but it was nowhere in the docs, and the app stores live Patreon / SubscribeStar / Pixiv session cookies on accounts that carry a payment method. Anyone standing this up from the README could reasonably have put it behind a TLS-terminating proxy and considered it handled. So the no-auth posture is now stated three times, in the three places someone decides where to bind the port: README has a "Before you expose it" section above the install instructions, .env.example explains why there is no auth variable in it, and the compose header says it before the first service. SECURITY.md claimed the opposite. It listed "a multi-user sharing ACL — instances can be shared" among the things worth protecting; there are no accounts to share between. That was rule 47 applied to a codebase that does not implement it, and it would have told a researcher FC holds a boundary it does not. Replaced with the real posture, including that TLS without an authenticating layer in front changes nothing. Also corrected, all of it stale rather than wrong-at-the-time: - EXTENSION_API_KEY was dead config. config.py read it into a field nothing consumed; the real key is generated into app_setting on first use and managed in the UI. Removed from config.py, compose and .env.example. - .env.example pointed at docs/superpowers/specs/… — there is no docs/ dir — and described the extension key as "lands in FC-3", closed 2026-05-21. - The /import mount comment described an FC-5 ImageRepo migration run from "Settings → Maintenance → Legacy migration", a surface with no frontend. - README said the extension installs from Settings → Maintenance. It is on Subscriptions → Settings. README is now split: running FC above the line, developing FC below it, with requirements, first run, the extension, upgrading and troubleshooting on the running side. First run documents the one real gap it found — a new installer with a library on disk has no button to import it, only POST /api/import/trigger, because the manual-scan UI was retired 2026-07-02 when that stopped mattering for an established install. Filed as #3367. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017QHszn9H8VBvx5Ke8x1hvw |
||
|
|
b93a5fc5b4 |
Merge pull request 'Collapse alembic 0001..0089 into one baseline' (#245) from dev into main
CI / lint (push) Successful in 5s
CI / extension-version (push) Successful in 5s
CI / frontend-build (push) Successful in 17s
CI / backend-lint-and-test (push) Successful in 32s
CI / integration (push) Successful in 3m40s
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 9s
Build images / build-web (push) Successful in 7s
Build images / build-ml (push) Successful in 14s
|
||
|
|
aa71cbbdbf |
db: the baseline was missing the three system-tag seeds (#3266)
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 4s
CI / frontend-build (push) Successful in 22s
CI / backend-lint-and-test (push) Successful in 30s
CI / integration (push) Successful in 3m41s
Build images / sign-extension (push) Successful in 3s
Build images / build-agent (push) Successful in 7s
Build images / build-web (push) Successful in 6s
Build images / build-ml (push) Successful in 27s
Integration caught it: 36 tests failing with NoResultFound, all on
_system_tag(db, "banner") and its siblings. 0075 seeds three hygiene
system tags — wip, banner, editor screenshot — and the first version of
the baseline carried only the two settings singletons.
This is the same defect class the baseline's own docstring warns about,
which I then walked into anyway. The reason is worth recording: my scan
for data statements used a regex requiring INSERT to sit immediately
after the opening quote, so it saw
op.execute("INSERT INTO ml_settings (id) VALUES (1)")
and missed 0075, which builds the statement through sa.text() across
several lines with bound parameters. The narrow pattern found two of
three seeds and reported itself complete.
The wider scan — grep for insert/bulk_insert across every revision in
|
||
|
|
973db73221 |
db: collapse alembic 0001..0089 into one baseline (#3266)
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
CI / extension-version (push) Successful in 3s
Build images / build-agent (push) Successful in 8s
CI / frontend-build (push) Successful in 17s
CI / backend-lint-and-test (push) Successful in 30s
Build images / build-web (push) Successful in 2m12s
Build images / build-ml (push) Successful in 2m52s
CI / integration (push) Failing after 3m42s
89 files and 6,300 lines become one file of 807. Nothing about the resulting schema changes; what goes away is the requirement that a new installation replay our development history to arrive at it. revision = "0089", down_revision = None. That pairing IS the migration strategy for existing installs, not a detail of it: a deployed database already has alembic_version = '0089' from running the real 0089, so alembic reads the version table, sees head reached, and does nothing. No stamp is required — which matters, because `alembic stamp` writes a version string without validating anything about the schema it is writing it against, and a wrong stamp is indistinguishable from a right one until the next migration fails. An empty database runs the file and records 0089. Both paths converge. The next migration is 0090, as it would have been; the numbering is continuous across the collapse on purpose. Autogenerate produced nearly all of this unaided, which was NOT true of the first attempt — that one was reverted because the generator silently dropped eleven indexes and three uniqueness guarantees. #3275 put those on the models first, so the HNSW index with its opclass, the COALESCE expression index, the partial uniques, 107 server_defaults and the enum CHECKs are all emitted now. Doing the reconciliation before the squash, rather than after, is what made this work. Hand-added, because none of it can live in a model: * CREATE EXTENSION vector / tsm_system_rows (0001, 0004) — database objects, not table metadata. * The pgvector import. Autogenerate writes qualified pgvector.sqlalchemy.vector.VECTOR references without importing the package, so its own output cannot run (run 4988). * THE TWO SEED ROWS. 0002 and 0003 did not only build schema — each inserted a settings singleton, and nothing in the app ever creates them: ImportSettings.load() and MLSettings.load() are select(...).scalar_one(), which RAISES NoResultFound rather than returning None. A models-only baseline would leave both tables empty and crash a fresh install on first settings access, while baseline.yml reported a perfect schema match. Only running the app against a new database finds that. Not carried over: 0023's DELETE FROM tag and 0047's series deletes, which are historical cleanups operating on rows an empty database lacks. downgrade() raises. A baseline's downgrade is "drop every table", which is a data-loss event wearing a migration as a disguise; offering it as one invites someone to run it. Restore from a backup. Also removed, per the plan: the 10 test_migration_*.py files (they assert intermediate states and backfills that no longer exist — a test that a column exists is already the model tests' job) and backend/app/utils/artist_backfill.py, whose only importer was 0008. Verified no other consumer anywhere in backend/ or tests/. baseline.yml changes with it. chain_ref now DEFAULTS to |
||
|
|
725bf15f88 |
Merge pull request 'AGPL-3.0, SECURITY and CONTRIBUTING; the install path pulls :latest, not :dev' (#244) from dev into main
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Successful in 5s
CI / extension-version (push) Successful in 5s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 9s
CI / frontend-build (push) Successful in 16s
Build images / build-web (push) Successful in 6s
CI / backend-lint-and-test (push) Successful in 31s
CI / integration (push) Successful in 3m43s
|
||
|
|
bc4eba636d |
fix: the documented install path pulled :dev, not :latest (#3270)
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 2s
Build images / build-agent (push) Successful in 7s
Build images / build-ml (push) Successful in 9s
Build images / build-web (push) Successful in 7s
CI / frontend-build (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 31s
CI / integration (push) Successful in 3m44s
docker-compose.yml pinned :dev on all five app services — web, worker, scheduler, maintenance-long, ml-worker. The README documents `docker compose -f docker-compose.yml up -d` as the production path, and -f means "use only this file", skipping the override and its build: directives. So Compose pulled image:, and image: was the rolling development channel. The documented way to install this product shipped development builds. It went unnoticed for a structural reason rather than a careless one: nobody who works on the project takes that path. The operator deploys from a swarm stack file; contributors get docker-compose.override.yml, which sets build: for all five services, and build: wins over image:. The broken path is reachable only by a stranger following the README — which is exactly the audience that did not exist until now. :latest, per rule 147: main IS production. It is also what the agent stack (agent/docker-compose.yml) already pinned, so this makes the two stacks agree rather than introducing a new convention. Both paths verified with `docker compose config`, which merges and prints without starting anything: dev path — build: present on all five, image: not pulled -f production — 0 build: directives, five :latest images resolved Also checked the base file for anything a stranger could not satisfy: no host-absolute volume paths, no operator-specific port bindings, no device mappings. The tag was the only defect in the consumer path. The comment on web.image is deliberately long (rule 32). A line reading :latest inside a file a developer is debugging with is exactly the line someone flips back to :dev to test something and then commits, and the consequence — strangers silently installing bleeding edge — is invisible to everyone who works here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017QHszn9H8VBvx5Ke8x1hvw |
||
|
|
dbc4e8b0c6 |
docs: AGPL-3.0, plus SECURITY and CONTRIBUTING (#3269)
CI / lint (push) Successful in 2s
CI / extension-version (push) Successful in 3s
CI / frontend-build (push) Successful in 20s
CI / integration (push) Successful in 3m52s
Build images / sign-extension (push) Successful in 4s
Build images / build-ml (push) Successful in 6s
Build images / build-agent (push) Successful in 8s
Build images / build-web (push) Successful in 7s
CI / backend-lint-and-test (push) Successful in 31s
The repo had no LICENSE, which meant all rights reserved by default: nobody could legally run or modify it, and "public availability" was a contradiction no amount of install documentation could fix. This is the hard blocker in milestone 328; everything else in it is quality. AGPL-3.0, at the operator's explicit choice, for the reason the operator gave: this should not become something another party runs as a hosted proprietary service. Section 13 is what makes it fit — for a self-hosted web app, distribution otherwise never happens, so the GPL's obligation would never actually bite. AGPL reaches the case that matters here: running a MODIFIED copy as a service for others. LICENSE is the FSF text fetched from gnu.org and verified byte-identical (34,523 bytes, 661 lines, §13 "Remote Network Interaction" present), not retyped. README's old "Personal project; use at your own discretion" said nothing legally and is replaced with what the licence actually asks — including the part worth being clear about, that running an unmodified copy for yourself carries no obligation whatsoever. SECURITY.md names what this software actually holds, because that is what makes a report serious here: live third-party session cookies for accounts with payment methods attached, the extension API key, the multi-user sharing ACL, and arbitrary downloaded media that gets decoded and fed to models. It also states the plain-HTTP posture up front, so "served over HTTP" and "no HSTS" are understood as the documented design rather than filed as findings. There is no private disclosure channel yet, so the reporting instruction is to open an issue containing NOTHING but the fact that a report exists, and wait for a private contact. Awkward on purpose: an issue tracker is public the moment it is written to, and every self-hosted instance stays vulnerable until its operator can update. Worth replacing with a real contact address — that decision is the operator's, since it publishes one. CONTRIBUTING records the two things that actually catch people: ruff's order-by-type import sorting, and that a model change and its migration belong in the same commit. The second is not style — the models and the chain silently diverged for a long time (#3275) and autogenerate was unsafe as a result. Pre-publication scan, since the repo is about to get attention: .env.example is placeholders only (`changeme_*`), `.env` is gitignored with an `!.env.example` exception, no credential-shaped literals are committed, and there are no private IPs or operator home paths. The only hostnames are the project's own forge, which is public by design. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017QHszn9H8VBvx5Ke8x1hvw |
||
|
|
b14818303c |
Merge pull request 'Schema reconciliation + index hygiene, and the weekly base-image refresh' (#243) from dev into main
CI / lint (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
CI / extension-version (push) Successful in 2s
Build images / build-agent (push) Successful in 8s
Build images / build-ml (push) Successful in 14s
CI / frontend-build (push) Successful in 19s
Build images / build-web (push) Successful in 14s
extension / lint (push) Successful in 18s
CI / backend-lint-and-test (push) Successful in 30s
CI / integration (push) Successful in 3m54s
|
||
|
|
08418d54a3 |
db: index the seven unindexed FKs, drop the seven redundant ones (#3300, #3301)
Build images / build-ml (push) Successful in 32s
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 11s
Build images / build-web (push) Successful in 26s
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
CI / integration (push) Successful in 3m44s
CI / frontend-build (push) Successful in 20s
CI / backend-lint-and-test (push) Successful in 30s
extension / lint (pull_request) Successful in 24s
A structural sweep of the deployed schema, run AFTER 0088 got the models and the chain to exact agreement. That agreement is what 0088 achieved, and it is worth naming what it does not prove: a models-vs-chain diff shows the two describe the same schema, not that the schema is right. Everything here was wrong in BOTH. The one that matters: image_tag has PRIMARY KEY (image_record_id, tag_id) and no other index, so tag_id is unindexed. That is the gallery's tag filter (tag_query.py builds `image_tag.c.tag_id == tid`) and the ON DELETE CASCADE from tag, both scanning the largest table in the schema. Six more FKs were unindexed on smaller tables; presentation_review.tag_id also CASCADEs. Dropped, on the other side: ix_image_record_sha256 was an exact duplicate of the index uq_image_record_sha256 already builds — two btrees on the same column of the highest-insert-rate table. The other six are single-column indexes a later composite superseded without the narrow one being retired; a btree on (a,b) already serves lookups on a. 0088 deliberately taught the models to declare BOTH sha256 indexes so they would describe reality. This changes the reality instead, and the models change with it — otherwise the next baseline.yml run reintroduces exactly the drift 0088 removed. CONCURRENTLY throughout, so building the image_tag index does not hold an ACCESS EXCLUSIVE lock over every write for the duration. The cost is that the migration cannot run in a transaction and so is not atomic: every statement is IF NOT EXISTS / IF EXISTS, making a re-run after a partial failure safe. The docstring carries the query for finding an INVALID index left by an interrupted CONCURRENTLY build. What the sweep found clean, for the record: all 43 tables have a primary key; all 51 FKs declare an explicit ON DELETE, so none silently blocks a delete; the three enum CHECKs match the code that writes them (rule 36). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017QHszn9H8VBvx5Ke8x1hvw |
||
|
|
b1bd2531ad |
style: sort JSON first in three sqlalchemy import blocks
CI / lint (push) Successful in 4s
Build images / sign-extension (push) Successful in 4s
CI / extension-version (push) Successful in 4s
Build images / build-agent (push) Successful in 8s
CI / frontend-build (push) Successful in 27s
CI / backend-lint-and-test (push) Successful in 33s
Build images / build-web (push) Successful in 43s
Build images / build-ml (push) Successful in 51s
CI / integration (push) Successful in 3m47s
ruff's isort runs with order-by-type, which sorts ALL_CAPS names ahead of CamelCase ones, so `JSON` belongs at the head of the list rather than between `Integer` and `String`. Two of these (backup_run.py, post.py) have been failing lint since |
||
|
|
389afe2f7b |
db: the doubled CHECK list was six, not four (#3275)
CI / extension-version (push) Successful in 6s
CI / lint (push) Failing after 6s
Build images / sign-extension (push) Successful in 6s
Build images / build-agent (push) Successful in 11s
CI / backend-lint-and-test (push) Successful in 31s
CI / frontend-build (push) Successful in 22s
Build images / build-ml (push) Successful in 52s
CI / integration (push) Successful in 3m43s
Build images / build-web (push) Successful in 41s
Run 5029 confirmed the four renames landed and surfaced two I had missed: external_link's host and status CHECKs are doubled the same way. They did not show in run 5026's diff because BOTH sides produced the doubled form back then — external_link.py pre-prefixed its names, so the models matched the chain's mistake. Switching all six models to bare names is what exposed the two the migration did not cover. The list in the file now comes from matching ck_(\w+?)_ck_\1_ against the chain's own pg_dump, rather than from reading migrations by eye. Reading by eye is what missed these, in the same way it earlier missed a UNIQUE constraint sitting two lines above the index being looked at. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017QHszn9H8VBvx5Ke8x1hvw |
||
|
|
b979062dd7 |
db: rename the four double-prefixed CHECK constraints (#3275)
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Failing after 5s
CI / extension-version (push) Successful in 5s
Build images / build-agent (push) Successful in 9s
CI / frontend-build (push) Successful in 20s
CI / backend-lint-and-test (push) Successful in 33s
Build images / build-ml (push) Successful in 44s
Build images / build-web (push) Successful in 41s
CI / integration (push) Successful in 3m52s
Run 5026 got the models-vs-chain diff to 7 lines. Three findings, and one
of them reverses an assumption I made in the previous commit.
The doubled CHECK names are what the DATABASE has, not what the generator
invented. base.py's convention is ck_%(table_name)s_%(constraint_name)s,
which — unlike uq/fk/ix — applies even to a constraint that already has a
name, so four migrations that passed an already-prefixed name got it
prefixed twice:
ck_import_settings_ck_import_settings_singleton
ck_ml_settings_ck_ml_settings_singleton
ck_post_ck_post_translation_override
ck_tag_ck_tag_fandom_requires_character
The workflow repair added last commit is still correct and still needed —
autogenerate really does re-double a name on the round trip — but it was
making the MODELS side clean against a chain that is dirty. The
comment in ml_settings.py claiming its bare name "matches migration 0003"
was simply false; 0003 produces the doubled form.
Nothing reads a CHECK constraint by name, so this has never done harm.
But it is precisely the development-era residue the collapsed baseline
exists to leave behind, and a public schema should not ship it — so 0088
renames the deployed constraints and all six models now declare bare
names. RENAME CONSTRAINT is catalog-only: no scan, no rewrite, no
revalidation, which is why this is safe on post and tag. Guarded on
pg_constraint scoped by conrelid, so it is a no-op on a database built
from the models.
ix_tag_fandom_id showed as a difference only because chain_ref was pinned
to
|
||
|
|
573228b9da |
db: finish reconciling the models with the deployed schema (#3275)
CI / lint (push) Failing after 3s
Build images / sign-extension (push) Successful in 4s
CI / extension-version (push) Successful in 3s
Build images / build-agent (push) Successful in 9s
CI / frontend-build (push) Successful in 34s
Build images / build-ml (push) Successful in 53s
Build images / build-web (push) Successful in 44s
CI / integration (push) Successful in 4m5s
CI / backend-lint-and-test (push) Successful in 1m6s
Closes the residue the first reconciliation pass left, and corrects a
factual error I put into the record.
sha256 was NOT missing a uniqueness guarantee. I read
`op.create_index("ix_image_record_sha256", ...)` at 0001 line 151 and
concluded duplicates were possible, without reading line 149 two lines
above it:
sa.UniqueConstraint("sha256", name="uq_image_record_sha256"),
Uniqueness has held since the initial schema. The database expresses it
as a CONSTRAINT plus a separate non-unique lookup index; the model said
`unique=True, index=True`, which is one UNIQUE index under a different
name. Same guarantee, different objects — which is exactly why the two
schemas did not line up. The model now declares both objects. No DDL.
0088's docstring, which repeated the claim, is corrected in place.
Two real divergences, both the MODEL over-claiming:
* source: uq_source_artist_platform_url (alembic 0010) was declared
nowhere in the models — source.py had no __table_args__ at all — so
autogenerate would have proposed DROPPING it.
* head_metrics_snapshot.tag_id: model said NOT NULL, 0060 created it
nullable. Left nullable; the FK already cascades.
Seven constraints renamed to what the chain actually created, rather than
what base.py's naming convention renders: uq_series_page_image,
uq_series_chapter_anchor_page, fk_series_chapter_anchor_page,
fk_image_record_artist_id, fk_image_provenance_from_attachment, and the
two hand-shortened fk_tsr_* names from 0003.
Float server_defaults now mirror their own migration, per column. The
chain is MIXED: a plain string renders DEFAULT '0.90'::double precision,
sa.text() renders DEFAULT 0.90, and the migrations used both. Seven
columns take text(); the rest stay strings. Two literals also disagreed
outright — process_{auto_apply,conflict}_threshold said 0.9/0.5 against
the migration's 0.90/0.50.
baseline.yml gains two things. A repair for a SECOND generator defect in
the same class as the missing pgvector import: base.py's ck convention
contains %(constraint_name)s, so it applies even to a NAMED
CheckConstraint — autogenerate writes the already-rendered name into the
migration and running it applies the convention again, yielding
ck_ml_settings_ck_ml_settings_singleton. That is round-tripping damage,
not a claim the models make, so it is undone rather than counted.
And the diff now runs twice. Column ORDER differs permanently between a
schema built by 87 ADD COLUMNs and one built in a single shot — the
operator's database keeps chain order forever, a fresh install gets model
order — so a check that failed on it could never pass. The second pass
SORTS column lines within each CREATE TABLE instead of deleting them,
which cannot hide a column present on one side only, or one whose type,
nullability or default differs. Ordered diff is reported as information;
the order-insensitive one is the verdict.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017QHszn9H8VBvx5Ke8x1hvw
|
||
|
|
d044e93bdb |
ci: repair autogenerate's missing pgvector import before applying (#3275)
CI / lint (push) Failing after 3s
Build images / sign-extension (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 8s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 38s
Build images / build-web (push) Successful in 13s
CI / integration (push) Successful in 3m52s
mode: models applies the raw autogenerated candidate, and it cannot run:
sa.Column('weights', pgvector.sqlalchemy.vector.VECTOR(dim=1152), ...)
NameError: name 'pgvector' is not defined
Alembic emits the qualified reference without emitting the import.
Observed on run 4988, which turns this from a thing I predicted by
reading the candidate into a thing demonstrated by executing it.
Repaired in the workflow rather than counted as a schema difference: the
comparison asks whether the MODELS describe the schema, and this is a
defect in the generator. The same fixup has to be applied by hand to any
baseline generated this way, which is why it is item 4 on the collapsed
baseline's hand-written list.
|
||
|
|
ed2b1adc2e |
ci: compare the schema the MODELS produce against the migrations (#3275)
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Failing after 4s
CI / extension-version (push) Successful in 4s
Build images / build-agent (push) Successful in 11s
Build images / build-ml (push) Successful in 47s
CI / frontend-build (push) Successful in 22s
CI / backend-lint-and-test (push) Successful in 35s
Build images / build-web (push) Successful in 37s
CI / integration (push) Successful in 3m53s
baseline.yml only ever compared migrations against migrations. The question #3275 exists because nobody had ever asked the other one: does a database built from the MODELS match the one the chain produces? `mode: models` answers it. It applies the candidate autogenerated from the models instead of this tree's revisions, and diffs that against the chain. A clean run means --autogenerate is trustworthy again, which it demonstrably has not been: against the pre-reconciliation models it would have proposed dropping eleven indexes and two uniqueness guarantees. The two extensions are created by hand in that mode. They are database objects rather than table metadata, so no model can carry them — their absence is outside what this comparison asks about, and silently tolerating it is correct rather than a filter that hides a defect. Also declares the HNSW index on the ImageRecord model. SQLAlchemy can express an hnsw access method with an operator class (postgresql_using + postgresql_ops), so there was never a reason for it to live only in 0036. That removes the last item from the list of things a generated baseline cannot reproduce, leaving only the two extensions. |
||
|
|
5e1996e77f |
db: reconcile the models with the deployed schema (#3275)
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Failing after 2s
CI / extension-version (push) Successful in 2s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 27s
Build images / build-ml (push) Successful in 48s
CI / backend-lint-and-test (push) Successful in 1m7s
Build images / build-web (push) Successful in 40s
CI / integration (push) Successful in 4m1s
Milestone 328's acceptance test compared a database built by the real
0001..0087 chain against one built from the models, and found ~130
places where they disagree. This closes them.
Almost all were the MODEL being wrong, so almost all of this is model
edits with no DDL — the database already had these things, nothing in it
changes, and no deploy is needed for this part:
* 92 columns gained server_default. The models carried Python-side
`default=` only, so the ORM filled the value and the column had no
database default. Anything inserting outside the ORM behaved
differently from production.
* Eleven indexes that existed only in migrations are now declared:
the three backup_run reporting indexes, the two date-ordered
image_record browse indexes, import_task and presentation_review,
and the three task_run history indexes. All use text() for their DESC
ordering and postgresql_where for the partial one.
* Two UNIQUE indexes that autogenerate silently proposed DROPPING,
because neither is expressible as a UniqueConstraint:
uq_tag_name_kind_fandom — an EXPRESSION index over
(name, kind, COALESCE(fandom_id, 0))
uq_post_artist_external_id_null_source — PARTIAL, WHERE source_id
IS NULL
post.py already had a comment describing the second one. The comment
was right; nothing declared it.
* The two external_link enum CHECKs (host, status) — rule 36 territory,
and absent from the model entirely.
* Two indexes were named explicitly. A bare index=True generated
ix_tag_alias_canonical_tag_id where the database has
ix_tag_alias_canonical, so autogenerate proposed a drop+create of an
index that was already there under another name. Same for
tag_suggestion_rejection.
Only ONE thing needed DDL, as 0088: tag.fandom_id is declared
index=True but no migration ever created that index.
Deliberately NOT here: image_record.sha256. The model says unique=True;
0001 created a plain index. Duplicates are possible today and the ORM
believes otherwise. The fix depends on whether duplicates already exist
— if they do, that is a dedupe decision, not a constraint — so it waits
on an answer about live data.
The real severity of #3275 is not the squash. It is that --autogenerate
has been unsafe on this project: run against the old models it would
have proposed dropping eleven indexes and two uniqueness guarantees.
|
||
|
|
98b56330d0 |
ci: emit the chain schema dump for local reconciliation work (#3275)
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 5s
Build images / build-agent (push) Successful in 8s
CI / frontend-build (push) Successful in 26s
Build images / build-ml (push) Successful in 44s
CI / backend-lint-and-test (push) Successful in 43s
Build images / build-web (push) Successful in 36s
CI / integration (push) Successful in 3m55s
Reconciling the models against the deployed schema needs the actual pg_dump, not an inference from the unified diff. Parsing table context out of diff hunks drops every table whose CREATE TABLE line falls outside a hunk — it under-reported 81 columns across 13 tables when the real figure spans more, missing artist, gpu_job, download_event and external_link entirely. Same checksummed-base64 transport as the candidate baseline, for the same reason: a plain cat of a file this size was silently truncated mid-line by the runner on run 4964. |
||
|
|
6959e1220c |
Revert "db: collapse alembic 0001..0087 into one baseline"
This reverts
|
||
|
|
2529b516e6 |
db: collapse alembic 0001..0087 into one baseline (milestone 328 step 1)
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 9s
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 5s
CI / frontend-build (push) Successful in 24s
Build images / build-ml (push) Successful in 42s
CI / backend-lint-and-test (push) Successful in 53s
Build images / build-web (push) Successful in 33s
CI / integration (push) Failing after 3m47s
87 revisions narrating this project's build-out become one file that creates the schema in a single step. They cost nothing at runtime — all 86 upgrade steps ran in 0.2s (note #3260) — so this is a presentation change, not a performance one: a new installer should not inherit our development history to stand up a database. Deleted: 87 revisions (6,052 lines), the 10 tests/test_migration_*.py files (483 lines) that asserted intermediate states and backfills which no longer exist, and backend/app/utils/artist_backfill.py — the only live module a migration imported, with no other consumer anywhere. That last one satisfies the operator's separate request to inline it into 0008 and delete the module; the squash removes both outright. THE REVISION ID IS "0087", NOT "0001", ON PURPOSE. It is the id of the last revision collapsed, so an existing database is already at head and `alembic upgrade head` does nothing. The alternative is `alembic stamp` against live data, and stamp validates NOTHING — it writes a version string whether or not the schema matches, so a wrong baseline surfaces later, via the next real migration, with no clean way back. This removes that operation rather than making it safe. Future revisions run from 0088. Four things are hand-written because SQLAlchemy metadata does not carry them, and none fail at generation time: 1. CREATE EXTENSION vector — the VECTOR columns cannot be created without it, so it is ordered first in upgrade(). 2. CREATE EXTENSION tsm_system_rows — surfaces only when the random sample query runs. 3. the HNSW index on image_record.siglip_embedding, raw SQL because create_index cannot express USING hnsw (... vector_cosine_ops). The quietest of the four: everything works, similarity search just stops using an index. 4. import pgvector.sqlalchemy.vector — autogenerate EMITS pgvector.sqlalchemy.vector.VECTOR references without importing it, so the generated file dies with NameError on first run. The candidate came out of CI (run 4967) as checksummed base64 rather than a plain cat, because run 4964's cat was truncated mid-line inside a column definition with the step still green — 29 tables instead of 42, and it looked entirely plausible. Verified here: 56,582 bytes, sha256 471acfca69c0…, 42 tables, 66 indexes, 42 drops. NOT YET PROVEN against the old chain. baseline.yml does that, and it is step 2's gate; this commit does not claim the schemas match. |
||
|
|
8f1ac0c96a |
ci: transport the candidate baseline as verifiable base64
CI / integration (push) Successful in 3m48s
CI / lint (push) Successful in 4s
Build images / sign-extension (push) Successful in 5s
CI / extension-version (push) Successful in 5s
Build images / build-ml (push) Successful in 8s
Build images / build-agent (push) Successful in 9s
Build images / build-web (push) Successful in 7s
CI / frontend-build (push) Successful in 19s
CI / backend-lint-and-test (push) Successful in 45s
Run 4964 passed the control (1121 normalised lines, schemas identical)
but its candidate print was silently truncated. `cat` of the ~33KB
generated file stopped mid-line inside
sa.Column('mime', sa.String(length=128)
and the runner carried straight on to the next traced command with the
step still green. The captured text was 484 lines and 29 tables, and
looked entirely plausible — which is exactly what makes it dangerous:
a schema definition cut in half is still syntactically suggestive, and
nothing in the log says it was cut.
Now emitted as base64 at a fixed 120-column width, followed by a
sha256, a byte count and a base64 line count. Short lines instead of
long ones, and more importantly the receiving end can PROVE it got the
whole file rather than trusting that it did.
Also found in that output, and the reason the candidate could never
have been committed as-is: it references
pgvector.sqlalchemy.vector.VECTOR(dim=1152)
for head_training_run.weights and image_record.siglip_embedding, but
autogenerate does not add the corresponding import. The file would die
with NameError on the first run. That is the fourth item on the list of
things the generator cannot be trusted with, alongside the two CREATE
EXTENSIONs and the HNSW index.
|
||
|
|
5fd171a544 |
ci: fix two things the baseline control run found
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 4s
Build images / sign-extension (push) Successful in 4s
Build images / build-web (push) Successful in 6s
CI / integration (push) Successful in 3m52s
Build images / build-ml (push) Successful in 8s
Build images / build-agent (push) Successful in 8s
CI / frontend-build (push) Successful in 23s
CI / backend-lint-and-test (push) Successful in 32s
Run 4960 was the control — the chain compared against itself, which must come back clean before a clean diff after the squash means anything. It did its job and failed on both counts. 1. The harness is sound. Both dumps came back 1123 normalised lines and differed on EXACTLY two, the \restrict / \unrestrict pair that newer pg_dump emits to fence a dump against injection during restore. It is a fresh random nonce per invocation, so it differs by construction and is noise by definition. Now filtered — and the control is what licenses that filter: it was OBSERVED to be the only false positive rather than assumed to be one, which matters for a check whose whole value is that its normalisation does not hide a real difference. 2. The candidate-baseline step never ran. `if: github.event.inputs .generate == 'true'` on a `type: boolean` input silently evaluated false — no diagnostic, step skipped, job carried on. The same `github.event.inputs` typing quirk build.yml already works around for force_build. Rather than fight the input typing, the gate is now the tree itself: skip if alembic/versions holds one file. That is the real question anyway — there is nothing to generate once the chain is collapsed — and it cannot be silently wrong the way an unevaluated expression can. Worth noting what the control also proved incidentally: the two schemas were byte-identical across 1123 lines despite being built by separate alembic runs into separate databases, so pg_dump's object ordering is stable enough to diff directly and no sort normalisation is needed. |
||
|
|
62583791d8 |
ci: a workflow that proves a collapsed alembic chain matches the old one
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 3s
Build images / build-agent (push) Successful in 9s
CI / frontend-build (push) Successful in 25s
Build images / build-ml (push) Successful in 7s
Build images / build-web (push) Successful in 8s
CI / backend-lint-and-test (push) Successful in 43s
CI / integration (push) Successful in 3m59s
Milestone 328 step 1 needs a baseline generated from the models, and
step 2 must not stamp the operator's live database until that baseline
is proven to reproduce what the 87-revision chain produced. `alembic
stamp` validates nothing, so an unproven baseline fails silently now and
loudly later, on real data.
There is no local Python environment and rules 10/12 point away from
standing one up, so the comparison runs in CI, where a pgvector Postgres
is already built from the chain on every integration run and nothing is
at risk.
It builds two databases and diffs their pg_dump --schema-only output:
one from `alembic upgrade head` on the revisions read out of git at
`chain_ref`, one from the current tree. Reading the chain from git via a
worktree — rather than from the working tree — is what keeps this usable
AFTER the old revisions are deleted, so it is the proof for step 1 and
the pre-flight for step 2 rather than a one-shot script.
Both sides use `alembic upgrade head`, never metadata.create_all, per
rule 82 — and that rule's reasoning is exactly the hazard here.
`create_all` emits plain CREATE TABLE and skips everything else, which is
why the optional autogenerated candidate CANNOT be trusted as the answer.
Three things in this schema are invisible to SQLAlchemy metadata:
CREATE EXTENSION vector (0001)
CREATE EXTENSION tsm_system_rows (0004)
the HNSW index on image_record.siglip_embedding, raw SQL because
alembic's create_index cannot express USING hnsw (...) (0036)
plus any CHECK constraint or server_default a migration added without the
model declaring it — 4 model files declare CheckConstraints against 6
migrations that touch them. The candidate is a starting point to hand
finish; the diff is what proves nothing was missed.
Results are printed to the job log rather than uploaded: ci-requirements
records that this runner cannot do actions/upload-artifact@v4+, and the
repo dropped the action entirely in 2026-05.
Run it first with the chain still present, as a control — the diff
compares the chain against itself and must come back clean. A clean diff
after the squash only means something if the harness was shown to be
capable of producing one beforehand.
Temporary. Delete once the baseline is stamped.
|
||
|
|
0a5bbe81dc |
docs: the scheduled refresh does NOT republish nothing (#3265)
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 5s
Build images / build-agent (push) Successful in 8s
CI / integration (push) Successful in 3m58s
Build images / build-ml (push) Successful in 7s
Build images / build-web (push) Successful in 8s
CI / frontend-build (push) Successful in 20s
extension / lint (push) Successful in 27s
CI / backend-lint-and-test (push) Successful in 52s
Step 4 asserted that when the base has not moved the refresh is "a ~13s
no-op that republishes nothing", and that this no-op was the point. The
first half is false and was written without being tested.
Run 4934, the first real fire: every content step reported CACHED and
both bases resolved to unchanged pinned digests, yet all three :latest
tags took a new manifest digest anyway.
fabledcurator 4ea5265ba017 -> 380e504de0fa
fabledcurator-ml 6e7cfc0c09fd -> 6b2eefc301d8
fabledcurator-agent 44920e0af1f3 -> 54accbeb52ed
buildkit mints a fresh image config per run, so identical layers get
republished under a new config blob. Storage cost is trivial; the cost
that matters is that a :latest digest change stops meaning "something is
different", and :c-<sha> is handed a new manifest to diverge from every
Sunday for no reason.
Corrects the workflow comment (x3) and ci-requirements.md to say what
actually happens. Filed as #3265 with the candidate fixes; the likely one
is a deterministic SOURCE_DATE_EPOCH off the value artifacts.sh already
derives, which would make "same source, same version" into "same source,
same bytes".
The rest of step 4 verified clean on the same run: the guard passed
(HEAD is main (
|
||
|
|
6663e06aa6 |
ci: assert the scheduled refresh actually checked out main
CI / extension-version (push) Successful in 5s
CI / lint (push) Successful in 6s
Build images / build-ml (push) Successful in 9s
Build images / build-web (push) Successful in 6s
CI / frontend-build (push) Successful in 18s
extension / lint (push) Successful in 19s
Build images / sign-extension (push) Successful in 6s
Build images / build-agent (push) Successful in 11s
CI / backend-lint-and-test (push) Successful in 39s
CI / integration (push) Successful in 3m48s
BUILD_REF is read through the `env` context inside `with:`, which this
runner is not known to evaluate. `${{ steps.* }}` and `${{ secrets.* }}`
in `with:`/`env:` are proven here; `env` is not, and run 4915's checkout
log (`git checkout -B dev refs/remotes/origin/dev`) cannot tell an
honoured `refs/heads/dev` from an empty value falling back to the same
place — the two are indistinguishable on every path except the one that
matters.
If it does resolve empty, the weekly refresh checks out dev and pushes
its source to :latest, which is production. Every lane stays green and
the first symptom is production running code that was never merged.
So each of the four jobs now asserts its own checkout before doing
anything, gated on `github.event_name` — the `github` context is
demonstrably evaluated in `if:`, so the guard cannot be disabled by the
same uncertainty it covers. A red weekly job is an acceptable outcome;
shipping dev to production is not.
|
||
|
|
63e0a423d7 |
ci: a weekly base-image refresh on the channel tags (milestone 326 step 4)
Build images / sign-extension (push) Successful in 4s
CI / extension-version (push) Successful in 3s
CI / lint (push) Successful in 3s
Build images / build-web (push) Successful in 6s
extension / lint (push) Successful in 20s
Build images / build-agent (push) Successful in 7s
Build images / build-ml (push) Successful in 8s
CI / frontend-build (push) Successful in 21s
CI / backend-lint-and-test (push) Successful in 31s
CI / integration (push) Successful in 3m50s
Skip-if-exists is keyed on our own source, so an artifact whose source stops moving stops picking up base-image updates. `agent/` last changed 2026-07-17; every push since has correctly declined to rebuild it, which also means it will serve that day's nvidia/cuda layers indefinitely. A `schedule:` trigger, Sunday 06:00 UTC, away from CI-runner's Monday security sweep so the two are never diagnosing each other. #3154's blocking open question is dissolved rather than answered. It was written when the identity was a `r-<revision>` TAG, and asked how the next ordinary push could avoid repointing :latest back off the refresh. Milestone 318 replaced that tag with a LABEL, and #3183 made the repoint step exclude its source tag so the label stays readable. Excluding the source is what also keeps a refresh from being undone: on the next main push the reuse check hits, :latest is not rewritten, and the new :c-<sha> is written FROM the refreshed :latest. To be verified by digest, not by this argument. Four decisions, each commented where it lives: * It builds `main`, not the branch that triggered it. Forgejo registers a cron from the default branch — `dev` here — so a scheduled run arrives with github.ref on dev, and a refresh of :dev would be refreshing the one channel that is rebuilt constantly anyway. The ref is decided once in a top-level `env: BUILD_REF` that all four checkouts take. Deriving it per job would let the halves disagree: sign-extension would derive dev's extension version while build-web bundled main's, and the release download would 404 on a version that exists perfectly well. * It publishes only the channel tag. :c-<sha> for main's HEAD already names the bytes that commit built; re-pushing it over refreshed layers would break the one tag rule 145 makes immutable, and it is the rollback unit — so the breakage would surface on the day somebody needed it. The repoint step needs no schedule case: the tag list is the channel tag alone, SOURCE is the only entry, it is excluded as always, and the step correctly does nothing. * It bypasses reuse by construction, since it rebuilds the same source and fc.revision always matches. Checked in the reuse step beside force_build, so one decision still drives both the build and the repoint. * `pull: true`, on the scheduled path only, is the actual mechanism. A moved base tag changes the FROM layer's cache key and everything above it rebuilds; an unmoved one is satisfied by the registry cache and the refresh is a ~13s no-op that republishes nothing. That no-op is the point — :latest should change when there is something new in it, not every Sunday. The known lag, left deliberately: an apt package update while the base tag stands still is not caught, and closing it needs no-cache: true, which buys weekly churn for it. |
||
|
|
499720d87e |
Merge pull request #242 from dev
Build images / sign-extension (push) Successful in 4s
CI / extension-version (push) Successful in 5s
CI / lint (push) Successful in 4s
Build images / build-agent (push) Successful in 10s
Build images / build-ml (push) Successful in 13s
Build images / build-web (push) Successful in 8s
extension / lint (push) Successful in 22s
CI / frontend-build (push) Successful in 25s
CI / backend-lint-and-test (push) Successful in 57s
CI / integration (push) Successful in 4m3s
ci: buildx driver, registry layer cache, and a force_build hatch (milestone 326 steps 1-3) |
||
|
|
5e72076298 |
ci: registry-backed layer cache for all three images (milestone 326 step 2)
Build images / sign-extension (push) Successful in 4s
CI / lint (push) Successful in 5s
CI / extension-version (push) Successful in 5s
Build images / build-ml (push) Successful in 8s
Build images / build-agent (push) Successful in 9s
extension / lint (push) Successful in 20s
Build images / build-web (push) Successful in 6s
CI / frontend-build (push) Successful in 17s
CI / backend-lint-and-test (push) Successful in 33s
CI / integration (push) Successful in 3m44s
extension / lint (pull_request) Successful in 29s
`cache-from`/`cache-to` on `<image>:buildcache`, `mode=max`, on all three
build steps. Closes the half of the driver change that step 1 left open.
Step 1 measured worse, not better, and that was expected but is worth stating
with numbers. Run 4896, first builds after moving to `docker-container`:
build-web 3m44s (cold baseline 2m23s)
build-ml 3m49s (cold baseline 3m20s)
build-agent 11m12s (cold baseline 9m26s)
The container driver gets a FRESH buildkit instance per job, so it has no
local layer store to fall back on — where the old docker driver at least
reused whatever the runner's dockerd happened to hold. That is why a registry
cache is the only cache this driver can have, and why step 1 on its own is a
regression rather than a win.
`mode=max` so intermediate stages cache too. The agent's two ~150s pip layers
and web's frontend-builder stage are the entire cost, and a min-mode cache
would drop exactly those.
A `:buildcache` tag is not the withdrawn tag scheme returning. Rule 145
narrowed against names NOTHING reads; this one is read by every build that
runs, is one moving ref per image rather than one per build, holds cache blobs
rather than a shippable artifact, and is overwritten in place rather than
accumulating. Closer to `:dev` than to the `:2026.8.28` tags 318 deleted — and
said so in the workflow, so it is not "cleaned up" by a later reader.
Expect the next build to be slower again, once: it is still cold AND now pays
the cache export. The measurement that matters is the one after that.
Scribe #3114.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
e21c9fdd34 |
ci: a force_build escape hatch for the path skip-if-exists hides (326 step 3)
CI / lint (push) Successful in 3s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 3s
Build images / build-ml (push) Successful in 7s
Build images / build-agent (push) Successful in 6s
CI / frontend-build (push) Successful in 20s
CI / backend-lint-and-test (push) Successful in 29s
Build images / build-web (push) Successful in 7s
extension / lint (push) Successful in 18s
CI / integration (push) Successful in 3m47s
`workflow_dispatch` with a `force_build` boolean, honoured inside each of the three reuse steps. It exists because skip-if-exists made its own build path untestable. `agent/` has not changed since 2026-07-17, so the agent build has correctly declined to run on every push since — which also means #3190, whose whole symptom lives on that path, cannot be reproduced on demand. Editing build.yml does not force a build either, and that is deliberate: the workflow is not shipped bytes, so it is in no artifact's path set, and putting it in one would re-version every artifact for a comment change. That is also why this lands before step 2 rather than after. Step 1 moved the builds onto a container driver and turned attestations off; the claim that `fc.revision` still reads back cannot be checked until something actually builds under that driver. Run 4887 confirmed only the cheaper half — `Set up buildx` succeeded on all three jobs, so the buildkit sibling container does start against the mounted socket. Details worth keeping: * FORCE is checked in the reuse step, not in the build step's `if:`. The repoint step keys off `hit` too, and a force that bypassed only the build would leave the two disagreeing about what had happened. * `github.event.inputs`, not the `inputs` context — release.yml already uses that form and it is the one this runner is known to evaluate. Read through env rather than interpolated into the run block, same as release.yml's TAG. * One input, not one per artifact. Three booleans is an interface nobody remembers. Scribe #3252, #3249. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |