From 0f98e46200870c469208a3cd5a24102c098a1e7a Mon Sep 17 00:00:00 2001 From: Bryan Van Deusen Date: Tue, 22 Sep 2026 09:02:47 -0400 Subject: [PATCH] =?UTF-8?q?docs:=20the=20merged=20image's=20cost,=20measur?= =?UTF-8?q?ed=20=E2=80=94=20and=20it=20corrects=20my=20own=20estimate=20(4?= =?UTF-8?q?296)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit I wrote in ffcd130's Dockerfile comment that merging ML in means "everyone pulls it, including the many who will never turn tagging on", framed as a real cost the milestone accepted. My working estimate behind that was ~4GB. Measured from run 7273's build-web log: torch 2.12.1+cpu wheel 192.3 MB torchvision 0.27.1+cpu 1.8 MB transformers / onnxruntime / opencv / sklearn and friends 62.0, 35.3, 23.6, 16.7, 12.3, 9.2, 6.9 MB largest newly-pushed layer 222.07 MB The ML code adds a few HUNDRED MB, not gigabytes. The `--index-url` CPU resolution is what makes that true — the default PyPI torch wheel carries the CUDA runtime and is ~2GB by itself, and the log confirms 2.12.1+cpu resolved, so it is working as intended rather than as intended-but-unverified. Why this matters beyond a comment being wrong: it settles the trade this step was explicitly asked to weigh and could not, and it reverses how close the call looked. Baking the weights in adds ~3.5GB to every pull for a feature many adopters never enable; shipping the code and fetching on demand adds ~350MB. An order of magnitude, where the estimate had them within 15% of each other. Off-by-default is not a judgement call here, it is arithmetic. The gigabytes were always in the MODEL, and the model is not in the image. Two things NOT measured, still: the total image size (the push only transfers layers the registry lacks, so a push log cannot give it) and the per-slot resident RAM, which stays flagged `measured=False` in the lane table and renders as "about" in the UI. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR --- Dockerfile | 40 ++++++++++++++++++++++++++++------------ 1 file changed, 28 insertions(+), 12 deletions(-) diff --git a/Dockerfile b/Dockerfile index 292a087..1cffc96 100644 --- a/Dockerfile +++ b/Dockerfile @@ -51,19 +51,35 @@ RUN pip install -r requirements.txt # in one process tree, a second image would mean the `ml` lane could never be # enabled from the UI — there would be no worker in this container to enable. # -# The COST, stated because it is real and falls on every adopter: this adds -# torch, torchvision, transformers, onnxruntime and opencv to an image that -# previously carried none of them. Everyone pulls it, including the many who -# will never turn tagging on. That is the trade the milestone accepted for -# being able to offer the lane as a switch rather than a second deployment. -# What it buys back is that nothing downloads a MODEL until the switch is -# thrown — the weights are not baked in, and rule 164 permits that only -# because the feature is optional and clearly off. +# THE COST, MEASURED from run 7273 rather than guessed — and it is far +# smaller than the estimate this comment first carried, which said "everyone +# pulls ~4GB": # -# CPU-only torch from the PyTorch CPU index. The default PyPI wheel bundles -# the NVIDIA CUDA runtime (~5.6GB of layer) and nothing here uses a GPU — the -# GPU agent is a separate service with its own image. `--index-url`, not -# `--extra-index-url`: the latter would let pip resolve a +cu wheel anyway. +# torch 2.12.1+cpu wheel 192.3 MB +# torchvision 0.27.1+cpu 1.8 MB +# transformers / onnxruntime / opencv / sklearn and friends +# 62.0, 35.3, 23.6, 16.7, 12.3, 9.2, 6.9 MB +# largest newly-pushed layer 222.07 MB +# +# So the ML code adds a few hundred MB to the pull, not gigabytes. The CPU +# index is what makes that true: the default PyPI torch wheel bundles the +# NVIDIA CUDA runtime and is ~2GB on its own. +# +# The GIGABYTES are in the MODEL — ~3.5GB of SigLIP weights — and those are +# NOT in this image. They arrive only when the operator enables the lane, +# which is what lets rule 164 permit a runtime fetch at all ("optional and +# clearly off"). That also settles the trade this step was asked to weigh: +# baking the weights in would add ~3.5GB to every pull for a feature many +# adopters never enable, against ~350MB for the code that makes the switch +# available. Off-by-default wins by an order of magnitude, which was NOT +# obvious before measuring — the estimate had the two costs within 15% of +# each other. +# +# `--index-url`, not `--extra-index-url`: the latter would let pip resolve a +# +cu wheel anyway, and the whole saving above depends on it not doing that. +# +# CPU-only torch from the PyTorch CPU index. Nothing here uses a GPU — the +# GPU agent is a separate service with its own image. RUN pip install --index-url https://download.pytorch.org/whl/cpu \ "torch>=2.12,<3.0" "torchvision>=0.27,<0.28" RUN pip install -r requirements-ml.txt