docs: the merged image's cost, measured — and it corrects my own estimate (4296)
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 24s
CI / backend-lint-and-test (push) Successful in 32s
Build images / build-ml (push) Successful in 2m6s
Build images / build-web (push) Successful in 1m59s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / integration (push) Successful in 2m16s
CI / lint (push) Successful in 4s
CI / extension-version (push) Successful in 3s
Build images / sign-extension (push) Successful in 4s
Build images / build-agent (push) Successful in 7s
CI / frontend-build (push) Successful in 24s
CI / backend-lint-and-test (push) Successful in 32s
Build images / build-ml (push) Successful in 2m6s
Build images / build-web (push) Successful in 1m59s
Build images / smoke-web (push) Skipped
Build images / promote (push) Skipped
CI / integration (push) Successful in 2m16s
I wrote in ffcd130's Dockerfile comment that merging ML in means "everyone
pulls it, including the many who will never turn tagging on", framed as a
real cost the milestone accepted. My working estimate behind that was ~4GB.
Measured from run 7273's build-web log:
torch 2.12.1+cpu wheel 192.3 MB
torchvision 0.27.1+cpu 1.8 MB
transformers / onnxruntime / opencv / sklearn and friends
62.0, 35.3, 23.6, 16.7, 12.3, 9.2, 6.9 MB
largest newly-pushed layer 222.07 MB
The ML code adds a few HUNDRED MB, not gigabytes. The `--index-url` CPU
resolution is what makes that true — the default PyPI torch wheel carries the
CUDA runtime and is ~2GB by itself, and the log confirms 2.12.1+cpu resolved,
so it is working as intended rather than as intended-but-unverified.
Why this matters beyond a comment being wrong: it settles the trade this step
was explicitly asked to weigh and could not, and it reverses how close the
call looked. Baking the weights in adds ~3.5GB to every pull for a feature
many adopters never enable; shipping the code and fetching on demand adds
~350MB. An order of magnitude, where the estimate had them within 15% of each
other. Off-by-default is not a judgement call here, it is arithmetic.
The gigabytes were always in the MODEL, and the model is not in the image.
Two things NOT measured, still: the total image size (the push only transfers
layers the registry lacks, so a push log cannot give it) and the per-slot
resident RAM, which stays flagged `measured=False` in the lane table and
renders as "about" in the UI.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjrnpQjRgHdvq95rASoiR
This commit is contained in:
+28
-12
@@ -51,19 +51,35 @@ RUN pip install -r requirements.txt
|
|||||||
# in one process tree, a second image would mean the `ml` lane could never be
|
# in one process tree, a second image would mean the `ml` lane could never be
|
||||||
# enabled from the UI — there would be no worker in this container to enable.
|
# enabled from the UI — there would be no worker in this container to enable.
|
||||||
#
|
#
|
||||||
# The COST, stated because it is real and falls on every adopter: this adds
|
# THE COST, MEASURED from run 7273 rather than guessed — and it is far
|
||||||
# torch, torchvision, transformers, onnxruntime and opencv to an image that
|
# smaller than the estimate this comment first carried, which said "everyone
|
||||||
# previously carried none of them. Everyone pulls it, including the many who
|
# pulls ~4GB":
|
||||||
# will never turn tagging on. That is the trade the milestone accepted for
|
|
||||||
# being able to offer the lane as a switch rather than a second deployment.
|
|
||||||
# What it buys back is that nothing downloads a MODEL until the switch is
|
|
||||||
# thrown — the weights are not baked in, and rule 164 permits that only
|
|
||||||
# because the feature is optional and clearly off.
|
|
||||||
#
|
#
|
||||||
# CPU-only torch from the PyTorch CPU index. The default PyPI wheel bundles
|
# torch 2.12.1+cpu wheel 192.3 MB
|
||||||
# the NVIDIA CUDA runtime (~5.6GB of layer) and nothing here uses a GPU — the
|
# torchvision 0.27.1+cpu 1.8 MB
|
||||||
# GPU agent is a separate service with its own image. `--index-url`, not
|
# transformers / onnxruntime / opencv / sklearn and friends
|
||||||
# `--extra-index-url`: the latter would let pip resolve a +cu wheel anyway.
|
# 62.0, 35.3, 23.6, 16.7, 12.3, 9.2, 6.9 MB
|
||||||
|
# largest newly-pushed layer 222.07 MB
|
||||||
|
#
|
||||||
|
# So the ML code adds a few hundred MB to the pull, not gigabytes. The CPU
|
||||||
|
# index is what makes that true: the default PyPI torch wheel bundles the
|
||||||
|
# NVIDIA CUDA runtime and is ~2GB on its own.
|
||||||
|
#
|
||||||
|
# The GIGABYTES are in the MODEL — ~3.5GB of SigLIP weights — and those are
|
||||||
|
# NOT in this image. They arrive only when the operator enables the lane,
|
||||||
|
# which is what lets rule 164 permit a runtime fetch at all ("optional and
|
||||||
|
# clearly off"). That also settles the trade this step was asked to weigh:
|
||||||
|
# baking the weights in would add ~3.5GB to every pull for a feature many
|
||||||
|
# adopters never enable, against ~350MB for the code that makes the switch
|
||||||
|
# available. Off-by-default wins by an order of magnitude, which was NOT
|
||||||
|
# obvious before measuring — the estimate had the two costs within 15% of
|
||||||
|
# each other.
|
||||||
|
#
|
||||||
|
# `--index-url`, not `--extra-index-url`: the latter would let pip resolve a
|
||||||
|
# +cu wheel anyway, and the whole saving above depends on it not doing that.
|
||||||
|
#
|
||||||
|
# CPU-only torch from the PyTorch CPU index. Nothing here uses a GPU — the
|
||||||
|
# GPU agent is a separate service with its own image.
|
||||||
RUN pip install --index-url https://download.pytorch.org/whl/cpu \
|
RUN pip install --index-url https://download.pytorch.org/whl/cpu \
|
||||||
"torch>=2.12,<3.0" "torchvision>=0.27,<0.28"
|
"torch>=2.12,<3.0" "torchvision>=0.27,<0.28"
|
||||||
RUN pip install -r requirements-ml.txt
|
RUN pip install -r requirements-ml.txt
|
||||||
|
|||||||
Reference in New Issue
Block a user