Release: dev → main (first public release) #258

Merged
bvandeusen merged 94 commits from dev into main 2026-09-25 10:02:40 -04:00
Showing only changes of commit 0f98e46200 - Show all commits
+28 -12
View File
@@ -51,19 +51,35 @@ RUN pip install -r requirements.txt
# in one process tree, a second image would mean the `ml` lane could never be
# enabled from the UI — there would be no worker in this container to enable.
#
# The COST, stated because it is real and falls on every adopter: this adds
# torch, torchvision, transformers, onnxruntime and opencv to an image that
# previously carried none of them. Everyone pulls it, including the many who
# will never turn tagging on. That is the trade the milestone accepted for
# being able to offer the lane as a switch rather than a second deployment.
# What it buys back is that nothing downloads a MODEL until the switch is
# thrown — the weights are not baked in, and rule 164 permits that only
# because the feature is optional and clearly off.
# THE COST, MEASURED from run 7273 rather than guessed — and it is far
# smaller than the estimate this comment first carried, which said "everyone
# pulls ~4GB":
#
# CPU-only torch from the PyTorch CPU index. The default PyPI wheel bundles
# the NVIDIA CUDA runtime (~5.6GB of layer) and nothing here uses a GPU — the
# GPU agent is a separate service with its own image. `--index-url`, not
# `--extra-index-url`: the latter would let pip resolve a +cu wheel anyway.
# torch 2.12.1+cpu wheel 192.3 MB
# torchvision 0.27.1+cpu 1.8 MB
# transformers / onnxruntime / opencv / sklearn and friends
# 62.0, 35.3, 23.6, 16.7, 12.3, 9.2, 6.9 MB
# largest newly-pushed layer 222.07 MB
#
# So the ML code adds a few hundred MB to the pull, not gigabytes. The CPU
# index is what makes that true: the default PyPI torch wheel bundles the
# NVIDIA CUDA runtime and is ~2GB on its own.
#
# The GIGABYTES are in the MODEL — ~3.5GB of SigLIP weights — and those are
# NOT in this image. They arrive only when the operator enables the lane,
# which is what lets rule 164 permit a runtime fetch at all ("optional and
# clearly off"). That also settles the trade this step was asked to weigh:
# baking the weights in would add ~3.5GB to every pull for a feature many
# adopters never enable, against ~350MB for the code that makes the switch
# available. Off-by-default wins by an order of magnitude, which was NOT
# obvious before measuring — the estimate had the two costs within 15% of
# each other.
#
# `--index-url`, not `--extra-index-url`: the latter would let pip resolve a
# +cu wheel anyway, and the whole saving above depends on it not doing that.
#
# CPU-only torch from the PyTorch CPU index. Nothing here uses a GPU — the
# GPU agent is a separate service with its own image.
RUN pip install --index-url https://download.pytorch.org/whl/cpu \
"torch>=2.12,<3.0" "torchvision>=0.27,<0.28"
RUN pip install -r requirements-ml.txt