TT Lab
Get started
Learn Learning paths Courses

Building Images

The Order in Which You Shrink an Image

Continue in TT Lab

One-line summary

The heart of multi-stage is not the syntax but what you copy into the final stage. The layers of the builder stage never even enter the final image's manifest.

Why this is needed

There was a 1.9 GB Node image. Each time nodes were added, it took 90 seconds for a Pod to become Ready, and another 90 seconds for each rollback. Instead of guessing, look at docker history first.

RUN npm install               901MB
COPY . .                      286MB
RUN apt-get update && ...     412MB
ADD file:07cf5b0f8bd1d5d5…     74.8MB

The source code cannot be 286 MB, so it means .git or a local node_modules went in whole. Measured once more with a tool, the waste is 408 MB and the efficiency is 71.8%. 408 MB is transferred and stored, yet it is not even visible from the container.

How it works

Here is the most counterintuitive fact. Even if you clean up like this, not a single byte shrinks.

0B      RUN rm -rf /var/lib/apt/lists/*
2.1MB   RUN apt-get purge -y build-essential && apt-get autoremove -y
18.4MB  RUN make -C /src all
396MB   RUN apt-get update && apt-get install -y build-essential

The purge layer not only failed to reduce the size but added 2.1 MB. This is because the delete marker and the updated package DB were recorded in the new layer. The 396 MB stays as it is.

The rule is one. Delete what you create within the same RUN that created it. But it must not lead to "so merge every RUN". That destroys cache reuse. What to merge is only the commands where creation and cleanup are paired.

The anti-pattern of multi-stage is also clear.

FROM node:22 AS builder
COPY . . ; RUN npm install && npm run build
FROM node:22-slim
COPY --from=builder /app /app     # devDeps, 소스, 테스트, 빌드 캐시 전부

It only swaps the base, so there is almost no saving. You have to pick and copy only the build outputs.

What it looks like in the field

If you choose a base by size alone, you lose out.

Base Size libc Caveats
debian:bookworm-slim about 75 MB glibc A safe default
python:3.12-slim about 130 MB glibc No compiler
alpine:3.20 about 8 MB musl Wheel rebuilds, DNS, malloc
distroless/static about 2 MB none Fails if you use cgo
scratch 0 B none Copy CA and tzdata yourself

A typical case where Alpine becomes a loss is Python wheels. The same pandas install takes 9.4 seconds on slim, but on alpine it falls back to compiling from source and takes 11 minutes 38 seconds. CI time is 70 times longer. And three things that are easy to forget when using scratch are CA certificates, timezone data and /etc/passwd. When TLS verification fails with x509: certificate signed by unknown authority, it is usually because of the first one.

Finally, we have to turn one direction around. What matters is layer reuse rate, not size. Of a 400 MB image, what you actually download may be 12 MB, and even a 200 MB "small" image transfers 200 MB every time if the dependency layer is invalidated on every commit. Verify the results not by image size but by the cold-start pull time.

Why a deleted file does not disappear from the image

Before learning multi-stage, a commonly used method is "install and then delete". But written this way, the image does not shrink by a single byte.

RUN apt-get install -y build-essential   # 레이어 A: 400MB 늘어남
RUN apt-get purge -y build-essential     # 레이어 B: 삭제 표시만 기록

Each RUN creates one layer, and layers only stack and cannot erase the earlier ones. B only holds a marker saying "make those files look absent", and A's 400 MB is deployed as it is and downloaded as it is. Moreover, those files can still be extracted — this is the path by which a secret that was briefly put in and deleted during a build remains in the image.

If you finish within one layer, it shrinks. If you bundle installation and deletion into a single RUN, all that is recorded in the layer is the state after that command finished.

RUN apt-get update  && apt-get install -y --no-install-recommends build-essential  && make install  && apt-get purge -y build-essential  && rm -rf /var/lib/apt/lists/*

Even so, multi-stage is better. The approach above pays the price of losing the entire cache. Even if the source changes by one character, it runs again from apt-get install. Multi-stage keeps the build-stage cache as it is, and puts into the final image only what you decided to bring with COPY --from. The point is that not the output but the boundary becomes clear.

Order the cache by how often things change. Copy the dependency list first and install, and copy the source after that. If you do the opposite, you download the dependencies anew every time you edit one line of source.

COPY go.mod go.sum ./
RUN go mod download        # go.mod 가 그대로면 캐시가 산다
COPY . .
RUN go build -o /app ./cmd/server

Choose the bottom image by what is needed to run. A statically linked Go binary runs on scratch too. If you use TLS you need ca-certificates, and if you look up users you need /etc/passwd. Alpine is small, but it uses musl libc, so a binary built assuming glibc can silently behave differently. For languages like Python that compile extension modules, it is common for Alpine to make the image larger and the build slower.

What you will do in the next lab

You will create the output in a builder stage and copy only that into the final stage to measure how much the size shrinks, and check that an image with just one static binary on top of scratch really works (and that it has no shell).