TT Lab
Get started
Learn Learning paths Courses

Container Internals

What Is Inside When You Open One Image

Continue in TT Lab

In one line

An OCI image consists of three layers of JSON — Image Index → Image Manifest → Image Config — and the layer blobs beneath them, and every piece is addressed by the SHA-256 of its own contents.

Why this matters

If you do not know what happens behind a single docker pull, you cannot explain why a pull pinned to a digest fails, why tags get tangled when you make the registry redundant, or why the disk grew only a little even though you uploaded ten images. If you open the three-layer structure just once, all these questions land on the same picture.

How it works

The Image Index (fat manifest) is a list of per-platform manifests. The accident where you pull on an arm64 Mac and get an amd64 image is decided right here.

The Image Manifest holds one config digest, a layer list whose order is guaranteed, and the mediaType of each entry. The mediaTypes you often see are these.

application/vnd.oci.image.index.v1+json
application/vnd.oci.image.manifest.v1+json
application/vnd.oci.image.layer.v1.tar+gzip

The Image Config is the configuration the runtime actually reads. It contains values such as Env, Entrypoint, Cmd, and WorkingDir, along with rootfs.diff_ids and the history.

This is where the most frequently confused distinction comes in.

The two are different values. What you see in RootFS.Layers of docker image inspect is the diff_id side, and the digest you see in a registry is the compressed side. If you do not know this, you fall into the illusion that "it is the same image but the hash is different".

The benefits of content addressing boil down to three: deduplication (identical content is stored only once), integrity verification (hash the received bytes and compare with the name), and immutability (the same hash means the same content).

What it looks like in the field

The difference between a digest and a tag is half of real-world work. A digest is a hash of the content, so there is no collision and no cache invalidation problem. A tag is a name a person attached, and it can change at any time to point to a different digest. This is where the state is, and every hard part of a distributed system is here. This is why, when you try to bind several registries together, it ends up converging on the problem of serializing tag updates.

You also have to be careful with format conversion. Moving from the Docker schema to OCI changes the bytes of the manifest, and when the bytes change, the digest changes. Then pulls pinned by digest and signature verification both break. That is the right answer to "the content is the same, so why doesn't it work?"

Layer sharing is clear in numbers. If 10 images use a base built with node on top of ubuntu:22.04, what gets stored is 1x the base + 10x the app layers. So unifying the base is in effect a registry capacity policy.

Fewer layers, in a stable order

Knowing the image structure explains why some Dockerfiles build in seconds and others take minutes every time. The cache works per layer, and when one layer changes, every layer after it is rebuilt.

That gives the principle of ordering. Put what rarely changes first and what changes often last. This is why the standard pattern is to copy only the dependency list first and install, and then copy the whole source. If you copy the source first, changing a single character reruns everything from the dependency install.

The number of layers is itself a cost. Each layer carries metadata, requests are split when pulling, and the number of stacked filesystem layers grows. However, if you blindly merge everything into one line, the cache is wholly invalidated, so the criterion is to group things that change together.

And a deleted file does not disappear from the image. If you delete in a later layer a file created in an earlier layer, it is gone in the overlaid view, but it remains as is in the earlier layer and counts toward the size. This is exactly the same structure as the secrets problem seen in the previous module, and in terms of size the same conclusion follows. You must create and delete within the same layer, or move only the output with a multi-stage build.

A multi-stage build solves this problem most cleanly. Keep the build tools and intermediate artifacts in the earlier stages, and copy only what is needed to run into the final stage. The compiler and caches drop out entirely, so the image becomes several times smaller and the attack surface shrinks as well — a tool that is not in the final image cannot be used by an intruder either.

Finally, you have to decide whether to pin the base by digest or leave it as a tag. Pinning by digest gives perfect reproducibility, but security updates do not arrive automatically. Leaving it as a tag brings updates along, but yesterday's and today's builds may use different things. The standard is to pin it and set up a separate mechanism that proposes updates automatically.

What you will do in the next lab

You confirm that the image ID is the content hash, export an image to a tar, and open the manifest and config blob inside yourself. You look at the same information again through the eyes of a registry tool with skopeo, and prove by comparing digests that two images built from the same base share their first layer.