TT Lab
Get started
Learn Learning paths Courses

Container Internals

Taking an Image Apart to Look Inside

Continue in TT Lab

This lab runs on a real VM

This box is not a Pod but a virtual machine started by KubeVirt. A Linux kernel of its own runs, systemd actually manages services, and docker is a real Docker engine, not an imitation. A container started with docker run becomes an actual process, and docker exec and docker logs work as usual.

This lab used to run inside a Pod. Since it was a box with all kernel capabilities dropped, the step of starting a container was blocked, so you learned through a workaround of unpacking the image archive yourself. The workaround is no longer needed.

There are two things to know.

Goal

You export an image as an archive, open the manifest and config blob inside yourself, and confirm by hand that RootFS.Layers (diff_id) and the layer digest in the manifest are values at different levels. At the end, you prove by comparing digests that two images built from the same base share their first layer.

Why it matters

Most image-related incidents happen "because you trusted a tag". A tag is a name a person attached and it can change at any time to point to a different digest. A digest, on the other hand, is a hash of the content, so the same hash guarantees the same content. If you confirm this distinction by hand, it becomes clear what to pin in a deployment pipeline. One more thing: the layer digest written in the manifest is the hash of the compressed blob, and the rootfs.diff_ids in the config is the hash of the uncompressed tar. If you do not know that the two differ, you spend a long time stuck in the illusion that "it is the same image but the hash is different". That is why this lab has you pull the two values out of different files.

Steps

  1. Create the /root/int3 directory and save the ID (content hash) of the alpine:3.20 image to /root/int3/image-id.txt. It must be 64 hexadecimal digits, and the sha256: prefix may be present.
  2. Save the rootfs layer digests of nginx:1.27-alpine to /root/int3/layers.txt, one per line. Every line must start with sha256: followed by 64 digits, and the number of lines must equal the actual number of layers.
  3. Export the alpine:3.20 image as an archive and save it as /root/int3/alpine.tar. docker save does not work on this box, so use skopeo copy oci-archive:/opt/images/alpine_3.20.tar oci-archive:/root/int3/alpine.tar:alpine:3.20 — it does the same job. The archive must contain manifest.json (or index.json/oci-layout) and hash-named blobs.
  4. Extract manifest.json from that archive and save it as /root/int3/manifest.json. It must be valid JSON and contain at least 1 layer in its layer list.
  5. Extract from the archive the config blob file that the Config field of the manifest from step 4 points to, and save it as /root/int3/config.json. This file must have architecture, rootfs.type must be layers, and rootfs.diff_ids must have at least 1 entry.
  6. Query the information for nginx:1.27-alpine in the local image storage with skopeo and save it to /root/int3/skopeo.json. The storage is empty, so first put it in with skopeo copy --insecure-policy oci-archive:/opt/images/nginx_1.27-alpine.tar containers-storage:docker.io/library/nginx:1.27-alpine — podman load is blocked on this box, but skopeo does not create namespaces, so it works as is. Name must contain nginx, Digest must be in the sha256: format, and Layers must have at least 1 entry.
  7. Build two images on top of alpine:3.20 as the base, each with one layer of different content added, tagged labhub/oci:a and labhub/oci:b. The two images must have the same first layer digest and different last layer digests.
  8. Write the following four lines in /root/int3/oci.md in exactly this format. All are based on nginx:1.27-alpine. Replace each placeholder in angle brackets with the actual value — the number of layers, the architecture, the operating system, and the first 12 characters of the first layer digest with the sha256: prefix removed.
    • layer_count=<레이어 개수>
    • architecture=<아키텍처>
    • os=<운영체제>
    • first_layer12=<첫 레이어 다이제스트에서 sha256: 을 뗀 앞 12자>

Notes

The image ID is the content hash

Create the /root/int3 directory and save the ID (content hash) of the alpine:3.20 image to /root/int3/image-id.txt. It must be 64 hexadecimal digits, and the sha256: prefix may be present.

The Id field of docker image inspect is that value. The sha256: prefix may or may not be there, but the hexadecimal part after it must be exactly 64 digits. This step is about saving a hash, not a tag name.

Extract the list of layer digests

Save the rootfs layer digests of nginx:1.27-alpine to /root/int3/layers.txt, one per line. Every line must start with sha256: followed by 64 digits, and the number of lines must equal the actual number of layers.

Unroll the RootFS.Layers array to one element per line and save it. If you expand the array with jq, it comes out without quotes. The number of lines must exactly equal the actual number of layers, so no blank lines or headers may be mixed in.

Export the image as an archive

Export the alpine:3.20 image as an archive and save it as /root/int3/alpine.tar. docker save does not work on this box, so use skopeo copy oci-archive:/opt/images/alpine_3.20.tar oci-archive:/root/int3/alpine.tar:alpine:3.20 — it does the same job. The archive must contain manifest.json (or index.json/oci-layout) and hash-named blobs.

The command that exports a single copy of a container's filesystem is different from the command that exports an image, including layers, tags, and history. You must use the latter for the manifest to be included.

Open the manifest inside the archive

Extract manifest.json from that archive and save it as /root/int3/manifest.json. It must be valid JSON and contain at least 1 layer in its layer list.

You can pull just one specific file out of a tar archive. Do not make up the JSON by hand; extract what is in the archive as it is. The extracted file must be valid JSON and must contain the layer list.

The config blob and diff_ids

Extract from the archive the config blob file that the Config field of the manifest from step 4 points to, and save it as /root/int3/config.json. This file must have architecture, rootfs.type must be layers, and rootfs.diff_ids must have at least 1 entry.

The Config field of the manifest is the name of another file in the archive. Only if you extract that file again do you get the real config that contains architecture and rootfs.diff_ids. If you copy the manifest itself, it fails.

Query the manifest with a registry tool

Query the information for nginx:1.27-alpine in the local image storage with skopeo and save it to /root/int3/skopeo.json. The storage is empty, so first put it in with skopeo copy --insecure-policy oci-archive:/opt/images/nginx_1.27-alpine.tar containers-storage:docker.io/library/nginx:1.27-alpine — podman load is blocked on this box, but skopeo does not create namespaces, so it works as is. Name must contain nginx, Digest must be in the sha256: format, and Layers must have at least 1 entry.

This environment is offline, so you cannot reach a remote registry. skopeo chooses its target by the transport prefix, and if you use the transport that points to the local image storage, the query works without a network. The output must show Name, Digest, and Layers together.

Prove that the base layer is shared

Build two images on top of alpine:3.20 as the base, each with one layer of different content added, tagged labhub/oci:a and labhub/oci:b. The two images must have the same first layer digest and different last layer digests.

The two images must start from the same base and add different content in one layer. If the bases differ, the first layers diverge, and if the content is the same, even the last layers become the same, so both fail.

Summarize the OCI metadata

Write the following four lines in /root/int3/oci.md in exactly this format. All are based on nginx:1.27-alpine. Replace each placeholder in angle brackets with the actual value — the number of layers, the architecture, the operating system, and the first 12 characters of the first layer digest with the sha256: prefix removed.

All four lines are in 키=값 format (key=value), and no spaces or quotes may be mixed in. first_layer12 is the first 12 characters of the first layer digest with sha256: removed. Everything is based on the nginx image.