TT Lab
Get started
Learn Learning paths Courses

CKAD — Kubernetes Application Developer

Why a Pod Is Not One Container

Continue in TT Lab

In one line

A Pod is not "a box that holds containers." It is the smallest deployment unit for a group of processes that run on the same node, share the same network namespace and volumes, and start and die together. Once you accept this definition, it explains at a glance why init containers and sidecars look the way they do.

Why this was needed

If you stick to the principle of one container = one process, you soon hit a wall. A web server writes its access log to a file, but the log collector is a separate process. An app needs a configuration file to start, but that file has to be downloaded right before startup. If you push all of this into one image, the image gets bloated, and you have to rebuild the app image every time you change the log collector.

If you split them into completely separate Pods instead, two things break. They cannot see the same filesystem, and their scheduling can diverge. If the log collector lands on a different node, there is no file for it to read.

So Kubernetes introduced an intermediate layer. It separated the unit of scheduling from the unit of execution: what gets scheduled is the Pod, and what gets run is the container. The containers in a Pod are placed on the same node, call each other through localhost, and see the same directory through an emptyDir volume.

How it works

A Pod spec has two places to put containers.

Slot When it runs If it fails
spec.initContainers One at a time, in order, until each finishes Neither the next init container nor the main containers start
spec.containers All at once, after every init container succeeds Restarted according to restartPolicy

An init container exists to "complete." It does things like downloading configuration, running a DB migration, or waiting for a prerequisite service. If you list three of them, the first must finish before the second starts. If any one of them never finishes, the Pod stays in a state such as Init:1/3 forever, and the main containers never even start.

There was a long-standing problem here. A sidecar (a log collector or a proxy) has to "stay up," so it could not go in the init slot. But if you put it in containers, it starts at the same time as the main container, which creates a race condition: the collector comes up only after the main container has already started writing logs. In Job Pods it was worse, because the sidecar kept running even after the main container finished, so the Job never completed.

Native sidecars cleaned this up. You put the container in the initContainers list and give only that container restartPolicy: Always.

spec:
  initContainers:
    - name: logger
      image: busybox:1.36
      restartPolicy: Always     # ← 이 한 줄이 네이티브 사이드카로 만든다
      command: ["sh", "-c", "tail -F /var/log/nginx/access.log"]
  containers:
    - name: web
      image: nginx:1.27

With this, three things hold at once. The container starts before the main container, keeps running, and, in a Job, terminates together with the main container when it finishes.

It also helps to pin down the pattern names. A sidecar assists the main container's function (logs, metrics). An ambassador is a proxy that goes out on the main container's behalf when it reaches outside (connect to localhost:6379 and it routes to the real cluster). An adapter converts the main container's output into the format the outside world wants (app logs → Prometheus metrics).

What it looks like in the field

This happened when I installed the NVIDIA GPU Operator on my seven-node homelab cluster. Right after the installation, every GPU-related Pod was in this state.

gpu-feature-discovery-fpw5l            0/1   Init:0/1
nvidia-dcgm-exporter-gmlvl             0/1   Init:0/1
nvidia-device-plugin-daemonset-k4ggb   0/1   Init:0/1
nvidia-operator-validator-krj69        0/1   Init:0/4

0/1 and 0/4 mean that none of the init containers had passed. The main containers are not even mentioned. When I looked at the events, the cause was not inside the containers but outside them.

Warning FailedCreatePodSandBox  desc = failed to get sandbox runtime:
        no runtime for "nvidia" is configured

It was blocked at the step that creates the sandbox to hold the Pod. The container did not fail; the step that creates the network and IPC namespaces the containers share was what failed. The definition of a Pod as a "group" shows up exactly here. If there is no place to create the group, none of the containers in it can start.

One more case. When I installed KubeVirt on the same cluster, every component status was AllComponentsReady, yet the VM would not start. When I took apart the virt-launcher Pod spec, the volume mount holding the binary the init container was supposed to run was missing. The init container could not finish, so the main container waited forever. This was the third time on this cluster alone that I confirmed that "the status is Ready" and "it actually works" are different claims.

What you will do in the next lab

In the ckad-design namespace, you create Pods, labels, annotations, command overrides, environment variables, and restartPolicy yourself, and fill in a Job's completions/parallelism/backoffLimit and a CronJob's schedule/concurrencyPolicy/startingDeadlineSeconds by hand. In the lab that follows, you build by hand, in the ckad-multi namespace, the init container order, a native sidecar, the ambassador and adapter patterns, a shared emptyDir, and terminationGracePeriodSeconds.