KCNA — Kubernetes and Cloud Native Associate
From One Pod to Workload Controllers
One-line summary
A Pod is not a unit of deployment but a unit of scheduling. You almost never create one by hand; in practice a controller creates it for you. Which controller you choose depends on "when does this work finish?"
Why this was needed
If a single container were the smallest unit, it would be unclear where to put helper processes such as log collectors or proxies. If you cram them into the same image, the 12-factor rule of "one concern per process" breaks, and if you put them on another machine, they cannot communicate over localhost.
A Pod is the answer in between: a bundle of containers that share a network namespace (the same IP and port space) and volumes. So a sidecar can reach the app over localhost, and the app and the sidecar are always scheduled together on the same node.
But a Pod itself does not bring itself back to life. If the node dies, the Pods on it disappear with it. That is why you need a controller on top of a Pod.
How it works
The chain of ownership
Deployment --(소유)--> ReplicaSet --(소유)--> Pod
When you create a Deployment, the controller manager creates a ReplicaSet, and the ReplicaSet creates the Pods. The parent is written into each object's metadata.ownerReferences, and this field is the basis of garbage collection. When you delete a Deployment, the ReplicaSet and Pods are cleaned up in a cascade following ownerReferences.
Why are there several ReplicaSets after you roll out a Deployment? Because the previous version's ReplicaSet is kept. That is why kubectl rollout undo is possible. Rolling back is nothing more than raising the replicas of the old ReplicaSet again.
Choosing a workload controller
| Controller | When to use | Does it finish? |
|---|---|---|
| Deployment | Stateless services. Web, API | It does not finish |
| StatefulSet | Things that need stable names, ordering, and dedicated storage. DBs, queues | It does not finish |
| DaemonSet | One per node. Log collectors, CNI, node exporters | It does not finish |
| Job | Work that runs once and finishes. Migrations, batches | It finishes |
| CronJob | Creates a Job at a set time | The Job finishes |
The key test is "does this process finish by itself?" If you use a Deployment for work that finishes, the container is restarted every time it exits, making an infinite loop. That is why a Job's Pod template cannot set restartPolicy to Always.
A DaemonSet has no replicas field. That is because a person does not decide the count; the number of nodes is the count. If you add a node, one more appears automatically.
Self-healing is not magic
The reason a Pod comes back when you delete it is that the ReplicaSet controller runs like this.
- How many Pods currently match my selector?
- How many is spec.replicas?
- If there are too few, create; if there are too many, delete
It is not a Pod with the same name coming back; a new Pod is created. The name is different every time. That is why a design that depends on Pod names always breaks.
Probes
- livenessProbe: If it fails, the container is restarted. "It's dead, so start it again."
- readinessProbe: If it fails, the Pod is taken out of the Service endpoints. "It's alive but can't take traffic right now."
- startupProbe: Delays the start of liveness/readiness for an app that starts slowly.
If you do not attach readiness, traffic gets driven into a Pod that is still booting, and if you set liveness too aggressively, an app that is only briefly slow keeps being restarted and the situation gets worse.
What it looks like in practice
Right after installing Cilium in the author's homelab, the hubble-relay and hubble-ui Pods stayed in Pending. The event was 0/1 nodes are available: 1 node(s) had untolerated taint(s).
The cause was the difference in controller type. The control plane node has the node-role.kubernetes.io/control-plane:NoSchedule taint, and the hubble components are Deployments, not DaemonSets, so they do not tolerate this taint. CoreDNS came up fine at the same time because CoreDNS has a control-plane toleration by default. This was not an error but normal behavior, and it was resolved as soon as a worker node joined.
One more thing. As the same cluster grew to 7 nodes, it came to have 4 GPU workers (an RTX 3090 24GB, an RTX 5090 32GB, and two RTX 4070 Laptop 8GB), and if a Pod requested just nvidia.com/gpu: 1, a training job that needed the 32GB 5090 could land on an 8GB laptop GPU. From Kubernetes' point of view, both are "1 GPU." In the end, the team attached meaning-based labels such as gpu.homelab/tier=xlarge|large|small by hand and had workloads pick them with nodeSelector. It is a case showing that the same resource name does not mean the same resource, and that labels are what bridge that gap.
What you will do in the next lab
In the next lab you start your first Pod, create and scale a Deployment, and check a ReplicaSet's ownerReferences yourself. You create a Job, a CronJob, and a DaemonSet one by one, and at the end you delete a Pod on purpose and prove by comparing the lists before and after deletion that self-healing really runs.