TT Lab
Get started
Learn Learning paths Courses

CKA — Kubernetes Administrator

Why the Control Plane Split Into Four

Continue in TT Lab

Summary

The Kubernetes control plane is split into four processes: etcd, kube-apiserver, kube-controller-manager, and kube-scheduler, and the only one that talks to etcd is the apiserver. That one line explains half of CKA troubleshooting.

Why this matters

Imagine the scheduler connecting directly to etcd, reading Pod objects, and writing in nodeName. Here is what would happen.

That is why Kubernetes made the apiserver the single gateway. The apiserver does not create state. It receives, validates, authorizes, stores, and notifies the watchers. All the real decisions are made by the controllers outside it.

How it works

A controller's reconcile loop looks like this.

  1. Read the declared state (spec) with a watch
  2. Observe the actual state (status)
  3. Compute the difference and call the API only for that difference
  4. Go back to step 1

The key point is that it reduces a difference rather than executing a command. So even if you miss one event, the next resync converges, and even if you restart a controller, it reconciles from scratch. The Deployment controller creates a ReplicaSet, and the ReplicaSet controller creates Pods. Each looks only at its own layer.

The scheduler works in two phases.

Phase What it does Result
filter (predicate) Removes nodes with insufficient resources, untolerated taints, nodeSelector mismatches, or volume zone mismatches A list of nodes the Pod can be placed on
score (priority) Scores the remaining nodes (resource balance, image locality, topology spread) The single highest-scoring node

A message such as 0/3 nodes are available: 3 Insufficient cpu is a tally of the reasons nodes were eliminated in the filter phase. If you can read this sentence, 90% of Pending Pods are solved on the spot.

What it looks like in the field

Case 1 — DHCP killed the cluster. The IP of a homelab control plane once changed from 10.0.0.111 to 10.0.0.120 because of a DHCP lease renewal. The symptom was dial tcp 10.0.0.111:6443: connect: no route to host, and the real cause was in the certificate. The apiserver certificate's SAN had only IP Address:10.0.0.111 and lacked .120. The interesting part was the difference in reaction between components. etcd and kube-apiserver tried to bind to the nonexistent .111 and fell into CrashLoopBackOff, while kube-scheduler and controller-manager bind to 127.0.0.1, so they were in a state where the process was alive but could do nothing. You can read this picture only if you know the design in which the four processes each attach to a different address.

Case 2 — etcd quorum and API availability are separate. While growing the same homelab from 3 nodes to 7, I made the control plane 3 machines. etcdctl member list showed exactly 3 members, and each node ran one apiserver, scheduler, and controller-manager. HA looked complete. It was not.

controlPlaneEndpoint: 10.0.0.120:6443     # cp-1 의 물리 IP

This value was the first node's real IP, not a VIP or DNS name. So when cp-1 died, etcd quorum stayed healthy at 2/3 and the apiservers on cp-2 and cp-3 worked normally, yet kubectl and the kubelets on all 7 nodes could no longer connect. On top of that, the certificate SAN lacked the other control plane IPs, so even connecting directly to cp-2 failed TLS verification. Data availability and access availability are different problems.

Note also that growing the control plane only to 2 machines is riskier than 1. The majority of 2 members is 2, so losing any one machine loses quorum. This is why odd numbers (1, 3, 5) are recommended, and why this expansion also had to go all the way to 3.

What to do in the next lab

In the first lab you investigate the cluster, work with namespaces, labels, and annotations, create and switch kubeconfig contexts yourself, and generate manifests with kubectl explain and --dry-run=client -o yaml. In the second lab you write a CustomResourceDefinition yourself to extend the API. The goal is to see with your own eyes the moment schema validation actually rejects a request.