Building clusters with Kubespray and Terraform
From one node to many — quorum, load balancers, and the order of adding and removing nodes
One-line summary
What changes in kubespray when you widen to several nodes is a few lines of names in the inventory and a few host_vars files, but behind that are design decisions: the etcd quorum, API server load balancing, and the order of adding and removing nodes.
Why this was needed
The labs in this course used a single VM as both the control node and the target node. A real cluster is different. If one control plane machine dies, the API stops, and if one etcd machine dies, the data stops. The high-availability topology section of the Kubernetes documentation describes two shapes — stacked, which puts etcd on the control plane nodes, and external, which puts etcd separately. Stacked needs fewer nodes but if you lose one node you lose a control plane and an etcd member together, while external splits that risk at the price of twice the nodes. In kubespray, this choice is the difference between [etcd:children] kube_control_plane and writing other names under [etcd].
How it works
etcd is an odd number. etcd needs the agreement of a majority (quorum, n//2+1) for every write. So the number of failures it can tolerate is n − majority.
멤버 과반 견디는 장애
1 1 0
2 2 0 ← 한 대보다 나을 것이 없다
3 2 1
5 3 2
7 4 3 ← 쓰기마다 넷을 기다린다
Two machines have the same number of tolerated failures as one while only the majority grows. kubespray blocks this with "Stop if even number of etcd hosts" in validate_inventory, and the documentation also says to have 3 or more for failure resilience. The etcd FAQ recommends 3 or 5 for production clusters and explains that write latency increases as you add members.
Workers attach to the API server through localhost. If there are several control planes, the kubelet and kube-proxy have to decide which API server to attach to. The kubespray default is that if you do not define an external load balancer (loadbalancer_apiserver), loadbalancer_apiserver_localhost becomes true, and on every node that is not a control plane it starts an nginx proxy (loadbalancer_apiserver_type: nginx), receives on localhost at kube_apiserver_port (6443), and spreads the requests across all the API servers. The HA documentation says this method is less efficient than a dedicated LB because it sends more health checks to the API servers, but it is practical where managing a VIP is a hassle. For external clients (the operator's kubectl and so on), you set up an LB or kube-vip separately.
Adding and removing nodes are dedicated playbooks. The order in the getting started documentation is this. To add a worker, add the name to an inventory group and run scale.yml — running cluster.yml again also works, but scale.yml does only what is needed to bring the kubelet up on the new worker. To remove one, remove-node.yml -e node=<이름> (the placeholder is the node name) does drain → stop services → clean up certificates → delete the node. The upgrade documentation says that before running with --limit to pick nodes, you should run playbooks/facts.yml once without a limit to refresh the facts cache. The repository's ansible.cfg caches facts in /tmp as jsonfile for a day (86400 seconds), because a --limit run does not gather the facts of other nodes and uses the cache.
You can check the inventory without nodes. boilerplate.yml does not gather facts and looks only at the inventory and variables. If you override the connection with -e ansible_connection=local, it runs even without real nodes. In measurement, it checked a five-node inventory in 2.4 seconds. --list-hosts changes nothing and only shows whom each play targets, so before running you can see whether scale.yml targets only the new node and whether remove-node.yml confirms only the node to be removed.
What it looks like in the field
There are three failures you often see. Increasing etcd to an even number (with 4, the majority is 3, so the number tolerated stays the same as with 3), increasing the control planes without deciding how the workers reach the API server, and running only the new node with --limit while the facts cache is empty, so the new node's /etc/hosts or etcd configuration gets the wrong addresses.
This module goes only as far as design and checking. When the feature of giving several VMs per session is ready, labs will follow that use this inventory as is to actually do node joining (scale.yml), three control planes and the etcd quorum, node failure and recovery (remove-node.yml, recover-control-plane.yml), and a rolling upgrade one at a time (serial=1). Moving the connection method out to host_vars in the single-machine labs was preparation for that time.
What you will do in the next lab
You copy the sample and make an inventory of 3 control planes and 2 workers with per-node host_vars, and pass kubespray's inventory check without ssh. You see where an inventory with etcd changed to two machines gets blocked, and calculate the quorum table. You read from the defaults how workers attach to the API server, and, after checking with --list-hosts whom the commands that add and remove one worker target, write them in a runbook.