Building clusters with Kubespray and Terraform
Operating means rerunning the same playbook
One-line summary
cluster.yml is 15 plays that run in the order host preparation → runtime → download → etcd → kubelet → control plane → CNI → add-ons, and it is built so that no matter how many times you run it again with the same inventory, you get the same cluster.
Why this was needed
Setting up one machine with kubeadm takes just a few command lines. What is hard is what comes after. Three months later, when someone has to change one containerd setting, that person needs to know what was done at the beginning and in what order, and has to produce the same result on every node. A shell script runs fine the first time, but the second time it runs into files that already exist and daemons that are already up. kubespray wrote installation as Ansible tasks that "look at the current state and fill in only what is missing", so the install command and the operations command are the same. Whether you change a setting or add a node, you run the same cluster.yml again.
How it works
When you lay it out with --list-tasks, the cluster.yml of v2.32.0 is 15 plays (measured: 4 seconds, and it changes nothing). There is a reason for the order.
#1-2 Ansible 판 확인, 인벤토리 검사(boilerplate) — 틀린 인벤토리는 여기서 1초 만에 멈춘다
#4-5 호스트 부트스트랩, 사실(facts) 수집
#6 etcd 준비: preinstall(스왑·sysctl·패키지) → 컨테이너 런타임 → 받기(download)
#8 etcd 설치 — 기본은 컨테이너가 아니라 호스트의 systemd 서비스(etcd_deployment_type: host)
#9 kubelet 설치
#10 컨트롤 플레인 — 안에서 kubeadm init 을 부른다
#11 kubeadm 마무리, 노드 라벨·테인트, CNI(기본 calico)
#14 애드온 — CoreDNS, nodelocaldns, (켰다면) metrics-server 등
#15 클러스터 DNS 가 뜬 뒤 resolv.conf 정리
etcd comes before the control plane because the API server does not come up without a datastore, and the CNI comes before the add-ons because CoreDNS does not come up without a Pod network. If you know this order, you can guess from the failure point alone what already exists and what does not.
Read the log from the end. One line of the PLAY RECAP is the verdict — ok, changed, unreachable, failed, skipped. The repository's ansible.cfg turns on the profile_tasks callback, so a TASKS RECAP is attached below it; the cumulative time on the header line is the total time, and the list below is the tasks in order of how long they took. A fatal: in the middle is not the verdict. For example, at the first installation Get currently-deployed etcd version is expected to fail because etcd does not exist yet, and kubespray passes that result over with ...ignoring.
Idempotent does not mean changed=0. Most Ansible modules work by "comparing the desired state with the current state and changing only if they differ", so in the second run most end as ok. But tasks invoked through command or shell have no way to compare, so they either report changed every time they run or are deliberately suppressed with changed_when. So the changed list of the second run is the list of "things that change every time", and you need to know it to tell whether a new changed in the third run is my change or not.
You can run only a part with tags. The tag table in the kubespray documentation has names such as containerd, etcd, network, apps, coredns, and metrics_server, and if you give --tags containerd, only the tasks with that tag and the always-tagged ones (inventory check, facts) run. The documentation warns you to use tags "only when you are 100% sure what they do." This is because dependencies between roles (for example, changing the containerd configuration calls the restart handler) are missed if they lie outside the tag.
What it looks like in the field
These are values measured on the medium VM (4 vCPU · 4 GiB) while building this course. On a clean VM, the first run of cluster.yml took 341 seconds (changed 132), and the second run of the same command took 183 seconds (changed 27). The memory use of the whole VM during installation (MemTotal − MemAvailable) peaked at 1.52GiB, and the Ansible processes accounted for up to 345MB of that. That means 4 GiB is enough. What was downloaded came from github.com releases, dl.k8s.io, registry.k8s.io, quay.io, and the Ubuntu apt repository, all fetched only over HTTPS (443) or HTTP (80).
I stumbled twice during measurement, and both are shapes you meet in the field just as they are. First, to save the root, I bind-mounted /usr/local wholesale to another disk, and /usr/local/share/ca-certificates was hidden, so the etcd certificate step stopped with Destination directory ... does not exist. Second, when I ran the playbook in an environment where HOME is empty (a systemd unit), kubespray's kube module called kubectl without --kubeconfig and went to localhost:8080, and the cluster_roles step retried for a minute and then stopped. This second failure was before the CNI installation, and when I ran cluster.yml again in that state, this time "Wait for new control plane nodes to be Ready" waited 130 seconds and failed. This is because on a node where kubeadm has already run, the task that waits for Ready without a CNI comes first. A half-built cluster may not be fixed by running again, and in that case it is faster to erase it with reset.yml and build from scratch (reset was measured at 78 seconds).
What you will do in the next lab
You first lay out the play order, and run cluster.yml to build the cluster. You read the result and time from the end of the log, run the same command once more, and collect the tasks the second run changed. Finally, you change one containerd setting value through group_vars and reapply only that part with --tags containerd, then check that both the configuration file and the daemon changed.