Building clusters with Kubespray and Terraform
The inventory assigns roles; group_vars decide what gets installed
One-line summary
kubespray is a set of Ansible playbooks that install Kubernetes, and the only two things a person writes are the inventory (who is in charge of what) and group_vars (what to install).
Why this was needed
kubeadm is a tool that assembles a control plane on a single node. If you have ten servers, someone has to install the runtime on all ten, match the kernel settings, run kubeadm init on the first one, and run join on the rest in order. If you start writing this job as a shell script, it quickly becomes "a script that breaks if you run it twice." kubespray moves that repetition into Ansible roles. The official comparison documentation explains that kubespray handles "general configuration management on the OS side", and that for knowledge about the cluster lifecycle, since v2.3 it calls kubeadm internally and borrows it. In other words, even a control plane set up with kubespray is, inside, static Pods created by kubeadm, and the difference is who does the host preparation, etcd, CNI, and add-ons before and after that.
How it works
The inventory is a role assignment table. The groups defined by the kubespray documentation are three.
kube_control_plane API 서버·스케줄러·컨트롤러 매니저가 뜨는 노드
kube_node 파드가 올라가는 노드(워커)
etcd etcd 멤버. 장애 대비에는 3대 이상, 반드시 홀수
k8s_cluster kube_node + kube_control_plane (+ calico_rr) — kubespray 가 실행 중에 만든다
There is a trap in k8s_cluster. The sample inventory.ini of v2.32.0 does not write this group, and the dynamic_groups role in boilerplate.yml creates it inside the playbook with group_by. So tools outside the playbook do not know this group. Measured on this VM, even if you write kube_version in group_vars/k8s_cluster/k8s-cluster.yml, ansible-inventory --host node1 shows that value as null, and if you ask with ansible ... -m debug, the value in group_vars/all wins. The installation itself works without problems, but a tool you tried to use to check values before installing lies. So this course writes kube_control_plane and kube_node under [k8s_cluster:children] — it is how the old sample did it, and even if kubespray creates the same group again during execution, the result is the same.
If one node is in both kube_control_plane and kube_node, the control plane also does work. If you do not want to place etcd separately, put kube_control_plane wholesale under [etcd:children] (stacked etcd). In this course there is only one VM, so the same single node goes into all three groups.
Group names are a contract down to the spelling. boilerplate.yml at the very start of the playbook uses the dynamic_groups role to move only the old names (kube-master, kube-node) to the new names, and it does not know any other names. Next, the validate_inventory role checks the inventory, and when measured, if you wrote [masters], the check "stop if kube_control_plane is empty" actually passes. This is because groups.get() returns None for a group that does not exist, and None is different from an empty list. The failure comes at the very next check, "stop if the number of etcd members is even", with object of type 'dict' has no attribute 'kube_control_plane'. The error does not tell you the group name, so looking at the tree with ansible-inventory --graph before running the installation is the cheapest check.
group_vars is the installation content. The sample is split in two branches.
group_vars/all/*.yml etcd 를 포함한 모든 노드 — 프록시, 오프라인 저장소, etcd 설정
group_vars/k8s_cluster/*.yml 클러스터 노드 — kube_version, container_manager, kube_network_plugin, 대역, 애드온
group_vars/kube_control_plane.yml 컨트롤 플레인만
The variable layer table in the kubespray documentation is short. The group_vars of the inventory are the most used, host_vars are per-node exceptions, and extra vars (-e) always win. And it says to use -e for overriding internal variables that kubespray does not promise to users. Under Ansible's rules, if the same name is in both group_vars/all and group_vars/k8s_cluster, the more specific child group (k8s_cluster) wins. So if someone writes a version number in all.yml, that line does nothing and misleads the next person.
From ansible-core 2.19 (Ansible 12), conditional expressions must be booleans. -e key=value always passes a string, so the v2.32.0 release notes say to pass booleans as JSON, like -e '{"drain_nodes": true}'.
The version is decided by checksums. The default of kube_version is the first key of the kubelet checksum list in roles/kubespray_defaults/vars/main/checksums.yml, and the lowest version accepted is the last key. In v2.32.0 the default is 1.36.4 and the minimum is 1.34.0. A version with no checksum cannot be installed, at the very least. And if you do not write kube_version in the inventory, Kubernetes goes up along with it the moment you move kubespray to a new tag. This course writes 1.35.8 and raises it one step to 1.36.4 in the upgrade module.
What it looks like in the field
The first is the directory trap. The ansible.cfg in the kubespray repository puts .ini in inventory_ignore_extensions. So if you pass a directory like -i inventory/mycluster, it skips inventory.ini, leaves only a "No inventory was parsed" warning, and runs with 0 hosts. This is a result measured on this VM, and it is also the reason why the official documentation examples always point to a file with -i inventory/mycluster/inventory.ini.
The second is where you write the connection method. When there is only one node, it runs the same even if you attach it to the inventory line, as in node1 ansible_connection=local. But on the day you add nodes, if you copy that line, the new node is also connected as local, and the incident occurs of installing twice on the control node itself. If you move the connection information out to host_vars/<노드>.yml (the placeholder is the node name), the inventory remains only a role assignment table, and when adding a node you just add one line with the name to a group and one host_vars file (ansible_host, ansible_user). Every lab in this course uses this shape.
The third is the Ansible version on the control node. kubespray pins Ansible with requirements.txt in every tag, and v2.32.0 requires ansible==12.3.0 (ansible-core 2.19). The documentation recommends installing it as is in a virtual environment, and if the version does not match, ansible_version.yml stops at the first task. It is common to get stuck running with the system package's Ansible.
What you will do in the next lab
You copy the sample, put one node into three groups, and move the connection method out to host_vars. You count how the number of hosts differs when you pass a directory versus a file, pin the version in group_vars, and then check which one wins among all, k8s_cluster, and -e. Finally, you see where and with what words an inventory with wrong group names stops, and confirm that your own inventory passes the boilerplate check.