TT Lab
Get started
Learn Learning paths Courses

Building clusters with Kubespray and Terraform

Design and check a five-node cluster on a single VM

Continue in TT Lab

Goal

You design an inventory of 3 control planes (also serving as etcd) and 2 workers, and, on one VM, check what can be checked without real nodes — group placement, kubespray's inventory check, the etcd odd rule and quorum, the API server load balancing method, and where the playbooks aim when adding and removing nodes.

Why it matters

What changes when you move from a single-machine cluster to several machines is not the installation command but the design. How many control planes to have, whether they also serve as etcd, which API server the workers attach to, and which playbook to narrow down and run where when adding and removing nodes are all decided by the inventory and a few variables. And these decisions are hard to change once the installation has started. If you check the inventory with tools before creating real nodes, you can prevent a good share of the cases of stopping midway through installation. The lab of doing node joining, failures, and rolling upgrades with several real VMs comes after this module when the feature of giving several VMs per session is ready.

Steps

  1. Copy /opt/ks/kubespray/inventory/sample to /root/ks/inventory/ha, and in /root/ks/inventory/ha/inventory.ini, put cp1, cp2, and cp3 in kube_control_plane (with etcd taking kube_control_plane as a child), put w1 and w2 in kube_node, and write it so that k8s_cluster takes the two groups as children. The control planes do not double as workers. For each node, put ansible_host and ip (cp1=192.0.2.11 through w2=192.0.2.15, in order) and ansible_user: ubuntu in /root/ks/inventory/ha/host_vars/<노드>.yml (the placeholder is the node name).
  2. In /opt/ks/kubespray, run ansible-playbook -i /root/ks/inventory/ha/inventory.ini playbooks/boilerplate.yml -e ansible_connection=local and leave the entire output in /root/ks/ha/validate.log. The PLAY RECAP must be failed=0 for all five nodes.
  3. Copy the inventory from step 1 to /root/ks/inventory/ha-even, but change it so that etcd has only the two machines cp1 and cp2 (two names under [etcd]), and run boilerplate the same way (-e ansible_connection=local). Write to /root/ks/ha/even.json failed_task (the name of the failed task, without the role prefix), etcd_members (the number of etcd hosts in that inventory), quorum (the majority of that number), and tolerated_failures (how many can be lost while writes still work).
  4. Write to /root/ks/ha/quorum.json, for each etcd member count of 1, 3, 5, and 7, the quorum (the majority) and tolerated_failures (the number of members you can lose while keeping writes), in the shape {"1": {"quorum": .., "tolerated_failures": ..}, "3": {...}, ...}.
  5. Read kubespray's roles/kubespray_defaults/defaults/main/main.yml and docs/operations/ha-mode.md and write to /root/ks/ha/lb.json localhost_lb (the result of loadbalancer_apiserver_localhost when no external load balancer is defined, a boolean), lb_type (the default of loadbalancer_apiserver_type), lb_port (the port that local proxy uses — the value it follows if loadbalancer_apiserver_port is absent, a number), and who_uses_it (the node group that uses that local proxy: whichever of "kube_node" or "kube_control_plane" the documentation says).
  6. Add w3 (192.0.2.16 in host_vars) to kube_node in the inventory from step 1. Then in /opt/ks/kubespray, run scale.yml --list-hosts --limit w3 and remove-node.yml --list-hosts -e node=w2 with the inventory /root/ks/inventory/ha/inventory.ini, and write to /root/ks/ha/plan.json scale_node_play_hosts (a sorted array of the hosts targeted by the play in scale.yml whose name ends with "...(node)") and remove_confirm_hosts (a sorted array of the hosts targeted by the "Confirm node removal" play in remove-node.yml). In /root/ks/ha/runbook.sh, write one line each for the command that adds w3, the command that removes w2, and the upgrade command that raises nodes one at a time (do not run them).
  7. Read /opt/ks/kubespray/ansible.cfg and write to /root/ks/ha/facts.json cache_plugin (the fact_caching value), cache_dir (the fact_caching_connection value), cache_timeout_sec (fact_caching_timeout, a number), and refresh_playbook (the path of the playbook that the documentation says to run without a limit before using --limit, relative to the kubespray repository).

Notes

3 control planes, 2 workers

Copy /opt/ks/kubespray/inventory/sample to /root/ks/inventory/ha, and in /root/ks/inventory/ha/inventory.ini, put cp1, cp2, and cp3 in kube_control_plane (with etcd taking kube_control_plane as a child), put w1 and w2 in kube_node, and write it so that k8s_cluster takes the two groups as children. The control planes do not double as workers. For each node, put ansible_host and ip (cp1=192.0.2.11 through w2=192.0.2.15, in order) and ansible_user: ubuntu in /root/ks/inventory/ha/host_vars/<노드>.yml (the placeholder is the node name).

It has the same shape as the single-machine inventory of module 1, with only more names in the groups and more host_vars files. This is where the reason for not putting the connection method on the inventory line shows. ip is the address kubespray attaches the API server and etcd to, and ansible_host is the address Ansible reaches over ssh — the two differ if the management network and the service network are different. 192.0.2.0/24 is a range reserved for documentation, so it is not actually reachable.

Only the inventory check, without ssh

In /opt/ks/kubespray, run ansible-playbook -i /root/ks/inventory/ha/inventory.ini playbooks/boilerplate.yml -e ansible_connection=local and leave the entire output in /root/ks/ha/validate.log. The PLAY RECAP must be failed=0 for all five nodes.

boilerplate does not gather facts and checks only the inventory and variables, so if you override the connection with local, you can run it without real nodes. -e takes precedence over host_vars, so the connection method changes only in this run. It is the cheapest check you can use for reviewing the inventory before starting an installation.

If etcd is two machines

Copy the inventory from step 1 to /root/ks/inventory/ha-even, but change it so that etcd has only the two machines cp1 and cp2 (two names under [etcd]), and run boilerplate the same way (-e ansible_connection=local). Write to /root/ks/ha/even.json failed_task (the name of the failed task, without the role prefix), etcd_members (the number of etcd hosts in that inventory), quorum (the majority of that number), and tolerated_failures (how many can be lost while writes still work).

etcd needs the agreement of a majority (quorum) for every write. The majority is the quotient of the member count divided by 2, plus 1. With two machines the majority is two, so losing just one stops it, and with one machine too, losing one stops it, so two machines are no better than one and only add members. kubespray blocks this with the inventory check.

How many can you lose

Write to /root/ks/ha/quorum.json, for each etcd member count of 1, 3, 5, and 7, the quorum (the majority) and tolerated_failures (the number of members you can lose while keeping writes), in the shape {"1": {"quorum": .., "tolerated_failures": ..}, "3": {...}, ...}.

The majority is n//2 + 1, and the failures it can tolerate are n - majority. If you add members the number tolerated grows, but every write has to wait for the agreement of more members. The reason the etcd documentation recommends 3 or 5 for production clusters and says not to exceed 7 is in this table.

Which API server do workers attach to

Read kubespray's roles/kubespray_defaults/defaults/main/main.yml and docs/operations/ha-mode.md and write to /root/ks/ha/lb.json localhost_lb (the result of loadbalancer_apiserver_localhost when no external load balancer is defined, a boolean), lb_type (the default of loadbalancer_apiserver_type), lb_port (the port that local proxy uses — the value it follows if loadbalancer_apiserver_port is absent, a number), and who_uses_it (the node group that uses that local proxy: whichever of "kube_node" or "kube_control_plane" the documentation says).

If there are several control planes, the workers' kubelet and kube-proxy have to decide which API server to attach to. The kubespray default is that if you do not define an external load balancer (loadbalancer_apiserver) separately, it starts an nginx proxy on each worker, receives on localhost, and spreads the requests across all the API servers. The documentation explains that this method is less efficient than a dedicated LB but practical where managing a VIP is a hassle.

Where the commands that add and remove nodes aim

Add w3 (192.0.2.16 in host_vars) to kube_node in the inventory from step 1. Then in /opt/ks/kubespray, run scale.yml --list-hosts --limit w3 and remove-node.yml --list-hosts -e node=w2 with the inventory /root/ks/inventory/ha/inventory.ini, and write to /root/ks/ha/plan.json scale_node_play_hosts (a sorted array of the hosts targeted by the play in scale.yml whose name ends with "...(node)") and remove_confirm_hosts (a sorted array of the hosts targeted by the "Confirm node removal" play in remove-node.yml). In /root/ks/ha/runbook.sh, write one line each for the command that adds w3, the command that removes w2, and the upgrade command that raises nodes one at a time (do not run them).

--list-hosts changes nothing and only shows whom each play targets. scale.yml installs only on the new node, so narrow it with --limit, and remove-node.yml takes the node to remove with -e node=. The kubespray documentation recommends running facts.yml once without a limit to refresh the facts cache before using --limit. You control upgrading one at a time with the serial variable.

Why you run facts.yml before --limit

Read /opt/ks/kubespray/ansible.cfg and write to /root/ks/ha/facts.json cache_plugin (the fact_caching value), cache_dir (the fact_caching_connection value), cache_timeout_sec (fact_caching_timeout, a number), and refresh_playbook (the path of the playbook that the documentation says to run without a limit before using --limit, relative to the kubespray repository).

If you aim only at the new node with --limit, that run does not gather the facts of other nodes and uses the cache. But the configuration files (for example /etc/hosts, the etcd member list, the load balancer backends) are built from the facts of all nodes. If the cache is empty or stale, the new node gets wrong addresses. See the Node-based upgrade section of the upgrade documentation.