Building clusters with Kubespray and Terraform
Run cluster.yml twice, get the same cluster
Goal
You run cluster.yml with the prepared inventory and set up Kubernetes 1.35.8 on one VM (4 vCPU · 4 GiB). You run the same playbook once more
to see what changes again, and then change one value and reapply only that part with tags.
Why it matters
The way kubespray is operated is "fix the declaration and run the same playbook again." Whether you add a node or change a setting, you run the same cluster.yml. So you need to know what the second run changes — idempotent does not mean changed is 0, and you have to know which tasks change every time to judge "did this run change only what was intended?" You also learn here how to read the log of a long playbook from the end (PLAY RECAP and TASKS RECAP). In this lab you wait about 15 minutes in total for two installations and one partial run. The session is 60 minutes, so extend it if you are short on time. When the VM ends, the cluster disappears too.
Steps
- In
/opt/ks/kubespray, look at the play list withansible-playbook -i /root/ks/inventory/lab/inventory.ini cluster.yml --list-tasksand write to/root/ks/install/plan.json, as numbers,play_count(the number of plays),etcd_play(the number of the "Install etcd" play),cni_play(the number of "Invoke kubeadm and install a CNI"), andapps_play(the number of "Install Kubernetes apps"). - In
/opt/ks/kubespray, runansible-playbook -i /root/ks/inventory/lab/inventory.ini cluster.ymland leave the entire output in/root/ks/logs/cluster-1.log(| teeor a redirect). It takes about 7 minutes. When it finishes, node1 in the PLAY RECAP must havefailed=0, andkubectl get node node1must be Ready. - Write to
/root/ks/install/run1.jsonthe first run'sok,changed, andfailed(the node1 values in the PLAY RECAP),elapsed_sec(the cumulative time printed on the TASKS RECAP header, in seconds, an integer),slowest_task(the name of the task at the top of the TASKS RECAP list, including the role prefix, as is), andnode_uid(the metadata.uid ofkubectl get node node1). - Run the same command once more and leave the entire output in
/root/ks/logs/cluster-2.log. The second run must also havefailed=0, and the node UID must equal the value you wrote in step 3 (meaning the cluster was not recreated). - Write the names of the tasks that reported
changed: [node1]in the second run, exactly as inside the brackets ofTASK [...]in the log (RUNNING HANDLER [...]for handlers), including the role prefix, one per line, to/root/ks/install/changed.txt. And write the second run'sok,changed,failed, andelapsed_secto/root/ks/install/run2.json. - Add
containerd_max_container_log_line_size: 32768to/root/ks/inventory/lab/group_vars/all/containerd.yml, rerun only the container runtime part withcluster.yml --tags containerd, and leave the entire output in/root/ks/logs/containerd.log. When it finishes,max_container_log_line_sizein/etc/containerd/config.tomlmust be 32768 and containerd must be running again with the new configuration, and the node must still be Ready. - Write to
/root/ks/install/report.jsonkubespray(the tag),server_version(the API server gitVersion),first_run_secandsecond_run_sec(the cumulative times of the two runs, integers),first_changed,second_changed,tags_run_changed(the changed of the run in step 6), andcontainerd_restarted(whether containerd came back up because of step 6, a boolean).
Notes
- kubespray v2.32.0 is ready at
/opt/ks/kubesprayand the inventory at/root/ks/inventory/lab/inventory.ini(one node, kube_version 1.35.8). Run the playbooks in/opt/ks/kubespray. - Files downloaded during installation pile up in
/tmp/releases, and images in the containerd store. This is why the second run is faster. - Common mistake: judging failure from a
fatal:in the log. A fatal with...ignoringattached is a check task that kubespray deliberately passes over. The verdict is the failed in the PLAY RECAP. - Common mistake: rerunning only cluster.yml after the first installation stopped midway. If it stopped before the CNI, the second run waits for the control plane to be Ready and fails. In that case, erase it with reset.yml and build from scratch (module 7).
- Documentation: Kubespray — Getting started · Kubespray — Ansible tags
Look at the order before running
In /opt/ks/kubespray, look at the play list with ansible-playbook -i /root/ks/inventory/lab/inventory.ini cluster.yml --list-tasks and write to /root/ks/install/plan.json, as numbers, play_count (the number of plays), etcd_play (the number of the "Install etcd" play), cni_play (the number of "Invoke kubeadm and install a CNI"), and apps_play (the number of "Install Kubernetes apps").
--list-tasks changes nothing and only lays out the playbook for you. Look at the lines of the form play #N (호스트 패턴): 이름 in the output (the placeholders are the host pattern and the play name). If you think about why etcd comes before the control plane and the CNI comes before the add-ons, it becomes easy to read the failure point later.
Run cluster.yml
In /opt/ks/kubespray, run ansible-playbook -i /root/ks/inventory/lab/inventory.ini cluster.yml and leave the entire output in /root/ks/logs/cluster-1.log (| tee or a redirect). It takes about 7 minutes. When it finishes, node1 in the PLAY RECAP must have failed=0, and kubectl get node node1 must be Ready.
If the console is cut off, the playbook may die with it. If you start it with systemd-run --unit=<이름> --setenv=HOME=/root ... (the placeholder is the unit name) or tmux, it runs regardless of the terminal, and you can watch it with tail -f. If HOME is empty, kubespray's kube module cannot find the kubeconfig and goes to localhost:8080. A fatal with ...ignoring in the middle of the log is normal.
What you read from the end of the log
Write to /root/ks/install/run1.json the first run's ok, changed, and failed (the node1 values in the PLAY RECAP), elapsed_sec (the cumulative time printed on the TASKS RECAP header, in seconds, an integer), slowest_task (the name of the task at the top of the TASKS RECAP list, including the role prefix, as is), and node_uid (the metadata.uid of kubectl get node node1).
The profile_tasks callback that ansible.cfg turned on attaches a TASKS RECAP at the end of the log. The last time on the first line is the total cumulative time, and the list below is in order of how long each took. The node UID becomes the basis in the next step for checking 'is it the same node even after running again?'
Run it once more
Run the same command once more and leave the entire output in /root/ks/logs/cluster-2.log. The second run must also have failed=0, and the node UID must equal the value you wrote in step 3 (meaning the cluster was not recreated).
kubespray is built so that you may run cluster.yml again on an already built cluster — running the same playbook again when adding nodes or changing settings is the basic way of operating. The second run is faster than the first because what it needs to download is in the cache.
It is idempotent but changed is not 0
Write the names of the tasks that reported changed: [node1] in the second run, exactly as inside the brackets of TASK [...] in the log (RUNNING HANDLER [...] for handlers), including the role prefix, one per line, to /root/ks/install/changed.txt. And write the second run's ok, changed, failed, and elapsed_sec to /root/ks/install/run2.json.
Idempotent means 'the resulting state is the same no matter how many times you run it', not 'no task reports changed.' Command and shell tasks that execute something every time, and tasks that rewrite without comparing values, report changed. In the log, the nearest TASK (or RUNNING HANDLER) header just above a changed: [node1] line is that task.
Change one value and rerun only that part
Add containerd_max_container_log_line_size: 32768 to /root/ks/inventory/lab/group_vars/all/containerd.yml, rerun only the container runtime part with cluster.yml --tags containerd, and leave the entire output in /root/ks/logs/containerd.log. When it finishes, max_container_log_line_size in /etc/containerd/config.toml must be 32768 and containerd must be running again with the new configuration, and the node must still be Ready.
If you give a tag, only the tasks with that tag (and the preparation tasks with the always tag) run. You can see which tags exist in the tag table of the kubespray documentation or with --list-tags. If only the configuration file changes and the daemon does not come up again, it keeps running with the old value — the role's handler takes care of the restart. As the documentation warns, use tags only when you know exactly what will run.
Installation report
Write to /root/ks/install/report.json kubespray (the tag), server_version (the API server gitVersion), first_run_sec and second_run_sec (the cumulative times of the two runs, integers), first_changed, second_changed, tags_run_changed (the changed of the run in step 6), and containerd_restarted (whether containerd came back up because of step 6, a boolean).
These are all values you can recalculate from the records of earlier steps, the three logs, and the current cluster. The grader recalculates from the same places and compares. Judge whether containerd came back up by whether the restart handler was printed as changed in containerd.log.