TT Lab
Get started
Learn Learning paths Courses

Building clusters with Kubespray and Terraform

Run cluster.yml twice, get the same cluster

Continue in TT Lab

Goal

You run cluster.yml with the prepared inventory and set up Kubernetes 1.35.8 on one VM (4 vCPU · 4 GiB). You run the same playbook once more to see what changes again, and then change one value and reapply only that part with tags.

Why it matters

The way kubespray is operated is "fix the declaration and run the same playbook again." Whether you add a node or change a setting, you run the same cluster.yml. So you need to know what the second run changes — idempotent does not mean changed is 0, and you have to know which tasks change every time to judge "did this run change only what was intended?" You also learn here how to read the log of a long playbook from the end (PLAY RECAP and TASKS RECAP). In this lab you wait about 15 minutes in total for two installations and one partial run. The session is 60 minutes, so extend it if you are short on time. When the VM ends, the cluster disappears too.

Steps

  1. In /opt/ks/kubespray, look at the play list with ansible-playbook -i /root/ks/inventory/lab/inventory.ini cluster.yml --list-tasks and write to /root/ks/install/plan.json, as numbers, play_count (the number of plays), etcd_play (the number of the "Install etcd" play), cni_play (the number of "Invoke kubeadm and install a CNI"), and apps_play (the number of "Install Kubernetes apps").
  2. In /opt/ks/kubespray, run ansible-playbook -i /root/ks/inventory/lab/inventory.ini cluster.yml and leave the entire output in /root/ks/logs/cluster-1.log (| tee or a redirect). It takes about 7 minutes. When it finishes, node1 in the PLAY RECAP must have failed=0, and kubectl get node node1 must be Ready.
  3. Write to /root/ks/install/run1.json the first run's ok, changed, and failed (the node1 values in the PLAY RECAP), elapsed_sec (the cumulative time printed on the TASKS RECAP header, in seconds, an integer), slowest_task (the name of the task at the top of the TASKS RECAP list, including the role prefix, as is), and node_uid (the metadata.uid of kubectl get node node1).
  4. Run the same command once more and leave the entire output in /root/ks/logs/cluster-2.log. The second run must also have failed=0, and the node UID must equal the value you wrote in step 3 (meaning the cluster was not recreated).
  5. Write the names of the tasks that reported changed: [node1] in the second run, exactly as inside the brackets of TASK [...] in the log (RUNNING HANDLER [...] for handlers), including the role prefix, one per line, to /root/ks/install/changed.txt. And write the second run's ok, changed, failed, and elapsed_sec to /root/ks/install/run2.json.
  6. Add containerd_max_container_log_line_size: 32768 to /root/ks/inventory/lab/group_vars/all/containerd.yml, rerun only the container runtime part with cluster.yml --tags containerd, and leave the entire output in /root/ks/logs/containerd.log. When it finishes, max_container_log_line_size in /etc/containerd/config.toml must be 32768 and containerd must be running again with the new configuration, and the node must still be Ready.
  7. Write to /root/ks/install/report.json kubespray (the tag), server_version (the API server gitVersion), first_run_sec and second_run_sec (the cumulative times of the two runs, integers), first_changed, second_changed, tags_run_changed (the changed of the run in step 6), and containerd_restarted (whether containerd came back up because of step 6, a boolean).

Notes

Look at the order before running

In /opt/ks/kubespray, look at the play list with ansible-playbook -i /root/ks/inventory/lab/inventory.ini cluster.yml --list-tasks and write to /root/ks/install/plan.json, as numbers, play_count (the number of plays), etcd_play (the number of the "Install etcd" play), cni_play (the number of "Invoke kubeadm and install a CNI"), and apps_play (the number of "Install Kubernetes apps").

--list-tasks changes nothing and only lays out the playbook for you. Look at the lines of the form play #N (호스트 패턴): 이름 in the output (the placeholders are the host pattern and the play name). If you think about why etcd comes before the control plane and the CNI comes before the add-ons, it becomes easy to read the failure point later.

Run cluster.yml

In /opt/ks/kubespray, run ansible-playbook -i /root/ks/inventory/lab/inventory.ini cluster.yml and leave the entire output in /root/ks/logs/cluster-1.log (| tee or a redirect). It takes about 7 minutes. When it finishes, node1 in the PLAY RECAP must have failed=0, and kubectl get node node1 must be Ready.

If the console is cut off, the playbook may die with it. If you start it with systemd-run --unit=<이름> --setenv=HOME=/root ... (the placeholder is the unit name) or tmux, it runs regardless of the terminal, and you can watch it with tail -f. If HOME is empty, kubespray's kube module cannot find the kubeconfig and goes to localhost:8080. A fatal with ...ignoring in the middle of the log is normal.

What you read from the end of the log

Write to /root/ks/install/run1.json the first run's ok, changed, and failed (the node1 values in the PLAY RECAP), elapsed_sec (the cumulative time printed on the TASKS RECAP header, in seconds, an integer), slowest_task (the name of the task at the top of the TASKS RECAP list, including the role prefix, as is), and node_uid (the metadata.uid of kubectl get node node1).

The profile_tasks callback that ansible.cfg turned on attaches a TASKS RECAP at the end of the log. The last time on the first line is the total cumulative time, and the list below is in order of how long each took. The node UID becomes the basis in the next step for checking 'is it the same node even after running again?'

Run it once more

Run the same command once more and leave the entire output in /root/ks/logs/cluster-2.log. The second run must also have failed=0, and the node UID must equal the value you wrote in step 3 (meaning the cluster was not recreated).

kubespray is built so that you may run cluster.yml again on an already built cluster — running the same playbook again when adding nodes or changing settings is the basic way of operating. The second run is faster than the first because what it needs to download is in the cache.

It is idempotent but changed is not 0

Write the names of the tasks that reported changed: [node1] in the second run, exactly as inside the brackets of TASK [...] in the log (RUNNING HANDLER [...] for handlers), including the role prefix, one per line, to /root/ks/install/changed.txt. And write the second run's ok, changed, failed, and elapsed_sec to /root/ks/install/run2.json.

Idempotent means 'the resulting state is the same no matter how many times you run it', not 'no task reports changed.' Command and shell tasks that execute something every time, and tasks that rewrite without comparing values, report changed. In the log, the nearest TASK (or RUNNING HANDLER) header just above a changed: [node1] line is that task.

Change one value and rerun only that part

Add containerd_max_container_log_line_size: 32768 to /root/ks/inventory/lab/group_vars/all/containerd.yml, rerun only the container runtime part with cluster.yml --tags containerd, and leave the entire output in /root/ks/logs/containerd.log. When it finishes, max_container_log_line_size in /etc/containerd/config.toml must be 32768 and containerd must be running again with the new configuration, and the node must still be Ready.

If you give a tag, only the tasks with that tag (and the preparation tasks with the always tag) run. You can see which tags exist in the tag table of the kubespray documentation or with --list-tags. If only the configuration file changes and the daemon does not come up again, it keeps running with the old value — the role's handler takes care of the restart. As the documentation warns, use tags only when you know exactly what will run.

Installation report

Write to /root/ks/install/report.json kubespray (the tag), server_version (the API server gitVersion), first_run_sec and second_run_sec (the cumulative times of the two runs, integers), first_changed, second_changed, tags_run_changed (the changed of the run in step 6), and containerd_restarted (whether containerd came back up because of step 6, a boolean).

These are all values you can recalculate from the records of earlier steps, the three logs, and the current cluster. The grader recalculates from the same places and compares. Judge whether containerd came back up by whether the restart handler was printed as changed in containerd.log.