TT Lab
Get started
Learn Learning paths Courses

Building clusters with Kubespray and Terraform

Terraform builds the inventory, calls Kubespray, then deploys on top

Continue in TT Lab

Goal

Within the same VM, OpenTofu makes the kubespray inventory from a template, wraps the kubespray run to set up a cluster, and declares a namespace, RBAC, and an application on the cluster it set up using the kubernetes and helm providers. You catch what was changed from outside with plan and revert it.

Why it matters

The cluster lifecycle in the field usually has three layers — the layer that creates nodes (cloud or virtualization), the layer that installs Kubernetes on the nodes (kubespray), and the layer that installs the team's space, permissions, and apps on the cluster. Terraform is strong at the first and third layers, and it is common to leave the second to Ansible. This lab joins that boundary together inside one VM: what to keep as Terraform state, what to leave to kubespray's idempotency, and, when the declaration and reality diverge, who notices it. The first layer, creating nodes, is not done on this VM — because having the lab call this platform's virtualization API would break isolation. You wait about 20 minutes including the installation.

Steps

  1. In /root/ks/tf/cluster/main.tf, declare three local_files with the hashicorp/local provider — inventory (/root/ks/inventory/lab/inventory.ini, from the /root/ks/tf/cluster/inventory.tftpl template and the nodes variable), host_vars (for each node, ansible_connection in /root/ks/inventory/lab/host_vars/<노드>.yml, where the placeholder is the node name), and version (kube_version: <변수> in /root/ks/inventory/lab/group_vars/k8s_cluster/zz-terraform.yml, where the placeholder is the variable). The default of the kube_version variable is 1.35.8, and the default of nodes is one node, node1 (both control plane and worker, local connection). tofu validate must pass.
  2. In /root/ks/tf/cluster, create the three files with tofu apply. The resulting /root/ks/inventory/lab/inventory.ini must equal the content of local_file.inventory in state, and when read with ansible-inventory, node1 must be in kube_control_plane, etcd, and kube_node, and kube_version must resolve to 1.35.8.
  3. Add terraform_data "kubespray" to /root/ks/tf/cluster/main.tf so that it is recreated only when the content of the inventory, version, and host_vars files changes (triggers_replace), and when it is created, have local-exec run ansible-playbook -i /root/ks/inventory/lab/inventory.ini cluster.yml in /opt/ks/kubespray and leave the output in /root/ks/logs/tf-cluster.log (pass HOME=/root). When tofu apply finishes up to the installation (about 7 minutes), terraform_data.kubespray must be in state, the PLAY RECAP in the log must be failed=0, and node1 must be Ready.
  4. In /root/ks/tf/cluster, run tofu plan -detailed-exitcode once as is, and once with -var kube_version=1.36.4 added (do not apply). Write to /root/ks/tf/plan.json steady_exit (the exit code of the first plan), bump_exit (the exit code of the second), and bump_replaces (a sorted array of the resource addresses that the second plan says it will change or recreate).
  5. In /root/ks/tf/apps/main.tf, configure the hashicorp/kubernetes (3.2.1) and hashicorp/helm (3.3.0) providers with /root/.kube/config, and declare a namespace team-a (label owner=platform), a service account deployer, a Role deployer that can create and edit deployments but only read pods and services, its RoleBinding, and a helm_release "hello" that installs the local chart /opt/ks/charts/hello into team-a with 2 replicas, and run tofu apply. deployer must be able to create deployments in team-a and must not be able to delete nodes, and both hello Pods must be Ready.
  6. After changing the namespace label from outside with kubectl label ns team-a owner=someone-else --overwrite, catch the difference with tofu plan -detailed-exitcode in /root/ks/tf/apps and leave the output in /root/ks/tf/drift-plan.txt. Write to /root/ks/tf/drift.json exit_code (the exit code of that plan) and drifted (a sorted array of the resource addresses that it says it will change). Do not apply yet.
  7. In /root/ks/tf/apps, revert the difference with tofu apply. When you are done, the owner label of team-a must be platform, and tofu plan -detailed-exitcode must be 0. And the plan of /root/ks/tf/cluster must also be 0 (the declaration on the cluster side is unchanged too).

Notes

Make the inventory from a template

In /root/ks/tf/cluster/main.tf, declare three local_files with the hashicorp/local provider — inventory (/root/ks/inventory/lab/inventory.ini, from the /root/ks/tf/cluster/inventory.tftpl template and the nodes variable), host_vars (for each node, ansible_connection in /root/ks/inventory/lab/host_vars/<노드>.yml, where the placeholder is the node name), and version (kube_version: <변수> in /root/ks/inventory/lab/group_vars/k8s_cluster/zz-terraform.yml, where the placeholder is the variable). The default of the kube_version variable is 1.35.8, and the default of nodes is one node, node1 (both control plane and worker, local connection). tofu validate must pass.

templatefile() repeats with the %{ for } … %{ endfor } directives, and ~ eats a line break. If you put the node roles in a variable (a map), then when you add nodes, adding one line to the map changes the inventory and host_vars together. The version line that the recipe used to write in k8s-cluster.yml has been removed — so that Terraform alone is the owner of the version. OpenTofu can also be called by the name terraform.

Apply only the files first

In /root/ks/tf/cluster, create the three files with tofu apply. The resulting /root/ks/inventory/lab/inventory.ini must equal the content of local_file.inventory in state, and when read with ansible-inventory, node1 must be in kube_control_plane, etcd, and kube_node, and kube_version must resolve to 1.35.8.

State holds the content of the files Terraform created as is. If someone edits inventory.ini by hand, the next plan says it will revert it — this means the owner of this file is now Terraform. You can see the content with tofu state show local_file.inventory.

Call kubespray with terraform_data

Add terraform_data "kubespray" to /root/ks/tf/cluster/main.tf so that it is recreated only when the content of the inventory, version, and host_vars files changes (triggers_replace), and when it is created, have local-exec run ansible-playbook -i /root/ks/inventory/lab/inventory.ini cluster.yml in /opt/ks/kubespray and leave the output in /root/ks/logs/tf-cluster.log (pass HOME=/root). When tofu apply finishes up to the installation (about 7 minutes), terraform_data.kubespray must be in state, the PLAY RECAP in the log must be failed=0, and node1 must be Ready.

local-exec runs the command on the machine that runs Terraform (this VM). If the command fails, the resource is left tainted and the next apply calls it again. kubespray's kube module cannot find the kubeconfig if HOME is empty, so pass it through environment. Since apply holds the terminal for 7 minutes, it is safer to start it with systemd-run or tmux.

If nothing changed it does nothing — what if you change the version?

In /root/ks/tf/cluster, run tofu plan -detailed-exitcode once as is, and once with -var kube_version=1.36.4 added (do not apply). Write to /root/ks/tf/plan.json steady_exit (the exit code of the first plan), bump_exit (the exit code of the second), and bump_replaces (a sorted array of the resource addresses that the second plan says it will change or recreate).

-detailed-exitcode returns 0 if nothing will change and 2 if something will. See what gets recreated when you change the version — if terraform_data is recreated, cluster.yml is called again. A kubespray upgrade is upgrade-cluster.yml, not cluster.yml. Terraform does not know that difference. The resource_changes of tofu plan -json tell you the address and the action (update, replace).

Deploy to the cluster you set up by declaration

In /root/ks/tf/apps/main.tf, configure the hashicorp/kubernetes (3.2.1) and hashicorp/helm (3.3.0) providers with /root/.kube/config, and declare a namespace team-a (label owner=platform), a service account deployer, a Role deployer that can create and edit deployments but only read pods and services, its RoleBinding, and a helm_release "hello" that installs the local chart /opt/ks/charts/hello into team-a with 2 replicas, and run tofu apply. deployer must be able to create deployments in team-a and must not be able to delete nodes, and both hello Pods must be Ready.

There is a reason to separate the root module that sets up the cluster from the root module that deploys on it — if you create the cluster in one module and configure a provider with that cluster in the same apply, the first plan has to read a kubeconfig that does not exist yet. In helm provider 3.x, the kubernetes setting is not a block but the kubernetes = {{ ... }} attribute, and set is also a list attribute. Check permissions with kubectl auth can-i ... --as=system:serviceaccount:team-a:deployer.

Someone changed it from outside

After changing the namespace label from outside with kubectl label ns team-a owner=someone-else --overwrite, catch the difference with tofu plan -detailed-exitcode in /root/ks/tf/apps and leave the output in /root/ks/tf/drift-plan.txt. Write to /root/ks/tf/drift.json exit_code (the exit code of that plan) and drifted (a sorted array of the resource addresses that it says it will change). Do not apply yet.

plan first reads the actual state (refresh) and compares it with state, and then compares with the declaration. If what was changed from outside differs from the declaration, a plan to revert it to the declaration comes out. This is drift detection, and a common way of operating is to run plan periodically and hook exit code 2 to an alert.

Revert it to the declaration

In /root/ks/tf/apps, revert the difference with tofu apply. When you are done, the owner label of team-a must be platform, and tofu plan -detailed-exitcode must be 0. And the plan of /root/ks/tf/cluster must also be 0 (the declaration on the cluster side is unchanged too).

apply does the same comparison as plan once more and brings the difference into line with the declaration. Whether to revert the drift, or to fix the declaration because what was changed from outside was right, is for a person to judge — this time we judge that the declaration is right and revert.