Building clusters with Kubespray and Terraform
Where Terraform meets Kubespray — templates before, wrapping in the middle, declarations after
One-line summary
Terraform is a tool that remembers "what should exist" in state and changes only the difference, so the boundary is this: in front of kubespray it makes the inventory, behind it it declares the resources on top of the cluster and catches drift, and the installation in the middle is left to kubespray.
Why this was needed
The input of kubespray is the inventory, and the addresses and names in the inventory are known by whoever created the nodes. If you created the VMs in the cloud with Terraform, it is natural to make the inventory from its outputs (IPs, names, roles), and the kubespray repository's contrib/terraform also has per-cloud examples such as aws, gcp, openstack, and vsphere. The other side is the same. After the cluster is up, setting up namespaces, permissions, and default apps for each team is repetitive work, and you have to notice what someone changed by hand. Terraform's state and plan do exactly that job. The problem is the middle. Installing Kubernetes is a procedure in which hundreds of tasks must keep their order, and kubespray already does it idempotently. Rewriting that as Terraform resources would be doing the work twice.
How it works
In front: the inventory as a template. templatefile() reads one file and fills it in with variables. If you put a node map (name → {control_plane, worker, connection}) in a variable and fill the groups with the %{ for } and %{ if } directives, adding a node becomes adding one line to the map. local_file writes the file, and its content remains in state. So if someone edits the file by hand, the next plan says it will revert it — this means Terraform has become the owner of this file, and it is also a way to settle on a single owner for values that tend to be written in many places, like the version (kube_version).
In the middle: only wrap it. terraform_data is a built-in resource that creates no infrastructure. It is recreated when the value you put in triggers_replace changes, and when it is created, the local-exec provisioner runs a command. If you put the inventory, version, and host_vars file contents in it and set the command to ansible-playbook cluster.yml, it becomes "call kubespray only when the declaration changes." If the run fails, the resource is left tainted and is called again at the next apply.
The limits show here too. If you change the version from 1.35.8 → 1.36.4, the plan says it will recreate the version file and terraform_data, and then what gets called is cluster.yml. A kubespray upgrade must be upgrade-cluster.yml (cordon, drain, one step at a time). Terraform knows "what changed" but does not know "with what procedure to apply that change." So it is safer to wrap the installation, but leave jobs with a different procedure, such as upgrades and node removal, to the pipeline or a person who picks and calls the kubespray playbook.
Behind: declare what is on top of the cluster. The kubernetes provider attaches to the API server with a kubeconfig and creates resources such as kubernetes_namespace_v1, kubernetes_service_account_v1, kubernetes_role_v1, and kubernetes_role_binding_v1, and the helm provider installs charts with helm_release. In helm provider 3.x, the Kubernetes connection setting is the kubernetes = { ... } attribute and the values are a set = [{ name, value }] list. It is good to separate the root module that creates the cluster from this module — if you create the cluster in one module and configure a provider with that cluster in the same run, the first plan has to read a kubeconfig that does not exist yet, and the order gets tangled.
Drift. tofu plan first reads the actual state to refresh the state, and compares it with the declaration. If someone changes a namespace label with kubectl label, the plan says "I will revert it to the declaration," and -detailed-exitcode returns 2. If you hook this exit code into a scheduled job, it becomes a drift alert. However, helm_release compares the release metadata (the chart and values), so it does not always catch an external edit of a Deployment that the chart created. You need to know which tool watches what.
What it looks like in the field
A common design is repositories split in three. Terraform that creates the nodes (per cloud account), a pipeline that takes the inventory and calls kubespray, and Terraform that declares what is on top of the cluster (per team). This lab compresses these three into one VM. The first layer, creating the nodes, is not done here — on this platform that layer is the virtualization API that creates VMs, and if the lab had you call that API, the isolation between students would break. In the field, a cloud provider or the vSphere or OpenStack provider takes that place.
State is also something to be careful with. This lab keeps the state as a local file, but when a team uses it, you need a remote backend and locking, and since file contents and resource attributes go into state in plaintext, you must not put secrets in it.
What you will do in the next lab
You make the inventory, host_vars, and version file from the node map and the template and apply them, and see Terraform become the owner of those files. You wrap kubespray with terraform_data to set up the cluster, and use plan to check what gets called again when you change the version. In a different root module you declare the namespace, RBAC, and a helm release, and catch a label changed from outside with plan and revert it.