Building clusters with Kubespray and Terraform
Evidence, not green lights — verify the install layer by layer
Goal
You verify a cluster built with kubespray in the order node → core Pods → DNS → certificates → leftover files → real workload, and leave "what to do next and when" as a one-page report.
Why it matters
failed=0 in the PLAY RECAP means "the playbook ran to the end", not "the cluster is usable." kubespray puts a few things of its own on top of kubeadm in its own way — it sets up etcd separately as a host service, puts a node-local cache in front of Pod DNS, and moves the certificate directory. So if you look at it with the same habits you used to check clusters you built directly with kubeadm, you miss the etcd certificates, or you judge that DNS works by looking only at CoreDNS. If you get used to the order of personally checking each layer once, you can check in the same order after installation and after an upgrade too.
Steps
- In
/opt/ks/kubespray, runansible-playbook -i /root/ks/inventory/lab/inventory.ini cluster.ymland leave the entire output in/root/ks/logs/cluster-1.log(about 7 minutes). node1 in the PLAY RECAP must havefailed=0. - Write to
/root/ks/verify/node.jsonready(the Ready condition status string),kubelet_version,container_runtime(nodeInfo.containerRuntimeVersion),internal_ip,roles(a sorted array of the role names from the labels that start withnode-role.kubernetes.io/), andtaints(an array of키:효과strings, that is, key:effect, or an empty array if none). - Look at the Pods in kube-system and write to
/root/ks/verify/pods.jsonapps(a deduplicated sorted array of thek8s-applabel of each Pod, or itscomponentlabel value if there is none),not_ready(an array of the names of the Pods that are not Ready, or an empty array if none),etcd_is_pod(whether etcd is running as a Pod, a boolean), andetcd_unit(the name of the systemd unit that started etcd, including .service). - Write to
/root/ks/verify/dns.jsonkubelet_cluster_dns(the first clusterDNS value in/var/lib/kubelet/config.yaml),coredns_service_ip(the ClusterIP of the CoreDNS Service in kube-system), andanswer_via_nodelocalandanswer_via_coredns(the A records obtained by asking each of the two addresses from the node forkubernetes.default.svc.cluster.local). Both answers must equal the ClusterIP of the kubernetes Service. - Write to
/root/ks/verify/certs.jsonapiserver_not_after,ca_not_after(apiserver.crt and ca.crt in/etc/kubernetes/ssl), andetcd_member_not_after(/etc/ssl/etcd/ssl/member-node1.pem) in the UTC2026-01-01T00:00:00Zformat,apiserver_valid_daysandetcd_member_valid_days(the number of days from notBefore to notAfter of each certificate, integers), andkubeadm_lists_etcd(whether the output ofkubeadm certs check-expirationhas an etcd certificate line, a boolean). - Write to
/root/ks/verify/files.jsonkubeadm_config_version(the ClusterConfiguration kubernetesVersion in/etc/kubernetes/kubeadm-config.yaml),etcd_data_dir(ETCD_DATA_DIR in/etc/etcd.env),kubeconfig_server(the server address in/root/.kube/config),containerd_version(the third column ofcontainerd --version), andreleases_mb(the size of/tmp/releases, an integer in MiB —du -sm). - In the default namespace, create a Deployment
web(imageregistry.k8s.io/e2e-test-images/agnhost:2.59, argumentsnetexec --http-port=8080, 2 replicas) and a ClusterIP Service of the same name (port 80 → 8080). Once both Pods are Ready, callcurl http://<web 서비스 IP>/hostname(the placeholder is the web Service IP) from the node several times to confirm that both Pod names come back, and write those two names, sorted, one per line, to/root/ks/verify/smoke.txt. - Write to
/root/ks/verify/report.jsonnode_ready(a boolean),core_apps(the number of apps from step 3, a number),dns_ok(whether both answers in step 4 equal the kubernetes Service IP, a boolean),first_cert_to_expire(which expires first of the kubeadm leaf certificate and the etcd member certificate:"kubeadm"or"etcd"),days_until_first_expiry(the number of days remaining until that certificate expires, an integer), andsmoke_pods(the number of Pods that responded in step 7, a number).
Notes
- kubespray v2.32.0 is ready at
/opt/ks/kubesprayand the inventory at/root/ks/inventory/lab/inventory.ini(kube_version 1.35.8). You do the installation yourself in step 1 (about 7 minutes). digis installed.kubectlandkubeadmappear in/usr/local/binonce the installation is done.- Common mistake: ending the certificate check after looking only at
kubeadm certs check-expiration. In this deployment the etcd certificates are not in that list. - Common mistake: saying you checked DNS just because the CoreDNS Pods are Running. The address Pods actually ask is separate.
- Documentation: Kubespray — DNS stack · Kubernetes — Certificate Management with kubeadm · Kubernetes — NodeLocal DNSCache
Set up the cluster to verify
In /opt/ks/kubespray, run ansible-playbook -i /root/ks/inventory/lab/inventory.ini cluster.yml and leave the entire output in /root/ks/logs/cluster-1.log (about 7 minutes). node1 in the PLAY RECAP must have failed=0.
It is the same as the installation you did in module 3. Start it with systemd-run or tmux so that it keeps running even if the console is cut off, and pass HOME=/root. While you wait, it is a good idea to read the next steps of this lab.
Node — Ready and taints
Write to /root/ks/verify/node.json ready (the Ready condition status string), kubelet_version, container_runtime (nodeInfo.containerRuntimeVersion), internal_ip, roles (a sorted array of the role names from the labels that start with node-role.kubernetes.io/), and taints (an array of 키:효과 strings, that is, key:effect, or an empty array if none).
kubeadm puts a NoSchedule taint on the control plane node. But this node is also in kube_node in the inventory — confirm from the taint list what kubespray did upon seeing that fact. The role name is the last piece of the label key.
Core Pods — what is running and what is not
Look at the Pods in kube-system and write to /root/ks/verify/pods.json apps (a deduplicated sorted array of the k8s-app label of each Pod, or its component label value if there is none), not_ready (an array of the names of the Pods that are not Ready, or an empty array if none), etcd_is_pod (whether etcd is running as a Pod, a boolean), and etcd_unit (the name of the systemd unit that started etcd, including .service).
The difference from a cluster you built directly with kubeadm shows here. Find in group_vars/all/etcd.yml what kubespray's default etcd arrangement (etcd_deployment_type) is, and confirm the actual unit with systemctl list-units.
DNS — the place Pods ask is not CoreDNS
Write to /root/ks/verify/dns.json kubelet_cluster_dns (the first clusterDNS value in /var/lib/kubelet/config.yaml), coredns_service_ip (the ClusterIP of the CoreDNS Service in kube-system), and answer_via_nodelocal and answer_via_coredns (the A records obtained by asking each of the two addresses from the node for kubernetes.default.svc.cluster.local). Both answers must equal the ClusterIP of the kubernetes Service.
kubespray turns on nodelocaldns (enable_nodelocaldns: true) by default. A caching DNS runs at a link-local address on each node, and the kubelet tells Pods that address. The name of the CoreDNS Service may differ from the kubeadm default, so find it by selector in the Service list. dig +short @<주소> <이름> returns only the answer (the placeholders are the address and the name).
Certificates — the expiration dates of two branches
Write to /root/ks/verify/certs.json apiserver_not_after, ca_not_after (apiserver.crt and ca.crt in /etc/kubernetes/ssl), and etcd_member_not_after (/etc/ssl/etcd/ssl/member-node1.pem) in the UTC 2026-01-01T00:00:00Z format, apiserver_valid_days and etcd_member_valid_days (the number of days from notBefore to notAfter of each certificate, integers), and kubeadm_lists_etcd (whether the output of kubeadm certs check-expiration has an etcd certificate line, a boolean).
kubespray puts kubeadm's certificate directory at /etc/kubernetes/ssl (the kubeadm default is pki). In the default arrangement that runs etcd as a host service, the etcd certificates are created not by kubeadm but by kubespray with openssl, and their duration is certificates_duration in kubespray_defaults. This is why you must not say 'certificate expiry check done' after looking at only one side.
What kubespray left behind
Write to /root/ks/verify/files.json kubeadm_config_version (the ClusterConfiguration kubernetesVersion in /etc/kubernetes/kubeadm-config.yaml), etcd_data_dir (ETCD_DATA_DIR in /etc/etcd.env), kubeconfig_server (the server address in /root/.kube/config), containerd_version (the third column of containerd --version), and releases_mb (the size of /tmp/releases, an integer in MiB — du -sm).
kubespray leaves on the node the configuration file it used when calling kubeadm and the etcd environment file. They are the primary sources for checking later what values the installation was done with. /tmp/releases is the binary cache kubespray downloaded (local_release_dir), and it is also why reinstalling or upgrading is fast.
The last evidence is a workload
In the default namespace, create a Deployment web (image registry.k8s.io/e2e-test-images/agnhost:2.59, arguments netexec --http-port=8080, 2 replicas) and a ClusterIP Service of the same name (port 80 → 8080). Once both Pods are Ready, call curl http://<web 서비스 IP>/hostname (the placeholder is the web Service IP) from the node several times to confirm that both Pod names come back, and write those two names, sorted, one per line, to /root/ks/verify/smoke.txt.
If this step passes, image pulling (containerd and registry egress), scheduling (taints), the Pod network (calico), and the Service (kube-proxy) are all confirmed at once. This is why we send a real request instead of looking at green lights. agnhost's netexec returns the Pod name at /hostname.
Verification report
Write to /root/ks/verify/report.json node_ready (a boolean), core_apps (the number of apps from step 3, a number), dns_ok (whether both answers in step 4 equal the kubernetes Service IP, a boolean), first_cert_to_expire (which expires first of the kubeadm leaf certificate and the etcd member certificate: "kubeadm" or "etcd"), days_until_first_expiry (the number of days remaining until that certificate expires, an integer), and smoke_pods (the number of Pods that responded in step 7, a number).
You can recalculate everything from the records of earlier steps and the current cluster. Compute the remaining days as an integer rounded down, from the current time to notAfter. The purpose of this report is to leave 'what to do next and when' in one line.