Building clusters with Kubespray and Terraform
Replacing the file vs. using the new certificate — renewal and reset
One-line summary
Certificate renewal is two steps: changing the file and restarting the process that reads that file. kubespray's timer takes over those two steps only when expiry is near, and reset.yml erases the cluster but leaves behind the downloaded files and a few other traces.
Why this was needed
According to the kubeadm certificate management documentation, client certificates created by kubeadm expire after a year, and kubeadm renews them all when you upgrade the control plane. The documentation therefore says "upgrading frequently is the best approach." But in the field it is common to keep a cluster's version pinned for more than a year, and such a cluster has authentication between the API server and the controllers fail all at once on the day the installation turns one year old. Without an alert, it takes a long time just to find the cause.
How it works
In kubespray the kubeadm certificates are in /etc/kubernetes/ssl. The kubeadm default is /etc/kubernetes/pki, but kubespray sets kube_cert_dir to ssl and links pki to it. kubeadm certs check-expiration shows the leaf certificates (apiserver, apiserver-kubelet-client, front-proxy-client, and the client certificates inside admin.conf, controller-manager.conf, scheduler.conf, and super-admin.conf) and the CA, and in this arrangement, which has an external etcd, it does not show the etcd certificates (module 4).
Renewal is two steps. kubeadm certs renew all re-signs the leaf certificates with the CA and changes the files and kubeconfigs. And it prints "you must restart kube-apiserver, kube-controller-manager, kube-scheduler, and etcd for them to use the new certificates." When measured, the API server's serving certificate came out on 6443 with a new serial even before the restart as soon as the file changed — because kube-apiserver rereads the serving certificate file. But controller-manager and scheduler read the client certificate inside the kubeconfig at startup. So after renewal, kubespray's k8s-certs-renew.sh deletes the sandboxes of the three static Pods with crictl rmp (the kubelet sees the manifests and starts them again), replaces /root/.kube/config with the new admin.conf, and waits until 6443 opens again. In measurement, it took 6 seconds from deleting the sandboxes until /readyz came back.
The timer moves only when expiry is near. If you set auto_renew_certificates: true, the control plane role installs k8s-certs-renew.timer. The default calendar is Mon *-*-1,2,3,4,5,6,7 03:00:00, that is, 3 a.m. on the first Monday of every month. The script renews and restarts only when there is a certificate that expires before "the next timer time + 7 days." If you run the service with a certificate that was just renewed, it ends with ## Skip cert renew and K8S container restart, since all residualTimes are beyond threshold ##. It is designed so that it restarts only once, about a month before expiry, without shaking the control plane every month.
reset.yml cannot be undone. So if you do not give reset_confirmation=yes, it stops at a prompt. The role stops the services (kubelet, containerd, etcd), deletes containers and Pods, clears the iptables and IPVS rules, and deletes a long list of files and directories — /etc/kubernetes, /var/lib/kubelet, the etcd data, the containerd store, /etc/cni, ~/.kube, and the binaries kubespray installed. This delete task is ignore_errors, so even if one fails it keeps deleting the rest. Conversely, what is not on the list remains. In measurement, the downloaded file cache /tmp/releases (543MB), /usr/local/bin/etcdutl, and the hostname that kubespray changed remained. The downloaded files remaining is an intended convenience that makes reinstallation faster (measured 326 seconds), and etcdutl is something that was left off the delete list.
What it looks like in the field
The typical sequence of a year-one outage is this. One day kubectl is rejected with x509: certificate has expired, and a little later the controller stops and new Pods are not created. In a hurry you run kubeadm certs renew all, and it is still strange — the files are new but controller-manager is running with the old kubeconfig, and the operator's ~/.kube/config still has the old certificate too. You are done only after the second step (restarting and distributing the kubeconfig). This course's recommendation is simple. Do an upgrade at least once within a year, and for a cluster where you cannot, turn on auto_renew_certificates and put the check-expiration result into monitoring.
You use reset in two cases. When the first installation stopped before the CNI and running again does not continue it, as seen in module 3, and when you must change a value that is hard to change after installation (the network plugin, the Pod and Service CIDR ranges). In both, "erase and do it again with the same inventory" is the fastest path, and it is a choice that is possible because the inventory remains as the declaration.
What you will do in the next lab
After building the cluster and recording the certificate serials, you renew by hand and restart the static Pods to confirm the new certificates are used. You turn on the automatic renewal timer and run it once now to see why the script does not renew. After erasing with reset.yml, you write down what remains, build it again with the same inventory, and check that the CA and the node are new.