TT Lab
Get started
Learn Learning paths Courses

Kubernetes Distributions — Build Them Yourself

Upgrade one minor version at a time

Continue in TT Lab

One-line summary

A Kubernetes upgrade is not "installing the latest version" but raising the minor version one step at a time while keeping the version difference (skew) between components within the allowed range. k3s is no exception to this rule.

Why this was needed

Because k3s is installed with a single line, it is easy to think that upgrading is a single line too. If you run the install script again, it downloads the new binary and brings the service back up. So people often do this: "What is the stable channel right now? Let's upgrade to that."

There is a trap here. A channel only points to the latest recommended version at that moment, and it does not know what version your cluster is on now. If you upgrade a cluster stuck on 1.34 with the stable channel, then on a day when stable points to 1.36, you skip two minor versions at once. The install script does not prevent this.

But the Kubernetes project policy is clear. The version skew policy says that when upgrading kube-apiserver you must not skip a minor version, and that is true even for a cluster with only one instance. The k3s manual upgrade documentation also says the same policy applies and tells you not to skip intermediate minor versions. The reason is that conversions of objects in the datastore, removal of deprecated APIs, and changes in controller behavior are tested only one version at a time. If you skip two, you are the first to walk a path nobody has tested.

How it works

The allowed range of version differences

The skew policy decides, for each component, "how many versions it may differ from the API server." The key point is that no component may be newer than the API server.

구성요소                  API 서버가 1.35 일 때 허용되는 판
kube-apiserver (HA)       서로 한 마이너 이내
controller-manager·scheduler  1.35, 1.34 (한 판 낮게까지)
kubelet·kube-proxy        1.35, 1.34, 1.33, 1.32 (세 판 낮게까지)
kubectl                   1.36, 1.35, 1.34 (위아래 한 판)

So the order of upgrading is decided too. The control plane (API server) comes first, and the kubelet comes later. If you do it the other way around, there is a moment when the kubelet is newer than the API server. This is why the k3s documentation requires the order of server nodes first, one at a time, then agent nodes. Also, thanks to the fact that the kubelet may lag up to three versions behind, you do not have to bring the workers up each time while you raise the control plane one step at a time, several times.

What actually happens in k3s

In k3s, the API server, kubelet, and controllers are a single binary. If you run the install script again with the version pinned via INSTALL_K3S_VERSION, it downloads the binary of that version and restarts the service. On a single node, the control plane and the kubelet go up together, so skew arises only between nodes.

There are two things to watch out for.

First, settings you gave by environment variable or argument at install time disappear if you do not give them again when you run it again. This is behavior the k3s documentation states explicitly. So it is safer to keep settings in /etc/rancher/k3s/config.yaml, because it remains regardless of the install script.

Second, even if you stop k3s, the Pods' containers keep running. So even if the API is briefly cut, services are usually still alive, but for workloads that become a problem while the API is cut, the documentation recommends draining first.

I will also write down what was seen by measurement. On the same VM as this lab, upgrading v1.34.11+k3s1 to v1.35.8+k3s1 with the install script took 11 seconds, and the two nginx Pods that had been brought up beforehand kept the same Pod UIDs and container IDs, with a restart count of 0. On the other hand, the kubeadm upgrade documentation says that the container spec hash changes so all containers restart after the upgrade, and tells you to always drain before a minor-version kubelet upgrade. Even for the same "Kubernetes upgrade", what workloads go through differs depending on the distribution and version combination, so do not assume it from a single line of documentation; measure.

What drain does and what it cannot do

kubectl drain cordons the node (forbids placing new Pods) and then sends Pods out through the eviction API. Eviction honors the PodDisruptionBudget. An eviction that would break the budget is rejected, and drain keeps retrying the rejected Pod.

This means that if there is nowhere to go, drain does not finish. If there is only one node, the replacement Pod for an evicted Pod cannot be placed on that cordoned node and becomes Pending, the number of ready Pods falls below the budget, and the next eviction is blocked. In measurement, drain repeated Cannot evict pod as it would violate the pod's disruption budget eight times at 5-second intervals and ended when it hit --timeout=40s. In this lab you see that scene for yourself.

Backup and deprecated API check

The default datastore of k3s is SQLite. The backup and restore documentation tells you to be sure to take the token file (/var/lib/rancher/k3s/server/token) together with /var/lib/rancher/k3s/server/db/. The token is used to encrypt the confidential data in the datastore, so if you have only the DB and not the token, you cannot restore.

For deprecated APIs, check what disappears in each version in the API migration guide, and use the API server's apiserver_requested_deprecated_apis metric to confirm who is actually calling that API. If you look only at the documentation, you come to believe "we don't use it", and if you look at the metric, it turns out that an old Helm chart or script is still calling the old version.

What it looks like in the field

It is common to neglect dozens of k3s machines scattered at the edge for a while and then upgrade them in a hurry because of a security notice. If you point at the stable channel at that time, thinking "we're going to the latest anyway," each node starts from a different version, so some nodes skip one step and some skip three. The upgrade succeeds, but afterward, objects created with old APIs and workloads that relied on removed features quietly break.

For automation, use Rancher's system-upgrade-controller. You write the target version (version) or channel (channel) in a Plan object and decide how many machines to upgrade at once and how with concurrency, cordon, and nodeSelector. Here too, pointing at a channel creates the same trap. Moreover, in measurement, the Plan CRD (v0.20.1) accepted a Plan with neither version nor channel in server validation as it was. That applying succeeded does not mean the version is right. Pinning the version and changing the Plan one step at a time matches the skew policy.

What really matters in practice

What you will do in the next lab

You start from a single k3s 1.34 in the VM. You record the starting version, bring up a workload with a PDB, and then prepare with the deprecated API metric and an SQLite backup. You see why drain stops on a single node, raise it one step to 1.35 with the install script, and then compare whether the workloads and the node are alive as the same objects. Finally, you write a system-upgrade-controller Plan for the next step and calculate how many steps you would have skipped had you followed the stable channel.