When the Rollout Stalls on One Node
In one line
A DaemonSet's maxUnavailable: 1 means that if one node cannot bring up the new Pod, it stops as it is without touching the remaining nodes. This is not a bug but a promise, and as a result the cluster is left split into nodes with the old configuration and nodes with the new one.
Why let one node stall the whole rollout
Think of the opposite and the answer comes out. If the new spec is wrong and the DaemonSet keeps pushing without stopping, what happens? All seven GPU nodes die at the same time. The purpose of a rolling update is not to finish quickly but to contain the blast radius of a bad change to one machine. So a DaemonSet changes one machine, waits until that Pod is ready, and if it does not become ready, waits forever.
The problem is that this stop is quiet. kubectl get ds looks like this.
NAME DESIRED CURRENT READY UP-TO-DATE AVAILABLE
nvidia-container-toolkit 3 3 2 1 2
It is not an error. There is no red text. READY 2/3 reads at a glance as "almost done". But what this table says is the fact that the cluster has split into two generations. One machine cannot come up with the new spec, and two are still on the old spec. If it is a DaemonSet that edits node files, as with GPU configuration, this state means "a cluster where the runtime configuration differs per node".
How it works
A DaemonSet rolling update goes node by node like this.
- Pick one node that is still on the old spec.
- Delete first that node's Pod.
- Create a Pod of the new spec.
- Wait until that Pod becomes available.
- When room of
maxUnavailableopens up again, go back to step 1.
If in step 3 the new Pod cannot be scheduled, step 4 does not finish and step 5 never comes. It matters that step 2 comes before step 3 — the stalled node becomes an empty state with neither the old Pod nor a new Pod up.
The reason the new Pod does not come up is usually one of three. It cannot pull the image, the node cannot supply the requested resources, or the toleration or nodeSelector is off. In the first two the Pod stays Pending or ImagePullBackOff, and the last is different in nature — if the toleration is off, the DaemonSet controller judges "this node must not have a Pod" and removes it from the targets, so the DESIRED number itself shrinks. It looks like the same symptom, but the column to look at is different.
kubectl rollout status ds/... just hangs in this state. If this command is in a CI pipeline without a timeout, the job stays stuck for an hour each time, so always give --timeout.
What it looks like in the field
First, set the alert on numberUnavailable. If a DaemonSet stays in the state desiredNumberScheduled != numberReady for more than a few minutes, a person must look. Pod restart counts or error logs do not catch this stop — because nothing restarts and no error appears.
Second, build the habit of checking the split generations. The fastest way to see which spec is actually running on each node is to extract the Pod's image and node together. If you look only at the UP-TO-DATE number, you cannot tell which node fell behind, and in GPU configuration that "which node" is the scope of the failure.
Third, rolling back is often faster than going forward. While you fix the cause of the new spec not coming up, one node stays empty. It is better to restore the old spec with kubectl rollout undo, bring the three nodes back to the same state, and find the cause afterwards. The first thing to do in incident response is not diagnosis but gathering the state into one.
What you will do in the next lab
In the next lab, you handle the three methods of sharing through configuration and objects. You hold several sets of device plugin configuration under names and have them chosen by node labels, confirm through scheduling that MIG nodes have different resource names so old manifests cannot use those nodes, and pin the share handed to a team with a quota. You meet the shape of a stalled rollout again in the last lab of this course.