TT Lab
Get started
Learn Learning paths Courses

Kubernetes Operations

Who Decided on Five Minutes After a Node Dies

Continue in TT Lab

Summary

When something goes wrong with a node, the node controller puts a NoExecute taint on it, and from then on when a Pod disappears is decided by the Pod's tolerationSeconds. Most Pods have never had this value written, and live with the 300 seconds that admission put in.

Why you need to know this value

A drain is typed by a person. We choose the time and the target. But the incidents that call you out at dawn are usually the opposite — nobody typed a command, yet Pods start moving around, some workloads move too fast and lose their entire local cache, and some stay attached to a dead node for too long so traffic goes to an empty place. The reason these two happen in the same cluster at the same time is simple. The two workloads are using the same value.

That value is 300 seconds. When it creates a Pod, Kubernetes automatically attaches NoExecute tolerations for the two keys node.kubernetes.io/not-ready and node.kubernetes.io/unreachable with a tolerationSeconds of 300. Only when you or a controller did not state those tolerations explicitly. As the official documentation puts it, because of these automatically added tolerations, a Pod stays bound to the node for 5 minutes after the problem is detected.

How it works

First, the node side. The kubelet periodically renews its own Lease to announce that it is alive. When this renewal stops, the node controller changes the node's Ready condition to Unknown and attaches the corresponding taint. The conditions and taints are paired like this.

Node condition Taint attached Default effect
Ready = False node.kubernetes.io/not-ready NoExecute
Ready = Unknown node.kubernetes.io/unreachable NoExecute
MemoryPressure node.kubernetes.io/memory-pressure NoSchedule
DiskPressure node.kubernetes.io/disk-pressure NoSchedule
PIDPressure node.kubernetes.io/pid-pressure NoSchedule
NetworkUnavailable node.kubernetes.io/network-unavailable NoSchedule

It is important that the scheduler looks at taints, not node conditions. If you made it look at conditions directly, placement decisions would branch as many ways as there are kinds of conditions, but once you translate them into taints, both placement and eviction are handled by a single rule.

Next is the difference in effects. NoSchedule blocks only new placements and does not touch Pods that are already running. PreferNoSchedule is its weak version. Only NoExecute evicts Pods that are already running. Here the tolerations split three ways. With no toleration that tolerates it, eviction is immediate; if it tolerates but has no tolerationSeconds, the Pod stays forever; and if seconds are written, it is evicted after holding out for that many seconds. If the taint is removed before that, the eviction does not happen.

A DaemonSet is an exception. Pods created by the DaemonSet controller get NoExecute tolerations for the two keys not-ready and unreachable without tolerationSeconds. That is because if the node agents also left whenever a node wobbled, the means to observe or recover that node would disappear with them.

One more thing. To keep evictions from piling up all at once, the control plane limits the rate at which it attaches new taints to nodes. It is a mechanism that prevents the whole cluster from being rescheduled in an instant during a large-scale network partition. And since 1.29, this eviction implementation has been split out of the node controller and is handled by a separate controller called taint-eviction-controller.

What it looks like in the field

The most common incident is a stateful workload moving wholesale because of a short network wobble. It is a 40-second switch restart, but after 300 seconds the Pod is already on another node, looking at an empty local disk and downloading data from scratch. For such workloads it is better to give tolerationSeconds a long value — the official documentation also notes that you may want an application with a lot of local state to stay bound to a node for a long time during a network partition, and gives 6000 seconds as an example.

The incident in the opposite direction comes from the same value. The node has really died, and for 300 seconds the service runs short of replicas. For a workload like a frontend that is the same wherever it runs, these 5 minutes are a total loss. However, in such places it is usually better to increase replicas and spread the topology than to just reduce the value — if you only reduce the value, rescheduling becomes frequent every time something wobbles and it actually becomes unstable.

The third is forgetting to remove the taint. If you put on NoExecute for maintenance and emptied the Pods, and do not delete the taint after the maintenance is over, that node stays empty. It is not rare for weeks to go by with the cluster capacity reduced by a third.

Limits of this lab environment

The nodes in this cluster are fake nodes made by kwok, so there is no real kubelet. So you cannot cut the heartbeat to make the node controller attach the taint itself. On top of that, I measured one thing here — even if you attach node.kubernetes.io/unreachable:NoExecute by hand to a Ready node, it disappears within a few seconds. The taints with those two keys are managed directly by the node controller based on the node conditions, so one attached to a node whose condition is healthy is treated as a wrong state and swept away (even if you directly edit the node status to change Ready to False, kwok immediately reverts it). So in the lab, you put on the same NoExecute taint with a custom key. Only the key differs; the eviction rules are literally the same, and the eviction that happens afterward — who disappears when — is real behavior by a real controller.

Also, a kwok Pod is not a container, so there are no container logs, no exec, and no OOM. The scene of a Pod moving and downloading data again cannot be seen in this environment, and what we can see is the time at which the object disappears. That time is exactly what this lab covers.

What to do in the next lab

You set up a control group on one node that you will not touch, and see with your own eyes the default toleration that nobody wrote. Then you put two Pods that differ only in how long they tolerate onto different nodes and see that nothing happens with NoSchedule, while with NoExecute one of them actually disappears. Next you put a NoExecute taint that imitates a broken link on another node to measure the difference between 20 seconds and 3600 seconds, and build a script that computes the eviction timetable for the whole namespace. At the end you decide the value to give a stateful workload, put it on the cluster, and leave the rationale.

References: