How to Read a Drain That Never Finishes
Summary
A drain does not delete Pods; it calls the eviction API, and the PodDisruptionBudget examines that request and returns 429 if the budget is short. The cause of a blocked drain is almost always one of three things: the budget is 0, the selectors overlap, or Pods that are not ready have already broken the budget.
Why diagnosis is needed separately
Creating a PDB is not hard. One page of documentation is enough. But what eats time in the field is not the creating side but when it gets blocked. The command you typed to empty one node keeps printing the same line for 30 minutes, and that line contains nothing but error when evicting pods. Which budget is blocking it, how many are short, and how many Pods can be evicted right now are not on that screen.
So you first need to know what a drain actually does. kubectl drain does not delete Pods. For each Pod it sends a POST to the eviction subresource (/api/v1/namespaces/<ns>/pods/<pod>/eviction). This request goes through PDB examination inside the API server, and if the budget is short, it is rejected with 429 Too Many Requests. And the drain does not give up but tries again. This is why a blocked drain does not end in failure but goes on forever. Conversely, kubectl delete pod skips this examination entirely — so you are tempted to use it in a hurry, but that is a declaration that you will ignore the budget.
How it works
A PDB's state is not one number but four. In kubectl get pdb <이름> -o json (where the placeholder is the PDB name), .status looks like this.
| Field | Meaning |
|---|---|
expectedPods |
How many Pods are expected to match the selector |
currentHealthy |
How many of them are healthy right now |
desiredHealthy |
The minimum number that must remain after eviction |
disruptionsAllowed |
How many can be evicted right now |
The word "healthy" here has a precise definition. The official documentation says it counts as healthy a Pod that has an item in .status.conditions with type=Ready and status=True. Being Running does not make it healthy.
You can use only one of minAvailable and maxUnavailable. The rounding direction when you use a percentage matters. If there are 7 Pods and minAvailable: "50%", it is 3.5, but Kubernetes rounds up and requires 4. That is the safe side. But if you write maxUnavailable as a percentage, the number that can be evicted is rounded up. So, as the official documentation puts it, the actual disruption can exceed the percentage you wrote, and when there is only one replica, maxUnavailable: 30% makes that one evictable, ending up as a 100% disruption.
The second trap is budgets with overlapping selectors. If a single Pod falls under two or more PDBs, eviction is rejected even when each budget has room. This is because the eviction subresource does not support that situation. The documentation also says to avoid overlapping selectors, and gives only the transition period of moving Pods from one budget to another as a legitimate use.
The third is unhealthyPodEvictionPolicy. It is a field that became stable in 1.31, and the default is IfHealthyBudget. With this default, to evict a Pod that is not yet healthy, that application must not already be breaking its budget. That is, if there is even one Pod in CrashLoopBackOff or one that cannot report Ready, even that broken Pod itself cannot be evicted. The drain stops dead right there, forever. If you change it to AlwaysAllow, Pods that are running but not healthy can be evicted regardless of the budget. This is why the official documentation says it recommends this value to support node drains.
Finally, you also need to know what a budget does not protect. A PDB blocks only voluntary disruptions. It cannot stop a node from simply dying, which is why the documentation states firmly that a budget does not always guarantee that number. And if you set maxUnavailable: 0, or set minAvailable equal to the replica count, you have made voluntary eviction 0, so the node holding that Pod can never finish draining. This is not a bug but the meaning exactly as the documentation writes it.
What it looks like in the field
The most common case is putting minAvailable: 1 on a workload with 1 replica. The intent was "this service must never be cut off," but the result is that the node that workload is on can never be emptied. The whole upgrade plan stalls.
The second is not restoring the node after a drain gets blocked. Before starting evictions, a drain first puts the node in an unschedulable state. Even if you interrupt the command because eviction is blocked, that mark remains. Without anyone noticing, cluster capacity is reduced by one node, and you pay for it weeks later when traffic surges.
The third is workloads with no budget. This is an incident in the opposite direction — the drain passes with no resistance and the replicas all go down at once. That is why the first material you should build when designing a maintenance window is a list of "workloads with 2 or more replicas that are not covered by any PDB."
Limits of this lab environment
A kwok Pod is not a real container, so you cannot really create CrashLoopBackOff. Instead, you directly edit the Pod's status.conditions to make Ready False — the very condition a PDB looks at, so the disruption controller really reduces currentHealthy by one, and the eviction verdict really changes. When I checked, it was 429 under the default policy and passed when changed to AlwaysAllow. Also, if you actually carry out an eviction, the Pods the later steps use disappear, so this lab gets most of its verdicts with ?dryRun=All. In practice too, you can ask ahead of time in the same way before setting a maintenance window.
What to do in the next lab
You gather a workload on one node, create a budget with 0 allowed disruptions, and call the eviction API directly to get a 429. Next you watch a drain actually stall, restore the node, then fix the budget and confirm the same request passes. Then you put a 50% budget on 7 replicas to confirm the rounding direction with numbers the cluster computed, create a deadlock with overlapping-selector budgets and read its message. You create a Pod that is not ready to measure the difference between the default policy and AlwaysAllow, and finally build an audit script that finds workloads with no budget across the whole cluster.
References:
- https://kubernetes.io/docs/concepts/workloads/pods/disruptions/
- https://kubernetes.io/docs/tasks/run-application/configure-pdb/
- https://kubernetes.io/docs/tasks/administer-cluster/safely-drain-node/