The Drain Has Been Running for Thirty Minutes
Goal
When a drain is blocked, you call the eviction API directly to read what is blocking it, confirm with numbers the cluster computed the rounding rule of percentage budgets, the deadlock of budgets with overlapping selectors, and the eviction policy for Pods that are not ready, and then find the workloads with no budget across the whole cluster and make a list.
Why it matters
kubectl drain does not delete Pods. It calls the Eviction subresource for each Pod, and the PodDisruptionBudget examines that request. If the budget is short, what comes back is not a deletion but an HTTP 429, and the drain tries again until it works. That is why a blocked drain does not end in failure but goes on forever. If you do not know this structure, you cannot find the cause from just the error when evicting pods on the screen. A budget is described not by one number but by three, currentHealthy, desiredHealthy, and disruptionsAllowed, and if you know how to read those numbers, the reason for being blocked usually narrows down to one of three — the budget is 0, the selectors overlap, or Pods that are not ready have already broken the budget.
Steps
- Create the namespace
ops-pdband, in/root/ops-disruption/api.yaml, write the Deploymentapi— 4 replicas, labelapp: api, imagenginx:1.27.3, andnodeSelectoriskubernetes.io/hostname: lab-node-0. Apply it and wait until all 4 Pods are Running on lab-node-0. - In
/root/ops-disruption/api-pdb.yaml, write the PodDisruptionBudgetapi-pdb— namespaceops-pdb,minAvailable: 4, selectorapp: api. After applying it, wait until the controller has computed the status, then save it as one line to/root/ops-disruption/pdb-status.tsv—api-pdbTAB<currentHealthy>TAB<desiredHealthy>TAB<disruptionsAllowed>TAB<expectedPods>. - Pick one
apiPod and write an Eviction object to/root/ops-disruption/eviction.json—apiVersionispolicy/v1,kindisEviction,metadata.nameis that Pod's name, andmetadata.namespaceisops-pdb. With that file, runkubectl create --raw /api/v1/namespaces/ops-pdb/pods/<파드이름>/eviction -f /root/ops-disruption/eviction.json(where the placeholder is the Pod name) and save the error output to/root/ops-disruption/eviction-denied.txt(standard error must be included too). - Run
kubectl drain lab-node-0 --ignore-daemonsets --delete-emptydir-data --force --timeout=20sand save the output to/root/ops-disruption/drain-blocked.txt(including standard error). A drain first puts the node in an unschedulable state, so after the command ends, be sure to put it back withkubectl uncordon lab-node-0. - In
/root/ops-disruption/api-pdb-fixed.yaml, write the PDBapi-pdbagain under the same name, and instead ofminAvailable, usemaxUnavailable: 1, and apply it (the two fields cannot be used together, so the new file must not haveminAvailable). Wait until the allowed disruptions become 1, save it to/root/ops-disruption/pdb-after.tsvin the same five columns as step 2, and send the request from step 3 again with?dryRun=Allattached, saving the output to/root/ops-disruption/eviction-allowed.txt. - In
/root/ops-disruption/batch.yaml, write the Deploymentbatch— namespaceops-pdb, 7 replicas, labelapp: batch, imagenginx:1.27.3. And in/root/ops-disruption/batch-pdb.yaml, write the PDBbatch-pdb—minAvailable: 50%and selectorapp: batch. Apply both, and once the status is computed, save it as one line to/root/ops-disruption/rounding.tsv—batch-pdbTAB7TAB50%TAB<desiredHealthy>TAB<disruptionsAllowed>. - In
/root/ops-disruption/batch-extra.yaml, write one more PDB,batch-extra— namespaceops-pdb,maxUnavailable: 1, and the selector is, exactly the same asbatch-pdb's,app: batch. After applying it, for onebatchPod, send an eviction request with?dryRun=Allattached and save the output to/root/ops-disruption/overlap.txt(put the request body in/root/ops-disruption/overlap-eviction.json). - In
/root/ops-disruption/flaky.yaml, write the Deploymentflaky(3 replicas, labelapp: flaky, imagenginx:1.27.3), and in/root/ops-disruption/flaky-pdb.yaml, write the PDBflaky-pdb(minAvailable: 3, selectorapp: flaky), and apply them. Then, for oneflakyPod, patchstatus.conditionsto makeReadyFalse(kubectl -n ops-pdb patch pod <이름> --subresource=status --type=merge -p ..., where the placeholder is the Pod name). Send an eviction with?dryRun=Allfor that Pod and save the rejection message to/root/ops-disruption/unhealthy-denied.txt, then changeflaky-pdb'sspec.unhealthyPodEvictionPolicytoAlwaysAllow, send the same request again, and save it to/root/ops-disruption/unhealthy-allowed.txt. - In
/root/ops-disruption/orphan.yaml, write the Deploymentorphan— namespaceops-pdb, 3 replicas, labelapp: orphan, imagenginx:1.27.3, and create no PDB. Then create/root/ops-disruption/pdb-audit.sh— among the Deployments in all namespaces, print those whosespec.replicasis 2 or more but that are not covered by any PDB in the same namespace as<네임스페이스>TAB<이름>(namespace, name) to standard output only, sorted by name. If the PDB'sspec.selector.matchLabelsis a subset of the Deployment'sspec.template.metadata.labels, it counts as covered. Save that output to/root/ops-disruption/pdb-audit.txt.
Reference
- You can call the eviction subresource directly with
kubectl create --raw /api/v1/namespaces/<ns>/pods/<pod>/eviction -f <파일>(where the placeholders are the namespace, the Pod, and the file). - With
?dryRun=Allattached, you can get only the verdict without deleting the Pod. - The PDB status is all in
kubectl get pdb <이름> -o json, under.status(where the placeholder is the PDB name). - Common mistake: putting
2>&1first when saving the output, so standard error does not end up in the file. - Common mistake: not uncordoning the node after a drain is blocked, so capacity quietly stays reduced.
- Reference: https://kubernetes.io/docs/concepts/workloads/pods/disruptions/
- Reference: https://kubernetes.io/docs/tasks/run-application/configure-pdb/
A workload gathered on one node
Create the namespace ops-pdb and, in /root/ops-disruption/api.yaml, write the Deployment api — 4 replicas, label app: api, image nginx:1.27.3, and nodeSelector is kubernetes.io/hostname: lab-node-0. Apply it and wait until all 4 Pods are Running on lab-node-0.
A drain is a per-node operation, so it is reproducible only if the target node is fixed. If you gather them on one node, it becomes clear later what blocks you when you try to empty that node. A new namespace takes a moment for its default service account to appear.
Create the number 0 allowed disruptions
In /root/ops-disruption/api-pdb.yaml, write the PodDisruptionBudget api-pdb — namespace ops-pdb, minAvailable: 4, selector app: api. After applying it, wait until the controller has computed the status, then save it as one line to /root/ops-disruption/pdb-status.tsv — api-pdb TAB <currentHealthy> TAB <desiredHealthy> TAB <disruptionsAllowed> TAB <expectedPods>.
Right after you create a PDB, the status may be empty. Wait until the disruption controller counts the Pods matching the selector and fills in the status — once expectedPods is no longer 0, the calculation is done. You asked to protect all four, so predict first what the allowed disruptions will be.
Ask the eviction API directly
Pick one api Pod and write an Eviction object to /root/ops-disruption/eviction.json — apiVersion is policy/v1, kind is Eviction, metadata.name is that Pod's name, and metadata.namespace is ops-pdb. With that file, run kubectl create --raw /api/v1/namespaces/ops-pdb/pods/<파드이름>/eviction -f /root/ops-disruption/eviction.json (where the placeholder is the Pod name) and save the error output to /root/ops-disruption/eviction-denied.txt (standard error must be included too).
A drain does not delete Pods; it calls the eviction subresource. That is why when a PDB blocks it, what comes back is an HTTP 429, not a deletion — the response body says why it was blocked. To keep the output in a file, the order must not be 2>&1 > 파일 but > 파일 2>&1 (where the placeholder is the file).
See for yourself that the drain never finishes
Run kubectl drain lab-node-0 --ignore-daemonsets --delete-emptydir-data --force --timeout=20s and save the output to /root/ops-disruption/drain-blocked.txt (including standard error). A drain first puts the node in an unschedulable state, so after the command ends, be sure to put it back with kubectl uncordon lab-node-0.
A drain command does two things — it leaves an unschedulable mark on the node, and it evicts the Pods on it one by one through the eviction API. If the former succeeds and the latter is blocked, the node stays blocked. This is a common path by which capacity quietly shrinks in the field.
Fix the budget and the same request passes
In /root/ops-disruption/api-pdb-fixed.yaml, write the PDB api-pdb again under the same name, and instead of minAvailable, use maxUnavailable: 1, and apply it (the two fields cannot be used together, so the new file must not have minAvailable). Wait until the allowed disruptions become 1, save it to /root/ops-disruption/pdb-after.tsv in the same five columns as step 2, and send the request from step 3 again with ?dryRun=All attached, saving the output to /root/ops-disruption/eviction-allowed.txt.
The eviction subresource supports dryRun — it asks only whether this Pod can be evicted right now, without actually deleting it. This is a way to check before setting a maintenance window in production, and here we use it so as not to delete the Pods the later steps use. On success, a Status containing code 201 comes back.
Which way are percentages rounded
In /root/ops-disruption/batch.yaml, write the Deployment batch — namespace ops-pdb, 7 replicas, label app: batch, image nginx:1.27.3. And in /root/ops-disruption/batch-pdb.yaml, write the PDB batch-pdb — minAvailable: 50% and selector app: batch. Apply both, and once the status is computed, save it as one line to /root/ops-disruption/rounding.tsv — batch-pdb TAB 7 TAB 50% TAB <desiredHealthy> TAB <disruptionsAllowed>.
50% of 7 is 3.5. Predict first which way Kubernetes rounds, and then compare it against the number the cluster computed. If your prediction was wrong, think about why that direction is the safe side. In YAML, it is safer to wrap a percentage in quotes.
Two budgets with overlapping selectors create a deadlock
In /root/ops-disruption/batch-extra.yaml, write one more PDB, batch-extra — namespace ops-pdb, maxUnavailable: 1, and the selector is, exactly the same as batch-pdb's, app: batch. After applying it, for one batch Pod, send an eviction request with ?dryRun=All attached and save the output to /root/ops-disruption/overlap.txt (put the request body in /root/ops-disruption/overlap-eviction.json).
Even though each of the two budgets has room, the request is blocked. That is because the eviction subresource does not support the situation of a Pod falling under two budgets at all. Read the response sentence as it is — someone who has seen this message once will immediately suspect overlapping selectors from then on.
A Pod that is not ready holds the drain hostage
In /root/ops-disruption/flaky.yaml, write the Deployment flaky (3 replicas, label app: flaky, image nginx:1.27.3), and in /root/ops-disruption/flaky-pdb.yaml, write the PDB flaky-pdb (minAvailable: 3, selector app: flaky), and apply them. Then, for one flaky Pod, patch status.conditions to make Ready False (kubectl -n ops-pdb patch pod <이름> --subresource=status --type=merge -p ..., where the placeholder is the Pod name). Send an eviction with ?dryRun=All for that Pod and save the rejection message to /root/ops-disruption/unhealthy-denied.txt, then change flaky-pdb's spec.unhealthyPodEvictionPolicy to AlwaysAllow, send the same request again, and save it to /root/ops-disruption/unhealthy-allowed.txt.
The healthy Pods a PDB counts are the Pods whose Ready condition is True. You asked to protect three, and if one is not Ready, the budget is already broken, and under the default policy even that broken Pod cannot be evicted — a common reason a drain never finishes. If you change the policy, the same request passes.
Find workloads with no budget across the whole cluster
In /root/ops-disruption/orphan.yaml, write the Deployment orphan — namespace ops-pdb, 3 replicas, label app: orphan, image nginx:1.27.3, and create no PDB. Then create /root/ops-disruption/pdb-audit.sh — among the Deployments in all namespaces, print those whose spec.replicas is 2 or more but that are not covered by any PDB in the same namespace as <네임스페이스> TAB <이름> (namespace, name) to standard output only, sorted by name. If the PDB's spec.selector.matchLabels is a subset of the Deployment's spec.template.metadata.labels, it counts as covered. Save that output to /root/ops-disruption/pdb-audit.txt.
This list is the first material you need when designing a maintenance window — a workload with no budget goes down all at once with no resistance in a drain. For the subset check, you can use jq to count whether each entry of the PDB selector is present in the Pod labels with the same value. If the script writes a file itself, it overwrites the student's artifact when the grader runs it again, so send it out to standard output only.