Nobody Ran Drain, Yet the Pods Are Gone
Goal
You confirm through actual eviction what NoSchedule and NoExecute each do to Pods that are already running, compute for each Pod the eviction time that tolerationSeconds sets and put it in a table, and then decide the value to give a stateful workload, with your reasoning.
Why it matters
Half of the incidents that call operators out are cases where Pods moved though nobody typed a command. When a node stops its heartbeat, the node controller attaches the node.kubernetes.io/not-ready or node.kubernetes.io/unreachable taint as NoExecute, and from that moment each Pod's tolerationSeconds starts the clock. The problem is that most teams have never written this value. If you do not write it, admission puts in 300 seconds — that is, every workload uses the same policy of moving after 5 minutes. It is too long for a stateless frontend and too short for a data workload with a large local state. Understanding this value and choosing it per workload is half of responding to node incidents.
Steps
- Create the namespace
ops-evict, and in/root/ops-eviction/web.yaml, write the Deploymentweb— 3 replicas, labelapp: web, imagenginx:1.27.3, and withnodeSelectormake it sit only onkubernetes.io/hostname: lab-node-0. Apply it and wait until all 3 Pods are Running. - You did not write a single line of toleration on the
webPods. Yet the actual objects have them. Pick any onewebPod and check it with-o json, and save those tolerations to/root/ops-eviction/default-tolerations.tsv, one line each of<키>TAB<효과>TAB<tolerationSeconds>(key, effect, tolerationSeconds), in ascending key name order. Do not put in any other lines. - In
/root/ops-eviction/brittle.yamland/root/ops-eviction/patient.yaml, write two Pods. Both use the namespaceops-evict, imagenginx:1.27.3, andnodeSelectorkubernetes.io/hostname: lab-node-1. The two Pods differ only in the effect of the toleration —brittlehas the labelrole: brittleand one toleration that tolerates the keymaintwithoperator: Existsandeffect: NoSchedule(notolerationSeconds), andpatienthas the labelrole: patientand one toleration that tolerates the keymaintwithoperator: Exists,effect: NoExecute, andtolerationSeconds: 3600. Apply both and get them Running on lab-node-1. - On
lab-node-1, put the taintmaint=planned:NoSchedule. Then, in/root/ops-eviction/noschedule.tsv, save the current state ofbrittleandpatient, one line each of<파드이름>TAB<노드이름>TAB<phase>(Pod name, node name, phase), in ascending Pod name order. - On the same node
lab-node-1, put one more taint,maint=planned:NoExecute. Wait forbrittleto disappear not with a fixed sleep but with a loop that waits for a condition, then save the result to/root/ops-eviction/noexecute.tsvas two lines —brittleTABgoneandpatientTABRunning. - Write
/root/ops-eviction/shortwait.yamland/root/ops-eviction/longwait.yaml. Both use the namespaceops-evict, imagenginx:1.27.3, andnodeSelectorkubernetes.io/hostname: lab-node-2. Both Pods tolerate the keylinkdownwithoperator: Existsandeffect: NoExecute, andshortwaitgetstolerationSeconds: 20andlongwaitgetstolerationSeconds: 3600. The labels arerole: shortwaitandrole: longwaitrespectively. Get both Running on lab-node-2. - You simulate a broken link — on
lab-node-2, put the taintlinkdown=yes:NoExecute. Wait with a condition loop untilshortwaitdisappears, then save it to/root/ops-eviction/partition.tsvas two lines —shortwaitTABgoneandlongwaitTABRunning. - First, in
/root/ops-eviction/forever.yaml, create the Podforever— namespaceops-evict, labelrole: forever, imagenginx:1.27.3,nodeSelectorkubernetes.io/hostname: lab-node-0, and the toleration is just one, with no key andoperator: Existsandeffect: NoExecute(notolerationSeconds). Then create/root/ops-eviction/evict-plan.sh— read all the Pods inops-evictand, for each Pod, print how many seconds it tolerateslinkdown'sNoExecutetaint as<파드이름>TAB<값>(Pod name, value) to standard output only. The value is0if there is no toleration that tolerates it; if there is one buttolerationSecondsis absent,never; otherwise that number of seconds. The output must be in ascending Pod name order. Save that output to/root/ops-eviction/evict-plan.tsv. - Create the namespace
ops-ledgerand in/root/ops-eviction/ledger.yamlwrite the Deploymentledger— namespaceops-ledger, 2 replicas, labelapp: ledger, imagenginx:1.27.3. This workload has a large node-local state, so it is better to ride out short disconnections. Tolerate the two keysnode.kubernetes.io/not-readyandnode.kubernetes.io/unreachablewithoperator: Existsandeffect: NoExecute, but set thetolerationSecondsof both tolerations to the same value, between 900 and 3600 inclusive. After applying so that 2 Pods are Running, write two lines to/root/ops-eviction/decision.tsv—tolerationSecondsTAB<고른 값>(chosen value) andreasonTAB<40자 이상의 근거 한 문장>(one sentence of rationale, 40 characters or more).
Reference
- NoSchedule blocks only new placements, and NoExecute also evicts Pods that are already running.
tolerationSecondshas meaning only on a NoExecute toleration.- You put on a taint with an empty value using
kubectl taint node <노드> <키>=:NoExecute(where the placeholders are the node and the key). - When waiting, use a condition loop instead of a fixed sleep — times differ from machine to machine.
- Common mistake: putting on NoSchedule and waiting for the Pod to disappear.
- Common mistake: writing a toleration yourself and thinking the admission default of 300 seconds is attached as well.
- Reference: https://kubernetes.io/docs/concepts/scheduling-eviction/taint-and-toleration/
Set up a control group on a node you will not touch
Create the namespace ops-evict, and in /root/ops-eviction/web.yaml, write the Deployment web — 3 replicas, label app: web, image nginx:1.27.3, and with nodeSelector make it sit only on kubernetes.io/hostname: lab-node-0. Apply it and wait until all 3 Pods are Running.
In this lab you deliberately break lab-node-1 and lab-node-2 later. Without a control group, you cannot tell whether a Pod disappeared because of a taint or because of something else. A new namespace takes a moment for its default service account to appear, so if it fails, apply again a few seconds later.
A toleration nobody wrote is already attached
You did not write a single line of toleration on the web Pods. Yet the actual objects have them. Pick any one web Pod and check it with -o json, and save those tolerations to /root/ops-eviction/default-tolerations.tsv, one line each of <키> TAB <효과> TAB <tolerationSeconds> (key, effect, tolerationSeconds), in ascending key name order. Do not put in any other lines.
The admission controller attaches them. If you do not write these two keys yourself when creating a Pod, they are filled in automatically, and if you write them yourself, they are not filled in. The answer is the same whichever web Pod you pick. The columns are separated by tabs, so jq's @tsv is handy.
Two Pods on the same node that differ only in how long they tolerate
In /root/ops-eviction/brittle.yaml and /root/ops-eviction/patient.yaml, write two Pods. Both use the namespace ops-evict, image nginx:1.27.3, and nodeSelector kubernetes.io/hostname: lab-node-1. The two Pods differ only in the effect of the toleration — brittle has the label role: brittle and one toleration that tolerates the key maint with operator: Exists and effect: NoSchedule (no tolerationSeconds), and patient has the label role: patient and one toleration that tolerates the key maint with operator: Exists, effect: NoExecute, and tolerationSeconds: 3600. Apply both and get them Running on lab-node-1.
A toleration is a declaration that 'even if this taint is attached, I will hold out.' If the operator is Exists, the value is not looked at, but the effect must match exactly — even with the same key, a different effect does not tolerate it. The difference between the two Pods is exactly that, and in the later steps you see what that difference produces.
NoSchedule does not touch Pods that are already running
On lab-node-1, put the taint maint=planned:NoSchedule. Then, in /root/ops-eviction/noschedule.tsv, save the current state of brittle and patient, one line each of <파드이름> TAB <노드이름> TAB <phase> (Pod name, node name, phase), in ascending Pod name order.
There is a reason the name means 'do not schedule.' This effect blocks only new placements. It is normal for patient to stay as it is even though it does not tolerate this effect — confirming that nothing happens is the purpose of this step.
Put on NoExecute and it really disappears
On the same node lab-node-1, put one more taint, maint=planned:NoExecute. Wait for brittle to disappear not with a fixed sleep but with a loop that waits for a condition, then save the result to /root/ops-eviction/noexecute.tsv as two lines — brittle TAB gone and patient TAB Running.
NoExecute also has an effect on Pods that are already running. brittle also has a maint toleration, but its effect is NoSchedule, so it does not tolerate this taint. The waiting loop looks like this — for i in $(seq 1 60); do kubectl -n ops-evict get pod brittle >/dev/null 2>&1 || break; sleep 2; done
Two Pods prepared for when the network is cut
Write /root/ops-eviction/shortwait.yaml and /root/ops-eviction/longwait.yaml. Both use the namespace ops-evict, image nginx:1.27.3, and nodeSelector kubernetes.io/hostname: lab-node-2. Both Pods tolerate the key linkdown with operator: Exists and effect: NoExecute, and shortwait gets tolerationSeconds: 20 and longwait gets tolerationSeconds: 3600. The labels are role: shortwait and role: longwait respectively. Get both Running on lab-node-2.
linkdown is a custom key we chose. There is a reason we do not use the built-in keys (not-ready and unreachable) — the taints with those two keys are managed directly by the node controller based on the node conditions, so if you attach one by hand to a Ready node, it is swept away immediately (I confirmed this directly). The eviction rules themselves are the same regardless of the key.
Make the node look like it stopped responding and measure the time
You simulate a broken link — on lab-node-2, put the taint linkdown=yes:NoExecute. Wait with a condition loop until shortwait disappears, then save it to /root/ops-eviction/partition.tsv as two lines — shortwait TAB gone and longwait TAB Running.
On a real cluster, after the node stops its heartbeat, the node controller attaches node.kubernetes.io/unreachable and from then on the same rule runs. The nodes here are fakes made by kwok, so you cannot cut the heartbeat, and if you attach a built-in key by hand the node controller sweeps it away immediately, so we do the same thing with a custom key. It disappears only after 20 seconds pass, so wait with a condition loop.
Have a machine compute the eviction timetable
First, in /root/ops-eviction/forever.yaml, create the Pod forever — namespace ops-evict, label role: forever, image nginx:1.27.3, nodeSelector kubernetes.io/hostname: lab-node-0, and the toleration is just one, with no key and operator: Exists and effect: NoExecute (no tolerationSeconds). Then create /root/ops-eviction/evict-plan.sh — read all the Pods in ops-evict and, for each Pod, print how many seconds it tolerates linkdown's NoExecute taint as <파드이름> TAB <값> (Pod name, value) to standard output only. The value is 0 if there is no toleration that tolerates it; if there is one but tolerationSeconds is absent, never; otherwise that number of seconds. The output must be in ascending Pod name order. Save that output to /root/ops-eviction/evict-plan.tsv.
A toleration with an empty key means it tolerates all keys, and an empty effect means it tolerates all effects — if you miss those two cases, forever will not come out as never. If the script writes a file itself, it overwrites the student's artifact when the grader runs it again, so send it out to standard output only.
Decide the value for a stateful workload and leave the rationale
Create the namespace ops-ledger and in /root/ops-eviction/ledger.yaml write the Deployment ledger — namespace ops-ledger, 2 replicas, label app: ledger, image nginx:1.27.3. This workload has a large node-local state, so it is better to ride out short disconnections. Tolerate the two keys node.kubernetes.io/not-ready and node.kubernetes.io/unreachable with operator: Exists and effect: NoExecute, but set the tolerationSeconds of both tolerations to the same value, between 900 and 3600 inclusive. After applying so that 2 Pods are Running, write two lines to /root/ops-eviction/decision.tsv — tolerationSeconds TAB <고른 값> (chosen value) and reason TAB <40자 이상의 근거 한 문장> (one sentence of rationale, 40 characters or more).
Increasing the value reduces pointless moving around on a node that was cut off briefly, but on a node that really died, the service stays empty for that much longer too. The rationale must include what you gain and what you lose. It passes only if the number you wrote in the file and the number actually deployed to the cluster are the same.