I Break It — A Chaos Lab Where the Hypothesis Comes First
Kill one — how many people hurt?
One-line summary
When you kill one Pod, how many people are hurt depends more on the shutdown procedure than on the number of replicas.
Why this is needed
The sentence "we added replicas, so there is no downtime" is only half right. Services are common where, even with three replicas, a few requests fail every time you deploy. It is always only a few, so nobody reports it, and so it has stayed the same for years. To find out where these leaking requests come from, you need to look at the process of a Pod dying, one step at a time.
How it works
When a Pod deletion request comes in, two things start at the same time. One is the control plane removing that Pod from the Service's EndpointSlice, and the other is the kubelet stopping the container. The Pod lifecycle documentation states clearly that these two run in parallel, not in sequence. The ready state of a terminating endpoint always becomes false, so it is dropped from load balancing, but it takes time for that fact to propagate to the node's iptables rules.
This is where the problem arises. During the short window before the rules propagate, new connections keep coming in to the dying Pod, whose process has already received TERM and closed its sockets. The request fails with a connection refused. The application is fine and Kubernetes is fine, but only the user fails.
The key to the solution is making the process hold on a little longer.
The preStop of a container lifecycle hook
runs before the TERM signal is sent to the container, and the signal goes out only after the hook finishes. If the hook holds
things for just a few seconds, endpoint removal propagates during that time, and new connections go only to Pods that are alive. The key point is that the application keeps
serving while the hook runs.
잘못된 기대 실제 순서
1. 엔드포인트에서 뺀다 엔드포인트 제거 ─┐ 동시에 시작
2. 그다음 TERM 을 보낸다 TERM 전송 ──────┘
preStop 이 있으면 TERM 만 뒤로 밀린다
The time budget is set by terminationGracePeriodSeconds. The default is 30 seconds,
and if the container is still alive after that time, the runtime ends it with KILL. And
the grace period countdown has already started before the preStop hook begins. If you plan to spend
5 seconds in the hook, the grace period must be more generous than that.
There is one more distinction to keep in mind.
The disruption documentation
divides disruptions into voluntary and involuntary ones. A PodDisruptionBudget intervenes
only in voluntary disruptions (draining a node, upgrading a cluster). When we manually run
kubectl delete pod or a process dies, the PDB cannot stop it.
That is why "we set a PDB, so we're safe" is dangerous.
If the client reuses connections, the story gets one layer more complicated. The endpoint list only decides where new connections go; it does not move connections that are already open. A client holding a connection open with keep-alive keeps talking to a Pod that has been removed from the endpoints, and fails the moment that Pod dies. So to do shutdown handling properly, in addition to buying time with preStop, you need code in the application that, on receiving TERM, cleans up open connections as it goes down. The load generator in this lab opens a new connection for every request in order to leave this variable out on purpose.
What you see in the field
Draining a node brings all three of these out at once. Without a PDB, all Pods of the same service are removed at the same time, causing a complete outage, and even with a PDB, if there is no shutdown handling, failures leak out each time a Pod is removed. The number of replicas decides how badly it hurts, and the shutdown procedure decides whether it hurts at all.
Read deployment strategies along the same axis. A rolling update is a device that adjusts the blast radius through how many Pods may be
down at once (maxUnavailable) and how many extra may be brought up temporarily (maxSurge).
By contrast, Recreate takes down all the old Pods before bringing up the new ones, so an outage necessarily occurs. Using Recreate in an experiment environment
is a control condition to leave only one cause, not a production recommendation. If old Pods
remain, the effect of the new settings is masked by the old Pods' responses and the numbers get smeared.
What to check in the next quiz
Check the order of endpoint removal and TERM delivery, what the preStop hook postpones, and which disruptions a PodDisruptionBudget applies to.