I Break It — A Chaos Lab Where the Hypothesis Comes First
CPU makes it slow, memory kills it
One-line summary
CPU limits and memory limits only sound alike; they bite in completely different ways. One slows you down, and the other kills you.
Why this is needed
Writing limits is usually a habit. You copy a value the previous person wrote, or you hear
"specify limits" in a review and put in a reasonable-looking number. After that, two kinds
of incidents arrive months later. One is "we didn't change anything and it suddenly got slow,"
and the other is "restarts keep happening and there are no clues in the logs."
The two incidents have different causes, yet both came from a single line of limits.
How it works
The documentation on managing resources for Pods and containers clearly distinguishes that the two limits are enforced in different ways.
A cpu limit is enforced by throttling. The kernel lets the container use only its set share of CPU, and once that share is used up, it holds that container until the next period arrives. Because it is a hard limit enforced by the kernel, the container cannot use more than the limit. But it does not die. This is misunderstood so often that the documentation states separately that "the runtime does not terminate a Pod or container for using a lot of CPU." The only symptom is slowness, and since the process and the probes are all alive, nothing is left in the logs.
A memory limit is enforced by an OOM kill. If a container goes over the limit, the kernel may
terminate it. However, the documentation says this is reactive — it happens when the kernel detects memory
pressure, so a container that has gone over the limit may not die immediately. When it does die,
exit code 137 and the reason OOMKilled are left behind.
집행 방식 증상 남는 흔적
cpu 스로틀링 지연만 늘어난다 없음(로그가 조용하다)
memory OOM kill 재시작·CrashLoop exit 137 · OOMKilled
There is one more lesser-known trap. An emptyDir volume backed by memory is counted
toward the container's memory usage. The same documentation says, "the kubelet tracks tmpfs
emptyDir volumes as container memory use, rather than as local ephemeral storage."
If you attach a memory volume to write temporary files quickly, code that thought it was writing files to disk
pushes up the memory limit and kills the container.
And the symptom is not "died while writing a file" but simply OOMKilled.
You should also remember the roles of requests and limits separately. requests is the value the scheduler uses to decide
where to place the Pod, and limits is the value the kubelet and kernel use to decide how much to
let it use. So if you write only requests, the container may use more than that when the node is idle,
and if you write only limits, Kubernetes copies the same value into requests.
One more thing. When lowering a limit, you must also look at its relationship to the request. A limit cannot be smaller than the request, so even if you try to lower only the memory limit to 40Mi, if the request is written as 64Mi, the API server rejects that change outright. When preparing an experiment, you often run into cases where "I expected a failure but the command is rejected," and that is usually because you tripped over this kind of consistency rule. The app in this lab has its requests set generously low so that you do not hit this wall during the experiment.
What you see in the field
The most common accident is a deployment where the cpu limit was lowered to 100m "to be safe." Nothing happens normally, but the moment traffic rises a little, response time jumps severalfold. The success rate stays the same, so no alert fires, and the CPU utilization graph sits flat right at the limit, so it actually "looks like there is headroom." Incidents like this have their causes stay hidden until you measure throttling directly.
There are accidents in the opposite direction too. If you give a generous memory limit, OOM disappears, but unless you raise the request along with it, that Pod is evicted first when the node gets tight. The documentation says that a Pod with a container that exceeds its request is likely to be evicted when the node runs short of memory. In other words, the limit sets the ceiling for an individual container, and the request sets the order when resources run short. Even when confirming by experiment, you must touch the two separately to be able to say what the cause was.
So when dealing with resource incidents, you first split into branches by symptom. If it is latency with quiet logs, look at cpu first; if it is restarts and exit code 137, look at memory first. Even just splitting these branches correctly greatly shortens the investigation. Conversely, a response like "it got slow, so let's increase memory" changes nothing and only raises cost. A change that touches limits triggers a rollout and creates new Pods, so just after the change there can even be an illusion that things look better for a moment. To avoid being fooled by that illusion, you need numbers measured the same way before and after the change.
What to check in the next quiz
Check the difference in how cpu limits and memory limits are enforced, what it means that OOM enforcement is reactive, and what a memory-backed emptyDir is counted as.