TT Lab
Get started
Learn Learning paths Courses

I Break It — A Chaos Lab Where the Hypothesis Comes First

CPU makes it slow, memory kills it

Continue in TT Lab

One-line summary

CPU limits and memory limits only sound alike; they bite in completely different ways. One slows you down, and the other kills you.

Why this is needed

Writing limits is usually a habit. You copy a value the previous person wrote, or you hear "specify limits" in a review and put in a reasonable-looking number. After that, two kinds of incidents arrive months later. One is "we didn't change anything and it suddenly got slow," and the other is "restarts keep happening and there are no clues in the logs." The two incidents have different causes, yet both came from a single line of limits.

How it works

The documentation on managing resources for Pods and containers clearly distinguishes that the two limits are enforced in different ways.

A cpu limit is enforced by throttling. The kernel lets the container use only its set share of CPU, and once that share is used up, it holds that container until the next period arrives. Because it is a hard limit enforced by the kernel, the container cannot use more than the limit. But it does not die. This is misunderstood so often that the documentation states separately that "the runtime does not terminate a Pod or container for using a lot of CPU." The only symptom is slowness, and since the process and the probes are all alive, nothing is left in the logs.

A memory limit is enforced by an OOM kill. If a container goes over the limit, the kernel may terminate it. However, the documentation says this is reactive — it happens when the kernel detects memory pressure, so a container that has gone over the limit may not die immediately. When it does die, exit code 137 and the reason OOMKilled are left behind.

             집행 방식        증상                  남는 흔적
  cpu        스로틀링        지연만 늘어난다        없음(로그가 조용하다)
  memory     OOM kill        재시작·CrashLoop       exit 137 · OOMKilled

There is one more lesser-known trap. An emptyDir volume backed by memory is counted toward the container's memory usage. The same documentation says, "the kubelet tracks tmpfs emptyDir volumes as container memory use, rather than as local ephemeral storage." If you attach a memory volume to write temporary files quickly, code that thought it was writing files to disk pushes up the memory limit and kills the container. And the symptom is not "died while writing a file" but simply OOMKilled.

You should also remember the roles of requests and limits separately. requests is the value the scheduler uses to decide where to place the Pod, and limits is the value the kubelet and kernel use to decide how much to let it use. So if you write only requests, the container may use more than that when the node is idle, and if you write only limits, Kubernetes copies the same value into requests.

One more thing. When lowering a limit, you must also look at its relationship to the request. A limit cannot be smaller than the request, so even if you try to lower only the memory limit to 40Mi, if the request is written as 64Mi, the API server rejects that change outright. When preparing an experiment, you often run into cases where "I expected a failure but the command is rejected," and that is usually because you tripped over this kind of consistency rule. The app in this lab has its requests set generously low so that you do not hit this wall during the experiment.

What you see in the field

The most common accident is a deployment where the cpu limit was lowered to 100m "to be safe." Nothing happens normally, but the moment traffic rises a little, response time jumps severalfold. The success rate stays the same, so no alert fires, and the CPU utilization graph sits flat right at the limit, so it actually "looks like there is headroom." Incidents like this have their causes stay hidden until you measure throttling directly.

There are accidents in the opposite direction too. If you give a generous memory limit, OOM disappears, but unless you raise the request along with it, that Pod is evicted first when the node gets tight. The documentation says that a Pod with a container that exceeds its request is likely to be evicted when the node runs short of memory. In other words, the limit sets the ceiling for an individual container, and the request sets the order when resources run short. Even when confirming by experiment, you must touch the two separately to be able to say what the cause was.

So when dealing with resource incidents, you first split into branches by symptom. If it is latency with quiet logs, look at cpu first; if it is restarts and exit code 137, look at memory first. Even just splitting these branches correctly greatly shortens the investigation. Conversely, a response like "it got slow, so let's increase memory" changes nothing and only raises cost. A change that touches limits triggers a rollout and creates new Pods, so just after the change there can even be an illusion that things look better for a moment. To avoid being fooled by that illusion, you need numbers measured the same way before and after the change.

What to check in the next quiz

Check the difference in how cpu limits and memory limits are enforced, what it means that OOM enforcement is reactive, and what a memory-backed emptyDir is counted as.