cgroup v2, and When Limits Look Like They Are Not Being Honoured
In one line
Namespaces handle what can be seen, and cgroups handle how much can be used.
And the fact that most often fools people in this course is that while a cgroup imposes limits,
it does not hide /proc.
Why this matters
cgroups started at Google in 2006 and were merged into kernel 2.6.24 in 2007. The original name was "process container", but it was renamed control group, or cgroup for short, because it was confused with the already widely used word container. The history of the name is itself an explanation of the role — it is not isolation but resource control.
If you put thirty containers on one host and one of them eats all the memory, the other twenty-nine die with it. Namespaces do not solve this problem at all. They only keep the containers from seeing each other, and they still share the same physical resources.
How it works
First, check the version. If stat -fc %T /sys/fs/cgroup prints cgroup2fs, it is v2, and
if it prints tmpfs, it is v1. The main files in v2 are as follows.
| File | Meaning |
|---|---|
cpu.max |
Two values, quota period. 50000 100000 means 0.5 CPU |
cpu.weight |
A relative weight from 1 to 10000, default 100 |
cpu.stat |
nr_periods, nr_throttled, throttled_usec |
memory.min/low/high/max |
Protection lines and upper limits |
memory.current |
Current usage |
memory.events |
Number of times the limit was hit |
pids.max |
Upper limit on the number of processes |
In v1 these were cpu.cfs_quota_us and memory.limit_in_bytes, respectively.
It is important to tell memory.high and memory.max apart. Exceeding high applies
reclaim pressure and slows allocation but does not kill anything. Exceeding max triggers an OOM
kill. So high means "start hurting from here" and max means "this is the wall".
CPU is not killed but put to sleep. If nr_throttled / nr_periods in cpu.stat is
over 5%, treat it as a problem. If 350 of 1000 periods were throttled, that is 35%, and then
the typical picture appears in which the CPU utilization graph looks idle while only the response time spikes.
It is convenient to memorize commonly used byte values. 64 MiB = 67108864, 192 MiB = 201326592, 256 MiB = 268435456, 512 MiB = 536870912.
Exit code 137 = 128 + 9 = SIGKILL. An OOM kill gives the app no exception whatsoever.
You cannot catch it with try/except and no log is left. You can only confirm it from the kernel log's
Memory cgroup out of memory: Killed process ..., and on the runtime side you must tell it apart with
State.OOMKilled.
What it looks like in the field
This is the most important trap. In a container started with --memory=256m --cpus=0.5,
free -m shows 64228 MB and nproc shows 32. This is because a cgroup imposes limits but
does not hide /proc. The JVM solved this problem with UseContainerSupport by reading the cgroup files
directly, but Node's os.cpus() and Go's runtime.NumCPU()
still see the host values. That is why you must set GOMAXPROCS explicitly for Go services.
If you start workers with 32 threads and assign 0.5 cores, only throttling piles up.
Kubernetes QoS classes come from here too. Guaranteed has an OOM score of -997, Burstable has 2 to 999, and BestEffort has 1000. That BestEffort dies first is not a policy statement but a direct result of cgroups and OOM scores.
One more thing. In a rootless environment, even if you set a limit, it may not actually be enforced.
Without systemd delegation (Delegate=cpu memory pids io), the request is
recorded but not applied to the kernel. That is why the next lab checks both the HostConfig
request values from docker inspect and the values the container actually sees. If the two
differ, that is itself something to learn.
Where v1 and v2 look different
cgroup v1 and v2 coexist, and their file names and meanings differ. If you do not check which one it is first, you end up looking for files that do not exist.
stat -fc %T /sys/fs/cgroup # cgroup2fs 면 v2, tmpfs 면 v1
| What | v1 | v2 |
|---|---|---|
| Memory limit | memory/memory.limit_in_bytes |
memory.max |
| Current usage | memory.usage_in_bytes |
memory.current |
| CPU limit | cpu.cfs_quota_us / cpu.cfs_period_us |
cpu.max (two values on one line) |
| Throttle statistics | cpu.stat |
cpu.stat |
| OOM record | (none) | memory.events |
v2's memory.events is especially useful. v1 did not have it, so "has this container ever
been killed by an OOM?" could be known only from the kernel log.
When there is no limit, it shows up as max. It is the string max, not a number, so the parsing
code throws an exception. You often run into this when building tools.
The memory limit also counts page cache. A container that reads many files accumulates cache and
hits the limit, but the kernel reclaims that cache, so it usually does not die. So
memory.current sticking to the limit is not a problem in itself. It is a real problem when oom_kill
in memory.events increases.
A CPU limit does not show up in utilization. It uses up its quota every 100 ms and waits for the rest,
so the average utilization is low while only responses get slow. nr_throttled and
throttled_usec in cpu.stat are the evidence.
A small limit is actually harmful. If you set the CPU limit to 0.1 core, it can run only 10 ms out of 100 ms, so even a short job finishes across several periods. Latency grows noticeably, so keeping the request low and the limit generous or absent is better for response time.
What you will do in the next lab
You check the cgroup version and the controller list, apply memory, CPU, and PID limits one at a time, and then compare the requested values with the observed values. You create a 137 yourself with SIGKILL and sort out how to tell it apart from an OOM kill.