TT Lab
Get started
Learn Learning paths Courses

Container Internals

cgroup v2, and When Limits Look Like They Are Not Being Honoured

Continue in TT Lab

In one line

Namespaces handle what can be seen, and cgroups handle how much can be used. And the fact that most often fools people in this course is that while a cgroup imposes limits, it does not hide /proc.

Why this matters

cgroups started at Google in 2006 and were merged into kernel 2.6.24 in 2007. The original name was "process container", but it was renamed control group, or cgroup for short, because it was confused with the already widely used word container. The history of the name is itself an explanation of the role — it is not isolation but resource control.

If you put thirty containers on one host and one of them eats all the memory, the other twenty-nine die with it. Namespaces do not solve this problem at all. They only keep the containers from seeing each other, and they still share the same physical resources.

How it works

First, check the version. If stat -fc %T /sys/fs/cgroup prints cgroup2fs, it is v2, and if it prints tmpfs, it is v1. The main files in v2 are as follows.

File Meaning
cpu.max Two values, quota period. 50000 100000 means 0.5 CPU
cpu.weight A relative weight from 1 to 10000, default 100
cpu.stat nr_periods, nr_throttled, throttled_usec
memory.min/low/high/max Protection lines and upper limits
memory.current Current usage
memory.events Number of times the limit was hit
pids.max Upper limit on the number of processes

In v1 these were cpu.cfs_quota_us and memory.limit_in_bytes, respectively.

It is important to tell memory.high and memory.max apart. Exceeding high applies reclaim pressure and slows allocation but does not kill anything. Exceeding max triggers an OOM kill. So high means "start hurting from here" and max means "this is the wall".

CPU is not killed but put to sleep. If nr_throttled / nr_periods in cpu.stat is over 5%, treat it as a problem. If 350 of 1000 periods were throttled, that is 35%, and then the typical picture appears in which the CPU utilization graph looks idle while only the response time spikes.

It is convenient to memorize commonly used byte values. 64 MiB = 67108864, 192 MiB = 201326592, 256 MiB = 268435456, 512 MiB = 536870912.

Exit code 137 = 128 + 9 = SIGKILL. An OOM kill gives the app no exception whatsoever. You cannot catch it with try/except and no log is left. You can only confirm it from the kernel log's Memory cgroup out of memory: Killed process ..., and on the runtime side you must tell it apart with State.OOMKilled.

What it looks like in the field

This is the most important trap. In a container started with --memory=256m --cpus=0.5, free -m shows 64228 MB and nproc shows 32. This is because a cgroup imposes limits but does not hide /proc. The JVM solved this problem with UseContainerSupport by reading the cgroup files directly, but Node's os.cpus() and Go's runtime.NumCPU() still see the host values. That is why you must set GOMAXPROCS explicitly for Go services. If you start workers with 32 threads and assign 0.5 cores, only throttling piles up.

Kubernetes QoS classes come from here too. Guaranteed has an OOM score of -997, Burstable has 2 to 999, and BestEffort has 1000. That BestEffort dies first is not a policy statement but a direct result of cgroups and OOM scores.

One more thing. In a rootless environment, even if you set a limit, it may not actually be enforced. Without systemd delegation (Delegate=cpu memory pids io), the request is recorded but not applied to the kernel. That is why the next lab checks both the HostConfig request values from docker inspect and the values the container actually sees. If the two differ, that is itself something to learn.

Where v1 and v2 look different

cgroup v1 and v2 coexist, and their file names and meanings differ. If you do not check which one it is first, you end up looking for files that do not exist.

stat -fc %T /sys/fs/cgroup      # cgroup2fs 면 v2, tmpfs 면 v1
What v1 v2
Memory limit memory/memory.limit_in_bytes memory.max
Current usage memory.usage_in_bytes memory.current
CPU limit cpu.cfs_quota_us / cpu.cfs_period_us cpu.max (two values on one line)
Throttle statistics cpu.stat cpu.stat
OOM record (none) memory.events

v2's memory.events is especially useful. v1 did not have it, so "has this container ever been killed by an OOM?" could be known only from the kernel log.

When there is no limit, it shows up as max. It is the string max, not a number, so the parsing code throws an exception. You often run into this when building tools.

The memory limit also counts page cache. A container that reads many files accumulates cache and hits the limit, but the kernel reclaims that cache, so it usually does not die. So memory.current sticking to the limit is not a problem in itself. It is a real problem when oom_kill in memory.events increases.

A CPU limit does not show up in utilization. It uses up its quota every 100 ms and waits for the rest, so the average utilization is low while only responses get slow. nr_throttled and throttled_usec in cpu.stat are the evidence.

A small limit is actually harmful. If you set the CPU limit to 0.1 core, it can run only 10 ms out of 100 ms, so even a short job finishes across several periods. Latency grows noticeably, so keeping the request low and the limit generous or absent is better for response time.

What you will do in the next lab

You check the cgroup version and the controller list, apply memory, CPU, and PID limits one at a time, and then compare the requested values with the observed values. You create a 137 yourself with SIGKILL and sort out how to tell it apart from an OOM kill.