TT Lab
Get started
Learn Learning paths Courses

Container Internals

Four Files That Prove a Container

Continue in TT Lab

In one line

There is no "container" data structure in the kernel. Instead, for each process, the namespaces, cgroup, mounts, and privileges are recorded under /proc/<pid>/, and reading those four files reveals everything about what isolation that process is inside.

Why this matters

To answer the question "is this process inside a container?" with a Docker command, you need Docker. But the situation you face in incident investigations is usually the opposite — you ran ps on the host and an unfamiliar process is eating CPU, and you need to know which container it belongs to and what is restricted.

/proc answers the same way whether you use Docker, containerd, or podman. It is because you are asking the kernel directly, not the runtime.

File 1 — /proc/PID/ns/: what can it see

$ ls -l /proc/1/ns/
mnt -> 'mnt:[4026532187]'
net -> 'net:[4026532190]'
pid -> 'pid:[4026532188]'

The number in the brackets is the inode number, and if it is the same, it is the same namespace. If it differs when compared with the host's /proc/1/ns/net, the network is separated, and if it is the same, the container was started with --network host. Comparing two containers also tells you which namespaces they share.

File 2 — /proc/PID/cgroup: how much can it use

0::/kubepods.slice/kubepods-burstable.slice/.../cri-containerd-9f3c1a....scope

In cgroup v2 it is a single line. This path contains the container ID. It is the fastest way to identify an unfamiliar process on the host. And if you follow this path under /sys/fs/cgroup/, you can read the actual limit values.

/sys/fs/cgroup/<경로>/memory.max      # 메모리 상한
/sys/fs/cgroup/<경로>/memory.current  # 현재 사용량
/sys/fs/cgroup/<경로>/cpu.max         # "200000 100000" = 2 코어

File 3 — /proc/PID/mountinfo: where do the files come from

A container's root filesystem is usually overlay.

... / overlay rw,lowerdir=/var/lib/.../l/ABC:/var/lib/.../l/DEF,upperdir=...

The colon-separated list in lowerdir is the image layers, and upperdir is the container's writable layer. Volume mounts show up here too — you use it to check which host path is attached where.

File 4 — /proc/PID/status: what can it do

CapEff:	00000000a80425fb
NoNewPrivs:	1
Seccomp:	2

CapEff is the effective capability bitmask. 0 means it has no privileges at all, and 0000003fffffffff means it has virtually everything — a trace of --privileged. NoNewPrivs: 1 is the result of allowPrivilegeEscalation: false, and Seccomp: 2 means a seccomp filter is applied.

You can decode it into human-readable form with capsh --decode=00000000a80425fb.

Summary — the questions the four files answer

File The question it answers
ns/ What it can see (PID, network, and mount isolation)
cgroup What it can use and how much + which container it is
mountinfo Where the files come from (layers, volumes)
status What it can do (capabilities, seccomp)

Real questions you can answer with the same files

Once you learn to read /proc, the number of questions you can answer grows even in environments without tools. Here are four commonly used ones.

"Do these two containers share a network?" — Compare the inode numbers of the namespaces. If they are the same, it is the same namespace.

readlink /proc/$A/ns/net /proc/$B/ns/net
# net:[4026532567] 이 같으면 같은 네트워크

The problem of a sidecar not seeing the main container's network, and the problem of an ephemeral container not seeing the processes, can all be confirmed with this one line.

"Is the memory limit really applied?" — In cgroup v2, you follow the path and read the files.

cat /proc/$PID/cgroup                       # 0::/kubepods/.../<id>
cat /sys/fs/cgroup/kubepods/.../memory.max  # 한도
cat /sys/fs/cgroup/kubepods/.../memory.current
cat /sys/fs/cgroup/kubepods/.../memory.events  # oom_kill 횟수가 여기 있다

If oom_kill in memory.events is not 0, that container has already died once. This is the first place to look when finding out why a Pod restarted.

"Where did this file come from?" — mountinfo shows where a volume is actually attached and whether it is read-only. The problem of a ConfigMap not being updated (because of the structure where it is mounted through symbolic links) shows up here.

grep -E ' /etc/config | /data ' /proc/$PID/mountinfo

"What can this process do?" — Decode CapEff in status into something human-readable.

grep CapEff /proc/$PID/status
capsh --decode=00000000a80425fb

If CapEff is 0000000000000000, all privileges have been dropped, and that is the state we want. Conversely, if you see 0000003fffffffff, that container is effectively running in privileged mode.

The opposite direction, finding the container from the host, is also used often. When you find a process using a lot of CPU on the host, which Pod it belongs to is written in the cgroup path.

What it looks like in the field

What you will look at in the next check

In the following quiz, you check what each of the namespace inodes in /proc, the cgroup path and limits, and the capability bitmask proves. Answer by distinguishing the kernel values you read in the previous lab from the container settings on paper.