TT Lab
Get started
Learn Learning paths Courses

Virtualisation with QEMU/KVM

Where Is the Isolation Boundary

Continue in TT Lab

In one line

A VM is isolated by a hardware boundary, and a container is isolated by kernel features (namespaces/cgroups). All the remaining differences follow from this one sentence.

Why you need this

"Should we use containers or VMs?" is still a frequent question. The answer is decided not by the workload but by what kind of isolation you need.

How it works

[VM]                          [컨테이너]
앱                             앱
게스트 라이브러리               라이브러리
게스트 커널      <- 별개!       (호스트 커널을 공유)
가상 하드웨어                   namespace + cgroup
하이퍼바이저                    컨테이너 런타임
호스트 커널                     호스트 커널

A container has no guest kernel. So it starts fast (tens of ms), has almost no memory overhead, and has high density. In exchange, a single kernel vulnerability can bring down the isolation altogether.

Item VM Container
Kernel Separate for each guest Shared with the host
Start time Tens of seconds Tens of ms
Memory overhead Guest kernel + hypervisor Almost none
Density Tens per host Hundreds per host
Isolation strength Strong (hardware boundary) Weak (kernel boundary)
Different OS/kernel Possible Not possible
Image size GB MB

When to use which

When you should choose a VM

When you should choose a container

In between — microVMs

Things like Firecracker and Kata Containers fill the gap between the two. They keep the VM's isolation boundary while reducing start time and overhead to something close to a container's. They cut the device model down to the extreme (just a few virtio devices) and shorten the boot path so that it comes up in around 100ms. AWS Lambda runs on Firecracker.

Kata Containers goes one step further and keeps the container interface (OCI/CRI) while running a lightweight VM inside. From Kubernetes's point of view it is just a Pod, but in reality it is isolated by a VM boundary. You can choose it per workload with a RuntimeClass, so a policy such as "strong isolation for only this namespace" is possible.

What it looks like in the field

"I need to change a kernel parameter in the container but it doesn't work." Of course it doesn't. The kernel is shared with the host, so running sysctl -w inside a container would affect the entire host. That is why most of them are blocked. If you really need it, set it at the node level (a DaemonSet, node bootstrap) or go with a VM.

This lab environment is exactly that example. It is a container with capabilities removed, so mount, tcpdump, and iptables do not work. So this curriculum did not work around that constraint but took the constraint itself as the teaching material — knowing why an operation requires privileges is itself understanding the isolation structure.

Why density and resource management differ

VMs and containers also handle resources differently, and that difference often surprises people in operations.

A VM's memory is reserved in advance. A VM given 4GB takes that much on the host no matter how much it uses inside (you can give some back with a mechanism such as a balloon driver, but it is not immediate). By contrast, a container's memory is taken only as much as it actually uses. The limit is only a ceiling, and if it is not used, other containers use it. So on the container side, it is normal for actual usage to be much smaller than the sum of requests, and this is the source of density.

In exchange, containers carry the risk of overcommitment. If the sum of limits is allocated as more than the node's memory and several use a lot at the same time, the node falls into memory pressure and the kernel picks processes to kill. What dies then is not the one that used the most but the one with the largest excess over its request, so workloads that have not written down a request are sacrificed first.

CPU is different in nature yet again. Unlike memory, CPU can be taken back later, so when a limit is exceeded, it is slowed by the throttling seen earlier instead of being killed. So the general advice becomes set CPU limits generously and memory limits accurately.

When you run containers inside a VM, two layers overlap. The memory given to the VM becomes the node's total memory, and containers share it again inside. If a mechanism for reclaiming memory is on at the VM side, the node keeps scheduling without knowing its own memory has shrunk, and unpredictable failures occur. This is why the general recommendation is to turn off memory reclaim on virtualized nodes.

The microVMs seen earlier are also a compromise between the two. The isolation is that of a VM, but the memory overhead is small, so you get a kernel boundary without losing much density.

What comes next

The next course covers rootless containers. If you dig into why containers run without privileges (user namespaces), the sentence "it is isolated by a kernel boundary" from this module turns into concrete system calls and files.