Why KVM Is Close to Native
In one line
Virtualization performance is determined by how many times the guest drops out to kernel mode (VM Exits). Hardware support and virtio are both techniques that reduce that count.
Why you need this
x86 was originally a poor architecture to virtualize. Some of its privileged instructions, when run in unprivileged mode, raise no exception and quietly return a different value (the so-called "unvirtualizable instructions"). So early virtualization split into two paths.
- Full virtualization — inspects and rewrites guest code right before it runs (binary translation). You do not need to modify the guest OS, but it is slow.
- Paravirtualization — modifies the guest OS so that it calls the hypervisor directly (a hypercall) instead of issuing privileged instructions. It is fast, but the guest has to be modified.
Then Intel VT-x and AMD-V appeared and changed the game. They added a guest-only execution mode to the CPU.
How it works
Hardware-assisted virtualization
VT-x creates two new modes in the CPU: VMX root (the hypervisor) and VMX non-root (the guest). The guest runs in its own ring 0 as it is, but when a certain event that the hypervisor has specified occurs, it automatically drops out to root mode. This is a VM Exit.
Hardware helps with memory too. EPT (Intel) and NPT (AMD) handle the guest physical address → host physical address translation with hardware page tables. Before this existed, the hypervisor had to maintain shadow page tables in software, and that ate a large share of performance.
Where KVM sits
KVM is not a separate hypervisor but a Linux kernel module. The moment it is loaded, the Linux kernel itself becomes the hypervisor. So KVM sits on the boundary between type 1 (bare metal) and type 2 (hosted) — it runs on top of a host OS, but that host OS is itself the hypervisor.
The execution flow is like this.
QEMU (유저 공간)
├─ ioctl(KVM_RUN)
↓
KVM 모듈 (커널)
├─ VMLAUNCH / VMRESUME
↓
게스트 (VMX non-root)
├─ VM Exit 발생
↓
KVM 이 처리 가능하면 여기서 끝 (빠름)
KVM 이 못 하면 QEMU 로 반환 (느림)
What KVM handles directly in the kernel: most MSR accesses, simple I/O ports, EPT violations, external interrupts. What goes up to QEMU: complex device I/O, MMIO, some CPUID.
That sets the direction of performance tuning. Reduce the Exits that go up to QEMU.
The role of QEMU
QEMU is the device model. It creates, in software, the virtual NICs, disk controllers, and timers that the guest sees. If you use QEMU alone without KVM, even the CPU instructions are all translated in software, and that engine is TCG (Tiny Code Generator).
TCG cuts guest instructions into basic blocks, converts them to TCG IR, then translates them into host instructions and caches them (Translation Blocks). Thanks to the cache and chaining it is much faster than a pure interpreter, but it is still 10–100 times slower than KVM.
| Item | Native | QEMU+KVM | QEMU (TCG) |
|---|---|---|---|
| CPU computation | 100% | About 98% | About 25% |
| Disk I/O (virtio) | 100% | About 90% | About 30% |
| Network (vhost-net) | 100% | About 95% | About 25% |
There is no /dev/kvm in this lab environment. This is because that device was not given to the container. So QEMU runs with -accel tcg. The slow boot is due to the environment, not a configuration mistake.
virtio — the return of paravirtualization
An emulated e1000 NIC causes a VM Exit every time the guest touches a register. virtio has the guest and host communicate through shared-memory ring buffers, which dramatically reduces Exits. The guest needs a virtio driver, but modern Linux includes it by default.
| Device | Relative performance |
|---|---|
| virtio-net | More than 10 times e1000 |
| virtio-blk | 2–5 times IDE emulation |
| virtio-scsi | Advantageous for attaching many disks |
| vhost-net | Handles ring buffer processing in the kernel → further improvement |
When a performance report comes in, the first thing to ask is "are you using virtio?" Surprisingly many VMs are running on IDE because the wrong template was chosen.
What it looks like in the field
The nested virtualization trap. To use KVM again inside a cloud VM, nested virtualization has to be enabled. If it is not, it quietly falls back to TCG and you get "why is this so slow?" You confirm it with whether /dev/kvm exists and the output of qemu-system-x86_64 -accel help.
Poor performance because the CPU model was not set to host-passthrough. The default CPU model hides the latest instruction extensions for compatibility. For a workload that uses AVX, it makes a big difference. In exchange, live migration compatibility is broken, so it is a trade-off.
What to look for in the next check
The quiz that follows first separates the conditions of KVM acceleration versus TCG fallback, virtio devices, and the performance and portability of CPU models. After you confirm that judgment, from the next module you practice qcow2 backing files and snapshots, and the QEMU serial console and monitor socket.