TT Lab
Get started
Learn Learning paths Courses

Virtualisation with QEMU/KVM

qcow2 and the Backing Chain

Continue in TT Lab

In one line

The real power of qcow2 is not compression or encryption but the backing file. You can create dozens of VMs by laying thin overlays on top of one golden image.

Why you need this

You have to bring up 100 VMs. Each needs a 20GB disk. As is, that is 2TB. But those 100 start from the same OS image, and the part that actually differs is a few hundred MB per VM.

The backing file of qcow2 solves this problem. Reads come from the backing file, and writes go only to the overlay. The 100 VMs share one 20GB golden image, and each holds only its own changes.

How it works

Basic operations

qemu-img create -f qcow2 base.qcow2 20G        # 씬 프로비저닝: 실제 점유는 거의 0
qemu-img info base.qcow2
qemu-img info --output=json base.qcow2         # 스크립트로 파싱하기 좋다
qemu-img convert -f raw -O qcow2 in.raw out.qcow2
qemu-img check base.qcow2

There are two values to look at in qemu-img info.

The difference between the two is thin provisioning. The disk size of a freshly created 20G image is about 200KB. If you look only at virtual size in monitoring, you miss overcommit. A situation where the host disk is 500GB while the sum of virtual sizes is 2TB holds up perfectly well, and it blows up when the guests actually start to fill it.

Backing files

qemu-img create -f qcow2 -b /var/lib/libvirt/base.qcow2 -F qcow2 vm01.qcow2
qemu-img info --backing-chain vm01.qcow2

You must specify the backing file's format explicitly with -F. It used to be possible to omit it, but now it gives a warning or is rejected. This is because letting it guess the format can be a security problem.

There are rules for backing chains.

  1. You must never modify a backing file. An overlay works on the premise that "this block of the backing has not changed," so if the backing changes, the overlays quietly break. It is the convention to keep a golden image read-only.
  2. The path is recorded inside the image. If you move the backing file, the overlay cannot find it. You can fix just the record with qemu-img rebase -u -b <새경로> <오버레이> (the placeholders are the new path and the overlay; -u means unsafe — it changes only the reference without actually moving data).
  3. Read performance drops as the chain gets longer. To find one block you have to trace back through several files. Flattening with qemu-img convert is a routine maintenance item.

Two kinds of snapshots

Internal snapshots — hold several points in time together inside the qcow2 file.

qemu-img snapshot -c before-upgrade disk.qcow2   # 생성
qemu-img snapshot -l disk.qcow2                  # 목록
qemu-img snapshot -a before-upgrade disk.qcow2   # 적용(되돌리기)
qemu-img snapshot -d before-upgrade disk.qcow2   # 삭제

They are convenient because they are managed as a single file, but the file grows and you must not touch it with qemu-img while the VM is running (changing a running image causes corruption). A snapshot while running has to go through the QEMU monitor or libvirt.

External snapshots — create a new overlay that uses the current image as its backing.

qemu-img create -f qcow2 -b disk.qcow2 -F qcow2 disk.snap1.qcow2

The original is frozen at that point and later writes go to the new file. A backup tool can copy the original safely, so this is the standard in backup workflows. Rolling back ends by discarding the overlay. In exchange, the number of files grows and you need to manage the chain.

Item Internal snapshot External snapshot
Number of files 1 1 per point in time
Rollback -a Delete the overlay
Backup friendliness Low High (the original is fixed)
Creation while running Needs monitor/libvirt Needs monitor/libvirt
Chain management Not needed Needed

cluster_size

The allocation unit of qcow2. 64KB by default.

qemu-img create -f qcow2 -o cluster_size=1M big.qcow2 100G

A larger value reduces metadata and helps sequential I/O, and a smaller value reduces waste in random writes. Raising it to 1M for large images is a common tuning.

What it looks like in the field

The accident where modifying a backing file breaks dozens of VMs. The moment you boot the golden image directly to apply a patch, every overlay stacked on it loses consistency. You have to make a copy of the golden image, apply the patch there, and deploy it as a new generation.

The accident where the host disk suddenly fills up. If thin-provisioned images grow at the same time, the host dies first. The VMs hit disk errors and their filesystems break. You have to monitor the sum of virtual sizes and the actual free space together.

What you will do in the next lab

You create a qcow2, read its information, convert it, set up a backing chain, and fix the path with rebase. In the lab that follows, you handle both internal and external snapshots and make a comparison table.