TT Lab
Get started
Learn Learning paths Courses

Storage in Practice — RAID, Snapshots, iSCSI, fio

Build the Way Back First — LVM Snapshots, Filesystem Choice, Thin Provisioning

Continue in TT Lab

In one line

An LVM snapshot taken right before a change creates "we can roll back" in a few seconds. In exchange, a snapshot becomes useless when its space fills up, it is not a backup, and it lives on the same disk as the original. When first choosing a filesystem, you must also look at whether it can be shrunk, and thin provisioning is a matter of promising space that does not exist, so monitoring comes with it.

Why you need this

Work such as package upgrades, bulk edits of configuration files, and schema changes is approved on the premise that "if it goes wrong, we roll back." But if the only way to roll back is "restore from backup," it takes an hour, and the data that came in during that time is lost. What was needed was a means of instantly freezing the state of a volume right before the work and returning to that moment if it fails. That is the snapshot.

One more thing. It is common to size a volume generously and then need space on another volume. If you had chosen a "filesystem that cannot be shrunk" then, your only option is to create a new one and move. The choice of filesystem is decided once at the start and lasts a long time.

How it works

COW snapshots. The -s of lvcreate(8) creates a snapshot of the origin LV. At the moment of creation, nothing is copied. It is a copy-on-write method: afterward, when any block of the origin is first changed, the old contents are moved to the snapshot space. So the size of the snapshot (-L) is set not to the origin's size but to "the amount that will change while the snapshot is kept." The Data% column of lvs shows how much of that space has been used, and when it reaches 100% the snapshot becomes invalid and is discarded. If a snapshot quietly becomes invalid during a big job, you keep working without knowing the way back is gone.

Rolling back is a merge. The --merge of lvconvert(8) writes the snapshot's contents back to the origin and removes the snapshot. If the origin is open (mounted), the merge is deferred until the next activation, so in the work plan you write the order "unmount → merge → mount again." When the merge finishes, the snapshot LV disappears from the list.

A snapshot is not a backup. It lives in the same VG as the origin, usually on the same disk. If the disk dies, the origin and the snapshot disappear together. And while a COW snapshot is attached, every first write to the origin causes one more copy, so writes get slower. So you keep a snapshot only for the duration of the work window and delete it when finished.

What to look at when choosing a filesystem. Both are journaling filesystems and both can be grown while mounted. The difference shows up when shrinking.

ext4 XFS
Online growth resize2fs xfs_growfs (only while mounted)
Shrinking Possible with resize2fs after unmounting Design as if it is outside the supported range — create a new one and move
Common default Debian and Ubuntu The RHEL family

resize2fs(8) can grow a mounted ext4, and shrinking is done only with it unmounted. xfs_growfs(8) is, as the name says, a tool for growing. A recent xfsprogs has gained an experimental feature that shrinks only the free space of the last allocation group, but in operational design you take "XFS cannot shrink" as a premise.

When shrinking, the order is your life. First shrink the filesystem to smaller than the target, then shrink the LV, and finally resize the filesystem again to fit the LV size. If you shrink the LV first, the tail end of the filesystem gets cut off. lvreduce -r hands this process to fsadm and does it in one go, but you must know what is happening before you use it.

Thin provisioning. The thin pool of lvmthin(7) is a pool that has real space, and a thin volume receives space from that pool when it writes. So you can create a 1GiB thin volume on a 400MiB pool. This is overprovisioning. It relies on the statistic that not all volumes fill up, and it is the same trade as memory overcommit in virtualization. lvmthin(7) warns that when the pool fills up, writes stop or become errors, and it describes a setting that monitors pool usage and extends it automatically (thin_pool_autoextend_threshold). If there is no spare VG space to extend the pool with, that setting does not help either.

What it looks like in the field

Many teams put a snapshot in the "pre-work" column of the change work plan. A good plan also writes the basis for the snapshot size (an estimate of how much the work will change) and a criterion such as "stop if snapshot usage exceeds some percent." It also writes a line to delete the snapshot when the work is finished. A snapshot that someone forgot to delete often hits 100% a few weeks later and becomes invalid, or keeps dragging down write performance.

I also often see a request like "let's shrink this volume by half and give the space to the neighboring volume" stop in front of XFS. It happens because the RHEL-family default is XFS, so people pick it without thinking at install time. For a volume that may need shrinking, go with ext4 from the start, or design it to start small and grow rather than sizing it generously.

What you will do in the next lab

You build a VG vg_sto on 3GiB loop devices and put a configuration file and an order file on an XFS volume lv_app. You take a snapshot, deliberately make a wrong change, record the snapshot usage, and then roll back with a merge. Next you grow the XFS online, and create an ext4 volume and shrink it offline while protecting the data. Finally, you create a 1GiB thin volume on a 400MiB thin pool to see overprovisioning with your own eyes.