Storage in Practice — RAID, Snapshots, iSCSI, fio
RAID Is Availability, Not Backup — The Rebuild Is the Riskiest Time
In one line
RAID is a mechanism that keeps a server from stopping when one disk dies. Deleted files, wrong changes, and encryption attacks are replicated identically to every disk, so RAID cannot stop them. And the rebuild time, during which you replace the dead disk and bring the array back into sync, is the most dangerous stretch in the whole array's life.
Why you need this
Disks are consumables. If a server has several disks, one of them dying is not a matter of "if" but of "when." If you reinstall the operating system and restore from backup every time a disk dies, the service stops for hours each time. RAID (Redundant Array of Independent Disks) ties several disks together as if they were one and absorbs the death of a single disk.
Linux provides software RAID with the kernel's md (multiple devices) driver, and mdadm(8) creates and manages the arrays. With a hardware RAID card, the controller does the same job and shows the operating system just one disk. So the state of the disks behind the card usually has to be viewed with the manufacturer's tools. The principle is the same, so this course teaches it with md.
How it works
The levels that md(4) describes, condensed into a table, look like this.
| Level | Method | Failures tolerated | Usable capacity |
|---|---|---|---|
| RAID0 | Striping | None — if even one disk dies, you lose everything | All N disks |
| RAID1 | Mirror | N−1 disks | One disk |
| RAID5 | One distributed parity | 1 disk | N−1 disks |
| RAID6 | Two parities | 2 disks | N−2 disks |
| RAID10 | Stripe of mirrors | 1 disk per mirror pair | Half |
RAID5 and RAID6 are cheap in capacity but expensive for small writes. To change one block, you have to read the old data and the old parity, compute the new parity, and write both again. This is why you choose RAID10 for a write-heavy database volume.
The state of an array is all on one line of /proc/mdstat. [2/2] [UU] means two of two disks are alive, and [2/1] [U_] is a degraded state where one disk is missing. A failed device gets (F) after its name. The State : line of mdadm --detail shows the same fact in words such as clean, degraded, and recovering.
Replacing a disk takes four operations — mark the failure with --fail, take it out of the array with --remove, and when you put in the new disk with --add, the kernel starts a recovery that copies the contents of the surviving side to the new disk. Only when this copy finishes is it clean again. mdadm --wait waits until that operation is done.
The reason a rebuild is dangerous is that during it, the one surviving disk is the only copy. A rebuild reads every block of that disk. If an unreadable sector was hiding in a corner nobody normally reads, it shows up right now, and there is nowhere to restore that block from. That is why md provides a scrub that reads everything ahead of time in normal operation. According to the kernel's md documentation, writing check to /sys/block/mdX/md/sync_action reads and compares all copies, leaves the number of differing blocks in mismatch_cnt, and when it finishes sync_action goes back to idle and last_sync_action becomes check. The mdadm package on Debian-family systems schedules this check to run periodically.
The array has to be reassembled at boot. mdadm knows which devices with which UUID form one array from the ARRAY lines in its configuration file. On Ubuntu and Debian this is /etc/mdadm/mdadm.conf, and to have it assembled early in the boot, you also put it into the initramfs with update-initramfs -u. mdadm --detail --brief produces that line for you. If you run the command twice and append with >>, the same line piles up twice; this file is also a document that people reread during an incident, so you leave just one line per array.
What it looks like in the field
The most common big accident is pulling a disk that is alive. When the alert says "one disk failed," if you trust only the slot number and pull out the healthy side, a RAID1 goes down entirely. Before replacing, look at the failed device name with mdadm --detail, check the serial number with lsblk -o NAME,SERIAL or smartctl -i, and match it against the label on the physical disk. Many teams write this check as a line in the work plan.
The second is the belief that "it is RAID, so we do not need backups." An rm -rf happens on both disks at once, and a wrong migration is replicated as faithfully as a mirror. RAID prevents just one kind of accident, disk failure. Undoing is the job of snapshots and backups, which are covered in the next module and in the capacity and change management course.
The third is that a rebuild slows the service. A rebuild is disk I/O too, so it competes with business traffic. md(4) lets you adjust the lower and upper bounds of the rebuild speed with /proc/sys/dev/raid/speed_limit_min and speed_limit_max. You operate by lowering the upper bound during business hours and raising it at night, but the price of going slower is that the dangerous stretch gets longer.
What you will do in the next lab
You create three 256MiB loop devices and build a RAID1 /dev/md/sto from two of them. You format it as ext4, leave a 20MiB file and its checksum, and register the array in mdadm.conf as a single line. Then you fail one disk and remove it to record the degraded state, rebuild with the third disk, and use the checksum to confirm the data is safe. Finally you run a check and record mismatch_cnt.