Storage in Practice — RAID, Snapshots, iSCSI, fio
Build a RAID1, Kill a Disk, Bring It Back
Goal
Build an md RAID1 on loop devices, fail one disk to see the degraded state yourself, rebuild with a new device, and then leave evidence that the data is safe, using a checksum and a check.
Why it matters
Replacing a disk is the hardware task infrastructure engineers do most often, yet most people learn it by going through that moment for the first time. If you go through it by hand beforehand — what [U_] in /proc/mdstat is, why --fail and --remove are separate, and why "the rebuild finished" and "the data is intact" are different pieces of evidence — you will not panic at a dawn alert.
This lab runs as root on an Ubuntu 24.04 VM. Since there are no empty disks, you imitate three disks with loop devices (losetup), which make a file look like a block device. This is exactly the method used in practice when testing RAID without disks. The session starts at 60 minutes and can be extended up to 180 minutes, and the VM disappears when the session ends.
Steps
- Create
/var/lib/sto/raid-a.img,/var/lib/sto/raid-b.img, and/var/lib/sto/raid-c.img, each 256MiB, and attach all three as loop devices. Save thelosetup -aoutput to/root/sto/raid/loops.txt. - Create a RAID1 array
/dev/md/stofrom the two loop devices ofraid-a.imgandraid-b.img(metadata 1.2). Do not addraid-c.imgyet. - Format
/dev/md/stoas ext4 and mount it at/srv/raid. Write 20MiB of random data to/srv/raid/ledger.dat, and save the output ofsha256sum /srv/raid/ledger.datto/root/sto/raid/ledger.sha256. - Add an
ARRAYline for/dev/md/stoto/etc/mdadm/mdadm.conf. TheARRAYline carrying that array's UUID must be exactly one line. - Save
/proc/mdstatfrom right after marking the loop device ofraid-b.imgas failed (fail) to/root/sto/raid/degraded.txt, and remove that device from the array. - Add the loop device of
raid-c.imgto the array to rebuild, and wait until it finishes. The state must beclean, with 2 active devices and 0 failed devices. - Check
ledger.datagainst the checksum from step 3, and save the output ofsha256sum -cto/root/sto/raid/verify.txt. - Run a check on the array, and when it finishes, write two lines,
last_sync_action=<값>andmismatch_cnt=<값>(each placeholder is the value you read), to/root/sto/raid/scrub.txt(you read the values from/sys/block/<md 장치>/md/, where the placeholder is the md device).
Notes
- Loop device:
truncate -s 256M 파일→losetup -f --show 파일; to find it again,losetup -j 파일(in each command, replace the last word with the path of the image file). - Array:
mdadm --create,mdadm --detail,cat /proc/mdstat,mdadm --wait. - Replacement:
mdadm <배열> --fail <장치>→--remove <장치>→--add <새 장치>(the placeholders are the array, the device, and the new device). - Common mistake 1: running
--detail --brieftwice and piling up the sameARRAYline twice inmdadm.conf. - Common mistake 2: recording "it is done" before the rebuild has finished — wait until
recoveringdisappears.
Create three loop devices
Create /var/lib/sto/raid-a.img, /var/lib/sto/raid-b.img, and /var/lib/sto/raid-c.img, each 256MiB (sparse files are fine), and attach all three as loop devices. Save the losetup -a output to /root/sto/raid/loops.txt.
Set the size with truncate -s and attach to a free loop device with losetup -f --show. If you write the number yourself, it will collide with a device that snap is using.
Create the RAID1 array /dev/md/sto
Create a RAID1 array /dev/md/sto from the two loop devices of raid-a.img and raid-b.img (metadata 1.2). Do not add raid-c.img yet.
Give mdadm --create the --level and --raid-devices options. If it asks a confirmation question that this is not suitable as a boot device, you can answer y in this lab. You find which loop device is which again with losetup -j .
Leave a filesystem and a checksum
Format /dev/md/sto as ext4 and mount it at /srv/raid. Write 20MiB of random data to /srv/raid/ledger.dat, and save the SHA-256 of that file to /root/sto/raid/ledger.sha256 exactly in the format of sha256sum /srv/raid/ledger.dat.
An array is also a block device, so mkfs.ext4 and mount work as they are. You can create random data with head -c 20M /dev/urandom. The checksum file must contain the absolute path so that you can check it with sha256sum -c from anywhere later.
Register the array in mdadm.conf as one line
Add an ARRAY line for /dev/md/sto to /etc/mdadm/mdadm.conf. The ARRAY line carrying that array's UUID must be exactly one line.
mdadm --detail --brief produces the ARRAY line for you. First check whether that UUID is already in the file, and append only if it is not. On a real server you would also have to run update-initramfs -u so that it is assembled early in the boot.
Fail one disk and remove it
Save /proc/mdstat from right after marking the loop device of raid-b.img as failed (fail) in the array to /root/sto/raid/degraded.txt, and remove that device from the array. /srv/raid/ledger.dat must still be readable.
The order is mdadm --fail and then --remove . In mdstat right after the failure is marked, you see (F) after the device name and [2/1] in the status column.
Rebuild with a new disk
Add the loop device of raid-c.img to the array to rebuild, and wait until the rebuild finishes. Afterward, the state in mdadm --detail /dev/md/sto must be clean, with 2 active devices and 0 failed devices.
When you put it in with mdadm --add , the kernel starts a recovery. mdadm --wait waits for it to finish. You can see the progress in cat /proc/mdstat.
Check the data with the checksum
Check /srv/raid/ledger.dat against the checksum you left in step 3, and save the output of sha256sum -c to /root/sto/raid/verify.txt.
sha256sum -c prints OK or FAILED, one line per file. That the array is clean and that the data is intact are different pieces of evidence.
Run the check and leave the result
Run a check on the array, wait until it finishes, and then write two lines to /root/sto/raid/scrub.txt: last_sync_action=<값> and mismatch_cnt=<값> (each placeholder is the value you read). You read both values from the files of the same name under /sys/block/<md 장치>/md/ (the placeholder is the md device).
/dev/md/sto is a link that points to /dev/mdNNN (readlink -f). If you write check to /sys/block/mdNNN/md/sync_action of that device, the check starts, and you can wait for it to finish with mdadm --wait.