Storage in Practice — RAID, Snapshots, iSCSI, fio
Turn "The Disk Is Slow" into Numbers — IOPS, Latency, Queue Depth and fio
In one line
Disk performance is not one number but three. How many per second (IOPS), how much per second (throughput), and how long one takes (latency). These three change with one another depending on block size and queue depth, so "how many IOPS does it give" has no meaning unless you state the conditions too. fio is a tool that measures with those conditions held fixed.
Why you need this
A report that "the storage is slow" mostly comes without evidence. Before buying new storage, before moving a LUN, before choosing a cloud volume tier, you need numbers measured under the same conditions to compare. But a number measured with dd if=/dev/zero of=파일 is almost always wrong. Writes go into the page cache first and show memory speed, and because it writes sequentially one at a time, it measures something completely different from the small random reads a database actually does.
How it works
The relationship of the three numbers. Throughput = IOPS × block size. 10,000 IOPS at 4KiB is about 40MiB/s, and 400 IOPS at 1MiB is 400MiB/s. A database's random reads are small blocks, so IOPS and latency matter, while backups or large copies are big blocks, so throughput matters. So in measuring, you first decide "what to imitate."
Queue depth and Little's law. The number of requests in flight at the same time is the queue depth. Little's law from queueing theory says "the average number in the system = the arrival rate × the average time spent." Applied to storage, the number of concurrent requests ≈ IOPS × average latency. At queue depth 1, IOPS cannot exceed 1 ÷ latency. With a latency of 0.2ms, the maximum is 5,000 IOPS. If you raise the queue depth, the device overlaps several requests and IOPS goes up, but each request waits in line, so latency goes up too. The big IOPS in a storage spec sheet is usually a value measured at a deep queue, and an application that waits for one request at a time will never see that number.
Percentiles, not averages. Even if the average latency is 1ms, if one request in 100 takes 50ms, a service that reads the disk several times per request will often step on that tail. That is why you also look at the p99 (99th percentile) latency.
fio's options. Among the options in the fio documentation, a few govern the result.
| Option | Meaning | Value in this lab |
|---|---|---|
rw |
Read/write, sequential/random | randread |
bs |
Block size | 4k |
ioengine |
How requests are issued | libaio (asynchronous, so queue depth is meaningful) |
iodepth |
Queue depth | 1 and 32 |
direct |
Bypass the page cache | 1 |
runtime, time_based |
Keep running for a set time | 10 seconds |
output-format |
Output format | json |
direct=1 opens with O_DIRECT and skips the page cache. If you leave it out, from the second read on the answers come from memory, and you measure RAM, not the disk. With a synchronous method (psync) instead of ioengine=libaio, even if you raise iodepth, only one request goes out at a time, so it is not a queue depth experiment.
How to read the JSON. Under jobs[0].read in the --output-format=json result, there are iops, the latency statistics lat_ns (the average mean), and the "99.000000" key of the completion latency clat_ns.percentile. The unit is nanoseconds, so to convert to microseconds you divide by 1,000. Do not copy the human-readable default output; take the numbers out of the JSON, so that the report and the reproduction match.
What it looks like in the field
Cloud block volumes usually sell their IOPS and throughput caps by tier. That cap is reached only at a sufficiently deep queue. If an application does synchronous writes at queue depth 1 (for example, a database log that calls fsync for every transaction), even an expensive tier gives you only as much as the latency allows. The number to look at then is not IOPS but latency.
In a test for adopting new storage, you fix two or three conditions that imitate the production load and always measure with the same command. If you keep the command and the JSON result together, then when someone says "it got slower than before" half a year later, you can run the same command again and answer with numbers. Do the measurement on a dedicated file on a filesystem that is in operation; running a write measurement on a raw device erases data, so do it only on an empty device. And during the measurement, keep iostat -x open next to it, so you can also check whether fio's numbers match the numbers the kernel sees for the device.
What you will do in the next lab
You set up a LIO target inside the same VM to provide a 512MiB LUN, discover and log in with the initiator, and check that the new disk shows up as iscsi. You format it as ext4, mount it with an fstab line that has _netdev, and turn on automatic login. On a dedicated file on it, you measure 4KiB random reads at queue depth 1 and 32, extract IOPS, p99, and average latency from the JSON, and check whether Little's law holds.