TT Lab
Get started
Learn Learning paths Courses

Computer Architecture

Storage Devices — From Spinning Platters to Parallel Queues

Continue in TT Lab

In one line

An HDD must physically move a head, so random access takes milliseconds, while an SSD has no such constraint but has a large erase unit, which gives writes their own distinctive cost structure.

Why this was needed

A large part of database and filesystem design was built on the premise of a "spinning platter." The fact that sequential access is a hundred times faster than random access gave rise to ideas such as the B-tree structure, write-ahead logging, and log-structured merge trees. When the storage medium changes, you need to know which of those premises survive and which collapse.

How it works

HDD. To read one block, the head must be moved to the right track (seek time, usually 5–10ms) and you must wait for the desired sector to come around under the head (rotational latency, on average about 4ms at 7200rpm). So random reads are limited to a few hundred per second. On the other hand, reading on without moving the head gives hundreds of MB per second. This property, a gap of more than a hundredfold between sequential and random, is the starting point of every classical design.

SSD. With no moving parts, random reads take microseconds. But writes are not symmetric. NAND flash is written in page units (a few KB) and can be erased only in block units (a few MB). Because you cannot overwrite a spot already written, the controller writes to a new spot and marks the old one invalid, then later gathers the valid pages, moves them, and erases the whole block (garbage collection). In this process more is actually written than the host requested, and that ratio is called write amplification. Also, each cell has a limit on how many times it can be erased, so the controller spreads writes evenly through wear leveling.

NVMe. SATA had one command queue with a depth of 32. NVMe can have tens of thousands of queues, each tens of thousands deep. An important shift happens here. The latency of an individual request does not drop by much, but you have to raise the number of requests thrown at once (the queue depth) to fill the bandwidth. IOPS measured at queue depth 1 shows only part of the device's performance.

What you see in the field

The default value of 4.0 for PostgreSQL's random_page_cost is based on spinning disks. If you leave this value as it is on SSD or NVMe, the optimizer prices random access at four times the real cost and undervalues indexes. A large share of the complaints "I created an index but it isn't used" are solved by lowering this one setting to around 1.1. It is a case where, once the medium has changed, the cost model built on that medium must change too.

Conversely, some premises are still valid. That sequential writes are better than random writes holds on an SSD as well. Only the reason has changed. On an HDD it was because the head did not move, and on an SSD it is because it reduces garbage collection and write amplification.

Getting a feel with numbers

The differences between storage devices are clear in a table. The orders of magnitude differ.

HDD SATA SSD NVMe SSD Memory
Random read latency 5–10 ms 0.1 ms 0.02 ms 0.0001 ms
IOPS (random 4K) 100–200 Tens of thousands Hundreds of thousands to a million —
Sequential bandwidth 150 MB/s 500 MB/s 3–7 GB/s 50 GB/s+
Queue depth 1 (effectively) 32 65,536 × several queues —

The reason an HDD's random IOPS is only 200 is physical. The head has to move and the platter has to spin, so each one takes 5–10ms. So if you put a database on an HDD, 200 random lookups per second is the upper limit.

The reason NVMe is faster than SATA SSD is not the medium but the interface. SATA has one queue with a depth of 32, whereas NVMe has multiple queues, each with a depth of 65,536. Each CPU core can have its own queue, so the parallelism is completely different.

Distinguishing sequential from random is the design

Even on the same device, performance differs by more than tenfold depending on the access pattern.

The third is often hit in practice. If you store image thumbnails or log fragments one file at a time, disk capacity remains but you run out of inodes and can write no more. Check with df -i.

When a write actually reaches the disk

Just because an application called write() does not mean the data is on the disk.

앱 버퍼 → 페이지 캐시(커널) → 장치 캐시 → 매체
              ↑ write() 는 여기까지만 보장한다
              ↑ fsync() 가 아래로 밀어낸다

If the power goes out, whatever was not fsynced is lost. This is why a database calls fsync on every commit, and that sets the upper limit on write performance.

Here comes the trap of disks with a write cache. If the device puts data only in its cache and returns "done," even fsync lies. Server SSDs carry a power-loss protection (PLP) capacitor to prevent this problem. This is why you should not use consumer SSDs in a DB server.

What to check in the quiz that follows

Check whether you can distinguish which design premises collapse and which remain when the medium changes.