TT Lab
Get started
Learn Learning paths Courses

Computer Architecture

I/O and Buses — Interrupts, DMA and PCIe

Continue in TT Lab

In one line

How you talk to a device comes down to three layers. When to notify (polling versus interrupts), who does the moving (CPU versus DMA), and where it passes through (buses and interconnects).

Why this was needed

A CPU operates on the scale of nanoseconds, while a disk or network card operates on the scale of microseconds to milliseconds. If this speed difference is left as it is, the CPU spends all its time waiting for devices. The history of I/O design is an accumulation of answers to the question of how to hide this difference.

How it works

Polling and interrupts. Polling is a method in which the CPU keeps reading the device's status register. It is simple and has low latency, but the CPU can do nothing else meanwhile. An interrupt is a method in which the device sends a signal to the CPU when it is ready. The CPU does other work in the meantime, but each interrupt costs saving the context and jumping to the handler. So on a network card receiving hundreds of thousands of packets per second, interrupts can actually paralyze the system (an interrupt storm), and a hybrid approach that switches to interrupts at first and to polling when busy, like Linux's NAPI, became standard.

DMA. If data is moved one word at a time through the CPU registers, the CPU is tied down to hauling data. A DMA controller moves data directly between the device and memory and raises just one interrupt when it is all done. In modern systems, large transfers are DMA without exception.

MMIO and port I/O. If you map device registers into the memory address space, you can access them with ordinary loads and stores (MMIO). The trap is that the syntax is the same as for memory. The cost is entirely different. An MMIO read is a non-posted transaction that must wait for a response, so the PCIe round-trip time becomes the execution time as it is, and it cannot be loaded into the cache. If driver code reads a device register inside a loop, that loop is not computation but a repetition of I/O round trips.

PCIe. Bandwidth is determined by the number of lanes and the generation. With each generation the per-lane speed roughly doubles, and lanes are bundled, as in x1/x4/x8/x16, to widen the link. If you plug a GPU into an x4 slot instead of an x16 slot, the compute performance stays the same but the data supply drops to a quarter. This is also the reason dedicated interconnects such as NVLink are used when training with several GPUs tied together. Communication between cards through PCIe becomes the bottleneck.

What you see in the field

The most common mistake when building a GPU node in a home lab is slot bandwidth. The second and third slots on a motherboard are often electrically x4, so if you plug in three cards, two of them quietly run at a quarter of the bandwidth. The training itself works, but the step time becomes inexplicably long. The habit of checking the negotiated link width with lspci -vv saves time.

Seeing where the bottleneck is in numbers

Bandwidth is easy to get a feel for when you remember it by orders of magnitude.

Path Bandwidth (approximate)
CPU ↔ L1 cache 1 TB/s or more
CPU ↔ memory (DDR5) 50–100 GB/s
PCIe 4.0 ×16 32 GB/s
PCIe 4.0 ×4 (one NVMe drive) 8 GB/s
10 GbE network 1.25 GB/s
SATA III 0.6 GB/s

PCIe becomes the bottleneck when you use several GPUs. 32GB/s is less than half of the memory bandwidth, so work that frequently moves data between cards (model parallelism) is slow. This is why NVLink exists.

Look at the lane count too. Even if you plug into a ×16 slot, if the motherboard connects it as only ×4, the bandwidth is a quarter.

lspci -vv -s <슬롯> | grep -E 'LnkCap|LnkSta'
LnkCap: Speed 16GT/s, Width x16      ← 장치가 할 수 있는 것
LnkSta: Speed 16GT/s, Width x8       ← 실제로 연결된 것   ← 절반이다

If these two lines differ, it is a problem with the slot or the bifurcation settings. If the GPU performance is half of what you expect, start looking here.

Interrupts and polling

There are two ways to notify the CPU when a device has finished its work.

High-speed networking and NVMe mix the two. Linux's NAPI turns off interrupts and switches to polling when packets pile up. So a high si (softirq) utilization means the network is busy.

The interrupt coalescing settings of ethtool -c eth0 adjust this balance. Raising the value increases throughput and increases latency.

What DMA does

The device reads and writes memory directly, without going through the CPU. Without this, reading 1GB from disk would make the CPU haul all of that 1GB.

What comes out of this is zero-copy. When sending a file to a socket, it is normally copied four times: disk → kernel buffer → user buffer → kernel socket buffer → NIC, but sendfile() and splice() skip user space. This is why static file servers are fast, and it is what nginx's sendfile on does.

What to check in the quiz that follows

Check whether you can explain why the same mov instruction takes nanoseconds at some addresses and milliseconds at others.