TT Lab
Get started
Learn Learning paths Courses

Computer Architecture

GPUs — Widening Throughput Instead of Cutting Latency

Continue in TT Lab

In one line

A CPU is built to finish a single thread as fast as possible, and a GPU is built to run tens of thousands of threads at once so that they hide each other's memory wait times.

Why this was needed

A CPU spends its transistors on reducing the latency of a single flow. The branch predictor, the out-of-order execution engine, and the large caches are all devices that keep "this thread from stalling right now." The problem is that the returns on this approach diminish. Even if you double the size of the predictor, performance rises by a few percent.

There is a way to use the same transistors differently. Simplify the control circuitry to the extreme, fill the chip with arithmetic units, and launch a great many threads. While one thread waits on memory, running another thread hides the wait. This is the GPU's design philosophy, which is why a GPU is not a device that reduces latency but one that hides latency.

How it works

A GPU consists of many streaming multiprocessors (SMs). Inside each SM, threads are grouped in 32s and scheduled in warp units, and the threads within a warp execute the same instruction on their own data (SIMT).

Two performance traps come out of this.

Branch divergence. If half of a warp goes down the if branch and half down the else branch, the hardware executes both paths sequentially, temporarily switching off the threads that do not apply. The result is correct, but the time becomes the sum of the two paths. This is why kernels whose conditions vary by data are slow.

Coalescing. When the 32 threads of a warp read consecutive addresses, it finishes with just a few memory transactions. When they read scattered addresses, in the worst case 32 transactions are needed. This is why the same computation differs by a multiple with only a change of indexing.

The memory hierarchy also differs from a CPU's. Inside an SM there is shared memory that the programmer manages explicitly. Instead of waiting for the cache to do it for you, you load the data to be reused yourself, and several threads share it. Most of the performance of a matrix multiplication kernel comes from this shared-memory tiling.

Occupancy is the number of warps actually resident relative to the number of warps an SM can keep at once. When occupancy is low, there are too few threads to hide with, so the memory wait shows through as it is. However, high occupancy does not always mean faster. For a kernel that uses many registers, it is sometimes better to increase the work per thread even at the cost of lower occupancy.

What you see in the field

It is common to hear that GPU utilization in AI inference never exceeds 30 percent. The cause is usually not a lack of compute but memory bandwidth. LLM decoding reads a huge set of weights, multiplies them by a small input, and discards them, so its arithmetic intensity is low, and it therefore sits on the left slope of the roofline. In this region, buying a chip with higher compute performance does not raise performance. The answer is to reduce the bytes you read (quantization) or to increase the batch so that several requests share the same weight reads.

This is also the point at which accelerators like the TPU gave a different answer. A systolic array flows data through the grid and reuses it, maximizing the number of operations per memory access. It is a design that gave up generality to buy arithmetic intensity.

What suits a GPU and what does not

Even for the same problem, some things gain from being put on a GPU and others actually lose. There are three criteria.

Is there enough work to do at once? A GPU earns its keep only by running tens of thousands of threads, so a computation with only a few thousand elements does not even recoup the cost of launching a kernel. Launching a kernel itself takes several microseconds, and in that time the CPU has often already finished the job.

Is what you move small relative to the amount of computation? CPU memory and GPU memory are separate, so data has to be copied, and that channel is much narrower than the bandwidth inside the GPU. If the time spent copying is longer than the computation time, the GPU is a pure loss. So the standard practice is to keep using data, once uploaded, on the GPU across several stages, and a structure that moves it down to the CPU and back up at every stage is almost always a wrong design.

Do all threads do the same thing? A computation that diverges because the condition differs per element loses much of its gain to the branch divergence seen earlier. So problems that suit a GPU are generally regular operations over dense arrays.

Invert these three and you get exactly what does not suit a GPU. Logic with many branches, data structures that follow pointers, sequential processing in which the next step can begin only when the previous result exists, and short jobs that compute only a little each time. The reason AI training and inference suit GPUs so well is not some mysterious property but that they satisfy these three conditions exactly.

One more practical sense to add. When performance is poor on a server with a GPU, people suspect the GPU first, but very often the cause is the side that supplies the data. If the pipeline that reads from disk, preprocesses, and uploads cannot fill the GPU, the expensive card spends most of its time waiting idle. If you plot utilization over time and it rises and falls like a sawtooth, that is the typical look of this situation.

What to check in the quiz that follows

Check whether you can explain why the statement "a GPU is faster than a CPU" is only half true, and the typical reasons performance does not materialize even after you buy a GPU.