One Instruction Can Take 62 Seconds — Latency Is a Property of the Path
In one line
The question "how many cycles is this instruction?" has no answer unless you pin down the surrounding state. Latency is not a property of an instruction but an observation of the combination of the instruction and the system state.
Why this was needed
When you learn optimization, you start by memorizing instruction latency tables. Addition 1 cycle, multiplication 3 cycles, division 20 cycles. This table is useful, but the moment you believe those numbers are fixed properties of an instruction, you can no longer explain real performance.
A leaderboard called the Assembly Hall of Shame, published in 2026, shows this point in the extreme. Its subtitle is "a race to the bottom of CPU performance," and it records a competition to make a single instruction as slow as possible. nop at the bottom takes 1 cycle, and the top entry, fxrstor64, takes 198 billion cycles, which is 62 seconds. If there is a 200-billion-fold difference within the same class, you should begin by doubting the premise that you are measuring that class with a single number.
How it works
Read that leaderboard from the bottom up and it becomes a list of the points where a modern CPU can stall. It divides into three zones.
First, the points where microcode steps in inside the core. If you give input that the hardware fast path cannot handle, control passes to a microcode routine. fadd takes a few cycles with normal operands, but if the operand is a denormal number, it takes 677 cycles. The instruction is the same and only the data changed, yet a two-digit multiplier appears.
Second, the points where coherency and the platform intervene. If the operand of an atomic operation straddles a cache line boundary (a split lock), the fast cache coherence path cannot be used and an external bus lock must be taken. It takes 865 cycles, and other cores are affected in the meantime. wbinvd, which flushes the entire cache, takes 1.6 million cycles, and that is not because the instruction is complex but because all the dirty data in the cache at that moment has to be pushed out to DRAM.
Third, the moment of leaving the die. The top ranks are all mov. An instruction that does nothing but move data takes more than a second. The reason lies in the address. It is reading a device register (MMIO) beyond the PCIe fabric, not DRAM. An MMIO read is a non-posted transaction that finishes only when a response arrives, so the round-trip time becomes the execution time as it is.
The strategy of the first-place entry is decisive. Having it restore 512 bytes of state from the slowest MMIO region gives 23 seconds. Adding to this making the other cores keep hammering other MMIO registers gives 62 seconds. That is, the execution time of a single instruction more than doubled depending on what the other cores were doing.
What you see in the field
The conclusion of this extreme game applies directly to ordinary performance work.
A microbenchmark measures system state, not an instruction. The same code gives completely different numbers when the cache is warm and when it is cold, when other cores are idle and when they are busy. That is why you must note the conditions when quoting a benchmark figure.
The worst case is not a multiple of the average. The zones above are not continuous but separated by cliffs. No matter how good the average latency, stepping off one cliff makes that one request a hundred times slower. When dealing with tail latency, it is more effective to find and remove cliffs than to improve the average.
There are three cliffs you actually step off in practice. The denormal numbers in signal-processing code where values converge to 0, the split lock in atomic operations where alignment was ignored (the Linux kernel's split_lock_detect prints a warning with its default value warn), and MMIO polling, which reads a device register inside a loop.
What to check in the quiz that follows
Check whether you can answer the question "why was this code fast yesterday and slow today?" with paths and state rather than a list of instructions.