TT Lab
Get started
Learn Learning paths Courses

LLM Serving

prefill and decode Are Different Jobs

Continue in TT Lab

In one line

LLM inference splits into two phases. Prefill, which processes the prompt all at once, is compute-bound, and decode, which produces tokens one at a time, is memory-bandwidth-bound.

Why this was needed

In training, you only need to watch one metric — tokens processed per second. In serving, that is not enough. There are two things users feel, and they come from different causes.

TTFT (Time To First Token) is the end-to-end time from sending a request until the first token arrives. For a conversational service you can set 200ms as an example SLO, but the real target must be decided based on the product experience, the model size, the prompt length and the deployment environment. TTFT includes the network, the request queue, scheduling, tokenization and prefill. For long prompts, prefill, which computes attention over the whole prompt, easily becomes a major component, but it does not always decide the whole of TTFT by itself.

ITL or TPOT (Inter-Token Latency, Time Per Output Token) is the interval at which tokens come out one at a time after that. It only needs to be faster than a person reads. This value is decided by the decode phase.

The key point is that the two phases have different bottlenecks. Prefill involves large matrix multiplications, so it is bound by compute. Decode has to read the entire set of model weights once for each token, so it is bound by memory bandwidth. That is why the time of decode does not grow much even when you increase the batch — the time to read the weights dominates anyway. This property is why batching is so dramatically effective.

How it works

The KV cache is the newly appearing resource. In each token, a transformer refers to the keys and values of all the previous tokens. Recomputing them every time would be O(n²), so the computed keys and values are piled up in memory. This is the KV cache.

The problem is the size. The formula is this.

KV 바이트 = 2 x 레이어 수 x 은닉 차원 x 시퀀스 길이 x 배치 x dtype 바이트

If you run a 7B model (32 layers, hidden size 4096) in fp16 with a 4K context and batch 1, it is 2 x 32 x 4096 x 4096 x 1 x 2 = 2GiB. If you extend the context to 128K, it is 64GiB. The model weights are 14GB, yet the KV cache eats 64GB. With batch 8 it is 512GiB, which can never fit on a single GPU.

That is why most of the design of a serving system becomes a story about KV cache management. PagedAttention splits the KV cache into fixed-size blocks and, like the virtual memory of an OS, allows non-contiguous allocation, which reduces internal fragmentation. GQA shares key/value heads in groups according to the model's number of KV heads, reducing the cache size. KV cache quantization (INT8, INT4) reduces the dtype bytes, but whether it is supported and its effect on accuracy must be checked per server and model.

What it looks like in the field

A common mistake in measurement is lumping TTFT and ITL together into an average. If you divide the total latency by the total number of tokens, the two metrics get mixed and you can see nothing. You must measure them separately and look at the p50, p95 and p99 of each.

And the GPU memory utilization setting (for example gpu_memory_utilization in vLLM) matters. In some dedicated GPU environments, around 0.90 is taken as a starting point, but the safe value varies with the model server version, quantization, CUDA graphs and co-resident processes. Set it too high and you get a CUDA OOM, and set it too low and the space for the KV cache shrinks and concurrent throughput drops, so you must verify with actual peak load.

Prompt length changes everything

The fact that sequence length enters the earlier formula as a multiplicative term keeps coming back in practice. The decision to write a long prompt is a decision that buys accuracy and at the same time sells throughput, and if you do not know that exchange rate, capacity planning goes off every time.

Three things get worse together.

TTFT gets longer. Prefill computes attention over the whole prompt, so the cost rises steeply with length. If the system prompt is a few thousand tokens, you pay that cost every time no matter what the user asks.

Concurrency drops. The cache is proportional to length, so if the prompt doubles, only half as many fit in the same memory. The preemption seen in the earlier module starts here.

Cost rises. Whether it is a service priced per token or your own operation, the amount you read and compute becomes the bill or the hardware as it is.

So the response in practice is settled. Fix the common prefix — if you keep the system prompt and examples exactly the same instead of changing them a little for each request, the prefix cache takes effect. Even one character's difference causes everything after it to be recomputed, so if you put something like a timestamp or a request ID at the front of the prompt, the cache is invalidated entirely. Sending what changes to the back is the basic rule for keeping this cache alive.

Choose the documents to include. Instead of putting in every document fetched by retrieval, put in only the top few, and cut each document to just the necessary part. Putting in twice as many documents does not make the answer twice as good, and on the contrary, accuracy often drops when irrelevant content gets mixed in.

Put a cap on the output length. As seen earlier, the cache keeps growing as generation proceeds. With no cap, a single request keeps eating memory and takes the place of other requests. Sending work that really needs a long output to a different instance from the conversational one is the answer here too.

What to look for in the next check

First, in the quiz you distinguish the bottlenecks of TTFT and ITL and the conditions on KV cache capacity. From the next module you build a token streaming server yourself, simulate a batching scheduler, and write a KV cache calculator.