Building the Memory Budget by Hand
In one line
The GPU memory budget has three items. The model weights, the KV cache, and the activations and overhead. The second is the largest and the most often ignored.
Why this was needed
To answer the question "Does a 70B model fit on a single A100 80GB?", you need a calculation. Looking at the weights alone, at fp16 it is 140GB, so it does not fit. If you quantize to 4 bits, it is about 35GB, so it seems it would fit. But the KV cache has been left out.
If you start without calculating, you meet an OOM on deployment day. And since an OOM occurs when the load piles up, it occurs at the worst possible time.
How it works
The model weights are the number of parameters times the dtype bytes. For 7B at fp16 it is 14GB, and for 70B at fp16 it is 140GB. Quantizing reduces this value — INT8 is half and INT4 is about a quarter.
The KV cache formula is this.
KV 바이트 = 2 x 레이어 수 x 은닉 차원 x 시퀀스 길이 x 배치 x dtype 바이트
The 2 at the front means the two items, key and value. Calculated for 7B (32 layers, hidden size 4096) at fp16, it goes like this.
| Context | Batch 1 | Batch 8 |
|---|---|---|
| 4K | 2 GiB | 16 GiB |
| 32K | 16 GiB | 128 GiB |
| 128K | 64 GiB | 512 GiB |
With a 128K context and batch 8, it is 512GiB. To run a model with 14GB of weights, the KV cache is 512GB. Why long context is hard is in this one table.
There are three ways to reduce it. GQA shares key/value heads in groups, reducing it to about one eighth with 8 groups. KV cache quantization reduces the dtype bytes, to half with INT8 and a quarter with INT4. PagedAttention does not reduce the size itself but eliminates fragmentation and raises effective efficiency.
Stacking these yields a large saving. In the author's example, the baseline of 1,280GB becomes 320GB with GQA-8, 160GB with INT8, about 128GB on an effective PagedAttention basis, and 80GB if you go as far as INT4.
What it looks like in the field
The calculation that decides the number of concurrent requests is what you need most often in practice. Subtracting the weights and overhead from the total GPU memory gives the KV cache budget, and dividing that by the KV size per request gives the maximum number of concurrent requests. This value is the upper limit of max_num_seqs and the starting point of capacity planning.
The choice of quantization also comes from this calculation. GPTQ and AWQ quantize only the weights, and KV cache quantization is a separate setting. If you confuse the two, you get "I switched to 4 bits but memory hardly shrank" — because in workloads with a long context the KV cache dominates.
The formula for building the budget
GPU memory splits into three shares. If you calculate in order, you get the number of requests that can be handled concurrently.
1. 가중치 = 파라미터 수 × (비트수 ÷ 8)
7B × 2바이트(FP16) = 14.0 GB
7B × 0.5바이트(INT4) = 3.5 GB
2. 여유 = 활성화 + 단편화 ≈ 전체의 5~10%
3. KV 캐시로 쓸 수 있는 몫 = 전체 − 가중치 − 여유
The size one token of KV cache takes is this.
토큰당 = 2(K와 V) × 레이어 수 × KV 헤드 수 × 헤드 차원 × 정밀도 바이트
예: Llama-3 8B (32레이어, KV헤드 8, 헤드차원 128, FP16)
= 2 × 32 × 8 × 128 × 2 = 131,072 바이트 ≈ 128 KB/토큰
GQA (Grouped-Query Attention) makes a big difference here. With 32 KV heads it is 512KB/token, but with 8 it is 128KB. This is why recent models use GQA.
24GB 카드, 8B 모델 FP16:
가중치 16GB + 여유 2GB → KV 로 6GB
6GB ÷ 128KB = 약 49,000 토큰
→ 문맥 4K 면 동시 12요청, 문맥 8K 면 6요청
What this calculation tells you
- If you open the context to twice the length, concurrent throughput halves. "Just in case, let's open it at 32K" is a decision that makes throughput one eighth.
- Quantization reduces only the weights. If you reduce the weights to 3.5GB with INT4, the share available for KV grows a lot and concurrent throughput becomes several times higher. This gain is often larger than the quality loss.
- The KV cache can be quantized too. An FP8 KV cache halves the capacity, and there are many reports that the effect on quality is smaller than with weight quantization.
What happens when it overflows
When KV runs short, vLLM does preemption. It picks one in-progress request, discards its cache, and recomputes it from scratch later.
로그: Sequence group ... is preempted by PreemptionMode.RECOMPUTE
If you see this log often, concurrency has exceeded capacity. Latency spikes while GPU utilization comes out high, so it is easy to misread as "it's being used well". Making the number of preemptions a metric is how to notice this state.
What you will do in the next lab
You implement the formula in code and calculate several scenarios. You verify that 7B at 4K is 2GiB, calculate 128K, apply GQA and INT8, and finally find the maximum number of concurrent requests possible on an 80GB GPU.