TT Lab
Get started
Learn Learning paths Courses

LLM Serving

Fix the Measurement Contract Before the Tool Name

Continue in TT Lab

In one line

There is no universal winner among model servers. You must measure with the same model, hardware and workload, and compare the features you need and the operating cost together.

Why this was needed

To the question "which server should I use?", you cannot answer with a single benchmark table. That is because the same tool can be the best for one workload and unsuitable for another.

You must set the criteria first. Write down whether it is conversational or batch, what the input and output length distributions are, whether prompts have many common prefixes, what the concurrency peak is, whether you need structured output or a specific quantization, and whether you can bear the cost of engine builds and version upgrades.

How it works

To make the comparison reproducible, first fix the measurement contract.

  1. Record not just the model name but the exact revision, the tokenizer, the dtype and quantization method, and the server version.
  2. Keep the same GPU model and count, driver and runtime versions, memory limit and container image.
  3. Extract the input and output token lengths, the common prefix ratio, and the concurrency and arrival intervals from real traffic, and replay them.
  4. After warmup, repeat several times and record together the p50/p95/p99 of TTFT, ITL and end-to-end latency, the number of tokens processed, the error rate, GPU memory and the preemption rate.
  5. Check that the quality settings are the same. If only one candidate uses more aggressive quantization or a shorter maximum output, it is not a speed comparison.

Then you look at the features and the operational constraints. vLLM's block-based KV cache management, SGLang's prefix reuse feature, TensorRT-LLM's pre-built engines, TGI's supported models and deployment integration, and Ollama's convenience of running locally are each only clues for narrowing down candidates and do not guarantee a performance ranking. Even with the same feature name, the gain varies with the release, the model and the request distribution, so decide with the documentation and measurement results of the combination you will actually use.

What it looks like in the field

What matters more than the choice is being able to reproduce the result. In a benchmark report, leave the run commands, the image and model revisions, how the workload data was generated, the concurrency, the warmup and repetition counts, and the raw results. Even if only the average throughput is good, if the p99 TTFT or the error rate exceeds the product SLO, it cannot be adopted.

And there is a decision that comes before the server choice. The model size and quantization. Changing only the server on a GPU that cannot hold 7B fp16 weights is meaningless. If you quantize the weights to 4 bits with AWQ or GPTQ, the weight memory is in theory about a quarter of fp16, but that does not make the total GPU memory, including the KV cache, the runtime workspace and the quantization metadata, a quarter. The actual savings and quality loss must be measured with a supported model and server combination.

What a server actually does

Even though the names of vLLM, TGI and TensorRT-LLM differ, the mechanisms that produce performance are largely the same. If you know them, you know what is being adjusted even when the setting names differ.

Continuous batching. It does not wait until the requests finish, and at every step takes out the finished ones and puts in new ones. Throughput rises several times over static batching. Whether this is turned on is the first thing to check.

PagedAttention. It manages the KV cache in units of pages to eliminate fragmentation. In the past, it reserved the maximum length in advance and wasted two or three times what was actually used.

Prefix cache. It reuses the KV of requests whose front part is the same, like a system prompt. It reduces TTFT with no effect on accuracy, so it is the first thing to turn on.

Quantization. It reduces the weights to 4–8 bits to save memory and bandwidth. There is some quality loss, so you check with an evaluation set before turning it on.

How to split memory

GPU memory splits into three shares.

전체 24GB
 ├ 가중치         16.2GB   ← 모델 크기 × 비트수/8
 ├ KV 캐시         6.5GB   ← 나머지의 대부분. 동시 요청 수를 정한다
 └ 활성화·여유     1.3GB

Raising gpu_memory_utilization enlarges the KV cache and raises concurrent throughput, but if you raise it too much, there is not enough activation headroom and an OOM occurs. 0.90–0.93 is the practical range.

The KV cache size is proportional to the product of the number of concurrent requests and the context length.

KV 바이트 ≈ 2 × 레이어 × 헤드 × 헤드차원 × 문맥길이 × 배치 × 정밀도바이트

So if you extend the context from 16K to 32K, the number of requests that can be handled concurrently halves. "Let's keep a long context open" is not free.

What to measure and how

Before choosing a tool, you fix how to measure. Otherwise the comparison does not hold.

Value Definition Pitfall
TTFT Request → first token It improves dramatically when the prefix cache is on
TPOT Average interval between tokens It gets worse when the batch is large
Throughput Output tokens per second (total) It is meaningful only if you also write the number of concurrent requests
p95 latency The top 5% You miss it if you look only at the average

The third one is important. "138 tok/s" is a completely different number depending on whether it is the sum of 4 concurrent requests or a single request. When writing down a benchmark, write the concurrency, the input length and the output length together. Without these three, it cannot be reproduced.

What to look for in the next check

After confirming in the quiz the criteria for choosing a server that fits the workload conditions, in the next module you build yourself, as a calculator, the GPU memory estimation that is the basis of that judgment.