TT Lab
Get started
Learn Learning paths Courses

LLM Serving

What Static Batching Throws Away

Continue in TT Lab

In one line

Static batching leaves the remaining slots idle until the longest request in the batch finishes. Continuous batching immediately fills a finished slot with a new request. That difference is a 2–5x throughput gain.

Why this was needed

In work like image classification, batching is simple. The input sizes are the same and so are the processing times, so you can gather 32 and run them at once.

LLM generation is different. The output length of each request varies. Some requests are 10 tokens, some are 1,000 tokens. If you bundle 32 with static batching, even if 31 finish at 10 tokens, those 31 slots stay empty until the 1,000-token one finishes. The GPU holds the resources for 32 and processes 1.

How it works

Continuous batching (also called in-flight batching) sees the batch not as a fixed bundle but as a fluid set of slots. At every decoding iteration the scheduler checks — if there is a finished sequence, it takes it out of its slot, and if there are requests in the wait queue, it puts them into that slot.

Then the GPU always keeps a batch close to the maximum. In benchmarks, throughput is reported to improve 2–5x over static batching.

There are two values you can adjust here. max_num_seqs is the upper limit on the number of sequences processed concurrently, and max_num_batched_tokens is the upper limit on the number of tokens processed in one iteration. The former is tied to KV cache memory and the latter to compute capability.

There is a trade-off. Raising concurrency raises throughput but increases the latency of individual requests. And when the KV cache runs short, the scheduler preempts some sequences, sends them out and puts them back in later. At that point it either throws away that sequence's KV cache and recomputes it or swaps it to the CPU, and both are expensive. If preemption happens often, throughput actually drops.

Another candidate is prefix caching. If several requests start with the same system prompt, you can reuse the KV cache of that part. The gain varies with the common prefix length, the reuse rate, and the implementation and memory policy, and if the prompts are completely different every time, there is almost no effect. You must measure the cache hit rate and TTFT together on the actual request distribution.

What it looks like in the field

Looking at the relationship between queue depth and tail TTFT builds operational intuition. As the inflow approaches the processing capacity, the wait queue grows and high-percentile latency gets worse before the average does. You choose a tail percentile such as p95 or p99 as the SLO metric to match the product's user experience and traffic scale.

And there is an order in setting the targets. First set a TTFT SLO that fits the product and workload (this lab uses 200ms as an example for the calculation), find the maximum concurrency that satisfies it, and use the throughput at that concurrency as the basis for capacity planning. If you maximize throughput first, it is easy to miss the tail latency SLO.

The KV cache is the real limit line

If you ask what limits the concurrency of continuous batching, people usually answer compute capability, but what actually runs out first is almost always memory. A sequence being generated must hold the keys and values for all the tokens so far, and that size keeps growing in proportion to the number of tokens. When one request reaches 4,000 tokens, that request's cache is several times what it was at the start.

So the calculation for deciding the batch size goes like this. What remains after the model weights is the share available for the cache, and dividing that by the cache size one request uses on average gives the number of sequences that can be held concurrently. Here you must look not at the average but at the long end. Even if most requests are short, if a few long requests occupy the cache, the slots shrink by that much.

When the cache runs short, the scheduler preempts sequences and sends them out, and this moment is the most dangerous point in operation. A sequence that was sent out comes back later and is recomputed from scratch, so it does again what it has already done. The cliff where throughput suddenly breaks down with only a slight rise in load is born here. There are three metrics to watch.

Metric What it tells you Bad sign
KV cache utilization Memory headroom Stays near 90% for a long time
Number of preemptions The amount of work undone If not 0, you have already gone past the limit
Wait queue length The amount accepted but not done If it keeps growing, you must block inflow

The response has three branches. Limit the maximum output length per request so the cache cannot grow without bound, lower max_num_seqs to a value at which preemption does not occur, and if that is still not enough, add instances. The choice to put up with preemption and keep concurrency high is almost always a loss. It looks on the surface as though you accept more, but in reality you are doing the same computation twice.

There is one more thing to point out. If you mix long requests and short requests on the same instance, while a long request occupies the cache, the latency of short requests gets worse along with it. Separating conversational responses from long generation of a batch nature onto different instances is not a waste of resources but the surest way to protect tail latency.

What you will do in the next lab

You implement static batching and continuous batching schedulers each as a simulation and compare their throughput. You apply a concurrency cap, observe the relationship between queue depth and p99 wait, model the costs of prefill and decode separately, and then search for the maximum concurrency that satisfies the lab's example SLO of TTFT 200ms.