Simulating a Continuous Batching Scheduler
Goal
Implement static batching and continuous batching each as a simulation, confirm the throughput difference in numbers, and learn how to search for the maximum concurrency that satisfies a TTFT SLO.
Why it matters
In LLM generation, the output length varies from request to request. Some requests are 10 tokens, some are 1,000 tokens. If you bundle 32 with static batching, even if 31 finish early, those slots stay empty until the longest one finishes. The GPU holds the resources for 32 and processes 1. Continuous batching fills the finished slots with requests from the wait queue at every decoding iteration. It is a one-line difference in code, but throughput differs by 2–5x. This lab reproduces just that mechanism with a simulation, without a GPU. The order in the last step, step 8, is the heart of practice — if you maximize throughput first and look at latency later, you usually cannot keep the SLO. Setting the TTFT target first and finding the maximum concurrency that satisfies it is the correct order.
Steps
- Simulate static batching with
/root/lb2/static.py. Use/opt/fixtures/llms/requests.json(100 requests, each with prompt_tokens and output_tokens) for the requests, with a batch size of 16. - Write
total_steps=<n> throughput_tps=<수> slot_util=<0~1 소수>in/root/lb2/static.txt(the placeholders are a count, a number, and a decimal between 0 and 1).slot_utilmust be less than 0.7. - Simulate continuous batching with
/root/lb2/continuous.py. At every iteration, take out the finished sequences and fill from the wait queue. - Write
static_tps=<수> continuous_tps=<수> speedup=<수>in/root/lb2/compare.txt(the placeholders are numbers).speedupmust be at least 1.8. - Apply an upper limit of
max_num_seqs=32. Writemax_num_seqs=32 peak_running=<n>in/root/lb2/maxseqs.txt, andpeak_runningmust be 32 or less. - Vary the arrival rate through 5, 10, 20 and 40 requests per second, and write the header
rps,avg_queue,p99_wait_msand 4 rows in/root/lb2/queue.csv. As rps rises,p99_wait_msmust increase monotonically. /root/lb2/cost.pyseparates the costs withprefill_ms = prompt_tokens * 0.05anddecode_ms = output_tokens * 8.0. Writetotal_prefill_ms=<수> total_decode_ms=<수> decode_share=<0~1 소수>in/root/lb2/cost.txt(the placeholders are numbers and a decimal between 0 and 1).- Write
target_ttft_ms=200 max_concurrency=<n> throughput_at_target=<수>in/root/lb2/tune.txt(the placeholders are a count and a number). At that concurrency, p99 TTFT must be 200 or less.
Notes
- The heart of continuous batching is the one line that refills the slots at every iteration.
- If you preempt a sequence because the KV cache is short, you have to throw away the cache and recompute or swap, which is expensive. If preemption is frequent, throughput actually drops.
- Prefix caching reduces TTFT by up to 8x when there is a common system prompt.
- Common mistake 1: comparing the two methods on different sets of requests.
- Common mistake 2: looking at TTFT as an average and judging that the SLO is kept — you must look at p99.
Build a static batching simulator
Simulate static batching with /root/lb2/static.py. Use /opt/fixtures/llms/requests.json (100 requests, each with prompt_tokens and output_tokens) for the requests, with a batch size of 16.
Fill the batch and wait until all of it finishes. The key is that the output length differs from request to request.
Measure the static batching metrics
Write total_steps=<n> throughput_tps=<수> slot_util=<0~1 소수> in /root/lb2/static.txt (the placeholders are a count, a number, and a decimal between 0 and 1). slot_util must be less than 0.7.
Look at throughput and slot utilization together. How many slots are idle is the size of the problem.
Build a continuous batching scheduler
Simulate continuous batching with /root/lb2/continuous.py. At every iteration, take out the finished sequences and fill from the wait queue.
At every iteration, take out the finished sequences and fill from the wait queue. This one line is the whole difference.
Compare the throughput of the two methods
Write static_tps=<수> continuous_tps=<수> speedup=<수> in /root/lb2/compare.txt (the placeholders are numbers). speedup must be at least 1.8.
The comparison is meaningful only if you use the same set of requests. Compute the improvement multiple.
Apply the concurrency cap
Apply an upper limit of max_num_seqs=32. Write max_num_seqs=32 peak_running=<n> in /root/lb2/maxseqs.txt, and peak_running must be 32 or less.
You cannot put in requests without limit. The KV cache sets the upper limit.
See the relationship between queue depth and p99 wait
Vary the arrival rate through 5, 10, 20 and 40 requests per second, and write the header rps,avg_queue,p99_wait_ms and 4 rows in /root/lb2/queue.csv. As rps rises, p99_wait_ms must increase monotonically.
If you measure while raising the arrival rate, you can see at which point p99 collapses. The average looks fine for a while.
Model the prefill and decode costs separately
/root/lb2/cost.py separates the costs with prefill_ms = prompt_tokens * 0.05 and decode_ms = output_tokens * 8.0. Write total_prefill_ms=<수> total_decode_ms=<수> decode_share=<0~1 소수> in /root/lb2/cost.txt (the placeholders are numbers and a decimal between 0 and 1).
The cost functions of the two differ. Separate the one proportional to prompt length from the one proportional to the number of tokens.
Find the maximum concurrency that satisfies the TTFT target
Write target_ttft_ms=200 max_concurrency=<n> throughput_at_target=<수> in /root/lb2/tune.txt (the placeholders are a count and a number). At that concurrency, p99 TTFT must be 200 or less.
Set the target first and find the maximum that satisfies it. If you maximize throughput first, you cannot keep the SLO.