TT Lab
Get started
Learn Learning paths Courses

LLM Serving

Simulating a Continuous Batching Scheduler

Continue in TT Lab

Goal

Implement static batching and continuous batching each as a simulation, confirm the throughput difference in numbers, and learn how to search for the maximum concurrency that satisfies a TTFT SLO.

Why it matters

In LLM generation, the output length varies from request to request. Some requests are 10 tokens, some are 1,000 tokens. If you bundle 32 with static batching, even if 31 finish early, those slots stay empty until the longest one finishes. The GPU holds the resources for 32 and processes 1. Continuous batching fills the finished slots with requests from the wait queue at every decoding iteration. It is a one-line difference in code, but throughput differs by 2–5x. This lab reproduces just that mechanism with a simulation, without a GPU. The order in the last step, step 8, is the heart of practice — if you maximize throughput first and look at latency later, you usually cannot keep the SLO. Setting the TTFT target first and finding the maximum concurrency that satisfies it is the correct order.

Steps

  1. Simulate static batching with /root/lb2/static.py. Use /opt/fixtures/llms/requests.json (100 requests, each with prompt_tokens and output_tokens) for the requests, with a batch size of 16.
  2. Write total_steps=<n> throughput_tps=<수> slot_util=<0~1 소수> in /root/lb2/static.txt (the placeholders are a count, a number, and a decimal between 0 and 1). slot_util must be less than 0.7.
  3. Simulate continuous batching with /root/lb2/continuous.py. At every iteration, take out the finished sequences and fill from the wait queue.
  4. Write static_tps=<수> continuous_tps=<수> speedup=<수> in /root/lb2/compare.txt (the placeholders are numbers). speedup must be at least 1.8.
  5. Apply an upper limit of max_num_seqs=32. Write max_num_seqs=32 peak_running=<n> in /root/lb2/maxseqs.txt, and peak_running must be 32 or less.
  6. Vary the arrival rate through 5, 10, 20 and 40 requests per second, and write the header rps,avg_queue,p99_wait_ms and 4 rows in /root/lb2/queue.csv. As rps rises, p99_wait_ms must increase monotonically.
  7. /root/lb2/cost.py separates the costs with prefill_ms = prompt_tokens * 0.05 and decode_ms = output_tokens * 8.0. Write total_prefill_ms=<수> total_decode_ms=<수> decode_share=<0~1 소수> in /root/lb2/cost.txt (the placeholders are numbers and a decimal between 0 and 1).
  8. Write target_ttft_ms=200 max_concurrency=<n> throughput_at_target=<수> in /root/lb2/tune.txt (the placeholders are a count and a number). At that concurrency, p99 TTFT must be 200 or less.

Notes

Build a static batching simulator

Simulate static batching with /root/lb2/static.py. Use /opt/fixtures/llms/requests.json (100 requests, each with prompt_tokens and output_tokens) for the requests, with a batch size of 16.

Fill the batch and wait until all of it finishes. The key is that the output length differs from request to request.

Measure the static batching metrics

Write total_steps=<n> throughput_tps=<수> slot_util=<0~1 소수> in /root/lb2/static.txt (the placeholders are a count, a number, and a decimal between 0 and 1). slot_util must be less than 0.7.

Look at throughput and slot utilization together. How many slots are idle is the size of the problem.

Build a continuous batching scheduler

Simulate continuous batching with /root/lb2/continuous.py. At every iteration, take out the finished sequences and fill from the wait queue.

At every iteration, take out the finished sequences and fill from the wait queue. This one line is the whole difference.

Compare the throughput of the two methods

Write static_tps=<수> continuous_tps=<수> speedup=<수> in /root/lb2/compare.txt (the placeholders are numbers). speedup must be at least 1.8.

The comparison is meaningful only if you use the same set of requests. Compute the improvement multiple.

Apply the concurrency cap

Apply an upper limit of max_num_seqs=32. Write max_num_seqs=32 peak_running=<n> in /root/lb2/maxseqs.txt, and peak_running must be 32 or less.

You cannot put in requests without limit. The KV cache sets the upper limit.

See the relationship between queue depth and p99 wait

Vary the arrival rate through 5, 10, 20 and 40 requests per second, and write the header rps,avg_queue,p99_wait_ms and 4 rows in /root/lb2/queue.csv. As rps rises, p99_wait_ms must increase monotonically.

If you measure while raising the arrival rate, you can see at which point p99 collapses. The average looks fine for a while.

Model the prefill and decode costs separately

/root/lb2/cost.py separates the costs with prefill_ms = prompt_tokens * 0.05 and decode_ms = output_tokens * 8.0. Write total_prefill_ms=<수> total_decode_ms=<수> decode_share=<0~1 소수> in /root/lb2/cost.txt (the placeholders are numbers and a decimal between 0 and 1).

The cost functions of the two differ. Separate the one proportional to prompt length from the one proportional to the number of tokens.

Find the maximum concurrency that satisfies the TTFT target

Write target_ttft_ms=200 max_concurrency=<n> throughput_at_target=<수> in /root/lb2/tune.txt (the placeholders are a count and a number). At that concurrency, p99 TTFT must be 200 or less.

Set the target first and find the maximum that satisfies it. If you maximize throughput first, you cannot keep the SLO.