TT Lab
Get started
Learn Learning paths Courses

LLM Serving

Implementing a Token Streaming Server Yourself

Continue in TT Lab

Goal

Build a server that streams tokens over SSE yourself, measure TTFT and the inter-token latency separately, and prove that generation actually stops when the client aborts.

Why it matters

Even if the total generation time is the same 6 seconds, getting it all at once after 6 seconds and having it flow out from 200ms are completely different products. That is why streaming in an LLM service is the default, and from that moment the way of measuring changes too — instead of a single response completion time, you must measure the time to the first chunk and the interval between chunks separately. What is especially easy to miss in this lab is step 7. It is very common for users to move on to another question in the middle of a long response, and without cancellation handling those requests keep occupying batch slots and produce tokens that no one is watching. Concurrent throughput drops for no reason and only the GPU bill grows. It is a quiet and expensive incident, and that is why a whole step is assigned to it.

Steps

  1. Create generate(prompt, n) in /root/ls/gen.py. The same prompt and n must always return the same sequence of tokens. Create /root/ls/gen1.txt and /root/ls/gen2.txt with the same arguments, and the contents of the two files must be the same.
  2. Start /root/ls/server.py on 127.0.0.1:8170. POST /generate takes {"prompt":"hello","max_tokens":20} and returns {"text":"...","usage":{"prompt_tokens":<n>,"completion_tokens":20}}.
  3. GET /stream?prompt=hello&max_tokens=20 responds with Content-Type: text/event-stream and sends each token as data: {"token":"..."} followed by a blank line. The response body must have at least 20 lines starting with data: .
  4. Send data: [DONE] at the end of the stream.
  5. Measure the time to the first chunk with /root/ls/ttft.py and write ttft_ms=<정수> in /root/ls/ttft.txt (the placeholder is an integer). It must be greater than 0 and less than 3000.
  6. Write count=<n> p50_ms=<수> p99_ms=<수> in /root/ls/itl.txt (the placeholders are a count and numbers). count must be at least 19.
  7. Open a stream with max_tokens=100, receive only 5 chunks, and then disconnect. Make GET /_debug/generated report the number of tokens the server actually produced for the immediately preceding stream request (it is convenient to also include a cumulative total), and write that value in /root/ls/cancel.txt as requested=100 generated=<n>. generated must be 30 or less.
  8. Write four lines in /root/ls/report.txt: ttft_ms=<n>, itl_p50_ms=<수>, tokens=<n> and throughput_tps=<수> (the placeholders are numbers).

Notes

Build a deterministic token generator

Create generate(prompt, n) in /root/ls/gen.py. The same prompt and n must always return the same sequence of tokens. Create /root/ls/gen1.txt and /root/ls/gen2.txt with the same arguments, and the contents of the two files must be the same.

The same prompt must always produce the same sequence of tokens for grading and evaluation to be possible. You can use a hash as the seed.

Build a non-streaming endpoint

Start /root/ls/server.py on 127.0.0.1:8170. POST /generate takes {"prompt":"hello","max_tokens":20} and returns {"text":"...","usage":{"prompt_tokens":<n>,"completion_tokens":20}}.

First build the form that returns everything at once to set a baseline. Return the number of tokens used as well.

Send chunks over SSE

GET /stream?prompt=hello&max_tokens=20 responds with Content-Type: text/event-stream and sends each token as data: {"token":"..."} followed by a blank line. The response body must have at least 20 lines starting with data: .

The content type and the event separation rule are fixed. If you leave out the blank line between events, the client does not recognize them.

Send the termination signal

Send data: [DONE] at the end of the stream.

There is a conventional marker that tells you when the stream has ended.

Measure the time to the first token

Measure the time to the first chunk with /root/ls/ttft.py and write ttft_ms=<정수> in /root/ls/ttft.txt (the placeholder is an integer). It must be greater than 0 and less than 3000.

It is between the start of the request and the arrival of the first chunk. Do not confuse it with the total completion time.

Produce the inter-token latency distribution

Write count=<n> p50_ms=<수> p99_ms=<수> in /root/ls/itl.txt (the placeholders are a count and numbers). count must be at least 19.

Record the arrival time of each chunk and take the differences. The average alone is not enough.

Stop generation when the client aborts

Open a stream with max_tokens=100, receive only 5 chunks, and then disconnect. Make GET /_debug/generated report the number of tokens the server actually produced for the immediately preceding stream request (it is convenient to also include a cumulative total), and write that value in /root/ls/cancel.txt as requested=100 generated=<n>. generated must be 30 or less.

You can stop only if you detect the disconnection. Prove it by counting the tokens the server produced.

Build a metrics report

Write four lines in /root/ls/report.txt: ttft_ms=<n>, itl_p50_ms=<수>, tokens=<n> and throughput_tps=<수> (the placeholders are numbers).

Gather the values you measured earlier into one file. Throughput is the number of generated tokens divided by the total time.