Implementing a Token Streaming Server Yourself
Goal
Build a server that streams tokens over SSE yourself, measure TTFT and the inter-token latency separately, and prove that generation actually stops when the client aborts.
Why it matters
Even if the total generation time is the same 6 seconds, getting it all at once after 6 seconds and having it flow out from 200ms are completely different products. That is why streaming in an LLM service is the default, and from that moment the way of measuring changes too — instead of a single response completion time, you must measure the time to the first chunk and the interval between chunks separately. What is especially easy to miss in this lab is step 7. It is very common for users to move on to another question in the middle of a long response, and without cancellation handling those requests keep occupying batch slots and produce tokens that no one is watching. Concurrent throughput drops for no reason and only the GPU bill grows. It is a quiet and expensive incident, and that is why a whole step is assigned to it.
Steps
- Create
generate(prompt, n)in/root/ls/gen.py. The same prompt and n must always return the same sequence of tokens. Create/root/ls/gen1.txtand/root/ls/gen2.txtwith the same arguments, and the contents of the two files must be the same. - Start
/root/ls/server.pyon 127.0.0.1:8170.POST /generatetakes{"prompt":"hello","max_tokens":20}and returns{"text":"...","usage":{"prompt_tokens":<n>,"completion_tokens":20}}. GET /stream?prompt=hello&max_tokens=20responds withContent-Type: text/event-streamand sends each token asdata: {"token":"..."}followed by a blank line. The response body must have at least 20 lines starting withdata:.- Send
data: [DONE]at the end of the stream. - Measure the time to the first chunk with
/root/ls/ttft.pyand writettft_ms=<정수>in/root/ls/ttft.txt(the placeholder is an integer). It must be greater than 0 and less than 3000. - Write
count=<n> p50_ms=<수> p99_ms=<수>in/root/ls/itl.txt(the placeholders are a count and numbers). count must be at least 19. - Open a stream with
max_tokens=100, receive only 5 chunks, and then disconnect. MakeGET /_debug/generatedreport the number of tokens the server actually produced for the immediately preceding stream request (it is convenient to also include a cumulative total), and write that value in/root/ls/cancel.txtasrequested=100 generated=<n>. generated must be 30 or less. - Write four lines in
/root/ls/report.txt:ttft_ms=<n>,itl_p50_ms=<수>,tokens=<n>andthroughput_tps=<수>(the placeholders are numbers).
Notes
- An SSE event is always followed by a blank line after the
data: <내용>line (the placeholder is the content). - You must turn off buffering in the framework or proxy for streaming to actually flow.
- An error in the middle of a stream cannot be reported by a status code — since 200 has already been sent, errors are also sent as events.
- Common mistake 1: leaving out the blank line between chunks so that the client cannot read anything.
- Common mistake 2: not detecting cancellation and so continuing to produce tokens for a dropped request.
Build a deterministic token generator
Create generate(prompt, n) in /root/ls/gen.py. The same prompt and n must always return the same sequence of tokens. Create /root/ls/gen1.txt and /root/ls/gen2.txt with the same arguments, and the contents of the two files must be the same.
The same prompt must always produce the same sequence of tokens for grading and evaluation to be possible. You can use a hash as the seed.
Build a non-streaming endpoint
Start /root/ls/server.py on 127.0.0.1:8170. POST /generate takes {"prompt":"hello","max_tokens":20} and returns {"text":"...","usage":{"prompt_tokens":<n>,"completion_tokens":20}}.
First build the form that returns everything at once to set a baseline. Return the number of tokens used as well.
Send chunks over SSE
GET /stream?prompt=hello&max_tokens=20 responds with Content-Type: text/event-stream and sends each token as data: {"token":"..."} followed by a blank line. The response body must have at least 20 lines starting with data: .
The content type and the event separation rule are fixed. If you leave out the blank line between events, the client does not recognize them.
Send the termination signal
Send data: [DONE] at the end of the stream.
There is a conventional marker that tells you when the stream has ended.
Measure the time to the first token
Measure the time to the first chunk with /root/ls/ttft.py and write ttft_ms=<정수> in /root/ls/ttft.txt (the placeholder is an integer). It must be greater than 0 and less than 3000.
It is between the start of the request and the arrival of the first chunk. Do not confuse it with the total completion time.
Produce the inter-token latency distribution
Write count=<n> p50_ms=<수> p99_ms=<수> in /root/ls/itl.txt (the placeholders are a count and numbers). count must be at least 19.
Record the arrival time of each chunk and take the differences. The average alone is not enough.
Stop generation when the client aborts
Open a stream with max_tokens=100, receive only 5 chunks, and then disconnect. Make GET /_debug/generated report the number of tokens the server actually produced for the immediately preceding stream request (it is convenient to also include a cumulative total), and write that value in /root/ls/cancel.txt as requested=100 generated=<n>. generated must be 30 or less.
You can stop only if you detect the disconnection. Prove it by counting the tokens the server produced.
Build a metrics report
Write four lines in /root/ls/report.txt: ttft_ms=<n>, itl_p50_ms=<수>, tokens=<n> and throughput_tps=<수> (the placeholders are numbers).
Gather the values you measured earlier into one file. Throughput is the number of generated tokens divided by the total time.