TT Lab
Get started
Learn Learning paths Courses

LLM Serving

Why Streaming Changes the UX

Continue in TT Lab

In one line

If you finish generating a 300-token response and send it all at once, the user waits 6 seconds at a blank screen. If you stream it, characters appear after 200ms. The total time is the same.

Why this was needed

There are few areas where perceived performance and actual performance part ways this much. Even if the total generation time is the same 6 seconds, getting it all at once after 6 seconds and having it flow out from 200ms are completely different products. In the latter, the user can start reading, and if the direction is wrong, can stop midway.

That is why streaming in an LLM service is not an option but the default.

How it works

The transport is one of three. WebSocket is bidirectional but heavy, long polling is outdated, and SSE (Server-Sent Events) fits this purpose best. It is one-way, works on top of HTTP, and has built-in browser support.

The SSE format is simple. You respond with Content-Type: text/event-stream and send each event separated by one data: <내용> line (the placeholder is the content) and a blank line.

data: {"token":"안녕"}

data: {"token":"하세요"}

data: [DONE]

It matters that the blank line is the event separator. If you leave it out, the client does not recognize the event and piles it up in the buffer.

There are four things you must take care of in the implementation.

First, turning off buffering. If the server framework or a proxy in front buffers the response, streaming becomes meaningless. If you put it behind NGINX, you need the X-Accel-Buffering: no header.

Second, the termination signal. You have to tell the client when the stream has ended. You use a conventional marker such as [DONE].

Third, cancellation handling. When the user closes the browser tab, the server must stop generating. If you do not, the GPU keeps producing tokens that no one is watching. The cost goes out just the same.

Fourth, delivering errors. If an error occurs in the middle of a stream, you cannot change the HTTP status code. That is because you have already sent the headers with 200. So errors must also be sent as events.

Things in the middle block the stream

Even when SSE is implemented properly, it is common for everything to arrive in the browser all at once. The cause is not the application but the things in between.

What blocks it Symptom How to fix
nginx proxy buffer Arrives lumped in 4KB pieces proxy_buffering off or X-Accel-Buffering: no on the response
Compression middleware Nothing comes out, then everything at the end at once Exclude only this path from gzip
Framework buffer The first piece is late Use the framework's streaming response type
CDN Receives the whole thing in order to cache it Take this path out of the cached targets

The most common is the first row. nginx buffers the upstream response by default, so if you put it in front without any settings, streaming is disabled entirely. One line in the response header (X-Accel-Buffering: no) can turn it off for just that request, so this method is safe.

Notice a dropped connection to save the GPU

Even if the user closes the tab, the server keeps producing tokens. The GPU generates for 6 seconds an answer that no one is watching. In a service with many concurrent requests, this waste is large.

So you detect the client disconnecting and stop generation. With FastAPI, you check await request.is_disconnected() for every token, and in vLLM, if you cancel the request, it takes that sequence out of the batch.

async def stream(request, prompt):
    async for chunk in llm.generate(prompt):
        if await request.is_disconnected():
            await llm.abort(chunk.request_id)   # GPU 를 놓아준다
            return
        yield f"data: {json.dumps(chunk.model_dump())}

"
    yield "data: [DONE]

"

Delivering errors inside the stream

This is the awkward part of streaming. If it fails after you have already sent 200 OK and the headers, you cannot change the status code. So you send errors as events too.

data: {"token":"안녕하"}

event: error
data: {"code":"context_length","message":"입력이 너무 깁니다"}

The client must also handle the case where the stream ends without [DONE]. That is usually the network being cut, and it is better to keep what has already been received and show "the response was cut off midway" than to show nothing at all.

What to measure

A streaming service has two metrics.

The total generation time is only a result of these two, so it is not made a target. Raising the batch size increases overall throughput but makes TTFT worse — throughput and perceived speed are a trade-off, and the product decides which to favor.

What it looks like in the field

Once you turn on streaming, the way of measuring changes too. What used to be only the response completion time must be measured split into the first chunk time and the interval between chunks. And since these values swing greatly depending on the batching situation, you must look at p99.

Incidents that miss cancellation handling are quiet and expensive. It is common for users to move on to another question in the middle of a long response, and without cancellation those requests keep occupying batch slots. Concurrent throughput drops for no reason, and only the GPU bill grows.

What you will do in the next lab

You build a deterministic token generator and implement a server that streams it over SSE yourself. You measure TTFT and the interval between chunks, and prove with a counter that generation actually stops when canceled.