TT Lab
Get started
Learn Learning paths Courses

SSE — How the Server Speaks First

What Kills a Stream Sits in the Middle

Continue in TT Lab

In one line

The place where streaming dies is usually not the application but the things in between — proxy buffers, read timeouts, and the browser's concurrent connection limit.

Why this was needed

Locally, tokens flow one character at a time, but once deployed the answer comes out all at once after it is finished. Not one line of application code differs. Even when you check the logs, they say the server sent on time.

That's because there is a proxy in the middle. nginx's proxy_buffering is on by default according to the docs, and when it is on, the response is collected in a buffer before being sent out. Streaming is all about "a little, often," while the purpose of buffering is "collect and send at once," so they are exactly opposite.

The second thing you meet often is the timeout. In the same doc, the default of proxy_read_timeout is 60s, and the description is clear — if the proxied server sends nothing within that time, the connection is closed. This is why a notification stream drops every minute during night hours when users ask no questions.

How it works

Two places to turn off buffering

You can turn it off in the configuration file with proxy_buffering off, but then every response on that path loses buffering. The same doc describes a narrower method — turning it off with a response header.

X-Accel-Buffering: no

If you send yes or no in this header, buffering is turned on or off only for that response. If you attach this header only on streaming endpoints, the other paths keep the benefits of buffering. It is a design that lets the application, which knows its own situation best, make the decision, so you can use it even when you don't have permission to edit the configuration file.

Note that this speaks only to nginx. If there is a CDN or another proxy in front, you have to say the same thing to that one separately. If you release one layer and move on saying "fixed," it stays blocked in another layer.

A heartbeat is not politeness but a requirement

The way to avoid the read timeout is to eliminate quiet stretches. The comment line defined by the specification fits that spot exactly — a line starting with a colon is ignored, so bytes flow with no effect on the client side.

: keep-alive

The interval must be comfortably shorter than what the proxy tolerates. Where the default is 60 seconds, a 55-second interval gets cut by a single bout of jitter. In practice, use half of that or less.

A heartbeat has a second use. It lets the client side measure "since when has nothing arrived." In a stream where events are rare by nature, there is no way to tell whether the quiet is normal or the stream is dead, but with a heartbeat that judgment becomes possible.

The browser's concurrent connection limit

Under HTTP/1.1, browsers limit the number of connections they open to one origin. MDN says the default that was once between 2 and 3 is now commonly 6 (HTTP/1.x connection management). An SSE connection is a response that never ends, so it keeps occupying its slot. Open six tabs and all six slots fill up, and not only the seventh tab but even ordinary API requests to that origin queue up. The screen looks frozen, yet the requests never even arrive in the server log, so finding the cause is especially hard.

There are two ways to solve it. One is to move up to HTTP/2 — it multiplexes several streams over a single TCP connection (RFC 9113), so it turns into a one-connection problem. The other is not to connect per tab. A shared worker (SharedWorker) holds one connection and shares it with the other tabs of the same origin via BroadcastChannel.

What it looks like in the field

A report that "it works locally, but only on staging everything comes out at once" is almost always buffering. The way to confirm is not the application log but the time to first byte. If you measure the first-byte time of the streaming path and of the path that returns everything at once side by side, it settles immediately — if the total time is the same and only the first byte differs, something in the middle is collecting it.

A notification stream that drops only at night is a case with no heartbeat or an interval too close to the timeout. It looks hard to reproduce, but the very fact that it happens only in quiet hours is already the clue that "something isn't being sent."

The concurrent connection limit is almost never caught in QA, because tests are done with a single tab. The report always comes in as "it gets slow when I keep several tabs open," and the moment you read that sentence as a performance problem, days disappear.

What to check in the next quiz

Six questions check the place where buffering is turned off with a response header, what the heartbeat interval is set against, and what symptom the concurrent connection limit shows up as.