TT Lab
Get started
Learn Learning paths Courses

SSR — The Server Draws First

The Missing Time Is in the Round Trips

Continue in TT Lab

In one line

Even when the server render is fast, if the order in which you call for data is wrong, round trips keep chaining one after another after the first screen. Where you put the boundary decides that order.

Why this was needed

It is common to hear that after adding SSR the perceived speed did not improve much. You measure the server render time and it is 30ms, yet it takes 2 seconds before the user can actually use the page. The lost time is not in the render but in the round trips.

The typical shape is this. The page receives the user info, receives the team with the team id inside it, receives that team's project list, and receives the recent activity of each project. In code it is a natural nesting, but in network terms it is four round trips lined up in a row. With 200ms of latency, 800ms simply disappears.

There is an even worse form. If each component fetches its own data, the child exists only once the parent has rendered, and that request goes out only once the child exists. It is a structure in which the depth of the component tree becomes the number of round trips.

How it works

Send together what has no reason to wait for each other

The first step to eliminating a waterfall is to ask, "does this request really use the value from the previous response?" Most of the time it does not. The team info and the notification list do not need to know about each other; they were lined up only because of the shape of the code. If you send such things out all at once, the round trips shrink to one.

If you truly need the previous response, go the way of merging the requests. If you assemble it once on the server and send it down, the client has one round trip, and the queries inside the server happen within the same data center.

The boundary decides what to wait for

In React's streaming render, the part outside the <Suspense> boundaries is called the shell. The shell goes down first, and what is inside the boundaries flows out afterwards as it becomes ready (renderToPipeableStream). So putting the slow data inside a boundary is the whole of the design. If a single slow piece remains in the shell, that one piece delays the first byte entirely.

There is one thing to watch out for. A boundary is activated only when you read a data source that activates the boundary. The React documentation gives something like a promise read with use as an example, and drives home that data fetched inside an effect or an event handler does not suspend the render (Suspense). If you wrap a boundary but fetch the data in an effect, the fallback is never seen, and the HTML goes out with an empty spot drawn on the first screen — and that empty spot becomes the raw material for a hydration mismatch.

The more finely you split the boundaries, the faster the first screen gets, but the screen jolts several times. If you make the boundary large, it is quiet but late. The criterion here is not taste but whether that spot is something the user uses the moment they see it.

Cache and personalization cannot live in the same response

HTML with the data already filled in is that user's data itself. As we saw in ssr-core, if you attach a shared-cache directive to a personalized response, the CDN gives someone else's page. But it is not only the response header that is the problem. If the same address sends out different content depending on a condition, the cache has to know that condition. That is the role of Vary — a header that states which part of the request affected the response content, beyond the method and the URL, and Vary: * implies that it cannot be cached, because it means a factor outside the request headers was involved (Vary).

The practical branch is this. Send down a shell that can be cached publicly first, and receive the pieces that differ from person to person separately inside a boundary. Then the cache applies only to the shell and the personalization is confined inside the boundary — the same line as the boundary that contains hydration mismatches.

What it looks like in the field

Most reports of "the server render is 30ms but the screen takes 2 seconds" are waterfalls. The way to check is not a profiler but the waterfall chart in the network tab. If the request bars are staggered like stairs, they are lined up in a queue, and if they start side by side, they are parallel. The number of steps in the staircase is exactly the number of round trips you can eliminate.

You also often see code that wraps a boundary but fetches the data in an effect. A spinner shows up nicely on screen, so it feels as if it is working correctly, but that spinner is client state, not a streaming fallback. The HTML the server sent down has only an empty spot, and the values of the first screen still come into being only after the browser has received them.

Finally, a page that sends out different HTML depending on whether the user is logged in but does not state Vary will surely cause an accident in front of the CDN someday. This is because the cache gives everyone whichever version it received first.

Choosing the boundary boils down, in the end, to one sentence. What the user uses the moment they see it goes in the shell; what can wait a little goes inside a boundary. Once you draw that line, what the first byte is tied to, how far you can cache publicly, and how far a mismatch spreads are all decided at once. That all three sit on the same line is the fact you realize last in SSR design.

What to check in the next quiz

Across six questions, you check how to recognize a waterfall, what the shell is, what activates a boundary, and why a personalized response needs Vary.