Real-Time Communication — WebSocket, gRPC Streaming and WebRTC
SSE, WebSocket, gRPC streaming or WebRTC?
In one line
If only the server talks, use SSE; if both sides talk whenever they like, WebSocket; if the contract between services matters, gRPC streaming; and to carry a person's voice within tens of ms, WebRTC.
Why this was needed
The first decision you make when building a real-time feature is the transport, and once you pick it, it is hard to change. That is because the client code, the authentication method, the proxy configuration, and the observability tools are all tied to that choice. Yet this decision is often made for reasons like "everyone seems to use WebSocket these days." If you lay side by side what each transport gives you for free and what it makes you build yourself, you get grounds for the choice.
How it works
Server-Sent Events in the HTML standard is a method of keeping an ordinary HTTP response open and continuing to write to it as text/event-stream. It is one direction, from server to client, and carries only text. In return, the browser's EventSource reconnects by itself when cut, and the standard includes a resume rule that reports the last received event number through the Last-Event-ID header. Being ordinary HTTP, it passes through proxies and authentication as they are. LLM token streaming APIs often use this method. However, EventSource cannot attach custom request headers, and over HTTP/1.1 it runs into the browser's per-domain connection limit (usually 6). Over HTTP/2 the streams are overlaid and this limit disappears.
WebSocket is bidirectional, carries binary, and preserves message boundaries. In return, the application has to build reconnection, resuming, backpressure, and the pairing of requests and responses entirely itself. The sequence numbers, ring buffer, and jittered backoff you built yourself in the earlier module are exactly that. It starts with HTTP/1.1 Upgrade, so intermediate devices have to pass the Upgrade along.
gRPC streaming gives you a contract defined in proto, deadlines, status codes, flow control, and a retry policy. For communication between services, it gives you the most for free. But browsers have no API for handling HTTP/2 trailers, so they cannot speak gRPC directly, and gRPC-Web, which puts a translation layer in between, does not support client streaming or bidirectional streaming.
WebRTC is at a different layer from the previous three. It can discard late data over UDP, and the Opus codec, the jitter buffer, and echo cancellation are built into the browser, so it carries a person's voice with the shortest delay. In exchange, you have to operate signaling, ICE, and TURN servers yourself.
| Criterion | SSE | WebSocket | gRPC streaming | WebRTC |
|---|---|---|---|---|
| Direction | Server → client | Bidirectional | Four shapes | Bidirectional (peer to peer) |
| Browser | Built-in support | Built-in support | Only partly through gRPC-Web | Built-in support |
| Reconnection and resuming | In the standard | Do it yourself | Do it yourself (retries only before commit) | ICE restart |
| Discarding late data | Not possible | Not possible (TCP) | Not possible (TCP) | Possible |
| Operational burden | Lowest | Medium | Medium (proxy HTTP/2) | Highest (TURN) |
I will also point out one thing not in the table. WebTransport over HTTP/3 is a new API that lets a browser use QUIC streams and unreliable datagrams together, aiming at both WebSocket's easy server structure and WebRTC's discarding of late data. It is not yet usable in every browser and server setup, so to adopt it you must first check in the target environment. Even when a new technology appears, the questions for choosing are the same — the direction, browser support, whether you can discard late data, and whether the people who will operate it can handle it.
What it looks like in the field
Voice AI services often mix more than one. They carry the voice between browsers or phones and the server with WebRTC, talk to the recognition and synthesis engines inside the server with gRPC bidirectional streaming or WebSocket, and carry the captions and tokens drawn on screen with SSE or WebSocket. The results are better if you choose, for each segment, the properties it needs — discarding late data, contracts and deadlines, browser support — than if you unify the transport into one.
Authentication and observability differ by transport too. SSE and gRPC have headers on every request, so you can use existing authentication middleware and access logs as they are. WebSocket has no headers after a single handshake, so on a long connection where the token expires, you have to re-authenticate with a message or cut the connection and have it reattach, and you have to build per-message metrics yourself.
There are also common reasons that force a choice to be reversed. A corporate proxy blocks the WebSocket Upgrade and you fall back to SSE, or because of a network where UDP is blocked, you end up adding TURN over TCP/TLS 443 to WebRTC. Checking "does it work on this network" from the start is the first step in choosing a transport.
What you will do in the next quiz
You choose the suitable transport for each situation, and tell apart what each transport gives you for free from what you have to build yourself.