Real-Time Communication — WebSocket, gRPC Streaming and WebRTC
Operating long-lived WebSockets
In one line
A WebSocket that stays attached for a long time needs limits, heartbeats, and resume rules so that it can withstand slow subscribers, idle timeouts, silent peers, and mass reconnects.
Why this was needed
Even if you implement the protocol correctly, a WebSocket service collapses as time passes. That is because four things inevitably happen while connections stay alive for hours. Someone slows down, intermediate devices cut quiet connections, a phone goes into a tunnel and vanishes without a word, and a deployment cuts every connection at once. In HTTP, where requests and responses end quickly, these four were almost invisible.
How it works
Slow subscribers. If the publishing loop calls await send for each subscriber in turn, when one person's send buffer fills up, everyone behind them waits. If you give each subscriber a separate queue and a loop that drains that queue, the waiting is cut loose person by person. But if the queue has no upper bound, the slow one's share piles up in server memory without end. When you hit the bound, there are three things you can do — discard the old ones, merge into a single latest value (quotes, cursor position), or cut the connection. For a stream where order and omissions matter, it is more honest to cut and have them reconnect than to keep sending with holes, and then you leave the reason with close code 1013 (try again later). The browser's WebSocket API has no backpressure and you can only see the amount backed up through bufferedAmount, so the sending side has to watch this number itself.
Idle timeouts. Load balancers and proxies cut quiet connections after a set time. The default of an AWS Application Load Balancer is 60 seconds, and the default of nginx's proxy_read_timeout is also 60 seconds. Linux TCP keepalive by default sends its first probe only after two hours, so it is useless for this purpose. So the application sends a ping more often than the shortest idle limit. A ping keeps the connection alive and, at the same time, confirms whether the other side is alive by whether the answer (pong) comes.
Silent peers. If the other side vanishes without a close, TCP knows nothing for a while. Sending a ping and cutting if there is no pong within a set time is the only way to check. There is one common defect here. Even if the library closes the connection, the coroutine that was waiting on the subscriber queue does not wake until a new message arrives. A dead connection stays in the list, inflating the connection count metric and receiving and piling up messages. You have to wait on the queue and the connection closing together.
Mass reconnects. When one server restarts, tens of thousands of that server's connections are cut at the same time, and the clients retrying with the same rules come back in a rush at the same moment. Mixing randomness into exponential backoff (the method the AWS architecture blog called full jitter: drawing uniformly between 0 and min(cap, base × 2^attempt)) spreads that wave out. Close codes are also used in the decision. Only cases where the other side's circumstances may change, such as 1001, 1006, 1011, 1012, and 1013, reconnect, and a connection cut with 1008 (policy violation) or 1002 (protocol error) will be cut for the same reason even if it reconnects.
Resume. SSE has a rule in its specification for announcing the place where it broke off with Last-Event-ID, but WebSocket does not. You attach sequence numbers to messages, the server keeps recent messages in a ring buffer, and when the client reports the last sequence number it has processed, the server resends what comes after it. If that place has already been pushed out of the buffer, do not quietly leave a hole; tell the client "start over from the beginning." What the client should report is not the sequence number it received but the sequence number it processed and recorded.
Limits on the number of connections. One connection takes up one file descriptor, a kernel buffer, and the application's queue. The Pod's descriptor limit (ulimit -n) and memory set the ceiling on concurrent connections, so in a load test you have to raise the connection count step by step and measure the memory per connection.
Scaling. A connection being tied to one server is also a property of WebSocket. When the publisher and the subscriber attach to different servers, you need somewhere for the servers to share messages (Redis pub/sub, a message broker), and the ring buffer for resuming also has to sit in that shared store and not in one server's memory, so that they can pick up where they left off even when reconnecting to another server. That is also why sequence numbers are assigned per channel, not per server.
What it looks like in the field
Services are common where every deployment sends server CPU soaring for a few minutes and errors pour out. Looking closely, reconnection has no jitter, and right after reconnecting, the full state is downloaded again. With jitter and sequence-number-based resume, those few minutes disappear. When taking a server down, it also helps to close with 1001 or 1012 so that clients know "it's fine to reconnect."
What you will do in the next lab
With websockets, you build a publish/subscribe hub and a reconnecting client. With one non-reading subscriber attached, you publish 64MB and check memory and 1013, you hold on for 6 seconds behind a relay with a 2.5-second idle limit, you clear out a subscriber that does not send a pong within 6 seconds, and you check that even if the connection is cut twice in the middle of publishing, sequence numbers from 1 to the end are recorded exactly once each with none missing.