Real-Time Communication — WebSocket, gRPC Streaming and WebRTC
gRPC deadlines, cancellation, flow control and keepalive
In one line
A deadline has to be passed down the chain while shrinking, cancellation has to be conveyed directly, flow control must not be bypassed, and keepalive, retries, and shutdown need the rules on both sides to match.
Why this was needed
Calls in a real-time service form a chain. It goes browser → gateway → speech recognition → language model → speech synthesis, and if any one place in the chain breaks the rules, the symptom shows up in an unexpected place. A request the user has given up on in front is processed to the end behind, one slow consumer fills the producer's memory, a quiet stream gets cut by an intermediate device, a retry makes a result twice, and streams break at every deployment.
How it works
Deadline propagation. When the front service passes a call it received with a 0.6-second deadline on to the back, it should give the back only the remaining time. In Go and Java, if you pass the incoming context along as it is, the deadline follows, but in Python you have to pass context.time_remaining() yourself as the timeout of the back call. On a call with no deadline, this value becomes a very large number, so do not pass it as it is; call without a deadline.
Cancellation propagation. As the cancellation guide explains, when a client cancels a call, the server-side context becomes inactive. But if that server is stuck waiting on a back service, the back call does not know about the cancellation. In Python, you have to make the back call as a future and use context.add_callback to cancel that future when your own call ends, so that the chain stops together.
Flow control. According to the flow control guide, gRPC uses HTTP/2 flow control to make the sender send only what the receiver can handle. A streaming generator in a Python server takes out the next value only when the window allows, so if the consumer stops, production stops too. If you keep a separate production thread and fill an unbounded queue ahead of time, you bypass this mechanism and the consumer's share piles up in server memory. In the lab, when you request 3,000 messages of 64KiB and stop after reading just one, this difference splits into a few MB versus hundreds of MB.
keepalive. The client settings in the keepalive guide are the ping interval (grpc.keepalive_time_ms), the time to wait for an answer (grpc.keepalive_timeout_ms), and whether to ping even when there are no calls (grpc.keepalive_permit_without_calls). The server has the right to reject pings that are too frequent. If you ping more often than the minimum interval the server allows (grpc.http2.min_recv_ping_interval_without_data_ms), it cuts the connection with a GOAWAY carrying the ENHANCE_YOUR_CALM error code and too_many_pings. So the interval has to be shorter than the idle limit of intermediate devices and longer than the interval the server allows. The two are set by different teams, so you must check that they match.
Retries. Retries in the retry design document (gRFC A6) are declared through retryPolicy in the service config — the maximum number of attempts, the first and maximum waits, the multiplier, and the list of status codes to retry. The important constraint is commit. Once you receive the server's response headers or the first message, the call is committed, and later failures are not retried even with the same status code. If you call a stream again from the beginning in the application, you receive the messages you already received twice. Errors that would give the same result if tried again, such as INVALID_ARGUMENT, are not put in the list.
Graceful shutdown. When Kubernetes takes down a Pod, it sends SIGTERM and then sends SIGKILL after terminationGracePeriodSeconds (default 30 seconds). Without a handler, a Python process dies immediately on SIGTERM, and all in-flight streams are cut with UNAVAILABLE. server.stop(grace) rejects new calls, gives in-flight calls a grace period, and then finishes.
What it looks like in the field
gRPC connections live long, so behind an L4 load balancer a connection stays on the server it first attached to. Even if you add servers, the load does not go to the new ones. If the server puts a maximum age on connections (grpc.max_connection_age_ms) and sends a GOAWAY periodically, clients reattach and the load spreads. This too is a "connection lifetime rule" of the same kind as keepalive.
What you will do in the next lab
You build a relay server that passes the deadline and cancellation on to the back service, a streaming server that respects flow control, a channel that declares keepalive and retry policies, and a server that shuts down gracefully on SIGTERM. The judgment is made from the deadline, the number of calls, and the number of completed steps the reference server received, and from your server's memory.