Real-Time Communication — WebSocket, gRPC Streaming and WebRTC
The four shapes of gRPC streaming
In one line
A gRPC call is one HTTP/2 stream, messages flow on it as pieces with a length in front, and the status code arrives in the trailers at the very end.
Why this was needed
If you build a real-time feature with unary calls alone, it becomes polling. If you gather results that are produced little by little, such as the partial results of speech recognition, an LLM's tokens, or quotes, and send them all at once, the time until you receive the first piece becomes the same as the total processing time. The gRPC core concepts documentation divides calls into four shapes — unary, server streaming, client streaming, and bidirectional streaming. The shape is decided by which of the request and the response carries stream in the proto's rpc declaration, and the generated code creates a stub that fits that shape.
How it works
Looking at the gRPC over HTTP/2 specification, one call is one HTTP/2 stream. The request starts with :path /패키지.서비스/메서드 (the placeholders are the package, the service, and the method), content-type: application/grpc, and, if there is a deadline, a grpc-timeout header, and messages are laid end to end as 1 byte of compression flag + 4 bytes of length + protobuf bytes. The response also ends with grpc-status and grpc-message in the trailers after messages of the same shape.
Three things follow from this structure.
First, streaming is streaming only when you send as things are produced. If you yield from a generator in a Python server, gRPC takes them out and sends them one at a time. If you collect all the results into a list and return it, the first message arrives at the same moment as the last message. The code has a streaming shape but the user experience is unary.
Second, the status comes at the end. If the server sends three messages and then ends with an error, the client sees the error only after receiving the three messages. In a Python client, grpc.RpcError pops out in the middle of iteration. If you receive with the single line list(stub.Method(...)), you lose the three you already received. For a stream where partial results are meaningful, you have to secure them one by one inside the loop. There are 17 status codes, and the ones often seen in real-time services are DEADLINE_EXCEEDED (4), CANCELLED (1), UNAVAILABLE (14), RESOURCE_EXHAUSTED (8), ABORTED (10), and INVALID_ARGUMENT (3).
Third, the two directions of a bidirectional stream are independent of each other. The server can answer before it has received all the requests, and as the core concepts documentation puts it, the two streams can read and write in any order. Because of that freedom, deadlocks where each waits on the other and stops arise easily. The client sends the next request after receiving an answer, but if the server first collects all the requests with list(request_iterator), the client waits for an answer and the server waits for the end of the requests (half-close). Without a deadline, they wait forever.
You put the deadline on the call. The deadlines guide says the default is effectively infinite and recommends setting a deadline explicitly on every call. The deadline is conveyed to the server through the grpc-timeout header, and when it passes, the client receives DEADLINE_EXCEEDED. But the server's processing thread does not stop by itself. The server code has to check for itself with context.is_active() or context.time_remaining(). If it does not check, it uses CPU and DB connections to the end for a result nobody will receive.
Message size has a limit too. Most implementations put the default upper bound for received messages at 4MB and end the call with RESOURCE_EXHAUSTED if it is exceeded. Another use of streaming is to cut a big result into a stream instead of putting it in one message. If you send it cut up, the receiving side can start processing from the first piece, and the memory that flow control holds at a time shrinks to the piece size. Conversely, if you cut pieces too fine, the 5-byte header attached to each message and the serialization cost grow, so for data with a cadence like voice, you cut at a natural unit such as 20ms.
What it looks like in the field
"We switched to streaming but it still comes all at once" is the case where the server gathers and sends, or where an intermediate proxy or load balancer buffers the response. To tell the two apart, measure the first message's arrival time once next to the server and once on the user's side. "The client timed out but server CPU keeps running" is server code that does not check the deadline. At the moment deadline overruns pile up, the server fills with work that has already been abandoned, and it spreads into a cascading failure where even new requests slow down.
What you will do in the next lab
You generate code from meter.proto and implement the server, the client, and bidirectional streaming. The grader checks whether it streams from the arrival time of the first message, and checks for a deadlock with a call with a 3-second deadline that exchanges one line at a time. You also build a client that keeps the values it received on a mid-stream error and a server that stops by itself once the deadline passes.