Real-Time Communication — WebSocket, gRPC Streaming and WebRTC
Measure real-time channels and send voice frames
In one line
You judge a real-time channel not by the average but by the tail of the latency distribution and the reconnect rate, send voice frames keeping a 20ms cadence, and trade breakups against delay with the playout buffer.
Why this was needed
Most failures of a real-time channel happen "occasionally." If 1% of frames are 300ms late, the average latency stays almost the same but the conversation breaks up. The connection count metric can look fine while thousands of reconnects happen per minute. For a voice AI service, cadence is added on top. The microphone makes a frame every 20ms, and the recognition engine and the player expect that cadence. If you break the cadence, the receiving side's buffer overflows or runs empty and the sound skips or cuts out.
How it works
Frame arithmetic. The format speech recognition commonly takes is 16kHz, 16-bit, mono. That is 32,000 bytes per second (256kbps), and one 20ms frame is 320 samples, 640 bytes. Encoding with Opus brings speech down to tens of kbps. When sending frames, you have to set the deadline as an absolute time (start + i × 20ms). The approach of sleeping 20ms each time adds the time spent on encoding every time, so it drifts by hundreds of ms per second.
The distribution of latency. Round-trip latency is measured by writing the monotonic clock at the moment of sending into the frame header and subtracting it at the moment the echo arrives. One-way latency is much harder because the clocks of the two machines have to agree. You summarize the collected samples as p50, p95, and p99. The nearest-rank method picks the ceil(p/100 × n)-th of the sorted values, so it reports only values that were actually observed. Averaging the p99 of several servers gives a meaningless number, so you have to collect and merge each server's histogram and then compute the percentile.
Jitter. Section 6.4.1 of RFC 3550 estimates interarrival jitter by repeating J = J + (|D| − J)/16 using the absolute value of D, the difference in transit time of two consecutive packets. The 1/16 is a gain that reduces noise. The browser's WebRTC statistics (getStats) give values such as jitter, packetsLost, jitterBufferDelay, and concealedSamples for each incoming RTP stream, and also the roundTripTime reported by the other side.
The playout buffer. The receiving side waits for the buffer time after the first frame arrives and then plays one frame every 20ms. If the i-th frame arrives later than its playback time, it is the same as not having it. If you increase the buffer, late frames decrease and every sound is delayed by that much. With WebSocket over TCP, nothing is lost, yet once it stops, the frames of that period all arrive late at once, so to keep the same breakup rate you need a buffer as long as the stall. A WebRTC data channel or RTP, which does not require order or retransmission, loses only what was lost.
Observing the channel. There are five numbers to look at on a real-time channel.
| Metric | Why |
|---|---|
| Connection setup time (handshake, ICE, DTLS) | The first time the user waits |
| Message latency p50 and p99 | The tail that hides behind the average |
| Reconnect rate and distribution of close codes | Where and why the breaks occur |
| Per-subscriber queue length and number dropped | Slow consumers |
| Ping round-trip time | Network condition and dead peers |
Pitfalls in how you measure. If a load generator sends the next request only after waiting for the response, the requests that should have been sent while the server was stalled drop out of the measurement entirely (coordinated omission). Only if you measure with a sender that keeps the cadence — the earlier sending without accumulated error — is the stall captured properly in the tail. Also, latency has to be measured with the monotonic clock of the same machine. A wall clock can even run backward when NTP adjusts it midway, producing negative latency.
Attach labels to the observed numbers. You have to record which transport (WebSocket, WebRTC), which region, and which network type (Wi-Fi, LTE) it was, so that when the overall p99 gets worse, you can tell which group's tail it is.
You look at the reconnect rate in the same way. Even if the connection count looks constant, if you do not separately count reconnects per minute and the distribution of close codes, a situation where 1% of users are cut and reattached every 30 seconds hides behind the metrics.
What it looks like in the field
The felt latency of a voice AI is the sum of a chain. The time to gather a 20ms frame, the network, end-of-utterance detection, the recognition's partial result, the language model's first token, the synthesis's first piece, and the playout buffer. The one-way 150ms that ITU-T G.114 speaks of is the standard for calls between people, and for a voice AI, processing time is added on top of it. So the tens of ms you can save on the transport matter, and you first need the tool that measures those tens of ms. The numbers in this lab become the baseline when you put recognition and synthesis on top of this transport in the voice AI course.
What you will do in the next lab
You cut a 16kHz WAV into 20ms frames, send them keeping the cadence without accumulated error, and compute percentiles and RFC 3550 jitter. You measure the round-trip latency of the same frames over WebSocket and over a WebRTC data channel, send sound on an Opus media track and confirm it is received at 48kHz, and then compute the playout buffer size from the arrival records.