Real-Time Communication — WebSocket, gRPC Streaming and WebRTC
Send audio frames over WebSocket and WebRTC and measure latency
Goal
Create the 16kHz 20ms frames a voice AI actually exchanges, send them over WebSocket and WebRTC, compute yourself the percentiles of round-trip latency, the jitter, and the playout buffer size, and compare the two transports with numbers.
Why it matters
The perceived quality of a voice conversation is decided not by the average latency but by the slowest 1% of frames and the playout buffer. WebSocket is easy to implement on the server side and passes through firewalls well, but being TCP, once it stops, every frame behind it is late. WebRTC can discard what is late and has Opus and a jitter buffer, but its connection setup is heavy. Before putting ASR and TTS on top of this transport in the voice AI course, you confirm here with numbers you measured yourself which one wins under which conditions.
Steps
- Cut a WAV into 20ms frames — In /root/rt/voice/voice.py, create frames(path, frame_ms=20). Read a 16-bit mono WAV with the wave module and return a list of bytes pieces of frame_ms length. The byte count of one piece is sample rate × frame_ms / 1000 × 2. If the last piece falls short, pad it with zeros. It is a ValueError if the file is not 16-bit mono. The material is /opt/fixtures/rt/voice/speech16k.wav (16kHz, 3 seconds).
- Send in step with real time — In /root/rt/voice/voice.py, add async send_paced(send, frames, frame_ms, prepare=None). Send the i-th frame in step with start time + i × frame_ms. When that time comes, if prepare is given, call prepare(frame) and pass its result, and otherwise pass the frame, on to await send(...). Return a list holding the time each frame was sent, in seconds relative to the start time. The grader gives a prepare that takes 5ms each time.
- Look at percentiles and jitter, not the average — In /root/rt/voice/voice.py, add percentiles(samples, ps=(50, 95, 99)) and jitter(transit_ms). percentiles uses the nearest-rank method, picking for each p the ceil(p/100 × n)-th of the sorted values, and returns a {p: value} dict, and it is a ValueError if samples is empty. jitter is RFC 3550's interarrival jitter estimator, and returns the final J after repeating J = J + (|D| - J) / 16 using the absolute value of D, the difference in transit time of two consecutive packets (J starts at 0, and with only one value it is 0).
- Send over WebSocket and measure the round trip — In /root/rt/voice/voice.py, add async ws_measure(url, frames, frame_ms). Attach to url, prepend to each frame the 12 bytes of struct.pack("!IQ", , time.monotonic_ns()), send them with send_paced, and each time an echo is received, count the round-trip milliseconds from the difference between the current time and the time in the header. Anything that has not arrived within 1 second after sending is counted as lost. Return {"sent": , "received": , "rtt_ms": , "p50": , "p99": }. Then run /opt/rt-lab/bin/python /opt/fixtures/rt/voice/bench.py ws and write the three output lines ws_p50_ms, ws_p99_ms, and ws_stall_p99_ms to /root/rt/voice/report.txt.
- Do the same thing over a WebRTC data channel — In /root/rt/voice/voice.py, add async rtc_measure(frames, frame_ms). Inside one process, create two aiortc peers with an empty ICE server list, hand over the offer and answer directly, and have the offering side open a data channel with ordered=False and maxRetransmits=0. The answering side sends back the received messages as they are, and the offering side measures with the same header and in the same way as ws_measure and returns a dict of the same shape. Then run /opt/rt-lab/bin/python /opt/fixtures/rt/voice/bench.py rtc and add the two lines rtc_p50_ms and rtc_p99_ms to report.txt.
- Send sound on an Opus media track — In /root/rt/voice/voice.py, add async rtc_audio(path, seconds=2.0). Make an audio track that inherits aiortc.MediaStreamTrack emit the first seconds seconds of frames(path, 20) every 20ms as av.AudioFrame (s16, mono, 16000Hz, pts in sample units), and send it with addTrack between two peers in the same process. The receiving side keeps calling recv on the track it received through the track event and counts the number of frames and the sample rate. Return {"codec": , "sent": , "received": , "sample_rate": , "duration": }.
- Compute how large to make the playout buffer — In /root/rt/voice/voice.py, add playout(arrivals_ms, frame_ms, buffer_ms) and min_buffer(arrivals_ms, frame_ms, max_late). arrivals_ms is the arrival time (in milliseconds, None if lost) of sequence number i. The playback reference point is the arrival time of the first frame k to arrive minus k × frame_ms, and the playback deadline of the i-th frame is the reference point + buffer_ms + i × frame_ms. playout returns {"late": , "lost": }. min_buffer increases from 0 in steps of 10ms and returns the smallest buffer_ms at which the number late ÷ the number arrived is at most max_late (None if there is none up to 1000).
Notes
- The working folder is /root/rt/voice. Create it first with mkdir -p /root/rt/voice.
- The measuring tools are /opt/rt-lab/bin/python /opt/fixtures/rt/voice/bench.py ws and rtc, and they call your functions. The measuring tool starts the echo server and the stalling relay itself.
- The material WAV is a 16kHz mono 3-second file made deterministically by /opt/fixtures/rt/voice/make_wav.py. It is not a human voice but a sound imitated with harmonics and a syllable envelope.
- Two common mistakes. Sleeping frame_ms every time you send so that the preparation time accumulates, and the audio track's recv not sleeping so that several seconds' worth go out in an instant.
- Always run Python with /opt/rt-lab/bin/python. This lab's libraries are only in that virtual environment, and if you run it with plain python3, you get a ModuleNotFoundError. It is convenient to shorten it with something like alias rpy=/opt/rt-lab/bin/python.
- The lab Pod blocks outbound connections. All communication happens on 127.0.0.1 inside the same Pod, and no installation or download is needed.
- The grader loads your code in a separate process and makes real connections. The example file is only a function skeleton, so it does not pass if left as is. Do not delete the functions you finished in earlier steps.
- When the lab session ends, the files in /root do not remain. Keep the code you need separately before you finish.
Cut a WAV into 20ms frames
In /root/rt/voice/voice.py, create frames(path, frame_ms=20). Read a 16-bit mono WAV with the wave module and return a list of bytes pieces of frame_ms length. The byte count of one piece is sample rate × frame_ms / 1000 × 2. If the last piece falls short, pad it with zeros. It is a ValueError if the file is not 16-bit mono. The material is /opt/fixtures/rt/voice/speech16k.wav (16kHz, 3 seconds).
Most voice codecs and ASR take sound in units of 10, 20, or 30ms. At 16kHz, 20ms is 320 samples, 640 bytes. If you drop the last piece, the end of the speech is cut off.
Send in step with real time
In /root/rt/voice/voice.py, add async send_paced(send, frames, frame_ms, prepare=None). Send the i-th frame in step with start time + i × frame_ms. When that time comes, if prepare is given, call prepare(frame) and pass its result, and otherwise pass the frame, on to await send(...). Return a list holding the time each frame was sent, in seconds relative to the start time. The grader gives a prepare that takes 5ms each time.
If you sleep frame_ms every time you send, the time spent on preparation adds up each time, and with 50 frames it drifts by hundreds of ms. Compute the deadline as an absolute time and sleep only for the remainder. You must not send everything in a burst either — the receiving side's buffer overflows.
Look at percentiles and jitter, not the average
In /root/rt/voice/voice.py, add percentiles(samples, ps=(50, 95, 99)) and jitter(transit_ms). percentiles uses the nearest-rank method, picking for each p the ceil(p/100 × n)-th of the sorted values, and returns a {p: value} dict, and it is a ValueError if samples is empty. jitter is RFC 3550's interarrival jitter estimator, and returns the final J after repeating J = J + (|D| - J) / 16 using the absolute value of D, the difference in transit time of two consecutive packets (J starts at 0, and with only one value it is 0).
What the user feels on a real-time channel is not the average but the occasional long delay. If 1% of frames are 300ms late, the average hardly changes but the conversation breaks up. Jitter expresses how uneven the arrivals are and is the basis for deciding the playout buffer size.
Send over WebSocket and measure the round trip
In /root/rt/voice/voice.py, add async ws_measure(url, frames, frame_ms). Attach to url, prepend to each frame the 12 bytes of struct.pack("!IQ", , time.monotonic_ns()), send them with send_paced, and each time an echo is received, count the round-trip milliseconds from the difference between the current time and the time in the header. Anything that has not arrived within 1 second after sending is counted as lost. Return {"sent": , "received": , "rtt_ms": , "p50": , "p99": }. Then run /opt/rt-lab/bin/python /opt/fixtures/rt/voice/bench.py ws and write the three output lines ws_p50_ms, ws_p99_ms, and ws_stall_p99_ms to /root/rt/voice/report.txt.
If you alternate sending and receiving in one loop, the next frame is late while you wait to receive. Put the receiving side in a separate asyncio task. The second measurement of bench is the case where TCP stalls once for 300ms. You should be able to explain why p99 spikes even though nothing was lost.
Do the same thing over a WebRTC data channel
In /root/rt/voice/voice.py, add async rtc_measure(frames, frame_ms). Inside one process, create two aiortc peers with an empty ICE server list, hand over the offer and answer directly, and have the offering side open a data channel with ordered=False and maxRetransmits=0. The answering side sends back the received messages as they are, and the offering side measures with the same header and in the same way as ws_measure and returns a dict of the same shape. Then run /opt/rt-lab/bin/python /opt/fixtures/rt/voice/bench.py rtc and add the two lines rtc_p50_ms and rtc_p99_ms to report.txt.
It is the same process, so you do not need a signaling server. Just pass one peer's localDescription straight to the other peer's setRemoteDescription. If you send before the channel opens, it disappears.
Send sound on an Opus media track
In /root/rt/voice/voice.py, add async rtc_audio(path, seconds=2.0). Make an audio track that inherits aiortc.MediaStreamTrack emit the first seconds seconds of frames(path, 20) every 20ms as av.AudioFrame (s16, mono, 16000Hz, pts in sample units), and send it with addTrack between two peers in the same process. The receiving side keeps calling recv on the track it received through the track event and counts the number of frames and the sample rate. Return {"codec": , "sent": , "received": , "sample_rate": , "duration": }.
A data channel carries bytes, and a media track carries sound through a codec, RTP timestamps, and SRTP. WebRTC's Opus is always written as opus/48000/2 in the SDP even if the input is 16kHz, and the receiving side decodes it at 48kHz. That means you have to downsample to 16kHz again before ASR. If recv does not sleep by itself, 3 seconds' worth goes out in an instant.
Compute how large to make the playout buffer
In /root/rt/voice/voice.py, add playout(arrivals_ms, frame_ms, buffer_ms) and min_buffer(arrivals_ms, frame_ms, max_late). arrivals_ms is the arrival time (in milliseconds, None if lost) of sequence number i. The playback reference point is the arrival time of the first frame k to arrive minus k × frame_ms, and the playback deadline of the i-th frame is the reference point + buffer_ms + i × frame_ms. playout returns {"late": , "lost": }. min_buffer increases from 0 in steps of 10ms and returns the smallest buffer_ms at which the number late ÷ the number arrived is at most max_late (None if there is none up to 1000).
A playout buffer waits for late frames but delays every sound by that much. If it is too small, it breaks up, and if it is too large, the conversation gets sluggish. The more jittery the network, the larger the buffer you need to keep the same breakup rate.