TT Lab
Get started
Learn Learning paths Courses

Real-Time Communication — WebSocket, gRPC Streaming and WebRTC

Send audio frames over WebSocket and WebRTC and measure latency

Continue in TT Lab

Goal

Create the 16kHz 20ms frames a voice AI actually exchanges, send them over WebSocket and WebRTC, compute yourself the percentiles of round-trip latency, the jitter, and the playout buffer size, and compare the two transports with numbers.

Why it matters

The perceived quality of a voice conversation is decided not by the average latency but by the slowest 1% of frames and the playout buffer. WebSocket is easy to implement on the server side and passes through firewalls well, but being TCP, once it stops, every frame behind it is late. WebRTC can discard what is late and has Opus and a jitter buffer, but its connection setup is heavy. Before putting ASR and TTS on top of this transport in the voice AI course, you confirm here with numbers you measured yourself which one wins under which conditions.

Steps

  1. Cut a WAV into 20ms frames — In /root/rt/voice/voice.py, create frames(path, frame_ms=20). Read a 16-bit mono WAV with the wave module and return a list of bytes pieces of frame_ms length. The byte count of one piece is sample rate × frame_ms / 1000 × 2. If the last piece falls short, pad it with zeros. It is a ValueError if the file is not 16-bit mono. The material is /opt/fixtures/rt/voice/speech16k.wav (16kHz, 3 seconds).
  2. Send in step with real time — In /root/rt/voice/voice.py, add async send_paced(send, frames, frame_ms, prepare=None). Send the i-th frame in step with start time + i × frame_ms. When that time comes, if prepare is given, call prepare(frame) and pass its result, and otherwise pass the frame, on to await send(...). Return a list holding the time each frame was sent, in seconds relative to the start time. The grader gives a prepare that takes 5ms each time.
  3. Look at percentiles and jitter, not the average — In /root/rt/voice/voice.py, add percentiles(samples, ps=(50, 95, 99)) and jitter(transit_ms). percentiles uses the nearest-rank method, picking for each p the ceil(p/100 × n)-th of the sorted values, and returns a {p: value} dict, and it is a ValueError if samples is empty. jitter is RFC 3550's interarrival jitter estimator, and returns the final J after repeating J = J + (|D| - J) / 16 using the absolute value of D, the difference in transit time of two consecutive packets (J starts at 0, and with only one value it is 0).
  4. Send over WebSocket and measure the round trip — In /root/rt/voice/voice.py, add async ws_measure(url, frames, frame_ms). Attach to url, prepend to each frame the 12 bytes of struct.pack("!IQ", , time.monotonic_ns()), send them with send_paced, and each time an echo is received, count the round-trip milliseconds from the difference between the current time and the time in the header. Anything that has not arrived within 1 second after sending is counted as lost. Return {"sent": , "received": , "rtt_ms": , "p50": , "p99": }. Then run /opt/rt-lab/bin/python /opt/fixtures/rt/voice/bench.py ws and write the three output lines ws_p50_ms, ws_p99_ms, and ws_stall_p99_ms to /root/rt/voice/report.txt.
  5. Do the same thing over a WebRTC data channel — In /root/rt/voice/voice.py, add async rtc_measure(frames, frame_ms). Inside one process, create two aiortc peers with an empty ICE server list, hand over the offer and answer directly, and have the offering side open a data channel with ordered=False and maxRetransmits=0. The answering side sends back the received messages as they are, and the offering side measures with the same header and in the same way as ws_measure and returns a dict of the same shape. Then run /opt/rt-lab/bin/python /opt/fixtures/rt/voice/bench.py rtc and add the two lines rtc_p50_ms and rtc_p99_ms to report.txt.
  6. Send sound on an Opus media track — In /root/rt/voice/voice.py, add async rtc_audio(path, seconds=2.0). Make an audio track that inherits aiortc.MediaStreamTrack emit the first seconds seconds of frames(path, 20) every 20ms as av.AudioFrame (s16, mono, 16000Hz, pts in sample units), and send it with addTrack between two peers in the same process. The receiving side keeps calling recv on the track it received through the track event and counts the number of frames and the sample rate. Return {"codec": , "sent": , "received": , "sample_rate": , "duration": }.
  7. Compute how large to make the playout buffer — In /root/rt/voice/voice.py, add playout(arrivals_ms, frame_ms, buffer_ms) and min_buffer(arrivals_ms, frame_ms, max_late). arrivals_ms is the arrival time (in milliseconds, None if lost) of sequence number i. The playback reference point is the arrival time of the first frame k to arrive minus k × frame_ms, and the playback deadline of the i-th frame is the reference point + buffer_ms + i × frame_ms. playout returns {"late": , "lost": }. min_buffer increases from 0 in steps of 10ms and returns the smallest buffer_ms at which the number late ÷ the number arrived is at most max_late (None if there is none up to 1000).

Notes

Cut a WAV into 20ms frames

In /root/rt/voice/voice.py, create frames(path, frame_ms=20). Read a 16-bit mono WAV with the wave module and return a list of bytes pieces of frame_ms length. The byte count of one piece is sample rate × frame_ms / 1000 × 2. If the last piece falls short, pad it with zeros. It is a ValueError if the file is not 16-bit mono. The material is /opt/fixtures/rt/voice/speech16k.wav (16kHz, 3 seconds).

Most voice codecs and ASR take sound in units of 10, 20, or 30ms. At 16kHz, 20ms is 320 samples, 640 bytes. If you drop the last piece, the end of the speech is cut off.

Send in step with real time

In /root/rt/voice/voice.py, add async send_paced(send, frames, frame_ms, prepare=None). Send the i-th frame in step with start time + i × frame_ms. When that time comes, if prepare is given, call prepare(frame) and pass its result, and otherwise pass the frame, on to await send(...). Return a list holding the time each frame was sent, in seconds relative to the start time. The grader gives a prepare that takes 5ms each time.

If you sleep frame_ms every time you send, the time spent on preparation adds up each time, and with 50 frames it drifts by hundreds of ms. Compute the deadline as an absolute time and sleep only for the remainder. You must not send everything in a burst either — the receiving side's buffer overflows.

Look at percentiles and jitter, not the average

In /root/rt/voice/voice.py, add percentiles(samples, ps=(50, 95, 99)) and jitter(transit_ms). percentiles uses the nearest-rank method, picking for each p the ceil(p/100 × n)-th of the sorted values, and returns a {p: value} dict, and it is a ValueError if samples is empty. jitter is RFC 3550's interarrival jitter estimator, and returns the final J after repeating J = J + (|D| - J) / 16 using the absolute value of D, the difference in transit time of two consecutive packets (J starts at 0, and with only one value it is 0).

What the user feels on a real-time channel is not the average but the occasional long delay. If 1% of frames are 300ms late, the average hardly changes but the conversation breaks up. Jitter expresses how uneven the arrivals are and is the basis for deciding the playout buffer size.

Send over WebSocket and measure the round trip

In /root/rt/voice/voice.py, add async ws_measure(url, frames, frame_ms). Attach to url, prepend to each frame the 12 bytes of struct.pack("!IQ", , time.monotonic_ns()), send them with send_paced, and each time an echo is received, count the round-trip milliseconds from the difference between the current time and the time in the header. Anything that has not arrived within 1 second after sending is counted as lost. Return {"sent": , "received": , "rtt_ms": , "p50": , "p99": }. Then run /opt/rt-lab/bin/python /opt/fixtures/rt/voice/bench.py ws and write the three output lines ws_p50_ms, ws_p99_ms, and ws_stall_p99_ms to /root/rt/voice/report.txt.

If you alternate sending and receiving in one loop, the next frame is late while you wait to receive. Put the receiving side in a separate asyncio task. The second measurement of bench is the case where TCP stalls once for 300ms. You should be able to explain why p99 spikes even though nothing was lost.

Do the same thing over a WebRTC data channel

In /root/rt/voice/voice.py, add async rtc_measure(frames, frame_ms). Inside one process, create two aiortc peers with an empty ICE server list, hand over the offer and answer directly, and have the offering side open a data channel with ordered=False and maxRetransmits=0. The answering side sends back the received messages as they are, and the offering side measures with the same header and in the same way as ws_measure and returns a dict of the same shape. Then run /opt/rt-lab/bin/python /opt/fixtures/rt/voice/bench.py rtc and add the two lines rtc_p50_ms and rtc_p99_ms to report.txt.

It is the same process, so you do not need a signaling server. Just pass one peer's localDescription straight to the other peer's setRemoteDescription. If you send before the channel opens, it disappears.

Send sound on an Opus media track

In /root/rt/voice/voice.py, add async rtc_audio(path, seconds=2.0). Make an audio track that inherits aiortc.MediaStreamTrack emit the first seconds seconds of frames(path, 20) every 20ms as av.AudioFrame (s16, mono, 16000Hz, pts in sample units), and send it with addTrack between two peers in the same process. The receiving side keeps calling recv on the track it received through the track event and counts the number of frames and the sample rate. Return {"codec": , "sent": , "received": , "sample_rate": , "duration": }.

A data channel carries bytes, and a media track carries sound through a codec, RTP timestamps, and SRTP. WebRTC's Opus is always written as opus/48000/2 in the SDP even if the input is 16kHz, and the receiving side decodes it at 48kHz. That means you have to downsample to 16kHz again before ASR. If recv does not sleep by itself, 3 seconds' worth goes out in an instant.

Compute how large to make the playout buffer

In /root/rt/voice/voice.py, add playout(arrivals_ms, frame_ms, buffer_ms) and min_buffer(arrivals_ms, frame_ms, max_late). arrivals_ms is the arrival time (in milliseconds, None if lost) of sequence number i. The playback reference point is the arrival time of the first frame k to arrive minus k × frame_ms, and the playback deadline of the i-th frame is the reference point + buffer_ms + i × frame_ms. playout returns {"late": , "lost": }. min_buffer increases from 0 in steps of 10ms and returns the smallest buffer_ms at which the number late ÷ the number arrived is at most max_late (None if there is none up to 1000).

A playout buffer waits for late frames but delays every sound by that much. If it is too small, it breaks up, and if it is too large, the conversation gets sluggish. The more jittery the network, the larger the buffer you need to keep the same breakup rate.