TT Lab
Get started
Learn Learning paths Courses

Real-Time Communication — WebSocket, gRPC Streaming and WebRTC

Connect two WebRTC peers from signaling to data channel

Continue in TT Lab

Goal

Build a WebSocket signaling server and two aiortc peers, exchange an offer and an answer, pass messages back and forth over a data channel with a chosen reliability and ordering, and then measure where ICE candidates come from and what gets slow when STUN is blocked.

Why it matters

WebRTC is the standard that carries voice between browsers and phones with the shortest delay. But the specification does not define how to establish the connection, so the service itself is responsible for signaling, ICE server configuration, and candidate gathering. Most of the failures "it can't connect only on the company network" and "it takes 5 seconds to connect" come from this part. Being in the same Pod, there is no NAT, but thanks to that you can see the roles of candidates and STUN without clutter.

Steps

  1. Build a signaling server — In /root/rt/webrtc/signaling.py, create the run(host, port) coroutine. Accept the /room/ path with a websockets server, and pass text messages through as they are between two connections attached to the same room. Messages that arrive while the other side is not there yet are held and passed on in order the moment the other side attaches. A third connection to attach is closed with close code 4001 and the reason "room full", and when one side leaves, send {"type": "bye"} to the remaining side.
  2. Create an offer SDP — In /root/rt/webrtc/peer.py, create the make_offer() coroutine. Create an RTCPeerConnection with RTCConfiguration(iceServers=[]), which leaves the ICE server list empty, create a "chat" data channel, and then go through createOffer and setLocalDescription and return the localDescription.sdp string. Close the connection before returning.
  3. Read an SDP — In /root/rt/webrtc/sdp.py, create summarize_sdp(sdp). Read an SDP string and return a dict. The keys are five: ufrag (the value of the first a=ice-ufrag), fingerprint (the hash name of the first a=fingerprint in lowercase, for example sha-256), setup (the value of the first a=setup), media (a list holding the medium type of each m= line in the order they appear), and candidates ({"host": n, "srflx": n, "relay": n} — the count per typ of the a=candidate lines, with kinds that are absent as 0). Line endings can be either \r\n or \n.
  4. The answering peer — In /root/rt/webrtc/peer.py, add the answer_peer(signal_url, room, out_path) coroutine. Attach to signal_url/room/ and wait for {"type": "offer", "sdp": ...}, and on receiving it, after setRemoteDescription → createAnswer → setLocalDescription, send {"type": "answer", "sdp": ...}. For messages on the data channel the other side opened, send them back with "echo:" prepended, and write the channel's label, ordered, and maxRetransmits as JSON to out_path. On receiving {"type": "bye"}, close the connection and finish.
  5. The offering peer — In /root/rt/webrtc/peer.py, add the offer_peer(signal_url, room, count) coroutine. Create a "chat" data channel and send the offer, receive the answer and call setRemoteDescription, and when the channel opens, send count messages one at a time starting from "m0", measuring the round-trip time each time an echo arrives. Return {"sent": count, "echoed": , "rtt_ms": }.
  6. A channel that discards late data — Create offer_peer's data channel with ordered=False and maxRetransmits=0. The grader's reference answerer records the properties of the channel it received and checks them.
  7. See where the candidates come from — In /root/rt/webrtc/peer.py, add the gather(stun=None) coroutine. If stun is given, use a single RTCIceServer(stun) as the ICE server, and otherwise an empty list, create one data channel, and measure the milliseconds spent in createOffer and setLocalDescription. Return {"elapsed_ms": , "candidates": [(typ, address, port), ...]}. The grader calls it three times: without STUN, with a STUN server in the same Pod, and with an unreachable STUN address.

Notes

Build a signaling server

In /root/rt/webrtc/signaling.py, create the run(host, port) coroutine. Accept the /room/ path with a websockets server, and pass text messages through as they are between two connections attached to the same room. Messages that arrive while the other side is not there yet are held and passed on in order the moment the other side attaches. A third connection to attach is closed with close code 4001 and the reason "room full", and when one side leaves, send {"type": "bye"} to the remaining side.

The WebRTC specification does not define how to hand over the SDP (JSEP, RFC 9429, also leaves that part to the application). So almost every service keeps a small relay server like this. 4000–4999 is the range of close codes RFC 6455 left for applications.

Create an offer SDP

In /root/rt/webrtc/peer.py, create the make_offer() coroutine. Create an RTCPeerConnection with RTCConfiguration(iceServers=[]), which leaves the ICE server list empty, create a "chat" data channel, and then go through createOffer and setLocalDescription and return the localDescription.sdp string. Close the connection before returning.

For an m= line to appear in the SDP, you have to decide what to send (a data channel or a media track) before offering. If you do not specify ICE servers, aiortc uses a public STUN server, and where outbound UDP is blocked, candidate gathering stalls for several seconds waiting on its response.

Read an SDP

In /root/rt/webrtc/sdp.py, create summarize_sdp(sdp). Read an SDP string and return a dict. The keys are five: ufrag (the value of the first a=ice-ufrag), fingerprint (the hash name of the first a=fingerprint in lowercase, for example sha-256), setup (the value of the first a=setup), media (a list holding the medium type of each m= line in the order they appear), and candidates ({"host": n, "srflx": n, "relay": n} — the count per typ of the a=candidate lines, with kinds that are absent as 0). Line endings can be either \r\n or \n.

host is the address of its own interface, srflx (server reflexive) is the external address a STUN server saw, and relay is the relay address a TURN server lent. The fingerprint is the hash of the DTLS certificate, and the connection is made only if this value that came over signaling matches the certificate received during the handshake.

The answering peer

In /root/rt/webrtc/peer.py, add the answer_peer(signal_url, room, out_path) coroutine. Attach to signal_url/room/ and wait for {"type": "offer", "sdp": ...}, and on receiving it, after setRemoteDescription → createAnswer → setLocalDescription, send {"type": "answer", "sdp": ...}. For messages on the data channel the other side opened, send them back with "echo:" prepended, and write the channel's label, ordered, and maxRetransmits as JSON to out_path. On receiving {"type": "bye"}, close the connection and finish.

The data channel is created by the offering side, and the answering side receives it through the datachannel event. Leave the ICE server list empty on the answering peer too. The grader's reference offerer attaches through your signaling server.

The offering peer

In /root/rt/webrtc/peer.py, add the offer_peer(signal_url, room, count) coroutine. Create a "chat" data channel and send the offer, receive the answer and call setRemoteDescription, and when the channel opens, send count messages one at a time starting from "m0", measuring the round-trip time each time an echo arrives. Return {"sent": count, "echoed": , "rtt_ms": }.

If you send before the channel's open event, the message disappears. Only after the ICE check, the DTLS handshake, and the SCTP connection have all finished is it open. The time taken up to here is WebRTC's connection setup cost.

A channel that discards late data

Create offer_peer's data channel with ordered=False and maxRetransmits=0. The grader's reference answerer records the properties of the channel it received and checks them.

For data that is useless after 20ms, like a voice frame or a cursor position, it is better to drop it than to resend one lost piece and hold up everything behind it. SCTP lets you set order and retransmission per channel, and this is something you cannot imitate with WebSocket over TCP.

See where the candidates come from

In /root/rt/webrtc/peer.py, add the gather(stun=None) coroutine. If stun is given, use a single RTCIceServer(stun) as the ICE server, and otherwise an empty list, create one data channel, and measure the milliseconds spent in createOffer and setLocalDescription. Return {"elapsed_ms": , "candidates": [(typ, address, port), ...]}. The grader calls it three times: without STUN, with a STUN server in the same Pod, and with an unreachable STUN address.

If you split an a=candidate line by spaces, the fifth field is the address, the sixth is the port, and what follows typ is the kind. There is no NAT inside the same Pod, so first predict what the srflx address will turn out to be the same as.