Real-Time Communication — WebSocket, gRPC Streaming and WebRTC
Connect two WebRTC peers from signaling to data channel
Goal
Build a WebSocket signaling server and two aiortc peers, exchange an offer and an answer, pass messages back and forth over a data channel with a chosen reliability and ordering, and then measure where ICE candidates come from and what gets slow when STUN is blocked.
Why it matters
WebRTC is the standard that carries voice between browsers and phones with the shortest delay. But the specification does not define how to establish the connection, so the service itself is responsible for signaling, ICE server configuration, and candidate gathering. Most of the failures "it can't connect only on the company network" and "it takes 5 seconds to connect" come from this part. Being in the same Pod, there is no NAT, but thanks to that you can see the roles of candidates and STUN without clutter.
Steps
- Build a signaling server — In /root/rt/webrtc/signaling.py, create the run(host, port) coroutine. Accept the /room/ path with a websockets server, and pass text messages through as they are between two connections attached to the same room. Messages that arrive while the other side is not there yet are held and passed on in order the moment the other side attaches. A third connection to attach is closed with close code 4001 and the reason "room full", and when one side leaves, send {"type": "bye"} to the remaining side.
- Create an offer SDP — In /root/rt/webrtc/peer.py, create the make_offer() coroutine. Create an RTCPeerConnection with RTCConfiguration(iceServers=[]), which leaves the ICE server list empty, create a "chat" data channel, and then go through createOffer and setLocalDescription and return the localDescription.sdp string. Close the connection before returning.
- Read an SDP — In /root/rt/webrtc/sdp.py, create summarize_sdp(sdp). Read an SDP string and return a dict. The keys are five: ufrag (the value of the first a=ice-ufrag), fingerprint (the hash name of the first a=fingerprint in lowercase, for example sha-256), setup (the value of the first a=setup), media (a list holding the medium type of each m= line in the order they appear), and candidates ({"host": n, "srflx": n, "relay": n} — the count per typ of the a=candidate lines, with kinds that are absent as 0). Line endings can be either \r\n or \n.
- The answering peer — In /root/rt/webrtc/peer.py, add the answer_peer(signal_url, room, out_path) coroutine. Attach to signal_url/room/ and wait for {"type": "offer", "sdp": ...}, and on receiving it, after setRemoteDescription → createAnswer → setLocalDescription, send {"type": "answer", "sdp": ...}. For messages on the data channel the other side opened, send them back with "echo:" prepended, and write the channel's label, ordered, and maxRetransmits as JSON to out_path. On receiving {"type": "bye"}, close the connection and finish.
- The offering peer — In /root/rt/webrtc/peer.py, add the offer_peer(signal_url, room, count) coroutine. Create a "chat" data channel and send the offer, receive the answer and call setRemoteDescription, and when the channel opens, send count messages one at a time starting from "m0", measuring the round-trip time each time an echo arrives. Return {"sent": count, "echoed": , "rtt_ms": }.
- A channel that discards late data — Create offer_peer's data channel with ordered=False and maxRetransmits=0. The grader's reference answerer records the properties of the channel it received and checks them.
- See where the candidates come from — In /root/rt/webrtc/peer.py, add the gather(stun=None) coroutine. If stun is given, use a single RTCIceServer(stun) as the ICE server, and otherwise an empty list, create one data channel, and measure the milliseconds spent in createOffer and setLocalDescription. Return {"elapsed_ms": , "candidates": [(typ, address, port), ...]}. The grader calls it three times: without STUN, with a STUN server in the same Pod, and with an unreachable STUN address.
Notes
- The working folder is /root/rt/webrtc. Create it first with mkdir -p /root/rt/webrtc.
- Start the signaling server with cd /root/rt/webrtc && /opt/rt-lab/bin/python -c "import asyncio, signaling; asyncio.run(signaling.run('127.0.0.1', 8765))". Do not name the file signal.py — it shadows the standard library's signal, and asyncio breaks first.
- The reference peer is /opt/fixtures/rt/webrtc/refpeer.py, and the STUN server in the Pod is /opt/fixtures/rt/webrtc/stunserver.py.
- Two common mistakes. Sending before the channel opens, and not specifying ICE servers so that getting the connection ready is delayed by several seconds waiting on a blocked public STUN.
- Always run Python with /opt/rt-lab/bin/python. This lab's libraries are only in that virtual environment, and if you run it with plain python3, you get a ModuleNotFoundError. It is convenient to shorten it with something like alias rpy=/opt/rt-lab/bin/python.
- The lab Pod blocks outbound connections. All communication happens on 127.0.0.1 inside the same Pod, and no installation or download is needed.
- The grader loads your code in a separate process and makes real connections. The example file is only a function skeleton, so it does not pass if left as is. Do not delete the functions you finished in earlier steps.
- When the lab session ends, the files in /root do not remain. Keep the code you need separately before you finish.
Build a signaling server
In /root/rt/webrtc/signaling.py, create the run(host, port) coroutine. Accept the /room/ path with a websockets server, and pass text messages through as they are between two connections attached to the same room. Messages that arrive while the other side is not there yet are held and passed on in order the moment the other side attaches. A third connection to attach is closed with close code 4001 and the reason "room full", and when one side leaves, send {"type": "bye"} to the remaining side.
The WebRTC specification does not define how to hand over the SDP (JSEP, RFC 9429, also leaves that part to the application). So almost every service keeps a small relay server like this. 4000–4999 is the range of close codes RFC 6455 left for applications.
Create an offer SDP
In /root/rt/webrtc/peer.py, create the make_offer() coroutine. Create an RTCPeerConnection with RTCConfiguration(iceServers=[]), which leaves the ICE server list empty, create a "chat" data channel, and then go through createOffer and setLocalDescription and return the localDescription.sdp string. Close the connection before returning.
For an m= line to appear in the SDP, you have to decide what to send (a data channel or a media track) before offering. If you do not specify ICE servers, aiortc uses a public STUN server, and where outbound UDP is blocked, candidate gathering stalls for several seconds waiting on its response.
Read an SDP
In /root/rt/webrtc/sdp.py, create summarize_sdp(sdp). Read an SDP string and return a dict. The keys are five: ufrag (the value of the first a=ice-ufrag), fingerprint (the hash name of the first a=fingerprint in lowercase, for example sha-256), setup (the value of the first a=setup), media (a list holding the medium type of each m= line in the order they appear), and candidates ({"host": n, "srflx": n, "relay": n} — the count per typ of the a=candidate lines, with kinds that are absent as 0). Line endings can be either \r\n or \n.
host is the address of its own interface, srflx (server reflexive) is the external address a STUN server saw, and relay is the relay address a TURN server lent. The fingerprint is the hash of the DTLS certificate, and the connection is made only if this value that came over signaling matches the certificate received during the handshake.
The answering peer
In /root/rt/webrtc/peer.py, add the answer_peer(signal_url, room, out_path) coroutine. Attach to signal_url/room/ and wait for {"type": "offer", "sdp": ...}, and on receiving it, after setRemoteDescription → createAnswer → setLocalDescription, send {"type": "answer", "sdp": ...}. For messages on the data channel the other side opened, send them back with "echo:" prepended, and write the channel's label, ordered, and maxRetransmits as JSON to out_path. On receiving {"type": "bye"}, close the connection and finish.
The data channel is created by the offering side, and the answering side receives it through the datachannel event. Leave the ICE server list empty on the answering peer too. The grader's reference offerer attaches through your signaling server.
The offering peer
In /root/rt/webrtc/peer.py, add the offer_peer(signal_url, room, count) coroutine. Create a "chat" data channel and send the offer, receive the answer and call setRemoteDescription, and when the channel opens, send count messages one at a time starting from "m0", measuring the round-trip time each time an echo arrives. Return {"sent": count, "echoed": , "rtt_ms": }.
If you send before the channel's open event, the message disappears. Only after the ICE check, the DTLS handshake, and the SCTP connection have all finished is it open. The time taken up to here is WebRTC's connection setup cost.
A channel that discards late data
Create offer_peer's data channel with ordered=False and maxRetransmits=0. The grader's reference answerer records the properties of the channel it received and checks them.
For data that is useless after 20ms, like a voice frame or a cursor position, it is better to drop it than to resend one lost piece and hold up everything behind it. SCTP lets you set order and retransmission per channel, and this is something you cannot imitate with WebSocket over TCP.
See where the candidates come from
In /root/rt/webrtc/peer.py, add the gather(stun=None) coroutine. If stun is given, use a single RTCIceServer(stun) as the ICE server, and otherwise an empty list, create one data channel, and measure the milliseconds spent in createOffer and setLocalDescription. Return {"elapsed_ms": , "candidates": [(typ, address, port), ...]}. The grader calls it three times: without STUN, with a STUN server in the same Pod, and with an unreachable STUN address.
If you split an a=candidate line by spaces, the fifth field is the address, the sixth is the port, and what follows typ is the kind. There is no NAT inside the same Pod, so first predict what the srflx address will turn out to be the same as.