TT Lab
Get started
Learn Learning paths Courses

Real-Time Communication — WebSocket, gRPC Streaming and WebRTC

Stack requests on one HTTP/2 connection and measure the cost

Continue in TT Lab

Goal

Check HTTP/2 frames and header compression at the byte level, build a client that overlays requests on one connection, and measure yourself the difference from HTTP/1.1 and TCP-level HOL blocking.

Why it matters

HTTP/2 is a protocol made to carry several requests at once over one connection. In exchange, you bet everything on one connection, so a single bug that does not return the flow-control window or a single segment loss stops every stream of that connection. gRPC runs on HTTP/2, and a good share of the failures where a streaming response suddenly stalls come from this window. This lab measures the benefit and the cost with the same tool and leaves in numbers which side wins and when.

Steps

  1. Read the 9-byte frame header — In /root/rt/h2/h2lab.py, create frame_header(data). Read the first 9 bytes of the bytes as an HTTP/2 frame header and return a tuple of four ints: (length, type, flags, stream id). The length is the unsigned big-endian integer of the first 3 bytes, and the stream id is the 31 bits of the last 4 bytes minus the leading reserved bit. It is a ValueError if shorter than 9 bytes, and bytes after the first 9 are ignored.
  2. Sending the same headers twice makes them smaller — In /root/rt/h2/h2lab.py, add hpack_sizes(headers, times). Encode headers, a list of (name, value) tuples, times times with hpack.Encoder, and return the byte count of each encoding result as a list. This imitates a connection sending requests several times, so create only one encoder and keep using it.
  3. Overlay requests on one connection — In /root/rt/h2/h2lab.py, add fetch_all(host, port, paths, timeout=10). Open one TCP connection to host:port, send the preface with the h2 library's client connection, and then send all the GET requests in paths without waiting for responses. Collect each as its response ends, and return a list of tuples (path, status int, body bytes, milliseconds float from start to finish) in the same order as paths. Give the window back for received DATA with acknowledge_received_data, and raise an exception if it does not finish within timeout seconds.
  4. Measure with the same ruler as HTTP/1.1 — Run /opt/rt-lab/bin/python /opt/fixtures/rt/h2/bench.py compare. It prints the time to send six requests that each take 300ms one after another on a single HTTP/1.1 keep-alive connection, and the time to send them with your fetch_all. Write the two output lines h1_total_ms and h2_total_ms as they are into /root/rt/h2/report.txt. The grader re-measures on the spot and compares.
  5. Respect the concurrent stream limit the server advertised — Fix fetch_all so that after receiving the server's SETTINGS, it opens streams without exceeding remote_settings.max_concurrent_streams. If the limit is full, the remaining requests wait and are opened as streams end. The grader sends five requests to a server that advertises a limit of 2, and checks that it never exceeds the limit while still sending two at a time overlapped.
  6. Measure where the flow-control window closes — In /root/rt/h2/h2lab.py, add window_probe(host, port, path, window, quiet=0.3). After the preface on a new connection, send SETTINGS announcing INITIAL_WINDOW_SIZE as window, and send GET path. Count the received DATA without giving the window back, and when nothing arrives for quiet seconds, record the bytes received until then, give back the window for the amount received, and read to the end. Return (bytes received before it stopped, total bytes).
  7. If one TCP line is blocked, every stream stops — Run /opt/rt-lab/bin/python /opt/fixtures/rt/h2/bench.py hol. With a relay in between that stops once for 800ms after 20000 bytes have passed downstream, it requests a large response and a small response together and measures the time the small response ends. HTTP/2 sends them on one connection with your fetch_all, and HTTP/1.1 sends them split over two connections. Add the two output lines hol_h2_ms and hol_h1_ms to /root/rt/h2/report.txt.

Notes

Read the 9-byte frame header

In /root/rt/h2/h2lab.py, create frame_header(data). Read the first 9 bytes of the bytes as an HTTP/2 frame header and return a tuple of four ints: (length, type, flags, stream id). The length is the unsigned big-endian integer of the first 3 bytes, and the stream id is the 31 bits of the last 4 bytes minus the leading reserved bit. It is a ValueError if shorter than 9 bytes, and bytes after the first 9 are ignored.

The shape of the frame header is exactly the figure in RFC 9113 section 4.1. The reserved bit must be 0 when sending, but it says to ignore it when receiving. Even if it arrives set, the stream id must not change.

Sending the same headers twice makes them smaller

In /root/rt/h2/h2lab.py, add hpack_sizes(headers, times). Encode headers, a list of (name, value) tuples, times times with hpack.Encoder, and return the byte count of each encoding result as a list. This imitates a connection sending requests several times, so create only one encoder and keep using it.

HPACK's dynamic table is one per connection and builds up while the connection is alive. The encoder object is that table. If the second result is the same as the first, look at what you are creating anew.

Overlay requests on one connection

In /root/rt/h2/h2lab.py, add fetch_all(host, port, paths, timeout=10). Open one TCP connection to host:port, send the preface with the h2 library's client connection, and then send all the GET requests in paths without waiting for responses. Collect each as its response ends, and return a list of tuples (path, status int, body bytes, milliseconds float from start to finish) in the same order as paths. Give the window back for received DATA with acknowledge_received_data, and raise an exception if it does not finish within timeout seconds.

h2 does not know about sockets. You take the bytes to send out with data_to_send() and send them yourself, and you put received bytes into receive_data() and get events back. If you don't give the window back, one large response stops the whole connection at 64KiB. The grading requests include a 200000-byte response.

Measure with the same ruler as HTTP/1.1

Run /opt/rt-lab/bin/python /opt/fixtures/rt/h2/bench.py compare. It prints the time to send six requests that each take 300ms one after another on a single HTTP/1.1 keep-alive connection, and the time to send them with your fetch_all. Write the two output lines h1_total_ms and h2_total_ms as they are into /root/rt/h2/report.txt. The grader re-measures on the spot and compares.

HTTP/1.1 handles requests on one connection only in response order. Pipelining is in the specification, but browsers and proxies effectively do not use it. So the six 300ms periods simply add up.

Respect the concurrent stream limit the server advertised

Fix fetch_all so that after receiving the server's SETTINGS, it opens streams without exceeding remote_settings.max_concurrent_streams. If the limit is full, the remaining requests wait and are opened as streams end. The grader sends five requests to a server that advertises a limit of 2, and checks that it never exceeds the limit while still sending two at a time overlapped.

Before receiving SETTINGS, you do not know the limit. By the specification, the default at that time is unlimited, so if you rush requests out right after sending the preface, the server refuses. The first SETTINGS arrives as a RemoteSettingsChanged event.

Measure where the flow-control window closes

In /root/rt/h2/h2lab.py, add window_probe(host, port, path, window, quiet=0.3). After the preface on a new connection, send SETTINGS announcing INITIAL_WINDOW_SIZE as window, and send GET path. Count the received DATA without giving the window back, and when nothing arrives for quiet seconds, record the bytes received until then, give back the window for the amount received, and read to the end. Return (bytes received before it stopped, total bytes).

There is one flow-control window per stream and one for the whole connection. What you can change with SETTINGS is only the stream's initial value, and the connection window starts at 65535. If you give window as 100000, work out first where it will stop.

If one TCP line is blocked, every stream stops

Run /opt/rt-lab/bin/python /opt/fixtures/rt/h2/bench.py hol. With a relay in between that stops once for 800ms after 20000 bytes have passed downstream, it requests a large response and a small response together and measures the time the small response ends. HTTP/2 sends them on one connection with your fetch_all, and HTTP/1.1 sends them split over two connections. Add the two output lines hol_h2_ms and hol_h1_ms to /root/rt/h2/report.txt.

The relay stopping imitates the way TCP does not pass later bytes up to the application while one lost segment waits for retransmission. A stream is an HTTP/2 concept, and TCP does not know about it.