TT Lab
Get started
Learn Learning paths Courses

Networking Fundamentals

Speak HTTP by hand

Continue in TT Lab

Goal

HTTP is not difficult. It is just exchanging agreed-upon text over TCP. But because libraries keep that agreement for you, when something goes out of line, you can no longer read what went out of line.

Here you keep that agreement by hand and look, one by one, at the things libraries hide.

Materials

python3 /opt/lab/netplan/httpsrv.py 8080 &

It is a small server that deliberately answers differently on each path.

/            Content-Length 로 길이를 알려 준다
/chunked     길이를 모른 채 조각으로 나눠 보낸다
/slow        첫 조각을 보내고 잠시 쉰다
/nolength    길이도 조각도 없이, 연결을 닫아 끝을 알린다
/moved  /seeother  /keepme    301 · 303 · 307

What to leave

01-raw.txt        손으로 친 요청과 그 응답
02-nohost.txt     Host 를 뺐을 때와 1.0 으로 보냈을 때
03-length.txt     본문의 끝을 알리는 세 방법
04-keepalive.txt  한 연결에 요청 두 번
05-redirect.txt   301 · 303 · 307 의 차이
06-timing.txt     구간별 시간
07-notes.md       다음에 읽을 사람에게

Type the request by hand

Without a library, write the HTTP request string yourself over a TCP connection, receive the response for /, and save it to 01-raw.txt. Save both what you sent and what you received.

Bash alone is enough.

exec 3<>/dev/tcp/127.0.0.1/8080
printf 'GET / HTTP/1.1\r\nHost: localhost\r\nConnection: close\r\n\r\n' >&3
timeout 3 cat <&3

Line endings must be \r\n, and there must be one more blank line at the end of the headers. That blank line is the signal that "the headers end here".

Try leaving out Host

Send an HTTP/1.1 request without the Host header, and send the same request as HTTP/1.0 too, and save the results to 02-nohost.txt. Write one line on why it became mandatory from 1.1.

In 1.1, 400 Bad Request is the rule. It is the specification, not the server's taste.

It became mandatory when hosting several sites on one address became possible. The server can tell which site a request was sent to only from Host.

Where does the body end

Receive each of /, /chunked, and /nolength, and save three ways of signaling the end of the body to 03-length.txt. Do not process the chunked one; save it as the original bytes exactly.

curl merges chunked for you automatically, so the original is not visible. Receive it directly with /dev/tcp.

Each chunk is preceded by a hexadecimal size line, and at the end a chunk of size 0 signals the end. Without that 0, the receiving side keeps waiting.

Twice on one connection

Keep one connection open, send two requests, receive the two responses, and save them to 04-keepalive.txt. Write one line on what is saved by reusing the connection.

If you leave out Connection: close, the server keeps the connection. After reading the first response for the length given in Content-Length, send the second request on the same fd.

If you do it in Python, it is easy to read exactly that length.

How do 301, 303, and 307 differ

Receive each of the three paths, save the responses to 05-redirect.txt, and write the differences among the three. In particular, write whether the method may be changed when requesting again.

All three give a Location, but the way of requesting again differs.

When a payment request meets a redirect, which of these it is separates an accident from normal operation.

Point at the slow segment with numbers

Measure each of / and /slow with curl's timing variables, save them to 06-timing.txt, and write which segment grew longer. Also write that these values are cumulative.

curl -sS -o /dev/null -w 'dns=%{time_namelookup} tcp=%{time_connect} ttfb=%{time_starttransfer} total=%{time_total}\n' http://127.0.0.1:8080/

The numbers are cumulative. If time_connect is 0.32, it means it took 0.32 seconds up to the connection, not that the connection itself took 0.32 seconds. You get the time per segment by subtracting.

/slow rests after sending the first byte. Then ttfb is fast and only total grows.

For the next person who reads this

Pick four or more of the things you saw here and summarize them in 07-notes.md. Write not what you did but why it is so.

Write as if your own self a few months later is reading. "I received /chunked" does not help, but "when the length is not known in advance, send in chunks and signal the end with a chunk of size 0 — without that 0, the receiving side keeps waiting" does.