Network Fundamentals — Hands-on in a Linux VM
HTTP Is Text — and Check the Layers in Order
In one line
An HTTP request is text in three parts, the request line, headers, and a blank line, so you can write it by hand with nc. The status code of the response tells you "how far it got" — 404 is an answer after both TCP and HTTP succeeded, and refused means there is no server.
Why this exists
It is convenient as long as browsers and SDKs hide HTTP from you. But when a proxy swaps out headers, when a wrong Host brings up a different site, or when a redirect loops A→B→A, you have to look at what was hidden. Someone who has written a request by hand once knows what the > and < in curl -v are. That is half of API debugging.
And network diagnosis is, in the end, a job of separating the layers in order. Link → address → neighbor → route → name → port → firewall → response. Each layer has a command that looks only at that layer. If you type commands in no particular order, you get trapped in "everything looks fine but it doesn't work".
How it works
A request is one line, GET /docs/guide.txt HTTP/1.1, a few header lines such as Host: localhost, and a blank line. Lines end in CRLF. The blank line means "end of headers", so if you leave it out, the server waits for more headers. Host is the only required header in HTTP/1.1 — one server hosts several sites, so this is how it picks which one. Connection: close means close after the response, and without it the connection is reused (keep-alive).
A response is a status line such as HTTP/1.0 200 OK, headers, a blank line, and a body. 2xx is success, 3xx means it is somewhere else (follow Location), 4xx means the request side is at fault, and 5xx means the server side is at fault. HEAD receives only the headers, and requesting a directory without a slash gets a 301 pointing to the path with a slash.
Written as commands, the diagnostic order looks like this. ip -br link (is the link UP) → ip -br addr (is there an address) → ip neigh (is the neighbor FAILED) → ip route get (where does it go out) → getent hosts (does the name resolve; ask the server separately with dig) → ss -ltn (is the port open) → the counters in nft list ruleset (is the firewall counting) → curl -v (what is the response). The later steps mean something only if the earlier steps are right.
What it looks like in the field
Ping works but it won't open. Ping is ICMP and the service is TCP. A firewall treats each protocol differently. Conversely, ping failing does not mean the connection is cut — the last lab of this course blocks ICMP by design. Diagnose with the same protocol and port as the real service.
dig works but curl doesn't. An old address remains in /etc/hosts. dig asks the server directly, while curl goes through libc and looks at hosts first. The same name has two answers.
Telling the layers apart by status code
When you get a report that "it doesn't work", a single returned result lets you judge how far it got. If you memorize this mapping, the scope of the investigation narrows immediately.
| Result | How far it got | Where to look next |
|---|---|---|
| Name not found | Did not get past DNS | getent hosts, dig, /etc/hosts |
| Connection refused | Reached the peer, but nobody is on that port | Whether the service is up, with ss -ltn |
| Timeout (no response) | Silently dropped somewhere | Firewall, security group, route |
| TLS handshake failure | TCP succeeded | Certificate expiry, name, intermediate certificate |
| 502 / 504 | Succeeded as far as the proxy; the problem is behind it | The upstream service and its timeout |
| 404 | Everything succeeded through HTTP | Path, Host header, routing rules |
| 401 / 403 | The server recognized me and refused | Credentials and permissions |
The difference between refused and a timeout is especially valuable. Refused means the peer host is alive and immediately answered "there is no such port", so the network path is fine, while a timeout means no answer came at all, so it was likely dropped along the way. The place to suspect a firewall is the latter, not the former.
You also have to tell 502 and 504 apart. 502 is when it connected to the upstream but got a strange answer or the connection was cut, and 504 is when the upstream did not answer in time. The former appears when the upstream is dead or restarting, and the latter when the upstream is alive but slow. If the proxy's timeout is shorter than the processing time of the service behind it, even normal requests become 504, so the basic practice is to tune the two values together.
Finally, when you test with curl, you have to recreate the same conditions as the real client. The Host header, the protocol (HTTP/1.1 or 2), and the proxy environment variables. A good share of situations where curl works on the server but the browser doesn't are due to one of these three.
What you will do in the next lab
You start python3 -m http.server, build GET, HEAD, 404, and 301 requests by hand with nc and receive the responses, and compare them with curl -v. Then in the comprehensive lab you set up two subnets, a router, dnsmasq, a forward firewall, and masquerade on one VM, and leave a report on the order in which you would examine an "it won't open" report.