Breaking "It Doesn't Work" Into Four Layers
In one line
90% of network outage diagnosis ends with deciding first "at which layer did it stop". If you turn on tools without deciding the layer, only time passes.
Why this exists
When a report comes in saying "the API won't connect", people usually do this. They try a ping, open up firewall rules, check DNS, and dig through application logs. There is no order. So they look at the same place twice and keep not looking at the places they have not looked at.
An experienced person does it differently. First they pin down the layer. And they use just one tool that fits that layer.
How it works
In practice, four layers are enough, not the OSI seven layers.
| Layer | Question | Tools to check | What failure looks like |
|---|---|---|---|
| Name | Which IP does this name resolve to | getent hosts, dig, resolvectl query |
Name or service not known |
| Reach | Do packets get to that IP | ping, traceroute -U, ip route |
Network is unreachable, no response |
| Port | Is that port on that IP open | nc -zv, ss -ltn, curl --connect-timeout |
Connection refused / timed out |
| Response | It is open but is the answer odd | curl -v, curl -w, application logs |
5xx, slowness, wrong body |
The key is that a failure at each layer has a different face. Just reading that face accurately decides the layer.
If you memorize the table that tells them apart by errno, the layer splits from the wording of the report alone.
| Message | errno | What actually happened | Where to look first |
|---|---|---|---|
Name or service not known |
EAI_NONAME | It failed at name resolution. It did not even open a socket | nsswitch.conf, resolv.conf, /etc/hosts |
Network is unreachable |
ENETUNREACH | There is no route in the routing table | ip route, especially IPv6 routes |
No route to host |
EHOSTUNREACH | Received an ICMP unreachable or no answer to ARP | Target power, ARP, REJECT rules |
Connection refused |
ECONNREFUSED | The destination sent back an RST | The process on the target host and its binding address |
Connection timed out |
ETIMEDOUT | No response even after all the SYN retransmissions | Firewall DROP, security group, accept queue |
The thing people most often get wrong here is blaming the firewall for refused. refused means an RST was received, and if an RST was received, it is evidence that the packet has already passed routing, the security group, NAT, and the network policy. Production firewalls almost always have a DROP policy, so a blocked port shows up as a timeout. Digging through the firewall after seeing refused is rechecking a gate already passed.
Time is evidence too. With the Linux default (net.ipv4.tcp_syn_retries=6), a connection with no response gives up after 1+2+4+8+16+32+64 = 127 seconds. If the application timeout is written as 30 seconds but it actually failed after 127 seconds, that is a signal the timeout setting is not being applied.
What it looks like in the field
"It works locally but only fails remotely." Nine times out of ten it is a binding address problem. If it is listening on 127.0.0.1:8080 in ss -ltnp, the kernel rejects a connection coming to the Pod IP with an RST — refused. Conversely, if it times out even to the Pod IP, then you start looking at network policy. This single fork halves the scope of the investigation.
What we will do next
We start with the lowest layer, name resolution. You check by hand why /etc/hosts and nsswitch.conf come before dig, and why dig does not look at either of them at all.