TT Lab
Get started
Learn Learning paths Courses

Network Troubleshooting

Breaking "It Doesn't Work" Into Four Layers

Continue in TT Lab

In one line

90% of network outage diagnosis ends with deciding first "at which layer did it stop". If you turn on tools without deciding the layer, only time passes.

Why this exists

When a report comes in saying "the API won't connect", people usually do this. They try a ping, open up firewall rules, check DNS, and dig through application logs. There is no order. So they look at the same place twice and keep not looking at the places they have not looked at.

An experienced person does it differently. First they pin down the layer. And they use just one tool that fits that layer.

How it works

In practice, four layers are enough, not the OSI seven layers.

Layer Question Tools to check What failure looks like
Name Which IP does this name resolve to getent hosts, dig, resolvectl query Name or service not known
Reach Do packets get to that IP ping, traceroute -U, ip route Network is unreachable, no response
Port Is that port on that IP open nc -zv, ss -ltn, curl --connect-timeout Connection refused / timed out
Response It is open but is the answer odd curl -v, curl -w, application logs 5xx, slowness, wrong body

The key is that a failure at each layer has a different face. Just reading that face accurately decides the layer.

If you memorize the table that tells them apart by errno, the layer splits from the wording of the report alone.

Message errno What actually happened Where to look first
Name or service not known EAI_NONAME It failed at name resolution. It did not even open a socket nsswitch.conf, resolv.conf, /etc/hosts
Network is unreachable ENETUNREACH There is no route in the routing table ip route, especially IPv6 routes
No route to host EHOSTUNREACH Received an ICMP unreachable or no answer to ARP Target power, ARP, REJECT rules
Connection refused ECONNREFUSED The destination sent back an RST The process on the target host and its binding address
Connection timed out ETIMEDOUT No response even after all the SYN retransmissions Firewall DROP, security group, accept queue

The thing people most often get wrong here is blaming the firewall for refused. refused means an RST was received, and if an RST was received, it is evidence that the packet has already passed routing, the security group, NAT, and the network policy. Production firewalls almost always have a DROP policy, so a blocked port shows up as a timeout. Digging through the firewall after seeing refused is rechecking a gate already passed.

Time is evidence too. With the Linux default (net.ipv4.tcp_syn_retries=6), a connection with no response gives up after 1+2+4+8+16+32+64 = 127 seconds. If the application timeout is written as 30 seconds but it actually failed after 127 seconds, that is a signal the timeout setting is not being applied.

What it looks like in the field

"It works locally but only fails remotely." Nine times out of ten it is a binding address problem. If it is listening on 127.0.0.1:8080 in ss -ltnp, the kernel rejects a connection coming to the Pod IP with an RST — refused. Conversely, if it times out even to the Pod IP, then you start looking at network policy. This single fork halves the scope of the investigation.

What we will do next

We start with the lowest layer, name resolution. You check by hand why /etc/hosts and nsswitch.conf come before dig, and why dig does not look at either of them at all.