TT Lab
Get started
Learn Learning paths Courses

Network Troubleshooting

A Successful ping Does Not Mean the Service Works

Continue in TT Lab

In one line

ping proves only reachability up to the IP layer. It does not prove the port above it, the application above that, or even whether a large packet can pass.

Why this exists

"Ping works but I can't connect" is not a contradiction but a normal state. ping merely exchanges ICMP echo. Whether TCP 8080 is open is a completely different question, and even if that port is open, the application can spit out 5xx.

The opposite direction matters too. Even if ping fails, the service can be perfectly fine. Many organizations block ICMP echo. So the judgment "ping doesn't work, so the server is dead" is dangerous.

How it works

The three tools have different roles.

Tool What it looks at Caveat in this environment
ping Whether an IP round trip to the destination works Works (ping_group_range allows it)
traceroute Up to which hop it goes Only UDP mode works (ICMP mode needs a raw socket)
mtr Continuously observing loss rate and latency per hop Needs --udp

The principle of traceroute is sending with a TTL raised from 1. The router at the point where the TTL hit 0 returns an ICMP time exceeded, and its source address is the identity of that hop. So traceroute is not a tool that "looks up the path" but one that "fails on purpose and collects the responses".

* * * is not an outage. It is just that the router at that hop limits or does not produce ICMP responses. The judgment rules are these.

What ping misses: path MTU

The nastiest case is an MTU black hole. The default ping payload is 56 bytes, so it always passes. But an API whose response gets just a bit bigger stalls. Not with an error — it just stalls. Because only packets over a threshold size disappear, the symptom splits by size.

You check with a ping that has the don't-fragment bit set.

ping -M do -s 1472 -c 1 10.20.0.5
# From 10.0.1.1 icmp_seq=1 Frag needed and DF set (mtu = 1420)

The -s value plus 28 (ICMP 8 + IP 20) is the actual MTU. If you find the maximum that passes by binary search, you get the path MTU.

If you use a tunnel, the effective MTU shrinks by the overhead.

Encapsulation Added header Effective MTU Recommended MSS clamp
PPPoE 8 1492 1452
GRE 24 1476 1436
VXLAN (IPv4) 50 1450 1410
WireGuard (IPv4) 60 1440 1400
IPsec ESP + NAT-T About 81 1419 1379

What it looks like in the field

"When I turn on the VPN, only certain sites won't open." This is the typical situation where the effective MTU shrank because of tunnel overhead and there is no MSS clamping. Small pages open and large pages stall. Ping succeeds 100%.

Running traceroute after blocking all ICMP in the cloud. Everything comes out as * * *, and it is easy to conclude "the path is broken". In that case, it is more accurate to use TCP-mode traceroute (-T -p 443) or to knock directly on the target port with nc -zv.

Where you measure changes the answer

Even if you measure the same target with the same tool, the conclusion differs depending on where you measured. If you are not conscious of this, you end up arguing over values measured in different places.

A value measured on my laptop is not a value for our service. The office line, the VPN, and the home router are all in the path. It differs from what the user experiences and from server-to-server communication. So the first principle is to measure from the same position as where the problem occurred, and when you cannot, write the measurement position together with the result.

Measuring from only one side is half. As we saw in the earlier course, it is common that there is a path going out but none coming back, and then a traceroute seen from one side looks normal. If you measure toward each other from both sides, the asymmetry shows up immediately.

The path is not the same every time. In segments with several paths, each flow takes a different route, so if you run traceroute twice, the hops come out different. A problem where loss shows up only in a particular flow is not caught with a single measurement.

The measurement itself creates load. If you attach mtr to a link that is already saturated, it adds to the loss. And as said in the earlier course, it really happens that a diagnostic command on an already ailing system makes the outage worse.

Finally, write down when you measured. Network problems appear and disappear over time, so a measurement without a timestamp is no evidence at all a few days later. Measuring values in advance when things are normal is valuable for the same reason. If you have nothing to compare against, you cannot judge whether the current value is bad.

What you will do in the next lab

You check interfaces and routing, read the loss rate and RTT from the ping statistics, and run UDP-mode traceroute and mtr. At the end you compare ICMP works, but how about TCP yourself and make a table.