A Successful ping Does Not Mean the Service Works
In one line
ping proves only reachability up to the IP layer. It does not prove the port above it, the application above that, or even whether a large packet can pass.
Why this exists
"Ping works but I can't connect" is not a contradiction but a normal state. ping merely exchanges ICMP echo. Whether TCP 8080 is open is a completely different question, and even if that port is open, the application can spit out 5xx.
The opposite direction matters too. Even if ping fails, the service can be perfectly fine. Many organizations block ICMP echo. So the judgment "ping doesn't work, so the server is dead" is dangerous.
How it works
The three tools have different roles.
| Tool | What it looks at | Caveat in this environment |
|---|---|---|
ping |
Whether an IP round trip to the destination works | Works (ping_group_range allows it) |
traceroute |
Up to which hop it goes | Only UDP mode works (ICMP mode needs a raw socket) |
mtr |
Continuously observing loss rate and latency per hop | Needs --udp |
The principle of traceroute is sending with a TTL raised from 1. The router at the point where the TTL hit 0 returns an ICMP time exceeded, and its source address is the identity of that hop. So traceroute is not a tool that "looks up the path" but one that "fails on purpose and collects the responses".
* * * is not an outage. It is just that the router at that hop limits or does not produce ICMP responses. The judgment rules are these.
- If loss is seen only at intermediate hops and there is no loss at the final destination, it is likely ICMP rate limiting
- If latency jumps from a particular hop, that segment is the bottleneck
- If loss increases in the last few hops, it is likely a real problem
What ping misses: path MTU
The nastiest case is an MTU black hole. The default ping payload is 56 bytes, so it always passes. But an API whose response gets just a bit bigger stalls. Not with an error — it just stalls. Because only packets over a threshold size disappear, the symptom splits by size.
You check with a ping that has the don't-fragment bit set.
ping -M do -s 1472 -c 1 10.20.0.5
# From 10.0.1.1 icmp_seq=1 Frag needed and DF set (mtu = 1420)
The -s value plus 28 (ICMP 8 + IP 20) is the actual MTU. If you find the maximum that passes by binary search, you get the path MTU.
If you use a tunnel, the effective MTU shrinks by the overhead.
| Encapsulation | Added header | Effective MTU | Recommended MSS clamp |
|---|---|---|---|
| PPPoE | 8 | 1492 | 1452 |
| GRE | 24 | 1476 | 1436 |
| VXLAN (IPv4) | 50 | 1450 | 1410 |
| WireGuard (IPv4) | 60 | 1440 | 1400 |
| IPsec ESP + NAT-T | About 81 | 1419 | 1379 |
What it looks like in the field
"When I turn on the VPN, only certain sites won't open." This is the typical situation where the effective MTU shrank because of tunnel overhead and there is no MSS clamping. Small pages open and large pages stall. Ping succeeds 100%.
Running traceroute after blocking all ICMP in the cloud. Everything comes out as * * *, and it is easy to conclude "the path is broken". In that case, it is more accurate to use TCP-mode traceroute (-T -p 443) or to knock directly on the target port with nc -zv.
Where you measure changes the answer
Even if you measure the same target with the same tool, the conclusion differs depending on where you measured. If you are not conscious of this, you end up arguing over values measured in different places.
A value measured on my laptop is not a value for our service. The office line, the VPN, and the home router are all in the path. It differs from what the user experiences and from server-to-server communication. So the first principle is to measure from the same position as where the problem occurred, and when you cannot, write the measurement position together with the result.
Measuring from only one side is half. As we saw in the earlier course, it is common that there is a path going out but
none coming back, and then a traceroute seen from one side looks normal.
If you measure toward each other from both sides, the asymmetry shows up immediately.
The path is not the same every time. In segments with several paths, each flow
takes a different route, so if you run traceroute twice, the hops come out different. A problem where loss shows up only in a particular
flow is not caught with a single measurement.
The measurement itself creates load. If you attach mtr to a link that is already saturated, it adds to
the loss. And as said in the earlier course, it really happens that a diagnostic command on an already ailing system
makes the outage worse.
Finally, write down when you measured. Network problems appear and disappear over time, so a measurement without a timestamp is no evidence at all a few days later. Measuring values in advance when things are normal is valuable for the same reason. If you have nothing to compare against, you cannot judge whether the current value is bad.
What you will do in the next lab
You check interfaces and routing, read the loss rate and RTT from the ping statistics, and run UDP-mode traceroute and mtr. At the end you compare ICMP works, but how about TCP yourself and make a table.