TT Lab
Get started
Learn Learning paths Courses

Networking Fundamentals

MTU and Paths — Communication That Fails Only Because of Size

Continue in TT Lab

In a nutshell

TCP decides the segment size by looking only at its own interface's MTU, so it does not know about narrow segments in the middle of the path, and if you block the ICMP that tells it about them, a black hole arises in which only large packets silently vanish.

Why this was needed

There is an outage with a particularly peculiar set of symptoms.

With this combination, there is no need to look at anything else. It is a path MTU problem. And it usually appears right after a VPN, tunnel, or overlay network has been newly attached or the path has changed.

How it works

The MTU is the maximum size of an IP packet that one interface can send out at one time. The Ethernet default is 1500 bytes. From this, TCP subtracts the 20-byte IPv4 header and the 20-byte TCP header to set an MSS of 1460 bytes, and announces it to the peer as an option in the handshake's SYN packet.

The decisive fact is this. Both sides set the MSS by looking only at their own interface's MTU. Nobody knows whether there is a narrower segment in the middle of the path. Even if two hosts agree to exchange 1460 bytes, if there is a 1420-byte tunnel in the middle, that agreement cannot be kept.

This is why small requests succeed. The health check response ends in one segment and passes through the narrow segment safely. Only packets that exceed the threshold vanish, so the symptom divides by size. The same goes for stalling at the TLS handshake. The ClientHello is small, but the certificate chain the server sends is several thousand bytes, so it immediately becomes maximum-size segments.

TCP originally has a mechanism to solve this problem. It is path MTU discovery (PMTUD). Linux sets the DF (don't fragment) bit on TCP packets, and when an intermediate router receives a DF packet larger than its next-hop MTU, it drops it and tells the next-hop MTU with ICMP type 3 code 4 (Fragmentation Needed). The sender caches that value, cuts the segments smaller, and resends.

The problem is when this ICMP is blocked. The sender, knowing nothing, keeps sending large packets, and those packets keep getting dropped. Retransmissions are the same size, so they are dropped again. No error occurs. The connection stays ESTABLISHED, and the application ends only when it hits its own timeout. This is the PMTUD black hole.

Here we must correct a widespread misconception. The policy "block all ICMP for security" is itself a cause of outages. What can be blocked is about echo request/reply, and type 3 code 4 must be let through. In IPv6 there is not even a choice. IPv6 routers do not fragment packets, so ICMPv6 Packet Too Big (type 2) is essential.

Memorizing the bytes that encapsulation takes makes diagnosis faster. VXLAN takes 50 bytes (effective MTU 1450), WireGuard 60–80 bytes (1440–1420), GRE 24 bytes (1476), and IPsec ESP up to about 73 bytes.

What it looks like in the field

In Kubernetes, this problem recurs regularly because of overlay networks. If the node interface is 1500 and the VXLAN interface and the Pod's eth0 are also set to 1500, it is an immediate black hole. VXLAN takes 50 bytes, so it must be 1450. The symptom is characteristic. Between Pods small requests work but only POSTs with large bodies stall, and kubectl exec works but kubectl logs of a Pod with a lot of logs stalls midway.

The realistic solution is to apply MSS clamping at the tunnel entry point. The router itself rewrites lower the value of the MSS option in passing SYN packets, so it does not depend on ICMP at all. However, it applies only to TCP, and UDP-based protocols such as QUIC must adjust their size themselves.

What to check in the quiz that follows

Check whether you can immediately think of the MTU from the symptom "small requests work but only large responses stall", and whether you can explain why it stalls without an error.