TT Lab
Get started
Learn Learning paths Courses

Kubernetes Networking — On a Real Cluster

The API server does not know Pod addresses

Continue in TT Lab

In one line

Kubernetes does not hand out Pod addresses itself. The API server manages only Service addresses, and the network plugin on the node gives Pods their addresses. That is why "the Pod will not start" and "Pods cannot reach each other" are failures in different layers.

Why this was needed

The first thing you run into when you put several containers on one machine is ports. If three teams each want to use 8080, somebody has to give way, the side that gives way has to turn the port into a variable in its configuration, and then service discovery has to know the port too. The official documentation sums up this path as one that becomes very hard to coordinate at scale and creates cluster-level problems the user cannot control.

Kubernetes did not go down that path. Instead it gives each Pod one address. Inside a Pod, containers meet over localhost, and outside the Pod they meet by address. The problem of coordinating ports disappears altogether.

The question is who gives out that address, and how. It differs by cloud, it differs again on-premises, and some places solve it with routing and some with tunnels. Kubernetes left this slot empty and defined only the specification. That is CNI (Container Network Interface).

How it works

The official documentation is clear: a CNI plugin is required to implement the Kubernetes network model. The plugin must be compatible with CNI specification v0.4.0 or later, and the project recommends compatibility with v1.0.0.

Who calls it. Before 1.24, the kubelet managed plugins with the cni-bin-dir and network-plugin options. Those two options were removed in 1.24. What calls CNI now is not the kubelet but the container runtime (containerd, CRI-O). So half of the reports saying "I changed the CNI configuration and it has no effect" come from someone who restarted the kubelet.

Where it reads. The configuration is read from /etc/cni/net.d by default, and the binaries from /opt/cni/bin by default. The format that lists several plugins in one configuration file is conflist, and the plugins are called one after another from the top.

Why several. Because no single plugin does everything. The official documentation itself gives two examples.

Feature Plugin responsible How to turn it on
hostPort portmap portMappings: true in capabilities
Bandwidth limit bandwidth bandwidth: true in capabilities, plus the kubernetes.io/ingress-bandwidth annotation on the Pod
Loopback lo loopback The runtime has to provide it for every sandbox

This leads to an operational fact. You can write hostPort and nothing happens. The manifest passes API validation and the Pod becomes Running, but the port is not opened. That is what happens if portmap is not in the chain. It is the kind of failure that is silently ignored.

Addresses come from IPAM. The ipam entry inside a main plugin such as bridge does that job. The most common one, host-local, manages addresses only within that node, as its name says. It does not ask what the other nodes handed out. So you have to carve out non-overlapping ranges for each node in advance, and that range is spec.podCIDR of the node object. That is what it means when the official example is written as "subnet": "usePodCidr".

Three kinds of addresses. The official documentation states firmly that the cluster must assign Pod, Service and node addresses without overlap, and it also lists who decides each one: for Pods, the network plugin; for Services, kube-apiserver; for nodes, the kubelet or the cloud controller. The three parties do not consult one another, so preventing overlap is up to the designer.

What it looks like in the field

"Only one node will not accept Pods." Kubernetes expresses this situation as a taint. The official documentation's list of built-in taints includes node.kubernetes.io/network-unavailable, meaning "the node's network is unusable". So the report arrives as "Pod is Pending", but the place to fix is not the scheduler but the plugin configuration on that node. The plugin configuration file is missing only on that node, or the binary directory is different, or the runtime is looking at another path.

"The Pod IP is the same on two nodes." host-local does not know about anything outside the node, so if you give two nodes the same range, it calmly hands out the same address. The symptom shows up as random connection failures, so it takes a long time to find the cause.

"We set the range wrong, so we will change it." Once a value is in a node's spec.podCIDR, it cannot be changed to another value. The hardest decision to undo in cluster design is the address range.

Official documentation: Network Plugins · Cluster Networking

What you will do in the next lab

You pin a /24 Pod range to each of three nodes, and calculate which comes first: the number of Pods that range can hold, or the number of Pods the node accepts. Next you write a conflist that chains the three plugins bridge, portmap and bandwidth, create a Pod with hostPort and a bandwidth annotation, and write down which plugin handles each job. Finally you imitate a node with no plugin to create a NotReady condition, and decide by calculation which of four range layouts overlap.