TT Lab
Get started
Learn Learning paths Courses

Kubernetes Networking — On a Real Cluster

Carve Pod CIDRs and design the plugin chain

Continue in TT Lab

Goal

Confirm where Pod addresses come from, from both the node object and the plugin configuration, and work out by hand how to carve ranges so they do not overlap.

Why it matters

The address range is the first thing you have to decide when you set up a new cluster, and the hardest to undo. Pod, Service and node addresses are each assigned by a different party, and they do not consult one another, so keeping them from overlapping is the designer's job. Also, some features, such as hostPort or bandwidth limits, are written in the manifest but silently ignored if the plugin is missing, so unless you sort out once who is responsible for what, you will wander for a long time with "the configuration is right but it does not work".

Steps

  1. Split the cluster Pod range 10.244.0.0/16 into /24 blocks and pin them to lab-node-0, lab-node-1 and lab-node-2 in order from the front (both spec.podCIDR and spec.podCIDRs). Then in /root/k8nd-cni/01-podcidr.txt, write three lines of 노드이름 대역 (the placeholders are the node name and its range), in node-name order.
  2. In /root/k8nd-cni/02-capacity.txt, write five lines — subnets= (the number of pieces when you cut 10.244.0.0/16 into /24), addresses_per_subnet= (the total number of addresses in one piece), usable_per_subnet= (the count after removing the network address and the gateway), node_pod_limit= (the status.allocatable.pods value the node object reports), binding_limit= (whichever of the two is hit first, as subnet or node).
  3. Create /root/k8nd-cni/10-k8nd.conflist — cniVersion is 1.0.0, name is k8nd-pod-network, and plugins has exactly three entries, in order bridge, portmap and bandwidth. The ipam of bridge is host-local and its subnet is the Pod range of lab-node-0, portmap must have capabilities.portMappings true, and bandwidth must have capabilities.bandwidth true.
  4. Write a Pod manifest to /root/k8nd-cni/shaped.yaml and apply it — namespace k8nd-cni, name shaped, annotations kubernetes.io/ingress-bandwidth: 1M and kubernetes.io/egress-bandwidth: 1M, and a container web (nginx:1.27-alpine) with containerPort 8080 and hostPort 8080. Then in /root/k8nd-cni/04-chain.txt, write four lines hostport_plugin=, bandwidth_plugin=, ingress_annotation= and egress_annotation= (for the plugins, the type of the plugin that handles that job in the step 3 conflist; for the annotations, the values you actually set).
  5. Write the node k8nd-cni-edge to /root/k8nd-cni/edge-node.yaml and apply it — set spec.unschedulable to true, and add one taint of your own to spec.taints with the key k8nd.example.com/cni-missing (effect NoSchedule). After you apply it, in /root/k8nd-cni/05-notready.txt write four lines: node=, custom_taint= (the key you added), controller_taint= (the built-in taint key the controller added on top), and unschedulable=.
  6. Create /root/k8nd-cni/06-node1.conflist and /root/k8nd-cni/06-node2.conflist. Each uses the range you pinned to lab-node-1 and lab-node-2 as ipam.subnet, ipam.gateway is the first usable address of that range, and ipam.routes must contain both 0.0.0.0/0 and the cluster Pod range 10.244.0.0/16. ipam.type is host-local.
  7. Judge four range layouts — p1 is Pod 10.244.0.0/16, Service 10.96.0.0/16, node 192.168.10.0/24; p2 is Pod 10.96.0.0/12, Service 10.96.0.0/16, node 192.168.10.0/24; p3 is Pod 10.244.0.0/16, Service 10.245.0.0/16, node 10.244.3.0/24; p4 is Pod 172.16.0.0/16, Service 10.96.0.0/16, node 192.168.10.0/24. In /root/k8nd-cni/07-overlap.txt, write five lines in all: four lines from p1= to p4= (the value is ok or overlap) and one line rule=pod,service,node.
  8. In /root/k8nd-cni/08-report.md, write five lines — cluster_pod_cidr=, per_node_prefix=, nodes_addressable= (the number of nodes that prefix can hold), notready_node= (the node name you created in step 5), chained_plugins= (the plugins chained after bridge in the step 3 conflist, separated by commas) — and below them write at least four lines that start with - .

Notes

Carve out a Pod range for each node

Split the cluster Pod range 10.244.0.0/16 into /24 blocks and pin them to lab-node-0, lab-node-1 and lab-node-2 in order from the front (both spec.podCIDR and spec.podCIDRs). Then in /root/k8nd-cni/01-podcidr.txt, write three lines of 노드이름 대역 (the placeholders are the node name and its range), in node-name order.

The API server does not hand out Pod addresses. Each node receives its own share of the range, and inside it the network plugin's IPAM hands out addresses to Pods one at a time. So even as nodes are added, address management stays within each node. Pin it with kubectl patch node <이름> --type=merge -p '{"spec":{"podCIDR":"...","podCIDRs":["..."]}}' and read it back with kubectl get node <이름> -o jsonpath='{.spec.podCIDR}' (in both commands the placeholder is the node name). Once a value has been set, it cannot be changed to another value.

How many Pods can that range hold

In /root/k8nd-cni/02-capacity.txt, write five lines — subnets= (the number of pieces when you cut 10.244.0.0/16 into /24), addresses_per_subnet= (the total number of addresses in one piece), usable_per_subnet= (the count after removing the network address and the gateway), node_pod_limit= (the status.allocatable.pods value the node object reports), binding_limit= (whichever of the two is hit first, as subnet or node).

Counting addresses and counting Pods are different things. One /24 block holds 256 addresses, but the network address and the gateway each take one away. However, the number of Pods a node accepts is set separately by a limit on the kubelet side, and usually that limit is hit first, not the addresses. Read the node limit with kubectl get node lab-node-0 -o jsonpath='{.status.allocatable.pods}'.

Chain the plugin configurations together

Create /root/k8nd-cni/10-k8nd.conflist — cniVersion is 1.0.0, name is k8nd-pod-network, and plugins has exactly three entries, in order bridge, portmap and bandwidth. The ipam of bridge is host-local and its subnet is the Pod range of lab-node-0, portmap must have capabilities.portMappings true, and bandwidth must have capabilities.bandwidth true.

conflist is the format that lists several plugins in one configuration file. They are called in order from the top, and each later plugin takes the result the earlier plugin produced and adjusts it. That is why attaching an address, opening hostPort and tightening bandwidth are split across different plugins. The default configuration directory is /etc/cni/net.d and the binary directory is /opt/cni/bin. When you finish writing, check the syntax with jq . <파일> (the placeholder is the file).

Who does hostPort and bandwidth limits

Write a Pod manifest to /root/k8nd-cni/shaped.yaml and apply it — namespace k8nd-cni, name shaped, annotations kubernetes.io/ingress-bandwidth: 1M and kubernetes.io/egress-bandwidth: 1M, and a container web (nginx:1.27-alpine) with containerPort 8080 and hostPort 8080. Then in /root/k8nd-cni/04-chain.txt, write four lines hostport_plugin=, bandwidth_plugin=, ingress_annotation= and egress_annotation= (for the plugins, the type of the plugin that handles that job in the step 3 conflist; for the annotations, the values you actually set).

The hostPort in a Pod spec and the bandwidth annotations are not something the API server handles. Both are read and executed by a plugin chained into the chain. So if that plugin is not installed on the node, the manifest passes and nothing happens — it is a failure that is silently ignored. Create the Pod with kubectl apply -f, and read the annotations back with kubectl -n <ns> get pod shaped -o jsonpath='{.metadata.annotations}'.

Without a plugin the whole node is blocked

Write the node k8nd-cni-edge to /root/k8nd-cni/edge-node.yaml and apply it — set spec.unschedulable to true, and add one taint of your own to spec.taints with the key k8nd.example.com/cni-missing (effect NoSchedule). After you apply it, in /root/k8nd-cni/05-notready.txt write four lines: node=, custom_taint= (the key you added), controller_taint= (the built-in taint key the controller added on top), and unschedulable=.

A node with no network plugin configured accepts no Pods at all. What actually creates that isolation is a taint, and the official documentation's list of built-in taints has node.kubernetes.io/network-unavailable (network is unusable) and node.kubernetes.io/unschedulable side by side. There is one important difference — built-in taints derived from conditions are managed by the controller. Even if you add one by hand, it is removed right away when the condition is not in that state, and conversely, if you set spec.unschedulable to true, the controller adds a taint by itself. So the only taints a person can add are ones that use their own key. Do not assume that the first item in the list is the one you wrote; check all of them with kubectl get node <이름> -o json | jq -r '.spec.taints[].key' (the placeholder is the node name).

Create a different IPAM configuration for each node

Create /root/k8nd-cni/06-node1.conflist and /root/k8nd-cni/06-node2.conflist. Each uses the range you pinned to lab-node-1 and lab-node-2 as ipam.subnet, ipam.gateway is the first usable address of that range, and ipam.routes must contain both 0.0.0.0/0 and the cluster Pod range 10.244.0.0/16. ipam.type is host-local.

host-local manages addresses only within that node, as its name says. It does not ask what the other nodes handed out — so if you give overlapping ranges, two nodes calmly hand out the same address. Carving ranges so they do not overlap is the job of a person or the control plane. The reason you write a separate route to the cluster range is to handle, inside the node, Pod-to-Pod traffic that would otherwise leave the node if it were sent to the default route.

When ranges overlap, where traffic goes is undefined

Judge four range layouts — p1 is Pod 10.244.0.0/16, Service 10.96.0.0/16, node 192.168.10.0/24; p2 is Pod 10.96.0.0/12, Service 10.96.0.0/16, node 192.168.10.0/24; p3 is Pod 10.244.0.0/16, Service 10.245.0.0/16, node 10.244.3.0/24; p4 is Pod 172.16.0.0/16, Service 10.96.0.0/16, node 192.168.10.0/24. In /root/k8nd-cni/07-overlap.txt, write five lines in all: four lines from p1= to p4= (the value is ok or overlap) and one line rule=pod,service,node.

The official Kubernetes documentation states firmly that the three kinds of addresses, Pod, Service and node, must not overlap one another. If they overlap, which side's rule is hit first depends on the implementation, and the difference shows up as "it does not work only on some nodes". Do not judge by eye; calculate — one line is enough: python3 -c "import ipaddress; print(ipaddress.ip_network('10.96.0.0/12').overlaps(ipaddress.ip_network('10.96.0.0/16')))". It is easy to miss that a shorter prefix means a wider range.

Leave a Pod network design memo

In /root/k8nd-cni/08-report.md, write five lines — cluster_pod_cidr=, per_node_prefix=, nodes_addressable= (the number of nodes that prefix can hold), notready_node= (the node name you created in step 5), chained_plugins= (the plugins chained after bridge in the step 3 conflist, separated by commas) — and below them write at least four lines that start with - .

Take the values from the files you created in the earlier steps, not from memory. In the explanation lines, write sentences you will really pull out the next time you design a cluster — things like "who is the party that hands out the address" and "what is silently ignored".