TT Lab
Get started
Learn Learning paths Courses

CCA — Cilium Certified Associate

Why Policy Does Not Wobble When the IP Changes

Continue in TT Lab

In one line

Cilium attaches a single number (an identity) to a Pod's set of labels and makes policy decisions with that number. Even if a Pod is rescheduled and its IP changes, the identity stays the same as long as the labels are the same, so you only need to update the ipcache in one place, without touching the policy maps.

Why this was needed

The fundamental problem with IP-based policy is that in Kubernetes an IP is a temporary value. Pods die and come back up all the time, and each time they receive a new address. If you express policy in terms of IPs, every time a single Pod restarts you have to recompute the rules on every node, and in the short window in between, traffic is wrongly allowed or wrongly blocked.

Cilium solves this problem one level up. Policy is expressed as a meaning, namely "app=frontend can reach port 8080 of app=backend," and that meaning is preserved as a number all the way down to the kernel. In the iptables era, this meaning was lost as it was translated into IP rules.

How it works

When a Pod comes up, the agent computes an identity from only the security-relevant labels. Which labels go in and which are left out is the important part.

포함되는 라벨 (예)          제외되는 라벨 (예)
k8s:app=frontend            k8s:pod-template-hash=7d4f9c
k8s:team=payments           k8s:controller-revision-hash=...
k8s:io.kubernetes.pod.       k8s:pod-template-generation=...
     namespace=shop

The reason labels such as pod-template-hash are excluded is clear. This value changes every time you roll out a new Deployment. If it were included in the identity, a single rollout would give every Pod a new identity, and all the policy maps would have to be recomputed. The advantage of the label-based model would vanish entirely.

Once an identity is created, an "IP → identity" mapping is propagated to the cilium_ipcache map on every node. The policy map of the receiving endpoint decides in O(1) using (소스 아이덴티티, 포트, 프로토콜, 방향) as the key (the tuple is source identity, port, protocol, direction). The key benefit of this design is that when Pods move around, only the ipcache changes and the policy map stays as it is.

kube-proxy replacement implements the Service abstraction with two layers of maps. It finds the frontend (IP, port) in cilium_lb4_services_v2, gets the actual Pod address from cilium_lb4_backends, and performs DNAT. On top of that, there is an optimization called socket load balancing. At the moment an application inside a Pod calls connect(), the socket hook in the kernel swaps the Service IP for a backend IP. Since the packet never leaves for the Service IP in the first place, neither per-packet NAT nor a conntrack entry is needed.

This is also why you must provide k8sServiceHost and k8sServicePort together when you turn on kubeProxyReplacement. Without kube-proxy, the one that has to translate the ClusterIP of the kubernetes Service (usually 10.96.0.1) into the actual API server address is Cilium itself, but at bootstrap time that Cilium is not up yet. To avoid this chicken-and-egg problem, you tell it the real address directly.

The criteria for choosing a routing mode are simple.

Criterion Tunnel (VXLAN/Geneve) Native routing
Network requirements Only a UDP port between nodes needs to be open The underlay must know the PodCIDR routes
Overhead About 50 bytes, reduced MTU None
Suitable environment Environments where you cannot control the underlay On-premises with BGP peering, cloud ENI

For load balancing algorithms, you should know Maglev consistent hashing. When backends are added or removed, it minimizes the rearrangement of the hash table so that most existing flows keep mapping to the same backend. It also has a large effect in setups where multiple nodes receive the same VIP via ECMP, making every node make the same choice.

What it looks like in the field

The author's homelab runs Cilium 1.20.1 on PodCIDR 10.244.0.0/16, and the datapath state is reported like this.

KubeProxyReplacement:   True      [enp2s0  10.0.0.117 (Direct Routing)]
Routing:                Network: Tunnel [vxlan]   Host: BPF
Masquerading:           BPF   [enp2s0]   10.244.2.0/24
Encryption:             Wireguard [cilium_wg0 (Port: 51871, Peers: 2)]
Modules Health:         Stopped(0) Degraded(0) OK(92)

There is a way to read it. Routing: Network Tunnel[vxlan] means encapsulation is used between nodes, and Host: BPF means eBPF also handles host routing. Masquerading: BPF means SNAT is also handled by eBPF rather than iptables. In other words, this node needs neither kube-proxy nor iptables rules for masquerading. In fact, the number of KUBE- chains was 0.

The actual address of the control plane node, 10.0.0.120, went into k8sServiceHost. How this address was settled is also a lesson. This cluster originally used .111 through a DHCP lease, and one day it received .120 and died. Because .120 was not in the apiserver certificate SAN, recovery was hard apart from reverting the IP, and in the end the address was pinned as static and the cluster was rebuilt from scratch. It is worth remembering that the bootstrap address is one of the few settings that is extremely painful to change later.

What you will do in the next lab

You will write a Cilium Helm values file yourself to declare kube-proxy replacement and the datapath options, put labels on nodes to pin workload placement, actually apply a Service and a Deployment to confirm that backends are picked up, and then build a script that verifies that kube-proxy is really absent.