TT Lab
Get started
Learn Learning paths Courses

CCA — Cilium Certified Associate

Where a Packet Is Caught and Where It Bounces Off

Continue in TT Lab

In one line

Cilium pushes Kubernetes semantics (Services, labels, policies) directly into in-kernel data structures. That is why lookup cost does not grow as Services multiply, but not everything gets faster.

Why this was needed

In kube-proxy's iptables mode, a single Service request walks the KUBE-SERVICES chain, goes through a probabilistic branch in KUBE-SVC-, and is DNATed in KUBE-SEP-. The number of rules grows in proportion to the number of Services and backends, and for every packet the kernel effectively traverses that list.

The more painful part is the update cost. To change even one rule, iptables effectively rewrites the entire table. In a cluster with frequent deployments, this shows up as CPU spikes and delays before rules take effect. In an environment with thousands of Services, the gap of a few seconds where "the Pod is up but traffic has not arrived yet" comes from here.

eBPF turns both axes into constant time. Lookup is a hash map query, so it is independent of the number of Services, and updates are incremental per map entry, so other entries are left untouched.

How it works

This is the journey of a packet that arrives from outside through a NodePort.

NIC 수신
  -> [XDP 훅]            드라이버 수준. sk_buff 할당 전이라 가장 빠름
  -> [tc ingress 훅]     서비스 변환(lb4_services -> lb4_backends), conntrack,
                         ipcache 룩업으로 목적지 아이덴티티 확인
  -> bpf_redirect_peer   호스트 veth 에서 파드 네임스페이스 안 피어로 직행
  -> [파드 lxc 인터페이스] 정책 평가: (소스 아이덴티티, 포트, 프로토콜, 방향)
  -> 애플리케이션 소켓

The eBPF hooks a packet arriving through a NodePort passes through: NIC receive, XDP hook, tc ingress hook, bpf_redirect_peer, the Pod's lxc interface, and the application socket, in that order. Service translation, conntrack, and policy evaluation happen at the tc hook, and only traffic with L7 policies leaves the kernel and goes through the userspace Envoy, which adds latency

Three points need emphasis. First, XDP runs before memory allocation, so it is used for DDoS filtering and NodePort acceleration, but it requires a supported NIC driver. Second, the core of the datapath is the tc hook, where Service translation, conntrack, and policy evaluation all happen. Third, bpf_redirect_peer skips one softIRQ rescheduling cycle that occurs when crossing a veth pair, and crosses the namespace boundary in a single step.

Policies are evaluated by identity, not by IP. A set of Pod labels maps to a single number, and that number becomes the key in the kernel map. The reserved identities, which appear on the exam as they are, are worth memorizing.

Number Name Meaning
0 unknown Unknown source
1 host The local node itself
2 world Everything outside the cluster
3 unmanaged Endpoints not managed by Cilium
4 health Health check endpoint
5 init Endpoints being initialized
6 remote-node The other nodes
7 kube-apiserver The API server
8 ingress Ingress

Here is the division of roles as well. The Cilium Agent runs on every node as a DaemonSet and handles loading eBPF programs, managing endpoints, enforcing policies, and conntrack. The Cilium Operator runs once per cluster and handles cluster-scoped IPAM, CRD garbage collection, and node discovery. The Agent directly touches each node's datapath, which is why it has to be a DaemonSet, while the Operator deals with cluster-wide state, so one instance is enough.

Finally, the honest trade-offs. The claim that switching to eBPF makes everything faster is not true.

Knowing "what cost each feature adds when you turn it on" is the heart of operations.

What it looks like in the field

The author runs a seven-node homelab (cp-1/2/3 + gpu-a/b/c/d) with Kubernetes v1.34.10, containerd 1.7.27, kernel 6.14, and Cilium 1.20.1. This cluster was built with kubeadm init --skip-phases=addon/kube-proxy, without ever installing kube-proxy. That is different from deleting it after installation. A kube-proxy that has run even once leaves KUBE-SERVICES / KUBE-SVC-* / KUBE-SEP-* chains on the node, and even if you delete the DaemonSet, those rules remain and conflict with the eBPF datapath.

These are results confirmed by measurement, not by assertion.

kube-proxy 파드 수: 0
노드의 iptables KUBE- 체인 개수: 0

And this is how the Services sat inside the kernel map.

SERVICE ADDRESS             BACKEND ADDRESS (REVNAT_ID) (SLOT)
10.96.0.10:53/TCP (1)       10.244.0.150:53/TCP (4) (1)
10.96.0.1:443/TCP (1)       10.0.0.120:6443/TCP (1) (1)

cilium status reported KubeProxyReplacement: True and Modules Health: OK 92 / Degraded 0, and the 14 basic verification items, covering node status, the control plane, CNI, CoreDNS, scheduling, and Service DNS, passed 14 out of 14. The reason these numbers matter is simple. The evidence that the eBPF datapath works is not the documentation but the chain count and the map dump.

What to check in the next quiz

This module covers concepts only. In the next module, which covers the identity model and kube-proxy replacement, you will write the values seen above into a Helm values file yourself and even build a script that verifies that state.