CCA — Cilium Certified Associate
We Removed kube-proxy and the Service Rules Vanished
Goal
On a real Cilium 1.20.1 running without kube-proxy, you read directly which BPF maps and which eBPF programs Service load balancing has moved to. Confirm that iptables is empty, then follow the Service map, the conntrack map, the program on the Pod device, NodePort, and even the response of a Service with no backends.
Why it matters
kube-proxy's iptables mode creates rules for every Service and endpoint, and the rules are evaluated one after another from the top. As Services grow, the rules grow, and every time something changes, the rules are rewritten. Cilium's kube-proxy replacement keeps each Service as a hash map entry, and the eBPF programs attached to Pod and node devices look up the map as packets pass and change the destination.
You need to know this difference to know where to look during an outage. If you dig through iptables-save for a report that "the Service isn't working," nothing shows up. Instead you have to look at the backend slots in the Service map, the entries in conntrack, and whether a program is attached to the Pod device.
Use only the personal k3s inside the VM (Cilium 1.20.1, kubeProxyReplacement=true). Environment preparation takes about 5 minutes. When the session ends, the files in /root/cca-datapath disappear.
Steps
- Run
kubectl apply -f /opt/fixtures/cca-datapath/web.yamlto start web (2 replicas) and a ClusterIP Service web in the cca-dp namespace. Once the Pods are Ready, readiptables-savefrom the VM shell and record the following in/root/cca-datapath/iptables.json—cluster_ip(the ClusterIP of the Service web),kube_svc_lines(the number of lines containing KUBE-SVC),cluster_ip_lines(the number of lines containing that ClusterIP),cilium_lines(the number of lines containing CILIUM), andkube_proxy_replacement(the KubeProxyReplacement value from the agent'scilium-dbg status). Write the counts as numbers. - In the agent's
cilium-dbg service list -o json, find the ClusterIP entry of cca-dp/web, and incilium-dbg bpf lb list, check the slots of the same frontend. In/root/cca-datapath/svc.json, recordservice_id(a number),frontend(ClusterIP:80),backends(a list of two backend ip:port values), andbackend_pods(the names of the two web Pods that have those IPs). - Scale web up to 4 replicas. Poll briefly until all four Pods are Ready and the web frontend in
bpf lb listhas 4 backend slots, then in/root/cca-datapath/scale.jsonrecordbackends(the four ip:port values currently in the map) andnew_backends(the two that were not in svc.json). They must also match the ready addresses of the EndpointSlice. - Run
kubectl apply -f /opt/fixtures/cca-datapath/holder.yamlto start the holder Pod. The holder opens one TCP connection to the web Service (ClusterIP:80) from source port 40404 and keeps holding it, and prints in its log the peer address the app saw through getpeername. Look at the same connection in three places — the holder log,netstat -tninside the holder, and the:40404entry of the agent'scilium-dbg bpf ct list global. In/root/cca-datapath/ct.json, recordclient(holderIP:40404),app_peer(the peer address ip:port in the log),socket_peer(the Foreign Address in the holder's netstat),backend(the destination ip:port of the conntrack OUT entry),backend_pod(the name of the web Pod with that IP),svc_entries(the number of TCP SVC entries that have :40404, as a number), andsocket_lb_coverage(the Socket LB Coverage fromcilium-dbg status --verbose). The holder's IP:40404 must also appear in netstat inside the backend Pod. - Using the holder's endpoint number (the status.id from
kubectl -n cca-dp get cep holder), read the agent'scilium-dbg endpoint get <번호> -o json(the placeholder is the endpoint number) to find the host-side device name and ifindex, and from the VM shell comparebpftool net show dev <장치>andtc filter show dev <장치> ingress(the placeholder is the device name). In/root/cca-datapath/prog.json, recordendpoint_id,interface,ifindex,attach(the attachment point shown by bpftool),program(the program name),policy_map(the name of this endpoint's policy map found incilium-dbg map list), andtc_filter_lines(the number of lines in the tc filter output, as a number). - In cca-dp, create a NodePort Service
web-np— selectorapp=web, port 80, targetPort 8080, nodePort30780. From the VM shell, send a request to port 30780 of the node's InternalIP and confirm 200, and in/root/cca-datapath/nodeport.jsonrecordnode_ip,node_port,http_code(a number),iptables_lines(the number of lines in iptables-save containing 30780),listen_sockets(the number of lines fromss -Htln 'sport = :30780'), andservice_id(the entry number of the 0.0.0.0:30780 NodePort entry in the agent's list). - In cca-dp, create a ClusterIP Service
ghostwith selectorapp=ghostand port 80 → targetPort 8080 (do not create any Pod with that label). Start a curl Podprobe(imagecurlimages/curl:8.10.1@sha256:d9b4541e214bcd85196d6e92e2753ac6d0ea699f0af5741f8c6cccbfcf00ef4b, commandsleep 86400) and from inside it send a request with a 5-second limit to the ClusterIP of ghost. In/root/cca-datapath/ghost.json, recordcluster_ip,backends(the number of backends in the agent's list, as a number),curl_exit(the curl exit code),seconds(curl's time_total, as a number), andno_backend_response(the ServiceNoBackendResponse value fromcilium-dbg config -a). - In
/root/cca-datapath/report.txt, write seven lines in키=값form (key=value) —kube_proxy_replacement,socket_lb_coverage(the Socket LB Coverage from status --verbose),web_backends(the current number of web backend slots in bpf lb list),holder_backend(the ip:port the holder connection currently points to in conntrack),holder_program(the program in prog.json),nodeport_iptables_lines(the current number of iptables lines containing 30780), andghost_no_backend_response. All values must match the current state and the earlier records.
Notes
- Kubernetes Without kube-proxy (socket LB, hostNamespaceOnly, NodePort): https://docs.cilium.io/en/v1.20/network/kubernetes/kubeproxy-free/
- List and sizes of eBPF maps: https://docs.cilium.io/en/v1.20/network/ebpf/maps/
- BPF and XDP reference guide: https://docs.cilium.io/en/v1.20/reference-guides/bpf/index.html
- Agent commands:
kubectl -n kube-system exec ds/cilium -c cilium-agent -- cilium-dbg service list | bpf lb list | bpf ct list global | map list | endpoint get <번호>(the placeholder is the endpoint number). - The VM shell has bpftool, tc, ss, iptables-save, and jq. Inside a Pod with the python image there is busybox netstat.
With no kube-proxy, who runs the Services?
Run kubectl apply -f /opt/fixtures/cca-datapath/web.yaml to start web (2 replicas) and a ClusterIP Service web in the cca-dp namespace. Once the Pods are Ready, read iptables-save from the VM shell and record the following in /root/cca-datapath/iptables.json — cluster_ip (the ClusterIP of the Service web), kube_svc_lines (the number of lines containing KUBE-SVC), cluster_ip_lines (the number of lines containing that ClusterIP), cilium_lines (the number of lines containing CILIUM), and kube_proxy_replacement (the KubeProxyReplacement value from the agent's cilium-dbg status). Write the counts as numbers.
kube-proxy (iptables mode) creates KUBE-SVC and KUBE-SEP chains for every Service to change the destination. Count the lines to see whether any trace of them exists. grep -c returns exit code 1 when there are no matches, so be careful under set -e. Run agent commands as kubectl -n kube-system exec ds/cilium -c cilium-agent -- cilium-dbg ....
Find the Service's ticket number in the BPF map
In the agent's cilium-dbg service list -o json, find the ClusterIP entry of cca-dp/web, and in cilium-dbg bpf lb list, check the slots of the same frontend. In /root/cca-datapath/svc.json, record service_id (a number), frontend (ClusterIP:80), backends (a list of two backend ip:port values), and backend_pods (the names of the two web Pods that have those IPs).
In the Service list JSON, the name, namespace, and type are in spec.flags, and the addresses are in spec.frontend-address and spec.backend-addresses. A single frontend in bpf lb list has a line with the number 0 in parentheses (the Service itself) and backend lines that start from 1. Match the Pod IPs against kubectl get pod -o wide.
When you add replicas, the map knows first
Scale web up to 4 replicas. Poll briefly until all four Pods are Ready and the web frontend in bpf lb list has 4 backend slots, then in /root/cca-datapath/scale.json record backends (the four ip:port values currently in the map) and new_backends (the two that were not in svc.json). They must also match the ready addresses of the EndpointSlice.
Kubernetes updates the EndpointSlice, and the agent watches it and rewrites the backend slots of the Service map. Instead of rewriting thousands of iptables rules, only a few map entries change. Use a loop that counts slots rather than a fixed sleep.
The app believes it connected to the Service, but the socket is already attached to a Pod
Run kubectl apply -f /opt/fixtures/cca-datapath/holder.yaml to start the holder Pod. The holder opens one TCP connection to the web Service (ClusterIP:80) from source port 40404 and keeps holding it, and prints in its log the peer address the app saw through getpeername. Look at the same connection in three places — the holder log, netstat -tn inside the holder, and the :40404 entry of the agent's cilium-dbg bpf ct list global. In /root/cca-datapath/ct.json, record client (holderIP:40404), app_peer (the peer address ip:port in the log), socket_peer (the Foreign Address in the holder's netstat), backend (the destination ip:port of the conntrack OUT entry), backend_pod (the name of the web Pod with that IP), svc_entries (the number of TCP SVC entries that have :40404, as a number), and socket_lb_coverage (the Socket LB Coverage from cilium-dbg status --verbose). The holder's IP:40404 must also appear in netstat inside the backend Pod.
With the socket LB of kube-proxy replacement, at the moment a Pod calls connect(), the eBPF attached to the cgroup changes the destination to a backend. So the packet never carries the Service address in the first place, and the app is shown the Service address again through getpeername. Try to explain, in this order, why the three observations report different addresses. Also check whether the source the backend sees has changed to the node IP.
The program on the device next to the Pod, and the policy map only that Pod has
Using the holder's endpoint number (the status.id from kubectl -n cca-dp get cep holder), read the agent's cilium-dbg endpoint get <번호> -o json (the placeholder is the endpoint number) to find the host-side device name and ifindex, and from the VM shell compare bpftool net show dev <장치> and tc filter show dev <장치> ingress (the placeholder is the device name). In /root/cca-datapath/prog.json, record endpoint_id, interface, ifindex, attach (the attachment point shown by bpftool), program (the program name), policy_map (the name of this endpoint's policy map found in cilium-dbg map list), and tc_filter_lines (the number of lines in the tc filter output, as a number).
Cilium attaches a program to the host-side veth (lxc…) of every Pod and keeps a separate policy map for every Pod. Compare the trailing number in the map name with the endpoint number. On recent kernels the program is attached through tcx, and the old tc command does not show a tcx attachment.
No process is listening, yet the NodePort answers
In cca-dp, create a NodePort Service web-np — selector app=web, port 80, targetPort 8080, nodePort 30780. From the VM shell, send a request to port 30780 of the node's InternalIP and confirm 200, and in /root/cca-datapath/nodeport.json record node_ip, node_port, http_code (a number), iptables_lines (the number of lines in iptables-save containing 30780), listen_sockets (the number of lines from ss -Htln 'sport = :30780'), and service_id (the entry number of the 0.0.0.0:30780 NodePort entry in the agent's list).
A NodePort is not received by a process that opens the port; an eBPF program attached to the node device receives it and sends it to the Service map. That is why a port answers that shows up in neither ss nor iptables. In the Service list JSON, pick the entry whose frontend-address ip is 0.0.0.0 and whose flags.type is NodePort.
A Service with no Pods at all does not wait
In cca-dp, create a ClusterIP Service ghost with selector app=ghost and port 80 → targetPort 8080 (do not create any Pod with that label). Start a curl Pod probe (image curlimages/curl:8.10.1@sha256:d9b4541e214bcd85196d6e92e2753ac6d0ea699f0af5741f8c6cccbfcf00ef4b, command sleep 86400) and from inside it send a request with a 5-second limit to the ClusterIP of ghost. In /root/cca-datapath/ghost.json, record cluster_ip, backends (the number of backends in the agent's list, as a number), curl_exit (the curl exit code), seconds (curl's time_total, as a number), and no_backend_response (the ServiceNoBackendResponse value from cilium-dbg config -a).
If the datapath silently drops the packet when there is no backend, the client waits until the timeout, and if it returns a rejection response, the client fails immediately. Compare what curl exit codes 7 and 28 mean. The -w '%{time_total}' of curl prints even on failure.
The new address book for what kube-proxy used to do
In /root/cca-datapath/report.txt, write seven lines in 키=값 form (key=value) — kube_proxy_replacement, socket_lb_coverage (the Socket LB Coverage from status --verbose), web_backends (the current number of web backend slots in bpf lb list), holder_backend (the ip:port the holder connection currently points to in conntrack), holder_program (the program in prog.json), nodeport_iptables_lines (the current number of iptables lines containing 30780), and ghost_no_backend_response. All values must match the current state and the earlier records.
Instead of rewriting the files from the earlier steps, re-read the current state and write that. If the holder restarted and reconnected, the backend may have changed, so ct.json must be regenerated too.