TT Lab
Get started
Learn Learning paths Courses

CCA — Cilium Certified Associate

We Removed kube-proxy and the Service Rules Vanished

Continue in TT Lab

Goal

On a real Cilium 1.20.1 running without kube-proxy, you read directly which BPF maps and which eBPF programs Service load balancing has moved to. Confirm that iptables is empty, then follow the Service map, the conntrack map, the program on the Pod device, NodePort, and even the response of a Service with no backends.

Why it matters

kube-proxy's iptables mode creates rules for every Service and endpoint, and the rules are evaluated one after another from the top. As Services grow, the rules grow, and every time something changes, the rules are rewritten. Cilium's kube-proxy replacement keeps each Service as a hash map entry, and the eBPF programs attached to Pod and node devices look up the map as packets pass and change the destination.

You need to know this difference to know where to look during an outage. If you dig through iptables-save for a report that "the Service isn't working," nothing shows up. Instead you have to look at the backend slots in the Service map, the entries in conntrack, and whether a program is attached to the Pod device.

Use only the personal k3s inside the VM (Cilium 1.20.1, kubeProxyReplacement=true). Environment preparation takes about 5 minutes. When the session ends, the files in /root/cca-datapath disappear.

Steps

  1. Run kubectl apply -f /opt/fixtures/cca-datapath/web.yaml to start web (2 replicas) and a ClusterIP Service web in the cca-dp namespace. Once the Pods are Ready, read iptables-save from the VM shell and record the following in /root/cca-datapath/iptables.json — cluster_ip (the ClusterIP of the Service web), kube_svc_lines (the number of lines containing KUBE-SVC), cluster_ip_lines (the number of lines containing that ClusterIP), cilium_lines (the number of lines containing CILIUM), and kube_proxy_replacement (the KubeProxyReplacement value from the agent's cilium-dbg status). Write the counts as numbers.
  2. In the agent's cilium-dbg service list -o json, find the ClusterIP entry of cca-dp/web, and in cilium-dbg bpf lb list, check the slots of the same frontend. In /root/cca-datapath/svc.json, record service_id (a number), frontend (ClusterIP:80), backends (a list of two backend ip:port values), and backend_pods (the names of the two web Pods that have those IPs).
  3. Scale web up to 4 replicas. Poll briefly until all four Pods are Ready and the web frontend in bpf lb list has 4 backend slots, then in /root/cca-datapath/scale.json record backends (the four ip:port values currently in the map) and new_backends (the two that were not in svc.json). They must also match the ready addresses of the EndpointSlice.
  4. Run kubectl apply -f /opt/fixtures/cca-datapath/holder.yaml to start the holder Pod. The holder opens one TCP connection to the web Service (ClusterIP:80) from source port 40404 and keeps holding it, and prints in its log the peer address the app saw through getpeername. Look at the same connection in three places — the holder log, netstat -tn inside the holder, and the :40404 entry of the agent's cilium-dbg bpf ct list global. In /root/cca-datapath/ct.json, record client (holderIP:40404), app_peer (the peer address ip:port in the log), socket_peer (the Foreign Address in the holder's netstat), backend (the destination ip:port of the conntrack OUT entry), backend_pod (the name of the web Pod with that IP), svc_entries (the number of TCP SVC entries that have :40404, as a number), and socket_lb_coverage (the Socket LB Coverage from cilium-dbg status --verbose). The holder's IP:40404 must also appear in netstat inside the backend Pod.
  5. Using the holder's endpoint number (the status.id from kubectl -n cca-dp get cep holder), read the agent's cilium-dbg endpoint get <번호> -o json (the placeholder is the endpoint number) to find the host-side device name and ifindex, and from the VM shell compare bpftool net show dev <장치> and tc filter show dev <장치> ingress (the placeholder is the device name). In /root/cca-datapath/prog.json, record endpoint_id, interface, ifindex, attach (the attachment point shown by bpftool), program (the program name), policy_map (the name of this endpoint's policy map found in cilium-dbg map list), and tc_filter_lines (the number of lines in the tc filter output, as a number).
  6. In cca-dp, create a NodePort Service web-np — selector app=web, port 80, targetPort 8080, nodePort 30780. From the VM shell, send a request to port 30780 of the node's InternalIP and confirm 200, and in /root/cca-datapath/nodeport.json record node_ip, node_port, http_code (a number), iptables_lines (the number of lines in iptables-save containing 30780), listen_sockets (the number of lines from ss -Htln 'sport = :30780'), and service_id (the entry number of the 0.0.0.0:30780 NodePort entry in the agent's list).
  7. In cca-dp, create a ClusterIP Service ghost with selector app=ghost and port 80 → targetPort 8080 (do not create any Pod with that label). Start a curl Pod probe (image curlimages/curl:8.10.1@sha256:d9b4541e214bcd85196d6e92e2753ac6d0ea699f0af5741f8c6cccbfcf00ef4b, command sleep 86400) and from inside it send a request with a 5-second limit to the ClusterIP of ghost. In /root/cca-datapath/ghost.json, record cluster_ip, backends (the number of backends in the agent's list, as a number), curl_exit (the curl exit code), seconds (curl's time_total, as a number), and no_backend_response (the ServiceNoBackendResponse value from cilium-dbg config -a).
  8. In /root/cca-datapath/report.txt, write seven lines in 키=값 form (key=value) — kube_proxy_replacement, socket_lb_coverage (the Socket LB Coverage from status --verbose), web_backends (the current number of web backend slots in bpf lb list), holder_backend (the ip:port the holder connection currently points to in conntrack), holder_program (the program in prog.json), nodeport_iptables_lines (the current number of iptables lines containing 30780), and ghost_no_backend_response. All values must match the current state and the earlier records.

Notes

With no kube-proxy, who runs the Services?

Run kubectl apply -f /opt/fixtures/cca-datapath/web.yaml to start web (2 replicas) and a ClusterIP Service web in the cca-dp namespace. Once the Pods are Ready, read iptables-save from the VM shell and record the following in /root/cca-datapath/iptables.json — cluster_ip (the ClusterIP of the Service web), kube_svc_lines (the number of lines containing KUBE-SVC), cluster_ip_lines (the number of lines containing that ClusterIP), cilium_lines (the number of lines containing CILIUM), and kube_proxy_replacement (the KubeProxyReplacement value from the agent's cilium-dbg status). Write the counts as numbers.

kube-proxy (iptables mode) creates KUBE-SVC and KUBE-SEP chains for every Service to change the destination. Count the lines to see whether any trace of them exists. grep -c returns exit code 1 when there are no matches, so be careful under set -e. Run agent commands as kubectl -n kube-system exec ds/cilium -c cilium-agent -- cilium-dbg ....

Find the Service's ticket number in the BPF map

In the agent's cilium-dbg service list -o json, find the ClusterIP entry of cca-dp/web, and in cilium-dbg bpf lb list, check the slots of the same frontend. In /root/cca-datapath/svc.json, record service_id (a number), frontend (ClusterIP:80), backends (a list of two backend ip:port values), and backend_pods (the names of the two web Pods that have those IPs).

In the Service list JSON, the name, namespace, and type are in spec.flags, and the addresses are in spec.frontend-address and spec.backend-addresses. A single frontend in bpf lb list has a line with the number 0 in parentheses (the Service itself) and backend lines that start from 1. Match the Pod IPs against kubectl get pod -o wide.

When you add replicas, the map knows first

Scale web up to 4 replicas. Poll briefly until all four Pods are Ready and the web frontend in bpf lb list has 4 backend slots, then in /root/cca-datapath/scale.json record backends (the four ip:port values currently in the map) and new_backends (the two that were not in svc.json). They must also match the ready addresses of the EndpointSlice.

Kubernetes updates the EndpointSlice, and the agent watches it and rewrites the backend slots of the Service map. Instead of rewriting thousands of iptables rules, only a few map entries change. Use a loop that counts slots rather than a fixed sleep.

The app believes it connected to the Service, but the socket is already attached to a Pod

Run kubectl apply -f /opt/fixtures/cca-datapath/holder.yaml to start the holder Pod. The holder opens one TCP connection to the web Service (ClusterIP:80) from source port 40404 and keeps holding it, and prints in its log the peer address the app saw through getpeername. Look at the same connection in three places — the holder log, netstat -tn inside the holder, and the :40404 entry of the agent's cilium-dbg bpf ct list global. In /root/cca-datapath/ct.json, record client (holderIP:40404), app_peer (the peer address ip:port in the log), socket_peer (the Foreign Address in the holder's netstat), backend (the destination ip:port of the conntrack OUT entry), backend_pod (the name of the web Pod with that IP), svc_entries (the number of TCP SVC entries that have :40404, as a number), and socket_lb_coverage (the Socket LB Coverage from cilium-dbg status --verbose). The holder's IP:40404 must also appear in netstat inside the backend Pod.

With the socket LB of kube-proxy replacement, at the moment a Pod calls connect(), the eBPF attached to the cgroup changes the destination to a backend. So the packet never carries the Service address in the first place, and the app is shown the Service address again through getpeername. Try to explain, in this order, why the three observations report different addresses. Also check whether the source the backend sees has changed to the node IP.

The program on the device next to the Pod, and the policy map only that Pod has

Using the holder's endpoint number (the status.id from kubectl -n cca-dp get cep holder), read the agent's cilium-dbg endpoint get <번호> -o json (the placeholder is the endpoint number) to find the host-side device name and ifindex, and from the VM shell compare bpftool net show dev <장치> and tc filter show dev <장치> ingress (the placeholder is the device name). In /root/cca-datapath/prog.json, record endpoint_id, interface, ifindex, attach (the attachment point shown by bpftool), program (the program name), policy_map (the name of this endpoint's policy map found in cilium-dbg map list), and tc_filter_lines (the number of lines in the tc filter output, as a number).

Cilium attaches a program to the host-side veth (lxc…) of every Pod and keeps a separate policy map for every Pod. Compare the trailing number in the map name with the endpoint number. On recent kernels the program is attached through tcx, and the old tc command does not show a tcx attachment.

No process is listening, yet the NodePort answers

In cca-dp, create a NodePort Service web-np — selector app=web, port 80, targetPort 8080, nodePort 30780. From the VM shell, send a request to port 30780 of the node's InternalIP and confirm 200, and in /root/cca-datapath/nodeport.json record node_ip, node_port, http_code (a number), iptables_lines (the number of lines in iptables-save containing 30780), listen_sockets (the number of lines from ss -Htln 'sport = :30780'), and service_id (the entry number of the 0.0.0.0:30780 NodePort entry in the agent's list).

A NodePort is not received by a process that opens the port; an eBPF program attached to the node device receives it and sends it to the Service map. That is why a port answers that shows up in neither ss nor iptables. In the Service list JSON, pick the entry whose frontend-address ip is 0.0.0.0 and whose flags.type is NodePort.

A Service with no Pods at all does not wait

In cca-dp, create a ClusterIP Service ghost with selector app=ghost and port 80 → targetPort 8080 (do not create any Pod with that label). Start a curl Pod probe (image curlimages/curl:8.10.1@sha256:d9b4541e214bcd85196d6e92e2753ac6d0ea699f0af5741f8c6cccbfcf00ef4b, command sleep 86400) and from inside it send a request with a 5-second limit to the ClusterIP of ghost. In /root/cca-datapath/ghost.json, record cluster_ip, backends (the number of backends in the agent's list, as a number), curl_exit (the curl exit code), seconds (curl's time_total, as a number), and no_backend_response (the ServiceNoBackendResponse value from cilium-dbg config -a).

If the datapath silently drops the packet when there is no backend, the client waits until the timeout, and if it returns a rejection response, the client fails immediately. Compare what curl exit codes 7 and 28 mean. The -w '%{time_total}' of curl prints even on failure.

The new address book for what kube-proxy used to do

In /root/cca-datapath/report.txt, write seven lines in 키=값 form (key=value) — kube_proxy_replacement, socket_lb_coverage (the Socket LB Coverage from status --verbose), web_backends (the current number of web backend slots in bpf lb list), holder_backend (the ip:port the holder connection currently points to in conntrack), holder_program (the program in prog.json), nodeport_iptables_lines (the current number of iptables lines containing 30780), and ghost_no_backend_response. All values must match the current state and the earlier records.

Instead of rewriting the files from the earlier steps, re-read the current state and write that. If the holder restarted and reconnected, the backend may have changed, so ct.json must be regenerated too.