TT Lab
开始
学习 学习路径 课程

CCA — Cilium 认证助理

删掉 kube-proxy 后,Service 规则不见了

在 TT Lab 中继续学习

目标

在不使用 kube-proxy 运行的真正 Cilium 1.20.1 中,亲自读出服务负载均衡转移到了哪个 BPF map 和哪个 eBPF 程序上。 确认 iptables 是空的,然后一路追踪服务 map、conntrack map、Pod 设备上的程序、NodePort,直到没有后端的服务的响应。

为什么重要

kube-proxy 的 iptables 模式会为每个服务和 endpoint 创建规则,规则从上到下依次评估。服务增多,规则就增多,每次变化都要重写规则。Cilium 的 kube-proxy 替代把服务放成哈希 map 条目,由挂在 Pod 和节点设备上的 eBPF 程序在数据包经过时查 map 并改写目的地。

了解这种区别,才知道出故障时该看哪里。收到“服务不通”的报告时去翻 iptables-save,什么也翻不出来。应该转而查看服务 map 的后端槽位、conntrack 的条目,以及 Pod 设备上是否挂了程序。

请只使用 VM 内的个人 k3s(Cilium 1.20.1,kubeProxyReplacement=true)。环境准备约需 5 分钟。会话结束后,/root/cca-datapath 中的文件会消失。

步骤

  1. 用 kubectl apply -f /opt/fixtures/cca-datapath/web.yaml 在 cca-dp 命名空间中启动 web(副本 2 个)和 ClusterIP 服务 web。Pod 变为 Ready 后,在 VM shell 中读取 iptables-save,记录到 /root/cca-datapath/iptables.json——cluster_ip(服务 web 的 ClusterIP)、kube_svc_lines(包含 KUBE-SVC 的行数)、cluster_ip_lines(包含该 ClusterIP 的行数)、cilium_lines(包含 CILIUM 的行数)、kube_proxy_replacement(agent 中 cilium-dbg status 的 KubeProxyReplacement 值)。数量写成数字。
  2. 在 agent 的 cilium-dbg service list -o json 中找到 cca-dp/web 的 ClusterIP 条目,并在 cilium-dbg bpf lb list 中确认同一 frontend 的槽位。在 /root/cca-datapath/svc.json 中记录 service_id(数字)、frontend(ClusterIP:80)、backends(后端 ip:port 共 2 个的列表)、backend_pods(拥有这些 IP 的 web Pod 的 2 个名称)。
  3. 把 web 扩到 4 个副本。短时间轮询,直到四个 Pod 都 Ready 且 bpf lb list 中 web frontend 的后端槽位变为 4 个,然后在 /root/cca-datapath/scale.json 中记录 backends(当前 map 中的 4 个 ip:port)、new_backends(svc.json 中没有的 2 个)。它还必须与 EndpointSlice 的 ready 地址一致。
  4. 用 kubectl apply -f /opt/fixtures/cca-datapath/holder.yaml 启动 holder Pod。holder 以源端口 40404 向 web 服务(ClusterIP:80)打开并持有一个 TCP 连接,并在日志中输出应用通过 getpeername 看到的对端地址。请在三个位置查看同一个连接——holder 的日志、holder 内部的 netstat -tn、agent 的 cilium-dbg bpf ct list global 中的 :40404 条目。在 /root/cca-datapath/ct.json 中记录 client(holderIP:40404)、app_peer(日志中的对端地址 ip:port)、socket_peer(holder netstat 的 Foreign Address)、backend(conntrack OUT 条目的目的 ip:port)、backend_pod(该 IP 对应的 web Pod 名称)、svc_entries(带有 :40404 的 TCP SVC 条目数,数字)、socket_lb_coverage(cilium-dbg status --verbose 的 Socket LB Coverage)。后端 Pod 内的 netstat 中也必须能看到 holder 的 IP:40404。
  5. 用 holder 的 endpoint 编号(kubectl -n cca-dp get cep holder 的 status.id)执行 agent 的 cilium-dbg endpoint get <번호> -o json(占位符为编号),找到宿主侧的设备名称和 ifindex,并在 VM shell 中比较 bpftool net show dev <장치> 与 tc filter show dev <장치> ingress(占位符为设备名)。在 /root/cca-datapath/prog.json 中记录 endpoint_id、interface、ifindex、attach(bpftool 显示的挂载位置)、program(程序名称)、policy_map(在 cilium-dbg map list 中找到的该 endpoint 的策略 map 名称)、tc_filter_lines(tc filter 输出的行数,数字)。
  6. 在 cca-dp 中创建 NodePort 服务 web-np——selector 为 app=web,port 80,targetPort 8080,nodePort 为 30780。在 VM shell 中向节点 InternalIP 的 30780 发请求并确认 200,然后在 /root/cca-datapath/nodeport.json 中记录 node_ip、node_port、http_code(数字)、iptables_lines(iptables-save 中包含 30780 的行数)、listen_sockets(ss -Htln 'sport = :30780' 的行数)、service_id(agent 列表中 0.0.0.0:30780 NodePort 条目的编号)。
  7. 在 cca-dp 中创建 selector 为 app=ghost、port 80 → targetPort 8080 的 ClusterIP 服务 ghost(不创建带该标签的 Pod)。启动 curl Pod probe(镜像 curlimages/curl:8.10.1@sha256:d9b4541e214bcd85196d6e92e2753ac6d0ea699f0af5741f8c6cccbfcf00ef4b,命令 sleep 86400),在其中向 ghost 的 ClusterIP 发送限时 5 秒的请求。在 /root/cca-datapath/ghost.json 中记录 cluster_ip、backends(agent 列表中的后端数量,数字)、curl_exit(curl 退出码)、seconds(curl 的 time_total,数字)、no_backend_response(cilium-dbg config -a 中 ServiceNoBackendResponse 的值)。
  8. 在 /root/cca-datapath/report.txt 中写入七行 키=값(占位符依次为键和值)形式的内容——kube_proxy_replacement、socket_lb_coverage(status --verbose 中的 Socket LB Coverage)、web_backends(当前 bpf lb list 中 web 的后端槽位数)、holder_backend(当前 conntrack 中 holder 连接所指向的 ip:port)、holder_program(prog.json 中的 program)、nodeport_iptables_lines(当前包含 30780 的 iptables 行数)、ghost_no_backend_response。所有值必须与当前状态和前面的记录一致。

参考

没有 kube-proxy,服务由谁来运转

用 kubectl apply -f /opt/fixtures/cca-datapath/web.yaml 在 cca-dp 命名空间中启动 web(副本 2 个)和 ClusterIP 服务 web。Pod 变为 Ready 后,在 VM shell 中读取 iptables-save,记录到 /root/cca-datapath/iptables.json——cluster_ip(服务 web 的 ClusterIP)、kube_svc_lines(包含 KUBE-SVC 的行数)、cluster_ip_lines(包含该 ClusterIP 的行数)、cilium_lines(包含 CILIUM 的行数)、kube_proxy_replacement(agent 中 cilium-dbg status 的 KubeProxyReplacement 值)。数量写成数字。

kube-proxy(iptables 模式)会为每个服务创建 KUBE-SVC、KUBE-SEP 链来改写目的地。请按行数统计是否留有这些痕迹。grep -c 在一个都没有时会返回退出码 1,因此在 set -e 之下要小心。agent 命令用 kubectl -n kube-system exec ds/cilium -c cilium-agent -- cilium-dbg ... 执行。

在 BPF map 中找到服务编号

在 agent 的 cilium-dbg service list -o json 中找到 cca-dp/web 的 ClusterIP 条目,并在 cilium-dbg bpf lb list 中确认同一 frontend 的槽位。在 /root/cca-datapath/svc.json 中记录 service_id(数字)、frontend(ClusterIP:80)、backends(后端 ip:port 共 2 个的列表)、backend_pods(拥有这些 IP 的 web Pod 的 2 个名称)。

服务列表 JSON 的 spec.flags 中有名称、命名空间和类型,spec.frontend-address 与 spec.backend-addresses 中有地址。bpf lb list 的一个 frontend 下,有括号中编号为 0 的行(服务本身)和从 1 开始的后端行。请把 Pod IP 与 kubectl get pod -o wide 对照。

扩容副本后,map 最先知道

把 web 扩到 4 个副本。短时间轮询,直到四个 Pod 都 Ready 且 bpf lb list 中 web frontend 的后端槽位变为 4 个,然后在 /root/cca-datapath/scale.json 中记录 backends(当前 map 中的 4 个 ip:port)、new_backends(svc.json 中没有的 2 个)。它还必须与 EndpointSlice 的 ready 地址一致。

Kubernetes 会更新 EndpointSlice,agent 看到后重写服务 map 的后端槽位。不必重写几千行 iptables 规则,只需改变几个 map 条目。请用统计槽位数量的循环,而不是固定的 sleep。

应用以为连上的是服务,但套接字其实已经连到 Pod 上了

用 kubectl apply -f /opt/fixtures/cca-datapath/holder.yaml 启动 holder Pod。holder 以源端口 40404 向 web 服务(ClusterIP:80)打开并持有一个 TCP 连接,并在日志中输出应用通过 getpeername 看到的对端地址。请在三个位置查看同一个连接——holder 的日志、holder 内部的 netstat -tn、agent 的 cilium-dbg bpf ct list global 中的 :40404 条目。在 /root/cca-datapath/ct.json 中记录 client(holderIP:40404)、app_peer(日志中的对端地址 ip:port)、socket_peer(holder netstat 的 Foreign Address)、backend(conntrack OUT 条目的目的 ip:port)、backend_pod(该 IP 对应的 web Pod 名称)、svc_entries(带有 :40404 的 TCP SVC 条目数,数字)、socket_lb_coverage(cilium-dbg status --verbose 的 Socket LB Coverage)。后端 Pod 内的 netstat 中也必须能看到 holder 的 IP:40404。

kube-proxy 替代的 socket LB 是在 Pod 调用 connect() 的瞬间,由挂在 cgroup 上的 eBPF 把目的地改成后端。所以数据包一开始就没有服务地址,而对应用,则会通过 getpeername 还原并显示服务地址。请按这个顺序解释三种观测为什么给出不同的地址。还要确认后端看到的源地址是否变成了节点 IP。

挂在 Pod 旁设备上的程序,以及只属于该 Pod 的策略 map

用 holder 的 endpoint 编号(kubectl -n cca-dp get cep holder 的 status.id)执行 agent 的 cilium-dbg endpoint get <번호> -o json(占位符为编号),找到宿主侧的设备名称和 ifindex,并在 VM shell 中比较 bpftool net show dev <장치> 与 tc filter show dev <장치> ingress(占位符为设备名)。在 /root/cca-datapath/prog.json 中记录 endpoint_id、interface、ifindex、attach(bpftool 显示的挂载位置)、program(程序名称)、policy_map(在 cilium-dbg map list 中找到的该 endpoint 的策略 map 名称)、tc_filter_lines(tc filter 输出的行数,数字)。

Cilium 会给每个 Pod 的宿主侧 veth(lxc…)挂载程序,并为每个 Pod 单独设置策略 map。请比较 map 名称末尾的数字和 endpoint 编号。在较新的内核上,程序以 tcx 方式挂载,而旧的 tc 命令不会显示 tcx 挂载。

没有监听的进程,NodePort 却有响应

在 cca-dp 中创建 NodePort 服务 web-np——selector 为 app=web,port 80,targetPort 8080,nodePort 为 30780。在 VM shell 中向节点 InternalIP 的 30780 发请求并确认 200,然后在 /root/cca-datapath/nodeport.json 中记录 node_ip、node_port、http_code(数字)、iptables_lines(iptables-save 中包含 30780 的行数)、listen_sockets(ss -Htln 'sport = :30780' 的行数)、service_id(agent 列表中 0.0.0.0:30780 NodePort 条目的编号)。

NodePort 不是由打开端口的进程接收的,而是由挂在节点设备上的 eBPF 程序接收后送入服务 map。所以会有 ss 和 iptables 都看不到的端口在响应。请在服务列表 JSON 中选出 frontend-address 的 ip 为 0.0.0.0 且 flags.type 为 NodePort 的条目。

完全没有 Pod 的服务不会让人等待

在 cca-dp 中创建 selector 为 app=ghost、port 80 → targetPort 8080 的 ClusterIP 服务 ghost(不创建带该标签的 Pod)。启动 curl Pod probe(镜像 curlimages/curl:8.10.1@sha256:d9b4541e214bcd85196d6e92e2753ac6d0ea699f0af5741f8c6cccbfcf00ef4b,命令 sleep 86400),在其中向 ghost 的 ClusterIP 发送限时 5 秒的请求。在 /root/cca-datapath/ghost.json 中记录 cluster_ip、backends(agent 列表中的后端数量,数字)、curl_exit(curl 退出码)、seconds(curl 的 time_total,数字)、no_backend_response(cilium-dbg config -a 中 ServiceNoBackendResponse 的值)。

没有后端时,如果数据路径悄悄丢弃数据包,客户端会一直等到超时;如果返回拒绝响应,则会立刻失败。请比较 curl 退出码 7 和 28 分别意味着什么。curl 的 -w '%{time_total}' 即使失败也会输出。

kube-proxy 过去所做工作的新地址簿

在 /root/cca-datapath/report.txt 中写入七行 키=값(占位符依次为键和值)形式的内容——kube_proxy_replacement、socket_lb_coverage(status --verbose 中的 Socket LB Coverage)、web_backends(当前 bpf lb list 中 web 的后端槽位数)、holder_backend(当前 conntrack 中 holder 连接所指向的 ip:port)、holder_program(prog.json 中的 program)、nodeport_iptables_lines(当前包含 30780 的 iptables 行数)、ghost_no_backend_response。所有值必须与当前状态和前面的记录一致。

不要改写前面步骤的文件,而是重新读取当前状态来填写。如果 holder 重启并重新连接,后端可能已经不同,所以 ct.json 也必须重新生成。