CCA — Cilium Certified Associate
Why is the service unreachable when the router is connected?
Goal
On real Cilium, FRR, and Linux routers, you distinguish the scope of evidence of the BGP connection, the RIB, the FIB, and HTTP.
Why it matters
If you mistake the success of route advertisement for the success of Service reachability, you end up looking for the cause of an outage in the wrong place. Preserve the failing comparisons too, contrast them with normal forwarding, and withdraw only the advertisement without shutting down the server or the peer. All work is done only inside the k3s on this personal VM. Preparation can take several minutes, and it uses Cilium 1.20.1, FRR 8.4.4, and k3s 1.35.8+k3s1. The original experiment is IPv4 and single-node and is not a substitute for failover across multiple physical nodes.
Attach metadata.labels.lab=cca-bgp to all task resources. Do not modify or delete the 5 initial PeerConfigs, the HTTP Pod, or the cca-retire pool and Service. Do not add a default route, a static bypass, or any other policy. Step preparation does not overwrite existing files. A file you saved wrongly has to be fixed directly, and if you changed an object, update the observation record too. When the session ends, the files disappear, so export the material you need first.
Steps
- Save the output of the observation tool's inventory to /root/cca-bgp/inventory.json. Read the names and UIDs of the Pod and CiliumNode, the real Pod IP and the allocated PodCIDR, and the netns, both-side IPs, and AS of the 5 routers. Do not write the whole 10.42.0.0/16 as the node allocation range instead.
- Read the initial CiliumBGPClusterConfig cca-bgp and save and apply a corrected version to /root/cca-bgp/peers.yaml. Preserve the nodeSelector kubernetes.io/os=linux, the instance name cca, and localASN=65001. Fix only the wrong peerASN=65003 of the empty peer to 65002. Preserve the 5 peers, including the remaining rib, pod, vip, and withdraw, along with their peerAddress and peerConfigRef. empty must be Established, but with no destination advertisement and the Pod connection must fail.
- In /root/cca-bgp/rib.yaml, write and apply the CiliumBGPAdvertisement cca-rib. The labels are lab=cca-bgp and advertise=rib, and spec.advertisements has a single item with advertisementType=PodCIDR. The rib router must have a valid best BGP route for the real node PodCIDR, but it must not be in the FIB. Keep FRR's initial bgp no-rib and compare it with the Pod HTTP connection failure.
- In /root/cca-bgp/pod.yaml, write and apply the CiliumBGPAdvertisement cca-pod. The labels are lab=cca-bgp and advertise=pod, and the advertisement item is a single PodCIDR. The pod router must have that PodCIDR in both the RIB and the FIB and receive a Pod HTTP 200. In the FIB, the protocol is bgp, the gateway is 192.0.2.9, and the dev is router. Preserve the comparison state of the rib router.
- In /root/cca-bgp/services.yaml, save and apply the student pool and 2 Services as a v1 List. The pool CiliumLoadBalancerIPPool cca-student has blocks=[{start: 203.0.113.10, stop: 203.0.113.11}] and serviceSelector.matchLabels={lab: cca-bgp, pool: student}. The Services cca-selected and cca-excluded have type=LoadBalancer, loadBalancerClass=io.cilium/bgp-control-plane, externalTrafficPolicy=Cluster, selector.app=cca-bgp-web, port and targetPort=8080, and protocol=TCP. Attach lab=cca-bgp and pool=student to both Services, and set publish to selected and excluded respectively. Check that different VIPs are allocated, and preserve the initial cca-retire.
- In /root/cca-bgp/vip.yaml, save and apply the CiliumBGPAdvertisement cca-vip. The labels are lab=cca-bgp and advertise=vip, and the single item in advertisements has advertisementType=Service, service.addresses=[LoadBalancerIP], and selector.matchLabels.publish=selected. The vip router must have only the exact /32 of selected in its RIB/FIB and return 200. excluded and direct Pod access must fail to connect.
- Read the advertisement and Service UIDs and the real initial 200 in the initial material /opt/fixtures/cca-bgp/baseline.json. In /root/cca-bgp/withdrawal.json, record advertisement=cca-withdraw, advertisement_uid=the initial advertisement UID, service=cca-retire, service_uid=the initial Service UID, and action=delete-advertisement. Delete only the CiliumBGPAdvertisement cca-withdraw. The cca-retire Service, the HTTP Pod, and the withdraw peer must remain, and only the 203.0.113.20/32 RIB/FIB on the withdraw router must disappear, making the connection fail.
- Save the result of the observation tool's report to /root/cca-bgp/report.json and fill in the diagnoses with your own judgment. empty is no-advertisement, rib is not-installed, pod is pod-forwarding, vip is selected-vip, and withdraw is withdrawn. Read the current UIDs and spec hashes, the RIB/FIB, and each request identifier and server response, and try to explain the grounds for this judgment. Do not copy different requests and fill them in with the same identifier. In the full grading, all the earlier steps must also pass in their current state.
Notes
- Observation tool: python3 /opt/fixtures/cca-bgp/runtime.py observe step-name. The step names are inventory, peer, rib, pod, ipam, vip, withdraw, and report.
- FRR: ip netns exec cca-rib vtysh -N cca-rib -c 'show bgp ipv4 unicast json'
- FIB: ip -n cca-rib -j route show. For the other comparisons, replace rib in the name with the corresponding name.
- Peers: cilium bgp peers. Advertisements: kubectl get ciliumbgpadvertisement -o yaml.
- Write files as a single JSON/YAML object. For multiple resources, use v1 List.items.
- curl 7/000 is a connection failure, and 000 is not an HTTP status code from the server.
Record the identities of the real server and the five link networks
Save the output of the observation tool's inventory to /root/cca-bgp/inventory.json. Read the names and UIDs of the Pod and CiliumNode, the real Pod IP and the allocated PodCIDR, and the netns, both-side IPs, and AS of the 5 routers. Do not write the whole 10.42.0.0/16 as the node allocation range instead.
Check which node's PodCIDR the real Pod IP is in. inventory does not change the configuration.
Fix just one AS to make a connection with no advertisements
Read the initial CiliumBGPClusterConfig cca-bgp and save and apply a corrected version to /root/cca-bgp/peers.yaml. Preserve the nodeSelector kubernetes.io/os=linux, the instance name cca, and localASN=65001. Fix only the wrong peerASN=65003 of the empty peer to 65002. Preserve the 5 peers, including the remaining rib, pod, vip, and withdraw, along with their peerAddress and peerConfigRef. empty must be Established, but with no destination advertisement and the Pod connection must fail.
Cilium's peerASN is the peer FRR's AS. A peer connection and an advertisement are separate, so do not add an advertisement to empty.
A router that learned the route but does not forward
In /root/cca-bgp/rib.yaml, write and apply the CiliumBGPAdvertisement cca-rib. The labels are lab=cca-bgp and advertise=rib, and spec.advertisements has a single item with advertisementType=PodCIDR. The rib router must have a valid best BGP route for the real node PodCIDR, but it must not be in the FIB. Keep FRR's initial bgp no-rib and compare it with the Pod HTTP connection failure.
After checking FRR's RIB, read ip -n cca-rib route separately. The failing connection in this step is the intended comparison.
Confirm real HTTP forwarding with the same PodCIDR
In /root/cca-bgp/pod.yaml, write and apply the CiliumBGPAdvertisement cca-pod. The labels are lab=cca-bgp and advertise=pod, and the advertisement item is a single PodCIDR. The pod router must have that PodCIDR in both the RIB and the FIB and receive a Pod HTTP 200. In the FIB, the protocol is bgp, the gateway is 192.0.2.9, and the dev is router. Preserve the comparison state of the rib router.
Do not work around it with a static route or a default route. Also look at who installed the route and who the real next hop is.
Separate Service address allocation from advertisement
In /root/cca-bgp/services.yaml, save and apply the student pool and 2 Services as a v1 List. The pool CiliumLoadBalancerIPPool cca-student has blocks=[{start: 203.0.113.10, stop: 203.0.113.11}] and serviceSelector.matchLabels={lab: cca-bgp, pool: student}. The Services cca-selected and cca-excluded have type=LoadBalancer, loadBalancerClass=io.cilium/bgp-control-plane, externalTrafficPolicy=Cluster, selector.app=cca-bgp-web, port and targetPort=8080, and protocol=TCP. Attach lab=cca-bgp and pool=student to both Services, and set publish to selected and excluded respectively. Check that different VIPs are allocated, and preserve the initial cca-retire.
Having received an address does not mean the router learned that address. Read the LoadBalancer status first.
Advertise only the selected VIP among the same servers
In /root/cca-bgp/vip.yaml, save and apply the CiliumBGPAdvertisement cca-vip. The labels are lab=cca-bgp and advertise=vip, and the single item in advertisements has advertisementType=Service, service.addresses=[LoadBalancerIP], and selector.matchLabels.publish=selected. The vip router must have only the exact /32 of selected in its RIB/FIB and return 200. excluded and direct Pod access must fail to connect.
The advertisement label that the PeerConfig selects and the Service label that the advertisement selects are different selectors. Also check the comparison target that should fail.
Withdraw only the route, leaving the Service
Read the advertisement and Service UIDs and the real initial 200 in the initial material /opt/fixtures/cca-bgp/baseline.json. In /root/cca-bgp/withdrawal.json, record advertisement=cca-withdraw, advertisement_uid=the initial advertisement UID, service=cca-retire, service_uid=the initial Service UID, and action=delete-advertisement. Delete only the CiliumBGPAdvertisement cca-withdraw. The cca-retire Service, the HTTP Pod, and the withdraw peer must remain, and only the 203.0.113.20/32 RIB/FIB on the withdraw router must disappear, making the connection fail.
A failure made by deleting the server or the BGP peer is not evidence of advertisement withdrawal. Check that the initial UIDs are maintained.
A comprehensive incident report from the five current states
Save the result of the observation tool's report to /root/cca-bgp/report.json and fill in the diagnoses with your own judgment. empty is no-advertisement, rib is not-installed, pod is pod-forwarding, vip is selected-vip, and withdraw is withdrawn. Read the current UIDs and spec hashes, the RIB/FIB, and each request identifier and server response, and try to explain the grounds for this judgment. Do not copy different requests and fill them in with the same identifier. In the full grading, all the earlier steps must also pass in their current state.
Do not look only at the 200 in the report; read together which address, which request, and which resource lifetime the observation belongs to.