TT Lab
Get started
Learn Learning paths Courses

CCA — Cilium Certified Associate

Four layers between a green BGP session and an HTTP response

Continue in TT Lab

In one line

A BGP peer being Established is evidence that a route-exchange connection was made, not evidence that the user's HTTP request succeeded. This module puts five routers in front of the same server and compares the peer connection, the receive path (the RIB), the kernel forwarding path (the FIB), and the actual response separately. The goal is to explain which observation supports which claim, without lumping different failures together as the same timeout.

Why this was needed

On the on-call screen, four of the five peers are green. Yet the service owner says they still cannot connect. If you restart all the agents here, the routes change briefly or past logs disappear, making the cause even harder to find. First you have to ask whether the success the operator confirmed and the success the user expected are the same. Making a TCP connection, learning a prefix, installing a route in the kernel, and receiving the server's response are different events.

The lab has a peer with the AS number written wrongly, a peer that only connects, a router that learns routes but does not install them in the kernel, and a router that forwards normally. So that fixing one does not make the other comparisons disappear, each router is placed separately in a Linux network namespace. The interfaces and routing tables are separated per namespace, but everything is inside the personal VM, so it does not touch your home router or the platform's production BGP. The addresses below are for internal use in the lab, and you should not carry them as they are into the routing configuration of a real organization.

How it works

Read the addresses first and get the connection direction right

Each router is directly connected to the VM with a veth. The AS of Cilium on the VM side is 65001, and the AS of FRR on the router side is 65002. Cilium's peerASN must contain the peer FRR's 65002. Conversely, FRR's remote-as contains Cilium's 65001. Local and remote are not directions permanently fixed in the documentation; they are names that change depending on which device's configuration you are reading right now.

Comparison router VM interface and IP Router IP Role
empty bgp-empty, 192.0.2.1 192.0.2.2 Keep the connection with no advertisements after fixing the wrong AS
rib bgp-rib, 192.0.2.5 192.0.2.6 Learns the PodCIDR but does not install it in the kernel
pod bgp-pod, 192.0.2.9 192.0.2.10 Learns the PodCIDR and forwards to the real Pod
vip bgp-vip, 192.0.2.13 192.0.2.14 Reaches only the selected Service VIP
withdraw bgp-withdraw, 192.0.2.17 192.0.2.18 Withdraws only the initial Service's advertisement

Each link network is a /30 and the router's interface name is router. FRR is in a passive wait state and Cilium initiates the connection. Only the empty peer has Cilium's peerASN wrongly set to 65003. The first fix is not to rebuild the whole network but just to set that value to 65002. The sourceInterface specifies the VM veth per peer, so you can also tell apart the mistake of trying to connect to an address on a different link network. Check each direction of the configuration in Cilium BGP resources and peer configuration.

PodCIDR advertisement and the kernel route are not one step

The PodCIDR of a CiliumBGPAdvertisement advertises the Pod network allocated to the node. The whole pool in this lab is 10.42.0.0/16, but the range actually allocated to the node must be read from the CiliumNode. The observation tool also checks whether the real Pod IP is within that range. If you memorize and write down the whole /16, or fill in the table with addresses that were not observed, you cannot tell which node the route belongs to.

The rib router is deliberately configured with FRR's bgp no-rib. Even in this state, the received best path may be visible in the BGP RIB, but that route is not installed in the kernel. What the student must do here is not to restore all the routers the same way, but to preserve this comparison state and explain why HTTP fails. On the pod router, the same PodCIDR is advertised but FRR is left to install the route in the kernel, and this is compared with HTTP 200. It is normal for the two routers to give different results. When you read FRR 8.4's explanation of no-kernel and no-rib, also check which layer's table the word RIB refers to.

In the FIB you do not look only at the destination prefix. You also check whether the protocol is bgp, whether the next hop is the VM address of that link network, and whether the output device is router. An answer made to succeed by adding a default route or a static /32 does not prove BGP forwarding success. iproute2 JSON may show a /32 destination as an IP without a suffix, so compare the normalized network, not the shape of the string.

IP allocation and Service advertisement are also separated

LoadBalancer IPAM allocates addresses for use by Services. The fact that an address exists does not mean the external router knows that address. In step 5, you create the student pool 203.0.113.10–11 and only confirm that the selected and excluded Services receive different addresses. Their backend selectors are the same and they point to the same HTTP Pod. In this state, do not conclude from the address alone that advertisement has happened. Read the address allocation of LB IPAM and the advertisement of BGP as separate responsibilities.

In step 6, you pick only publish=selected with the Service selector of the CiliumBGPAdvertisement. The vip router must have the exact /32 of selected, and you do not add excluded's route or the PodCIDR. Even though the same backend is alive, excluded must not be reachable. If you check only one normal response, you miss a configuration that accidentally published both Services, so you test the target that should succeed and the target that should fail together. You must also distinguish that selectors exist in two layers: the layer where the PeerConfig selects advertisement resources, and the layer where the advertisement resource selects Services.

A withdrawal is not shutting down the server

The Service cca-retire for the withdrawal comparison is pre-allocated 203.0.113.20, separately from the student pool. Initialization confirms the actual advertisement UID and Service UID, the RIB/FIB, and HTTP 200, and then leaves baseline.json. The student reads that evidence and deletes only the cca-withdraw advertisement. The HTTP server and the Service remain and the peer is still Established, but you must create a state in which the destination route on the withdraw router disappears and the connection fails.

Creating a failure by deleting the Service or the peer is a different experiment. If the destination server itself disappears, you cannot isolate the effect of withdrawing the advertisement. So grading checks that the initial Pod, node, and Service UIDs are maintained, and also compares the current peers and routes and a new HTTP request. Changes to Cilium advertisements are propagated asynchronously, so do not assume the route disappeared as soon as the delete API finished; grade after the observation command confirms convergence. The BGP operation guide deals together with control plane changes and the continuity of forwarding.

Interpret the same connection failure together with the route table

You cannot conclude that a route is absent from curl exit code 7 alone. Even if the route exists, the same code can appear when the destination port refuses the connection. In this lab you use together the comparison that the same server responds normally on the pod router and the observation that there is no destination route in the RIB/FIB of the failing router. Conversely, the rib comparison has a RIB and no FIB, so even if the HTTP result is the same as empty, the layer that stopped is different. Not merging results with the same number into one cause is the core of an incident report.

Even if the observation tool summarizes the reason for a failure, read the original output yourself. First look at the intended AS, peers, and status conditions with kubectl get ciliumbgpclusterconfig cca-bgp -o yaml, and look at the connection state with cilium bgp peers. Then compare show bgp ipv4 unicast json and ip -n cca-pod -j route show on each router. If you have only evidence that the configuration was saved and the router has no receive route, the advertisement may not have converged yet or the selector may not match. Fixing the Service deployment first at that point would be touching a different layer.

If NoMatchingNode appears in the status conditions, first check the node selector, and if MissingPeerConfigs appears, first check the existence of the referenced PeerConfig. These conditions narrow the scope of the cause, but the fact that a condition disappeared does not prove HTTP success. In a task like this lab, which preserves the initial peers and servers, you must record what you changed even when you roll back a wrong experiment. If you create a new object with the same name, the UID differs, and the earlier record is no longer evidence from the same environment.

What it looks like in the field

In LabHub's recipe probe, it came to light that a router namespace on a single VM is not exactly the same as a real external machine. When socket-LB swaps the VIP for the Pod IP at connect time, you could think you were testing the external router's VIP route while actually testing the Pod route. This environment turns on socketLB.hostNamespaceOnly from the initial installation and compares using packet handling on the veth. Among several veths, it also specifies the direct routing device. This is not a cure-all setting to apply to every production outage but a condition of the measurement environment. Read together the prerequisites of socket-LB bypass and device selection.

The name of a failure must also be written accurately. If curl ends with 7 and the output is 000, this lab records it as a connection failure. 000 is not an HTTP code sent by a server. Do not write it as the same thing as an application's HTTP 500 or a NetworkPolicy timeout. Conversely, an HTTP 200 is also insufficient if it is the response of another server, so the real server returns this request's identifier in its response. The report bundles the current resource UIDs and configuration hashes, each router's routes, and the responses to different requests.

What you will do in the next lab

You record the identities of the initial environment and fix one wrong peer AS. Then you create a comparison that has only the RIB and real Pod forwarding, and separate address allocation from selective advertisement. Finally, after withdrawing the advertisement of the initial Service, you write a report in which the five states remain at the same time. After a configuration change or a Pod recreation, do not reuse a stale report. This lab is an IPv4, single-node comparison on a personal VM, and you should not extend the claim to say that it verified ECMP distribution across multiple physical nodes or failover.