CCA — Cilium Certified Associate
Gateway API and One Envoy per Node — How the Cilium Service Mesh Is Built
In one line
Cilium translates Gateway API resources into CiliumEnvoyConfig and loads them onto the one Envoy on each node, and eBPF intercepts traffic arriving at a Service port and redirects it to that Envoy. This structure, in which no sidecar is attached to each Pod, is the sidecarless mesh, and encryption is applied transparently between nodes with WireGuard or IPsec.
Why this was needed
The Ingress resource was standardized only as far as splitting by host and path. To split traffic by weight, modify headers, or rewrite URLs, you had to use annotations that differed from controller to controller, and those annotations did not work when you moved to another controller. It handled only HTTP and HTTPS, TCP and UDP were not in the standard, and it did not suit clusters where several teams share one load balancer. The Cilium documentation sums these up as the limits of Ingress and puts the Gateway API, which Kubernetes SIG-Network designed as the successor, as the answer.
The mesh side had the same kind of problem. Attaching a sidecar proxy to every Pod gives you L7 features, but the proxies multiply as fast as the Pods. From the start, Cilium chose a structure in which IP, TCP, and UDP handling is done in the kernel with eBPF, and only application protocols such as HTTP, gRPC, and DNS are handed to Envoy. The Service Mesh domain of the CCA (16%) asks about the reasons for this structure, the advantages of the Gateway API, and the encryption options.
How it works
Gateway API resources and the separation of roles
The Gateway API is designed to be role-oriented. The three roles the documentation cites are the Infrastructure Provider, the Cluster Operator, and the Application Developer. GatewayClass says which implementation handles the gateway (for Cilium, cilium), Gateway states the listeners (protocol and port) and which namespaces' Routes it accepts, and HTTPRoute states the matching rules and backends. You can split permissions so that developers create only Routes in their own namespaces and cannot touch the Gateway configuration.
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: example-route-1
spec:
parentRefs:
- name: cilium-gw # 어느 Gateway 에 붙는가
rules:
- matches:
- path: {type: PathPrefix, value: /echo}
backendRefs:
- {kind: Service, name: echo-1, port: 8080, weight: 50}
- {kind: Service, name: echo-2, port: 8090, weight: 50}
The weight in backendRefs is the traffic split. The documentation's example starts at 50/50 and changes to 99/1, and explains that it is used for canary and A/B scenarios. It is a standard field, not an annotation, so it works unchanged even if you move to another implementation.
Cilium supports GatewayClass, Gateway, HTTPRoute, GRPCRoute, TLSRoute, BackendTLSPolicy, and ReferenceGrant of Gateway API v1.6.1 and has passed the Core conformance tests, while TCPRoute, UDPRoute, and ListenerSet are turned on only when their CRDs are installed. The prerequisites are kubeProxyReplacement=true and l7Proxy=true (the default), and you turn on the controller with gatewayAPI.enabled=true. By default it creates a Service of type LoadBalancer, and from 1.16 it can also be exposed directly on the host network.
How it differs from other Ingress controllers
What the documentation calls "the biggest difference" is how closely the implementation is tied to the CNI. An ordinary controller runs separately as a Deployment or DaemonSet and is exposed through a LoadBalancer Service. Cilium's Ingress and Gateway API are part of the network stack, so when traffic arrives at a Service port, eBPF code intercepts it and hands it to Envoy via TPROXY. As a result, CiliumNetworkPolicy also applies to traffic entering and leaving through the ingress.
Traffic that arrives at Envoy at this point receives a special ingress identity in the policy engine. Traffic from outside usually has the world identity, so there are two policy decision points: the stage of entering the ingress from world, and the stage of leaving the ingress toward the backend identity. If you applied a default deny policy and the gateway is blocked, you have not allowed one of these two stages.
The internal operation is split into two parts. The Cilium operator watches the Gateway API resources, validates them, marks them as accepted, and translates them into CiliumEnvoyConfig. The Cilium agent reads that CiliumEnvoyConfig and gives the configuration to the built-in Envoy or the Envoy DaemonSet. Envoy handles the traffic.
A mesh without sidecars
The structure of using one Envoy per node applies in the same way to east-west traffic. GAMMA sets the parent of an HTTPRoute to a Service rather than a Gateway, and Cilium intercepts the L7 traffic going to that Service and routes it to the per-node Envoy. Cilium currently supports only producer Routes that are in the same namespace as the Service. Traffic that does not use L7 goes straight to the eBPF datapath without passing through Envoy. Compared with the sidecar approach, the number of proxies is proportional to the number of nodes, not the number of Pods, and you can turn L7 features on and off without restarting application Pods.
Encryption options — WireGuard, IPsec, mTLS
Cilium transparently encrypts traffic between endpoints managed by Cilium with IPsec, WireGuard, or ztunnel (beta). When you turn on WireGuard, the agent on each node establishes a WireGuard tunnel with every other node, creates a key pair per node, and distributes the public key through the network.cilium.io/wg-pub-key annotation on the CiliumNode resource. The tunnel endpoint is UDP 51871, and traffic within the same node is not encrypted (because it can be seen on the node anyway, so there is no benefit). The kernel must support WireGuard, and you turn it on with encryption.enabled=true and encryption.type=wireguard. IPsec distributes keys as a Kubernetes Secret, and from 1.18 it encrypts after tunnel encapsulation, hiding even the identities used for policy.
Mutual Authentication (beta) issues workload identities with SPIFFE/SPIRE and performs the mTLS handshake out-of-band, outside the ordinary connection. The documentation states that to get both authentication and confidentiality, you must turn on encryption (WireGuard or IPsec). In other words, Cilium's mTLS is identity verification, and the actual encryption is handled by the tunnels between nodes.
What it looks like in the field
The first thing a team moving from NGINX Ingress to the Cilium Gateway API runs into is annotations. The documentation says that carrying implementation-specific annotations over as they are is rare, and guides you to replace request and response manipulation, traffic splitting, and routing based on headers, queries, and methods with the standard fields of the Gateway API. The recommended order is to decide the scope and stages, classify the existing annotations, create Gateway API resources with the same meaning and verify them in parallel, move traffic gradually, and delete the Ingress once things are stable.
There are also reports that "the first few packets went out in plaintext" after turning on WireGuard. That is exactly the known-issues item in the documentation. Before the information that the destination IP is a remote Cilium endpoint has propagated, it is treated as outside the cluster and sent in plaintext. The responses the documentation suggests are to turn on encryption.strictMode or to block egress to world with a policy.
What to look at in the next reading
In the reading that follows right away, you will look at how routers outside the cluster learn the LoadBalancer IP that a Gateway received and the Pod CIDRs, that is, the BGP Control Plane and the Egress Gateway. References: Gateway API Support, Traffic Splitting Example, Migrating from Ingress to Gateway, WireGuard Transparent Encryption, Mutual Authentication.