CCA — Cilium Certified Associate
The Design You Must Settle Before Adding Clusters
In one line
ClusterMesh synchronizes the identities and Services of multiple clusters to keep the same policy model as a single cluster. However, the CIDR and CA design is almost impossible to change later, so it must be settled at the start.
Why this was needed
As soon as you have more than one cluster, questions arise right away. Can you fail over to another cluster when a region fails? Can you apply the same security policy to calls between clusters? You could solve it with a service mesh, but that adds another proxy layer.
ClusterMesh takes a different approach. Without creating an extra hop in the data plane, it connects only the control plane. Each cluster exposes its own Services, identities, and endpoints through clustermesh-apiserver, and the agents of the other clusters watch them. Once synchronization is done, packets flow directly in the existing routing mode. There is no central gateway and no single point of failure.
How it works
The prerequisites are themselves design decisions.
- Each cluster needs a unique name and ID (1–255). The ID is 8 bits wide, so you can connect up to 255 clusters.
- The PodCIDR and node IP ranges must not overlap. If they overlap, the same address points to different Pods in two clusters, and the ipcache mapping cannot hold. This is the most common design mistake, and fixing it in production is effectively a rebuild.
- All clusters must share the same CA. Communication between clusters is also mutually authenticated with mTLS, and if the roots of trust differ, you cannot verify the other side's certificate. That is why, before attaching a second cluster, you copy in the CA of the first cluster.
What gets synchronized are the Services (those marked as global), the identities, and the ipcache. The key is that identities keep their meaning across the boundary. The identity of an app=backend Pod in cluster B is also propagated to the ipcache of cluster A, so multi-cluster policy works with exactly the same label syntax as a single cluster.
A global Service is created with annotations.
| Annotation | Value | Behavior |
|---|---|---|
| service.cilium.io/global | "true" | Sums up backends across clusters |
| service.cilium.io/affinity | local | Prefers local backends, goes remote if they are all gone |
| service.cilium.io/affinity | remote | Prefers remote (drain, canary) |
| service.cilium.io/affinity | none | Distributes evenly across all clusters |
The practical standard is affinity: local. In normal times it handles traffic inside the cluster to save latency, and when all local backends become unhealthy, it moves to remote. But there is a hidden dependency here. The "unhealthy" judgment is readiness-based. If the readinessProbe is sloppy (for example, Ready as long as the process is alive), a backend that is in fact broken keeps staying Ready and failover does not happen. A sloppy readinessProbe makes for sloppy failover.
There is also a pitfall on the policy side. In a ClusterMesh environment, if you leave io.cilium.k8s.policy.cluster out of a label selector, Pods with the same label in both clusters all match. This is where the incident comes from: you intended "the DB is reachable only from apps in its own cluster," but workloads with the same name in the remote cluster get opened too.
There are also a few operating rules. A Cilium version skew between connected clusters is supported only up to one minor version, so the principle is "upgrade one cluster at a time, and do not move to the next minor until all are on the same version." When you remove a cluster, you must first clear the global Service dependencies. If there is a Service that depended on remote backends with affinity: none, its backends shrink at the moment of disconnect. And even if clustermesh-apiserver dies, traffic keeps flowing with the Services and identities that have already been synchronized. What stops is the propagation of new changes.
The External Workloads feature installs an agent on a VM or bare metal so that it participates in the cluster. That workload also receives an identity, so the same label-based policy applies.
What it looks like in the field
The author's homelab is still a single cluster, but while expanding from 3 nodes to 7 nodes it learned exactly the same kind of lesson.
Having grown the control plane to 3 machines and secured 3 etcd members, I thought it was HA, but it was not. The controlPlaneEndpoint in kubeadm-config was not a VIP or DNS but cp-1's physical IP 10.0.0.120, and that address was baked into the kubelet settings of the 7 nodes, all the kubeconfigs, and the apiserver certificate SAN. If cp-1 dies, the etcd quorum is fine at 2/3 and the other apiserver processes are also healthy, but it becomes a state where nobody can find the door.
This case maps directly onto ClusterMesh design. Data availability and access availability of the control plane are separate problems, and values such as the endpoint, the CIDR, and the CA are among the few settings that are extremely painful to change after the cluster is created. So even if you plan to use only one cluster, choosing non-overlapping PodCIDRs and doing the CA design in advance will save your future self. In the same homelab, the WireGuard encryption has node-to-node tunnels established with Peers: 2 and port 51871, and this transport encryption layer extends in the same way across clusters.
What to check in the next quiz
ClusterMesh can be practiced only with two or more clusters, so this module ends with concepts and a quiz. Check the cluster ID range, the CIDR and CA prerequisites, global Service affinity and its readiness dependency, and the pitfall of leaving out the cluster label in a policy.