CCA — Cilium Certified Associate
Sorting Out the Observability Pipeline and Identity
Goal
Declare the Hubble observation pipeline in a values file, sort out explicitly which labels go into the identity and which are excluded, and build the queries, diagnosis table, and alert rules you use straight away in practice.
Why it matters
For the loop of "observe → write policy → audit → enforce" to close in Cilium operations, the observation axis has to stand first. But if you turn Hubble on without attaching an export, the evidence disappears the moment the ring buffer rotates. The default of 4095 flows per node can be only a few seconds' worth on a busy node.
There is also a practical reason that sorting out identity labels matters. If a label that changes with every rollout, such as pod-template-hash, is included in the identity, a single deployment gives every Pod a new identity and the policy maps are recomputed wholesale. The team has to agree on which labels carry meaning and which are noise.
Alerts are the same. Right after enforcing policies, the textbook practice is two-stage operation: set the threshold deliberately low to find missing allow rules quickly, then raise it to the operating level once things have stabilized.
Write the Hubble settings and documents under /root/cca-hubble/, and actually apply the basic Kubernetes resources.
Steps
- Create the namespace
cca-obs. - Create
/root/cca-hubble/hubble-values.yamland writehubble.enabled: true,hubble.relay.enabled: true,hubble.ui.enabled: true, andhubble.eventBufferCapacity: 16383. - In the same file, add a
hubble.metrics.enabledlist that includes all four familiesdrop,dns,flow, andhttpV2. Then put a dynamic export configuration underhubble.export, and include bothDROPPEDandAUDITin the verdict of includeFilters. - In the namespace
cca-obs, deploy the Deploymentcheckoutwith 2 replicas. The Pod template labels are the two labelsapp=checkoutandteam=payments, and the image isnginx:1.27-alpine. - In the same namespace, create the ConfigMap
identity-labels. Putapp,team, andio.kubernetes.pod.namespacein the keyincluded, andpod-template-hash,controller-revision-hash, andpod-template-generationin the keyexcluded. - In
/root/cca-hubble/hubble-queries.txt, write at least 5 lines of queries that start withhubble observe. Among them there must be at least one query each with--verdict DROPPED,--verdict AUDIT,--namespace cca-obs,--protocol dns, and one that specifies a particular Pod with--from-pod. Then create a markdown table in/root/cca-hubble/drop-reasons.mdand write the four reasonsPOLICY_DENIED,CT_MAP_INSERT_FAILED,UNSUPPORTED_L3_PROTOCOL, andSTALE_OR_UNROUTABLE_IP, each with its first point of suspicion. - Write a Prometheus alert rule in
/root/cca-hubble/hubble-alerts.yaml.groups[0].nameiscilium-policyand there are two rules. The first is namedPolicyDropSpike; its expr uses therateofhubble_drop_totalnarrowed to thePOLICY_DENIEDreason, withfor: 10mand severitywarning. The second is namedCiliumAgentDownwith severitycritical.
Notes
- The quickest way to create the ConfigMap is
kubectl create configmap identity-labels -n cca-obs --from-literal=included=... --from-literal=excluded=.... You can join the values with commas. - The lines of a markdown table start with
|. Counting the header line and the separator line, it comes to 6 or more lines. - Common mistake 1: attaching the labels only to the Deployment's
metadata.labels. The identity is computed from the Pod labels, so they must be inspec.template.metadata.labels. - Common mistake 2: using the cumulative counter as is in the alert expr. You have to look at the rate of increase to catch a spike.
foris a YAML key, so you can write it as is, without quotes.
Create the observation lab namespace
Create the namespace cca-obs.
This is a first step where you just need to get the name exactly right.
Write the Hubble enablement values
Create /root/cca-hubble/hubble-values.yaml and write hubble.enabled: true, hubble.relay.enabled: true, hubble.ui.enabled: true, and hubble.eventBufferCapacity: 16383.
You have to turn on the server, the aggregator, and the UI separately. The default ring buffer is 4095 flows per node and rotates quickly with heavy traffic, so also specify the enlarged value.
Export metrics and security events
In the same file, add a hubble.metrics.enabled list that includes all four families drop, dns, flow, and httpV2. Then put a dynamic export configuration under hubble.export, and include both DROPPED and AUDIT in the verdict of includeFilters.
Continue writing in the same file. Turn on the four families of drop, DNS, flow, and HTTP, and in the export filter include both what was actually blocked and what would have been blocked in audit mode.
Deploy a workload with labels
In the namespace cca-obs, deploy the Deployment checkout with 2 replicas. The Pod template labels are the two labels app=checkout and team=payments, and the image is nginx:1.27-alpine.
The identity is computed from the set of Pod labels, so the labels must be in the Pod template, not on the Deployment. Once deployed, the controller automatically attaches one more label.
Sort out the included and excluded labels
In the same namespace, create the ConfigMap identity-labels. Put app, team, and io.kubernetes.pod.namespace in the key included, and pod-template-hash, controller-revision-hash, and pod-template-generation in the key excluded.
If a label whose value changes with every redeployment goes into the identity, a single rollout means recomputing all the policy maps. Pick out such labels and put them in the excluded list.
Build the frequently used queries and the drop reason table
In /root/cca-hubble/hubble-queries.txt, write at least 5 lines of queries that start with hubble observe. Among them there must be at least one query each with --verdict DROPPED, --verdict AUDIT, --namespace cca-obs, --protocol dns, and one that specifies a particular Pod with --from-pod. Then create a markdown table in /root/cca-hubble/drop-reasons.md and write the four reasons POLICY_DENIED, CT_MAP_INSERT_FAILED, UNSUPPORTED_L3_PROTOCOL, and STALE_OR_UNROUTABLE_IP, each with its first point of suspicion.
For the set to be usable in practice, it must contain a verdict filter, narrowing by namespace, the DNS protocol, and targeting a specific Pod. In the diagnosis table, write what to suspect first for each reason.
Create the drop spike and agent down alerts
Write a Prometheus alert rule in /root/cca-hubble/hubble-alerts.yaml. groups[0].name is cilium-policy and there are two rules. The first is named PolicyDropSpike; its expr uses the rate of hubble_drop_total narrowed to the POLICY_DENIED reason, with for: 10m and severity warning. The second is named CiliumAgentDown with severity critical.
You must not use the cumulative counter as is; you have to look at the rate of increase. If you do not distinguish the drop reasons, the noise gets severe. Also give a duration so that it does not react to momentary blips.