TT Lab
Get started
Learn Learning paths Courses

CCA — Cilium Certified Associate

Sorting Out the Observability Pipeline and Identity

Continue in TT Lab

Goal

Declare the Hubble observation pipeline in a values file, sort out explicitly which labels go into the identity and which are excluded, and build the queries, diagnosis table, and alert rules you use straight away in practice.

Why it matters

For the loop of "observe → write policy → audit → enforce" to close in Cilium operations, the observation axis has to stand first. But if you turn Hubble on without attaching an export, the evidence disappears the moment the ring buffer rotates. The default of 4095 flows per node can be only a few seconds' worth on a busy node.

There is also a practical reason that sorting out identity labels matters. If a label that changes with every rollout, such as pod-template-hash, is included in the identity, a single deployment gives every Pod a new identity and the policy maps are recomputed wholesale. The team has to agree on which labels carry meaning and which are noise.

Alerts are the same. Right after enforcing policies, the textbook practice is two-stage operation: set the threshold deliberately low to find missing allow rules quickly, then raise it to the operating level once things have stabilized.

Write the Hubble settings and documents under /root/cca-hubble/, and actually apply the basic Kubernetes resources.

Steps

  1. Create the namespace cca-obs.
  2. Create /root/cca-hubble/hubble-values.yaml and write hubble.enabled: true, hubble.relay.enabled: true, hubble.ui.enabled: true, and hubble.eventBufferCapacity: 16383.
  3. In the same file, add a hubble.metrics.enabled list that includes all four families drop, dns, flow, and httpV2. Then put a dynamic export configuration under hubble.export, and include both DROPPED and AUDIT in the verdict of includeFilters.
  4. In the namespace cca-obs, deploy the Deployment checkout with 2 replicas. The Pod template labels are the two labels app=checkout and team=payments, and the image is nginx:1.27-alpine.
  5. In the same namespace, create the ConfigMap identity-labels. Put app, team, and io.kubernetes.pod.namespace in the key included, and pod-template-hash, controller-revision-hash, and pod-template-generation in the key excluded.
  6. In /root/cca-hubble/hubble-queries.txt, write at least 5 lines of queries that start with hubble observe. Among them there must be at least one query each with --verdict DROPPED, --verdict AUDIT, --namespace cca-obs, --protocol dns, and one that specifies a particular Pod with --from-pod. Then create a markdown table in /root/cca-hubble/drop-reasons.md and write the four reasons POLICY_DENIED, CT_MAP_INSERT_FAILED, UNSUPPORTED_L3_PROTOCOL, and STALE_OR_UNROUTABLE_IP, each with its first point of suspicion.
  7. Write a Prometheus alert rule in /root/cca-hubble/hubble-alerts.yaml. groups[0].name is cilium-policy and there are two rules. The first is named PolicyDropSpike; its expr uses the rate of hubble_drop_total narrowed to the POLICY_DENIED reason, with for: 10m and severity warning. The second is named CiliumAgentDown with severity critical.

Notes

Create the observation lab namespace

Create the namespace cca-obs.

This is a first step where you just need to get the name exactly right.

Write the Hubble enablement values

Create /root/cca-hubble/hubble-values.yaml and write hubble.enabled: true, hubble.relay.enabled: true, hubble.ui.enabled: true, and hubble.eventBufferCapacity: 16383.

You have to turn on the server, the aggregator, and the UI separately. The default ring buffer is 4095 flows per node and rotates quickly with heavy traffic, so also specify the enlarged value.

Export metrics and security events

In the same file, add a hubble.metrics.enabled list that includes all four families drop, dns, flow, and httpV2. Then put a dynamic export configuration under hubble.export, and include both DROPPED and AUDIT in the verdict of includeFilters.

Continue writing in the same file. Turn on the four families of drop, DNS, flow, and HTTP, and in the export filter include both what was actually blocked and what would have been blocked in audit mode.

Deploy a workload with labels

In the namespace cca-obs, deploy the Deployment checkout with 2 replicas. The Pod template labels are the two labels app=checkout and team=payments, and the image is nginx:1.27-alpine.

The identity is computed from the set of Pod labels, so the labels must be in the Pod template, not on the Deployment. Once deployed, the controller automatically attaches one more label.

Sort out the included and excluded labels

In the same namespace, create the ConfigMap identity-labels. Put app, team, and io.kubernetes.pod.namespace in the key included, and pod-template-hash, controller-revision-hash, and pod-template-generation in the key excluded.

If a label whose value changes with every redeployment goes into the identity, a single rollout means recomputing all the policy maps. Pick out such labels and put them in the excluded list.

Build the frequently used queries and the drop reason table

In /root/cca-hubble/hubble-queries.txt, write at least 5 lines of queries that start with hubble observe. Among them there must be at least one query each with --verdict DROPPED, --verdict AUDIT, --namespace cca-obs, --protocol dns, and one that specifies a particular Pod with --from-pod. Then create a markdown table in /root/cca-hubble/drop-reasons.md and write the four reasons POLICY_DENIED, CT_MAP_INSERT_FAILED, UNSUPPORTED_L3_PROTOCOL, and STALE_OR_UNROUTABLE_IP, each with its first point of suspicion.

For the set to be usable in practice, it must contain a verdict filter, narrowing by namespace, the DNS protocol, and targeting a specific Pod. In the diagnosis table, write what to suspect first for each reason.

Create the drop spike and agent down alerts

Write a Prometheus alert rule in /root/cca-hubble/hubble-alerts.yaml. groups[0].name is cilium-policy and there are two rules. The first is named PolicyDropSpike; its expr uses the rate of hubble_drop_total narrowed to the POLICY_DENIED reason, with for: 10m and severity warning. The second is named CiliumAgentDown with severity critical.

You must not use the cumulative counter as is; you have to look at the rate of increase. If you do not distinguish the drop reasons, the noise gets severe. Also give a duration so that it does not react to momentary blips.