CCA — Cilium Certified Associate
Where Flows Are Kept, and How Much
In one line
Hubble puts the events made by the eBPF datapath into a node-local ring buffer, and relay merges them at the cluster level. The ring buffer rotates, so export is essential for long-term analysis.
Why this was needed
Once you start enforcing policies, the questions pour in. "Why did the payment API just get cut off?" "Is this drop because of a policy or because of routing?" As the datapath has moved into the kernel, the means of observation must come from the same depth.
The old way attached a sidecar proxy to observe requests. A proxy attached to every Pod doubles the resources, and changing the Pod spec means a restart is required. Hubble reads the events made by the eBPF programs already attached to the datapath as they are, so there is no extra proxy and no Pod change.
How it works
The structure has three layers.
[eBPF 데이터패스] --perf 링버퍼--> [cilium-agent 안의 Hubble 서버]
gRPC :4244
|
[hubble-relay] :4245
|
+-------------+-------------+
v v
[hubble CLI] [hubble-ui]
Three properties that matter operationally come out of this.
First, flows exist only briefly in node-local memory. The default buffer is 4095 flows per node, and a commonly used value when you enlarge it is 16383. A flow takes up about 500 bytes, so 16383 comes to roughly 8MiB per node. On a busy node it can rotate within seconds, so you must attach a file export for audit evidence and post-incident analysis.
Second, relay is only an aggregator, not a store. Even if relay dies, the datapath and node-local observation are not affected at all. The ports are 4245 for relay, 4244 for the node's Hubble server, and 9965 for the metrics endpoint.
Third, the payload is not stored. A flow is a structured event that carries the source and destination identities and labels, the verdict, the drop reason, and HTTP/DNS metadata. If you need the packet contents, you have to use a different tool.
Knowing how to read the verdict values is the core of practice.
| verdict | Meaning |
|---|---|
| FORWARDED | Allowed and forwarded |
| DROPPED | Blocked (a drop reason is attached) |
| AUDIT | Audit mode was on, and this is traffic that would have been blocked if enforced |
| ERROR | An error during processing |
It also saves time to have the first suspect for each drop reason sorted out. POLICY_DENIED means the policy did not allow it, so recheck the identity and the port. CT_MAP_INSERT_FAILED means the conntrack map is full, so adjust the map size. UNSUPPORTED_L3_PROTOCOL is non-IP traffic. STALE_OR_UNROUTABLE_IP is an ipcache mismatch, so check the synchronization between nodes.
Finally, something you must distinguish. An L7 denial looks different in the log from an L4 drop. When traffic is blocked at L4, the connection is not established and it ends with a single request line, but when an L7 rule is violated, the proxy generates a 403 on an already established connection and returns it, so the request is logged as DROPPED and the response as FORWARDED. Once you know this asymmetry, you can tell which layer the problem is in from just two log lines.
What it looks like in the field
When the author put Hubble on the homelab, hubble-relay and hubble-ui did not move from Pending. The event was this.
Warning FailedScheduling 0/1 nodes are available: 1 node(s) had untolerated taint(s).
The cause was scheduling. The control plane node has the taint node-role.kubernetes.io/control-plane:NoSchedule, and hubble-relay and hubble-ui are Deployments, not DaemonSets, so they had no toleration. This contrasts with CoreDNS, which has the control-plane toleration by default and started normally. It was resolved as soon as a worker joined, and this was normal behavior, not an error.
There is a sense to take away from this. The Cilium Agent and the Hubble server are DaemonSets because they must exist on every node, and relay and the UI are Deployments because one per cluster is enough. When the deployment shape differs, the scheduling constraints differ too. If "all the agents are up but only the hubble CLI cannot connect," the right order is to look at the relay Pod's state first.
The logs from applying an L7 policy in the same cluster are just as seen in the previous module. The GET request is FORWARDED, the POST request is DROPPED, and the 403 response to that POST is FORWARDED. Those three log lines were the evidence that the policy worked as intended.
What you will do in the next lab
You will write the values that enable Hubble along with the metrics and export settings, put labels on a real workload and sort out, in a ConfigMap, which labels go into the identity and which are excluded, and build frequently used flow filter queries, a drop reason diagnosis table, and even an alert rule for drop spikes.