CCA — Cilium Certified Associate
See Who Would Be Blocked Before Enforcing a New Policy
Goal
On a real Cilium 1.20.1 (enable-policy=default), you confirm when policies start blocking, and what. You look at per-direction default deny, trying a policy out in advance with endpoint audit mode, the real blocking after audit is turned off, the opposite meaning of ingress: [{}], enableDefaultDeny=false, and how two kinds of policy end up in a single store, using the endpoint enforcement columns, real requests, and Hubble verdicts together.
Why it matters
The moment you first apply an egress policy to a service in production is the most dangerous one. If there is a dependency that was not in the documentation (the legacy in this lab), that call is cut immediately. Cilium's audit mode computes the verdict but does not deny, and leaves it in the flow as AUDITED, so you can see who would be blocked first and only then enforce.
However, audit mode works per endpoint, so it also loosens the directions that were already being enforced. The same-looking YAML can also mean different things in a Kubernetes NetworkPolicy and in a CiliumNetworkPolicy. Even if the enforcement column says Enabled, there is no guarantee that default deny is on. Differences like these are easy to get confused about just from reading the documentation, so you check by sending real requests.
If you change the cluster-wide mode to always, even endpoints without a policy (such as coredns) get blocked. On the development VM it takes about 70 seconds after reverting for the service to come back, so this lab does not change it. Audit mode is also turned on for only one endpoint.
Environment preparation takes about 5 minutes. When the session ends, the files in /root/cca-audit disappear.
Steps
- Run
kubectl apply -f /opt/fixtures/cca-audit/workload.yamlto start the Pods api, ledger, legacy, frontend, and batch (all 8080 servers) and Services with the same names in cca-audit. Before applying any policy, create/root/cca-audit/baseline.json— underpods, for each of the five app names,uid(the Pod),endpoint_id,identity(from the status of the CiliumEndpoint), andingressandegress(the POLICY column values from the agent'scilium-dbg endpoint list), and at the top levelbatch_to_api(the HTTP code string of a request to http://api:8080/ from the batch Pod; "000" if there is no response) andrecorded_at(the UTC time of the moment you record it, fromdate -u +%Y-%m-%dT%H:%M:%SZ). - In cca-audit, create a CiliumNetworkPolicy
api-ingress— endpointSelector app=api, one ingress entry (fromEndpoints app=frontend, toPorts TCP "8080"), and no egress. Once the policy takes effect, record in/root/cca-audit/direction.jsonfrontend_to_api,batch_to_api, andapi_to_legacy(the HTTP code string of each request; "000" for no response),api_ingressandapi_egress(the POLICY column values of the api endpoint), andrecorded_at(the UTC time of the moment you record it). You must record this before the egress policy of step 3. - Turn on audit mode on the api endpoint (
cilium-dbg endpoint config <api 번호> PolicyAuditMode=Enabled; the placeholder is the api endpoint number). Then apply the CNPapi-egress— endpointSelector app=api, two egress entries: (1) toEndpoints app=ledger, TCP "8080" and (2) toEndpointsk8s:io.kubernetes.pod.namespace: kube-systemandk8s:k8s-app: kube-dns, port "53" protocol ANY. After you request api→ledger, api→legacy, and batch→api, read the cca-audit flows withhubble observeinside the agent and record them in/root/cca-audit/audit.json—endpoint_id,audit_mode,api_ingressandapi_egress(the POLICY columns),api_to_legacyandbatch_to_api(code strings), andflows(the list of compact output lines that contain the AUDITED flows of api→legacy and batch→api). - Turn off audit mode on the api endpoint (PolicyAuditMode=Disabled). Send the same requests again and record in
/root/cca-audit/enforce.jsonaudit_mode,api_ingressandapi_egress,api_to_ledger,api_to_legacy, andbatch_to_api(code strings), andflows(the list of compact lines that contain the DROPPED flows of api→legacy and batch→api after audit was turned off). - In cca-audit, create two policies — the standard NetworkPolicy
np-empty(podSelector app=batch, policyTypes [Ingress],ingress: [{}]) and the CiliumNetworkPolicycnp-empty(endpointSelector app=frontend,ingress: [{}]). Send requests from the ledger Pod to batch and to frontend, and record in/root/cca-audit/empty.jsonbatch_from_ledgerandfrontend_from_ledger(code strings), andbatch_ingressandfrontend_ingress(the ingress POLICY column of each endpoint). - In cca-audit, create the CNP
ledger-observe— endpointSelector app=ledger,enableDefaultDeny: {ingress: false}, and one ingress entry (fromEndpoints app=api, TCP "8080"). Send requests api→ledger and batch→ledger, read the flows, and record in/root/cca-audit/observe.jsonledger_ingress(the POLICY column),api_to_ledgerandbatch_to_ledger(code strings), andflows(the list of compact lines that contain the ledger-side INGRESS policy-verdict lines of the two requests). - Read
cilium-dbg policy get -o jsonin the agent and find the labelsio.cilium.k8s.policy.nameandio.cilium.k8s.policy.derived-fromfor each rule of the cca-audit namespace. Record in/root/cca-audit/sources.jsonderived_from(a dictionary from policy name to derived-from value, five entries) andrevision(the revision number in the output). - In
/root/cca-audit/report.txt, write seven lines in키=값form (key=value) —api_audit_mode(the current PolicyAuditMode of api),api_policy_columns(the ingress column/egress column of api, in the form A/B for example),api_to_legacy(allowed or denied),np_empty_ingress_ruleandcnp_empty_ingress_rule(allow-all or deny-all for each),ledger_ingress_column, andledger_default_deny_ingress(the enableDefaultDeny.ingress value of ledger-observe, true or false). All values must match the current state.
Notes
- Policy enforcement modes and the endpoint's default policy (enableDefaultDeny): https://docs.cilium.io/en/v1.20/security/policy/intro/
- L3 rules and Ingress/Egress Default Deny (the
- {}example): https://docs.cilium.io/en/v1.20/security/policy/layer3/ - Creating policies from verdicts (audit mode, per-endpoint settings): https://docs.cilium.io/en/v1.20/security/policy-creation/
- Kubernetes NetworkPolicy: https://kubernetes.io/docs/concepts/services-networking/network-policies/
- Hubble flows:
kubectl -n kube-system exec ds/cilium -c cilium-agent -- hubble observe --since <시각> --namespace cca-audit -o compact(the placeholder is the start time). - Enforcement columns:
kubectl -n kube-system exec ds/cilium -c cilium-agent -- cilium-dbg endpoint list. Get endpoint numbers withkubectl -n cca-audit get cep.
Write down the enforcement state first, when there are no policies at all
Run kubectl apply -f /opt/fixtures/cca-audit/workload.yaml to start the Pods api, ledger, legacy, frontend, and batch (all 8080 servers) and Services with the same names in cca-audit. Before applying any policy, create /root/cca-audit/baseline.json — under pods, for each of the five app names, uid (the Pod), endpoint_id, identity (from the status of the CiliumEndpoint), and ingress and egress (the POLICY column values from the agent's cilium-dbg endpoint list), and at the top level batch_to_api (the HTTP code string of a request to http://api:8080/ from the batch Pod; "000" if there is no response) and recorded_at (the UTC time of the moment you record it, from date -u +%Y-%m-%dT%H:%M:%SZ).
When enable-policy in cilium-config is default, an endpoint that no policy selects is not blocked. The two POLICY columns of endpoint list show, per direction, whether default deny is on. The Pods have no curl, so send requests with python3's urllib. This record must be made before the later policies exist — the grader compares recorded_at with the creation time of the policies (it does not look at the file's modification time, so it is fine to fix and save it later).
Only the incoming side was locked, and the outgoing side is unchanged
In cca-audit, create a CiliumNetworkPolicy api-ingress — endpointSelector app=api, one ingress entry (fromEndpoints app=frontend, toPorts TCP "8080"), and no egress. Once the policy takes effect, record in /root/cca-audit/direction.json frontend_to_api, batch_to_api, and api_to_legacy (the HTTP code string of each request; "000" for no response), api_ingress and api_egress (the POLICY column values of the api endpoint), and recorded_at (the UTC time of the moment you record it). You must record this before the egress policy of step 3.
In default mode, default deny is turned on per direction. If a rule selects an endpoint and has an ingress part, only that direction becomes an allowlist. A blocked request gets no response rather than a rejection, so set a short timeout. It can take a few seconds to take effect, so polling is safe.
Apply the new egress policy in audit mode first to see who would be blocked
Turn on audit mode on the api endpoint (cilium-dbg endpoint config <api 번호> PolicyAuditMode=Enabled; the placeholder is the api endpoint number). Then apply the CNP api-egress — endpointSelector app=api, two egress entries: (1) toEndpoints app=ledger, TCP "8080" and (2) toEndpoints k8s:io.kubernetes.pod.namespace: kube-system and k8s:k8s-app: kube-dns, port "53" protocol ANY. After you request api→ledger, api→legacy, and batch→api, read the cca-audit flows with hubble observe inside the agent and record them in /root/cca-audit/audit.json — endpoint_id, audit_mode, api_ingress and api_egress (the POLICY columns), api_to_legacy and batch_to_api (code strings), and flows (the list of compact output lines that contain the AUDITED flows of api→legacy and batch→api).
Audit mode computes the policy verdict but lets the traffic through instead of denying it, and leaves it in the flow as AUDITED. Since it is a per-endpoint option, compare with the result of step 2 that it applies to all directions of that endpoint. The ring buffer is pushed out within a few minutes, so save the flows to a file right after observing, and if you give --since the time just before the requests, only the lines you need are collected. Do not forget that once you lock egress, name resolution is also egress traffic.
When audit is lifted, the same two requests really fail
Turn off audit mode on the api endpoint (PolicyAuditMode=Disabled). Send the same requests again and record in /root/cca-audit/enforce.json audit_mode, api_ingress and api_egress, api_to_ledger, api_to_legacy, and batch_to_api (code strings), and flows (the list of compact lines that contain the DROPPED flows of api→legacy and batch→api after audit was turned off).
Check whether the AUDITED that audit mode showed becomes DROPPED as is, and whether the allowed destination (ledger) still gets through. The end of a flow line is the reason for the verdict. So that they do not get mixed with old DROPPED lines left over from the previous step, read from the time just before you turn audit off.
Both were written as ingress: [{}], yet one opened and one closed
In cca-audit, create two policies — the standard NetworkPolicy np-empty (podSelector app=batch, policyTypes [Ingress], ingress: [{}]) and the CiliumNetworkPolicy cnp-empty (endpointSelector app=frontend, ingress: [{}]). Send requests from the ledger Pod to batch and to frontend, and record in /root/cca-audit/empty.json batch_from_ledger and frontend_from_ledger (code strings), and batch_ingress and frontend_ingress (the ingress POLICY column of each endpoint).
The two APIs interpret an empty rule entry differently. In a Kubernetes NetworkPolicy, an ingress entry with no from means all sources. In Cilium, a rule entry with no source field such as fromEndpoints allows nobody and only turns on default deny. Also check that the enforcement columns look the same for both.
The enforcement column says Enabled, yet nobody is blocked
In cca-audit, create the CNP ledger-observe — endpointSelector app=ledger, enableDefaultDeny: {ingress: false}, and one ingress entry (fromEndpoints app=api, TCP "8080"). Send requests api→ledger and batch→ledger, read the flows, and record in /root/cca-audit/observe.json ledger_ingress (the POLICY column), api_to_ledger and batch_to_ledger (code strings), and flows (the list of compact lines that contain the ledger-side INGRESS policy-verdict lines of the two requests).
A rule with enableDefaultDeny turned off is left out when the endpoint's default mode is decided. So the rule is computed but does not lead to a denial. In the policy-verdict lines, compare how the match type right after policy-verdict: differs between the two requests. This is why you must not judge whether default deny is on from the POLICY column alone.
Two kinds of policy go into one store
Read cilium-dbg policy get -o json in the agent and find the labels io.cilium.k8s.policy.name and io.cilium.k8s.policy.derived-from for each rule of the cca-audit namespace. Record in /root/cca-audit/sources.json derived_from (a dictionary from policy name to derived-from value, five entries) and revision (the revision number in the output).
Cilium also translates a standard NetworkPolicy into its own rule format, puts it in the same store as the CNPs, and attaches where it came from as a label. Look at the key/value pairs in the Labels list of each rule. The revision goes up every time the store changes.
The enforcement mode incident report
In /root/cca-audit/report.txt, write seven lines in 키=값 form (key=value) — api_audit_mode (the current PolicyAuditMode of api), api_policy_columns (the ingress column/egress column of api, in the form A/B for example), api_to_legacy (allowed or denied), np_empty_ingress_rule and cnp_empty_ingress_rule (allow-all or deny-all for each), ledger_ingress_column, and ledger_default_deny_ingress (the enableDefaultDeny.ingress value of ledger-observe, true or false). All values must match the current state.
Do not copy the records from the earlier steps; re-read the current state. The grader also resends the requests to judge. If audit mode has been turned back on, several values in the report will differ.