CNPA — Cloud Native Platform Engineering Associate
We turned on the policy, and the hotfix got blocked
Goal
On a real k3s with Kyverno, you roll out a single platform rule (an owner label and resource requests) to three teams. You first count with audit and find violations with the report, experience what gets blocked when you switch to blocking, then give a time-limited exception to a target that cannot be fixed, and calculate the compliance rate.
Why it matters
A policy engine is a tool that enforces the platform's guardrails as code. But if you turn a rule on as blocking from the start, even emergency patches to workloads that already break the rule get blocked, and incident response stops. That is why governance usually runs in this order: audit → understand the current state from reports → fixes by each team → blocking → narrow, time-limited exceptions. In this lab you confirm how the cluster actually reacts at each step of that order, and what the mechanism is that keeps an exception from turning into a silent permanent allowance.
Steps
- Write three namespaces and four Deployments in
/root/cnpa-pol/tenants.yamland apply them. Attach the labelplatform.labhub.io/tenantto the namespacesteam-pay,team-legacy, andteam-vendorwith the valuespay,legacy, andvendorrespectively. All Deployments have replicas 1, the container nameapp, the imageregistry.k8s.io/pause:3.10, and the metadata and selector labelapp.kubernetes.io/name: <이름>(where the placeholder is the Deployment's name).team-pay/checkouthas both the metadata labelapp.kubernetes.io/owner: payand requests (cpu10m, memory16Mi),team-legacy/report-genhas only requests,team-legacy/batch-synchas only the owner label (legacy), andteam-vendor/vendor-agenthas neither. - Write and apply a ValidatingPolicy (
policies.kyverno.io/v1) namedtenant-deploy-baselinein/root/cnpa-pol/policy.yaml. It usesvalidationActions: [Audit]andevaluation.background.enabled: true, andmatchConstraintscovers CREATE and UPDATE of apps/v1deployments, with anamespaceSelectorthat selects only namespaces that have (Exists) the labelplatform.labhub.io/tenant. Put two validations in this order. ① The metadata labels containapp.kubernetes.io/owner, with the messageDeployment 에 app.kubernetes.io/owner 라벨이 필요합니다(meaning: the Deployment needs the owner label). ② Every container has cpu and memory in its requests, with the message모든 컨테이너에 cpu·memory requests 가 필요합니다(meaning: every container needs cpu and memory requests). - Read the PolicyReports that Kyverno left, and write the Deployments whose
tenant-deploy-baselineresult isfailas an array into/root/cnpa-pol/violations.json. Each element hasnamespace,name,uid(the report's scope.uid), andmessage(the message of that result). - Change the policy's
validationActionsto[Deny]. Then, imitating the legacy team's emergency patch, runkubectl -n team-legacy set image deploy/batch-sync app=registry.k8s.io/pause:3.9and save its output (including standard error) to/root/cnpa-pol/blocked.txt. Next, runkubectl -n team-legacy scale deploy/batch-sync --replicas=2, and writeimage_update_denied,scale_allowed(a boolean), andbatch_sync_image(the current image of batch-sync) into/root/cnpa-pol/deny-effects.json. - Bring the legacy team's two Deployments in line with the rule. Put the metadata label
app.kubernetes.io/owner: legacyonreport-gen, and put requests (cpu50m, memory64Mi) on the containerappofbatch-sync. Then run again theset image ... app=registry.k8s.io/pause:3.9that was blocked in step 4 and let it pass, and wait until the PolicyReport results of the two Deployments change topass. - vendor-agent is a manifest supplied by a vendor, so it cannot be fixed right away. First change the argument
--enablePolicyException=falseof thekyverno-admission-controllerDeployment to--enablePolicyException=trueand wait for the rollout to finish. Then create a PolicyException (policies.kyverno.io/v1) namedvendor-agent-requestsin theteam-vendornamespace in/root/cnpa-pol/exception.yaml.policyRefsis the ValidatingPolicytenant-deploy-baseline,matchConditionsmatches only the object namedvendor-agent, andexpiresAtis 30 days from now (RFC3339, UTC). After applying it, attach the annotationplatform.labhub.io/reviewed=trueto vendor-agent. - Create the PolicyException
old-cron-requestsinteam-legacyin/root/cnpa-pol/expired.yaml. The target is the object namedold-cron, the policy istenant-deploy-baseline, andexpiresAtis one day before now. Then try to create, as a server dry-run inteam-legacy, the Deploymentold-cron(imageregistry.k8s.io/pause:3.10) that has neither an owner label nor requests, and save its output (including standard error) to/root/cnpa-pol/expired.txt. - Count the
tenant-deploy-baselineresults in the PolicyReports of the three tenant namespaces and write them into/root/cnpa-pol/compliance.json. The keys arepass,fail,skip(counts),rate_pct(pass / (pass + fail) × 100, to one decimal place, 100.0 if the denominator is 0), andexceptions(the PolicyExceptions that have not expired as of now, as네임스페이스/이름strings in namespace/name form, as a sorted array).
Notes
- k3s and Kyverno 1.19 are installed inside the VM. Workload images use only
registry.k8s.io/pause. - Checking that the policy is ready:
kubectl get validatingpolicy tenant-deploy-baseline -o jsonpath='{.status.conditionStatus.ready}'. - Reports:
kubectl get policyreport -Aand-o json. The report name is the uid of the target resource. - Common mistake: creating a PolicyException while the exception feature is off.
kubectl applysucceeds and prints only one warning line, and admission keeps rejecting. - To see the description of the expiry field of an exception:
kubectl explain policyexception.spec.expiresAt --api-version=policies.kyverno.io/v1. - Common mistake: creating an exception without
matchConditions. Every Deployment in that namespace then escapes the rule. - Kyverno ValidatingPolicy · Kyverno Policy Exceptions · Kyverno Reporting · Kubernetes CEL · CNCF Platforms White Paper
Three teams before the policy
Write three namespaces and four Deployments in /root/cnpa-pol/tenants.yaml and apply them. Attach the label platform.labhub.io/tenant to the namespaces team-pay, team-legacy, and team-vendor with the values pay, legacy, and vendor respectively. All Deployments have replicas 1, the container name app, the image registry.k8s.io/pause:3.10, and the metadata and selector label app.kubernetes.io/name: <이름> (where the placeholder is the Deployment's name). team-pay/checkout has both the metadata label app.kubernetes.io/owner: pay and requests (cpu 10m, memory 16Mi), team-legacy/report-gen has only requests, team-legacy/batch-sync has only the owner label (legacy), and team-vendor/vendor-agent has neither.
A cluster before you introduce a policy engine already has workloads that break the rule mixed in. Nothing is blocked at this step, so all four must be created. You can put several objects into one file, joined with ---.
Count first, without blocking
Write and apply a ValidatingPolicy (policies.kyverno.io/v1) named tenant-deploy-baseline in /root/cnpa-pol/policy.yaml. It uses validationActions: [Audit] and evaluation.background.enabled: true, and matchConstraints covers CREATE and UPDATE of apps/v1 deployments, with a namespaceSelector that selects only namespaces that have (Exists) the label platform.labhub.io/tenant. Put two validations in this order. ① The metadata labels contain app.kubernetes.io/owner, with the message Deployment 에 app.kubernetes.io/owner 라벨이 필요합니다 (meaning: the Deployment needs the owner label). ② Every container has cpu and memory in its requests, with the message 모든 컨테이너에 cpu·memory requests 가 필요합니다 (meaning: every container needs cpu and memory requests).
A ValidatingPolicy validation is a CEL expression, and object is the requested object. Check whether a key exists with '키' in 맵 (a key-in-map test), and check a field that may be absent first with has(). A condition on every element of a list is .all(c, ...). Audit is the behavior that does not reject and only leaves a result. Wait until the policy's status.conditionStatus.ready becomes true. The grader also accepts the case where the policy was changed to Deny in a later step.
Violations we did not know about showed up in the report
Read the PolicyReports that Kyverno left, and write the Deployments whose tenant-deploy-baseline result is fail as an array into /root/cnpa-pol/violations.json. Each element has namespace, name, uid (the report's scope.uid), and message (the message of that result).
Background evaluation creates a PolicyReport per namespace within a few seconds after the policy is ready. One report corresponds to one resource, with the target in scope and the per-policy results in results. See the summary with kubectl get policyreport -A and read the details with -o json. For a resource that breaks both, the message is that of the validation that failed first.
When the policy was turned on, the emergency patch was blocked
Change the policy's validationActions to [Deny]. Then, imitating the legacy team's emergency patch, run kubectl -n team-legacy set image deploy/batch-sync app=registry.k8s.io/pause:3.9 and save its output (including standard error) to /root/cnpa-pol/blocked.txt. Next, run kubectl -n team-legacy scale deploy/batch-sync --replicas=2, and write image_update_denied, scale_allowed (a boolean), and batch_sync_image (the current image of batch-sync) into /root/cnpa-pol/deny-effects.json.
Deny evaluates not only newly created objects but also update requests for objects that already break the rule. The whole updated object must follow the rule to be accepted. Scale is a request that goes to the scale subresource, not to the Deployment itself, so it differs from the resource the rule selected. Pods that are already running do not go through admission again.
Once the rule was met, the same patch went through
Bring the legacy team's two Deployments in line with the rule. Put the metadata label app.kubernetes.io/owner: legacy on report-gen, and put requests (cpu 50m, memory 64Mi) on the container app of batch-sync. Then run again the set image ... app=registry.k8s.io/pause:3.9 that was blocked in step 4 and let it pass, and wait until the PolicyReport results of the two Deployments change to pass.
A rejected request was not stored, so a change that brings the object in line with the rule must go in first. That change itself produces an object that follows the rule, so it is accepted even under Deny. The shortest way to add requests is kubectl set resources. Reports are re-evaluated when a resource changes. The name is the resource uid, so the same report is updated.
A time-limited exception for an external manifest that cannot be fixed
vendor-agent is a manifest supplied by a vendor, so it cannot be fixed right away. First change the argument --enablePolicyException=false of the kyverno-admission-controller Deployment to --enablePolicyException=true and wait for the rollout to finish. Then create a PolicyException (policies.kyverno.io/v1) named vendor-agent-requests in the team-vendor namespace in /root/cnpa-pol/exception.yaml. policyRefs is the ValidatingPolicy tenant-deploy-baseline, matchConditions matches only the object named vendor-agent, and expiresAt is 30 days from now (RFC3339, UTC). After applying it, attach the annotation platform.labhub.io/reviewed=true to vendor-agent.
In this installation the exception feature is off by default. If you create an exception while it is off, only a warning appears and admission still rejects. You can change that element of the args array with a JSON patch (you can find the index with jq). An exception must select only the targets it needs, narrowly, and have an expiry date, so that the principle does not quietly collapse. You can produce the date with date -u -d '+30 days' +%Y-%m-%dT%H:%M:%SZ. The grader also checks that other names in the same namespace are still rejected.
An expired exception protects nothing
Create the PolicyException old-cron-requests in team-legacy in /root/cnpa-pol/expired.yaml. The target is the object named old-cron, the policy is tenant-deploy-baseline, and expiresAt is one day before now. Then try to create, as a server dry-run in team-legacy, the Deployment old-cron (image registry.k8s.io/pause:3.10) that has neither an owner label nor requests, and save its output (including standard error) to /root/cnpa-pol/expired.txt.
Expiry does not mean the exception object disappears; it means admission no longer uses that exception. The object remains, leaving a record of who allowed what until when. A dry-run does not store anything, but it does go through the admission webhook. old-cron must not actually be created at this step.
Calculate the compliance rate from the reports
Count the tenant-deploy-baseline results in the PolicyReports of the three tenant namespaces and write them into /root/cnpa-pol/compliance.json. The keys are pass, fail, skip (counts), rate_pct (pass / (pass + fail) × 100, to one decimal place, 100.0 if the denominator is 0), and exceptions (the PolicyExceptions that have not expired as of now, as 네임스페이스/이름 strings in namespace/name form, as a sorted array).
A resource skipped because of an exception is reported as skip, not fail. How to treat skip in the compliance rate is up to the organization, but here you exclude it from the denominator. Expiry is judged by comparing expiresAt with the current time. Reports may be updated late, so wait until vendor-agent changes to skip before counting. The grader counts the same reports again.