TT Lab
Get started
Learn Learning paths Courses

CNPA — Cloud Native Platform Engineering Associate

We turned on the policy, and the hotfix got blocked

Continue in TT Lab

Goal

On a real k3s with Kyverno, you roll out a single platform rule (an owner label and resource requests) to three teams. You first count with audit and find violations with the report, experience what gets blocked when you switch to blocking, then give a time-limited exception to a target that cannot be fixed, and calculate the compliance rate.

Why it matters

A policy engine is a tool that enforces the platform's guardrails as code. But if you turn a rule on as blocking from the start, even emergency patches to workloads that already break the rule get blocked, and incident response stops. That is why governance usually runs in this order: audit → understand the current state from reports → fixes by each team → blocking → narrow, time-limited exceptions. In this lab you confirm how the cluster actually reacts at each step of that order, and what the mechanism is that keeps an exception from turning into a silent permanent allowance.

Steps

  1. Write three namespaces and four Deployments in /root/cnpa-pol/tenants.yaml and apply them. Attach the label platform.labhub.io/tenant to the namespaces team-pay, team-legacy, and team-vendor with the values pay, legacy, and vendor respectively. All Deployments have replicas 1, the container name app, the image registry.k8s.io/pause:3.10, and the metadata and selector label app.kubernetes.io/name: <이름> (where the placeholder is the Deployment's name). team-pay/checkout has both the metadata label app.kubernetes.io/owner: pay and requests (cpu 10m, memory 16Mi), team-legacy/report-gen has only requests, team-legacy/batch-sync has only the owner label (legacy), and team-vendor/vendor-agent has neither.
  2. Write and apply a ValidatingPolicy (policies.kyverno.io/v1) named tenant-deploy-baseline in /root/cnpa-pol/policy.yaml. It uses validationActions: [Audit] and evaluation.background.enabled: true, and matchConstraints covers CREATE and UPDATE of apps/v1 deployments, with a namespaceSelector that selects only namespaces that have (Exists) the label platform.labhub.io/tenant. Put two validations in this order. ① The metadata labels contain app.kubernetes.io/owner, with the message Deployment 에 app.kubernetes.io/owner 라벨이 필요합니다 (meaning: the Deployment needs the owner label). ② Every container has cpu and memory in its requests, with the message 모든 컨테이너에 cpu·memory requests 가 필요합니다 (meaning: every container needs cpu and memory requests).
  3. Read the PolicyReports that Kyverno left, and write the Deployments whose tenant-deploy-baseline result is fail as an array into /root/cnpa-pol/violations.json. Each element has namespace, name, uid (the report's scope.uid), and message (the message of that result).
  4. Change the policy's validationActions to [Deny]. Then, imitating the legacy team's emergency patch, run kubectl -n team-legacy set image deploy/batch-sync app=registry.k8s.io/pause:3.9 and save its output (including standard error) to /root/cnpa-pol/blocked.txt. Next, run kubectl -n team-legacy scale deploy/batch-sync --replicas=2, and write image_update_denied, scale_allowed (a boolean), and batch_sync_image (the current image of batch-sync) into /root/cnpa-pol/deny-effects.json.
  5. Bring the legacy team's two Deployments in line with the rule. Put the metadata label app.kubernetes.io/owner: legacy on report-gen, and put requests (cpu 50m, memory 64Mi) on the container app of batch-sync. Then run again the set image ... app=registry.k8s.io/pause:3.9 that was blocked in step 4 and let it pass, and wait until the PolicyReport results of the two Deployments change to pass.
  6. vendor-agent is a manifest supplied by a vendor, so it cannot be fixed right away. First change the argument --enablePolicyException=false of the kyverno-admission-controller Deployment to --enablePolicyException=true and wait for the rollout to finish. Then create a PolicyException (policies.kyverno.io/v1) named vendor-agent-requests in the team-vendor namespace in /root/cnpa-pol/exception.yaml. policyRefs is the ValidatingPolicy tenant-deploy-baseline, matchConditions matches only the object named vendor-agent, and expiresAt is 30 days from now (RFC3339, UTC). After applying it, attach the annotation platform.labhub.io/reviewed=true to vendor-agent.
  7. Create the PolicyException old-cron-requests in team-legacy in /root/cnpa-pol/expired.yaml. The target is the object named old-cron, the policy is tenant-deploy-baseline, and expiresAt is one day before now. Then try to create, as a server dry-run in team-legacy, the Deployment old-cron (image registry.k8s.io/pause:3.10) that has neither an owner label nor requests, and save its output (including standard error) to /root/cnpa-pol/expired.txt.
  8. Count the tenant-deploy-baseline results in the PolicyReports of the three tenant namespaces and write them into /root/cnpa-pol/compliance.json. The keys are pass, fail, skip (counts), rate_pct (pass / (pass + fail) × 100, to one decimal place, 100.0 if the denominator is 0), and exceptions (the PolicyExceptions that have not expired as of now, as 네임스페이스/이름 strings in namespace/name form, as a sorted array).

Notes

Three teams before the policy

Write three namespaces and four Deployments in /root/cnpa-pol/tenants.yaml and apply them. Attach the label platform.labhub.io/tenant to the namespaces team-pay, team-legacy, and team-vendor with the values pay, legacy, and vendor respectively. All Deployments have replicas 1, the container name app, the image registry.k8s.io/pause:3.10, and the metadata and selector label app.kubernetes.io/name: <이름> (where the placeholder is the Deployment's name). team-pay/checkout has both the metadata label app.kubernetes.io/owner: pay and requests (cpu 10m, memory 16Mi), team-legacy/report-gen has only requests, team-legacy/batch-sync has only the owner label (legacy), and team-vendor/vendor-agent has neither.

A cluster before you introduce a policy engine already has workloads that break the rule mixed in. Nothing is blocked at this step, so all four must be created. You can put several objects into one file, joined with ---.

Count first, without blocking

Write and apply a ValidatingPolicy (policies.kyverno.io/v1) named tenant-deploy-baseline in /root/cnpa-pol/policy.yaml. It uses validationActions: [Audit] and evaluation.background.enabled: true, and matchConstraints covers CREATE and UPDATE of apps/v1 deployments, with a namespaceSelector that selects only namespaces that have (Exists) the label platform.labhub.io/tenant. Put two validations in this order. ① The metadata labels contain app.kubernetes.io/owner, with the message Deployment 에 app.kubernetes.io/owner 라벨이 필요합니다 (meaning: the Deployment needs the owner label). ② Every container has cpu and memory in its requests, with the message 모든 컨테이너에 cpu·memory requests 가 필요합니다 (meaning: every container needs cpu and memory requests).

A ValidatingPolicy validation is a CEL expression, and object is the requested object. Check whether a key exists with '키' in 맵 (a key-in-map test), and check a field that may be absent first with has(). A condition on every element of a list is .all(c, ...). Audit is the behavior that does not reject and only leaves a result. Wait until the policy's status.conditionStatus.ready becomes true. The grader also accepts the case where the policy was changed to Deny in a later step.

Violations we did not know about showed up in the report

Read the PolicyReports that Kyverno left, and write the Deployments whose tenant-deploy-baseline result is fail as an array into /root/cnpa-pol/violations.json. Each element has namespace, name, uid (the report's scope.uid), and message (the message of that result).

Background evaluation creates a PolicyReport per namespace within a few seconds after the policy is ready. One report corresponds to one resource, with the target in scope and the per-policy results in results. See the summary with kubectl get policyreport -A and read the details with -o json. For a resource that breaks both, the message is that of the validation that failed first.

When the policy was turned on, the emergency patch was blocked

Change the policy's validationActions to [Deny]. Then, imitating the legacy team's emergency patch, run kubectl -n team-legacy set image deploy/batch-sync app=registry.k8s.io/pause:3.9 and save its output (including standard error) to /root/cnpa-pol/blocked.txt. Next, run kubectl -n team-legacy scale deploy/batch-sync --replicas=2, and write image_update_denied, scale_allowed (a boolean), and batch_sync_image (the current image of batch-sync) into /root/cnpa-pol/deny-effects.json.

Deny evaluates not only newly created objects but also update requests for objects that already break the rule. The whole updated object must follow the rule to be accepted. Scale is a request that goes to the scale subresource, not to the Deployment itself, so it differs from the resource the rule selected. Pods that are already running do not go through admission again.

Once the rule was met, the same patch went through

Bring the legacy team's two Deployments in line with the rule. Put the metadata label app.kubernetes.io/owner: legacy on report-gen, and put requests (cpu 50m, memory 64Mi) on the container app of batch-sync. Then run again the set image ... app=registry.k8s.io/pause:3.9 that was blocked in step 4 and let it pass, and wait until the PolicyReport results of the two Deployments change to pass.

A rejected request was not stored, so a change that brings the object in line with the rule must go in first. That change itself produces an object that follows the rule, so it is accepted even under Deny. The shortest way to add requests is kubectl set resources. Reports are re-evaluated when a resource changes. The name is the resource uid, so the same report is updated.

A time-limited exception for an external manifest that cannot be fixed

vendor-agent is a manifest supplied by a vendor, so it cannot be fixed right away. First change the argument --enablePolicyException=false of the kyverno-admission-controller Deployment to --enablePolicyException=true and wait for the rollout to finish. Then create a PolicyException (policies.kyverno.io/v1) named vendor-agent-requests in the team-vendor namespace in /root/cnpa-pol/exception.yaml. policyRefs is the ValidatingPolicy tenant-deploy-baseline, matchConditions matches only the object named vendor-agent, and expiresAt is 30 days from now (RFC3339, UTC). After applying it, attach the annotation platform.labhub.io/reviewed=true to vendor-agent.

In this installation the exception feature is off by default. If you create an exception while it is off, only a warning appears and admission still rejects. You can change that element of the args array with a JSON patch (you can find the index with jq). An exception must select only the targets it needs, narrowly, and have an expiry date, so that the principle does not quietly collapse. You can produce the date with date -u -d '+30 days' +%Y-%m-%dT%H:%M:%SZ. The grader also checks that other names in the same namespace are still rejected.

An expired exception protects nothing

Create the PolicyException old-cron-requests in team-legacy in /root/cnpa-pol/expired.yaml. The target is the object named old-cron, the policy is tenant-deploy-baseline, and expiresAt is one day before now. Then try to create, as a server dry-run in team-legacy, the Deployment old-cron (image registry.k8s.io/pause:3.10) that has neither an owner label nor requests, and save its output (including standard error) to /root/cnpa-pol/expired.txt.

Expiry does not mean the exception object disappears; it means admission no longer uses that exception. The object remains, leaving a record of who allowed what until when. A dry-run does not store anything, but it does go through the admission webhook. old-cron must not actually be created at this step.

Calculate the compliance rate from the reports

Count the tenant-deploy-baseline results in the PolicyReports of the three tenant namespaces and write them into /root/cnpa-pol/compliance.json. The keys are pass, fail, skip (counts), rate_pct (pass / (pass + fail) × 100, to one decimal place, 100.0 if the denominator is 0), and exceptions (the PolicyExceptions that have not expired as of now, as 네임스페이스/이름 strings in namespace/name form, as a sorted array).

A resource skipped because of an exception is reported as skip, not fail. How to treat skip in the compliance rate is up to the organization, but here you exclude it from the denominator. Expiry is judged by comparing expiresAt with the current time. Reports may be updated late, so wait until vendor-agent changes to skip before counting. The grader counts the same reports again.