TT Lab
Get started
Learn Learning paths Courses

KCA — Kyverno Certified Associate

The Policy Remains, but an Invalid Pod Was Created

Continue in TT Lab

Goal

Even if the policy object remains, if the call registration is removed, you have not tested Fail's handling of call failures. After a graceful scale-down, you compare the recovery times of the registration and the response and the object UIDs, and distinguish validation from storage with a server dry-run.

Why it matters

If you lump a policy's explicit denial and a webhook call error together under the single word "failure," you pick the wrong response. You must observe the policy, the registration, the response, and the already running Pod separately. Each of the two labs starts on a new VM and needs no materials from the other lab. Do not create outages outside the student's dedicated VM.

The graceful scale-down uses a 55-second recovery watcher, and the response halt uses a pidfd and a 20-second automatic resume timer. This lab is expected to take 55 minutes. The default session is 60 minutes, and if you need more before it expires, extend it with +time. The maximum is 180 minutes, and when it ends, the VM and files disappear. Download the materials you need first.

All student files are under /root/kca-webhook. The raw data of each act N is saved in evidence/NN.json, and its contents are read from facts. References such as 02.json in the descriptions are this evidence path. The canonical JSON hash is digest(read(path)) of /opt/fixtures/kca_webhook_common.py, and it differs from the file byte hash of sha256sum. In Python, add /opt/fixtures to sys.path and use it.

Steps

  1. With inspect, check this VM's vm.node_uid and namespaces. In scope.json, record target=kca-webhook-target, control=kca-webhook-control, node_uid, and namespaces. For namespaces, use the actual names as keys and the UIDs as values. Preserve this VM's scope with act 1.
  2. In policy.json, write a ValidatingPolicy of policies.kyverno.io/v1. The name is kca-webhook-label, validationActions=[Deny], failurePolicy=Fail, webhookConfiguration.timeoutSeconds=3, and evaluation.background.enabled=false. matchConstraints.namespaceSelector.matchLabels is kubernetes.io/metadata.name=kca-webhook-target, and resourceRules is apiGroups=[an empty string], apiVersions=[v1], operations=[CREATE], and resources=[pods]. The expression of validations is "'environment' in object.metadata.?labels.orValue({})", and the message is KCA_ENVIRONMENT_REQUIRED. Write it without unnecessary fields and check the actual registration with act 2.
  3. In registration.json, record service=kyverno-svc, namespace=kyverno, path=/vpol/kca-webhook-label, selector_target=kca-webhook-target, selector_source=namespace-label, operation=CREATE, resource=pods, failure_policy=Fail, timeout_seconds=3, and policy_uid=facts.policy.metadata.uid of 02.json. With act 3, observe the registration and the stored UID of an unlabeled request outside the scope.
  4. Run act 4. In evidence/04.json, facts.good must be the stored UID of a Pod with the label, bad must be an explicit denial with no storage for a missing label, and existing must be Running with the same UID. Read the original error text and the storage status together.
  5. In shutdown-plan.json, record action=observe-graceful-shutdown, the controller.metadata.uid from inspect as deployment_uid, replicas_before=1, replicas_during=0, replicas_after=1, and restore_deadline_sec=55. act 5 gracefully scales down and restores the dedicated VM's controller. In evidence/05.json, distinguish the same policy UID, the absent registration, the stored violation, and the denial after recovery and the new controller Pod.
  6. In recovery.json, move facts.replicas_restored_at, facts.recovery.registered_at, and facts.recovery.response_at of 05.json under keys with the same names. Write replicas_mean=desired-count-not-policy-readiness, controller_pod=replaced, policy=same-uid, and shutdown_evidence_sha256=the canonical JSON hash of 05.json. act 6 rechecks the denial of a new violating request and the survival of the baseline Pod.
  7. In dryrun.json, record mode=server, admission=evaluated, good=accepted-not-stored, bad=explicit-deny, existing=same-uid-running, and evidence_sha256=the canonical JSON hash of 06.json. act 7 performs a server dry-run with new names. Compare the normal Pod response and the absence of storage in evidence/07.json with the explicit denial of the violating request.
  8. In report.json, record shutdown=registration-removed-not-fail-bypass, policy=same-uid, controller_pod=replaced, existing=same-uid-running, and dryrun=admission-without-storage. For shutdown_evidence_sha256 and dryrun_evidence_sha256, put the canonical JSON hashes of 05.json and 07.json. After act 8, run the full grading again.

Notes

The commands are python3 /opt/fixtures/kca_webhook_lab.py inspect, act 1 through act 8, and grade 1 through grade 8. grade reads the input and the preserved real observations and does not recreate the outage. A completed act preserves its materials and resources. Even partial input is not overwritten automatically. The results of an interrupted run are uncertain, so download the raw failure data and reproduce it in a new lab. The policy target is CREATE pods. Do not generalize to all APIs, all installed versions, or high availability. Kubernetes server dry-run

The VM and the target and control scopes

With inspect, check this VM's vm.node_uid and namespaces. In scope.json, record target=kca-webhook-target, control=kca-webhook-control, node_uid, and namespaces. For namespaces, use the actual names as keys and the UIDs as values. Preserve this VM's scope with act 1.

Look at the vm.node_uid and namespaces of inspect. A name and a UID are different.

Write the Deny and Fail policy yourself

In policy.json, write a ValidatingPolicy of policies.kyverno.io/v1. The name is kca-webhook-label, validationActions=[Deny], failurePolicy=Fail, webhookConfiguration.timeoutSeconds=3, and evaluation.background.enabled=false. matchConstraints.namespaceSelector.matchLabels is kubernetes.io/metadata.name=kca-webhook-target, and resourceRules is apiGroups=[an empty string], apiVersions=[v1], operations=[CREATE], and resources=[pods]. The expression of validations is "'environment' in object.metadata.?labels.orValue({})", and the message is KCA_ENVIRONMENT_REQUIRED. Write it without unnecessary fields and check the actual registration with act 2.

The namespaceSelector selects namespace labels. Distinguish the validation action Deny from the call-failure handling Fail.

Confirm the scope of the call registration with requests

In registration.json, record service=kyverno-svc, namespace=kyverno, path=/vpol/kca-webhook-label, selector_target=kca-webhook-target, selector_source=namespace-label, operation=CREATE, resource=pods, failure_policy=Fail, timeout_seconds=3, and policy_uid=facts.policy.metadata.uid of 02.json. With act 3, observe the registration and the stored UID of an unlabeled request outside the scope.

Besides the policy declaration, also cross-check the actual service path, namespaceSelector, and rules in 02.json. A request outside the scope does not mean Fail was bypassed.

Normal allow and deny, and the existing Pod baseline

Run act 4. In evidence/04.json, facts.good must be the stored UID of a Pod with the label, bad must be an explicit denial with no storage for a missing label, and existing must be Running with the same UID. Read the original error text and the storage status together.

AlreadyExists is not a policy denial. In a later step you compare the existing Pod's UID with this baseline.

The different lifetimes of the policy and the registration after a graceful shutdown

In shutdown-plan.json, record action=observe-graceful-shutdown, the controller.metadata.uid from inspect as deployment_uid, replicas_before=1, replicas_during=0, replicas_after=1, and restore_deadline_sec=55. act 5 gracefully scales down and restores the dedicated VM's controller. In evidence/05.json, distinguish the same policy UID, the absent registration, the stored violation, and the denial after recovery and the new controller Pod.

Even if the policy remains, the call registration can disappear. Do not assume that the policy has also been restored right after replicas=1.

Verify the recovery times and the real denial separately

In recovery.json, move facts.replicas_restored_at, facts.recovery.registered_at, and facts.recovery.response_at of 05.json under keys with the same names. Write replicas_mean=desired-count-not-policy-readiness, controller_pod=replaced, policy=same-uid, and shutdown_evidence_sha256=the canonical JSON hash of 05.json. act 6 rechecks the denial of a new violating request and the survival of the baseline Pod.

The times are not the exact moments the events occurred but the observation times the runner confirmed. Do not lump the registration check and the real denial response into the same state.

Distinguish server validation success from storage success

In dryrun.json, record mode=server, admission=evaluated, good=accepted-not-stored, bad=explicit-deny, existing=same-uid-running, and evidence_sha256=the canonical JSON hash of 06.json. act 7 performs a server dry-run with new names. Compare the normal Pod response and the absence of storage in evidence/07.json with the explicit denial of the violating request.

A client dry-run does not check the server webhook. Do not interpret the presence of an object in the normal response of a server dry-run as it actually having been stored.

An incident report based on the lifetimes of different objects

In report.json, record shutdown=registration-removed-not-fail-bypass, policy=same-uid, controller_pod=replaced, existing=same-uid-running, and dryrun=admission-without-storage. For shutdown_evidence_sha256 and dryrun_evidence_sha256, put the canonical JSON hashes of 05.json and 07.json. After act 8, run the full grading again.

Distinguish the recreated controller Pod from the baseline Pod that stays alive. Matching only the report hashes will not pass if the real raw data from the earlier steps is missing.