KCA — Kyverno Certified Associate
Webhook Failure Lab: Fail, Ignore, and Recovery
Goal
You halt only the responses of a webhook whose registration remains and compare the difference between Fail and Ignore with real requests. You check separately the explicit denial of a healthy server, the timeout and storage during the outage, the control outside the scope, and the recovery.
Why it matters
If you lump a policy's explicit denial and a webhook call error together under the single word "failure," you pick the wrong response. You must observe the policy, the registration, the response, and the already running Pod separately. Each of the two labs starts on a new VM and needs no materials from the other lab. Do not create outages outside the student's dedicated VM.
The graceful scale-down uses a 55-second recovery watcher, and the response halt uses a pidfd and a 20-second automatic resume timer. This lab is expected to take 55 minutes. The default session is 60 minutes, and if you need more before it expires, extend it with +time. The maximum is 180 minutes, and when it ends, the VM and files disappear. Download the materials you need first.
All student files are under /root/kca-webhook. The raw data of each act N is saved in evidence/NN.json, and its contents are read from facts. References such as 02.json in the descriptions are this evidence path. The canonical JSON hash is digest(read(path)) of /opt/fixtures/kca_webhook_common.py, and it differs from the file byte hash of sha256sum. In Python, add /opt/fixtures to sys.path and use it.
Steps
- With inspect, check this VM's vm.node_uid and namespaces. In scope.json, record target=kca-webhook-target, control=kca-webhook-control, node_uid, and namespaces. For namespaces, use the actual names as keys and the UIDs as values. Preserve this VM's scope with act 1.
- In policy.json, write a ValidatingPolicy of policies.kyverno.io/v1. The name is kca-webhook-label, validationActions=[Deny], failurePolicy=Fail, webhookConfiguration.timeoutSeconds=3, and evaluation.background.enabled=false. matchConstraints.namespaceSelector.matchLabels is kubernetes.io/metadata.name=kca-webhook-target, and resourceRules is apiGroups=[an empty string], apiVersions=[v1], operations=[CREATE], and resources=[pods]. The expression of validations is "'environment' in object.metadata.?labels.orValue({})", and the message is KCA_ENVIRONMENT_REQUIRED. Write it without unnecessary fields and check the actual registration with act 2.
- With act 3, observe the storage, UID, and Running of a Pod with the label and the explicit denial and non-storage of a Pod without the label. Compare facts.good, bad, and existing in evidence/03.json. Do not bring over files or a VM from the previous lab.
- In freeze-plan.json, record action=freeze-one-controller-process, the real node_uid and deployment_uid, resume_after_sec=20, signal_identity=pidfd, preserve_webhook=true, and evidence_sha256=the canonical JSON hash of 03.json. act 4 only investigates the CRI container, the Pod UID, and the PID start time, and does not stop anything yet.
- Run act 5. In evidence/05.json, compare the preserved registration, the timeouts and non-storage of normal and violating requests within the pause window, the storage of the control request outside the scope, and the denial of the violation after recovery and the existing Pod UID. Distinguish an explicit denial from context deadline exceeded.
- Based on policy.json, write policy-ignore.json, changing only spec.failurePolicy to Ignore. act 6 first confirms the normal violation denial on the same policy UID and then briefly halts only the responses. In evidence/06.json, compare the storage of normal and violating requests during the outage, the denial after recovery, and the survival of the existing Pod.
- In recovery.json, put mode=Ignore, healthy_bad=explicit-deny, existing=same-uid-running, policy_uid=facts.policy.metadata.uid of 02.json, and ignore_evidence_sha256=the canonical JSON hash of 06.json. act 7 rejects a violating request with a new name and confirms the same UID and Running of the baseline Pod from step 03.
- In report.json, record fail=matching-requests-timeout, ignore=unvalidated-request-stored, healthy_ignore=explicit-deny, control=outside-selector, and existing=same-uid-running. fail_evidence_sha256, ignore_evidence_sha256, and recovery_evidence_sha256 are the canonical JSON hashes of 05, 06, and 07.json respectively. After act 8, run the full grading again.
Notes
The commands are python3 /opt/fixtures/kca_webhook_lab.py inspect, act 1 through act 8, and grade 1 through grade 8. grade reads the input and the preserved real observations and does not recreate the outage. A completed act preserves its materials and resources. Even partial input is not overwritten automatically. The results of an interrupted run are uncertain, so download the raw failure data and reproduce it in a new lab. The policy target is CREATE pods. Do not generalize to all APIs, all installed versions, or high availability. Kubernetes server dry-run
An independent experiment scope on a new VM
With inspect, check this VM's vm.node_uid and namespaces. In scope.json, record target=kca-webhook-target, control=kca-webhook-control, node_uid, and namespaces. For namespaces, use the actual names as keys and the UIDs as values. Preserve this VM's scope with act 1.
Look at the vm.node_uid and namespaces of inspect. A name and a UID are different.
Write the Deny and Fail policy yourself
In policy.json, write a ValidatingPolicy of policies.kyverno.io/v1. The name is kca-webhook-label, validationActions=[Deny], failurePolicy=Fail, webhookConfiguration.timeoutSeconds=3, and evaluation.background.enabled=false. matchConstraints.namespaceSelector.matchLabels is kubernetes.io/metadata.name=kca-webhook-target, and resourceRules is apiGroups=[an empty string], apiVersions=[v1], operations=[CREATE], and resources=[pods]. The expression of validations is "'environment' in object.metadata.?labels.orValue({})", and the message is KCA_ENVIRONMENT_REQUIRED. Write it without unnecessary fields and check the actual registration with act 2.
The namespaceSelector selects namespace labels. Distinguish the validation action Deny from the call-failure handling Fail.
The allow and deny baseline of a healthy server
With act 3, observe the storage, UID, and Running of a Pod with the label and the explicit denial and non-storage of a Pod without the label. Compare facts.good, bad, and existing in evidence/03.json. Do not bring over files or a VM from the previous lab.
This VM's policy must work from the start for you to interpret the results during the later outage.
The CRI, Pod, and PID identities and the automatic resume plan
In freeze-plan.json, record action=freeze-one-controller-process, the real node_uid and deployment_uid, resume_after_sec=20, signal_identity=pidfd, preserve_webhook=true, and evidence_sha256=the canonical JSON hash of 03.json. act 4 only investigates the CRI container, the Pod UID, and the PID start time, and does not stop anything yet.
If you trust only the PID number, you may stop a different process that reused it. This lab does not do a scale-down.
Keep the registration and observe the Fail timeout
Run act 5. In evidence/05.json, compare the preserved registration, the timeouts and non-storage of normal and violating requests within the pause window, the storage of the control request outside the scope, and the denial of the violation after recovery and the existing Pod UID. Distinguish an explicit denial from context deadline exceeded.
If even a normal labeled request is blocked, it may be because the call got no response, not because the validation expression was False.
Ignore's normal denial and storage during the outage
Based on policy.json, write policy-ignore.json, changing only spec.failurePolicy to Ignore. act 6 first confirms the normal violation denial on the same policy UID and then briefly halts only the responses. In evidence/06.json, compare the storage of normal and violating requests during the outage, the denial after recovery, and the survival of the existing Pod.
Ignore is not disabling the policy. You must first confirm the explicit denial in the healthy state to interpret the storage during the outage.
Recheck the recovery with a new request and the earlier UID
In recovery.json, put mode=Ignore, healthy_bad=explicit-deny, existing=same-uid-running, policy_uid=facts.policy.metadata.uid of 02.json, and ignore_evidence_sha256=the canonical JSON hash of 06.json. act 7 rejects a violating request with a new name and confirms the same UID and Running of the baseline Pod from step 03.
Do not conclude that even the policy response has recovered merely from the fact that a resume signal was sent to the process.
A report linking the two outage results to their evidence
In report.json, record fail=matching-requests-timeout, ignore=unvalidated-request-stored, healthy_ignore=explicit-deny, control=outside-selector, and existing=same-uid-running. fail_evidence_sha256, ignore_evidence_sha256, and recovery_evidence_sha256 are the canonical JSON hashes of 05, 06, and 07.json respectively. After act 8, run the full grading again.
Explain the difference between Fail and Ignore as call-failure handling, not as the normal violation verdict. Complete it with only this VM's observations, without the materials of lab A.