KCA — Kyverno Certified Associate
The Policy Remains, but the Webhook Is Gone
In one line
That the policy file remains, that the webhook is registered with the API server, and that the webhook responds are different facts. In a webhook outage experiment, you must observe the three separately.
Why this was needed
A deployment owner says this: "I set the security policy to Fail, but even when I turned the engine off, a violating Pod was created." Can we conclude right away that it is a Kubernetes bug? The call registration may have disappeared during the process of shutting the controller down gracefully. If there is no webhook to call in the first place, you have not tested the failurePolicy that decides how to handle a call failure of that webhook.
This situation was also observed on LabHub's dedicated verification VM on 2026-09-12. The controller was scaled down to replicas=0 and it was confirmed that the Pod was gone, yet the request passed. The list of ValidatingWebhookConfigurations saved at that time had no webhook for that resource. It was not used as evidence that "Fail is broken." Afterward, when only the responses of the same controller process were briefly stopped while leaving the call registration in place, matching requests were rejected after 3 seconds due to a timeout. A graceful shutdown and a response halt are not the same experiment.
How it works
This section deals with phenomena reproduced on Kyverno v1.19.1 and k3s v1.35.8+k3s1. We do not generalize that the webhook disappears in the same way on graceful shutdown in other versions or installation options. When carrying the experiment over, you must recheck the installed version and the actual registration.
First, let us separate three settings.
- Deny in validationActions: the behavior of rejecting the request when the validation result is a violation.
- Fail and Ignore in failurePolicy: decides how to handle errors where the policy cannot be evaluated normally. Here we observe a timeout of the actual webhook call.
- webhookConfiguration.timeoutSeconds: the limit on waiting for the call. It is not a timer that waits for a normal violation response and turns it into an allow.
Ignore does not mean "do not validate." If the server judged a violation normally and sent back a denial response, that denial is maintained. So an Ignore experiment too must first send a violating request to a healthy server and confirm that it is rejected. Without this baseline, if you only look at success during the outage, you cannot tell it apart from the policy never having been applied in the first place.
You can check the policy fields and the automatic generation behavior in the official explanation of Kyverno ValidatingPolicy. This API is policies.kyverno.io/v1. Do not carry over the field locations of old ClusterPolicy examples as they are; cross-check them with the installed CRDs and kubectl explain. The path of running with Kubernetes' own ValidatingAdmissionPolicy is also separate, so do not mix it into the engine process halt experiment.
The namespace is the experiment scope
In the target namespace you apply a rule for Pod creation that requires an environment label, and you keep the control namespace outside that selection scope. The namespaceSelector selects namespace labels, not Pod labels. The environment attached to a Pod and the kubernetes.io/metadata.name that selects a namespace are at different layers.
Reading only the policy declaration is not enough. In the generated webhook, check the actual service path, failurePolicy, timeoutSeconds, namespaceSelector, and the CREATE pods rule. If there are additional objectSelector or matchConditions, requests may bypass validation, so look at those together. Normal behavior outside the scope is a control, not evidence that there was no outage inside the scope.
The official Kubernetes dynamic admission documentation distinguishes request matching from the handling of webhook call failures. It is accurate to explain it as "requests that match that webhook are blocked because of a call failure" rather than "all create requests are blocked."
Two identical requests and one control
Use a new Pod name for each experiment. If you reuse an existing name, AlreadyExists occurs and gets mixed with the policy's denial or the timeout. Besides the API response, also query and keep the actual storage status and the UID. Do not infer storage from a single line of success text.
| Condition | Normal-label request to the target | Violating request to the target | Violating request outside the scope |
|---|---|---|---|
| Webhook healthy, Deny | Stored | Explicit policy denial | Separate scope check |
| Registration kept, responses halted, Fail | Timeout, not stored | Timeout, not stored | Stored |
| Registration kept, responses halted, Ignore | Stored after the limit | Stored after the limit | Stored |
| Webhook recovered | Check the existing Pod separately | Explicitly denied again | Check according to the recovery scope |
The two outage rows of the table are results observed on the dedicated VM above. The actual error text was kept for each request. The error for Fail is context deadline exceeded occurring while calling that webhook, and the error for a normal violation is that policy returning denied the request. Both look like failures in the terminal, but the causes differ. A DNS error or connection refused is yet another type of fault, and it is not counted as a success of this experiment in which only the responses halt.
Why already running Pods remain
Admission is a gate before a request is stored. The target of this rule is CREATE pods, and background evaluation is turned off. It does not mean a controller was built that automatically deletes Pods that were allowed earlier and are running. So you separately observe the same UID and Running of the existing Pod. You must not count a Pod newly created under the same name as having "survived."
What it looks like in the field
Suppose the security engine's responses got slow right before a Friday deployment. Changing to Ignore may improve deployment availability, but it creates the risk of accepting unevaluated changes. Keeping Fail may stop those requests, but it closes the path by which things are stored without validation. Either way, you have to decide the purpose, scope, and tolerable outage time first. It is not something to decide with a single sentence like "Ignore is better most of the time."
The fail-open recommendation in the Kubernetes webhook design guidelines is in the context of looking at mutating webhooks together with final-state validation. Do not read it as an unconditional recommendation of Ignore for all security validation webhooks. You have to design together high availability, narrowing the call scope, sufficient resources, and outage alerts.
Restoring replicas and restoring the policy do not happen at the same time
Scaling the controller back up to 1 does not mean validation returns immediately. You have to check separately the new Pod starting, the leader taking over, the webhook registration, and the actual policy response. In this section's first student-path test, replicas were restored at 23:49:53 UTC, but the new process acquired the leader Lease at 23:50:27. The 20-second registration wait had ended before that. This was not a policy-violation verdict of Fail but a failure of the lab runner's readiness wait.
Kyverno v1.19.1's leader election implementation uses a Lease and is configured not to release it immediately on shutdown. The actual wait time varies with the installation and load. Do not use the observed 34 seconds as a constant for every environment; check the logs and the registration state. Even if the lab fails while waiting during recovery, it preserves the policy, registration, and request results observed during the scale-down. Before pressing the retry button, you must first distinguish what has been restored and what has not yet been confirmed.
Even the tool that creates the outage needs a recovery design
The pause is applied only to the single Kyverno container inside the student's dedicated VM. It checks the container ID, the Pod UID, and the PID start time, and binds the signal to a Linux pidfd so that reuse of a numeric PID does not touch another process. A separate watcher automatically resumes within 20 seconds, and the stop and the recovery are not split across clicks of the next step. The watcher tries to resume the target even on paths where the experiment runner is forcibly terminated or disappears before its readiness response.
This mechanism does not replace graceful shutdown procedures, high availability tests, or production outage drills. Situations that kill the watcher itself or shut down the VM are a separate failure area. The grader does not recreate the outage; it reads the preserved results. File-based grading is also not a security certification that prevents a malicious root from forging all evidence. It is used in learning to catch cases where the observation grounds are empty or contradict each other.
What you will do in the next lab
The first lab deals with graceful scale-down and the absence of registration. You record the target and control scopes, write a ValidatingPolicy yourself, and then observe the normal baseline and the registration state during a graceful scale-down. You distinguish the times of the recovery request, the registration check, and the real denial response, and confirm with a server dry-run that validation and storage are different stages.
The second lab starts independently on a new VM. The files from the first lab are not needed. After recreating the policy and the normal baseline, you compare Fail and Ignore with a short response halt that keeps the call registration. In both labs, you write the report by separating policy existence, registration, response, and survival of the existing Pod. Raw data must come before conclusions.
Check where the dry-run was done too
According to Kubernetes' official explanation of server dry-run, a server dry-run runs the validation path for that request without storing the change. A client dry-run, on the other hand, builds the output locally and so does not prove whether the server webhook will reject. Checking the file syntax, the server accepting the request, and the object actually being stored are different verifications.
In this lab, you server dry-run a normal request and a violating request with unique names and separately query a Pod with the same name. For the normal response, the kind, name, namespace, and labels must be right and there must be no stored result. The violating request must be an explicit denial with a policy marker. Do not record that it was stored just because a Pod object shows up in a success output, and do not record AlreadyExists as a policy denial. The webhook that receives the request must also support dry-run and side-effect handling, so when carrying this to another installation, you must also check the actual webhook's sideEffects setting.