TT Lab
Get started
Learn Learning paths Courses

KCA — Kyverno Certified Associate

Where a Request Gets Caught, and Where Kyverno Stands

Continue in TT Lab

In one line

A single kubectl apply passes through six stops inside the API server: authentication → authorization → mutating admission → schema validation → validating admission → storage in etcd. Kyverno's mutate sits at the third stop, and validate at the fifth. If you do not know this order, you get caught by your own policies.

Why this was needed

There were originally two ways to enforce rules on a cluster. One is to use RBAC to block "who can do what," and the other is for a person to look at the manifest in a review. But most real-world requirements fall in between. You have to allow a developer to create a Deployment itself, but you have to block that Deployment from coming up without resource limits, using the latest tag, or running as privileged. RBAC looks only at verbs and resource kinds and not at the contents, so it cannot express this requirement. That is why the API server opened, after authorization, an extension point that "looks at the contents of the request and makes a judgment," and that is the admission webhook.

There is a reason the stops have an order, too. Authentication and authorization come first so that heavy processing is not done for requests whose identity is not even known. Mutating comes before validating because a flow of "fix it up, then check it" is natural. If a webhook that fills in default values came after the checks, nothing would get through.

How it works

There is a practical conclusion that follows directly from this order. The object a validate rule sees is an object that mutate rules have already touched. If you apply a mutate policy that injects a sidecar together with a validate policy that requires resource limits on every container, the injected sidecar also becomes a target of the limits check. If you did not put limits in the injected spec, the deployment is blocked, and the log then shows the name of a container the user never wrote. Putting resource limits into the sidecar injection policy is not a matter of taste but a requirement.

Kyverno consists of three controllers, each doing a different job.

Controller When it runs What it does
admission In real time when a request comes in Evaluates mutate and validate in the webhook and responds
background Periodically, and when an UpdateRequest is created Scans existing resources and performs the actual creation and synchronization of generate rules
reports When evaluation results accumulate Creates and aggregates PolicyReport / ClusterPolicyReport objects

The three are separated because their characteristics differ. admission has to answer within milliseconds, background can take several minutes, and reports do a lot of writing. If you put them in one process, deployments get delayed along with it when report aggregation falls behind. This is also why a generate rule does not create resources directly in admission but leaves an UpdateRequest that background then processes.

Finally, failurePolicy. The default is Fail, which means that if the webhook cannot respond in time, the request is rejected. From a security standpoint it is the right default, but from an availability standpoint what it means is heavy. If you apply a policy that matches all resources with a wildcard, every request goes through the webhook and puts a latency tax on the whole cluster, and the moment it fails to answer within the default timeout, requests are rejected. That is, when Kyverno gets slow, it is not that the cluster gets slow but that the cluster stops. Narrowing the match to the necessary kinds and namespaces is not performance tuning but availability work. The default webhook timeout is 10 seconds, and only values from 1 to 30 seconds are allowed.

What it looks like in the field

The author's homelab is a 7-node cluster made of 3 control plane machines and 4 GPU workers, and its CNI is Cilium 1.20 eBPF, running without kube-proxy. The lesson confirmed here repeatedly is that "a state being Ready and something actually working are different propositions." KubeVirt had all components AllComponentsReady but the VM would not come up, and the cause was a missing volume mount in the virt-launcher Pod spec. With the Gateway API, when the CRDs were left at v1.2, the controller refused to start because tlsroutes and referencegrants were not v1, and it had to be raised to v1.6.1.

Admission policies have exactly the same trap. That a policy object exists and its status looks normal does not mean that the policy is actually being applied to requests. If the webhook configuration is not registered or the match block selects nothing, the policy shows up in the list looking fine while doing nothing. And a policy that matches nothing cannot be told apart by eye from a policy that is running well. This is the diagnostic principle this course will emphasize repeatedly, and the local verification habit you learn later is its answer.

What to check in the next quiz

This module covers concepts only. From the next module you start writing ClusterPolicy yourself, and you also set up the same rule with Kubernetes' built-in ValidatingAdmissionPolicy and compare the two approaches side by side.