TT Lab
Get started
Learn Learning paths Courses

KCA — Kyverno Certified Associate

Policy Is Promoted, Not Deployed

Continue in TT Lab

In one line

The core procedure of policy operations is one thing. Observe with Audit for a few days, confirm that the list of what gets caught matches expectations, and then raise it to Enforce. And leave that judgment not to people's memory but to PolicyReports and a CI gate.

Why this was needed

The failure pattern of applying policies as Enforce from the start on a development cluster always looks the same. If you turn on enforcement mode while the signing pipeline or the label rules have not yet been rolled out to every team, deployments are blocked wholesale, and in the end it is not the policy but the person who made the policy who becomes the bottleneck. What happens next is worse. Because it is urgent, they punch wide exceptions, and those exceptions become permanent.

Conversely, if you leave everything as Audit, nobody looks at the reports. That is why a promotion procedure is needed. It means deciding in advance what to look at, which conditions must be met, and which rule to raise.

How it works

The tools of observation are PolicyReport and ClusterPolicyReport. A namespace-level report contains a results array and a summary (the counts of pass/fail/warn/error/skip). The habit of pulling out only the failures is important, and you have to look at the policy name and the message together to judge which rule was caught because of what.

There are two ways to apply differently. Per-rule failureAction decides "which rule," and validationFailureActionOverrides decides "in which namespace." The typical configuration keeps development namespaces on Audit and raises only production to Enforce, and by combining the two axes you can run it differently both per rule and per namespace.

You manage legitimate exceptions with PolicyException. You point to the policy name and the rule name and narrow the target with match. The key here is to write it narrowly. Instead of excluding a whole namespace, you must write down even the name pattern, and once the exception list goes beyond half of the targets that actually run, that policy gives not security but the illusion that there is security.

You should also know the three policy-level flags. background turns on and off the behavior of scanning existing resources to create reports, and the default is true. admission is whether to apply rules at the admission stage, the default is true, and if you set it to false it becomes a background-only policy. applyRules is how many rules to apply to a matched resource; with One it stops at the first match, and All is the default. If you have listed the rules in order and are looking for why the later rules do not run, look at this field first.

The CI gate is made by kyverno apply. If you give it the policy files and the manifests to check with --resource, it evaluates them against each other locally, and if there is a failure or an error the exit code is 1, so it hooks straight into a CI job. The accidents this one habit prevents are large: writing a match block wrongly so that you create a policy that matches nothing and mistakenly believe it passed. A policy that matches nothing quietly just lets everything pass in the cluster, so it cannot be told apart by human eyes from a policy that is running well.

And there is one principle to repeat. A policy that matches all resources with a wildcard puts a latency tax on every request in the cluster. The default failurePolicy is Fail, so the moment it exceeds the timeout, the request is rejected. Narrowing the match to the necessary kinds and namespaces is not performance tuning but availability work.

Finally, when not to use it. If what you want to check is the value of one field and nothing more is needed, Kubernetes' built-in ValidatingAdmissionPolicy and CEL are enough. Operating one fewer component means one fewer upgrade target, one fewer place to suspect during an outage, and one fewer thing to worry about for webhook certificate renewal. Going further, for a standardized check such as the Pod security level, it is finished with two namespace label lines (Pod Security Admission), without a policy engine. What justifies Kyverno are the things the built-in features cannot do, such as generate, mutate, and image verification.

What it looks like in the field

The rule the author set up while handling image scan reports applies just the same to policy operations. It is that every exception must have an expiry date. In the scan exception file, you write the reason and the expiry date together, and once the expiry date passes, the scanner makes it fail again. This was the only way to stop exceptions from quietly becoming permanent. PolicyException has no expiry field, so to get the same effect it has to be managed by people. It is better to write the expiry date and reason in the annotations of the exception object and put in a procedure that sweeps through them periodically.

One more thing: the sense of priority learned from scanning is valid here too. Even if there are 1,247 findings in a report, what you act on today is usually a single digit. Keep only those with a fix available, and look first at those with confirmed real exploitation. Policy reports are the same. Even if there are hundreds of fails, it is not that all of them are to be fixed today; group them by rule and start by looking at which rule produces the most failures. That one rule is usually either a wrongly written rule or a rule the organization is not yet ready for.

The author's homelab is a configuration with Cilium eBPF, ArgoCD, Harbor, and Gitea on 7 nodes, and even at this scale, if you turn policies on fully as Enforce, system components such as KubeVirt and the GPU Operator get caught first. That is why excluding system namespaces with exclude is not an option but the default.

What you will do in the next lab

In /root/kca-ops/, you write a policy file containing per-namespace differential application, webhookConfiguration, and the policy flags, a PolicyException file, and a promotion checklist as an executable shell script. After that, you actually create a namespace that enforces the Pod security level with only PSA labels, and check by hand how far you can get without a policy engine.