validate — Patterns, deny, and Match Scope
In one sentence
The quality of a validating policy is determined not by the condition expression but by the match scope and the failure message.
Why it was needed
People writing policies for the first time start by worrying about how to express the condition. But what actually causes incidents is not the condition. If the condition is wrong, testing catches it, but if the match scope is wrong, nothing happens, so nobody knows. A policy applied to the whole cluster blocks the system Pods in kube-system and interferes with node boot, or conversely, by missing a single namespace, leaves only that place outside the rules for months on end.
So the order for writing a validating policy starts not with the condition but with the scope. Which kinds (kinds), which namespaces (namespaces), which labels (selector) to apply it to. And there is a question that must always go with it. What to leave out (exclude). Cluster components, the policy engine itself, and legacy namespaces that have not yet been cleaned up are better excluded from the start.
How it works
A validation rule has several ways to express a condition, and expressive power is inversely proportional to readability.
| Method | Where it is used | Character |
|---|---|---|
validate.pattern |
A field must exist / a value must have a certain shape | Easy to read because it is written in the same shape as the object |
validate.deny.conditions |
Conditions that a pattern cannot express, such as list comparison and string matching | Explicit denial. any is true if even one is true, all if all are true |
validate.cel |
Conditions that need computation | High expressiveness, but the whole team must be able to read it |
validate.foreach |
For each item of a list, such as containers | Judges per item and also produces the message per item |
Patterns come with operators. ?* means "any non-empty value," * means "any value (may be absent)," X|Y is a choice, !X is negation, and numeric comparisons such as >=256Mi also work. If you write team: "?*" under metadata.labels, it becomes "there must be a team label and it must not be empty."
An anchor is a prefix that changes the meaning of a pattern. The conditional anchor () means "check what is below only if this value matches," the equality anchor =() means "if this key exists, the value must be equal," and the negation anchor X() means "this key must not exist." The mistake of confusing the negation anchor with a value comparison is especially frequent. X(privileged) does not mean privileged must be false; it means that key must not exist.
preconditions and deny.conditions look similar in shape and are often confused. The difference is not the result but whether it is evaluated. If preconditions is false, the rule is not executed at all and is skipped. If deny.conditions is true, the rule was executed and its result is a denial. What remains in the report differs too — the former leaves no trace and the latter remains as fail. So something like "this rule applies only to CREATE requests" is written with preconditions, and "this image is not allowed" is written with deny.
Finally, the enforcement level. Previously, the policy-level spec.validationFailureAction alone governed all the rules in a policy. Now it is moving to the rule-level validate.failureAction, and the values are two: Enforce (blocks violating requests) and Audit (lets them pass but records them in the report). With rule-level control, it became possible within one policy for some rules to be enforced already and some to only observe. It means you can quietly layer a new rule on and watch it for a few days without splitting the policy.
What you see in the field
First, the message is half of the policy. The person who is denied cannot read the policy YAML. All they see is the single line kubectl apply spat out. A message like "Validation error" creates inquiries, while a message like "A Pod needs a team label — please write the owning team" lets people fix it themselves. Message quality is the policy's operating cost.
Second, a policy that matches nothing. It is common for a policy to be entirely empty because of a typo in the match block or a wrong namespace. There is only one way to catch this. Try both a resource that should pass and a resource that should be blocked. If you check only passes, you cannot tell "a policy that matches nothing" from "a working policy."
Third, sometimes the exception is needed before the policy. When you apply a new rule to workloads that are already running, some will certainly get caught. If you wait until you fix them all, the policy is never introduced. It is better to leave the exception as a document, narrowed down to the policy name, rule name, and even resource name with PolicyException, and to manage that list shrinking. Applying an exception to the whole policy is the same as turning the policy off.
What you will do in the next lab
You write a ClusterPolicy that requires mandatory labels in /root/policy/validate/require-labels.yaml, stating the match and exclude explicitly. Then you feed in a passing Pod and a failing Pod and check that the decisions diverge. Next you write a deny rule that restricts the registry, along with preconditions, and finally you build a PolicyReport that checks several resources at once and a narrowly scoped PolicyException. In this environment the cluster does not block requests, so you confirm the decisions by running kyverno apply locally.