TT Lab
Get started
Learn Learning paths Courses

Policy as Code

ValidatingAdmissionPolicy: blocking at the API server without an engine

Continue in TT Lab

In one sentence

ValidatingAdmissionPolicy (VAP, below) is a mechanism that checks requests with CEL expressions inside the API server, with no Pod, no certificates, and no network round trip, and the heart of its design is that it deliberately splits the rule (the policy) and its scope of application (the binding) into two objects.

Why it was needed

A webhook-based policy engine is powerful, but it is expensive. You have to run the engine Pods, issue TLS certificates and stick the CA bundle into the webhook configuration, secure availability with replicas and a PodDisruptionBudget, and from the moment you turn on failurePolicy: Fail, that engine's availability becomes the cluster's availability. There was a long-standing question of whether it is right to bear all of this just to apply a one-line rule like "reject if the image tag is latest."

Kubernetes answered this question with "let the API server judge simple rules directly." The official documentation states plainly that VAP is an in-process alternative to validating webhooks. If you write a rule as a CEL (Common Expression Language) expression, the API server evaluates it on the spot. There is no network call, so there is no timeout, and there are no engine Pods, so the cluster cannot halt because the engine died. It became a stable feature in Kubernetes v1.30, and the API server in this lab environment is exactly that v1.30.

How it works

A single policy consists of up to three objects.

Object What it holds
ValidatingAdmissionPolicy The abstract logic of the rule. Which resources it looks at (matchConstraints), and what must be true (validations)
Parameter resource (optional) The values the rule uses, such as the list of allowed registries or the maximum replica count
ValidatingAdmissionPolicyBinding Ties the policy and parameters together and gives the scope of where to apply it

There is one sentence the official documentation states plainly. A policy has an effect only when both the policy and its corresponding binding exist. A policy without a binding is neither a syntax error nor a warning; it simply blocks nothing. It is the place where first-time users get caught the most.

apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicyBinding
metadata:
  name: demo-binding-test.example.com
spec:
  policyName: demo-policy.example.com
  validationActions: [Deny]
  matchResources:
    namespaceSelector:
      matchLabels:
        environment: test

Because they are separated, the same policy can be turned on several times with different scopes. With a maximum replica count of 3 in the test namespace and 100 in the production namespace. The policy YAML is one, and only the bindings and parameters are two. Because the structure means that when you fix a rule you touch only the logic and not the values, review becomes much lighter.

validationActions attaches to the binding. There are three supported values.

Value On violation
Deny Rejects the request
Warn Lets the request pass and reports a warning to the client
Audit Records it in the audit event

Deny and Warn cannot be used together (because a rejected request already tells the reason through the response body and the HTTP warning header). So the actual adoption procedure is to turn it on with [Warn, Audit], count for a few days, and after everything that gets caught is cleaned up, change only the binding to [Deny]. Not a single character of the policy body changes. This is the second reason for splitting the policy and binding.

failurePolicy attaches to the policy side and defaults to Fail. It decides whether to reject or pass when the expression evaluation itself ends in an error (for example, accessing a field without has() because it does not exist). There is one subtle rule written in the documentation. The failure that failurePolicy defines follows validationActions only when failurePolicy is Fail. With Ignore, that failure is simply ignored.

Producing a rejection message a person can read is also the policy's job. If you do nothing, the rejection message prints the expression text as it is, like failed expression: object.spec.replicas <= 5. If you give a CEL expression to messageExpression, you can produce a sentence with the actual values embedded, and if you name a long expression with spec.variables, you can reuse it in several places as variables.<이름> (variables dot the variable name). Variables are evaluated only when needed, so there is also the effect of computing an expensive expression only once.

spec:
  variables:
    - name: environment
      expression: "has(namespaceObject.metadata.labels) ? namespaceObject.metadata.labels['environment'] : 'prod'"
  validations:
    - expression: "..."
      messageExpression: "'only ' + variables.environment + ' images are allowed'"

matchConditions is one step earlier. If the CEL condition written there is false, the API server does not evaluate the policy at all. You use it to exclude requests from system service accounts or to look only at things with a particular label. If condition evaluation ends in an error, Fail rejects the request without evaluating the policy, and Ignore skips the policy and lets it pass.

Pulling values out is done by paramKind (on the policy side) and paramRef (on the binding side). In paramRef you can use only one of name or selector, and parameterNotFoundAction is required. With Allow, not finding the parameter counts as a pass, and with Deny, it follows the policy's failurePolicy. The documentation warns that a binding that omits this field is ignored or behaves in unexpected ways.

What you gain and lose compared with a webhook engine

Axis VAP Webhook engine
Installation None. Built into the API server Pods, certificates, CA bundle, upgrades
Availability The API server is the engine If the engine dies, writes stop under Fail
Latency In-process evaluation A network round trip per request
Expressiveness CEL. Sees only the incoming request and parameters Arbitrary code. Queries the cluster and also mutates and generates
Reporting Audit events and warnings Dedicated report CRDs such as PolicyReport

To sum up, move validation that CEL can express down into VAP, and leave in the engine what needs to query the cluster or needs mutation and generation. It is not a matter of choosing one of the two but of deciding the place for each rule.

What you see in the field

First, nothing is blocked. Nine times out of ten, the binding was not created or matchConstraints does not match the actual request. A policy that matches nothing cannot be told apart from a working policy on the surface. That is why, when you turn on a policy, you must try both a sample that should be blocked and a sample that should pass.

Second, the rejection message is the raw expression, so inquiries come in. A developer receives object.spec.template.spec.containers.all(c, ...) and has no idea what to fix. Producing a sentence such as "The image nginx:latest of container web is outside the allowed registries" with messageExpression is not kindness but a device that makes the policy actually followed.

Third, type checking is for reference. When you create a policy, the API server checks the expression in advance and leaves the result in status.typeChecking. But the documentation states the limits clearly. It does not check a matchConstraints with wildcards, it ignores from the eleventh onward when there are many combinations, it does not apply to CRDs, and the type-checking result does not change the policy's behavior. No errors here does not mean the policy is correct.

Fourth, the honest limits of this environment. The API server of the kwok cluster is real, so VAP's Deny and Warn both actually work, and dividing scope by namespaceSelector works too. But Pods are not run and are faked as Ready, so whether a blocked Pod truly did not start is confirmed not from containers but from the API response.

References

What you will do in the next lab

You attach VAP directly to the v1.30 API server that kwok runs. First you create only the policy and see with your own eyes that nothing is blocked, then you attach a binding and, moving validationActions from [Warn, Audit] to [Deny], see how the response to the same request changes. You make the rejection message readable with spec.variables and messageExpression, remove certain requests from evaluation entirely with matchConditions, pull the allowlist out of the policy with paramKind and paramRef, and then attach two bindings to the same policy to put different limits on each namespace.