TT Lab
Get started
Learn Learning paths Courses

KCA — Kyverno Certified Associate

Some Controllers Don't Get Faster With More Replicas

Continue in TT Lab

In one line

Kyverno is not one program but a structure in which four controllers, admission, background, reports, and cleanup, each run as their own Deployment. So high availability (HA) is also decided per controller, an upgrade must not be done by only raising the image tag, you test policies with the CLI before putting them in the cluster, and after putting them in you watch them with the kyverno_* metrics. This article explains together the installation methods documentation, the high availability guide, the upgrade documentation, the CLI reference, and the metrics reference.

Why this was needed

An admission webhook is fail closed by default. The installation documentation describes this risk at length — if the API server cannot reach Kyverno, a request to create a resource that falls under a policy fails on the grounds that the policy cannot be evaluated. In a cluster with a single policy that enforces Pods to be non-root, if all the Kyverno Pods die, you cannot create a single new Pod. That is why a production environment must be installed with HA, and Kyverno's own namespace must be excluded from the webhook (the default configuration excludes kyverno and kube-system).

Three operational questions follow from this. How many of which controllers do you run, how do you upgrade to a new version, and how do you see what the policies are actually blocking.

How it works

The four controllers and the Helm installation

The installation documentation divides the controllers like this. The admission controller is mandatory; it receives the API server's webhook callbacks and handles validate, mutate, image verification, and PolicyException. The background controller handles generate and mutate-existing rules, the reports controller handles PolicyReports, and the cleanup controller handles CleanupPolicy. Each controller has its own ServiceAccount, so permissions are separated, and the default installation has 1 replica of each.

helm repo add kyverno https://kyverno.github.io/kyverno/
helm repo update
helm install kyverno kyverno/kyverno -n kyverno --create-namespace \
  --set admissionController.replicas=3 \
  --set backgroundController.replicas=2 \
  --set cleanupController.replicas=2 \
  --set reportsController.replicas=2

The above is the documentation's HA installation example. As a "complete" HA deployment, a values example where all four controllers have replicas: 3 is also included. You can also install with YAML manifests, in the form of kubectl create -f on the install.yaml of a tagged release, and the documentation states firmly that direct upgrades are not supported with this method. If you need the PSS (Pod Security Standards) policy bundle, you install the separate chart kyverno/kyverno-policies.

What replicas do differs by controller

This is the fact the HA guide stresses most. The admission controller does not use leader election for webhook requests, so all replicas share and handle the requests — replicas are used for both availability and throughput. Only the certificate and webhook management is handled by a single leader. The minimum number of replicas recognized as HA is 3. The reports controller and the background controller, on the other hand, are stateful services and so use leader election, and only one works no matter how many replicas there are. So for these two, replicas help only availability, and to raise throughput you have to increase the resources of the individual Pod (vertical scaling), not the replica count. This is what the installation documentation's sentence "more replicas do not mean higher performance in every controller" means.

There are also values to know about the webhook's own configuration. The default failurePolicy is Fail, and you can change it per policy or change everything with --forceFailurePolicyIgnore. The default webhookTimeout is 10 seconds (1–30 seconds). The default resourceFilters exclude Event, Node, and resources in the kube-system, kube-public, kube-node-lease, and kyverno namespaces.

The kinds of CRDs

The CRD documentation says you can see all the types with kubectl explain. The kinds are laid out in the v1.19 section of the upgrade documentation.

Kind API group/version Role
Policy / ClusterPolicy kyverno.io/v1 Traditional validate, mutate, generate, and verifyImages policies
CEL policies such as ValidatingPolicy policies.kyverno.io/v1 ValidatingPolicy, MutatingPolicy, GeneratingPolicy, DeletingPolicy, and ImageValidatingPolicy
CleanupPolicy / ClusterCleanupPolicy kyverno.io/v2 Schedule-based cleanup
PolicyException kyverno.io/v2(legacy) or policies.kyverno.io Exceptions
GlobalContextEntry kyverno.io/v2 (v2alpha1 is deprecated) Cached external data
PolicyReport / ClusterPolicyReport wgpolicyk8s.io/v1alpha2 The final report
EphemeralReport / ClusterEphemeralReport reports.kyverno.io/v1 Intermediate products of reports
UpdateRequest Internal type Intermediate products of generate and mutate-existing

From v1.19, the CRDs are managed as a chart dependency called kyverno-api, and the crds.install value turns it on and off. The same documentation announces that in v1.19 ClusterPolicy, Policy, CleanupPolicy, and the legacy PolicyException are deprecated and will be removed in v1.20.

Upgrades — why you cannot just raise the tag

The first sentence of the upgrade documentation is the principle. New versions change many supported resources, including the CRDs, so you cannot upgrade just by raising the image tag. To do a Helm upgrade from a version before 1.10 to 1.10 or later, a direct upgrade is impossible and you must follow the chart v2→v3 migration guide. If you skip minor versions, you must read the release notes of every minor version in between.

The v1.13 section is a good example. The wildcard view permission was removed, so mutate and generate policies and reports that looked at custom resources were affected, exceptions (PolicyException) that were allowed in all namespaces by default were changed because of a security issue (CVE-2024-48921) so that the features.policyExceptions.namespace value had to be specified explicitly, and as old CRD API versions were removed, a Helm hook handled the migration automatically. In v1.19, CLI commands such as kyverno migrate --resource policyexceptions.kyverno.io were added for migrating the storage version. The point is one — an upgrade is not a code swap but an event in which the CRDs, permissions, and defaults change together, so the release notes are the procedure manual.

The CLI — testing without a cluster

The Kyverno CLI is an executable separate from the controllers, and it is used without a cluster, as the reference's description of --kubeconfig says "needed only when running outside the cluster." In CI, you install it with the GitHub Action kyverno/action-install-cli, as the policy testing guide shows. There are three core commands.

kyverno apply policies/ -r resources/ applies policies to resource manifests and shows the results. You use it when you know the policies but not the resources — for example, when checking manifests that came in on a development team's PR.

kyverno test <디렉터리 또는 git 저장소> (the placeholder is a directory or a git repository), conversely, compares the actual results against a test manifest (kyverno-test.yaml, file name changeable with -f) in which the policies, resources, and expected results are written in advance. The expected results are pass, fail, and skip, and you can create a test file with kyverno create test -p policy.yaml -r resource.yaml --pass 정책이름,규칙이름,리소스이름,네임스페이스,종류 (the placeholders are the policy name, rule name, resource name, namespace, and kind). You test a branch of a remote repository with --git-branch, pick only some with --test-case-selector "policy=..., rule=..., resource=...", and change the output format as in -o junit.

kyverno jp is a JMESPath command line with Kyverno's custom functions added. You evaluate an expression against a file with kyverno jp query -i object.yaml '식' (the placeholder is the expression), look at the function list with kyverno jp function, see the description of a particular function as in kyverno jp function truncate, and look at the expression's syntax tree with kyverno jp parse. As seen in the earlier module, the textbook practice is to check the result of an apiCall in advance with kubectl get --raw ... | kyverno jp query "items | length(@)".

Metrics — what to look at

According to the monitoring guide, a Helm installation creates a metricsService for each controller and serves metrics at /metrics on port 8000. The default service type is ClusterIP, so only a Prometheus inside the cluster can scrape it, and to scrape from outside you change it to NodePort or LoadBalancer. You adjust the exposure scope with the kyverno-metrics ConfigMap — namespaces.include/exclude (exclude takes precedence), the histogram bucketBoundaries, and in metricsExposure you turn off metrics individually (enabled: false), drop label dimensions (disabledLabelDimensions), or change buckets. The documentation notes that narrowing the namespaces noticeably reduces memory use.

The main metrics the metrics reference lays out are these.

Metric Kind What it shows
kyverno_policy_rule_info_total Gauge (1 if the rule is active) Which policies and rules are in the cluster now. policy_type (cluster/namespaced), policy_validation_mode (enforce/audit), rule_type, status_ready
kyverno_policy_results Counter Rule execution results. rule_result (PASS/FAIL), rule_execution_cause (admission_request/background_scan), resource_kind
kyverno_policy_execution_duration_seconds Histogram The execution latency of a single rule
kyverno_admission_review_duration_seconds Histogram The overall admission latency for one request (summed over all policies)
kyverno_admission_requests_total Counter The number of admission requests and request_allowed
rate(kyverno_policy_results{resource_kind="Pod", rule_execution_cause="admission_request"}[1m])*60

The above is a query from the reference, the number of rule executions per minute caused by Pod requests. There are two things you look at first in practice — whether the admission latency histogram is approaching webhookTimeout, and in which policies and namespaces rule_result="FAIL" is growing. The former is a warning before the fail-closed webhook stops the cluster, and the latter is the list of what the policies are actually blocking. The Grafana dashboard JSON is included in the chart and can be deployed with the grafana.enabled value.

What it looks like in the field

An organization increased the reports controller to 5 replicas because it was slow, and nothing got faster. That is because only the leader works. The answer was not replicas but the CPU and memory of the leader Pod, and at the same time they reduced the metrics load by excluding namespaces in kyverno-metrics.

A common incident in upgrades is skipping two minors at once with Helm without reading the release notes. Because of the permission change in 1.13, a generate policy that targeted custom resources quietly stopped, and they discovered it late through status_ready="false" of kyverno_policy_rule_info_total.

What to check in the next quiz

The quiz asks which controller is mandatory, which controllers benefit from replicas in throughput and which do not, why you cannot just raise the tag in an upgrade, the difference between kyverno test and kyverno apply, the purpose of kyverno jp, and the labels of kyverno_policy_results.