TT Lab
Get started
Learn Learning paths Courses

Policy as Code

From Audit to Enforce — The Order in Which You Actually Turn Policy On

Continue in TT Lab

In one sentence

A policy is harder to turn on than to write, and the order of turning it on is always three stages: observe → clean up exceptions → enforce.

Why it was needed

What happens to a team that writes one policy well and immediately raises it to Enforce is usually the same. Deployments are blocked en masse, inquiries pile up in Slack, and a few hours later someone says, "Please just take the policy down for now." The policy comes down and never goes back up. The cause of the failure is not the quality of the policy. It is that nobody knew whether the things already running were following the rule.

A cluster does not hold only newly deployed workloads. A batch job that went up two years ago, a Helm chart a vendor provided, and a Pod some team started temporarily and forgot all live there together. You cannot know how many of them a new rule will catch until you apply it. That is why a policy engine always has a mode that counts without blocking.

How it works

There are two key mechanisms.

First, Audit and Enforce.

Mode Violating request What remains
Audit Lets it pass Recorded as fail in the report
Enforce Rejects it Remains in the report, and the user knows immediately

Audit is not "a state where the policy is not turned on" but "a state with only enforcement turned off." The decisions still happen and the results pile up. So if you leave it in Audit for a few days, you can answer the question "how many would be blocked if we turned on this rule?" with a number. Since the enforcement level moved down to the rule level (validate.failureAction), within one policy you can raise only the verified rules to Enforce and layer new rules on as Audit.

Second, background scanning and PolicyReport.

The admission stage sees only what is coming in from now on. Nobody rechecks objects already inside the cluster. So for a policy with spec.background: true, the background controller periodically sweeps existing resources and judges them with the same rules. The results pile up as reports.

A report holds per-resource decisions (policy, rule, result, message) in the results array, and the aggregate counts of pass, fail, warn, error, and skip in summary. This summary.fail is the progress metric of the adoption work. If it was 47 on the day you turned the policy on and is 3 two weeks later, those 3 are the remaining exception candidates.

Third, PolicyException.

Among what remains, there are things that truly cannot be fixed. A vendor image that cannot take limits, a node agent that needs privileges. In that case, instead of rolling the policy back, you leave the exception as a document. An exception has three layers of scope.

  1. spec.exceptions[].policyName — an exception for which policy
  2. spec.exceptions[].ruleNames — exempt from which rules of that policy only
  3. spec.match — which resources it applies to (kinds, namespaces, and names)

It is important to narrow all three layers. If you omit ruleNames and exempt the whole policy, that resource will also be left out of every rule added to that policy in the future. If you write only a namespace without names, even workloads newly created in it are permanently left outside the rules. An exception must not be a switch that turns the policy off but a debt list. When the list grows longer it becomes visible, and what is visible can be reduced. If you attach an expiry date or an owning team to each exception as an annotation, the list manages itself.

Using these three mechanisms in order gives you an adoption procedure.

1) 정책을 Audit 으로 배포 + background: true
2) 며칠~몇 주 리포트의 summary.fail 을 관찰
3) 고칠 수 있는 것은 팀과 함께 고친다 (fail 이 줄어드는 것을 본다)
4) 못 고치는 것만 좁은 PolicyException 으로 남긴다
5) fail 이 예외 건수까지 내려오면 그 규칙을 Enforce 로 올린다
6) 예외 목록을 주기적으로 다시 본다

What you see in the field

First, when Audit is on and nobody looks. Reports pile up, but without a dashboard or alerts, only CRD objects increase. Audit is observation, not a person who reads the observation results. The moment you deploy a policy in Audit, you must also decide "who looks at this number, and when."

Second, when reports pile up and press on the cluster. In a cluster with many resources, if you turn on a policy that matches everything with background, report objects are created in bulk. If the report controller cannot keep up with the aggregation rate, a backlog begins. Narrowing the match scope is the answer here too.

Third, when exceptions quietly become permanent. Clusters where an exception made with "we'll fix it next sprint" is still there two years later are common. You must write an expiry date and an owner on each exception and create a procedure to sweep the list every quarter. It is harder and more important to build a procedure for deleting exceptions than to create them.

Fourth, the honest limits of this environment. The background controller does not run here, so PolicyReports do not pile up in the cluster on their own. Instead, you can build a report of the same format locally with kyverno apply --policy-report, and the pass/fail numbers in its summary mean the same thing as those in a real report. The report you built in the pol-validate lab was exactly that.

What to look for in the next check

In the lab that follows, you build by hand the last two paragraphs of this text. You set up an exception register with owners, expiry dates, and justifications, write a checker that finds and blocks expired exceptions, confirm that an exception really narrows the policy's scope, and then gather the violations in one place and build a report that answers with numbers "is it okay to raise it to Enforce now?" There you will experience that building the procedure for deleting exceptions is harder than creating them. In the quiz after the lab, you confirm how the three layers of policyName, ruleNames, and resource names in /root/policy/validate/exception.yaml narrow the exception scope, and why you manage it as a debt list together with expiry dates and owners.