TT Lab
Get started
Learn Learning paths Courses

Policy as Code

Testing policies: pinning both the passes and the rejections

Continue in TT Lab

In one sentence

The minimum unit of a policy test is a three-column table of input sample, expected decision, and expected message, and if this table has no deny samples, that test protects nothing.

Why it was needed

Saying policy is code is not a metaphor. It goes into a repository, goes through review, gets deployed, and gets fixed. But there is one difference from application code. The most common way a policy breaks is not an error but silence. If the match scope is off and it sees nothing, the policy quietly lets everything through. Nothing happens on screen, and there is no red line on the dashboard. So "a policy that is running well" and "a dead policy" cannot be told apart on the surface.

A second force overlaps here. Policies are fixed only in the direction of loosening. When someone is blocked and files an inquiry, one condition is relaxed and one more exempt namespace is added. Each change is reasonable at the time. After half a year, nobody knows what that policy was originally meant to block. A regression test is the nail driven in at exactly this point. If the test has a deny sample, the test goes red on the very day the rule is loosened.

How it works

Table-based (golden case) tests. One test consists of three columns.

Column Content
Input sample A manifest fragment that could really go to the cluster. Both what should pass and what should be blocked
Expected decision pass / fail / skip
Expected message The sentence a person receives on denial

Many teams leave out the third column, but the message is also a contract. The denial message is the only channel through which the policy speaks to developers, and when that sentence changes, the pipelines that grabbed it and automated on it change too. If you put the message in the test, the inquiry "I don't know why it was blocked" surfaces at the review stage.

It is good to make samples different in only one way at a time. If you keep several pairs made by copying a passing sample and changing just one violating field, when a test breaks, what changed can be read straight from the file name.

Three places to run tests.

  1. Before going onto the cluster, an offline engine. You run the decision with only the policy file and the sample files. It needs neither network nor cluster, so it runs fastest in CI. Kyverno CLI's kyverno apply and kyverno test, Conftest, and OPA belong here.
  2. Before going onto the cluster, a server dry-run. A request sent with ?dryRun=All passes the normal path as it is except for the storage step. The documentation says that at that time all the relevant admission controllers run, validating admission sees the object after mutation is finished, and defaults are filled in and schema validation also happens. So this is not "imitation" but a real decision. What an offline engine cannot see (interactions with other policies, the shape after mutation) is visible only here.
  3. After going up. You sweep the resources that already exist with background scanning and reports, because admission sees only requests coming in from now on.

There is a companion fact to server dry-run. A request that an admission controller with side effects catches is, if it is a dry-run, rather failed. So a webhook must declare sideEffects of its configuration object as None or NoneOnDryRun to be evaluated normally on the dry-run path. All the built-in admission plugins support dry-run.

Common causes of a false pass. When the test is green but the cluster does not block, the cause is usually one of three.

The CI of a policy repository usually goes in this order.

1) 정책 파일 스키마·문법 검사
2) 오프라인 엔진으로 골든 케이스 표 실행 (pass·fail·메시지)
3) skip 이 0 인지 확인 — 매치가 어긋나지 않았다는 증거
4) 거부 기대 건수가 0 이 아닌지 확인
5) 스테이징 클러스터에 서버 dry-run 으로 대표 표본 몇 개
6) 배포

Steps 3 and 4 are the ones most often missing from this list, and they betray most quietly.

What you see in the field

First, a policy that is green but does not block in production. It is common for every sample to have been skip because the sample's apiVersion was one version lower. Someone looked only at the pass count on the test output's summary line and not at the skip count. A single line that forces skip to 0 in the summary blocks this incident entirely.

Second, the day a rule was loosened. "Please exclude just this namespace" repeats, and one day the exception condition becomes so broad that even what it originally blocked passes. At that moment, if the test has a deny sample, red appears in the PR, and if not, you find out half a year later as an incident.

Third, the day a message was quietly changed. While refactoring the policy, you polished the wording of the denial message, and a pipeline that grabbed that string to make Slack alerts silently stops. If you pin the message as an expected value, the person refactoring finds out on the spot.

Fourth, the honest limits of this environment. The lab Pod has no conftest or opa. Instead, there is the kyverno CLI, so you can run a table-based test of the same shape, and there is the real API server run by kwok, so the server dry-run actually works too. Even if the tool names differ, what you learn is the same.

References

What you will do in the next lab

You build from scratch a test suite to attach to one policy. You put passing samples and deny samples in pairs and write the expected decisions and expected messages in a table, and then write a runner that uses the second place above (the server dry-run) as its decision engine. You make a false pass by hand by putting into the deny samples a resource the policy does not even look at, and add a check that catches it. Then you loosen the rule by one notch and see a deny sample pass, and confirm the scene where the regression test catches that change. Finally, you set an exit code convention so that a test with no samples at all does not go green.