TT Lab
Get started
Learn Learning paths Courses

Policy as Code

Reading the plan as policy: catching destroys, replacements and exposure before apply

Continue in TT Lab

In one sentence

The plan of Terraform and OpenTofu is a structured change specification you can extract as JSON, and if you apply policy to that JSON, you can catch destruction, replacement, and exposure scope before resources are created.

Why it was needed

Admission control is powerful, but it looks at requests that have already been made. And quite a few of the world's dangerous changes never pass through the Kubernetes API in the first place. Opening a security group to 0.0.0.0/0, replacing a database, switching a storage bucket to public. These go straight to the cloud API. They do not come to a place admission can see.

But these changes have something in common. There is a planning stage before they are applied. tofu plan and terraform plan first compute "what will be created, what will be deleted, and what will be changed." The output is made for people to read, but it can also be extracted in a machine-readable form. Then this plan becomes a document you can check with policy. If admission is the gatekeeper of the cluster, plan checking is the gatekeeper of the whole cloud. And it is much cheaper — because there is not yet anything to revert.

How it works

The procedure is two lines. You save the plan to a file, and convert it to JSON.

tofu plan -out=tfplan.binary
tofu show -json tfplan.binary > plan.json

Terraform is the same (terraform show -json). In the resulting JSON, the place a policy looks is mostly the resource_changes array, and each item holds address, type, name, and a change object. The actions inside change is the center of the decision. The valid values the documentation states plainly are these.

actions Meaning
["no-op"] Nothing changes
["create"] Creates it new
["read"] Reads a data source
["update"] Fixes it in place
["delete"] Deletes it
["delete", "create"] Replacement. Delete and then recreate
["create", "delete"] Replacement, but create first and delete afterward

That replacement is expressed as a two-element array is deliberate. The documentation explains that if it is expressed this way, a caller can catch all three cases where a resource disappears just by scanning whether delete is in the list. The policy's first rule comes from here. If "delete" in actions, require human approval. This is especially so for stateful resources such as databases or volumes.

Destruction is not the only thing you can ask about.

A plan has values it does not yet know. This is the biggest constraint of plan policies. Values such as resource IDs, generated ARNs, and random suffixes are decided only when you apply. The JSON expresses these as after_unknown, and the documentation's description is precise. It is an object with the same structure as after in which unknown leaf values are true and known leaf values are omitted entirely. So when you write a policy, the rules become these.

Turning the decision into a pipeline result — the exit code is not always the truth. This is the most practical part of this module. When putting a policy tool into CI, people naturally write if ! tool scan ...; then exit 1; fi. But there really are tools that find violations and still exit with 0. The kyverno json scan in this lab environment is like that. Violations are printed as FAILED in the human-readable output and also remain in Violations of --output json, but the exit code is 0 whether or not there are violations (confirmed directly with the kyverno CLI 1.13.2 in the lab image). If you do not know that and trust only the exit code, the pipeline is green forever.

So when attaching a policy gate, the order is this.

1) 도구를 일부러 실패할 입력으로 돌려 본다
2) 종료 코드를 확인한다 (echo $?)
3) 0 이면 종료 코드를 쓰지 않고 보고서를 파싱한다
4) 파싱 결과가 비어 있지 않은지도 확인한다 (형식이 바뀌면 0건으로 읽힌다)
5) 그 판정으로 파이프라인을 세운다

Skipping step 2 is where incidents start, and skipping step 4 makes the gate quietly turn off on the day of the tool version upgrade.

The engines used in the field are things like Conftest (which checks arbitrary JSON with OPA's Rego), OPA itself, and kyverno-json. This lab Pod has no conftest or opa. Instead, it does the same job with the kyverno CLI's kyverno json scan — the structure of giving an arbitrary JSON document as a payload and judging it with policy is the same. Even if the tool differs, what you learn (the shape of the plan JSON, unknown values, the exit code trap) is the same.

What you see in the field

First, the day the approval gate really stopped something. A plan with ["delete", "create"] attached to an RDS instance came up in a PR, and the gate caught it so a person checked. Whether changing one field triggers a replacement cannot be known from the code alone and shows up only in the plan.

Second, a false positive caused by an unknown value. A required-tag check was written using only after.tags, and in a module that computes tags with a local variable, that value did not yet exist at the plan stage, so everything showed as a violation. The gate was turned off within days. It should have been fixed into a rule that also looks at after_unknown.

Third, a gate that was quietly off. After the tool was upgraded, one key in the report JSON changed, so the parse result was always an empty array. It read as 0 violations and stayed green for weeks. A single line that also checks "is the result non-empty?" prevents this.

Fourth, a plan depends on state. Even with the same code, the plan differs if the state file differs. If you apply as it is, several days later, a plan made in a PR, the changes in between are not reflected. Plan checking is right when done with a plan made again just before deployment.

References

What you will do in the next lab

You build a plan with an offline provider, extract it as JSON, and apply policy to that JSON. You sweep resource_changes to find the items containing delete, tell replacements from simple updates, and write a rule that catches resources missing the required tags. You deliberately create a place where the value is not yet in the plan, confirm that a rule that does not look at after_unknown produces a false positive, and then fix it. Finally, you feed the tool an input with a clear violation and print the exit code yourself to confirm that 0 comes out, and then build a gate that parses the report and sets up the pipeline.