TT Lab
Get started
Learn Learning paths Courses

Policy as Code

The exception someone label-patched in a hurry was still there two years later

Continue in TT Lab

Goal

Manage policy exceptions as a register with deadlines and owners, and build for yourself a script that turns the register into the cluster's actual scope of application, an expiry checker that takes the reference date as an argument, a per-namespace violation tally, and a gate that fails when debt grows.

Why it matters

The second wall a team runs into after turning on a policy is not the quality of the rules but the lifetime of exceptions. When you turn on a rule, things that cannot be fixed right now inevitably come out, and then there are only two choices — take the policy down, or create an exception. If you take the policy down, everyone is released and it never goes back up. So you create exceptions, but if an exception has no deadline and no owner, an exception made with "we'll fix it next sprint" is still there two years later and makes the whole rule untrustworthy. This is why you should design the procedure for deleting exceptions before the procedure for creating them. The core of that design is to make the register the source of truth — a marker attached by hand in the cluster is something nobody knows who attached or why, but a register in the repository goes through review, carries a deadline, and even keeps a record of what was deleted. The cluster is merely a reflection of that register, and a marker not in the register cannot be explained, so it is deleted. And only with numbers can you reduce it. How many exceptions, how many expired, how many unauthorized violations nobody yet knows about — if you put these three numbers on one screen and set a ceiling, a change that increases debt stops in the pipeline rather than in people's memory.

Steps

  1. Do all the work in /root/polexc (export KUBECONFIG=/root/.kube/config, kubectl config use-context kwok-lab). First build the scenario — create the namespaces polexc-pay, polexc-legacy, and polexc-sandbox, and bring up six with kubectl create deployment: polexc-pay/pay-api (registry.internal/pay:1.4), polexc-pay/pay-batch (vendor.example/paybatch:latest), polexc-legacy/billing-api (vendor.example/billing:latest), polexc-legacy/report-gen (vendor.example/report:latest), polexc-legacy/promo-web (vendor.example/promo:latest), and polexc-sandbox/scratch-job (vendor.example/scratch:latest). Then write a ValidatingAdmissionPolicy no-latest-tag in /root/polexc/policy.yaml — catch CREATE and UPDATE of deployments in apps/v1, block it if a container image in the Pod template ends with :latest, but let it pass if the workload's namespace has the label polexc.io/exempt-<워크로드이름> (where the placeholder is the workload name; look at namespaceObject). In /root/polexc/binding.yaml, write a ValidatingAdmissionPolicyBinding no-latest-deny — validationActions is ["Deny"] and matchResources.namespaceSelector.matchLabels is polexc.io/enforce: "on". After applying both, attach the label polexc.io/enforce=on only to polexc-pay and polexc-legacy (do not attach it to polexc-sandbox). Now save the output of kubectl -n polexc-pay patch deployment pay-batch --type merge -p '{"metadata":{"annotations":{"polexc.io/probe":"1"}}}' --dry-run=server, including standard error, to /root/polexc/01-blocked.txt. Finally, delete only the binding (kubectl delete -f /root/polexc/binding.yaml), send the same request once each to pay-batch and polexc-legacy/promo-web, collect the two outputs in /root/polexc/01-off.txt, and then reapply the binding.
  2. Write the exception register in /root/polexc/exceptions.txt. One line is one entry with exactly six fields, and the delimiter is | — in the order namespace|workload|owner|expires|ticket|reason, where expires is YYYY-MM-DD. Put a header starting with # on the first line so that people can read it (the tool skips lines that start with # and blank lines). There are three entries — polexc-pay|pay-batch|team-pay|2026-09-10|OPS-2201|근거, polexc-legacy|billing-api|team-billing|2026-10-20|OPS-2188|근거, and polexc-legacy|report-gen|team-report|2027-03-31|OPS-2245|근거 (in each entry, replace the last field with your own justification). Write the justification yourself, at least four characters long. Next, in /root/polexc/scope.txt, write the scope of application of this policy, one per line — three lines: polexc-pay, polexc-legacy, and polexc-sandbox. polexc-sandbox has no enforcement label but is included in the scope where violations are counted.
  3. Write /root/polexc/apply-exceptions.sh. It takes two arguments — the register and the scope file (in that order, as <등록부> <적용범위파일>). It does two things. (1) For each entry in the register, attach the label polexc.io/exempt-<workload>=<ticket> to that namespace and print one line, GRANT <namespace>/<workload> <ticket> (in the same order as written in the register). (2) Sweep the namespaces in the scope file, delete the labels starting with polexc.io/exempt- that are not in the register, and print REVOKE <namespace>/<workload> (gather these lines sorted lexicographically by <namespace>/<workload> and output them after the GRANT lines). First reproduce an exception that someone hurriedly attached by hand — kubectl label ns polexc-legacy polexc.io/exempt-promo-web=MANUAL --overwrite. Then run bash /root/polexc/apply-exceptions.sh /root/polexc/exceptions.txt /root/polexc/scope.txt and save the output to /root/polexc/03-apply.txt. Three GRANT lines and one line REVOKE polexc-legacy/promo-web must come out.
  4. Check for yourself that the exception is narrow. polexc-legacy now has billing-api (has a registered exception) and promo-web (its marker was deleted in step 3) together. Send both workloads the same kind of request as in step 1, kubectl -n polexc-legacy patch deployment <이름> --type merge -p '{"metadata":{"annotations":{"polexc.io/probe":"4"}}}' --dry-run=server (replacing the placeholder with the workload name), one after the other, and collect the two outputs, including standard error, in /root/polexc/04-narrow.txt — billing-api must pass and promo-web must be rejected even though it is in the same namespace.
  5. Write /root/polexc/expiry.sh. It takes two arguments, <등록부> <기준일 YYYY-MM-DD> (the register, then the reference date). It prints one line per entry, in the same order as written in the register — <상태> <namespace>/<workload> <expires> <owner> <ticket> (status first). There are three statuses: EXPIRED if the expiry date is before the reference date, DUE if it is from the reference date through the reference date +30 days (both ends inclusive), and OK if it is later than that. The exit code is 1 if there is even one EXPIRED and 0 if there is none — this code becomes the gate. Run bash /root/polexc/expiry.sh /root/polexc/exceptions.txt 2026-10-01 and save the output to /root/polexc/05-expiry.txt (one EXPIRED, one DUE, and one OK come out).
  6. Actually retire the expired entry that step 5 found. First append that entry's line (polexc-pay|pay-batch|...) to /root/polexc/retired.txt and delete it from /root/polexc/exceptions.txt — only if a record of the deletion remains can you later answer "why did this disappear?" Then rerun bash /root/polexc/apply-exceptions.sh /root/polexc/exceptions.txt /root/polexc/scope.txt and save the output to /root/polexc/06-apply.txt (REVOKE polexc-pay/pay-batch must come out). Finally, send the same request as in step 4 to polexc-pay/pay-batch and save the output, including standard error, to /root/polexc/06-revoked.txt — it must be rejected now. The exit code of bash /root/polexc/expiry.sh /root/polexc/exceptions.txt 2026-10-01 must also be 0.
  7. Write /root/polexc/scan.sh. It takes one argument, the scope file (as <적용범위파일>). It prints one line per namespace, in the order written in the file — <namespace> <위반수> <예외수> <무단수> (namespace, then the violation count, exception count, and unmanaged count). A violation is a Deployment whose container image ends with :latest, and among those, the ones whose namespace has the label polexc.io/exempt-<이름> are the exception count and the ones without are the unmanaged count (a namespace with no violations is also printed as 0 0 0). Run bash /root/polexc/scan.sh /root/polexc/scope.txt and save the output to /root/polexc/scan.txt. Three lines come out: polexc-pay 1 0 1, polexc-legacy 3 2 1, and polexc-sandbox 1 0 1.
  8. Write /root/polexc/debt.sh. It takes four arguments, <등록부> <적용범위파일> <기준일> <부채상한> (register, scope file, reference date, and debt ceiling), and calls and uses expiry.sh and scan.sh in the same directory (find them with $(dirname "$0")). The output is exactly six lines — EXCEPTIONS <등록부 항목 수> (the number of register entries), EXPIRED <만료 수> (the number expired), DUE <임박 수> (the number due soon), UNMANAGED <무단 위반 합계> (the total of unmanaged violations), DEBT <EXCEPTIONS + UNMANAGED>, and finally GATE OK or GATE FAIL. If there is even one EXPIRED or DEBT exceeds the ceiling, print GATE FAIL and end with exit code 1. Otherwise it is GATE OK with 0. Run bash /root/polexc/debt.sh /root/polexc/exceptions.txt /root/polexc/scope.txt 2026-10-01 5 and save the output to /root/polexc/debt.txt (DEBT 5, GATE OK, and exit code 0 come out). Run the same command once more with ceiling 4 and see with your own eyes GATE FAIL and exit code 1 too.

Notes

Turning on the rule blocked things that cannot be fixed right now

Do all the work in /root/polexc (export KUBECONFIG=/root/.kube/config, kubectl config use-context kwok-lab). First build the scenario — create the namespaces polexc-pay, polexc-legacy, and polexc-sandbox, and bring up six with kubectl create deployment: polexc-pay/pay-api (registry.internal/pay:1.4), polexc-pay/pay-batch (vendor.example/paybatch:latest), polexc-legacy/billing-api (vendor.example/billing:latest), polexc-legacy/report-gen (vendor.example/report:latest), polexc-legacy/promo-web (vendor.example/promo:latest), and polexc-sandbox/scratch-job (vendor.example/scratch:latest). Then write a ValidatingAdmissionPolicy no-latest-tag in /root/polexc/policy.yaml — catch CREATE and UPDATE of deployments in apps/v1, block it if a container image in the Pod template ends with :latest, but let it pass if the workload's namespace has the label polexc.io/exempt-<워크로드이름> (where the placeholder is the workload name; look at namespaceObject). In /root/polexc/binding.yaml, write a ValidatingAdmissionPolicyBinding no-latest-deny — validationActions is ["Deny"] and matchResources.namespaceSelector.matchLabels is polexc.io/enforce: "on". After applying both, attach the label polexc.io/enforce=on only to polexc-pay and polexc-legacy (do not attach it to polexc-sandbox). Now save the output of kubectl -n polexc-pay patch deployment pay-batch --type merge -p '{"metadata":{"annotations":{"polexc.io/probe":"1"}}}' --dry-run=server, including standard error, to /root/polexc/01-blocked.txt. Finally, delete only the binding (kubectl delete -f /root/polexc/binding.yaml), send the same request once each to pay-batch and polexc-legacy/promo-web, collect the two outputs in /root/polexc/01-off.txt, and then reapply the binding.

The policy (what it judges) and the binding (where and with what strength) are different objects. So if you delete only the binding, the rule is still there but nothing is blocked — this is the reality behind "if you turn off the policy, everyone is released," and it is the reason exceptions are needed. In a CEL expression, you read the request object as object and its namespace object as namespaceObject. Check fields that may be absent first with has(...), and use in to see whether a key built by concatenating strings is among the labels. --dry-run=server runs admission as it is but leaves no object. The denial message comes out on standard error, so you need 2>&1.

Write exceptions in a file, not in your head

Write the exception register in /root/polexc/exceptions.txt. One line is one entry with exactly six fields, and the delimiter is | — in the order namespace|workload|owner|expires|ticket|reason, where expires is YYYY-MM-DD. Put a header starting with # on the first line so that people can read it (the tool skips lines that start with # and blank lines). There are three entries — polexc-pay|pay-batch|team-pay|2026-09-10|OPS-2201|근거, polexc-legacy|billing-api|team-billing|2026-10-20|OPS-2188|근거, and polexc-legacy|report-gen|team-report|2027-03-31|OPS-2245|근거 (in each entry, replace the last field with your own justification). Write the justification yourself, at least four characters long. Next, in /root/polexc/scope.txt, write the scope of application of this policy, one per line — three lines: polexc-pay, polexc-legacy, and polexc-sandbox. polexc-sandbox has no enforcement label but is included in the scope where violations are counted.

What an exception register must have is more than "what is being lifted." Without an owner, there is nobody to ask; without an expiry date, there is no basis for deleting; and without a ticket number, you cannot trace why it was lifted. Only with these four does an exception become a debt list rather than a switch. The format should be easy for people to read while splitting in one go with awk -F'|' or while IFS='|' read. The reason for writing the scope separately shows up later — to delete markers not in the register, "how far to sweep" must be decided.

An exception attached by hand was deleted in front of the register

Write /root/polexc/apply-exceptions.sh. It takes two arguments — the register and the scope file (in that order, as <등록부> <적용범위파일>). It does two things. (1) For each entry in the register, attach the label polexc.io/exempt-<workload>=<ticket> to that namespace and print one line, GRANT <namespace>/<workload> <ticket> (in the same order as written in the register). (2) Sweep the namespaces in the scope file, delete the labels starting with polexc.io/exempt- that are not in the register, and print REVOKE <namespace>/<workload> (gather these lines sorted lexicographically by <namespace>/<workload> and output them after the GRANT lines). First reproduce an exception that someone hurriedly attached by hand — kubectl label ns polexc-legacy polexc.io/exempt-promo-web=MANUAL --overwrite. Then run bash /root/polexc/apply-exceptions.sh /root/polexc/exceptions.txt /root/polexc/scope.txt and save the output to /root/polexc/03-apply.txt. Three GRANT lines and one line REVOKE polexc-legacy/promo-web must come out.

What you decide here is not technology but which side is the source of truth. If the register is the source, the cluster is merely a reflection of it, so a marker not in the register is something that cannot be explained and must be deleted. Conversely, if you make the cluster the source, nobody will know who lifted what, when, and why. If you put the ticket number in the label value, you can trace the justification from the cluster alone. To delete a label, append - after the key. You can get a namespace's list of label keys by sweeping kubectl get ns <n> -o json with jq. Running the script twice must give the same result.

Check that the exception does not lift the whole namespace

Check for yourself that the exception is narrow. polexc-legacy now has billing-api (has a registered exception) and promo-web (its marker was deleted in step 3) together. Send both workloads the same kind of request as in step 1, kubectl -n polexc-legacy patch deployment <이름> --type merge -p '{"metadata":{"annotations":{"polexc.io/probe":"4"}}}' --dry-run=server (replacing the placeholder with the workload name), one after the other, and collect the two outputs, including standard error, in /root/polexc/04-narrow.txt — billing-api must pass and promo-web must be rejected even though it is in the same namespace.

Most incidents where the exception scope widens are quiet. An exception lifted at the namespace level permanently leaves even workloads newly created in it in the future outside the rules, and because no error appears at all, it comes to light only months later. That is why, right after creating an exception, poking a neighbor that must not be lifted once should be part of the procedure. The two outputs must go into one file, so pass them bundled together or append the second.

An expiry date that nobody reads is the same as none

Write /root/polexc/expiry.sh. It takes two arguments, <등록부> <기준일 YYYY-MM-DD> (the register, then the reference date). It prints one line per entry, in the same order as written in the register — <상태> <namespace>/<workload> <expires> <owner> <ticket> (status first). There are three statuses: EXPIRED if the expiry date is before the reference date, DUE if it is from the reference date through the reference date +30 days (both ends inclusive), and OK if it is later than that. The exit code is 1 if there is even one EXPIRED and 0 if there is none — this code becomes the gate. Run bash /root/polexc/expiry.sh /root/polexc/exceptions.txt 2026-10-01 and save the output to /root/polexc/05-expiry.txt (one EXPIRED, one DUE, and one OK come out).

Taking the reference date as an argument is the key point of this step. If you read today's date inside the script, the same register is suddenly judged differently one day, and you can neither trace a past state nor test next quarter in advance. If you supply the reference date from outside, you can also answer "how many were there as of last month?" For date comparison, convert to seconds with date -d '<날짜>' +%s (replacing the placeholder with the date) and compare as integers, and the digit-count and month-end problems disappear. 30 days is 30*86400. You cannot change the system time in this Pod (date -s fails for lack of permission). So the reference date argument is the only way to test.

Once the expired exception was actually retired, that workload was blocked again

Actually retire the expired entry that step 5 found. First append that entry's line (polexc-pay|pay-batch|...) to /root/polexc/retired.txt and delete it from /root/polexc/exceptions.txt — only if a record of the deletion remains can you later answer "why did this disappear?" Then rerun bash /root/polexc/apply-exceptions.sh /root/polexc/exceptions.txt /root/polexc/scope.txt and save the output to /root/polexc/06-apply.txt (REVOKE polexc-pay/pay-batch must come out). Finally, send the same request as in step 4 to polexc-pay/pay-batch and save the output, including standard error, to /root/polexc/06-revoked.txt — it must be rejected now. The exit code of bash /root/polexc/expiry.sh /root/polexc/exceptions.txt 2026-10-01 must also be 0.

The procedure for deleting exceptions is harder and more important than the procedure for creating them. If the deleting side does not run automatically, exceptions quietly become permanent, and from then on that rule becomes "a rule that is on but nobody trusts." Thanks to making the register the source of truth in step 3, all you do here is delete one line from the register and rerun the same script — you do not write separate revocation logic. When you edit the register in place, write to a temporary file and then move it. Even if you run this step twice, the same line must not go into retired.txt twice.

When the violations were counted, unmanaged ones outnumbered exceptions

Write /root/polexc/scan.sh. It takes one argument, the scope file (as <적용범위파일>). It prints one line per namespace, in the order written in the file — <namespace> <위반수> <예외수> <무단수> (namespace, then the violation count, exception count, and unmanaged count). A violation is a Deployment whose container image ends with :latest, and among those, the ones whose namespace has the label polexc.io/exempt-<이름> are the exception count and the ones without are the unmanaged count (a namespace with no violations is also printed as 0 0 0). Run bash /root/polexc/scan.sh /root/polexc/scope.txt and save the output to /root/polexc/scan.txt. Three lines come out: polexc-pay 1 0 1, polexc-legacy 3 2 1, and polexc-sandbox 1 0 1.

Admission sees only what is coming in from now on. Nobody rechecks what is already inside the cluster, so even after turning on the rule, "how many are violating right now" can be known only by sweeping separately. And that number has meaning only if you split it in two — what is covered by an exception is debt with a deadline, and unmanaged violations are holes nobody yet knows about. If you merge them, neither is visible. You must count namespaces without an enforcement label, like polexc-sandbox, too — enforcement being off does not mean there are no violations. You can get the image list in one go by sweeping kubectl get deploy -o json with jq.

Put the debt on one screen and make it fail when it grows

Write /root/polexc/debt.sh. It takes four arguments, <등록부> <적용범위파일> <기준일> <부채상한> (register, scope file, reference date, and debt ceiling), and calls and uses expiry.sh and scan.sh in the same directory (find them with $(dirname "$0")). The output is exactly six lines — EXCEPTIONS <등록부 항목 수> (the number of register entries), EXPIRED <만료 수> (the number expired), DUE <임박 수> (the number due soon), UNMANAGED <무단 위반 합계> (the total of unmanaged violations), DEBT <EXCEPTIONS + UNMANAGED>, and finally GATE OK or GATE FAIL. If there is even one EXPIRED or DEBT exceeds the ceiling, print GATE FAIL and end with exit code 1. Otherwise it is GATE OK with 0. Run bash /root/polexc/debt.sh /root/polexc/exceptions.txt /root/polexc/scope.txt 2026-10-01 5 and save the output to /root/polexc/debt.txt (DEBT 5, GATE OK, and exit code 0 come out). Run the same command once more with ceiling 4 and see with your own eyes GATE FAIL and exit code 1 too.

This one screen is everything the meeting must answer — how many exceptions are there, is anything past its deadline now, how many will pass it soon, and how many violations does nobody know about. If the numbers are scattered, nobody looks, and if you merge them into one line, you cannot tell what to fix. The reason to set a ceiling is that "let's reduce it" does not keep itself. With a ceiling, a change that increases debt stops in the pipeline. It matters to reuse the two scripts from the earlier steps — if the counting rule is written in two places, one of them inevitably goes stale first. expiry.sh ends with exit code 1 if anything has expired. Make sure the side that uses its output does not die on that code.