CNPA — Cloud Native Platform Engineering Associate
Telling Them Before Blocking Them Is Developer Experience
In one sentence
Guardrails must be placed so that "block" and "inform" are separated. Define a rule in one place and run it in several places, and each rule must carry an identifier and a way to fix it so that it works both as a metric and as guidance.
Why this was needed
When a platform team decides to enforce rules, the answer that usually comes first is an admission webhook. But admission is the very last place. A developer writes code, commits, runs the pipeline, builds the image, and presses deploy, and only then hears "resource requests are missing." It takes 30 minutes to find out about a problem that takes 1 minute to fix.
Conversely, if you block everything up front, another problem arises. When every single rule becomes mandatory, a team trying to get a prototype running in a day looks for a way around the platform. Once workarounds start, the number of rules grows while the actual compliance rate falls, and the platform team does not even know it.
So you have to split the question in two: what is the danger when this rule is broken, and where is it cheapest to tell people.
How it works
A rule comes with an identifier, a severity, and a fix
If you write a rule only as a sentence, you can neither count it nor give guidance on it. Define each rule together with these three things.
| Item | Why it is needed |
|---|---|
Identifier (example: PR001) |
You can count compliance per rule, and you can grant exceptions per rule |
| Severity (block · warn) | Separating what to block from what only to inform about reduces workarounds |
| How to fix | The person who reads the diagnostic result knows what to do next |
The criterion for separating severity is not taste. You look at whether breaking it harms other people. An image whose tag is latest makes rollback impossible and breaks reproducibility across the whole cluster, so it is worth blocking. On the other hand, a missing readiness probe is a problem of that service's own availability, so it is better to inform and let the team decide.
Run the same rule in several places
Keep the rule definition in one place and run it in several places.
- Editor and local command — the cheapest and fastest. But whether it runs depends on the person.
- Commit hook — runs right before the commit. Because it is a local file it can be bypassed, so it cannot be the last line of defense.
- CI — the first place that runs outside the repository. From here on, bypasses are left on record.
- Admission — the last gate into the cluster. If you block only here, feedback is the latest.
The earlier places take charge of fast feedback and the later places take charge of guarantees. With only the earlier places, rules are not kept; with only the later places, you find out late.
Produce output for both humans and machines
If a diagnostic tool gives a human-readable line to people and JSON to programs, the same tool can be used in several places. The exit code is also a convention. If there is even one violation, it must end with a non-zero value so that the pipeline can use it as a signal. Without such a contract, you end up building a different tool for each place, and the rules differ slightly from place to place.
The compliance rate becomes an adoption metric
If you run the same diagnostic tool over the whole repository, you get a per-rule compliance rate. This number becomes the basis for deciding what the platform team builds next. If only a particular rule has an unusually low compliance rate, it is very likely not that those teams are lazy but that we made the rule hard to follow. It is a signal to put it in as a default, pre-fill it in the scaffold, or rewrite the wording.
But be careful about the denominator. If the list of diagnosed targets changes, the compliance rate moves even if nobody does anything. When you look at the number, you should also record what the denominator was.
What it looks like in the field
LabHub's graders have exactly this structure. A rule is set so that when a grading script fails, it does not write "the check failed" but states together what the current value is and what it should be. The reason is the same. For someone who does not know what is wrong, the only option left is guessing, and if the guesses miss two or three times, that person leaves.
There was also an incident in the opposite direction. This repository once had several graders that "pass no matter what answer you enter." They stayed around for a long time because nobody reported them. The same goes for a diagnostic tool. A diagnostic tool that catches nothing is worse than none. The compliance rate looks like 100%, but in fact no check is being done. That is why a diagnostic tool must also be tested in the opposite direction with deliberately rule-breaking samples.
How to keep guardrails from becoming walls
A platform's guardrails are devices that prevent incidents, but if they are built wrong, people learn how to get around them first. Then you lose control and you lose trust too.
When you block, give the alternative on the same screen. If it ends with "privileged is not allowed," the user goes looking for another cluster. It must go as far as "use this capability instead of privileged, and if you really need it, apply for an exception through this procedure." A rejection message must be the next action, not a link to a document.
Start with a warning. Run a new policy in audit mode first to count what it catches, show that list to the teams, and then block. If you block right away, all of that day's deployments fail, and the memory of it lasts a long time.
Make exceptions a procedure. Without exceptions the policy collapses, and if exceptions are easy it is the same as having no policy. The answer is an exception with a deadline. It expires automatically after 90 days, and a notification goes out before it expires.
Self-service comes before control. If people can easily get what they want, there is no reason to get around it. In an organization where creating one namespace takes three days of approval, everyone piggybacks on someone else's namespace.
Give diagnosis back to the user. If people have to ask the platform team "why won't my Pod start," that team becomes the bottleneck. If you gather events, policy rejection reasons, and resource shortages on one screen, most people solve it themselves.
kubectl get events -n <ns> --sort-by=.lastTimestamp | tail -20
kubectl describe pod <파드> | sed -n '/Events:/,$p'
The metric that measures success is not "the number of times we blocked." That number only measures how annoying the policy is. What you should measure is the share of users who, after a rejection, fixed it themselves and succeeded. If that share is low, the message is bad.
What to do in the next lab
You define five golden path rules as values with identifiers and severities, and build a diagnostic tool that actually evaluates those rules. You create one compliant sample and one sample that breaks only a particular rule, test in both directions, export the results as JSON too, and then calculate the compliance rate for the whole service set. Finally, you attach it to a commit hook and confirm that only block rules are actually blocked.