TT Lab
Get started
Learn Learning paths Courses

CNPA — Cloud Native Platform Engineering Associate

Where Must a Baseline Be Nailed Down to Hold

Continue in TT Lab

In one sentence

A platform's minimum baseline is kept only when it exists in a form the API server can reject, not as a wiki document. And when you raise the baseline, you first announce it with warnings, investigate the impact, and give those who cannot comply an exception with an expiry date.

Why this was needed

The first deliverable a platform team builds is usually a document. You compile rules such as "No Pod runs as root," "Pin image tags," and "Keep the metrics port open," and post them to the internal wiki. Then half a year later you look over the cluster and fewer than half of the workloads follow those rules.

The reason is simple. A document rejects nothing. A deployment that breaks the rules also succeeds, and nobody rolls back a successful deployment. By contrast, when the API server rejects something, you find out on the spot, and if you do not fix it, the deployment does not go through. The substance of a baseline is not a document but the ability to reject.

But if you block everything from the start, another problem arises. If a deployment that worked until yesterday is suddenly blocked today, the platform is remembered as the cause of an outage. That is why a baseline needs both levels and a procedure.

How it works

The three levels and three modes of Pod Security Admission

Pod Security Admission, built into Kubernetes, works purely from namespace labels. There are three levels, privileged, baseline, and restricted, and you can set a mode separately for each level.

Mode What it does Where to use it
enforce Prevents Pods that violate the level from being created The level you can comply with now
audit Records only in the audit log When you look at statistics later
warn Shows a warning to the person creating the Pod The level you will raise to next

There is one trick here. Put enforce on the level you can comply with now, and put warn on the level you will raise to next. Then developers read "this will be blocked in the next stage" in advance every time they deploy, and are not surprised on the day you raise the level.

In the label you write not only the level but also the version (enforce-version: v1.30). The definition of a level broadens slightly with each Kubernetes version, and if you do not write the version, the content of the policy silently changes the moment you upgrade the cluster. If deployments are blocked that day even though nobody touched the policy, finding the cause takes a day.

How to investigate before raising

Do not guess whether you can raise the level; ask the server. If you send the request that changes the namespace label as a server dry-run, the API server evaluates the Pods currently running in that namespace against the new level and returns the result as warnings. Which Pod is caught, and why, comes out as is. It is a way to get the answer without creating anything.

Conformance is portability

The conformance that the CNCF talks about is not a "certification stamp" but the promise that if you use only the upper-level APIs, it behaves the same on any cluster. That is why leaving deprecated APIs behind breaks conformance. A manifest written with extensions/v1beta1 is rejected on a recent cluster not as a syntax error but as a mapping failure, because that group is not served at all. The message you get then is "no matches for kind," and if you can read this sentence, you know the cause within a minute.

Sometimes fields increase when you migrate. When you move an Ingress to networking.k8s.io/v1, you must write a pathType for each path. The API now makes explicit something whose interpretation used to differ from controller to controller.

Observability is a contract

For a platform to say "we collect metrics automatically," there must be an object that records what it collects from. The Prometheus Operator's ServiceMonitor plays that role. A ServiceMonitor selects targets by the Service's labels and selects scrape targets by the Service's port name. If either one is off, the object is created perfectly well and only the metrics are silently empty. So after creating this relationship, you must ask back "what did it actually select?" to confirm.

Give exceptions with an expiry date

There will always be teams that cannot meet the baseline right now. You have three options here: lower the baseline, block that team, or grant an exception with an expiry date. The first two respectively make the platform meaningless or make the platform an enemy. For an exception, write why it is an exception and until when in the object. An exception without that written becomes a permanent exemption, and when permanent exemptions pile up, the baseline goes back to being a document.

What it looks like in the field

The same thing happened in the homelab where this course runs. Lab Pods start without privileges (with all capabilities dropped), and if that premise had been written only in a document, it would have been opened up one by one for convenience. Now the Pod spec itself enforces it, so if a lab does not work, you change the lab design instead of opening up permissions. When the baseline is in the object, the argument moves into design.

The same kind of incident happened on the observability contract side. A metrics scrape target failed to attach because of a single character in a label, but the object was healthy and only the dashboard was empty, so for a while nobody noticed. "It was created" and "it is actually selecting something" are different claims, and this distinction can be confirmed only by asking the cluster back.

What to do in the next lab

You set up a baseline in one namespace and confirm that it actually rejects. After seeing with your own eyes why a deprecated API is rejected and moving to the upper-level API, you establish a scrape contract with a ServiceMonitor and ask back whether its selector really selects that Service. Finally, you ask the server about the impact of raising the baseline by one level, and leave an exception with an expiry date on the namespace that cannot comply.