TT Lab
Get started
Learn Learning paths Courses

CNPA — Cloud Native Platform Engineering Associate

Setting a Platform Baseline and Raising It

Continue in TT Lab

Goal

You pin the minimum baseline the platform requires onto a namespace, confirm that it actually rejects, and then run the process of raising it by one level all the way around, from impact investigation to exception handling.

Why it matters

If you keep a baseline only as a document, it rejects nothing, and a rule that cannot reject is half broken six months later. Conversely, if you block everything from the start, a deployment that worked yesterday is blocked today and the platform is remembered as the cause of an outage. So in practice the order is "enforce the level you can comply with now, announce the next level in advance with warnings, investigate the impact before raising, and give those who cannot comply an exception with an expiry date." This lab walks through that order by hand. The reason for covering conformance as well is the same. If you leave deprecated APIs behind, then one day the moment you upgrade the cluster, deployments stop altogether, and having the eyes to read that signal in advance is a basic skill for a platform team.

Steps

  1. Create the namespace platform-baseline and attach labels. They are pod-security.kubernetes.io/enforce=baseline, pod-security.kubernetes.io/enforce-version=v1.30, pod-security.kubernetes.io/warn=restricted, pod-security.kubernetes.io/warn-version=v1.30, and platform.labhub.io/owner=platform-team.
  2. Write the Pod bad-probe in the platform-baseline namespace in /root/cnpa-base/bad-probe.yaml. It has spec.hostPID: true and a container with the name probe and the image ghcr.io/labhub/probe:1.0.0. Try to apply it, and save the failure output, including standard error, to /root/cnpa-base/reject.txt. The Pod bad-probe must not remain in the cluster.
  3. Write the Ingress orders-legacy (namespace platform-baseline) with apiVersion: extensions/v1beta1 in /root/cnpa-base/legacy.yaml. Try to apply it, and save the failure output to /root/cnpa-base/legacy-reject.txt.
  4. Write and apply the apps/v1 Deployment orders in /root/cnpa-base/app.yaml. It has spec.replicas: 2, the metadata labels app.kubernetes.io/name: orders and app.kubernetes.io/part-of: cnpa-platform, the container name app, the image ghcr.io/labhub/orders:1.4.0, and two container ports, named http (8080) and metrics (9102). In /root/cnpa-base/edge.yaml, write and apply the Service orders (selector app.kubernetes.io/name: orders, the port named http from 80 to 8080, and the port named metrics from 9102 to 9102) and the networking.k8s.io/v1 Ingress orders (host orders.labhub.internal, path /, pathType: Prefix, with the backend being the port named http of the Service orders).
  5. Write and apply the ServiceMonitor orders (namespace platform-baseline) in /root/cnpa-base/servicemonitor.yaml. spec.selector.matchLabels is app.kubernetes.io/name: orders, spec.endpoints[0].port is metrics, and interval is 30s.
  6. Write and apply the namespace platform-legacy in /root/cnpa-base/exception.yaml. The labels are pod-security.kubernetes.io/enforce: baseline and pod-security.kubernetes.io/enforce-version: v1.30, and the annotations are platform.labhub.io/exception-expires (an expiry date in YYYY-MM-DD format) and platform.labhub.io/exception-reason (the reason). Also put the Deployment legacy-batch (replicas 1, image ghcr.io/labhub/legacy-batch:0.9.0, no security settings) in the same file. Then send the label change that raises this namespace to restricted as a server dry-run, and save its output, including standard error, to /root/cnpa-base/upgrade-check.txt. Do not fix legacy-batch; leave it as it is.
  7. Bring the Pod template of orders in line with the restricted standard. At the Pod level put runAsNonRoot: true, runAsUser: 10001, and seccompProfile.type: RuntimeDefault, and at the container level put allowPrivilegeEscalation: false and capabilities.drop: [ALL]. After you apply it again and the Pods have started anew, raise the enforce label of platform-baseline to restricted. Leave enforce-version as v1.30.
  8. Record the current state in /root/cnpa-base/baseline.json. The keys are namespace, enforce, enforce_version, deployments (the number of Deployments in that namespace), servicemonitor (the name), exception_namespace, and exception_expires.

Notes

Baseline namespace

Create the namespace platform-baseline and attach labels. They are pod-security.kubernetes.io/enforce=baseline, pod-security.kubernetes.io/enforce-version=v1.30, pod-security.kubernetes.io/warn=restricted, pod-security.kubernetes.io/warn-version=v1.30, and platform.labhub.io/owner=platform-team.

Pod Security Admission works purely from namespace labels. Put the level you can comply with now on enforce, and the level you will raise to next on warn. Pin the version label together for each level.

Confirm that it really rejects

Write the Pod bad-probe in the platform-baseline namespace in /root/cnpa-base/bad-probe.yaml. It has spec.hostPID: true and a container with the name probe and the image ghcr.io/labhub/probe:1.0.0. Try to apply it, and save the failure output, including standard error, to /root/cnpa-base/reject.txt. The Pod bad-probe must not remain in the cluster.

Attaching the label does not mean the policy is running. Try to create a Pod that requests the host namespace, and keep the rejection message that appears in a file, including standard error.

Why a deprecated API is rejected

Write the Ingress orders-legacy (namespace platform-baseline) with apiVersion: extensions/v1beta1 in /root/cnpa-base/legacy.yaml. Try to apply it, and save the failure output to /root/cnpa-base/legacy-reject.txt.

A manifest written with a deprecated group fails as a mapping failure, not a syntax error. Keep that message as it is, and check with the cluster whether that group is really not served.

Move to the upper-level API and apply

Write and apply the apps/v1 Deployment orders in /root/cnpa-base/app.yaml. It has spec.replicas: 2, the metadata labels app.kubernetes.io/name: orders and app.kubernetes.io/part-of: cnpa-platform, the container name app, the image ghcr.io/labhub/orders:1.4.0, and two container ports, named http (8080) and metrics (9102). In /root/cnpa-base/edge.yaml, write and apply the Service orders (selector app.kubernetes.io/name: orders, the port named http from 80 to 8080, and the port named metrics from 9102 to 9102) and the networking.k8s.io/v1 Ingress orders (host orders.labhub.internal, path /, pathType: Prefix, with the backend being the port named http of the Service orders).

A Deployment is apps/v1 and an Ingress is networking.k8s.io/v1. A v1 Ingress requires a pathType for each path, and the backend points to the Service name and port separately. If you refer to ports by name instead of by number, they will not drift from the Service.

Establish the scrape contract

Write and apply the ServiceMonitor orders (namespace platform-baseline) in /root/cnpa-base/servicemonitor.yaml. spec.selector.matchLabels is app.kubernetes.io/name: orders, spec.endpoints[0].port is metrics, and interval is 30s.

A ServiceMonitor selects a Service by label, and it refers to a port not by number but by the Service's port name. After creating it, ask back with kubectl what that selector actually selected to confirm.

Exception and upgrade impact investigation

Write and apply the namespace platform-legacy in /root/cnpa-base/exception.yaml. The labels are pod-security.kubernetes.io/enforce: baseline and pod-security.kubernetes.io/enforce-version: v1.30, and the annotations are platform.labhub.io/exception-expires (an expiry date in YYYY-MM-DD format) and platform.labhub.io/exception-reason (the reason). Also put the Deployment legacy-batch (replicas 1, image ghcr.io/labhub/legacy-batch:0.9.0, no security settings) in the same file. Then send the label change that raises this namespace to restricted as a server dry-run, and save its output, including standard error, to /root/cnpa-base/upgrade-check.txt. Do not fix legacy-batch; leave it as it is.

If you send the label change as a server dry-run, warnings come back that evaluate the Pods currently running against the new level. Warnings come out on standard error, so save them together. Deliberately leave the workload in the exception namespace unfixed, as it is.

Bring the workload in line and raise the level

Bring the Pod template of orders in line with the restricted standard. At the Pod level put runAsNonRoot: true, runAsUser: 10001, and seccompProfile.type: RuntimeDefault, and at the container level put allowPrivilegeEscalation: false and capabilities.drop: [ALL]. After you apply it again and the Pods have started anew, raise the enforce label of platform-baseline to restricted. Leave enforce-version as v1.30.

Restricted requires four things: not running as root, the default seccomp profile, no privilege escalation, and dropping all capabilities. The first two go at the Pod level, and the last two at the container level. Fix the workload first and raise the level afterward.

Record the baseline status as values

Record the current state in /root/cnpa-base/baseline.json. The keys are namespace, enforce, enforce_version, deployments (the number of Deployments in that namespace), servicemonitor (the name), exception_namespace, and exception_expires.

Every number in the report must be something you can recount from the cluster. Fill them in by querying the namespace labels, the Deployment count, the ServiceMonitor name, and the exception expiry date each.