CNPA — Cloud Native Platform Engineering Associate
Why the Kubernetes API Became the Platform's Shared Language
In one sentence
What Kubernetes won was not the container orchestration race but the API specification race. The pairing of declarative resources plus reconciling controllers spread beyond containers, which is why CRDs became the default choice when building a new platform API.
Why this was needed
When a platform team tries to make "spinning up a service" easy, it usually walks this path.
- Write the procedure on a wiki → nobody keeps it up to date.
- Write a shell script → results differ across execution environments, and a failure leaves an intermediate state behind.
- Build an internal web app → you have to implement state storage, authentication, auditing, retries, and concurrency all by yourself. And that app becomes a new SPOF.
If you build option 3 all the way through, you eventually realize what you are rebuilding. A state store, optimistic concurrency control, authentication and authorization, audit logs, watch, and a reconciliation loop — the Kubernetes API server already has all of these.
So the direction flips. Instead of building a new platform API, you extend the Kubernetes API. Here is everything that comes for free the moment you register a CRD.
- It is stored in etcd, with optimistic locking through the version and
resourceVersion - The existing RBAC applies as is (
kubectl auth can-i create webservicesworks right away) - It appears in the audit log
kubectl get/describe/edit,-o yaml, and--watchjust work- The OpenAPI schema rejects bad values immediately and fills in defaults
- GitOps tools handle it exactly like any other resource
This is the substance of the claim that "the Kubernetes API is the common language of platforms." You define a new vocabulary (the CRD) but use a grammar (the API conventions) that everyone already knows.
How it works
CRD + controller = platform API
You need two pieces.
The CRD defines the vocabulary: group, version, kind, scope (Namespaced or Cluster), and an OpenAPI v3 schema. The schema does more than you might expect.
spec:
versions:
- name: v1alpha1
served: true
storage: true
schema:
openAPIV3Schema:
type: object
properties:
spec:
type: object
required: [image]
properties:
image: { type: string }
replicas: { type: integer, default: 2, minimum: 1, maximum: 10 }
public: { type: boolean, default: false }
requiredimmediately rejects a request that omits a field.minimum/maximumenforces a range. If you sendreplicas: 20, the API server rejects it with a sentence a person can read.defaultis filled in by the server. If the user does not write it, the value is inserted at storage time. This is the cheapest implementation of "safe defaults."additionalPrinterColumnsshows the columns you want in thekubectl getoutput. It looks small, but it contributes a lot to developer experience.
The controller gives that vocabulary its meaning. It is a reconciliation loop that looks at a WebService CR and creates a Deployment, Service, HPA, and NetworkPolicy. With a CRD but no controller, the CR stays a "configuration file whose structure is validated" — that has its uses too, but it is not a platform API.
Scope selection is also an exam favorite. A Namespaced resource is confined within a tenant boundary, and namespace RBAC applies as is. A Cluster-scoped resource has a global name, so names collide between tenants, and permissions can only be granted with a ClusterRole, not a Role. A resource that tenants create should almost always be Namespaced.
The moment the abstraction leaks
A good platform API "asks only for what it needs." WebService asks only for the image and, if needed, replicas, and the controller fills in the rest — label conventions, the security context, resource defaults, observability annotations, and network policies.
But the day will surely come when it leaks.
- A team needs to attach a sidecar.
- A workload must run only on a specific node (for example, a 32GB VRAM GPU).
- A service uses a probe path different from the standard.
If you answer "that is not supported" at this point, that team abandons the platform and goes back to raw YAML. Once they leave, they do not come back. So you must build an escape hatch into the design in advance.
| Escape hatch | Form | Risk |
|---|---|---|
| Partial override | A free-form field such as spec.podOverrides |
If anything can go in, the abstraction becomes meaningless |
| Extension points | Allow only extraEnv, extraVolumes, and nodeSelector |
Cost of maintaining the list |
| Leaving after render | Copy the generated manifest and manage it directly | You no longer receive later platform improvements |
The balance point is this: an escape hatch must exist, but its use must be visible. If you leave a label or a status condition on services that use an override, the platform team receives the signal that "five teams are working around this feature with overrides" and can promote it to an official feature. This is the feedback loop that runs a platform as a product.
What it looks like in the field
In the author's homelab, this problem showed up in exactly this form. GPU Feature Discovery automatically labels nodes with card information — an RTX 3090 (24576MB, ampere), a 5090 (32607MB, blackwell), and two 4070 Laptops (8188MB, ada-lovelace). But if a Pod requests only nvidia.com/gpu: 1, a training job that needs 32GB can land on an 8GB laptop GPU. That is because to Kubernetes, both are "one GPU."
The gpu.memory label that GFD attaches is a string, so a comparison selector such as "24GB or more" is impossible. So semantic labels were added by hand — gpu.homelab/tier=xlarge|large|small and gpu.homelab/vram=32g|24g|8g. Now a workload chooses its own weight class with nodeSelector: gpu.homelab/tier: xlarge.
This one line is the archetype of platform API design. It did not expose the underlying physical facts (card model names, memory bytes) as they are, but translated them into a vocabulary that users can make decisions with. At the same time, an escape hatch remains — if you really need a specific card, you can select directly with the original GFD labels. A good abstraction does not hide the layer below; it covers it but leaves it open.
What to do in the next lab
You will create the CRD webservices.platform.labhub.io on a real cluster and see for yourself that a schema violation is rejected immediately and that defaults are filled in by the server. Next you will draw the tenant boundary with a namespace + ResourceQuota + LimitRange, grant self-service permissions with RBAC, and then prove with kubectl auth can-i that nothing in the other tenant can be touched.