TT Lab
Get started
Learn Learning paths Courses

CNPA — Cloud Native Platform Engineering Associate

Why the Kubernetes API Became the Platform's Shared Language

Continue in TT Lab

In one sentence

What Kubernetes won was not the container orchestration race but the API specification race. The pairing of declarative resources plus reconciling controllers spread beyond containers, which is why CRDs became the default choice when building a new platform API.

Why this was needed

When a platform team tries to make "spinning up a service" easy, it usually walks this path.

  1. Write the procedure on a wiki → nobody keeps it up to date.
  2. Write a shell script → results differ across execution environments, and a failure leaves an intermediate state behind.
  3. Build an internal web app → you have to implement state storage, authentication, auditing, retries, and concurrency all by yourself. And that app becomes a new SPOF.

If you build option 3 all the way through, you eventually realize what you are rebuilding. A state store, optimistic concurrency control, authentication and authorization, audit logs, watch, and a reconciliation loop — the Kubernetes API server already has all of these.

So the direction flips. Instead of building a new platform API, you extend the Kubernetes API. Here is everything that comes for free the moment you register a CRD.

This is the substance of the claim that "the Kubernetes API is the common language of platforms." You define a new vocabulary (the CRD) but use a grammar (the API conventions) that everyone already knows.

How it works

CRD + controller = platform API

You need two pieces.

The CRD defines the vocabulary: group, version, kind, scope (Namespaced or Cluster), and an OpenAPI v3 schema. The schema does more than you might expect.

spec:
  versions:
    - name: v1alpha1
      served: true
      storage: true
      schema:
        openAPIV3Schema:
          type: object
          properties:
            spec:
              type: object
              required: [image]
              properties:
                image:    { type: string }
                replicas: { type: integer, default: 2, minimum: 1, maximum: 10 }
                public:   { type: boolean, default: false }

The controller gives that vocabulary its meaning. It is a reconciliation loop that looks at a WebService CR and creates a Deployment, Service, HPA, and NetworkPolicy. With a CRD but no controller, the CR stays a "configuration file whose structure is validated" — that has its uses too, but it is not a platform API.

Scope selection is also an exam favorite. A Namespaced resource is confined within a tenant boundary, and namespace RBAC applies as is. A Cluster-scoped resource has a global name, so names collide between tenants, and permissions can only be granted with a ClusterRole, not a Role. A resource that tenants create should almost always be Namespaced.

The moment the abstraction leaks

A good platform API "asks only for what it needs." WebService asks only for the image and, if needed, replicas, and the controller fills in the rest — label conventions, the security context, resource defaults, observability annotations, and network policies.

But the day will surely come when it leaks.

If you answer "that is not supported" at this point, that team abandons the platform and goes back to raw YAML. Once they leave, they do not come back. So you must build an escape hatch into the design in advance.

Escape hatch Form Risk
Partial override A free-form field such as spec.podOverrides If anything can go in, the abstraction becomes meaningless
Extension points Allow only extraEnv, extraVolumes, and nodeSelector Cost of maintaining the list
Leaving after render Copy the generated manifest and manage it directly You no longer receive later platform improvements

The balance point is this: an escape hatch must exist, but its use must be visible. If you leave a label or a status condition on services that use an override, the platform team receives the signal that "five teams are working around this feature with overrides" and can promote it to an official feature. This is the feedback loop that runs a platform as a product.

What it looks like in the field

In the author's homelab, this problem showed up in exactly this form. GPU Feature Discovery automatically labels nodes with card information — an RTX 3090 (24576MB, ampere), a 5090 (32607MB, blackwell), and two 4070 Laptops (8188MB, ada-lovelace). But if a Pod requests only nvidia.com/gpu: 1, a training job that needs 32GB can land on an 8GB laptop GPU. That is because to Kubernetes, both are "one GPU."

The gpu.memory label that GFD attaches is a string, so a comparison selector such as "24GB or more" is impossible. So semantic labels were added by hand — gpu.homelab/tier=xlarge|large|small and gpu.homelab/vram=32g|24g|8g. Now a workload chooses its own weight class with nodeSelector: gpu.homelab/tier: xlarge.

This one line is the archetype of platform API design. It did not expose the underlying physical facts (card model names, memory bytes) as they are, but translated them into a vocabulary that users can make decisions with. At the same time, an escape hatch remains — if you really need a specific card, you can select directly with the original GFD labels. A good abstraction does not hide the layer below; it covers it but leaves it open.

What to do in the next lab

You will create the CRD webservices.platform.labhub.io on a real cluster and see for yourself that a schema violation is rejected immediately and that defaults are filled in by the server. Next you will draw the tenant boundary with a namespace + ResourceQuota + LimitRange, grant self-service permissions with RBAC, and then prove with kubectl auth can-i that nothing in the other tenant can be touched.