TT Lab
Get started
Learn Learning paths Courses

CNPA — Cloud Native Platform Engineering Associate

Platform Maturity and the Components of an IDP

Continue in TT Lab

In one sentence

An internal developer platform (IDP) is not a single portal but a collection of five planes. And maturity rises not with the number of tools but with "who decides how, and what it is measured by."

Why this was needed

When someone says "let's adopt an IDP," the conversation usually begins with a meeting to choose a portal product. But a portal is only the surface of an IDP. Without a layer beneath it that actually builds and runs things, the portal becomes a collection of links. You have to draw which pieces are needed first in order to see what is missing.

How it works

The five planes of an IDP

Restated in working language, the breakdown laid out in the CNCF platforms white paper looks like this.

Plane What it does Common implementations
Developer Control Plane The surface developers touch Portal (Backstage), CLI, repository templates, manifest abstractions
Integration & Delivery Builds and deploys CI, image registry, GitOps agent, Secrets Operator
Monitoring & Logging Shows what is happening Metrics, logs, traces, dashboards, alerts
Security Identity and policy Authentication and authorization, policy engine, image signing and scanning
Resource Plane The actual resources Cluster, nodes, network, storage, database

Two axes cut across these. They are APIs and interfaces (how each plane is called) and capabilities. If the exam asks "Is a portal the IDP itself?", the answer is no; a portal is only one implementation of the Developer Control Plane.

Maturity levels

For solving exam questions, it helps to memorize the symptoms of each level rather than the numbers.

  1. Provisional — each team has its own scripts. Standing up one new service takes days. Knowledge lives in people's heads.
  2. Operational — common scripts and documentation appear. The platform team still takes tickets and does the work on others' behalf. The bottleneck is people.
  3. Scalable — self-service exists. Developers provision on their own without tickets. The platform team builds features instead of processing requests.
  4. Optimizing — the platform itself is improved using usage data and user feedback. Deprecation and migration happen in a planned way.

The real criterion that separates the levels is "is a ticket needed?" There is a cliff between level 2 and level 3, and most organizations stop here. If you have automated but only the platform team can press the run button, you are still at level 2.

What it takes for self-service to hold

Self-service is not just "granting permissions." Three things are needed at the same time.

The third one matters especially. In Kubernetes this is schema validation. If you put maximum: 10 in a CRD's OpenAPI schema, the API server rejects the request immediately. A policy engine's webhook plays the same role.

What it looks like in the field

What the author is building on top of the homelab is exactly an attempt to move toward level 3 — "a system where, when a learner presses a button, a lab Pod starts, gets graded, and disappears when finished." Here the platform requirements show up directly. A lab Pod must not be able to see other Pods (boundary), must not be able to use unlimited resources (guardrails), must already contain the needed tools without any configuration (safe defaults), and must disappear on its own when finished (lifecycle).

And the cluster underneath already has constraints like these. With 3 control plane nodes the etcd quorum is in place, but controlPlaneEndpoint is pinned to the first node's physical IP, so if that node dies, API access is cut off. The GPUs are 24GB, 32GB, and two 8GB cards, each in a different weight class, so requesting only nvidia.com/gpu: 1 lands on the wrong card. That is why a person had to define and attach semantic labels such as gpu.homelab/tier.

This is the starting point of platform API design. Do not expose the physical fact (the GPU card model) as it is; translate it into a vocabulary a developer can understand (tier=xlarge) — that is abstraction, and it is the topic of the next module.

What to check in the next quiz

In the next module, you will build a platform API called WebService with a CRD. You will complete, on a real cluster, a full set: the schema immediately rejects bad values, the server fills in defaults, namespaces, ResourceQuota, and LimitRange draw the tenant boundaries, and RBAC grants self-service permissions.