CNPA — Cloud Native Platform Engineering Associate
What It Means to See the Platform as a Product
In one sentence
The goal of platform engineering is not to assemble many tools but to reduce the cognitive load on stream-aligned teams. That is why a platform is a product rather than a project, its users are other development teams, and its success metric is not the number of deployments but the rate of voluntary adoption.
Why this was needed
As "You build it, you run it" became popular, many organizations pushed operational responsibility onto development teams. The intent was good, but the result was this: what a single developer has to know exploded — languages and frameworks, domain logic, and on top of that Kubernetes, Helm, Terraform, CI syntax, secrets management, the observability stack, network policies, and cost optimization.
Team Topologies calls this cognitive overload. There are three kinds of cognitive load.
| Kind | What it is | Response |
|---|---|---|
| Intrinsic | Fundamentals such as languages and data structures | Through training and hiring |
| Extraneous | The procedure of writing eight YAML files in the right order | What the platform should eliminate |
| Germane | The domain problem itself | This is where people should spend their mental effort |
Platform engineering absorbs extraneous load to create room for germane load. So the answer to "Is our platform doing well?" is not the number of tools but "how much time developers spend on domain problems".
How it works
The platform as a product
If you treat a platform as a project, it goes like this: you gather requirements, build for 6 months, deploy, and disband the team. Treating it as a product is different.
- It has users. Those users are internal developers, and if they don't want to use it, they won't.
- It has a roadmap. It also decides what not to do.
- It has onboarding documentation and a support channel.
- It has versions and a deprecation policy. If something that worked yesterday breaks today, trust disappears.
- It has success metrics. Adoption rate, time to first deployment, number of support tickets.
There is one critical test: are people using it even though nobody forced them to? A platform that reached 100% through a company mandate is measuring compliance, not adoption. If compliance is high and satisfaction is low, the platform is failing while the metrics hide it.
The four team types in Team Topologies
CNPA asks about this classification directly.
| Team type | What it does |
|---|---|
| Stream-aligned | Owns a single flow of value (a product, feature, or user journey) end to end. It should make up the majority of the organization |
| Platform | Provides the internal services that stream-aligned teams use through self-service |
| Enabling | Temporarily fills capability gaps in other teams. It does not stay; it leaves |
| Complicated-subsystem | Owns the parts that require deep specialist knowledge (video codecs, payment settlement, math engines) |
There are also three interaction modes — Collaboration (working closely together for a limited time to discover), X-as-a-Service (a consumption relationship with clear boundaries), and Facilitating (teaching, then stepping back). The normal state for a platform team is X-as-a-Service. If a platform team is in permanent Collaboration with every stream-aligned team, that signals the platform is not yet self-service.
DevOps · SRE · platform engineering
The three are not competing concepts; they have different focuses.
- DevOps is a culture and a set of principles — tear down the wall between development and operations. It is not a team name. If you create a "DevOps team" and it starts deploying on others' behalf, you have rebuilt the very wall you meant to remove under a different name.
- SRE is a concrete practice that treats reliability as an engineering problem — SLIs/SLOs, error budgets, toil reduction, blameless postmortems. Its target is service reliability.
- Platform engineering is the work of packaging those practices as a self-service product that many teams can adopt. Its target is developer experience.
There is a lot of overlap. A platform should also have an SLO (an SRE method), and a platform team should run its own platform itself (a DevOps principle). Where they diverge is "who you optimize for" — SRE optimizes the end user's experience, while a platform engineer optimizes the internal developer's experience.
What it looks like in the field
The author's 7-node homelab is a miniature of this story. Cilium 1.20 eBPF CNI, MetalLB L2 (10.0.0.200-215), the Harbor registry (10.0.0.202), Gitea (10.0.0.200), Argo CD (10.0.0.201), kube-prometheus-stack and Grafana (10.0.0.203), CloudNativePG, 4 GPUs via the GPU Operator, KubeVirt, csi-driver-nfs — listing the stack alone makes it look like an excellent platform.
But the list of incidents that came up while bringing up each component shows the real value of a platform. Gateway API needed CRD v1.6.1, and with v1.2 the controller refused to start. In KubeVirt every component status was AllComponentsReady, yet the VM would not start; the cause was a missing volume mount on the virt-launcher Pod. The containerDisk path had a defect, so the DataVolume (PVC) path had to be used as a workaround.
All of this knowledge is extraneous cognitive load. It is knowledge a developer building a service has no reason to need. The platform team's job is to step on each of these traps once, then package the result as "one golden path" so nobody else has to step on them again. And the record of confirming for the third time on the same cluster the lesson that "status is Ready" and "it actually works" are different claims also means that what a platform must provide is not an installation but a verified path.
What to read next
The next reading summarizes the platform maturity levels and the components of an internal developer platform (IDP). In the following module, you will build a Kubernetes API extension into a platform API on a real cluster, using CRDs, CRs, ResourceQuota, and RBAC.