TT Lab
Get started
Learn Learning paths Courses

CNPA — Cloud Native Platform Engineering Associate

What It Means to See the Platform as a Product

Continue in TT Lab

In one sentence

The goal of platform engineering is not to assemble many tools but to reduce the cognitive load on stream-aligned teams. That is why a platform is a product rather than a project, its users are other development teams, and its success metric is not the number of deployments but the rate of voluntary adoption.

Why this was needed

As "You build it, you run it" became popular, many organizations pushed operational responsibility onto development teams. The intent was good, but the result was this: what a single developer has to know exploded — languages and frameworks, domain logic, and on top of that Kubernetes, Helm, Terraform, CI syntax, secrets management, the observability stack, network policies, and cost optimization.

Team Topologies calls this cognitive overload. There are three kinds of cognitive load.

Kind What it is Response
Intrinsic Fundamentals such as languages and data structures Through training and hiring
Extraneous The procedure of writing eight YAML files in the right order What the platform should eliminate
Germane The domain problem itself This is where people should spend their mental effort

Platform engineering absorbs extraneous load to create room for germane load. So the answer to "Is our platform doing well?" is not the number of tools but "how much time developers spend on domain problems".

How it works

The platform as a product

If you treat a platform as a project, it goes like this: you gather requirements, build for 6 months, deploy, and disband the team. Treating it as a product is different.

There is one critical test: are people using it even though nobody forced them to? A platform that reached 100% through a company mandate is measuring compliance, not adoption. If compliance is high and satisfaction is low, the platform is failing while the metrics hide it.

The four team types in Team Topologies

CNPA asks about this classification directly.

Team type What it does
Stream-aligned Owns a single flow of value (a product, feature, or user journey) end to end. It should make up the majority of the organization
Platform Provides the internal services that stream-aligned teams use through self-service
Enabling Temporarily fills capability gaps in other teams. It does not stay; it leaves
Complicated-subsystem Owns the parts that require deep specialist knowledge (video codecs, payment settlement, math engines)

There are also three interaction modes — Collaboration (working closely together for a limited time to discover), X-as-a-Service (a consumption relationship with clear boundaries), and Facilitating (teaching, then stepping back). The normal state for a platform team is X-as-a-Service. If a platform team is in permanent Collaboration with every stream-aligned team, that signals the platform is not yet self-service.

DevOps · SRE · platform engineering

The three are not competing concepts; they have different focuses.

There is a lot of overlap. A platform should also have an SLO (an SRE method), and a platform team should run its own platform itself (a DevOps principle). Where they diverge is "who you optimize for" — SRE optimizes the end user's experience, while a platform engineer optimizes the internal developer's experience.

What it looks like in the field

The author's 7-node homelab is a miniature of this story. Cilium 1.20 eBPF CNI, MetalLB L2 (10.0.0.200-215), the Harbor registry (10.0.0.202), Gitea (10.0.0.200), Argo CD (10.0.0.201), kube-prometheus-stack and Grafana (10.0.0.203), CloudNativePG, 4 GPUs via the GPU Operator, KubeVirt, csi-driver-nfs — listing the stack alone makes it look like an excellent platform.

But the list of incidents that came up while bringing up each component shows the real value of a platform. Gateway API needed CRD v1.6.1, and with v1.2 the controller refused to start. In KubeVirt every component status was AllComponentsReady, yet the VM would not start; the cause was a missing volume mount on the virt-launcher Pod. The containerDisk path had a defect, so the DataVolume (PVC) path had to be used as a workaround.

All of this knowledge is extraneous cognitive load. It is knowledge a developer building a service has no reason to need. The platform team's job is to step on each of these traps once, then package the result as "one golden path" so nobody else has to step on them again. And the record of confirming for the third time on the same cluster the lesson that "status is Ready" and "it actually works" are different claims also means that what a platform must provide is not an installation but a verified path.

What to read next

The next reading summarizes the platform maturity levels and the components of an internal developer platform (IDP). In the following module, you will build a Kubernetes API extension into a platform API on a real cluster, using CRDs, CRs, ResourceQuota, and RBAC.