CNPE — Cloud Native Platform Engineer
What Does Platform Architecture Decide First
One-line summary
Design is not about picking tools. It is about deciding which failures you will tolerate and what you will guarantee. In particular, review the address space, data ownership, and the capacity that must remain during an outage before you review deployment tooling.
Why this was needed
The following is a practice scenario. You are replacing an order service with a new version, and one Pod stays in Pending. One owner wants to add nodes, and another wants to switch the canary to blue/green. But nobody has yet checked the Pod events or where the volume lives. If the cause is a node constraint on the volume, adding CPU will not fix it. If the problem is that two versions must not write the same data at the same time, changing how you switch traffic will not fix it either.
Start by writing down the design assumptions. Can the data be regenerated? May the new and old versions write at the same time? Must the service keep running if one node disappears? Without answers, adding more boxes to the architecture diagram will not produce a recovery procedure.
How it works
Address space and candidate nodes
For example, if you split a /16 for Pods into a /24 per node, the arithmetic gives 256 blocks. The real upper limit on node count depends on the CNI, reservations, and cluster settings, so do not write this figure down as the number of nodes you can operate. Also check whether Service, Pod, corporate network, and VPN addresses overlap. Depending on the implementation, changing the CIDR or the CNI may require a migration or a rebuild, so it is worth investigating early.
If a Pod is Pending, look at the candidates the scheduler allows and at the events before you look at total CPU. Not only nodeSelector but also affinity, taint/toleration, volume topology, and requests together narrow the candidates. Matching labels alone does not guarantee that the Pod can be placed.
kubectl get nodes --show-labels
kubectl -n team-a describe pod orders-canary
kubectl -n team-a get events --sort-by=.metadata.creationTimestamp
kubectl -n team-a get pvc
In the commands above, team-a and orders-canary are names used for illustration. In a real lab, replace them with your target names, and tell apart whether the source of an error event is the scheduler or the volume attach and mount process.
The "Once" in RWO does not mean one Pod
ReadWriteOnce (RWO) is an access mode that allows reading and writing from a single node. Several Pods on the same node can use the same volume. ReadWriteOncePod (RWOP), which allows only one Pod, is a separate mode, and you must check support conditions such as CSI. Even when RWX is possible, that does not make concurrent writes by the application safe. Read the access modes in the official PV documentation and mark the number of nodes, the number of Pods, and data consistency separately.
| Observed situation | What to check first | What you cannot conclude yet |
|---|---|---|
| Two Pods reference an RWO PVC | The nodes of both Pods and driver constraints | The second Pod necessarily fails |
| A connection error occurs on another node | Events, existing attachments, topology | Adding only node CPU will fix it |
| Both are Ready on the same node | The app's locking, single writer, and schema compatibility | The data is safe under concurrent writes |
| You switched to blue/green | Writes and background jobs of the preview version | Switching traffic alone leaves a single writer |
Blue/green is not data access control. Even if the Service sends no traffic, the new version's batch jobs can still write. You need to design a single writer in the application, replication and verification to a new volume, schema compatibility, and an explicit stop if necessary. Which choice is right depends on RPO, RTO, and the data model.
What it looks like in the field
The order team in the practice scenario needs two different questions. "Can the new Pod start?" is answered by scheduling and volume events. "Is it OK for the new Pod to start?" is answered by the concurrent-write contract and compatibility tests. Do not skip the second question because the first one passed.
In the design memo, keep four columns: decision, assumption, verification command, and fallback if it fails. For example, next to the decision "share RWO," write the assumption "can be placed on the same node" and the separate condition "there is one writer." This is the practice of turning terms memorized for the exam into operational judgment.
What to do in the next lesson
Next you will see how the quota you give a tenant differs from an actual capacity reservation. In the final quiz, you will judge cases where the conclusion changes depending on conditions even though the volume is RWO in all of them. The reading examples here alone do not let you claim that you have verified a CSI multi-node failure or data recovery.