TT Lab
Get started
Learn Learning paths Courses

CNPE — Cloud Native Platform Engineer

What Does Platform Architecture Decide First

Continue in TT Lab

One-line summary

Design is not about picking tools. It is about deciding which failures you will tolerate and what you will guarantee. In particular, review the address space, data ownership, and the capacity that must remain during an outage before you review deployment tooling.

Why this was needed

The following is a practice scenario. You are replacing an order service with a new version, and one Pod stays in Pending. One owner wants to add nodes, and another wants to switch the canary to blue/green. But nobody has yet checked the Pod events or where the volume lives. If the cause is a node constraint on the volume, adding CPU will not fix it. If the problem is that two versions must not write the same data at the same time, changing how you switch traffic will not fix it either.

Start by writing down the design assumptions. Can the data be regenerated? May the new and old versions write at the same time? Must the service keep running if one node disappears? Without answers, adding more boxes to the architecture diagram will not produce a recovery procedure.

How it works

Address space and candidate nodes

For example, if you split a /16 for Pods into a /24 per node, the arithmetic gives 256 blocks. The real upper limit on node count depends on the CNI, reservations, and cluster settings, so do not write this figure down as the number of nodes you can operate. Also check whether Service, Pod, corporate network, and VPN addresses overlap. Depending on the implementation, changing the CIDR or the CNI may require a migration or a rebuild, so it is worth investigating early.

If a Pod is Pending, look at the candidates the scheduler allows and at the events before you look at total CPU. Not only nodeSelector but also affinity, taint/toleration, volume topology, and requests together narrow the candidates. Matching labels alone does not guarantee that the Pod can be placed.

kubectl get nodes --show-labels
kubectl -n team-a describe pod orders-canary
kubectl -n team-a get events --sort-by=.metadata.creationTimestamp
kubectl -n team-a get pvc

In the commands above, team-a and orders-canary are names used for illustration. In a real lab, replace them with your target names, and tell apart whether the source of an error event is the scheduler or the volume attach and mount process.

The "Once" in RWO does not mean one Pod

ReadWriteOnce (RWO) is an access mode that allows reading and writing from a single node. Several Pods on the same node can use the same volume. ReadWriteOncePod (RWOP), which allows only one Pod, is a separate mode, and you must check support conditions such as CSI. Even when RWX is possible, that does not make concurrent writes by the application safe. Read the access modes in the official PV documentation and mark the number of nodes, the number of Pods, and data consistency separately.

Observed situation What to check first What you cannot conclude yet
Two Pods reference an RWO PVC The nodes of both Pods and driver constraints The second Pod necessarily fails
A connection error occurs on another node Events, existing attachments, topology Adding only node CPU will fix it
Both are Ready on the same node The app's locking, single writer, and schema compatibility The data is safe under concurrent writes
You switched to blue/green Writes and background jobs of the preview version Switching traffic alone leaves a single writer

Blue/green is not data access control. Even if the Service sends no traffic, the new version's batch jobs can still write. You need to design a single writer in the application, replication and verification to a new volume, schema compatibility, and an explicit stop if necessary. Which choice is right depends on RPO, RTO, and the data model.

What it looks like in the field

The order team in the practice scenario needs two different questions. "Can the new Pod start?" is answered by scheduling and volume events. "Is it OK for the new Pod to start?" is answered by the concurrent-write contract and compatibility tests. Do not skip the second question because the first one passed.

In the design memo, keep four columns: decision, assumption, verification command, and fallback if it fails. For example, next to the decision "share RWO," write the assumption "can be placed on the same node" and the separate condition "there is one writer." This is the practice of turning terms memorized for the exam into operational judgment.

What to do in the next lesson

Next you will see how the quota you give a tenant differs from an actual capacity reservation. In the final quiz, you will judge cases where the conclusion changes depending on conditions even though the volume is RWO in all of them. The reading examples here alone do not let you claim that you have verified a CSI multi-node failure or data recovery.