CNPE — Cloud Native Platform Engineer
CNPE Mock Exam A
Goal
You solve 17 tasks within 120 minutes under the same conditions as the real CNPE. The passing score is 64%, and because it uses partial scoring, passing 11 of the 17 is treated as complete.
This is a practice exam. Do not look at the hints or the answer sheet; try to solve everything through to the end first. It is better to mark the tasks where you get stuck, move on, and come back with the remaining time. You can press grading at any time, and pressing it several times does not change the result.
Why it matters
CNPE is an expert-level certification, and what it asks about is judgment, not operation. Each task contains one place where you decide whether to block the same rule in the pipeline or at the API server, whether to delete a constraint or satisfy it when an incident happens, and what to open for a tenant and what to keep closed. Several tools appear, but none of them is asked about in depth. Instead it asks which tool to put where.
Exam environment (facts confirmed in the real exam)
- The documents you can view are
kubernetes.io/docs,kubernetes.io/blog, and the Quick Reference link given separately for each task. That link is added to the allowlist. - Inside the remote desktop, you use a terminal and web interfaces together. Depending on the task, opening a console in the browser can be faster.
- The tools that can appear in the exam are Argo, Crossplane, Flagger, Flux, Gatekeeper, Grafana, Istio, Jaeger, Kyverno, Linkerd, OPA, OpenCost, OpenTelemetry, Prometheus, and Tekton.
- The official curriculum states that even if an unfamiliar tool appears, you must be able to handle it by looking at the
documentation provided during the exam. So rather than memorizing the tool list, the skill of reading
the schema of a CRD you are seeing for the first time with
kubectl explainmatters more. - Copying in the terminal is
Ctrl+Shift+C, and pasting isCtrl+Shift+V.
What is different in this practice exam environment
This lab's cluster is a single-user cluster running inside a Pod. The kube-apiserver, the controller manager, and the scheduler are real, so schema validation, admission policy, RBAC decisions, quotas, scheduling, aggregated ClusterRoles, and the eviction API actually work and actually reject. However, there is no runtime that actually runs containers, so the following differ.
- Argo CD and Argo Rollouts have only their CRDs registered and no controllers. The AppProject in task 3 is evaluated only as a file.
exec,logs, andport-forwarddo not work for workloads. The Pods come up but no process runs inside them.- There is no Prometheus server. Tasks 9 and 10 are evaluated with
promtooland CRDs. No actual scraping takes place.
There are 3 nodes, each with 8 CPU cores and 32Gi of memory, and each is in a different zone: zone-0, zone-1, and
zone-2.
Steps
GitOps and Continuous Delivery
- Create a Helm chart in
/root/exam/chart. The chart name ispaved-app, and the template is a single Deployment whose name ispaved-app. The container name isapp, the replica count is.Values.replicas, and the image is.Values.image.repositoryand.Values.image.digestjoined with@. Withvalues.schema.json, allow only an integer from 2 to 10 forreplicas, one ofbronze,silver, orgoldfortier, and only a string ofsha256:followed by 64 hexadecimal digits forimage.digest, and make all three values required. The defaults are 3,silver, and an arbitrary valid digest, respectively. - Create a kustomize directory in
/root/exam/portal. The Deploymentportalhas 2 replicas, the container nameapp, and an image pinned by digest, and it reads the ConfigMapportal-configin its entirety withenvFrom. Do not keep that ConfigMap as a file; create it withconfigMapGenerator, containing the two entriesLOG_LEVEL=infoandFEATURE_FLAGS=beta, and do not turn off the name hash suffix. In the render result, the name the Deployment references and the name of the generated ConfigMap must be the same. - In
/root/exam/gitops/appproject.yaml, write the Argo CD AppProjectplatform. The namespace isargocd. InsourceRepos, write the real repository address and do not use*. Fordestinations, pin both the server and the namespace, and do not use*for the namespace. LeaveclusterResourceWhitelistempty, and put the core group'sResourceQuotaandLimitRangeinnamespaceResourceBlacklist. - Make
/root/exam/promotea git repository. It has the two filesenvs/staging/deployment.yamlandenvs/prod/deployment.yaml, both are the Deploymentledger, the container name isapp, and the image is pinned by digest. In the first commit the two digests are different from each other. Then stack one promotion commit that raises the staging digest to production. In that commit, the staging file must not change, the commit message must contain the digest that was promoted, and the working tree must be clean.
Platform APIs and Self-Service Capabilities
- Write a CRD in
/root/exam/api/environment-crd.yamland apply it. The group isplatform.labhub.io, the kind isEnvironment, the plural isenvironments, the short name isenv, and the scope is Namespaced. There are two versions.v1is both served and the storage version, andspec.owneris a required string,spec.tieris one ofbronze,silver, orgoldwith a default ofbronze, andspec.retentionDaysis an integer from 1 to 90 with a default of7.v1alpha1is served but not the storage version and has onlyspec.owner; mark it deprecated and write the deprecation warning text yourself. The server must reject fields that are not in the schema. - In
/root/exam/api/env-defaults.yaml, write the Kyverno ClusterPolicyenv-defaults. It is a mutate rule that applies only to Environment resources, filling the labelplatform.labhub.io/ownerwith the value ofspec.ownerand the annotationplatform.labhub.io/requested-tierwith the value ofspec.tier. Do not write the values as fixed strings; read them from the request to fill them in. - In
/root/exam/api/env-rbac.yaml, write the aggregated permissions and apply them. The ClusterRoleplatform-env-authorhas no rules of its own and collects the ClusterRoles that carry the labelplatform.labhub.io/aggregate-to-env-author: "true". Attach that label to the ClusterRoleplatform-env-author-baseand allow get, list, watch, create, update, and patch on Environment, but do not give delete. Create the namespacetenant-amberand the ServiceAccountenv-author, and bind that aggregated ClusterRole with the RoleBindingenv-author. Do not use*in any rule. - In
/root/exam/api/env-status-rbac.yaml, separate the permissions for spec and status. Intenant-amber, create the ServiceAccountenv-controller, the Roleenv-controller, and the RoleBindingenv-controller. This account can read Environment and update and patchenvironments/status, but cannot create, update, patch, or delete the Environment itself. Conversely, theenv-authorfrom task 7 must not be able to updateenvironments/status.
Observability and Operations
- In
/root/exam/ops/platform-rules.yaml, write the Prometheus rules. The group name isplatform-api, the recording ruleplatform:request_error:ratio5muses division and a 5-minute range, and the alert rulePlatformApiErrorBudgetBurnmust havefor,labels.severity,annotations.summary, andannotations.runbook_url. And in/root/exam/ops/platform-rules-test.yaml, write the rule unit test. Inrule_files, write only the relative nameplatform-rules.yaml, and have two or morealert_rule_testentries, one of which checks a time when the alert fires and the other a time when it does not fire (exp_alerts: []). Bothpromtool check rulesandpromtool test rulesmust pass. - Create the namespace
platform-systemand, in/root/exam/ops/servicemonitor.yaml, write the ServiceMonitorplatform-apiand apply it. The target selector isapp: platform-api, and thenamespaceSelectorisplatform-system. There is one endpoint, the port name ismetrics, the scrape interval is30s, and the scrape timeout must be shorter than the interval. InmetricRelabelings, there must be at least one rule that looks at__name__anddrops a specific metric, and at least one rule that drops a high-cardinality label withlabeldrop. - Create the namespace
tenant-ochre, attach the labelplatform.labhub.io/tenant: ochre, and limit it with the ResourceQuotaochre-quotatorequests.cpuof 2 andrequests.memoryof 4Gi. Inside it, create the Deploymentsearchwith 4 replicas. The container name isapp, the image is pinned by digest, and the requests are CPU 700m and memory 512Mi. At this point not all the Pods can start. Then, in/root/exam/ops/triage.txt, write four lines,quota_cpu_hard,quota_memory_hard,desired_replicas, andmax_cpu_per_pod, in the키=값format (key=value). Write CPU as an integer in millicores, memory as an integer in Mi, and for the last value, write as a millicore integer the maximum CPU request one Pod may have for all the replicas to start within this quota. Do not attach unit characters. - Without raising the quota, make all 4 Pods of
searchRunning. Leave the replica count and the memory request as they are, and lower only the CPU request to the value you computed in task 11. No Pod that uses the old request value may remain.
Platform Architecture and Infrastructure
- Put the label
platform.labhub.io/pool=platformand the taintplatform.labhub.io/pool=platform:NoScheduletogether on exactly two nodes. Then create the Deploymentportalinplatform-systemwith 3 replicas. The container name isapp, the image is pinned by digest, and the requests are CPU 100m and memory 128Mi. This workload must select only that pool, tolerate that taint, and spread based ontopology.kubernetes.io/zonewithmaxSkew: 1andwhenUnsatisfiable: DoNotSchedule. All 3 Pods must be Running on pool nodes. - Create the PriorityClasses
platform-critical(value 100000 or more) andtenant-batch(value 1000 or less,preemptionPolicy: Never). For both,globalDefaultis false. Apply the ResourceQuotaamber-prioritytotenant-amberso that Pods that use theplatform-criticalpriority cannot be created in that namespace at all. And make theportalfrom task 13 useplatform-critical. - In
platform-system, create the PodDisruptionBudgetportal. The selector selects theportalPods,minAvailableis 2, and do not usemaxUnavailable. It must be possible to drain only one node at a time during maintenance.
Security and Policy Enforcement
- Create the namespace
tenant-tealand apply Pod Security admissionenforce,audit, andwarnall asrestricted, but pin the version labels of the three modes tov1.31, not tolatest. Inside it, create the Deploymentcheckoutwith 2 replicas. The container name isapp, the image is pinned by digest, and the Pods must passrestricted. 2 Pods must actually be Running. - Enforce one rule in two places. The rule is "a Deployment must have the label
platform.labhub.io/owner." In/root/exam/security/owner-policy.yaml, write the Kyverno ClusterPolicyrequire-ownerwithEnforce. It is the gate the pipeline will run withkyverno apply. In/root/exam/security/owner-vap.yaml, write the ValidatingAdmissionPolicyrequire-ownerand a binding with the same name, and apply it.validationActionsisDeny, and the scope of application is only the namespaces that carry the labelplatform.labhub.io/gate=owner. Create the namespacedeliveryand attach that label. A Deployment without the label must be rejected indeliveryand pass indefault.
Reference
- To see what the server actually does, use
kubectl apply --dry-run=server. It runs through the admission chain as is while leaving nothing in the cluster. If you add-o json, you can also see the result with the defaults filled in. - Ask about subresource permissions with an option, as in
kubectl auth can-i update environments.platform.labhub.io --subresource=status처럼--subresource(do not append the subresource after the resource name). If you write a slash after the resource name, you get a wrong decision. - Right after you apply a CRD, wait for the registration to finish with
kubectl wait --for=condition=Established crd/.... - If you start a rolling update while the quota is full, there is no room to create new Pods and the rollout stalls. The old Pods must be emptied first for new Pods to get in.
- Three common mistakes. Do not write
sourceLabelsin alabeldroprule. If you create only a ClusterRole without a RoleBinding, no permissions arise. And therulesof an aggregated ClusterRole is not a place for people to fill in but a place the controller fills in.
Make the chart reject wrong values by itself
If there is a values.schema.json inside the chart, helm validates the values before template and install, and if they do not conform, it rejects the render itself. Use the JSON schema's minimum, maximum, enum, pattern, and required. You confirm that the schema takes effect by deliberately passing a nonconforming value, as in helm template ... --set replicas=1.
Make the deployment change when the configuration changes
configMapGenerator appends a hash of the content to the name and automatically rewrites the places that reference that name within the same kustomization. In the Deployment, it is enough to write the original name without the suffix. If you turn on disableNameSuffixHash, this link breaks, and Pods do not restart when only the configuration changes.
Pin down what the project allows
An AppProject is the fence that decides what an Application may deploy and where. If you leave asterisks in sourceRepos and destinations, it is the same as having no fence. If you leave clusterResourceWhitelist empty, cluster-scoped resources cannot be created at all, and namespaceResourceBlacklist is the place to list the kinds that must not be touched even inside a namespace.
Leave the promotion as a single commit
In GitOps, a deployment is a commit. So what went up and when must remain in the log, and a rollback must be a revert. Move the digest you verified in staging to production as is, but do not touch the staging file. When you are done, git status must be empty. A change you have not committed is not deployed.
Make the schema a contract
If you write a default in a structural schema, the server fills in the value, and the server rejects fields that are not written. Only when the two properties come together does the schema become a contract. Mark deprecation with deprecated and deprecationWarning in the versions entry, and that text actually appears as a Warning on the kubectl screen. Confirm with kubectl apply --dry-run=server -o json.
Have the platform fill in what the request is missing
The less a user has to write in self-service, the better. What can be derived from values the user already wrote must be filled in by the platform. In a Kyverno mutate rule, you can read the request body with {{ request.object... }}. If you write the values as fixed strings, it is exposed when the grader asks with a different request. Confirm with kyverno apply --resource -o .
Collect permissions in pieces
An aggregated ClusterRole does not write its own rules. If you write only a label selector, the controller collects and fills in the rules of the other ClusterRoles that carry that label. So when you increase permissions later, you do not edit this role; you just attach one more piece. Confirm that the aggregation actually happened with kubectl get clusterrole -o jsonpath='{.rules}'. If it is empty, the label is mismatched.
Separate who writes from who fills in
status is the place the controller fills in, and spec is the place the user writes. In RBAC these two are separate resource names, so you write environments and environments/status separately. When checking, do not append a slash after the resource name; use --subresource=status. If you write it with a slash, kubectl interprets it wrongly and gives the opposite result.
Pin down the alert rule with a test
promtool test rules feeds in fake time series and actually evaluates the rules. Write the values of input_series as a start value, an increment, and a repeat count, like '0+3x30'. If you check only the time it fires, even a rule that always fires passes, so also write the time when it must not fire, with exp_alerts: []. rule_files is looked up relative to the directory where the test file is.
Decide what not to measure
Observability cost is determined by the number of time series collected, and that number explodes with combinations of labels. So a scrape configuration must include what to drop as much as what to measure. drop discards the entire metric if the value picked by sourceLabels matches the regex, and labeldrop deletes label names that match the regex. You must not write sourceLabels in labeldrop.
Reproduce quota exhaustion and write it down as numbers
When a quota blocks, the Deployment does not report an error. The ReplicaSet carries a ReplicaFailure condition instead, and the Pods are not created in the first place. Look at that condition with kubectl describe rs or kubectl get rs -o yaml. Classification is not a feeling but a division. Dividing the quota cap by the desired replica count gives the maximum a single Pod can use.
Recover without raising the quota
Lowering the request alone is not enough. When the quota is full, there is no room to create new Pods, so the rolling update stops in the middle. As long as Pods that use the old request value remain, the room does not open up. Scale the replicas down to 0 and back up, or empty the old Pods first. Confirm by the request value of each Pod, not by the number of Pods.
Divide the pool and spread the workload
A label attracts and a taint repels. You must apply both together for that pool to be dedicated. With only the label, anyone can come in, and with only the taint, there is no way in. The labelSelector of topologySpreadConstraints must point to the Pod labels, and it is actually enforced only if you set DoNotSchedule. Confirm by which node each Pod sits on.
Decide who may use a priority
A PriorityClass is cluster-scoped, so once you create it, anyone can take it and use it by name. So you must block separately not the priority itself but the right to use that priority. If you apply the PriorityClass scope in a ResourceQuota's scopeSelector and set the cap to 0, Pods of that class cannot be created at all in that namespace. preemptionPolicy: Never does not mean it is not itself pushed out; it means it does not push others out.
Decide how many must remain during maintenance
A PodDisruptionBudget is looked at by the eviction API, not the scheduler. kubectl drain and node upgrades go through that path. With 3 replicas and minAvailable 2, you can drain only one at a time. The result the controller computed appears in the ALLOWED DISRUPTIONS column of kubectl get pdb, and if this value is 0, maintenance is blocked entirely, and if it equals the replica count, it protects nothing.
Apply a baseline to the tenant namespace
Pod security admission does not block the Deployment but blocks the Pods that the Deployment creates. If the Deployment was created but there are no Pods, look at the reason with kubectl describe rs. restricted requires runAsNonRoot and seccompProfile at the Pod level and allowPrivilegeEscalation: false and capabilities.drop [ALL] at the container level. If you leave the version label at latest, the baseline silently becomes stricter when you upgrade the cluster.
Enforce the same rule in two places
A pipeline gate gives quick feedback but cannot block changes that did not go through the pipeline, and an admission policy blocks whatever comes in but the developer learns of it only after pressing deploy. So you put the same rule in two places. The scope of application of a VAP is decided by the binding and can be narrowed with the namespaceSelector in matchResources. If validationActions has no Deny, it leaves only an audit record and lets it through.