TT Lab
Get started
Learn Learning paths Courses

CNPE — Cloud Native Platform Engineer

CNPE Mock Exam A

Continue in TT Lab

Goal

You solve 17 tasks within 120 minutes under the same conditions as the real CNPE. The passing score is 64%, and because it uses partial scoring, passing 11 of the 17 is treated as complete.

This is a practice exam. Do not look at the hints or the answer sheet; try to solve everything through to the end first. It is better to mark the tasks where you get stuck, move on, and come back with the remaining time. You can press grading at any time, and pressing it several times does not change the result.

Why it matters

CNPE is an expert-level certification, and what it asks about is judgment, not operation. Each task contains one place where you decide whether to block the same rule in the pipeline or at the API server, whether to delete a constraint or satisfy it when an incident happens, and what to open for a tenant and what to keep closed. Several tools appear, but none of them is asked about in depth. Instead it asks which tool to put where.

Exam environment (facts confirmed in the real exam)

What is different in this practice exam environment

This lab's cluster is a single-user cluster running inside a Pod. The kube-apiserver, the controller manager, and the scheduler are real, so schema validation, admission policy, RBAC decisions, quotas, scheduling, aggregated ClusterRoles, and the eviction API actually work and actually reject. However, there is no runtime that actually runs containers, so the following differ.

There are 3 nodes, each with 8 CPU cores and 32Gi of memory, and each is in a different zone: zone-0, zone-1, and zone-2.

Steps

GitOps and Continuous Delivery

  1. Create a Helm chart in /root/exam/chart. The chart name is paved-app, and the template is a single Deployment whose name is paved-app. The container name is app, the replica count is .Values.replicas, and the image is .Values.image.repository and .Values.image.digest joined with @. With values.schema.json, allow only an integer from 2 to 10 for replicas, one of bronze, silver, or gold for tier, and only a string of sha256: followed by 64 hexadecimal digits for image.digest, and make all three values required. The defaults are 3, silver, and an arbitrary valid digest, respectively.
  2. Create a kustomize directory in /root/exam/portal. The Deployment portal has 2 replicas, the container name app, and an image pinned by digest, and it reads the ConfigMap portal-config in its entirety with envFrom. Do not keep that ConfigMap as a file; create it with configMapGenerator, containing the two entries LOG_LEVEL=info and FEATURE_FLAGS=beta, and do not turn off the name hash suffix. In the render result, the name the Deployment references and the name of the generated ConfigMap must be the same.
  3. In /root/exam/gitops/appproject.yaml, write the Argo CD AppProject platform. The namespace is argocd. In sourceRepos, write the real repository address and do not use *. For destinations, pin both the server and the namespace, and do not use * for the namespace. Leave clusterResourceWhitelist empty, and put the core group's ResourceQuota and LimitRange in namespaceResourceBlacklist.
  4. Make /root/exam/promote a git repository. It has the two files envs/staging/deployment.yaml and envs/prod/deployment.yaml, both are the Deployment ledger, the container name is app, and the image is pinned by digest. In the first commit the two digests are different from each other. Then stack one promotion commit that raises the staging digest to production. In that commit, the staging file must not change, the commit message must contain the digest that was promoted, and the working tree must be clean.

Platform APIs and Self-Service Capabilities

  1. Write a CRD in /root/exam/api/environment-crd.yaml and apply it. The group is platform.labhub.io, the kind is Environment, the plural is environments, the short name is env, and the scope is Namespaced. There are two versions. v1 is both served and the storage version, and spec.owner is a required string, spec.tier is one of bronze, silver, or gold with a default of bronze, and spec.retentionDays is an integer from 1 to 90 with a default of 7. v1alpha1 is served but not the storage version and has only spec.owner; mark it deprecated and write the deprecation warning text yourself. The server must reject fields that are not in the schema.
  2. In /root/exam/api/env-defaults.yaml, write the Kyverno ClusterPolicy env-defaults. It is a mutate rule that applies only to Environment resources, filling the label platform.labhub.io/owner with the value of spec.owner and the annotation platform.labhub.io/requested-tier with the value of spec.tier. Do not write the values as fixed strings; read them from the request to fill them in.
  3. In /root/exam/api/env-rbac.yaml, write the aggregated permissions and apply them. The ClusterRole platform-env-author has no rules of its own and collects the ClusterRoles that carry the label platform.labhub.io/aggregate-to-env-author: "true". Attach that label to the ClusterRole platform-env-author-base and allow get, list, watch, create, update, and patch on Environment, but do not give delete. Create the namespace tenant-amber and the ServiceAccount env-author, and bind that aggregated ClusterRole with the RoleBinding env-author. Do not use * in any rule.
  4. In /root/exam/api/env-status-rbac.yaml, separate the permissions for spec and status. In tenant-amber, create the ServiceAccount env-controller, the Role env-controller, and the RoleBinding env-controller. This account can read Environment and update and patch environments/status, but cannot create, update, patch, or delete the Environment itself. Conversely, the env-author from task 7 must not be able to update environments/status.

Observability and Operations

  1. In /root/exam/ops/platform-rules.yaml, write the Prometheus rules. The group name is platform-api, the recording rule platform:request_error:ratio5m uses division and a 5-minute range, and the alert rule PlatformApiErrorBudgetBurn must have for, labels.severity, annotations.summary, and annotations.runbook_url. And in /root/exam/ops/platform-rules-test.yaml, write the rule unit test. In rule_files, write only the relative name platform-rules.yaml, and have two or more alert_rule_test entries, one of which checks a time when the alert fires and the other a time when it does not fire (exp_alerts: []). Both promtool check rules and promtool test rules must pass.
  2. Create the namespace platform-system and, in /root/exam/ops/servicemonitor.yaml, write the ServiceMonitor platform-api and apply it. The target selector is app: platform-api, and the namespaceSelector is platform-system. There is one endpoint, the port name is metrics, the scrape interval is 30s, and the scrape timeout must be shorter than the interval. In metricRelabelings, there must be at least one rule that looks at __name__ and drops a specific metric, and at least one rule that drops a high-cardinality label with labeldrop.
  3. Create the namespace tenant-ochre, attach the label platform.labhub.io/tenant: ochre, and limit it with the ResourceQuota ochre-quota to requests.cpu of 2 and requests.memory of 4Gi. Inside it, create the Deployment search with 4 replicas. The container name is app, the image is pinned by digest, and the requests are CPU 700m and memory 512Mi. At this point not all the Pods can start. Then, in /root/exam/ops/triage.txt, write four lines, quota_cpu_hard, quota_memory_hard, desired_replicas, and max_cpu_per_pod, in the 키=값 format (key=value). Write CPU as an integer in millicores, memory as an integer in Mi, and for the last value, write as a millicore integer the maximum CPU request one Pod may have for all the replicas to start within this quota. Do not attach unit characters.
  4. Without raising the quota, make all 4 Pods of search Running. Leave the replica count and the memory request as they are, and lower only the CPU request to the value you computed in task 11. No Pod that uses the old request value may remain.

Platform Architecture and Infrastructure

  1. Put the label platform.labhub.io/pool=platform and the taint platform.labhub.io/pool=platform:NoSchedule together on exactly two nodes. Then create the Deployment portal in platform-system with 3 replicas. The container name is app, the image is pinned by digest, and the requests are CPU 100m and memory 128Mi. This workload must select only that pool, tolerate that taint, and spread based on topology.kubernetes.io/zone with maxSkew: 1 and whenUnsatisfiable: DoNotSchedule. All 3 Pods must be Running on pool nodes.
  2. Create the PriorityClasses platform-critical (value 100000 or more) and tenant-batch (value 1000 or less, preemptionPolicy: Never). For both, globalDefault is false. Apply the ResourceQuota amber-priority to tenant-amber so that Pods that use the platform-critical priority cannot be created in that namespace at all. And make the portal from task 13 use platform-critical.
  3. In platform-system, create the PodDisruptionBudget portal. The selector selects the portal Pods, minAvailable is 2, and do not use maxUnavailable. It must be possible to drain only one node at a time during maintenance.

Security and Policy Enforcement

  1. Create the namespace tenant-teal and apply Pod Security admission enforce, audit, and warn all as restricted, but pin the version labels of the three modes to v1.31, not to latest. Inside it, create the Deployment checkout with 2 replicas. The container name is app, the image is pinned by digest, and the Pods must pass restricted. 2 Pods must actually be Running.
  2. Enforce one rule in two places. The rule is "a Deployment must have the label platform.labhub.io/owner." In /root/exam/security/owner-policy.yaml, write the Kyverno ClusterPolicy require-owner with Enforce. It is the gate the pipeline will run with kyverno apply. In /root/exam/security/owner-vap.yaml, write the ValidatingAdmissionPolicy require-owner and a binding with the same name, and apply it. validationActions is Deny, and the scope of application is only the namespaces that carry the label platform.labhub.io/gate=owner. Create the namespace delivery and attach that label. A Deployment without the label must be rejected in delivery and pass in default.

Reference

Make the chart reject wrong values by itself

If there is a values.schema.json inside the chart, helm validates the values before template and install, and if they do not conform, it rejects the render itself. Use the JSON schema's minimum, maximum, enum, pattern, and required. You confirm that the schema takes effect by deliberately passing a nonconforming value, as in helm template ... --set replicas=1.

Make the deployment change when the configuration changes

configMapGenerator appends a hash of the content to the name and automatically rewrites the places that reference that name within the same kustomization. In the Deployment, it is enough to write the original name without the suffix. If you turn on disableNameSuffixHash, this link breaks, and Pods do not restart when only the configuration changes.

Pin down what the project allows

An AppProject is the fence that decides what an Application may deploy and where. If you leave asterisks in sourceRepos and destinations, it is the same as having no fence. If you leave clusterResourceWhitelist empty, cluster-scoped resources cannot be created at all, and namespaceResourceBlacklist is the place to list the kinds that must not be touched even inside a namespace.

Leave the promotion as a single commit

In GitOps, a deployment is a commit. So what went up and when must remain in the log, and a rollback must be a revert. Move the digest you verified in staging to production as is, but do not touch the staging file. When you are done, git status must be empty. A change you have not committed is not deployed.

Make the schema a contract

If you write a default in a structural schema, the server fills in the value, and the server rejects fields that are not written. Only when the two properties come together does the schema become a contract. Mark deprecation with deprecated and deprecationWarning in the versions entry, and that text actually appears as a Warning on the kubectl screen. Confirm with kubectl apply --dry-run=server -o json.

Have the platform fill in what the request is missing

The less a user has to write in self-service, the better. What can be derived from values the user already wrote must be filled in by the platform. In a Kyverno mutate rule, you can read the request body with {{ request.object... }}. If you write the values as fixed strings, it is exposed when the grader asks with a different request. Confirm with kyverno apply --resource -o .

Collect permissions in pieces

An aggregated ClusterRole does not write its own rules. If you write only a label selector, the controller collects and fills in the rules of the other ClusterRoles that carry that label. So when you increase permissions later, you do not edit this role; you just attach one more piece. Confirm that the aggregation actually happened with kubectl get clusterrole -o jsonpath='{.rules}'. If it is empty, the label is mismatched.

Separate who writes from who fills in

status is the place the controller fills in, and spec is the place the user writes. In RBAC these two are separate resource names, so you write environments and environments/status separately. When checking, do not append a slash after the resource name; use --subresource=status. If you write it with a slash, kubectl interprets it wrongly and gives the opposite result.

Pin down the alert rule with a test

promtool test rules feeds in fake time series and actually evaluates the rules. Write the values of input_series as a start value, an increment, and a repeat count, like '0+3x30'. If you check only the time it fires, even a rule that always fires passes, so also write the time when it must not fire, with exp_alerts: []. rule_files is looked up relative to the directory where the test file is.

Decide what not to measure

Observability cost is determined by the number of time series collected, and that number explodes with combinations of labels. So a scrape configuration must include what to drop as much as what to measure. drop discards the entire metric if the value picked by sourceLabels matches the regex, and labeldrop deletes label names that match the regex. You must not write sourceLabels in labeldrop.

Reproduce quota exhaustion and write it down as numbers

When a quota blocks, the Deployment does not report an error. The ReplicaSet carries a ReplicaFailure condition instead, and the Pods are not created in the first place. Look at that condition with kubectl describe rs or kubectl get rs -o yaml. Classification is not a feeling but a division. Dividing the quota cap by the desired replica count gives the maximum a single Pod can use.

Recover without raising the quota

Lowering the request alone is not enough. When the quota is full, there is no room to create new Pods, so the rolling update stops in the middle. As long as Pods that use the old request value remain, the room does not open up. Scale the replicas down to 0 and back up, or empty the old Pods first. Confirm by the request value of each Pod, not by the number of Pods.

Divide the pool and spread the workload

A label attracts and a taint repels. You must apply both together for that pool to be dedicated. With only the label, anyone can come in, and with only the taint, there is no way in. The labelSelector of topologySpreadConstraints must point to the Pod labels, and it is actually enforced only if you set DoNotSchedule. Confirm by which node each Pod sits on.

Decide who may use a priority

A PriorityClass is cluster-scoped, so once you create it, anyone can take it and use it by name. So you must block separately not the priority itself but the right to use that priority. If you apply the PriorityClass scope in a ResourceQuota's scopeSelector and set the cap to 0, Pods of that class cannot be created at all in that namespace. preemptionPolicy: Never does not mean it is not itself pushed out; it means it does not push others out.

Decide how many must remain during maintenance

A PodDisruptionBudget is looked at by the eviction API, not the scheduler. kubectl drain and node upgrades go through that path. With 3 replicas and minAvailable 2, you can drain only one at a time. The result the controller computed appears in the ALLOWED DISRUPTIONS column of kubectl get pdb, and if this value is 0, maintenance is blocked entirely, and if it equals the replica count, it protects nothing.

Apply a baseline to the tenant namespace

Pod security admission does not block the Deployment but blocks the Pods that the Deployment creates. If the Deployment was created but there are no Pods, look at the reason with kubectl describe rs. restricted requires runAsNonRoot and seccompProfile at the Pod level and allowPrivilegeEscalation: false and capabilities.drop [ALL] at the container level. If you leave the version label at latest, the baseline silently becomes stricter when you upgrade the cluster.

Enforce the same rule in two places

A pipeline gate gives quick feedback but cannot block changes that did not go through the pipeline, and an admission policy blocks whatever comes in but the developer learns of it only after pressing deploy. So you put the same rule in two places. The scope of application of a VAP is decided by the binding and can be narrowed with the namespaceSelector in matchResources. If validationActions has no Deny, it leaves only an audit record and lets it through.