TT Lab
Get started
Learn Learning paths Courses

CNPE — Cloud Native Platform Engineer

The Order in Which You Classify an Incident

Continue in TT Lab

One-line summary

When a Pod does not start, what you should count first is not the remaining resources but the number of candidate nodes. If there are 0 candidates, the remaining resources mean nothing.

Why the order is fixed

First separate whether the Pod object exists. A quota violation, as described in the official ResourceQuota documentation, can reject a creation request. In that case you must look not for a Pending Pod but for FailedCreate in the Deployment and ReplicaSet events and for API errors. A LimitRange violation can also be an admission-stage problem.

If the Pod exists, check the conditions and events following the Pod lifecycle. Pending covers not only the period before scheduling but also waits such as image preparation. Look together at the PodScheduled condition, nodeName, and the container waiting reason. Do not interpret every Pending as a resource shortage in the same order.

The hypothetical incident in this lab is the case where there are 0 candidates that match the nodeSelector. You confirm that condition from the events and then count the candidates. Even when there are candidates, the Pod may not be placed because of taints, affinity, volumes, or the actual request volume. Distinguish the required selection conditions from the preferred conditions in the node assignment documentation.

How it works

The result of classification must be a number

"It seems resources are short" is not a classification. A classification is writing down the following four values.

Item Where it comes from
Selector key and value The workload's nodeSelector
Replica count The Deployment's spec.replicas
Total CPU required Container request × replicas
Candidate node count The value from counting nodes with that selector

If these four values are written down, the next person starts from the same place. If they are not, the next person starts over from the beginning. The value of an incident record lies not in sentences but in these numbers.

Recovery and removing a requirement are different things

The fastest way to get a Pod to start is to delete the nodeSelector. If there are no other constraints, it can be placed. But someone wrote that selector for a reason. It may need a GPU, need to attach to particular storage, or be required to run only on specific nodes for regulatory reasons.

If you delete the constraint without checking the basis of the requirement and the change approval, it can become removing a requirement. And this difference does not show on the screen. The Pod is Running, the alert goes quiet, and the dashboard is green. The problem comes back weeks later in an unexpected form.

Recovery is on the side of satisfying the constraint. It means attaching the pool label to a node, putting a node into that pool, or, if you cannot, writing down why you cannot and having a person make the decision to change the workload's requirement.

Right-sizing has a fixed direction

When you are asked to cut costs, you feel tempted to reduce nodes first. First collect usage, peaks, failures, and latency for a representative period, and review the slack in requests and the recovery capacity. If you lower things looking only at average usage, you miss load spikes. Request adjustments and node reductions should be tested in a small scope with the conditions for reverting defined. Scale-to-zero is likewise a possible choice depending on startup delay and workload requirements, and you cannot assert that it always means interruption or always means savings.

What it looks like in the field

The following is a hypothetical case for the lab. Even after adding three nodes, the Pod stays Pending. The pool label was not attached to the new nodes, and the workload was requiring that label. The capacity graph went up, and the candidate nodes stayed at 0. It is the kind of incident you would never see by looking only at a dashboard.

Another is the case where someone attached the label, checked immediately as if grading, and concluded "it doesn't work." The scheduler retries periodically, so you must wait a few seconds. If you check and judge immediately, you end up reverting a correct action.

What to do in the next lab

You create three tenants and divide resources with quotas and LimitRanges, compute the overcommit ratio yourself, reproduce a Pending incident and classify it, recover without deleting the constraint, and create and upload the platform's own recording rule and alert.