CKAD — Kubernetes Application Developer
The Batch Succeeded, but the Parcel Shipped Twice
In one line
A Job being Complete means that the specified execution completion condition was met; it is not proof that the business task in an external system was processed exactly once.
Why this was needed
There is a batch that takes one order and creates a shipping request. The developer set both completions and parallelism to 1 and chose Never for restartPolicy as well. The run screen shows one successful Pod and Complete. Yet the ledger has two shipments for the same order. You wrote the number 1 in three places, so how could this happen?
The last line of a program and the moment an external system stores something are not the same event. The first Pod sent a request to the shipping API, and the server stored a new shipment. After receiving the response, the program ends with an error before it can leave exit code 0. The Job controller sees the failed run and creates a new Pod. The new Pod processes the same order again and this time exits normally. From the controller's point of view, it has finished the work it had to do, but the external ledger may be left with two stores.
The shipping in this lab is a synthetic service that adds a row to SQLite. It uses no real address, personal information, or shipping carrier API. You have to use a small, controllable side effect so that you can focus on practicing reading the cause of the failure.
How it works
First, distinguish who is executing. A Job is the target of the controller that manages the required number of completions. A Pod is the execution unit that holds the work program. When the process in the container ends, an exit code is left behind. The external business system leaves the result of the request in its own storage. If you read the success of one layer as the success of another, you will miss the cause.
parallelism relates to the number of Pods you want to run at the same time, and completions relates to the number of successes required. Neither counts the number of transactions in the external API. restartPolicy: Never is a choice that the kubelet does not rerun a failed container within the same Pod. It is not a switch that also forbids the Job controller from creating a replacement Pod.
When observing, do not look only at names; check the UID and the ownership relationship. If two Pods point to the same Job UID through ownerReferences, each has a restartCount of 0, and one is Failed while the other is Succeeded, that is evidence for telling an in-Pod restart from a Job retry. You also have to preserve the failed Pod's exit code and logs to narrow down whether the failure was before or after the store.
kubectl get jobs,pods -n ckad-job-effects
kubectl get pod -n ckad-job-effects <파드이름> -o json
kubectl logs -n ckad-job-effects <파드이름>
But you do not count shipments from a single success sentence in the log. You query the list of request attempts and the list of shipment rows separately on the server. The request attempts record which Pod sent which business key, and the shipment rows record the actual side effect. You must be able to compare whether two requests created two shipments, or whether the same shipment number was returned.
What it looks like in the field
Retry is not a feature for hiding errors. It is needed to withstand transient failures, but the responsibility to judge whether the work is safe to run again does not go away. Sending email, registering the result of a file conversion, deducting inventory, and accepting an external order all meet the same boundary problem. The stretch where it is hard to confirm whether the other system stored it matters especially.
If you change backoffLimit to 0, you can reduce the number of retries. That does not mean a shipment that was already recorded is canceled. If the first Pod fails after storing, the Job is failed yet one shipment may exist. If an operator looks at the failed Job and manually creates a new Job, there remains the possibility of creating the same side effect again. That is why, before unconditionally rerunning after seeing a failed state, you query the ledger by business key.
activeDeadlineSeconds is likewise the time allowed for execution, not a validity period for external storage. Even if a program waiting after storing is terminated by the time limit, the rows already created do not disappear automatically. Cancelling an execution, cancelling a business transaction, and compensation are different designs. In the lab you only observe this difference and do not perform any arbitrary compensating deletion.
The observation timing matters too. If the controller deletes the work Pod after the deadline is exceeded, running only get pods later may not find that run. Connect the Pod UID, the owning Job, and the container execution information collected during the experiment with the requests in the ledger, and check the current Job's failure reason separately. Do not record a Running observation from before the deletion as if it were an exit code, and do not conclude that there was no request just because the Pod is not visible. Regrading needs preserved material that distinguishes the collection time.
What you will do in the next lesson
In the lesson that follows, instead of preventing reruns, you design a contract that lets the same work be repeated safely. First distinguish the execution identifier from the business identifier. In the lab after that, you leave a normal Job as the comparison group and run a Job that fails right after the store, and record the Job condition, the actual Pod UID, the exit code, the request count, and the number of shipment rows together. The goal is to explain each of the two cases: Complete but two shipments, and Failed but one shipment.
Official documentation and scope
The ledger structure and fault injection in this reading are an independent learning example. It does not assume that every external system has the same storage model, and it is neither a copy of exam questions nor a substitute for the whole CKAD scope.