TT Lab
Get started
Learn Learning paths Courses

CKAD — Kubernetes Application Developer

Duplicate Shipments: Job Success Is Not Business Success

Continue in TT Lab

Goal

You reproduce the incident in which a Job that processes one order succeeds but the shipment is created twice, and explain it on the basis of retries, idempotency keys, and the execution deadline.

Why it matters

Even if completions and parallelism are 1, it does not guarantee that the external work is processed only once. That is because a failed process may already have finished storing. You can judge whether to rerun only by observing separately the execution lifetime Kubernetes manages and the business lifetime the ledger manages.

It uses a real k3s in a personal VM and a synthetic HTTP and SQLite ledger. There is no real shipping, payment, or personal information. The ledger's safe mode merges duplicate rows with the server's write transaction and business-key uniqueness. This lab deals with applying and verifying that contract, and does not guarantee exactly-once processing across a whole distributed system. If you delete the ledger Pod, the data disappears too.

The initial setup is allowed up to 6 minutes. Put all files under /root/ckad-job-effects. The work disappears when the session ends, so download the records you need in advance. The lab is based on 55 minutes, and if you need more, extend the time before it expires.

The template and the run helper

job-template.json is a Job JSON that includes the real worker command, a pinned image, and permission restrictions. For each step, copy it and change metadata.name, spec.backoffLimit, and the CASE_ID, MODE, and FAULT values in the container env. Also change labhub.io/case in the Pod template's spec.template.metadata.labels to the same value as that CASE_ID. Keep WORK_ID as order-001 and POD_UID via the Downward API. Keep parallelism and completions at 1 and restartPolicy at Never. Do not change any other security or app settings. Add spec.activeDeadlineSeconds only in the steps that need a deadline.

Do not run it with a direct kubectl create; after writing the file, run python3 /opt/fixtures/ckad_job_effects_lab.py run <단계> (the placeholder is the step number). The helper submits the student file to the API as is and preserves the Pod observations during the run. If you run the same completed step again, it does not create a new Job and only grades. If the wait for observation is interrupted, it continues with the same command and does not delete or recreate the Job. If the creation response itself was lost, it asks for a new lab instead of creating a duplicate.

Steps

  1. baseline.json: name=baseline, CASE_ID=baseline, MODE=unsafe, FAULT=none, backoffLimit=0. Observe the normal comparison group with run 1.
  2. unsafe-retry.json: name=unsafe-retry, CASE_ID=unsafe-retry, MODE=unsafe, FAULT=after-commit, backoffLimit=1. Observe the failure after the store and the replacement Pod with run 2.
  3. Read observation-2.json and write job_uid, failed_pod_uid, succeeded_pod_uid, request_count, shipment_count, and retry_layer in retry-analysis.json. retry_layer is whichever of job-controller or kubelet-container fits the observation. Record the analysis with run 3.
  4. safe-retry.json: name=safe-retry, CASE_ID=safe-retry, MODE=safe, FAULT=after-commit, backoffLimit=1. With run 4, confirm the idempotent result from the same failure.
  5. safe-replay.json: name=safe-replay, CASE_ID=safe-retry, MODE=safe, FAULT=none, backoffLimit=0. With run 5, see whether the same business key carries through in a new Job.
  6. no-retry.json: name=no-retry, CASE_ID=no-retry, MODE=unsafe, FAULT=after-commit, backoffLimit=0. With run 6, see whether a shipment remains even after the failure.
  7. deadline.json: name=deadline, CASE_ID=deadline, MODE=safe, FAULT=deadline, backoffLimit=0, activeDeadlineSeconds=20. With run 7, observe the deadline overrun and the ledger. Even if the current Pod list is empty, do not erase the past observations in observed_pods.
  8. Compare the six runs and write final-analysis.json. request_count and shipment_count are totals for the whole ledger, and safe_replay_shipment_id is the shipment number preserved in the rerun. timeout_cancels_effect and ledger_survives_ledger_pod_replacement are booleans for whether exceeding the deadline cancels the work and whether the data is preserved even when the ledger Pod is replaced. idempotency_scope lists, joined with +, the names of the fields actually used to decide duplicates. In outcomes, use the six Job names as keys and write condition (Complete/Failed) and shipments (the number of shipments for that work right after each run). Record the overall analysis with run 8.

Notes

Normal shipment comparison group

Write /root/ckad-job-effects/baseline.json from the template and run run 1 with the name, business key, mode, fault condition, and retry budget from the step 1 instructions.

Even in a normal run, keep the Job observation and the ledger observation separately.

It succeeded, yet there are two

Declare a failure after the store and one retry in /root/ckad-job-effects/unsafe-retry.json, and run run 2.

Never is the restart policy for containers within the same Pod. Look at the UID of the replacement Pod.

Which retry is the culprit

From observation-2.json, read the Job, the failed Pod and succeeded Pod UIDs, and the request and shipment counts, record them in retry-analysis.json, and run run 3.

Compare the same ownerReferences, the different Pod UIDs, and restartCount together.

Same failure, one shipment

Declare safe mode and a stable business key in /root/ckad-job-effects/safe-retry.json, and run run 4.

See whether the requests, even when repeated, link to the same shipment number.

Rerun with a different Job

Write /root/ckad-job-effects/safe-replay.json with a separate Job name while keeping CASE_ID=safe-retry, and run run 5.

The execution ID must be different and the business ID must be kept.

What remains even when you turn retries off

Declare backoffLimit=0 and a failure after the store in /root/ckad-job-effects/no-retry.json, and run run 6.

A Failed condition does not mean the store was rolled back.

What remains even when time runs out

Declare activeDeadlineSeconds=20 and waiting after the store in /root/ckad-job-effects/deadline.json, and run run 7.

Distinguish the end of execution from the cancellation of work, and preserve the Pod observations from before deletion.

Shipping incident summary report

From the observations of steps 1, 2, 4, 5, 6, and 7, write the totals, the reused shipment number, the cancellation and preservation judgments, the business key scope, and the six outcomes in final-analysis.json, and run run 8.

The total number of requests, the number of shipments per work item, and the execution success condition are different metrics.