TT Lab
Get started
Learn Learning paths Courses

Integration and Deployment

Facing an API That Loses Its Replies

Continue in TT Lab

Goal

You build yourself how a retry after a timeout creates a duplicate payment, and block it with an idempotency key. And you also confirm the cases where it fails even though a key is attached.

Environment

/opt/app/pay.py imitates a customer's payment API. It is not something to fix; it is the other party.

mkdir -p /root/idem
nohup python3 /opt/app/pay.py > /tmp/pay.log 2>&1 &
sleep 1
curl -s http://127.0.0.1:8021/health
POST /payments    결제 하나를 기록한다
                  Idempotency-Key 헤더가 있고 이미 본 키면
                  새로 기록하지 않고 저장해 둔 결과를 그대로 돌려준다
GET  /payments    지금까지 기록된 결제 전부
POST /reset       상태 초기화

Every even-numbered request finishes processing but delays only the response by 6 seconds. If you set a 2-second timeout, the client sees it as a failure, but it is already recorded on the server. This is an "indeterminate result."

What to build

All of it is under /root/idem/.

naive.sh      멱등키 없이. 타임아웃이면 다시 보낸다
naive.txt     그 결과와 왜 그런지
safe.sh       비즈니스 행위 하나에 키 하나. 재시도에도 같은 키
newkey.sh     시도마다 새 키를 만든다 (일부러 틀린 형태)
newkey.txt    그 결과와 키를 무엇 단위로 만들어야 하는지
backoff.sh    대기 시간 간격을 낸다
retryable.sh  상태를 받아 재시도 여부를 답한다
payments.csv  조사할 결제 기록
forensics.txt 중복을 찾고 원인을 지목한 결과
report.md     정리

How it is graded

The grader resets the server every time and runs your scripts directly, and then counts the records. It looks not at the numbers you wrote down but at the records actually left.

naive.sh   기록 > 주문  이어야 한다 (중복이 생겨야 한다)
safe.sh    기록 = 주문  이어야 한다 (정확히 한 번)
newkey.sh  기록 > 주문  이어야 한다 (키가 무력해진다)
backoff.sh 간격이 늘고, 두 번 돌리면 달라야 한다

Steps

  1. Start the payment API.
  2. naive.sh — pay 4 or more orders without an idempotency key, and on a timeout send again. Write the result to naive.txt.
  3. safe.sh — decide one key per order, and when retrying use the same key again.
  4. newkey.sh — attach a key, but create a new one on every attempt. Write in newkey.txt why it is useless.
  5. backoff.sh — output 4 or more retry intervals. They must grow exponentially and have randomness mixed in.
  6. retryable.sh — take a status (500, 429, 400, timeout, and so on) as an argument and output yes or no.
  7. Create payments.csv and find the duplicates. The time intervals of the duplicate records tell you the cause.
  8. Wrap up.

Notes

The material for step 7 is made like this.

cat > /root/idem/payments.csv <<'CSV'
payment_id,order_id,created_at
PAY-1,ORD-1,2026-09-07T10:00:00Z
PAY-2,ORD-2,2026-09-07T10:00:05Z
PAY-3,ORD-2,2026-09-07T10:00:06Z
PAY-4,ORD-2,2026-09-07T10:00:08Z
PAY-5,ORD-2,2026-09-07T10:00:12Z
PAY-6,ORD-3,2026-09-07T10:01:00Z
PAY-7,ORD-4,2026-09-07T10:02:00Z
PAY-8,ORD-4,2026-09-07T10:02:01Z
PAY-9,ORD-4,2026-09-07T10:02:03Z
PAY-10,ORD-4,2026-09-07T10:02:07Z
CSV

If you see the intervals doubling as 1 second, 2 seconds, 4 seconds, it is almost certain.

Start an API that loses responses

Start the payment API.

Start /opt/app/pay.py in the background and check that /health is 200. Clear the state with /reset before starting.

If you treat a timeout as a failure

naive.sh — pay 4 or more orders without an idempotency key, and on a timeout send again. Write the result to naive.txt.

Set a timeout with curl --max-time 2, and on failure send once more with ||. 4 or more orders. Then count the actual records with GET /payments.

Exactly once with an idempotency key

safe.sh — decide one key per order, and when retrying use the same key again.

Decide one key per order and use the same key for the retry too. Attach -H "Idempotency-Key: $K" to both calls.

When it fails even though a key is attached

newkey.sh — attach a key, but create a new one on every attempt. Write in newkey.txt why it is useless.

Try creating a new key on every attempt. The server sees them as different requests and records again. A key is made for "one business action," not "one HTTP request."

Widen the intervals and scatter them

backoff.sh — output 4 or more retry intervals. They must grow exponentially and have randomness mixed in.

The intervals must grow each time (exponential backoff), and running it twice must give different values (jitter). Without jitter, all clients retry at the same moment.

What to retry and what not to

retryable.sh — take a status (500, 429, 400, timeout, and so on) as an argument and output yes or no.

A wrong request (400, 422) and a credentials problem (401, 403) are the same however many times you send them. Retry a 429, but observe Retry-After.

The intervals tell the cause

Create payments.csv and find the duplicates. The time intervals of the duplicate records tell you the cause.

Find the duplicated orders and look at the time differences of those records. If they double like 1 second, 2 seconds, 4 seconds, it is client retries.

Wrap-up

Wrap up.

Write what a timeout means, what unit a key is made for, and what each of backoff and jitter prevents.