Facing an API That Loses Its Replies
Goal
You build yourself how a retry after a timeout creates a duplicate payment, and block it with an idempotency key. And you also confirm the cases where it fails even though a key is attached.
Environment
/opt/app/pay.py imitates a customer's payment API. It is not something to fix;
it is the other party.
mkdir -p /root/idem
nohup python3 /opt/app/pay.py > /tmp/pay.log 2>&1 &
sleep 1
curl -s http://127.0.0.1:8021/health
POST /payments 결제 하나를 기록한다
Idempotency-Key 헤더가 있고 이미 본 키면
새로 기록하지 않고 저장해 둔 결과를 그대로 돌려준다
GET /payments 지금까지 기록된 결제 전부
POST /reset 상태 초기화
Every even-numbered request finishes processing but delays only the response by 6 seconds. If you set a 2-second timeout, the client sees it as a failure, but it is already recorded on the server. This is an "indeterminate result."
What to build
All of it is under /root/idem/.
naive.sh 멱등키 없이. 타임아웃이면 다시 보낸다
naive.txt 그 결과와 왜 그런지
safe.sh 비즈니스 행위 하나에 키 하나. 재시도에도 같은 키
newkey.sh 시도마다 새 키를 만든다 (일부러 틀린 형태)
newkey.txt 그 결과와 키를 무엇 단위로 만들어야 하는지
backoff.sh 대기 시간 간격을 낸다
retryable.sh 상태를 받아 재시도 여부를 답한다
payments.csv 조사할 결제 기록
forensics.txt 중복을 찾고 원인을 지목한 결과
report.md 정리
How it is graded
The grader resets the server every time and runs your scripts directly, and then counts the records. It looks not at the numbers you wrote down but at the records actually left.
naive.sh 기록 > 주문 이어야 한다 (중복이 생겨야 한다)
safe.sh 기록 = 주문 이어야 한다 (정확히 한 번)
newkey.sh 기록 > 주문 이어야 한다 (키가 무력해진다)
backoff.sh 간격이 늘고, 두 번 돌리면 달라야 한다
Steps
- Start the payment API.
naive.sh— pay 4 or more orders without an idempotency key, and on a timeout send again. Write the result tonaive.txt.safe.sh— decide one key per order, and when retrying use the same key again.newkey.sh— attach a key, but create a new one on every attempt. Write innewkey.txtwhy it is useless.backoff.sh— output 4 or more retry intervals. They must grow exponentially and have randomness mixed in.retryable.sh— take a status (500,429,400,timeout, and so on) as an argument and outputyesorno.- Create
payments.csvand find the duplicates. The time intervals of the duplicate records tell you the cause. - Wrap up.
Notes
The material for step 7 is made like this.
cat > /root/idem/payments.csv <<'CSV'
payment_id,order_id,created_at
PAY-1,ORD-1,2026-09-07T10:00:00Z
PAY-2,ORD-2,2026-09-07T10:00:05Z
PAY-3,ORD-2,2026-09-07T10:00:06Z
PAY-4,ORD-2,2026-09-07T10:00:08Z
PAY-5,ORD-2,2026-09-07T10:00:12Z
PAY-6,ORD-3,2026-09-07T10:01:00Z
PAY-7,ORD-4,2026-09-07T10:02:00Z
PAY-8,ORD-4,2026-09-07T10:02:01Z
PAY-9,ORD-4,2026-09-07T10:02:03Z
PAY-10,ORD-4,2026-09-07T10:02:07Z
CSV
If you see the intervals doubling as 1 second, 2 seconds, 4 seconds, it is almost certain.
Start an API that loses responses
Start the payment API.
Start /opt/app/pay.py in the background and check that /health is 200. Clear the state with /reset before starting.
If you treat a timeout as a failure
naive.sh — pay 4 or more orders without an idempotency key, and on a timeout
send again. Write the result to naive.txt.
Set a timeout with curl --max-time 2, and on failure send once more with ||. 4 or more orders. Then count the actual records with GET /payments.
Exactly once with an idempotency key
safe.sh — decide one key per order, and when retrying use the same key again.
Decide one key per order and use the same key for the retry too. Attach -H "Idempotency-Key: $K" to both calls.
When it fails even though a key is attached
newkey.sh — attach a key, but create a new one on every attempt. Write in newkey.txt
why it is useless.
Try creating a new key on every attempt. The server sees them as different requests and records again. A key is made for "one business action," not "one HTTP request."
Widen the intervals and scatter them
backoff.sh — output 4 or more retry intervals. They must grow exponentially
and have randomness mixed in.
The intervals must grow each time (exponential backoff), and running it twice must give different values (jitter). Without jitter, all clients retry at the same moment.
What to retry and what not to
retryable.sh — take a status (500, 429, 400, timeout, and so on) as an argument and
output yes or no.
A wrong request (400, 422) and a credentials problem (401, 403) are the same however many times you send them. Retry a 429, but observe Retry-After.
The intervals tell the cause
Create payments.csv and find the duplicates. The time intervals of the duplicate records
tell you the cause.
Find the duplicated orders and look at the time differences of those records. If they double like 1 second, 2 seconds, 4 seconds, it is client retries.
Wrap-up
Wrap up.
Write what a timeout means, what unit a key is made for, and what each of backoff and jitter prevents.