TT Lab
Get started
Learn Learning paths Courses

Integration and Deployment

What Happens When You Retry After a Timeout

Continue in TT Lab

Summary

A timeout does not mean the request failed but we do not know the result. If you miss this difference, retries create duplicate payments and duplicate sends.

Why this was needed

The most common incident in payment integration goes like this.

  1. We send a payment request
  2. A 3-second timeout — no response comes
  3. We conclude "it must have failed" and retry
  4. At the payment provider, it is approved twice

The judgment in step 3 was wrong. Not receiving a response and not having been processed are different. The request may have arrived, been processed, and only the response lost.

In distributed systems this is called an indeterminate result. It is a third state, neither success nor failure, and there are separate ways to handle this state.

What may and may not be retried

The key question is "is this operation idempotent?" If sending the same request twice gives the same result as sending it once, it is idempotent.

Operation Idempotent? Retry
GET /orders/123 Yes Freely
PUT /orders/123 {status:"PAID"} Yes (writes the same value) Safe
DELETE /orders/123 Yes (the second time, it is already gone) Safe
POST /orders {…} No Risky — needs an idempotency key
POST /payments {amount:1000} No Risky
Deducting a balance balance -= 1000 No Risky

The idempotency of an HTTP method is a convention and not a guarantee. If the server increments a counter on PUT, that PUT is not idempotent. You must check the behavior, not the documentation.

Idempotency keys

To safely retry a non-idempotent operation, the client creates and sends a unique key with each request.

POST /payments
Idempotency-Key: 3f9a1c7e-2b44-4c9d-a1f8-0d6e2b7c5a11
{"order_id": 123, "amount": 1000}

The server stores this key, and when the same key arrives again, it does not process it anew but returns the stored result as is. However many retries there are, the payment happens once.

There is something to watch out for when creating keys. You must use the same key when retrying. If you create a new UUID on every attempt, it means nothing. A key is one per "this business action" and not one per "this HTTP request."

Retry policy

Retrying blindly makes an outage bigger. If the downstream is slow and timeouts occurred, but everyone retries immediately, the load doubles and everything collapses entirely.

And distinguish which errors to retry.

Response Retry
Connection refused, timeout Yes (when idempotent or when there is an idempotency key)
429 Too Many Requests Yes — observing Retry-After
500, 502, 503, 504 Usually yes
400, 422 (the request is wrong) No — it is the same however many times
401, 403 No — you have to fix the credentials

Signals when investigating

If you suspect duplicates, you find them like this.

SELECT order_id, count(*) FROM payments
GROUP BY 1 HAVING count(*) > 1;

Then look at the interval between the created_at of those duplicates. If the intervals are similar to the retry backoff (1 second, 2 seconds, 4 seconds), the cause is almost certain.

How retries make outages bigger

Used well, retries hide transient failures, and used badly, they grow a small outage into a big one. The paths by which they grow it are fixed.

If every layer retries, it multiplies. If the client retries 3 times, the gateway 3 times, and the service 3 times, one user request arrives upstream as 27 times the load. What began because upstream was slow makes upstream slower. The principle is to retry in only one layer, and usually the layer closest to the user is the place.

Retrying at the same interval makes a wave. Requests that failed during the outage come back all at once exactly 1 second later. The service that was about to recover falls over again under that wave. This is why you mix random jitter into exponential backoff.

delay = min(cap, base * 2 ** attempt) * (0.5 + random.random() * 0.5)

Without a circuit breaker, retries do not stop. When upstream is completely dead, retries are meaningless and only burn resources. When the failure rate exceeds a threshold, stop even trying for a moment, and let through just one request now and then to check for recovery. Failing fast is better for the user too — an error that comes immediately is easier to retry than an error after waiting 30 seconds.

Propagate the deadline. If the user waits 5 seconds, every segment the request passes through must know the remaining time. Starting a new 3-second call when 200ms remain is waste. Each layer checks gRPC's deadline or an expiry time passed in a header.

Do not retry errors that cannot be retried. 400, 401, 404, and 409 are the same however many times you resend. What is worth retrying is only 429, 502, 503, 504, and connection errors, and a 429 comes with a Retry-After, so observe that value.

Leave a mark on retried requests. If you put the attempt number in the logs and headers, then in an outage investigation "requests tripled" is immediately split into whether users increased or retries did.

What it looks like in the field

What you will do in the following lab

You pay three times against a payment API that processes the request but loses only the response.

멱등키 없이             주문 4건 → 기록 7건
비즈니스 행위당 키 하나   주문 4건 → 기록 4건
시도마다 새 키           주문 4건 → 기록 7건   ← 키를 붙였는데도 소용없다

The third is the core of this lab. Merely having attached a key guarantees nothing.