Adding Timeouts and Retries to Service-to-Service REST Calls
Goal
Against an unstable downstream, implement timeouts, exponential backoff, jitter, no-retry conditions, and a fallback yourself, completing the fundamentals of synchronous calls.
Why it matters
Incidents in synchronous calls usually happen "because there was no timeout" or "because retries were done wrongly". Without a timeout, when the downstream slows down, the upstream's connection pool dries up first, and even requests unrelated to the downstream fail. Conversely, if you apply unlimited fixed-interval retries, the same wave hits every time the downstream tries to recover. In a real case, when 20 order instances each handling 100 requests per second retried with maxAttempts 5, the payment service received 10,000 requests per second. Retries do not add load; they multiply it. So this lab treats the order in which you put in the values as important — first the timeout, then the retry conditions, then backoff and jitter, and lastly the fallback.
Steps
- Start
/opt/app/flaky.pyon 127.0.0.1:8110.GET /fastmust be 200. - Create
/root/rest/call.py, call/fast, and save the response body to/root/rest/fast.out. - Create
/root/rest/timeout.pyand call/slowwith a 1.0-second timeout. On the first line of/root/rest/timeout.outwriteTIMEOUT, and on the second lineelapsed=<초>(the elapsed seconds). The elapsed time must be under 2.0 seconds. - Create
/root/rest/retry.py, and retry/flaky?key=lab&fail=2with exponential backoff until it finally succeeds. In/root/rest/retry.log, leave one line per attempt in the formattempt=<n> wait=<초>(the wait in seconds), for a total of 3 lines. - Change the wait calculation of
retry.pyto full jitter. Draw the wait time for attempt number 2 twenty times and write them to/root/rest/jitter.txt, one per line. There must be at least 15 distinct values. - With
/root/rest/retry400.py, call/bad. A 400 is not retried, so/root/rest/retry400.logmust be exactly 1 line. - In
/root/rest/budget.txt, write four lines:instances=20,rps=100,max_attempts=5, andworst_rps=10000. - Start
/root/rest/gateway.pyon 127.0.0.1:8111.GET /ordercalls the downstream/always500, applies timeouts and retries, and returns 200 and{"degraded":true}on final failure.
Notes
- Start the timeout value from the other side's p99. If you set it to 10 times the p99, it is the same as having none.
- full jitter:
wait = random.uniform(0, min(cap, base * 2 ** attempt)) - Common mistake 1: logging only the successful attempts in the retry log — you must log the failed attempts too to be able to calculate the budget.
- Common mistake 2: not catching the timeout as an exception, so only a stack trace remains and you cannot measure the elapsed time.
Start an unstable downstream
Start /opt/app/flaky.py on 127.0.0.1:8110. GET /fast must be 200.
If you run /opt/app/flaky.py, it comes up on 8110. First see with your own eyes how the four paths /fast, /slow, /flaky, and /bad behave.
Set a baseline with a normal call
Create /root/rest/call.py, call /fast, and save the response body to /root/rest/fast.out.
Call /fast with httpx or urllib and save the body to a file. You do not need to worry about the timeout yet.
Cut off a slow response with a timeout
Create /root/rest/timeout.py and call /slow with a 1.0-second timeout. On the first line of /root/rest/timeout.out write TIMEOUT, and on the second line elapsed=<초> (the elapsed seconds). The elapsed time must be under 2.0 seconds.
/slow answers after 3 seconds. If you set the client timeout to 1 second, an exception occurs. Catch the exception and record the result together with the elapsed time.
Implement exponential backoff retries
Create /root/rest/retry.py, and retry /flaky?key=lab&fail=2 with exponential backoff until it finally succeeds. In /root/rest/retry.log, leave one line per attempt in the form attempt=<n> wait=<초> (the wait in seconds), for a total of 3 lines.
Double the interval on each attempt. You must log the number and wait time of each attempt, one line each, to be graded.
Scatter the retry timing with jitter
Change the wait calculation of retry.py to full jitter. Draw the wait time for attempt number 2 twenty times and write them to /root/rest/jitter.txt, one per line. There must be at least 15 distinct values.
Full jitter waits a random number between 0 and the computed backoff. The value must differ each time even for the same attempt number.
Do not retry definitive errors
With /root/rest/retry400.py, call /bad. A 400 is not retried, so /root/rest/retry400.log must be exactly 1 line.
A 400 is a 400 however many times you send it. Add a branch that decides whether to retry by the status code.
Calculate the retry budget
In /root/rest/budget.txt, write four lines: instances=20, rps=100, max_attempts=5, and worst_rps=10000.
Number of instances x requests per second x maximum attempts is the worst-case requests per second the downstream receives. Write the three values and the result each.
Bring it together with a gateway that has a fallback
Start /root/rest/gateway.py on 127.0.0.1:8111. GET /order calls the downstream /always500, applies timeouts and retries, and returns 200 and {"degraded":true} on final failure.
Even if the downstream dies, give the user a 200, but leave a quality-degradation marker in the response. Use the timeout and retries from the previous steps together.