TT Lab
Get started
Learn Learning paths Courses

Microservice Architecture

Adding Timeouts and Retries to Service-to-Service REST Calls

Continue in TT Lab

Goal

Against an unstable downstream, implement timeouts, exponential backoff, jitter, no-retry conditions, and a fallback yourself, completing the fundamentals of synchronous calls.

Why it matters

Incidents in synchronous calls usually happen "because there was no timeout" or "because retries were done wrongly". Without a timeout, when the downstream slows down, the upstream's connection pool dries up first, and even requests unrelated to the downstream fail. Conversely, if you apply unlimited fixed-interval retries, the same wave hits every time the downstream tries to recover. In a real case, when 20 order instances each handling 100 requests per second retried with maxAttempts 5, the payment service received 10,000 requests per second. Retries do not add load; they multiply it. So this lab treats the order in which you put in the values as important — first the timeout, then the retry conditions, then backoff and jitter, and lastly the fallback.

Steps

  1. Start /opt/app/flaky.py on 127.0.0.1:8110. GET /fast must be 200.
  2. Create /root/rest/call.py, call /fast, and save the response body to /root/rest/fast.out.
  3. Create /root/rest/timeout.py and call /slow with a 1.0-second timeout. On the first line of /root/rest/timeout.out write TIMEOUT, and on the second line elapsed=<초> (the elapsed seconds). The elapsed time must be under 2.0 seconds.
  4. Create /root/rest/retry.py, and retry /flaky?key=lab&fail=2 with exponential backoff until it finally succeeds. In /root/rest/retry.log, leave one line per attempt in the form attempt=<n> wait=<초> (the wait in seconds), for a total of 3 lines.
  5. Change the wait calculation of retry.py to full jitter. Draw the wait time for attempt number 2 twenty times and write them to /root/rest/jitter.txt, one per line. There must be at least 15 distinct values.
  6. With /root/rest/retry400.py, call /bad. A 400 is not retried, so /root/rest/retry400.log must be exactly 1 line.
  7. In /root/rest/budget.txt, write four lines: instances=20, rps=100, max_attempts=5, and worst_rps=10000.
  8. Start /root/rest/gateway.py on 127.0.0.1:8111. GET /order calls the downstream /always500, applies timeouts and retries, and returns 200 and {"degraded":true} on final failure.

Notes

Start an unstable downstream

Start /opt/app/flaky.py on 127.0.0.1:8110. GET /fast must be 200.

If you run /opt/app/flaky.py, it comes up on 8110. First see with your own eyes how the four paths /fast, /slow, /flaky, and /bad behave.

Set a baseline with a normal call

Create /root/rest/call.py, call /fast, and save the response body to /root/rest/fast.out.

Call /fast with httpx or urllib and save the body to a file. You do not need to worry about the timeout yet.

Cut off a slow response with a timeout

Create /root/rest/timeout.py and call /slow with a 1.0-second timeout. On the first line of /root/rest/timeout.out write TIMEOUT, and on the second line elapsed=<초> (the elapsed seconds). The elapsed time must be under 2.0 seconds.

/slow answers after 3 seconds. If you set the client timeout to 1 second, an exception occurs. Catch the exception and record the result together with the elapsed time.

Implement exponential backoff retries

Create /root/rest/retry.py, and retry /flaky?key=lab&fail=2 with exponential backoff until it finally succeeds. In /root/rest/retry.log, leave one line per attempt in the form attempt=<n> wait=<초> (the wait in seconds), for a total of 3 lines.

Double the interval on each attempt. You must log the number and wait time of each attempt, one line each, to be graded.

Scatter the retry timing with jitter

Change the wait calculation of retry.py to full jitter. Draw the wait time for attempt number 2 twenty times and write them to /root/rest/jitter.txt, one per line. There must be at least 15 distinct values.

Full jitter waits a random number between 0 and the computed backoff. The value must differ each time even for the same attempt number.

Do not retry definitive errors

With /root/rest/retry400.py, call /bad. A 400 is not retried, so /root/rest/retry400.log must be exactly 1 line.

A 400 is a 400 however many times you send it. Add a branch that decides whether to retry by the status code.

Calculate the retry budget

In /root/rest/budget.txt, write four lines: instances=20, rps=100, max_attempts=5, and worst_rps=10000.

Number of instances x requests per second x maximum attempts is the worst-case requests per second the downstream receives. Write the three values and the result each.

Bring it together with a gateway that has a fallback

Start /root/rest/gateway.py on 127.0.0.1:8111. GET /order calls the downstream /always500, applies timeouts and retries, and returns 200 and {"degraded":true} on final failure.

Even if the downstream dies, give the user a 200, but leave a quality-degradation marker in the response. Use the timeout and retries from the previous steps together.