A Generous Timeout Makes the Outage Bigger
Summary
A timeout is not a number set separately per call but handing down to the lower levels the budget given to one upstream request, and when the budget is short, failing right away without calling is the cheapest choice.
Why this was needed
The statement "let's set the timeout generously" seems safe, but it is the opposite. You can see by following what happens when the other side gets slow.
The other side's response is normally 200ms, but one day it becomes 30 seconds. If our timeout is 60 seconds, our worker is tied to that call for 30 seconds. If there are 50 workers, throughput drops sharply, and from the point of view of whoever calls us, we have become slow. If their timeout is generous too, the same thing happens again one layer up. Because of this phenomenon, where one system's delay spreads upward, a timeout is not "a courtesy of waiting for the other side" but a circuit breaker that protects us.
The second common mistake here is setting the timeout as a single value. The advanced usage documentation of Python requests says the timeout can be given as two values, (connect, read). The reason to separate them is that their meanings are completely different.
- The connect limit is the time it takes to reach the other side. Within the same data center it should be in milliseconds, and if it takes several seconds, that is not slow but not working. So set the connect limit short.
- The read limit is the time the other side spends working. It differs by business, and in some cases it must be generous.
The shape of a connection failure also differs. Connection refused ends almost immediately because the packet went to the destination and came back, and when nobody answers, it waits up to the limit and then times out. For experiments, use an address like 192.0.2.0/24, which RFC 5737 reserves for documentation — since it is routed nowhere, packets silently vanish, and you can safely see what a connect timeout looks like.
How it works
If you think in terms of a budget, the rules reduce to three lines.
First, set a total budget for the upstream request. If a screen must be drawn within 2 seconds, that request's budget is 2000ms.
Second, each sub-call's read limit is the budget remaining at that moment. 2000ms for the first call, and if it uses 300ms, 1700ms for the next. This way no call can exceed the total budget.
Third, if the remaining budget is smaller than the lower bound, do not call. Making a call that normally takes 400ms when 200ms remain only postpones the failure by 200ms. If you fail on the spot, you can build even a partial response with that 200ms.
Retries eat the budget. A great many implementations miss this calculation. With a 2-second read limit and 3 retries, the worst case is 6 seconds plus the wait times. Waits should grow exponentially with randomness mixed in as the standard (it prevents clients that failed at the same moment from piling back at the same moment), but that wait also comes out of the budget. So right before a retry you ask "if I wait now and call once more, will it fit within the budget?" and stop if not.
Pass the remaining budget down. If we have only 1200ms left and a lower service waits 5 seconds by its own standard, those 4 seconds are spent producing an answer nobody is looking at. So send the remaining budget in a header and have the receiving side use it as its own limit. On the status code side, the 504 (Gateway Timeout) defined by RFC 9110 means "the upstream waited and cut off," and the 429 (Too Many Requests) of RFC 6585 means "slow down." They are different signals and call for different responses.
예산 2000ms
├─ 호출 A 남은 2000 → 300ms 사용
├─ 호출 B 남은 1700 → 250ms 사용
├─ 재시도 남은 1450 → 대기 200 + 호출 400
└─ 호출 C 남은 850 → 하한 1000 미만이면 걸지 않고 즉시 실패
What it looks like in the field
First, it is common that the timeout you set does not actually apply. If the app setting is 30 seconds but the failure happened at 127 seconds, that value is not applied and the kernel's SYN retransmissions were exhausted. Before fixing the setting, measure at how many seconds it actually cuts off.
Second, timeouts and connection pools move together. If the pool's wait time has no limit either, even if the call's own timeout is short, the total time gets long waiting in the pool.
Third, you do not decide in advance what to do when the budget is exceeded. There are usually three options. Give a partial response, give the old value from a cache, or give a failure. It does not matter which you choose, but if you do not choose, the code chooses the worst one — waiting to the end.
Fourth, you do not record the basis of the number. If it is not written where "read 3 seconds" came from, nobody can change it. It is enough to write one line saying it was set by multiplying the other side's observed 99th percentile by a margin.
What you will do in the next lab
You start a partner server that gets slow and measure for yourself the three failure shapes: connection refused, connect timeout, and read timeout. You build a caller that sets the connect and read limits separately, put a caller that keeps a total budget on top of it so that the remaining budget becomes the next call's limit. If the budget falls short of the lower bound, it skips without calling, and you include the calculation of retries eating the budget. Finally you pass the remaining budget down in a header and write the budget table of one screen in numbers.