A Network Call Is Not a Function Call
Summary
The three things you must set as defaults for a synchronous call are the timeout, the conditions for retrying, and the conditions for not retrying.
Why this was needed
A function call inside a monolith had one failure mode. It throws an exception or returns a value. A network call has five failure modes. A connection failure, connected but no response, a slow response, a response that only half arrived, and the nastiest one — the other side processed it but only the response was lost.
The reason the last case is nasty is that the caller has no way to tell "it did not happen" from "it happened but I don't know". This indistinguishability is the starting point of the idempotency and sagas that come later.
How it works
Let us start with the timeout. Many HTTP clients have no default, and without one it is effectively an infinite wait. An infinite wait is dangerous because threads or connections get tied up. If the downstream slows down by 5 seconds each, the upstream's connection pool dries up first. Then even requests entirely unrelated to the downstream fail. This is the most common starting point of a cascading failure.
Do not guess at the timeout value; start from the other side's p99. If the other side's p99 is 120ms, something around 300–500ms is reasonable. If you set it to 10 times the p99, the timeout might as well not exist.
Retries have conditions. Retrying is meaningful only for transient errors — a connection failure, a timeout, 503, 429. Definitive errors such as 400, 401, and 404 give the same answer however many times you send them. And retries must always use exponential backoff with jitter. Fixed-interval retries create a thundering herd.
It is good to remember one number here. If the order service has 20 instances each handling 100 requests per second and you set maxAttempts to 5, then the moment the downstream wobbles, the payment service receives up to 10,000 requests per second. 100 x 20 x 5. Retries multiply the load.
Does it change if you use gRPC? Serialization gets faster and streaming appears, but the three problems above remain as they are. However, it is better than REST in that the deadline is a first-class concept of the protocol and propagates across hops.
What you meet in the field
When the call chain is A → B → C, a common mistake is to set the timeouts to 3 seconds each. Then A waits at most 3 seconds, but B waits 3 seconds for C, so even if A's timeout fires first, B and C keep working. They spend resources on work that will be discarded. Timeouts must be large on the outside and small on the inside. And it is even better to pass the remaining budget down through a header or a deadline.
Timeouts differ by layer
When a request passes through several layers, the timeout must get shorter from the outside to the inside. If it is the reverse, the outside gives up first, and the inside keeps producing answers nobody is waiting for.
브라우저 30s
게이트웨이 10s ← 바깥보다 짧다
API 6s
결제 서비스 3s
DB 1s
If each layer has retries, they multiply. If the API has a 3-second timeout with 2 retries, the worst case is 9 seconds, using almost all of the gateway's 10 seconds. Timeout × (retries + 1) must be smaller than the outer timeout.
There is also a way to pass the deadline in a header (deadline propagation). gRPC supports it by default, and in HTTP you create a header such as X-Request-Deadline yourself. When the remaining time is close to 0, not starting at all saves resources.
The conditions under which a circuit breaker opens
Retries alone cannot revive a collapsed downstream. They actually add load. A circuit breaker stops the attempts themselves when failures continue.
닫힘(정상) ──실패율 임계 초과──> 열림(즉시 실패)
↑ │
└──성공── 반열림(몇 개만 보내 봄) <──일정 시간 뒤──┘
You set the threshold not by count but by rate and a minimum sample. "5 failures out of 10" is meaningful, but "1 failure out of 2" is coincidence.
# 예: 20건 이상일 때, 실패율 50% 넘으면 30초 열림
minimumNumberOfCalls: 20
failureRateThreshold: 50
waitDurationInOpenState: 30s
permittedNumberOfCallsInHalfOpenState: 3
What to return when it is open is the design. A cached old value, a reduced response, or a clear error. If you just throw a 500 with no plan, the circuit breaker might as well not exist.
Putting partial failure in the response
When you call several services to build one screen, you do not fail the whole thing because one failed. A response that tells what is missing is better.
{
"order": {"id": 1043, "amount": 52000},
"recommendations": null,
"degraded": ["recommendations"]
}
The screen empties only the recommendations area and shows the rest. To do this, you must separate required from optional in advance. Order information is required, and recommendations are optional. Without this distinction, every call becomes required and availability gets multiplied.
What you will do in the next lab
Against a deliberately slow and deliberately failing downstream, you set a timeout, implement exponential backoff and jitter, make 4xx errors not be retried, and finally calculate the retry budget.