TT Lab
Get started
Learn Learning paths Courses

Integration and Deployment

They Slowed Down and We Died First

Continue in TT Lab

Goal

You treat the time given to one upstream request as a budget, and build a client that divides that budget among connect, read, retry, and sub-calls. When the budget is short, it fails fast without calling, and it passes the remaining budget down in a header.

Why it matters

The habit of setting timeouts generously seems safe but is the opposite. If the other side slows to 30 seconds and our limit is 60 seconds, our worker is tied up for 30 seconds, and from the point of view of whoever calls us, it is we who have slowed. That delay spreads another layer up. So a timeout is not a courtesy of waiting for the other side but a circuit breaker that protects us. And it is not one value but two — connect is the time to reach, so if it takes several seconds it is not slow but not working, and read is the time the other side works, so it differs by business. If you think in terms of a budget, the rules are three lines. Set a total budget for the upstream request, make each call's read limit the budget remaining at that moment, and do not call if the remaining budget is smaller than the lower bound. Retries and their wait times come out of the same budget. The grader does not believe your sentences. It starts a slowing partner directly on a port the grader chooses, runs your caller for real, and measures the elapsed time and the phase names in the output together.

Steps

  1. Create /root/budget/partner.py, run it on port 8014, measure the three failure shapes, and write them in /root/budget/probe.json.
  2. Create /root/budget/call.py so that it sets the connect and read limits separately and says in which phase it was cut.
  3. Create /root/budget/deadline.py so that it makes several calls with one total budget and the remaining budget becomes the next call's read limit.
  4. Attach a lower bound (--min-ms) to deadline.py so that if the remaining budget is smaller than the lower bound, it does not call and skips as no_budget.
  5. Create /root/budget/retry.py so that the retry wait grows exponentially and is scattered randomly, but that wait is also subtracted from the budget.
  6. Create /root/budget/fanout.py so that it passes the remaining budget at that time to sub-calls in the X-Budget-Ms header.
  7. Write the budget table of one screen in numbers in /root/budget/budget_plan.json.
  8. Report in four sections in /root/budget/budget_report.md.

Notes

Measure the three shapes of failure

Create /root/budget/partner.py, run it on port 8014, measure connection refused, connect timeout, and read timeout yourself, and write them in /root/budget/probe.json as three entries with kind, target, elapsed_ms, and error.

The partner imitates the slow response with sleep. Make the connection refusal with a 127.0.0.1 port nobody is listening on, and the connect timeout with an address routed nowhere. How the elapsed times of the three cases differ from each other is the core of this step.

Set connect and read separately

Create /root/budget/call.py so that it takes --connect and --read separately and, on failure, writes in phase whether it was connect or read. If a response was received, the phase is done even if the status code is 4xx or 5xx.

The timeout of requests takes a pair holding two values. The exceptions are divided by kind as well, so you can distinguish the connect side from the read side — but a connection refusal is not a timeout exception, so you must catch it separately, and it too is something that ended before reaching the other side, so it is the connect phase.

The remaining budget is the next call's limit

Create /root/budget/deadline.py so that with a single --budget-ms it calls several --url values in turn. Each call's read limit must be the budget remaining at that moment, and you must write that value in read_timeout_ms.

If you import the previous step's call.py as a module, you do not have to write the same code twice. Take the start time once and compute "budget minus time used so far" just before each call. This way, however many calls there are, the whole cannot exceed the budget.

Do not make a hopeless call

Attach --min-ms to deadline.py so that if the remaining budget is smaller than that value, it does not make the call, records phase as no_budget, and increments skipped. The elapsed_ms of a skipped call is 0.

Making a call that normally takes 400ms when 200ms remain only postpones the failure by 200ms. If you fail on the spot, you can build even a partial response with the remaining time. The lower bound must be taken as an argument so it can be given differently in each situation.

Retries come out of the budget too

Create /root/budget/retry.py so that it calls a failed call again, growing the wait limit as base-ms * 2^(시도-1) (the exponent being the attempt number minus 1) and choosing the actual wait between 0 and that limit. If after waiting there is no room in the budget to call once more, set stopped_reason to budget and stop.

If you do not mix in randomness, clients that failed at the same moment pile back at the same moment. And since the wait time comes out of the budget too, you must ask before sleeping whether there will be room left to make the call after waking. If you write the reason for stopping as three kinds, later you know the cause from the log alone.

Pass the remaining budget down

Create /root/budget/fanout.py so that it calls <BASE>/budget?ms=200 --n times, each time carrying the remaining budget at that time in the X-Budget-Ms header. In each call of the output, write sent_budget_ms, received, and elapsed_ms.

If we have only 1200ms left and a lower service waits 5 seconds by its own standard, those 4 seconds are spent producing an answer nobody is looking at. The header name is a convention we decide, and the receiving side uses that value as its own limit. The partner's /budget returns the received header value as is, so you can check that it was actually passed on.

The budget table of one screen

Write the budget table in /root/budget/budget_plan.json. It needs total_ms, reserve_ms, worst_case_ms, and 3 or more calls (name, budget_ms, max_attempts). reserve_ms must be 100 or more, and worst_case_ms is the sum of the call budgets plus the reserve and must not exceed total_ms.

The reserve is time that is not calls, such as our own serialization and template rendering. If you leave this out and divide the whole budget among the calls, the screen is always on the edge of the limit. A call that has retries means it must finish its retries within that call's budget.

Budget inspection report

In /root/budget/budget_report.md, write four sections, ## 지금 무엇이 시간을 먹는가, ## 구간을 어떻게 나눴나, ## 재시도가 먹는 몫, and ## 예산을 넘겼을 때 무엇을 하나 (in order: what eats the time now, how the segments were divided, the share retries eat, what to do when the budget is exceeded). The numbers from probe.json and budget_plan.json must be in the body.

The reader is someone who said "please increase the timeout a bit." Show with numbers what gets worse if you increase it. For the retry share, multiply the number of attempts by the wait limit and write the worst case.