TT Lab
Get started
Learn Learning paths Courses

It Wasn't One Request - Everything Got Slow

Slice It or Move It, and What Each Costs

Continue in TT Lab

In one line

There are only two ways to deal with a computation that holds the loop for a long time: split it into pieces and hand the turn back between them, or move the whole thing to another thread. Neither is free, and if you measure the values, "where do we put it?" becomes a calculation rather than a matter of taste.

Why this was needed

Finding the blocking code does not end the problem. That computation is still needed, and it has to be done somewhere. "Send it to a worker" is common advice, but people who follow it as is often have to undo it, because they miss two things.

One is that memory is not shared. A worker thread cannot see this request's objects as they are. Values are copied going in and copied coming back, so work that has to touch the connection or cache you are holding cannot be moved in the first place. If you pass a large object, the copying itself becomes a new cost.

The other is that starting one takes time. Measured with an empty job on Node 22.11.0 in the lab image, starting one worker and getting an answer back took 22ms to 31ms. If you move a 5ms computation, you lose five times over. That is why real services start a few workers ahead of time and reuse them, and this number is why that design is needed.

Splitting has the opposite properties. It stays on the same thread, so it sees state as it is and there is no copying. In exchange the total time grows a little, and more importantly, the time of one slice becomes the floor of the latency.

How it works

For splitting, "divide" is half and "yield" is the other half. If you only divide the loop into slices and do nothing in between, from the loop's point of view it is the same as running it all at once. The yield is one line.

for (const { from, to } of chunks(total, sliceSize)) {
  result = hashRange(from, to, result);          // 앞 조각의 값을 이어받는다
  await new Promise((resolve) => setImmediate(resolve));   // 여기서 차례를 넘긴다
}

setImmediate is called in the check phase, so between these lines the timers that were waiting and the completed I/O callbacks get processed. On Node 22.11.0, running a 40-million-iteration computation in one go stalls the loop for 290ms, and splitting it into 40 slices brings that down to 8ms. The total time grew by a factor of 1.03, which is the price of going around the loop once per yield.

This gives you one design knob. While a slice is running, no one can cut in, so once you set a target latency, the slice size follows from it. In the measurement above, one slice took 7.9ms and the maximum measured latency was 8ms. If you want to respond within 50ms, you just make sure one slice does not exceed 50ms. It is not "split it reasonably"; you can decide it by calculation.

Moving has a different shape. The worker_threads module's Worker takes a file, runs that module in a new thread, passes values in with workerData, and gets the result back with postMessage. While the computation runs there, the loop on this side is completely free. When I sent the same 40-million-iteration computation to a worker, the maximum latency on this side was 9ms.

const worker = new Worker(new URL("./hash-worker.mjs", import.meta.url),
                          { workerData: { from: 0, to: total } });
worker.once("message", (result) => { /* 복사되어 건너온 값 */ });

There is one trap. A worker inherits the execution arguments of the parent process. So if you start a file-based worker inside a program run with node --input-type=module -e "...", it dies with ERR_INPUT_TYPE_NOT_ALLOWED. It is safer to keep a program that uses workers in a file and run it as is.

What it looks like in the field

The decision rule ends up simpler than you might expect. If it takes under 1ms at a time, just leave it: the cost of measuring is greater than the value of fixing. If it is heavier than that but needs to see the state of the current request, it cannot be moved, so you split it. If it is a pure computation independent of state and heavy enough to pay the startup cost, you move it. The boundary for "heavy enough" comes from the startup cost, so on a machine where startup takes 22ms, somewhere around 50ms is reasonable, and on another machine it will be a different value. The rule is tied to measurement; it is not a constant to memorize.

A failure you often see in the field is code that only imitates splitting. Someone changes a loop from for to map, or declares a function async, and thinks they have split it, but neither hands the turn back to the loop. Even inside an async function, the code runs synchronously to the end until it meets an await. There is one way to check whether you yielded: count how many times a short timer fired while the computation was running.

What you will do in the next lab

You will run the same computation in three ways: all at once, split into slices, and in a worker thread. All three must give the same answer (splitting does not change the answer), and the latency distributions should differ greatly.

Then you will calculate the value of each: the increase in total time from splitting, the time of one slice, and the startup cost of the worker. Finally you will harden those numbers into a rule called place(job), so that the next person can make the same judgment.