It Wasn't One Request - Everything Got Slow
What to Record, and What to Wake Someone For
In one line
Loop delay and heap usage are visible accurately only from inside the process. CPU utilization seen from outside cannot tell a blocked process from a busy one. So extract the metrics from inside and keep them, and put alerts on the tail and the duration, not the average.
Why this was needed
While the loop is blocked, CPU utilization is close to 100%. But when the process is working very well, it is also 100%. Metrics seen only from outside cannot tell the two apart, so a "CPU is high" alert tells you nothing about this incident. Conversely, there are cases where the loop is blocked and cannot respond but the CPU is idle: when the thread pool is full and everything is waiting.
Latency logs fall short in the same way. Response time is the result of something that already happened, and to trace the cause back from that result, a person has to match up by eye "were other requests slow at the same time?" If you measure the time the loop was stopped itself inside the process, you skip that step. Instead of a list of slow requests, you are left with a single line: "at 12:04:20 the loop stalled for 410ms."
How it works
Node gives you two tools for this purpose. One is the
perf_hooks.monitorEventLoopDelay() we saw earlier, and the other is
performance.eventLoopUtilization(). They answer different questions.
monitorEventLoopDelay() gives you how late it was as a distribution. Because it is a histogram, you can pull out
percentile(99) and max directly; the values are in nanoseconds, and the default sampling interval is
10ms (the floor you saw in module 1 comes from here). If you read it periodically and call
reset(), it becomes a distribution per interval.
eventLoopUtilization() gives you how busy it was. The official documentation describes this value
as "the ratio of time the event loop spent outside the event provider (for example, epoll_wait)"
(perf_hooks).
It looks similar to CPU utilization but is a different value: it looks only at loop statistics and not at the CPU.
If you pass the result of a previous call as an argument, it returns the change in between, so
you do not have to do the subtraction yourself to get the utilization for an interval. This API was introduced in Node 14.10.0.
On the memory side, process.memoryUsage() takes care of it. rss is the physical memory the operating system has
attached to this process, and heapUsed is the amount V8 is actually using. In a backpressure
incident the two rise together, and in particular the Buffers piled up in a stream's buffer also show up
under external. In the measurement in module 4, the side that ignored backpressure grew RSS by
131MB and the side that respected it did not: same code, same input, a one-line difference.
const h = monitorEventLoopDelay({ resolution: 20 });
h.enable();
setInterval(() => {
const p99ms = h.percentile(99) / 1e6; // 나노초로 나온다
const maxMs = h.max / 1e6;
const rssMB = process.memoryUsage().rss / 1048576;
h.reset(); // 다음 구간을 위해 비운다
// 여기서 남긴다 — 로그 한 줄이든 지표 노출이든
}, 10000);
What it looks like in the field
Where to put the alert is the truly hard part. If you keep to three rules, you generally get an alert that is quiet yet useful.
First, do not put it on the average. As you saw in module 1, a single 300ms stall barely moves the average. Use p99 and max.
Second, do not wake someone up for a single spike. A large value shows up once in a while from compilation right after startup, a big GC that runs occasionally, and cache warming right after a deployment. Put duration in the condition, such as "if p99 stays over the threshold for 5 minutes or more." A single spike is kept in the record but does not wake anyone.
Third, decide the threshold by measuring. That is why each of you measured a baseline in this course's labs. If you do not know what the p50 is in a quiet state or where the p99 sits under normal load, the threshold is a guess. Only when there is a basis, such as a multiple of the baseline or a fraction of the response time budget, can a person later fix that number.
The same principle applies to memory alerts. The rate of increase is better than the absolute value. A backpressure incident climbs linearly toward the limit, so if you calculate "how many minutes from now at this rate" with the arithmetic you built in module 4, you can act before hitting the limit. An alert that arrives after hitting the limit is closer to an obituary for a process that has already been restarted.
Finally, these metrics must be looked at per process. If you start several workers or run a cluster, the overall average looks fine even when only one process is blocked. If you do not attach a process identifier to the metrics, you get the same kind of illusion you saw in module 4.
What you will check in the next quiz
You will check which question loop delay and utilization each answer, why CPU utilization seen from outside cannot single out this incident, and why an alert goes on the tail and the duration rather than the average. The numbers you measured yourself in the previous four modules serve directly as the basis.