Turning "It's Slow" Into a Number
One-line summary
You cannot fix anything with "it's slow." The moment you turn it into what · when · how much · what percentage is slow, the scope of the investigation shrinks to a tenth.
Why this is needed
The report comes in like this. "The system has been slow lately."
The most common mistake here is to log into the server right away and look at CPU. If the CPU is 30%, it ends with "it looks normal," and if it is 90%, it wrongly ends with "it's the CPU." Neither is an answer.
There are four things to ask first.
| Question | Why | What the answer narrows down |
|---|---|---|
| Which screen/feature? | It is rare for everything to be slow | Code path |
| Since when? | Compare with deployments and configuration changes | Point in time of the cause |
| How much? (before 2 seconds, now 20 seconds) | Whether it is 10x or 10% | Nature of the problem |
| Always? Sometimes? | Whether it is an average problem or a tail problem | Investigation method |
Especially the fourth. Always slow and sometimes slow have different causes. If always, it is structure (queries, N+1, synchronous calls), and if sometimes, it is contention (locks, connection pools, GC, disk). If you look at the two mixed together, they get buried in the average and nothing shows.
Split the segments
You divide the path a single request travels and measure the time of each segment.
클라이언트 → [네트워크] → 웹서버 → [앱] → [DB] → 앱 → 웹서버 → 클라이언트
Even without measurement tools, curl alone can split the segments.
curl -o /dev/null -s -w \
'dns=%{time_namelookup} conn=%{time_connect} tls=%{time_appconnect} ttfb=%{time_starttransfer} total=%{time_total}\n' \
https://example.com/slow-page
dnsis large → name resolution. Resolver configuration, cache misses.connis large → TCP connection. Path, firewall, SYN waiting.tlsis large → handshake. Certificate chain, OCSP lookup.ttfbis large andtotal - ttfbis small → the time the server spends thinking. App/DB.total - ttfbis large → body transfer. The response is large or there is a bandwidth problem.
These five numbers alone decide "network or server." Most arguments end here.
Do not trust the average
It is common for users to say "it's slow" on a service with an average response of 200 ms. If 950 out of 1000 requests take 50 ms and 50 take 3 seconds, the average is 197 ms. The average is normal, but 1 in 20 people waits 3 seconds.
That is why you look at percentiles.
| Metric | Meaning | Use |
|---|---|---|
| p50 (median) | Half are faster than this | Typical experience |
| p95 | The value 1 in 20 people experiences | Where perceived complaints begin |
| p99 | 1 in 100 | The tail. Contention, GC, retries |
| max | The single worst one | Tracking outliers |
If p50 stays the same and only p99 rose, it is not capacity but contention. If it rose starting from p50, it is structure or capacity.
Load and latency are different axes
60% CPU utilization does not mean "40% headroom." In queueing theory, as utilization rises, wait time grows not linearly but sharply. Past 70%, a little more traffic multiplies the latency several times. So "but we still have CPU headroom" is no rebuttal to a latency problem.
What you see in the field
- "The network seems slow" → If you measure
ttfb, most of it is server processing time. - Slow only at 9 a.m. → Simultaneous logins at the start of the workday + a time when caches are empty.
- Slow only for a particular customer → That customer's data volume is different. The query does a full scan.
What to measure after narrowing down
If you have split the segments and learned that it is on the server side, next you look at what is waiting in there. A common mistake here is judging by CPU utilization alone, and as seen in the previous course, most of the time a server waits is not waiting for its turn on the CPU.
You rule out four things in order.
First, waiting to run. See whether more things are ready to run than there are cores. If the load average is high but CPU utilization is low, it is not this. For containers, you must also look at the throttling seen earlier. It is the classic cause where average utilization is low but only tail latency spikes.
Second, waiting for I/O. This is time spent waiting for disk or network. A slow database is also usually here, and what to fix then is not the application but the query or the index.
Third, locks and wait queues. Connection pools, thread pools, application locks, and database row locks. If it slows down only when concurrent requests increase, it is almost always here, and a hallmark is that it does not reproduce when tested alone.
Fourth, external calls. Another service we call is slow. Then the only things to fix on our side are timeouts, retries, and fallbacks, so cause and response are separated.
The cheapest way to tell these four apart is to change the load. If it is slow even with a single request, it is a structural problem (the first or second), and if it slows down only when several are sent at once, it is contention (the third). This one comparison cuts the scope of the investigation in half.
And after the fix, measure again the same way. Believing you fixed it is different from the numbers having changed, and without the value measured at the start, you have no basis to say it got better.
What you will do in the following lab
You receive 340 real request log lines and answer the four questions yourself. Since it is data that contains all the answers, just cutting it in order turns the one line "it's slow" into a report that points to a single version.