TT Lab
Get started
Learn Learning paths Courses

Load Testing

Sizing Is About Getting the Order of Magnitude Right

Continue in TT Lab

In one line

The purpose of capacity estimation is not an exact number but getting the order of magnitude right. Whether you need 3 instances or 30 changes the whole design.

Why this matters

System design proceeds in four steps. Clarify requirements → estimate scale → high-level design → detailed design (finding and solving bottlenecks). If you skip the second step, everything else becomes fantasy.

Take a simple example. If writes are 100 million per day, QPS is about 1,160 (100 million ÷ 86,400). If reads are 10 times that, QPS is about 11,600. Five years of 180 billion records at 500 bytes per record is about 90 TB. Once these three numbers come out, you can answer questions like "is a single instance enough", "is sharding needed", and "is it possible without a cache" with evidence.

How it works

On the road from measurement to arithmetic, there are two safeguards.

The first is headroom. You must not use the measured maximum throughput as it is. Take the safe throughput as about 70% of the measured value. The remaining 30% is room for traffic fluctuation, reduced instances during a deployment, and unexpectedly slow requests. If you measured 190 RPS per instance, the safe throughput is 133 RPS.

The second is rounding up. Dividing the target of 300 RPS by 133 gives 2.25 instances, and 2 is not enough, so it is 3. In a capacity calculation, the moment you round down the fractional part, the plan fails.

The same thinking is needed in pool sizing. If you set the recommended pool size per instance to 60, you must multiply it out at the fleet level. 40 instances × 60 = 2,400 connections, and if the database's max_connections is 200, every instance's calculation was right but the system dies right after deployment. It is a classic case where the local optimum is not the global optimum.

What it looks like in the field

The most common flaw in capacity documents is the absence of conditions. The sentence "our service handles 500 requests per second" has no concurrency, no p95, no error rate, and no time of measurement. If the same service delivers 500 RPS at concurrency 2 with a p95 of 20ms and 520 RPS at concurrency 50 with a p95 of 900ms, the second number is not capacity but an observation from an already saturated state.

So capacity must always be written together with the SLA. It should be written like "133 RPS per instance under the condition of p95 200ms or less and zero errors" so that the next person can reproduce it with the same criteria.

The four kinds of load tests

Distinguishing the names makes it clear what is being measured.

Type What it does What you learn
Load Sustain the expected traffic Latency and error rate at the target load
Stress Keep raising the load The point where it breaks and how it breaks
Spike Suddenly tenfold Whether autoscaling keeps up
Soak Usual load for several hours Memory leaks, connection leaks, disks filling

The value of a stress test is not the limit figure but how it breaks. Whether it rejects gracefully (429), everything slows down, or it dies shows the quality of the design. "It handles up to 5,000 requests per second" is far less useful than "above 5,000 it returns 429, and when it comes back down to 4,800 it recovers."

If you skip the soak test, you miss problems that appear an hour later. Connection pools leaking, file descriptors piling up, and GC getting longer and longer do not show up in a 5-minute test.

Closed loop and open loop

There are two ways a load tool creates requests, and the results are completely different.

폐 루프(closed): 가상 사용자 N 명이 응답을 받아야 다음 요청을 보낸다
                 → 서버가 느려지면 부하도 저절로 줄어든다
개 루프(open):   초당 N 건을 서버 상태와 무관하게 보낸다
                 → 서버가 느려지면 요청이 쌓인다. 현실에 가깝다

Real users are closer to an open loop. Because people do not reduce their requests just because the server is slow. If you test only with a closed loop, good numbers like "latency 200ms at 100 virtual users" come out, but in reality it has already collapsed at that point.

k6 builds an open loop with constant-arrival-rate, and Gatling with constantUsersPerSec. To test autoscaling, it must be an open loop.

Coordinated omission

This is a more subtle problem of closed-loop tools. When a response is delayed, the requests that could not be sent during that time never enter the measurement at all, so p99 comes out much better than reality.

서버가 10초 멈췄다
  폐 루프: 그동안 요청을 안 보냄 → 느린 요청 1건만 기록 → p99 정상
  현실:    그동안 요청이 계속 옴 → 수천 건이 10초를 기다림 → p99 폭발

Either use a tool that corrects for this (--latency-correction, based on HdrHistogram), or test with an open loop. If p99 looks suspiciously good, suspect this first.

The order of judgment in the field

  1. Decide the target traffic (peak-based, not average).
  2. Measure one instance's safe throughput under the SLA conditions.
  3. Reflect headroom by reducing it to 70%.
  4. Divide and round up.
  5. Check again by multiplication against the limits of shared resources (DB connections, cache, queues).

What you will look at in the next check

This module ends with a quiz that checks concepts and calculations, without a lab. You check the process of converting the numbers from the baseline, ladder, and bottleneck report built in the previous three modules into average QPS, peak multiple, safe throughput reflecting headroom, and the required number of instances.