TT Lab
Get started
Learn Learning paths Courses

Load Testing

The Load Ladder and the Saturation Point

Continue in TT Lab

Goal

You measure the relationship between throughput and latency while raising concurrency like a ladder, and find the saturation point and the maximum point that still meets the SLA, in numbers. And you use Little's law to verify whether the measurement itself was valid. Put the outputs under /root/lt2/.

Why it matters

The quietest failure in load testing is "the load generator did not produce the target load". If one response takes 2 seconds, a synchronous generator does not send the requests it should have sent during those 2 seconds. The samples from the period when the system was slowest vanish entirely, and the generator ends up cooperating with the backpressure it created itself. As a result, raising the load does not move p99, and only production has a bad tail. So you attach a one-line acceptance check to every run. Does the reported actual throughput match the configured target throughput? If it does not, the latency distribution of that run cannot be trusted. Steps 6 and 7 of this lab are exactly that check.

Preparation — start the load target

This lab runs on an Ubuntu 24.04 VM, and Docker is installed as well. Even so, you start the target not as a container but as a Python process. Not because you cannot, but because that is the right way.

You must not start just any load target. It must be a target where the cost of one request is distinct for the ladder to look like a ladder and for the Little's law check of step 6 (rps × 평균 지연 ≈ 동시성, that is, rps times average latency is roughly equal to the concurrency) to hold. If you serve only static files with nginx, one request takes less than 1ms, so what you measure is not the server but the load generator itself — even if you climb the ladder, no knee appears, and if one does appear, it is not the target's limit but the limit of hey.

The target below is built to spend 50ms on each request. Only with that 50ms can you see the point where throughput rises and then stops as you raise the concurrency.

mkdir -p /root/lt2
cat > /root/lt2/target.py <<'PY'
import time
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
BODY = b"labhub load-test target
"
class H(BaseHTTPRequestHandler):
    protocol_version = "HTTP/1.1"
    disable_nagle_algorithm = True     # 헤더와 본문이 따로 나가면 지연 ACK 로 40ms 가 얹힌다
    def do_GET(self):
        time.sleep(0.05)               # 요청 하나의 처리 비용
        self.send_response(200); self.send_header("Content-Length", str(len(BODY)))
        self.end_headers(); self.wfile.write(BODY)
    def log_message(self, *a): pass
ThreadingHTTPServer(("127.0.0.1", 8085), H).serve_forever()
PY
setsid nohup python3 /root/lt2/target.py >/dev/null 2>&1 &
curl -s -o /dev/null -w '%{http_code}
' http://127.0.0.1:8085/    # 200

If you do want to move it into a container, you must put the same Python target in it, as in docker run -d --name lt-web -p 127.0.0.1:8085:8080 -v /root/lt2:/app python:3.12-alpine python /app/target.py. For the reason in the paragraph above, if you switch to nginx, the numbers of all the later steps lose their meaning.

Steps

  1. Write the ladder plan to /root/lt2/plan.md. It must include: five concurrency steps 1, 2, 5, 10, 20, the duration of each step, the warm-up plan, and the success criterion (for example, a p95 threshold).
  2. Create /root/lt2/ramp.sh and chmod +x it. It must take the output directory as the first argument ($1), run concurrency 1, 2, 5, 10, 20 in order, and save the results to $1/c<동시성>.txt (for example c10.txt; the placeholder is the concurrency).
  3. Run the script against /root/lt2/out. Five files, /root/lt2/out/c1.txt, c2.txt, c5.txt, c10.txt, and c20.txt, must be created, each file must have 100 or more total responses, and the Requests/sec and 95% in lines must remain.
  4. Create /root/lt2/ramp.csv. The format is concurrency,rps,p95,errors and you write five rows for concurrency 1/2/5/10/20. rps and p95 are the values read from the corresponding cN.txt, and errors is the number of non-2xx responses in the status code distribution (an integer, which must match exactly).
  5. Write one line, knee_concurrency=, to /root/lt2/knee.txt. The decision rule is as follows. Going through the concurrency in the order 1 → 2 → 5 → 10 → 20, find the first point where the RPS growth rate versus the previous step is under 10%, and write the concurrency just before it as the answer. If it grew by 10% or more all the way to the end, the answer is 20.
  6. Write four lines to /root/lt2/little.txt: concurrency=10, rps= (the Requests/sec of out/c10.txt), latency_avg_s= (the Average of the same file), and product= (rps × average latency, to two decimal places). If this product differs greatly from the concurrency of 10, that run did not produce the target load.
  7. Run one more run with a fixed arrival rate (open) and save it to /root/lt2/open.txt. Then write the following in /root/lt2/model.md: an explanation of the closed (fixed concurrency) model, an explanation of the open (fixed arrival rate) model, an explanation of coordinated omission, and two lines, target_rps= (the target throughput you set) and actual_rps= (the actual throughput read from open.txt).
  8. Write two lines to /root/lt2/max-safe.txt: max_safe_concurrency= and max_safe_rps=. They are the largest concurrency among the rows of ramp.csv where p95 is 0.2 seconds or less and errors is 0, and the RPS at that concurrency.

Notes

Write the load ladder plan

Write the ladder plan to /root/lt2/plan.md. It must include: five concurrency steps 1, 2, 5, 10, 20, the duration of each step, the warm-up plan, and the success criterion (for example, a p95 threshold).

Write the five concurrency steps, the duration of each step, the warm-up, and the success criterion in /root/lt2/plan.md. If you do not decide in advance what counts as a pass, you end up making the criterion after seeing the results.

Write the run script

Create /root/lt2/ramp.sh and chmod +x it. It must take the output directory as the first argument ($1), run concurrency 1, 2, 5, 10, 20 in order, and save the results to $1/c<동시성>.txt (for example c10.txt; the placeholder is the concurrency).

/root/lt2/ramp.sh takes the output directory as its first argument, runs the five steps in order, and saves the results as c.txt. If you hardcode the path, you cannot run it again into a different directory. Do not forget chmod +x.

Run the five steps

Run the script against /root/lt2/out. Five files, /root/lt2/out/c1.txt, c2.txt, c5.txt, c10.txt, and c20.txt, must be created, each file must have 100 or more total responses, and the Requests/sec and 95% in lines must remain.

The results for c1 c2 c5 c10 c20 must all be under /root/lt2/out, and each step must have 100 or more requests. The summary and latency distribution sections must remain in each file as they are so that you can read the values in the next step.

Make the results table

Create /root/lt2/ramp.csv. The format is concurrency,rps,p95,errors and you write five rows for concurrency 1/2/5/10/20. rps and p95 are the values read from the corresponding cN.txt, and errors is the number of non-2xx responses in the status code distribution (an integer, which must match exactly).

Organize rps, p95, and errors by concurrency in /root/lt2/ramp.csv. The values must be ones read from the original output, and errors is the number of non-2xx responses.

Find the saturation point

Write one line, knee_concurrency=, to /root/lt2/knee.txt. The decision rule is as follows. Going through the concurrency in the order 1 → 2 → 5 → 10 → 20, find the first point where the RPS growth rate versus the previous step is under 10%, and write the concurrency just before it as the answer. If it grew by 10% or more all the way to the end, the answer is 20.

Write one line, knee_concurrency=, to /root/lt2/knee.txt. The answer is the concurrency just before the point where throughput starts to stop growing meaningfully, and the decision rule is written exactly in the instructions.

Verify the run with Little's law

Write four lines to /root/lt2/little.txt: concurrency=10, rps= (the Requests/sec of out/c10.txt), latency_avg_s= (the Average of the same file), and product= (rps × average latency, to two decimal places). If this product differs greatly from the concurrency of 10, that run did not produce the target load.

Write in /root/lt2/little.txt the rps, the average latency, and their product for the concurrency 10 run. If the product differs greatly from the concurrency, the load generator did not actually produce the target load.

Compare the open and closed models

Run one more run with a fixed arrival rate (open) and save it to /root/lt2/open.txt. Then write the following in /root/lt2/model.md: an explanation of the closed (fixed concurrency) model, an explanation of the open (fixed arrival rate) model, an explanation of coordinated omission, and two lines, target_rps= (the target throughput you set) and actual_rps= (the actual throughput read from open.txt).

Run a separate run with a fixed arrival rate and save it to /root/lt2/open.txt, and explain in model.md the difference between the two models and coordinated omission. Write the target throughput and the actual throughput side by side to compare them.

Work out the maximum point that meets the SLA

Write two lines to /root/lt2/max-safe.txt: max_safe_concurrency= and max_safe_rps=. They are the largest concurrency among the rows of ramp.csv where p95 is 0.2 seconds or less and errors is 0, and the RPS at that concurrency.

Write to /root/lt2/max-safe.txt the largest concurrency that satisfies the conditions and the throughput at that point. The decision conditions are the two in the instructions, and you must pick them directly from the table.