The Load Ladder and the Saturation Point
Goal
You measure the relationship between throughput and latency while raising concurrency like a ladder, and find the saturation point and the maximum point that still meets the SLA, in numbers. And you use Little's law to verify whether the measurement itself was valid. Put the outputs under /root/lt2/.
Why it matters
The quietest failure in load testing is "the load generator did not produce the target load". If one response takes 2 seconds, a synchronous generator does not send the requests it should have sent during those 2 seconds. The samples from the period when the system was slowest vanish entirely, and the generator ends up cooperating with the backpressure it created itself. As a result, raising the load does not move p99, and only production has a bad tail. So you attach a one-line acceptance check to every run. Does the reported actual throughput match the configured target throughput? If it does not, the latency distribution of that run cannot be trusted. Steps 6 and 7 of this lab are exactly that check.
Preparation — start the load target
This lab runs on an Ubuntu 24.04 VM, and Docker is installed as well. Even so, you start the target not as a container but as a Python process. Not because you cannot, but because that is the right way.
You must not start just any load target. It must be a target where the cost of one request is distinct for the ladder to look like a ladder and for the Little's law check of step 6 (rps × 평균 지연 ≈ 동시성, that is, rps times average latency is roughly equal to the concurrency) to hold. If you serve only static files with nginx, one request takes less than 1ms, so what you measure is not the server but the load generator itself — even if you climb the ladder, no knee appears, and if one does appear, it is not the target's limit but the limit of hey.
The target below is built to spend 50ms on each request. Only with that 50ms can you see the point where throughput rises and then stops as you raise the concurrency.
mkdir -p /root/lt2
cat > /root/lt2/target.py <<'PY'
import time
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
BODY = b"labhub load-test target
"
class H(BaseHTTPRequestHandler):
protocol_version = "HTTP/1.1"
disable_nagle_algorithm = True # 헤더와 본문이 따로 나가면 지연 ACK 로 40ms 가 얹힌다
def do_GET(self):
time.sleep(0.05) # 요청 하나의 처리 비용
self.send_response(200); self.send_header("Content-Length", str(len(BODY)))
self.end_headers(); self.wfile.write(BODY)
def log_message(self, *a): pass
ThreadingHTTPServer(("127.0.0.1", 8085), H).serve_forever()
PY
setsid nohup python3 /root/lt2/target.py >/dev/null 2>&1 &
curl -s -o /dev/null -w '%{http_code}
' http://127.0.0.1:8085/ # 200
If you do want to move it into a container, you must put the same Python target in it, as in docker run -d --name lt-web -p 127.0.0.1:8085:8080 -v /root/lt2:/app python:3.12-alpine python /app/target.py. For the reason in the paragraph above, if you switch to nginx, the numbers of all the later steps lose their meaning.
Steps
- Write the ladder plan to
/root/lt2/plan.md. It must include: five concurrency steps 1, 2, 5, 10, 20, the duration of each step, the warm-up plan, and the success criterion (for example, a p95 threshold). - Create
/root/lt2/ramp.shandchmod +xit. It must take the output directory as the first argument ($1), run concurrency 1, 2, 5, 10, 20 in order, and save the results to$1/c<동시성>.txt(for examplec10.txt; the placeholder is the concurrency). - Run the script against
/root/lt2/out. Five files,/root/lt2/out/c1.txt,c2.txt,c5.txt,c10.txt, andc20.txt, must be created, each file must have 100 or more total responses, and theRequests/secand95% inlines must remain. - Create
/root/lt2/ramp.csv. The format isconcurrency,rps,p95,errorsand you write five rows for concurrency 1/2/5/10/20.rpsandp95are the values read from the correspondingcN.txt, anderrorsis the number of non-2xx responses in the status code distribution (an integer, which must match exactly). - Write one line,
knee_concurrency=, to/root/lt2/knee.txt. The decision rule is as follows. Going through the concurrency in the order 1 → 2 → 5 → 10 → 20, find the first point where the RPS growth rate versus the previous step is under 10%, and write the concurrency just before it as the answer. If it grew by 10% or more all the way to the end, the answer is 20. - Write four lines to
/root/lt2/little.txt:concurrency=10,rps=(the Requests/sec ofout/c10.txt),latency_avg_s=(the Average of the same file), andproduct=(rps × average latency, to two decimal places). If this product differs greatly from the concurrency of 10, that run did not produce the target load. - Run one more run with a fixed arrival rate (open) and save it to
/root/lt2/open.txt. Then write the following in/root/lt2/model.md: an explanation of the closed (fixed concurrency) model, an explanation of the open (fixed arrival rate) model, an explanation of coordinated omission, and two lines,target_rps=(the target throughput you set) andactual_rps=(the actual throughput read from open.txt). - Write two lines to
/root/lt2/max-safe.txt:max_safe_concurrency=andmax_safe_rps=. They are the largest concurrency among the rows oframp.csvwhere p95 is 0.2 seconds or less and errors is 0, and the RPS at that concurrency.
Notes
- An open model approximation:
-qofheylimits requests per second per worker.hey -z 30s -c 50 -q 4 http://127.0.0.1:8085/means a target of 200 RPS (50 × 4), so writetarget_rps=200and compare it with the actual value. - Script skeleton:
for c in 1 2 5 10 20; do hey -n 300 -c "$c" http://127.0.0.1:8085/ > "$1/c$c.txt" 2>&1; done. Putmkdir -p "$1"before it. - Count
errorswithawk '/Status code distribution/{f=1;next} f && /responses/ {gsub(/[][]/," "); if($1<200||$1>=300) e+=$2} END{print e+0}'. If there are no errors, write0. - Common mistake 1: hardcoding the output path inside
ramp.sh. If you do not use the first argument, it is caught in grading. - Common mistake 2: writing "the concurrency with the highest RPS" as the knee in step 5. The knee is not the peak but the point just before growth starts to stop.
- Common mistake 3: running each step too short. With fewer than 100 requests, percentiles have no meaning.
Write the load ladder plan
Write the ladder plan to /root/lt2/plan.md. It must include: five concurrency steps 1, 2, 5, 10, 20, the duration of each step, the warm-up plan, and the success criterion (for example, a p95 threshold).
Write the five concurrency steps, the duration of each step, the warm-up, and the success criterion in /root/lt2/plan.md. If you do not decide in advance what counts as a pass, you end up making the criterion after seeing the results.
Write the run script
Create /root/lt2/ramp.sh and chmod +x it. It must take the output directory as the first argument ($1), run concurrency 1, 2, 5, 10, 20 in order, and save the results to $1/c<동시성>.txt (for example c10.txt; the placeholder is the concurrency).
/root/lt2/ramp.sh takes the output directory as its first argument, runs the five steps in order, and saves the results as c.txt. If you hardcode the path, you cannot run it again into a different directory. Do not forget chmod +x.
Run the five steps
Run the script against /root/lt2/out. Five files, /root/lt2/out/c1.txt, c2.txt, c5.txt, c10.txt, and c20.txt, must be created, each file must have 100 or more total responses, and the Requests/sec and 95% in lines must remain.
The results for c1 c2 c5 c10 c20 must all be under /root/lt2/out, and each step must have 100 or more requests. The summary and latency distribution sections must remain in each file as they are so that you can read the values in the next step.
Make the results table
Create /root/lt2/ramp.csv. The format is concurrency,rps,p95,errors and you write five rows for concurrency 1/2/5/10/20. rps and p95 are the values read from the corresponding cN.txt, and errors is the number of non-2xx responses in the status code distribution (an integer, which must match exactly).
Organize rps, p95, and errors by concurrency in /root/lt2/ramp.csv. The values must be ones read from the original output, and errors is the number of non-2xx responses.
Find the saturation point
Write one line, knee_concurrency=, to /root/lt2/knee.txt. The decision rule is as follows. Going through the concurrency in the order 1 → 2 → 5 → 10 → 20, find the first point where the RPS growth rate versus the previous step is under 10%, and write the concurrency just before it as the answer. If it grew by 10% or more all the way to the end, the answer is 20.
Write one line, knee_concurrency=, to /root/lt2/knee.txt. The answer is the concurrency just before the point where throughput starts to stop growing meaningfully, and the decision rule is written exactly in the instructions.
Verify the run with Little's law
Write four lines to /root/lt2/little.txt: concurrency=10, rps= (the Requests/sec of out/c10.txt), latency_avg_s= (the Average of the same file), and product= (rps × average latency, to two decimal places). If this product differs greatly from the concurrency of 10, that run did not produce the target load.
Write in /root/lt2/little.txt the rps, the average latency, and their product for the concurrency 10 run. If the product differs greatly from the concurrency, the load generator did not actually produce the target load.
Compare the open and closed models
Run one more run with a fixed arrival rate (open) and save it to /root/lt2/open.txt. Then write the following in /root/lt2/model.md: an explanation of the closed (fixed concurrency) model, an explanation of the open (fixed arrival rate) model, an explanation of coordinated omission, and two lines, target_rps= (the target throughput you set) and actual_rps= (the actual throughput read from open.txt).
Run a separate run with a fixed arrival rate and save it to /root/lt2/open.txt, and explain in model.md the difference between the two models and coordinated omission. Write the target throughput and the actual throughput side by side to compare them.
Work out the maximum point that meets the SLA
Write two lines to /root/lt2/max-safe.txt: max_safe_concurrency= and max_safe_rps=. They are the largest concurrency among the rows of ramp.csv where p95 is 0.2 seconds or less and errors is 0, and the RPS at that concurrency.
Write to /root/lt2/max-safe.txt the largest concurrency that satisfies the conditions and the throughput at that point. The decision conditions are the two in the instructions, and you must pick them directly from the table.