TT Lab
Get started
Learn Learning paths Courses

Load Testing

Reproducing and Removing the Bottleneck

Continue in TT Lab

This lab runs on a real VM

This box is not a Pod but a virtual machine started by KubeVirt. A Linux kernel of its own runs, systemd actually manages services, and docker is a real Docker engine, not an imitation. A container started with docker run becomes an actual process, and docker exec and docker logs work as usual.

This lab used to run inside a Pod. Since it was a box with all kernel capabilities dropped, the step of starting a container was blocked, so you learned through a workaround of unpacking the image archive yourself. The workaround is no longer needed.

There are two things to know.

Goal

You deliberately create and reproduce a bottleneck, narrow it down with hypotheses by layer, measure the improvement in numbers after removing it, and even calculate the number of instances needed to handle the target traffic. Put the outputs under /root/lt3/.

Why it matters

Finding a bottleneck is not a matter of intuition but of sweeping a ladder in order. Socket buffer → thread pool queue → connection pool wait → DB lock wait → processing. The bottleneck this lab creates is the very first of these, namely the concurrency constraint of 'handling only one request at a time'. The signal that appears then is important. Even if you raise the concurrency tenfold, throughput stays almost the same and only latency grows tenfold. Once you have seen with your own eyes this relationship, that when throughput is blocked the wait time grows instead, you will know what to look at first when you meet a log like Connection is not available, request timed out after 30000ms in production. And what you learn in the last step is the principle of not using a measured value as it is. You use only 70% of the measured maximum throughput and leave the rest as headroom.

Preparation

Run mkdir -p /root/lt3. The only images you can use are the pre-pulled python:3.12-alpine, nginx:1.27-alpine, alpine:3.20, and busybox:1.36.

Steps

  1. Create /root/lt3/slow.py. Conditions: use the standard library http.server's HTTPServer (handles one request at a time), add a 50ms delay in each request with time.sleep(0.05), and include the string slow-app in the response body. Run this script with the python:3.12-alpine image, starting a container named lt-slow on host port 127.0.0.1:8086. slow-app must appear in the response of curl http://127.0.0.1:8086/.
  2. Measure with concurrency 1 and save it to /root/lt3/slow-c1.txt. If it is normal, p95 is 0.04 seconds or more and throughput is 40 RPS or less (since the 50ms delay is processed serially, about 20 RPS in theory).
  3. Measure with concurrency 10 and save it to /root/lt3/slow-c10.txt. What to check: throughput is less than 2 times that at concurrency 1 (almost the same) and p95 increases by 2 times or more. When throughput is blocked, the wait time grows instead.
  4. In /root/lt3/hypothesis.md, write 3 or more bottleneck candidates as items starting with - or 1.. The candidates must be in different layers (at least three layers among concurrency/threads/workers, CPU, connection pool, lock contention, and IO/network). Also write how you will verify each candidate.
  5. Create /root/lt3/fast.py. Conditions: it must be able to handle requests concurrently with ThreadingHTTPServer (or ThreadingMixIn), keep the 50ms delay as it is, and include fast-app in the response body. Start it as a container named lt-fast on host port 127.0.0.1:8087. Measure with concurrency 10 and save it to /root/lt3/fast-c10.txt, and the throughput must be 3 times or more that of slow-c10.txt.
  6. Also measure the fast app with concurrency 1 to create /root/lt3/fast-c1.txt, and then organize four rows in /root/lt3/compare.csv. The format is app,concurrency,rps,p95 and the rows are slow,1, slow,10, fast,1, and fast,10. The values must be read from each result file.
  7. Write four lines to /root/lt3/capacity.txt.
    • target_rps=300 (the target peak traffic)
    • measured_rps= — the rps of the fast/10 row of compare.csv
    • safe_rps= — 70% of the measured value, rounded down (an integer)
    • instances= — ceil(300 ÷ safe_rps), that is, the rounded-up integer
  8. Write a report in /root/lt3/report.md. It needs four sections: Cause of the bottleneck, Evidence (measured values), Action, and Capacity conclusion. The body must contain the three numbers throughput before the improvement, throughput after the improvement, and the required number of instances as they are.

Notes

Start a slow application

Create /root/lt3/slow.py. Conditions: use the standard library http.server's HTTPServer (handles one request at a time), add a 50ms delay in each request with time.sleep(0.05), and include the string slow-app in the response body. Run this script with the python:3.12-alpine image, starting a container named lt-slow on host port 127.0.0.1:8086. slow-app must appear in the response of curl http://127.0.0.1:8086/.

/root/lt3/slow.py must be a server that adds an artificial delay to each request and handles only one at a time. The string slow-app must be in the response body, the container name is lt-slow, and the host port is 8086.

Baseline measurement at concurrency 1

Measure with concurrency 1 and save it to /root/lt3/slow-c1.txt. If it is normal, p95 is 0.04 seconds or more and throughput is 40 RPS or less (since the 50ms delay is processed serially, about 20 RPS in theory).

Save it to /root/lt3/slow-c1.txt. If you added a 50ms delay, p95 should be at least 40ms and throughput around 20 requests per second to be normal. If the values look strange, check the target port first.

Raise the concurrency tenfold

Measure with concurrency 10 and save it to /root/lt3/slow-c10.txt. What to check: throughput is less than 2 times that at concurrency 1 (almost the same) and p95 increases by 2 times or more. When throughput is blocked, the wait time grows instead.

Save it to /root/lt3/slow-c10.txt. With serial processing, throughput stays almost the same and the wait time grows instead. If both are observed at the same time, it is a signal that the bottleneck is in the concurrency.

Set up bottleneck hypotheses

In /root/lt3/hypothesis.md, write 3 or more bottleneck candidates as items starting with - or 1.. The candidates must be in different layers (at least three layers among concurrency/threads/workers, CPU, connection pool, lock contention, and IO/network). Also write how you will verify each candidate.

Write three candidates from different layers and how to verify each in /root/lt3/hypothesis.md. If you sweep the queues of the request handling path in order, the candidates do not overlap. A hypothesis with no way to verify it is not a hypothesis.

Switch to concurrent processing

Create /root/lt3/fast.py. Conditions: it must be able to handle requests concurrently with ThreadingHTTPServer (or ThreadingMixIn), keep the 50ms delay as it is, and include fast-app in the response body. Start it as a container named lt-fast on host port 127.0.0.1:8087. Measure with concurrency 10 and save it to /root/lt3/fast-c10.txt, and the throughput must be 3 times or more that of slow-c10.txt.

/root/lt3/fast.py must be able to handle requests concurrently. Leave the latency itself as it is. The point of this improvement is not to reduce latency but to eliminate the wait. The container name is lt-fast and the port is 8087.

Before-and-after comparison table

Also measure the fast app with concurrency 1 to create /root/lt3/fast-c1.txt, and then organize four rows in /root/lt3/compare.csv. The format is app,concurrency,rps,p95 and the rows are slow,1, slow,10, fast,1, and fast,10. The values must be read from each result file.

Measure the fast app once more with concurrency 1 and then organize four rows in /root/lt3/compare.csv. The values must be read from each result file.

Calculate the number of instances needed

Write four lines to /root/lt3/capacity.txt.

Write the target, measured, and safe throughput and the number of instances in /root/lt3/capacity.txt. If you use the measured maximum as it is, there is no headroom, and the number of instances must be rounded up.

Write the analysis report

Write a report in /root/lt3/report.md. It needs four sections: Cause of the bottleneck, Evidence (measured values), Action, and Capacity conclusion. The body must contain the three numbers throughput before the improvement, throughput after the improvement, and the required number of instances as they are.

Write the cause, evidence, action, and capacity conclusion in /root/lt3/report.md. The evidence must be numbers, not sentences, so write the values from the previous steps into the body as they are.