Reproducing and Removing the Bottleneck
This lab runs on a real VM
This box is not a Pod but a virtual machine started by KubeVirt. A Linux kernel of its own runs,
systemd actually manages services, and docker is a real Docker engine, not an imitation.
A container started with docker run becomes an actual process, and docker exec and docker logs work as usual.
This lab used to run inside a Pod. Since it was a box with all kernel capabilities dropped, the step of starting a container was blocked, so you learned through a workaround of unpacking the image archive yourself. The workaround is no longer needed.
There are two things to know.
- The first start takes a little over a minute. That is because the VM boots and installs Docker. It is slower than a Pod lab (usually 40 seconds).
- There is no browser preview. Only a single grading port is open for connections into the VM.
If you start a web server, check it with
curlfrom inside the VM.
Goal
You deliberately create and reproduce a bottleneck, narrow it down with hypotheses by layer, measure the improvement in numbers after removing it, and even calculate the number of instances needed to handle the target traffic. Put the outputs under /root/lt3/.
Why it matters
Finding a bottleneck is not a matter of intuition but of sweeping a ladder in order. Socket buffer → thread pool queue → connection pool wait → DB lock wait → processing. The bottleneck this lab creates is the very first of these, namely the concurrency constraint of 'handling only one request at a time'. The signal that appears then is important. Even if you raise the concurrency tenfold, throughput stays almost the same and only latency grows tenfold. Once you have seen with your own eyes this relationship, that when throughput is blocked the wait time grows instead, you will know what to look at first when you meet a log like Connection is not available, request timed out after 30000ms in production. And what you learn in the last step is the principle of not using a measured value as it is. You use only 70% of the measured maximum throughput and leave the rest as headroom.
Preparation
Run mkdir -p /root/lt3. The only images you can use are the pre-pulled python:3.12-alpine, nginx:1.27-alpine, alpine:3.20, and busybox:1.36.
Steps
- Create
/root/lt3/slow.py. Conditions: use the standard libraryhttp.server'sHTTPServer(handles one request at a time), add a 50ms delay in each request withtime.sleep(0.05), and include the stringslow-appin the response body. Run this script with thepython:3.12-alpineimage, starting a container namedlt-slowon host port 127.0.0.1:8086.slow-appmust appear in the response ofcurl http://127.0.0.1:8086/. - Measure with concurrency 1 and save it to
/root/lt3/slow-c1.txt. If it is normal, p95 is 0.04 seconds or more and throughput is 40 RPS or less (since the 50ms delay is processed serially, about 20 RPS in theory). - Measure with concurrency 10 and save it to
/root/lt3/slow-c10.txt. What to check: throughput is less than 2 times that at concurrency 1 (almost the same) and p95 increases by 2 times or more. When throughput is blocked, the wait time grows instead. - In
/root/lt3/hypothesis.md, write 3 or more bottleneck candidates as items starting with-or1.. The candidates must be in different layers (at least three layers among concurrency/threads/workers, CPU, connection pool, lock contention, and IO/network). Also write how you will verify each candidate. - Create
/root/lt3/fast.py. Conditions: it must be able to handle requests concurrently withThreadingHTTPServer(orThreadingMixIn), keep the 50ms delay as it is, and includefast-appin the response body. Start it as a container namedlt-faston host port 127.0.0.1:8087. Measure with concurrency 10 and save it to/root/lt3/fast-c10.txt, and the throughput must be 3 times or more that ofslow-c10.txt. - Also measure the fast app with concurrency 1 to create
/root/lt3/fast-c1.txt, and then organize four rows in/root/lt3/compare.csv. The format isapp,concurrency,rps,p95and the rows areslow,1,slow,10,fast,1, andfast,10. The values must be read from each result file. - Write four lines to
/root/lt3/capacity.txt.target_rps=300(the target peak traffic)measured_rps=— the rps of the fast/10 row ofcompare.csvsafe_rps=— 70% of the measured value, rounded down (an integer)instances=— ceil(300 ÷ safe_rps), that is, the rounded-up integer
- Write a report in
/root/lt3/report.md. It needs four sections: Cause of the bottleneck, Evidence (measured values), Action, and Capacity conclusion. The body must contain the three numbers throughput before the improvement, throughput after the improvement, and the required number of instances as they are.
Notes
- Container run example:
docker run -d --name lt-slow -p 127.0.0.1:8086:8080 -v /root/lt3:/app python:3.12-alpine python /app/slow.py(inside the script, bind to 8080 inside the container). - If the server does not come up, check with
docker logs lt-slow. If you leave out theContent-Lengthheader in the Python handler, the client waits for the connection to close and the measurement is distorted. - Measurement examples:
hey -n 100 -c 1 http://127.0.0.1:8086/ > /root/lt3/slow-c1.txt 2>&1,hey -n 200 -c 10 ..., and since fast is fast, about-n 1000 -c 10is good. - Rounding-up calculation:
awk -v t=300 -v s="$SAFE" 'BEGIN{n=t/s; r=int(n); if(n>r) r=r+1; print r}'. - Common mistake 1: understanding the improvement as 'reducing latency'. Leave the 50ms latency as it is and eliminate the wait. The right answer is that latency stays the same and only throughput grows greatly.
- Common mistake 2: rounding
safe_rps. It must be rounded down, and converselyinstancesmust be rounded up. - Common mistake 3: writing only sentences like "throughput grew greatly" in the step 8 report. The evidence must be numbers.
Start a slow application
Create /root/lt3/slow.py. Conditions: use the standard library http.server's HTTPServer (handles one request at a time), add a 50ms delay in each request with time.sleep(0.05), and include the string slow-app in the response body. Run this script with the python:3.12-alpine image, starting a container named lt-slow on host port 127.0.0.1:8086. slow-app must appear in the response of curl http://127.0.0.1:8086/.
/root/lt3/slow.py must be a server that adds an artificial delay to each request and handles only one at a time. The string slow-app must be in the response body, the container name is lt-slow, and the host port is 8086.
Baseline measurement at concurrency 1
Measure with concurrency 1 and save it to /root/lt3/slow-c1.txt. If it is normal, p95 is 0.04 seconds or more and throughput is 40 RPS or less (since the 50ms delay is processed serially, about 20 RPS in theory).
Save it to /root/lt3/slow-c1.txt. If you added a 50ms delay, p95 should be at least 40ms and throughput around 20 requests per second to be normal. If the values look strange, check the target port first.
Raise the concurrency tenfold
Measure with concurrency 10 and save it to /root/lt3/slow-c10.txt. What to check: throughput is less than 2 times that at concurrency 1 (almost the same) and p95 increases by 2 times or more. When throughput is blocked, the wait time grows instead.
Save it to /root/lt3/slow-c10.txt. With serial processing, throughput stays almost the same and the wait time grows instead. If both are observed at the same time, it is a signal that the bottleneck is in the concurrency.
Set up bottleneck hypotheses
In /root/lt3/hypothesis.md, write 3 or more bottleneck candidates as items starting with - or 1.. The candidates must be in different layers (at least three layers among concurrency/threads/workers, CPU, connection pool, lock contention, and IO/network). Also write how you will verify each candidate.
Write three candidates from different layers and how to verify each in /root/lt3/hypothesis.md. If you sweep the queues of the request handling path in order, the candidates do not overlap. A hypothesis with no way to verify it is not a hypothesis.
Switch to concurrent processing
Create /root/lt3/fast.py. Conditions: it must be able to handle requests concurrently with ThreadingHTTPServer (or ThreadingMixIn), keep the 50ms delay as it is, and include fast-app in the response body. Start it as a container named lt-fast on host port 127.0.0.1:8087. Measure with concurrency 10 and save it to /root/lt3/fast-c10.txt, and the throughput must be 3 times or more that of slow-c10.txt.
/root/lt3/fast.py must be able to handle requests concurrently. Leave the latency itself as it is. The point of this improvement is not to reduce latency but to eliminate the wait. The container name is lt-fast and the port is 8087.
Before-and-after comparison table
Also measure the fast app with concurrency 1 to create /root/lt3/fast-c1.txt, and then organize four rows in /root/lt3/compare.csv. The format is app,concurrency,rps,p95 and the rows are slow,1, slow,10, fast,1, and fast,10. The values must be read from each result file.
Measure the fast app once more with concurrency 1 and then organize four rows in /root/lt3/compare.csv. The values must be read from each result file.
Calculate the number of instances needed
Write four lines to /root/lt3/capacity.txt.
target_rps=300(the target peak traffic)measured_rps=— the rps of the fast/10 row ofcompare.csvsafe_rps=— 70% of the measured value, rounded down (an integer)instances=— ceil(300 ÷ safe_rps), that is, the rounded-up integer
Write the target, measured, and safe throughput and the number of instances in /root/lt3/capacity.txt. If you use the measured maximum as it is, there is no headroom, and the number of instances must be rounded up.
Write the analysis report
Write a report in /root/lt3/report.md. It needs four sections: Cause of the bottleneck, Evidence (measured values), Action, and Capacity conclusion. The body must contain the three numbers throughput before the improvement, throughput after the improvement, and the required number of instances as they are.
Write the cause, evidence, action, and capacity conclusion in /root/lt3/report.md. The evidence must be numbers, not sentences, so write the values from the previous steps into the body as they are.