TT Lab
Get started
Learn Learning paths Courses

Load Testing

Establishing the Baseline

Continue in TT Lab

This lab runs on a real VM

This box is not a Pod but a virtual machine started by KubeVirt. A Linux kernel of its own runs, systemd actually manages services, and docker is a real Docker engine, not an imitation. A container started with docker run becomes an actual process, and docker exec and docker logs work as usual.

This lab used to run inside a Pod. Since it was a box with all kernel capabilities dropped, the step of starting a container was blocked, so you learned through a workaround of unpacking the image archive yourself. The workaround is no longer needed.

There are two things to know.

Goal

You put load on a real container, extract RPS, p50, p95, p99, and the error rate, and repeat three times to build a reproducible baseline that includes the spread in /root/lt1/.

Why it matters

Without a baseline, "it got slower" is just an impression. And a baseline is not enough with numbers alone. Without the concurrency, the number of requests sent, and the time of measurement, you cannot compare next month. The criterion for choosing metrics is also clear. When 9,900 of 10,000 requests take 50ms and 100 take 3,000ms, the average is 79.5ms, and not a single request actually got a response in 79.5ms. Moreover, even if the slow 100 get twice as bad, the average moves only from 79.5 to 109.5ms. So instead of judging by a single average, you look at p50, p95, and p99 together, and in the end you measure three times and call only a difference larger than the spread a 'change'.

How to read the load tool output

This image contains hey. The output has three sections, and all the later steps read their values from these sections.

Save the whole output as it is. The grader checks the values you wrote against this original again.

Steps

First run mkdir -p /root/lt1.

  1. Start the load target container. The name must be exactly lt-web, the port is host 127.0.0.1:8085 → container 80, and curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:8085/ must return 200.\n The target must have a tail in its response time. Build in Python a server that is mostly fast and slow about once in twelve requests, and put it in a python:3.12-alpine container. If you serve a static file with nginx, one request takes less than 1ms, so the average and p95 become the same value, and the grounds you will use in step 6 disappear.
  2. Save the smoke run result to /root/lt1/smoke.txt. The total responses must be 50 or more, and all three sections, Requests/sec, Latency distribution, and Status code distribution, must be in the file.
  3. Separate the warm-up and the main measurement. Save /root/lt1/warmup.txt (the warm-up) and /root/lt1/run1.txt (the main measurement) separately, and the total responses in run1.txt must be 200 or more.
  4. Write five fields to /root/lt1/metrics.json: rps, p50, p95, p99, and error_rate. The first four values are the values read from run1.txt as they are (in seconds), and error_rate is the ratio (0–1) of the number of non-2xx responses ÷ the total number of responses.
  5. Write three lines to /root/lt1/errors.txt: non2xx= (an integer), total= (an integer), and rate= (a percentage, to two decimal places). The first two values must exactly equal the values you counted in the status code distribution.
  6. Write three lines and an explanation to /root/lt1/why.md: average_s= (the Average from run1.txt), p95_s= (the 95% percentile from run1.txt), and ratio= (p95 ÷ average, to two decimal places). Then explain what you miss by looking only at the average, including the word 'average' in the body.
  7. Measure twice more under the same conditions to create /root/lt1/run2.txt and /root/lt1/run3.txt, and organize them in /root/lt1/runs.csv. The format is a header line run,rps,p95 on the first line, followed by exactly 3 data rows (in the form 1,<rps>,<p95>), and a spread= line at the end. spread is (max RPS − min RPS) ÷ min RPS × 100, to one decimal place.
  8. Write six fields to /root/lt1/baseline.json: rps, p95, error_rate, concurrency, requests, and measured_at. rps must be the median of the 3 measurements (the middle value), and error_rate must be 0.01 or less (a number from a state with errors cannot be a baseline).

Notes

Start the load target container

Start the load target container. The name must be exactly lt-web, the port is host 127.0.0.1:8085 → container 80, and curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:8085/ must return 200.\n The target must have a tail in its response time. Build in Python a server that is mostly fast and slow about once in twelve requests, and put it in a python:3.12-alpine container. If you serve a static file with nginx, one request takes less than 1ms, so the average and p95 become the same value, and the grounds you will use in step 6 disappear.

The name must be exactly lt-web, and you must connect port 8085 on host 127.0.0.1 to container port 80. A host-side port below 1024 cannot be bound, so use a high port, and inside the container you are root, so you can use 80 as it is. The target must not be a static file server but a server with a tail in its response time — only then can you see in your own measurements, in step 6, the average and p95 diverging.

Check the tool output with a smoke run

Save the smoke run result to /root/lt1/smoke.txt. The total responses must be 50 or more, and all three sections, Requests/sec, Latency distribution, and Status code distribution, must be in the file.

Save the result to /root/lt1/smoke.txt in full. The summary, the latency distribution, and the status code distribution sections must all remain, so do not cut out only part of it when redirecting. Send at least 50 requests.

Separate the warm-up from the main measurement

Separate the warm-up and the main measurement. Save /root/lt1/warmup.txt (the warm-up) and /root/lt1/run1.txt (the main measurement) separately, and the total responses in run1.txt must be 200 or more.

Create warmup.txt and run1.txt separately. The first requests measure a state with no cache and no connections, so if you mix them into the main measurement, the distribution is contaminated. The main measurement must be 200 or more requests for you to see the distribution.

Organize the five key metrics as JSON

Write five fields to /root/lt1/metrics.json: rps, p50, p95, p99, and error_rate. The first four values are the values read from run1.txt as they are (in seconds), and error_rate is the ratio (0–1) of the number of non-2xx responses ÷ the total number of responses.

The values in /root/lt1/metrics.json must be the values read from run1.txt as they are. If you make them up, you get caught when they are checked against the original. The error rate is a ratio (0–1) and you get it by counting the non-2xx entries in the status code distribution.

Calculate the error rate by hand

Write three lines to /root/lt1/errors.txt: non2xx= (an integer), total= (an integer), and rate= (a percentage, to two decimal places). The first two values must exactly equal the values you counted in the status code distribution.

Write three lines, non2xx / total / rate, to /root/lt1/errors.txt. The first two values must match exactly as integers and rate is a percentage. If you add up all the counts in the status code distribution section, you get total.

Compare the multiple of p95 to the average

Write three lines and an explanation to /root/lt1/why.md: average_s= (the Average from run1.txt), p95_s= (the 95% percentile from run1.txt), and ratio= (p95 ÷ average, to two decimal places). Then explain what you miss by looking only at the average, including the word 'average' in the body.

Write three lines, average_s / p95_s / ratio, and an explanation of why the average is not enough, to /root/lt1/why.md. ratio is p95 divided by the average. Read the values from run1.txt.

Three repeated measurements and the spread

Measure twice more under the same conditions to create /root/lt1/run2.txt and /root/lt1/run3.txt, and organize them in /root/lt1/runs.csv. The format is a header line run,rps,p95 on the first line, followed by exactly 3 data rows (in the form 1,<rps>,<p95>), and a spread= line at the end. spread is (max RPS − min RPS) ÷ min RPS × 100, to one decimal place.

Create run2.txt and run3.txt under the same conditions and organize them in runs.csv. The spread is the difference between the maximum and the minimum divided by the minimum, as a percentage. A difference smaller than this value cannot be called an improvement.

Fix the baseline

Write six fields to /root/lt1/baseline.json: rps, p95, error_rate, concurrency, requests, and measured_at. rps must be the median of the 3 measurements (the middle value), and error_rate must be 0.01 or less (a number from a state with errors cannot be a baseline).

/root/lt1/baseline.json must contain not only the numbers but also the measurement conditions so that you can compare later. The representative value is the middle of the 3 measurements, and a number from a state with errors cannot be a baseline.