TT Lab
Get started
Learn Learning paths Courses

Load Testing

Averaging the shard p95s came out 108 ms above the real p95

Continue in TT Lab

Goal

You compute percentiles by hand with two methods and confirm that the values diverge, measure how the simple average and the weighted average of per-shard p95s deviate from the overall p95, measure the method of combining by adding histogram buckets and its error, and then build and submit a tool that correctly combines the results of two real load runs.

Why it matters

Unlike an average, a percentile cannot build the whole from its parts. An average is made of a sum and a count, so adding the sums of the parts gives the whole, but a percentile is a sorted position, so adding the positions of the parts does not give the position of the whole. Yet dashboards and reports casually average per-shard p95s. That value can be higher or lower than the true value, so even the intuition that 'an average would err on the safe side' does not hold. On top of that there are two calculation methods, so when you change tools, p95 looks like it got worse even though nothing happened. The only right ways to combine are to sort the raw data again or to add histogram buckets, and with buckets the boundaries decide the error. If you do these three things by hand once, from then on you will stop throwing away raw output files.

Steps

  1. In /opt/lab/lt/lt-percentiles/ there are three per-shard latency samples (one millisecond value per line). Create /root/lt-percentiles/shards.tsv. It has three lines, and each line has three tab-separated columns, <파일이름> <표본 수> <p95> (file name, number of samples, and p95). Write the names in the order shard-a, shard-b, shard-c, and compute p95 by nearest-rank (the ceil(0.95 x n)th value after sorting), writing it to one decimal place. Then write two lines to /root/lt-percentiles/01-note.txt: total_n=<세 파일의 표본 수 합> (the sum of the sample counts of the three files) and slowest=<p95 가 가장 큰 구간 이름> (the name of the shard with the largest p95).
  2. Create /root/lt-percentiles/pct.py. When called as python3 pct.py <표본파일> <분위> <nearest|linear> (the sample file, the quantile, and the method), it prints one percentile on one line to four decimal places. nearest uses the ceil(분위 x n)th smallest value as it is (that is, ceil(quantile x n)), and linear makes the position h = (n - 1) x 분위 (that is, h = (n - 1) x quantile) proportionally between the samples before and after it. With that tool, measure /opt/lab/lt/lt-percentiles/tiny.txt (20 lines) and write three lines to /root/lt-percentiles/methods.tsv. Each line has three tab-separated columns, <분위> <nearest> <linear> (quantile, nearest value, and linear value), the quantiles are in the order 0.50, 0.95, and 0.99, and the values are to four decimal places. Then write one line to /root/lt-percentiles/02-note.txt, gap_p95=<0.95 에서 linear 빼기 nearest> (at 0.95, linear minus nearest), to four decimal places.
  3. Write five lines to /root/lt-percentiles/combine.txt. mean_of_p95= is the simple average of the three p95s from step 1, weighted_mean_of_p95= is the average using sample counts as weights, and true_p95= is the p95 recomputed by nearest-rank after combining the raw data of the three files. Then write in error_mean= the mean_of_p95 minus true_p95, and in error_weighted= the weighted_mean_of_p95 minus true_p95. All five values are to one decimal place, with a minus sign in front if negative.
  4. Write two lines to /root/lt-percentiles/bound.tsv. Each line has three tab-separated columns, <이름> <표본 수> <p95> (name, number of samples, and p95), and the first line has the name all (all three shards combined) and the second line no-c (only shard-a and shard-b combined). p95 by nearest-rank to one decimal place. Then write three lines to /root/lt-percentiles/04-note.txt: min_shard_p95=, max_shard_p95=, and inside=<yes|no>. inside is yes if the p95 of all lies between the minimum and maximum of the per-shard p95s.
  5. Create /root/lt-percentiles/hist.py. When called as python3 hist.py <경계를 쉼표로> <표본파일...> (the boundaries separated by commas, then the sample files), it prints a table with two tab-separated columns. Each line is <경계> <그 경계 이하인 표본의 누적 개수> (the boundary and the cumulative count of samples at or below it) and the last line is +Inf <전체 개수> (the total count). If you give several files, it counts per file and then adds the same boundaries together. With this tool, give the boundaries 50,100,250,500,1000 and all three shard files and save the resulting table to /root/lt-percentiles/hist-coarse.tsv. Then write three lines, est_p95=, true_p95=, and abs_error=, to /root/lt-percentiles/hist-p95.txt to one decimal place. est_p95 is the p95 estimated from that table by linear interpolation — find the first boundary where the cumulative count becomes 0.95 x 전체 (that is, 0.95 x total) or more, and divide between the previous boundary and that boundary in proportion to the cumulative counts (if there is no previous boundary, treat it as 0).
  6. Count the same samples again with the dense boundaries 50,75,100,125,150,200,250,300,400,500,750,1000,1500 and save to /root/lt-percentiles/hist-fine.tsv. Then write two lines to /root/lt-percentiles/error.tsv. Each line has three tab-separated columns, <이름> <추정 p95> <절대 오차> (name, estimated p95, and absolute error), the names are in the order coarse, fine, and the values are to one decimal place. The error is the absolute value of the difference from true_p95 in step 3. Finally write two lines to /root/lt-percentiles/06-note.txt: better=<coarse|fine> and reason=<40자 이상> (at least 40 characters). In reason, also write why dense boundaries are not free.
  7. Start /opt/lab/lt/lt-percentiles/target.py on port 8080 (/fast takes 30 milliseconds and /slow takes 250 milliseconds). Run hey twice and keep the raw output — hey -n 300 -c 10 -o csv http://127.0.0.1:8080/fast into /root/lt-percentiles/run-fast.csv, and hey -n 100 -c 10 -o csv http://127.0.0.1:8080/slow into /root/lt-percentiles/run-slow.csv. Then write four lines to /root/lt-percentiles/runs.txt — p95_fast=, p95_slow=, mean_of_p95= (the simple average of the two p95s), and merged_p95= (the p95 recomputed after combining the raw response times of both runs). All are in seconds to four decimal places, and only lines whose status-code is 200 are counted.
  8. Create /root/lt-percentiles/merge_p95.py. When called as python3 merge_p95.py <분위> <hey CSV...> (the quantile and the hey CSV files), it gathers the first column (the response time, in seconds) of the lines whose status-code is 200 in the CSVs all into one pot, computes the percentile by nearest-rank, and prints it on one line to four decimal places. You must not compute a percentile per run and average them. Run that tool on the two CSVs from step 7 and write four lines to /root/lt-percentiles/merged.txt — q=0.95, merged_p95=, mean_of_p95=, and gap= (merged minus mean). Finally write one line starting with rule= of at least 60 characters to /root/lt-percentiles/policy.txt, stating the rule the team will follow from now on when combining the results of multiple load tests.

Notes

How many requests per shard, and what is p95

In /opt/lab/lt/lt-percentiles/ there are three per-shard latency samples (one millisecond value per line). Create /root/lt-percentiles/shards.tsv. It has three lines, and each line has three tab-separated columns, <파일이름> <표본 수> <p95> (file name, number of samples, and p95). Write the names in the order shard-a, shard-b, shard-c, and compute p95 by nearest-rank (the ceil(0.95 x n)th value after sorting), writing it to one decimal place. Then write two lines to /root/lt-percentiles/01-note.txt: total_n=<세 파일의 표본 수 합> (the sum of the sample counts of the three files) and slowest=<p95 가 가장 큰 구간 이름> (the name of the shard with the largest p95).

Sorting is sort -n and the line count is wc -l. The position number for nearest-rank is ceil(0.95 x n), and in awk you make it with int(0.95*n) + (0.95*n > int(0.95*n) ? 1 : 0). Notice that the three shards have different sample counts — in the later steps that difference changes the answer.

The same samples, two calculation methods

Create /root/lt-percentiles/pct.py. When called as python3 pct.py <표본파일> <분위> <nearest|linear> (the sample file, the quantile, and the method), it prints one percentile on one line to four decimal places. nearest uses the ceil(분위 x n)th smallest value as it is (that is, ceil(quantile x n)), and linear makes the position h = (n - 1) x 분위 (that is, h = (n - 1) x quantile) proportionally between the samples before and after it. With that tool, measure /opt/lab/lt/lt-percentiles/tiny.txt (20 lines) and write three lines to /root/lt-percentiles/methods.tsv. Each line has three tab-separated columns, <분위> <nearest> <linear> (quantile, nearest value, and linear value), the quantiles are in the order 0.50, 0.95, and 0.99, and the values are to four decimal places. Then write one line to /root/lt-percentiles/02-note.txt, gap_p95=<0.95 에서 linear 빼기 nearest> (at 0.95, linear minus nearest), to four decimal places.

In the sorted list, nearest is index k-1 and linear is values[lo] + (h-lo) x (values[hi]-values[lo]). The 20-sample set was chosen so that the difference between the two methods is visible — in particular the top two values are far apart. The grader runs this tool directly with sample files it made itself.

The average of per-shard p95s is not the overall p95

Write five lines to /root/lt-percentiles/combine.txt. mean_of_p95= is the simple average of the three p95s from step 1, weighted_mean_of_p95= is the average using sample counts as weights, and true_p95= is the p95 recomputed by nearest-rank after combining the raw data of the three files. Then write in error_mean= the mean_of_p95 minus true_p95, and in error_weighted= the weighted_mean_of_p95 minus true_p95. All five values are to one decimal place, with a minus sign in front if negative.

To combine the raw data, one cat /opt/lab/lt/lt-percentiles/shard-*.txt is enough. Use the pct.py you made in step 2 as it is. That the two errors have different signs is the heart of this step — an average is not wrong in only one direction.

The fence within which a combined p95 can lie

Write two lines to /root/lt-percentiles/bound.tsv. Each line has three tab-separated columns, <이름> <표본 수> <p95> (name, number of samples, and p95), and the first line has the name all (all three shards combined) and the second line no-c (only shard-a and shard-b combined). p95 by nearest-rank to one decimal place. Then write three lines to /root/lt-percentiles/04-note.txt: min_shard_p95=, max_shard_p95=, and inside=<yes|no>. inside is yes if the p95 of all lies between the minimum and maximum of the per-shard p95s.

See how much the overall p95 drops when you remove the one slow shard that has only 10% of traffic. And think about the fence — if 95% of every shard is at or below some value, then even after mixing, 95% is at or below that value. So a report that says 'after combining, it came out larger than every shard' should make you suspect the calculation, not the data.

Buckets can be added

Create /root/lt-percentiles/hist.py. When called as python3 hist.py <경계를 쉼표로> <표본파일...> (the boundaries separated by commas, then the sample files), it prints a table with two tab-separated columns. Each line is <경계> <그 경계 이하인 표본의 누적 개수> (the boundary and the cumulative count of samples at or below it) and the last line is +Inf <전체 개수> (the total count). If you give several files, it counts per file and then adds the same boundaries together. With this tool, give the boundaries 50,100,250,500,1000 and all three shard files and save the resulting table to /root/lt-percentiles/hist-coarse.tsv. Then write three lines, est_p95=, true_p95=, and abs_error=, to /root/lt-percentiles/hist-p95.txt to one decimal place. est_p95 is the p95 estimated from that table by linear interpolation — find the first boundary where the cumulative count becomes 0.95 x 전체 (that is, 0.95 x total) or more, and divide between the previous boundary and that boundary in proportion to the cumulative counts (if there is no previous boundary, treat it as 0).

Buckets are cumulative. le=100 means 'how many are 100 or less', not 'between 50 and 100'. The estimation formula is the same as Prometheus's histogram_quantile — it looks proportionally at where in the bucket the target count falls. The grader runs this tool directly with samples and boundaries it made itself.

The boundaries decide the error

Count the same samples again with the dense boundaries 50,75,100,125,150,200,250,300,400,500,750,1000,1500 and save to /root/lt-percentiles/hist-fine.tsv. Then write two lines to /root/lt-percentiles/error.tsv. Each line has three tab-separated columns, <이름> <추정 p95> <절대 오차> (name, estimated p95, and absolute error), the names are in the order coarse, fine, and the values are to one decimal place. The error is the absolute value of the difference from true_p95 in step 3. Finally write two lines to /root/lt-percentiles/06-note.txt: better=<coarse|fine> and reason=<40자 이상> (at least 40 characters). In reason, also write why dense boundaries are not free.

The error of a bucket estimate is decided by 'the width of the bucket that contains the target'. Look for the same target in a table where one bucket spans 250 to 500 and in a table where one bucket spans 250 to 300. In exchange, if you add boundaries, the time series grow by that much — in Prometheus, one boundary is one time series.

Combine two real runs

Start /opt/lab/lt/lt-percentiles/target.py on port 8080 (/fast takes 30 milliseconds and /slow takes 250 milliseconds). Run hey twice and keep the raw output — hey -n 300 -c 10 -o csv http://127.0.0.1:8080/fast into /root/lt-percentiles/run-fast.csv, and hey -n 100 -c 10 -o csv http://127.0.0.1:8080/slow into /root/lt-percentiles/run-slow.csv. Then write four lines to /root/lt-percentiles/runs.txt — p95_fast=, p95_slow=, mean_of_p95= (the simple average of the two p95s), and merged_p95= (the p95 recomputed after combining the raw response times of both runs). All are in seconds to four decimal places, and only lines whose status-code is 200 are counted.

In hey -o csv, the first column is the response time (seconds) and the seventh column is the status code. Skip the one header line. To use the pct.py from step 2 as it is, extract only the numbers with cut -d, -f1 and gather them into a temporary file. Note that the two runs have different counts — you can see which way the combined result leans.

A tool that combines multiple runs correctly

Create /root/lt-percentiles/merge_p95.py. When called as python3 merge_p95.py <분위> <hey CSV...> (the quantile and the hey CSV files), it gathers the first column (the response time, in seconds) of the lines whose status-code is 200 in the CSVs all into one pot, computes the percentile by nearest-rank, and prints it on one line to four decimal places. You must not compute a percentile per run and average them. Run that tool on the two CSVs from step 7 and write four lines to /root/lt-percentiles/merged.txt — q=0.95, merged_p95=, mean_of_p95=, and gap= (merged minus mean). Finally write one line starting with rule= of at least 60 characters to /root/lt-percentiles/policy.txt, stating the rule the team will follow from now on when combining the results of multiple load tests.

Skip the one header line, and discard lines with fewer than seven columns. If you do not filter by status code, the times of failed requests get mixed in and the answer changes. The grader runs this tool with two CSVs it made itself at two different quantiles — an implementation that averages per run or does not filter by status code fails right there.