Same average rate, but only one side blew up the queue
Goal
You create and send to the same server three arrivals with the same average request rate yourself — constant interval, exponential, and bunched — and measure how the queue and latency differ. Then you compare against measurement the law by which a closed loop's think time decides the effective request rate, show in numbers how many times off a capacity estimate that assumes even arrivals is, and decide this service's arrival model and number of workers together with the evidence.
Why it matters
Capacity estimation almost always starts with an average — so many per second, so many milliseconds each, so that many workers. Hidden in that calculation is the assumption that 'requests arrive evenly', but real traffic arrives in bunches because of cron, notifications, and retries. A queue reacts not to the average but to the moment, so even if the average utilization is 25%, at the moment a bunch arrives the workers fall short, and that wait creates timeouts and retries again. On the other side is the trap of the closed loop. If you test with a number of virtual users, the load drops along with it when the server slows down, so the report's 'it held up to so many per second' becomes not the server's limit but the limit of the test configuration. If you measure these two things by hand once, from then on you will write the arrival shape next to the average when writing down load.
Steps
- Start
/opt/lab/lt/lt-arrival-process/server.pywithpython3 /opt/lab/lt/lt-arrival-process/server.py 8080 40 4(port 8080, service 40 milliseconds, 4 workers). Measure the response time when idle six times withcurl -s -o /dev/null -w '%{time_total}\n'and leave the resulting numbers in seconds as they are, one per line, in/root/lt-arrival-process/01-probe.txt. Then write five lines to/root/lt-arrival-process/01-target.txt—service_ms=(the median of the six measurements, rounded to milliseconds),workers=(the number of workers),max_rate=(the number of workers divided by the service time, requests per second),chosen_rate=25, andutilization=(chosen_rate divided by max_rate, to two decimal places). - Create
/root/lt-arrival-process/gen.py. It is called aspython3 gen.py <even|poisson|burst> <요청수> <초당요청> <시드> <출력CSV> [URL](the shape, the number of requests, the requests per second, the seed, the output CSV, and an optional URL). If you do not give a URL, it does not send but only writes the plan to the CSV — after the headeri,intended, write for each request its number and scheduled time (seconds from the start).evenhas an interval always 1/request rate,poissonhas an exponential distribution of intervals with mean 1/request rate, andburstsends 12 requests at the same time and then pauses for 12/request rate. Take the seed and fix the random numbers. For checking, runpython3 gen.py even 300 25 7 /root/lt-arrival-process/sched-even.csvonce. - Run the generator three times as a dry run to create
/root/lt-arrival-process/sched-even.csv,/root/lt-arrival-process/sched-poisson.csv, and/root/lt-arrival-process/sched-burst.csv(all with300 25 7). Then write three lines to/root/lt-arrival-process/schedules.tsv. Each line has four tab-separated columns,<모양> <요청 수> <평균 간격> <변동계수>(shape, number of requests, mean interval, and coefficient of variation), the shapes are in the ordereven,poisson,burst, the mean interval is in seconds to four decimal places, and the coefficient of variation is the standard deviation of the intervals divided by the mean, to three decimal places. - Actually send the same number of requests to the same server with the three arrival methods. For each method, clear the records with
curl -s http://127.0.0.1:8080/reset, send withpython3 gen.py <모양> 150 25 7 /root/lt-arrival-process/run-<모양>.csv http://127.0.0.1:8080/work(where the placeholder is the shape), and then get the server records withcurl -s http://127.0.0.1:8080/dump > /root/lt-arrival-process/srv-<모양>.csv(the shapes areeven,poisson, andburst). Then write three lines to/root/lt-arrival-process/latency.tsv— four tab-separated columns,<모양> <성공 건수> <p50> <p95>(shape, number of successes, p50, and p95), where latency isreceivedminussentinrun-*.csvconverted to milliseconds and computed by nearest-rank to one decimal place, counting only lines whosecodeis 200. - The
/root/lt-arrival-process/srv-*.csvfiles the server left contain, for each request,qlen(the number of requests in the system at the moment that request arrived). Write three lines to/root/lt-arrival-process/queue.tsv— four tab-separated columns,<모양> <기록 수> <최대 대기열> <평균 대기열>(shape, number of records, maximum queue, and average queue), the shapes are in the ordereven,poisson,burst, and the average is to two decimal places. Then write two lines to/root/lt-arrival-process/05-note.txt:worst=<최대 대기열이 가장 큰 모양>(the shape with the largest maximum queue) andratio=<그 최대값을 even 의 최대값으로 나눈 값, 소수 첫째 자리까지>(that maximum divided by the maximum of even, to one decimal place). - Create
/root/lt-arrival-process/closed.py. When called aspython3 closed.py <URL> <사용자수> <생각시간초> <지속초> <출력CSV>(the URL, number of users, think time in seconds, duration in seconds, and output CSV), each user repeats sending one request and pausing for the think time for the duration. The CSV has the headeruser,sent,received,codefollowed by seconds from the start. Run it once withpython3 closed.py http://127.0.0.1:8080/work 8 0.2 10 /root/lt-arrival-process/closed.csv, and write six lines to/root/lt-arrival-process/think.txt—users=8,think_s=0.200,mean_r_s=(the mean response time),predicted_rate=(the number of users divided by (mean_r_s plus think_s)),measured_rate=(the number of successes divided by the time from the first sent to the last received), andnaive_rate=(the number of users divided by think_s, treating the response time as 0). Request rates are to three decimal places and times are to four decimal places. - Write five lines to
/root/lt-arrival-process/capacity.txt.formula_workers=(the calculation assuming even arrivals:chosen_rate x service_srounded up to an integer),observed_even=andobserved_burst=(the maximum number of requests in the system at the same time fromsrv-even.csvandsrv-burst.csvrespectively),underestimate_factor=(observed_burst divided by formula_workers, to one decimal place), andverdict=<under|ok>(under if formula_workers is smaller than observed_burst). For the number of simultaneous requests, sweep in time order adding 1 atarriveand subtracting 1 atend, and take the maximum. - Write four lines to
/root/lt-arrival-process/arrival-model.txt.model=<even|poisson|burst>is the arrival model you will use from now on when setting this service's capacity,workers=is the number of workers decided by that model (an integer),evidence=is why that model, citing file names and numbers you made in the previous steps, in at least 60 characters, andrisk=is what collapses first if that model is wrong, in at least 40 characters. The grader checks whetherworkersis at least the maximum number of simultaneous requests observed in step 7 and whether it is greater thanformula_workersassuming even arrivals.
Notes
- The working directory is
/root/lt-arrival-process. If it does not exist, create it first. - The material is just
/opt/lab/lt/lt-arrival-process/server.py./workis the path that does work,/resetclears the records, and/dumpreturnsseq,arrive,start,end,qlenper request as CSV. It spends time withtime.sleep, not CPU. - The load target is not a container but a Python standard library server. You cannot start containers in this lab environment. Start it with
(nohup python3 /opt/lab/lt/lt-arrival-process/server.py 8080 40 4 >/dev/null 2>&1 &), check once withcurl, and then apply the load. - Common mistake: not calling
/resetbefore changing methods. The records of the previous run remain and the queue comparison gets mixed up. - Common mistake: reading the latency of bunched arrivals as 'the server got slower'. The service time is the same and what grew is the wait time —
startminusarriveinsrv-*.csvis that wait. - k6 open vs closed models · k6 executors · k6 constant-arrival-rate · k6 ramping-arrival-rate · Service Level Objectives (SRE Book)
How many workers, and how long does one request take
Start /opt/lab/lt/lt-arrival-process/server.py with python3 /opt/lab/lt/lt-arrival-process/server.py 8080 40 4 (port 8080, service 40 milliseconds, 4 workers). Measure the response time when idle six times with curl -s -o /dev/null -w '%{time_total}\n' and leave the resulting numbers in seconds as they are, one per line, in /root/lt-arrival-process/01-probe.txt. Then write five lines to /root/lt-arrival-process/01-target.txt — service_ms= (the median of the six measurements, rounded to milliseconds), workers= (the number of workers), max_rate= (the number of workers divided by the service time, requests per second), chosen_rate=25, and utilization= (chosen_rate divided by max_rate, to two decimal places).
/work is the path that does work, and /reset and /dump are the paths that handle the records. Idle means a state with no load applied — send one at a time. If you calculate how many requests per second 4 workers, each spending 40 milliseconds, can handle, you will see how much headroom the 25 requests per second you will apply is.
A generator where you can choose the arrival shape
Create /root/lt-arrival-process/gen.py. It is called as python3 gen.py <even|poisson|burst> <요청수> <초당요청> <시드> <출력CSV> [URL] (the shape, the number of requests, the requests per second, the seed, the output CSV, and an optional URL). If you do not give a URL, it does not send but only writes the plan to the CSV — after the header i,intended, write for each request its number and scheduled time (seconds from the start). even has an interval always 1/request rate, poisson has an exponential distribution of intervals with mean 1/request rate, and burst sends 12 requests at the same time and then pauses for 12/request rate. Take the seed and fix the random numbers. For checking, run python3 gen.py even 300 25 7 /root/lt-arrival-process/sched-even.csv once.
Make exponential intervals with random.Random(seed).expovariate(rate). The scheduled time of burst is one line, (i // 12) * (12 / rate). In all three methods the scheduled time of the last request should be almost the same — meaning the same average request rate. The grader runs this tool directly with arguments it chooses itself.
The same average, different shapes
Run the generator three times as a dry run to create /root/lt-arrival-process/sched-even.csv, /root/lt-arrival-process/sched-poisson.csv, and /root/lt-arrival-process/sched-burst.csv (all with 300 25 7). Then write three lines to /root/lt-arrival-process/schedules.tsv. Each line has four tab-separated columns, <모양> <요청 수> <평균 간격> <변동계수> (shape, number of requests, mean interval, and coefficient of variation), the shapes are in the order even, poisson, burst, the mean interval is in seconds to four decimal places, and the coefficient of variation is the standard deviation of the intervals divided by the mean, to three decimal places.
An interval is the difference between adjacent scheduled times (299 of them). Compute the standard deviation on a population basis. What this step wants to show is that the three mean intervals are almost the same while only the coefficient of variation splits into 0, near 1, and much greater than 1. The intervals of burst contain lots of zeros.
Hit the same server in three ways
Actually send the same number of requests to the same server with the three arrival methods. For each method, clear the records with curl -s http://127.0.0.1:8080/reset, send with python3 gen.py <모양> 150 25 7 /root/lt-arrival-process/run-<모양>.csv http://127.0.0.1:8080/work (where the placeholder is the shape), and then get the server records with curl -s http://127.0.0.1:8080/dump > /root/lt-arrival-process/srv-<모양>.csv (the shapes are even, poisson, and burst). Then write three lines to /root/lt-arrival-process/latency.tsv — four tab-separated columns, <모양> <성공 건수> <p50> <p95> (shape, number of successes, p50, and p95), where latency is received minus sent in run-*.csv converted to milliseconds and computed by nearest-rank to one decimal place, counting only lines whose code is 200.
The three methods have the same average request rate, so the total time ends up similar too. But the latency distributions are not the same — in particular, look at the gap between p50 and p95. With constant intervals the two are almost stuck together, and with bunches they are far apart. Sending all three takes about 20 seconds.
The queue the server counted
The /root/lt-arrival-process/srv-*.csv files the server left contain, for each request, qlen (the number of requests in the system at the moment that request arrived). Write three lines to /root/lt-arrival-process/queue.tsv — four tab-separated columns, <모양> <기록 수> <최대 대기열> <평균 대기열> (shape, number of records, maximum queue, and average queue), the shapes are in the order even, poisson, burst, and the average is to two decimal places. Then write two lines to /root/lt-arrival-process/05-note.txt: worst=<최대 대기열이 가장 큰 모양> (the shape with the largest maximum queue) and ratio=<그 최대값을 even 의 최대값으로 나눈 값, 소수 첫째 자리까지> (that maximum divided by the maximum of even, to one decimal place).
The line counts of the three records must be the same — because you sent the same number of requests. But the maximum queues are not the same. With constant intervals it does not exceed the number of workers, but with bunches, as many as arrived at once pile up as they are. It means that even with the same average request rate, the number of simultaneous requests at a moment is decided by the arrival shape.
Think time decides the request rate
Create /root/lt-arrival-process/closed.py. When called as python3 closed.py <URL> <사용자수> <생각시간초> <지속초> <출력CSV> (the URL, number of users, think time in seconds, duration in seconds, and output CSV), each user repeats sending one request and pausing for the think time for the duration. The CSV has the header user,sent,received,code followed by seconds from the start. Run it once with python3 closed.py http://127.0.0.1:8080/work 8 0.2 10 /root/lt-arrival-process/closed.csv, and write six lines to /root/lt-arrival-process/think.txt — users=8, think_s=0.200, mean_r_s= (the mean response time), predicted_rate= (the number of users divided by (mean_r_s plus think_s)), measured_rate= (the number of successes divided by the time from the first sent to the last received), and naive_rate= (the number of users divided by think_s, treating the response time as 0). Request rates are to three decimal places and times are to four decimal places.
Applying Little's law to one user's round trip, the effective request rate is N / (R + Z). naive_rate is the prediction when you leave out R, and if you compare it with the measurement, you see why you must not leave it out. This step takes some ten-odd seconds.
How many times off is a capacity that assumes even arrivals
Write five lines to /root/lt-arrival-process/capacity.txt. formula_workers= (the calculation assuming even arrivals: chosen_rate x service_s rounded up to an integer), observed_even= and observed_burst= (the maximum number of requests in the system at the same time from srv-even.csv and srv-burst.csv respectively), underestimate_factor= (observed_burst divided by formula_workers, to one decimal place), and verdict=<under|ok> (under if formula_workers is smaller than observed_burst). For the number of simultaneous requests, sweep in time order adding 1 at arrive and subtracting 1 at end, and take the maximum.
Check whether using the qlen column directly as the maximum gives the same number — it is usually the same because it is the value the server counted at the moment of arrival. But if you sweep it yourself, you can see with your own eyes why 'even with low average utilization, the moment's concurrency can be much larger'. The denominator of the division can be 1, so calculate with decimals.
What to take as this service's arrival process
Write four lines to /root/lt-arrival-process/arrival-model.txt. model=<even|poisson|burst> is the arrival model you will use from now on when setting this service's capacity, workers= is the number of workers decided by that model (an integer), evidence= is why that model, citing file names and numbers you made in the previous steps, in at least 60 characters, and risk= is what collapses first if that model is wrong, in at least 40 characters. The grader checks whether workers is at least the maximum number of simultaneous requests observed in step 7 and whether it is greater than formula_workers assuming even arrivals.
The coefficient of variation from step 3 and the maximum queue from step 5 are the evidence that this service's arrivals are not even. If you accept that evidence, the number of workers is decided not by the average but by the size of the bunch. How much margin to leave is up to you, but if you go below the observed maximum number of simultaneous requests, a queue forms every time that bunch arrives.