FDE Capstone: The Warehouse Got the Same Order Three Times
The customer only wrote 'fast'
In one line
Acceptance criteria turn the customer's adjectives into four pieces: a metric, a calculation method, a threshold, and measurement conditions, and the acceptance test runner reads that criteria table from outside the code, sends the requests, and leaves the evidence and the verdict together.
Why this was needed
The purchasing team's email had just one line: "Quote lookups have to be fast." The development team said the average came out to 80ms, so it was fast, and the operations team said the lookup screen froze for 3 seconds yesterday afternoon. Both are true. A service that freezes once in twenty times is fast if you look at the average. If this difference does not come out at the acceptance meeting, then after delivery "But you said it was fast" and "But it isn't fast" collide.
For an FDE, acceptance criteria are not paperwork but the product of a negotiation. You turn the experience the customer wants into numbers, but you also have to agree on what the number is a measurement of and how it was measured, so that later you do not say different things about the same number. And if you hard-code the agreed numbers into the code, you have to fix the runner every time the criteria change, and it becomes blurry which version of the runner judged against which criteria. That is why you keep the criteria table in a file.
How it works
The service level objectives chapter of the Google SRE book warns that gathering an SLI only as an average hides a situation where most are fast and a long tail is much slower, and recommends looking at the shape of the distribution with percentiles. The same chapter writes the natural form of an SLO as "SLI ≤ target". One row of this lab's criteria table is exactly that shape.
{"id": "quote-latency", "metric": "p95_ms", "op": "<=", "threshold": 250,
"source": "견적 조회가 빠르게 되어야 합니다"}
You also have to write down the calculation method. "p95" is not one value. Python's statistics.quantiles defaults to the exclusive method and interpolates linearly between two samples, and there is a separate inclusive method. We compared the three methods by measurement when 38 of 40 requests were 5ms and 2 were 400ms. nearest-rank (the ceil(0.95×n)-th value after sorting) was 5.0, and the 95th cut point of quantiles(n=100) was 380.25 with exclusive and 24.75 with inclusive. If the threshold is 250ms, the same evidence can be both a pass and a failure depending on the method. That is why you state nearest-rank in the contract and make the evidence such that anyone can recompute it.
The measurement conditions are a criterion too. Latency is measured from sending the request until the whole body is received, with time.perf_counter. The documentation explains this clock as the most precise clock for measuring short intervals, and says that its reference point is undefined so only the difference between two calls is meaningful. In CPython it is the same clock as the monotonic clock that never goes backward. On the other hand, the same documentation says that time.time() can return a smaller value if the system clock is adjusted backward between two calls, so you do not use it for interval measurement. For the error rate, "what do we count as an error" comes first. With this customer we decided to count responses that are not 200 and responses that do not arrive within 1000ms. The urllib.request documentation explains the timeout of urlopen as a limit on each blocking operation, such as a connection attempt. That means it is not a deadline for the whole request, so with a server that trickles data out little by little, it can take longer than the limit. The runner leaves the measured time in the evidence as it is to expose such cases.
Accuracy and reproducibility are measured separately. Accuracy is compared with the expected values in the case table the customer gave, and reproducibility is compared by querying the same case table twice and comparing the results with each other. In the second comparison you have to exclude fields that are normally different each time, such as request_id or a creation time, and which fields to exclude is also written in the criteria table. If you write it in the code, a normal build fails the moment the server renames a field.
| Customer sentence | Metric | Criterion | Measurement condition |
|---|---|---|---|
| Fast | p95_ms (nearest-rank) | 250 or less | The whole request, until the body is received |
| Without errors | error_rate | 0.01 or less | Not 200 + over 1000ms |
| With correct amounts | mismatches | 0 | First pass, compared with the case table's expected values |
| The same on a rerun | rerun_diffs | 0 | Compare the first and second passes, excluding fields that change |
What it looks like in the field
The most common failure is a report with only a verdict. The customer's operations team, handed a single line "acceptance test passed", has no way to verify it, and when an outage happens two months later, they start by re-arguing what the test measured. If there is evidence that leaves the latency, status, and response for each request, the customer can recompute in their own way. If the value recomputed from the evidence differs from the value in the report, you do not trust the report. That is exactly what this lab's grader does.
The second is changing the criteria after the fact. It is the scene where a candidate build fails the amount criterion, and someone says "a 1-won difference is rounding, so let's allow it", changes the criteria table, and runs again. A tolerance may be needed. But that is something to newly agree on with the customer, not something the person who ran the runner decides after seeing the result. This is also why you copy the thresholds as they are into the verdict report.
What really matters in practice
- For each adjective, write the metric, the calculation method, the threshold, and the measurement conditions, and leave the customer sentence that was the basis together with them.
- Read the threshold from the criteria table. If the number is hard-coded in the runner, it is silently wrong on the day the criteria change.
- Keep the number of requests exact. If you mix in warm-up or retries, the denominators of the error rate and the percentile change.
- Leave the evidence first and compute the verdict from the evidence. It is a report only if the recomputation is the same.
What you will do in the next lab
After turning the customer email and the meeting minutes into a criteria table, you add latency, error rate, accuracy, and reproducibility to the runner in turn. The grader starts a fake quote server with a build with a slow tail, a build that sometimes freezes or returns 500, a build in which some amounts are off by 1 won, and a build whose re-queries are unstable, and runs your runner changing the thresholds and the case table each time. At the end, you start the customer's candidate rc2 yourself and judge whether to accept it.