FDE Capstone: The Warehouse Got the Same Order Three Times
Average was 80ms, yet the lookup screen froze for 3 seconds
Goal
You turn ambiguous customer sentences into a measurable criteria table, and build an acceptance test runner that reads the criteria table, sends requests, and leaves evidence and a verdict. It must accept only a normal build and fail a defective build on the corresponding criterion.
Why it matters
"Fast" can be measured by the average or by p95, and even p95 gives a different value depending on the calculation method. If you do not agree on what to measure and how, after delivery you end up saying different things about the same number. If you write the thresholds in the code, the runner is silently wrong on the day the criteria change. A report with only a verdict and no evidence cannot be verified by the customer. The grader starts a fake quote server on its own port, makes builds with a slow tail, that sometimes freeze or return 500, whose amounts are off by 1 won, and whose re-queries are unstable, runs your runner changing the thresholds, the case table, and the names of the fields that change each time, and then recomputes the metrics from the evidence and compares them with the report.
The expected time is 60 minutes. Extend with +time before the default session ends (up to 180 minutes). When the session ends, the files in /root disappear, so keep the code separately.
Steps
- Read the customer email and the meeting minutes and write the agreed criteria into /root/accept/criteria.json. Write not the purchasing manager's first hoped-for number but the numbers and measurement conditions agreed at the meeting.
- Make /root/accept/accept.py query the case table as many times as passes in the criteria table and leave samples.jsonl, then judge the p95_ms (nearest-rank) criterion and produce report.json and exit codes 0 and 1.
- Add error_rate to accept.py. Count responses that are not 200 and responses that do not arrive within timeout_ms as errors, and leave "timeout" in the evidence's status.
- Add mismatches to accept.py. Compare with expected_total only the cases that got 200 in the first pass, and leave the ids of the wrong cases as mismatched_cases.
- Add rerun_diffs to accept.py. Compare the cases that were 200 in both the first and second passes by the body excluding the criteria table's ignore_fields, and leave rerun_diff_cases.
- In a criteria table that uses all four criteria, follow op (<, <=, ==) as it is, and make the report's observed, pass, and metrics equal the values recomputed from the evidence.
- Run the five candidates (normal, slow tail, occasional 500, wrong amounts, unstable re-queries) with a newly changed criteria table and check that only the normal one is accepted.
- Start the customer candidate rc2 with
python3 /opt/lab/p1a-criteria/quote_server.py --port 8095 --build rc2, run with criteria.json and the provided case table, and leave report.json and samples.jsonl in /root/accept/rc2/.
Notes
- Materials: the customer request /opt/lab/p1a-criteria/customer-request.md, the execution contract /opt/lab/p1a-criteria/CONTRACT.md, the case table /opt/lab/p1a-criteria/cases.csv, and the fake server /opt/lab/p1a-criteria/quote_server.py
- Practice server: python3 /opt/lab/p1a-criteria/quote_server.py --port 18095 --build rc1 & (a normal candidate). When you are done, pick out and stop only that PID.
- Direct run: python3 /root/accept/accept.py --criteria /root/accept/criteria.json --cases /opt/lab/p1a-criteria/cases.csv --base-url http://127.0.0.1:18095 --out /tmp/run1; echo $?
- Common mistakes: computing p95 with the average or the default of statistics.quantiles, mixing in warm-up or retry requests, also counting an error response as a wrong amount, writing the names of fields that change in the code, and writing the thresholds in the code.
- The grader does not use this server address. The grader starts its own server anew each time.
Four adjectives into four numbers
For each of the four customer sentences, decide the metric, op, threshold, and source and write them into /root/accept/criteria.json together with timeout_ms, passes, and ignore_fields.
The minutes mix the first hoped-for numbers with the final agreement. Write percentages as a ratio (a decimal) and times as integer ms. In source, write the customer sentence from which that criterion came.
Catch the slow tail with p95, not the average
Make /root/accept/accept.py query the case table passes times to leave samples.jsonl and judge the p95_ms criterion.
Leave the ms measured with time.perf_counter for each request, sort the ms of all requests, and pick the ceil(0.95 × n)-th value. The default of statistics.quantiles is interpolation, so it gives a different value. The number of requests is exactly passes × the number of cases.
Count a response that never came as an error too
Add error_rate to /root/accept/accept.py so that it counts responses that are not 200 and timeout_ms overruns as errors.
Pass the criteria table's timeout_ms, in seconds, as the timeout argument of urlopen. An HTTPError is a response that has a status code, and a timeout arrives as TimeoutError or URLError. Leave timeout as a string in the evidence's status.
Find the amounts that are off by 1 won
Add mismatches to /root/accept/accept.py so that it compares the total of the first-pass 200 responses with expected_total and leaves mismatched_cases.
The CSV's expected_total is a string, so convert it to an integer before comparing. An error response was already counted in the error rate, so leave it out of the amount comparison. If you count one problem twice under two criteria, it becomes blurry which criterion is the cause.
Query twice and find the cases that wobble
Add rerun_diffs to /root/accept/accept.py so that it compares the first-pass and second-pass bodies excluding the criteria table's ignore_fields and leaves rerun_diff_cases.
If you write a field name like request_id in the code, a normal build fails the moment the grader changes the field name. Compare after removing only the criteria table's fields from the dictionary.
Match the report to the evidence yourself
Make /root/accept/accept.py judge the four criteria with the criteria table's op as it is, and make report.json's observed, pass, and metrics equal the values recomputed from samples.jsonl.
Handle op with a table that converts the string into an operation. When the error rate is exactly equal to the threshold, < fails and <= passes. Copy the threshold from the criteria table as it is, and write the evidence one line at a time right after each request.
Separate five candidates with a changed criteria table
Check that /root/accept/accept.py accepts only the normal candidate even with new thresholds, a new case table, and new names of fields that change, and fails a defective candidate on only its corresponding criterion.
If one defect fails several criteria at the same time, you cannot separate the cause. Look for any places where you wrote thresholds, field names, or cases in the code.
Judge the acceptance of the customer candidate rc2
Start the rc2 server on port 8095, run with criteria.json and the provided case table, and leave /root/accept/rc2/report.json and /root/accept/rc2/samples.jsonl.
Do not fix the criteria table to fit the result. The grader checks whether criteria.json is the same as the step 1 agreement, whether the report is the same as the evidence, and whether the evidence is the actual rc2 responses. When you are done with the server, pick out and stop only that PID.