Working customers and failing ones — erasing candidates with a difference table
Goal
You split successful and failing samples by attribute to build a difference table, and keep as candidates only the values that are in every failure and in no success. You knock out perfectly correlated false candidates with new samples and with an intervention experiment respectively, and also find the interaction that does not split on a single attribute.
Why it matters
"Some customers work and some don't" sounds like bad news but is actually good news. That there is a side that works means there is a control group. If you dig only into failure logs, every line looks suspicious, but if the same line is also in the success logs, it is not a cause but background. What is hard in this lab is not the code that builds the table but what you do when two candidates remain. If two things deployed on the same day always travel together, you cannot tell them apart from the samples alone. If you pick the plausible one and report it then, you waste days half the time. There are only two ways to separate them. Get more samples in which the two separate, or change just one attribute and hold the rest fixed and see whether the result flips. The former is observation and the latter is experiment, and only an experiment can speak of causation. The grader does not trust your conclusions. It builds samples in which the cause and the false candidates are planted in a different place each time, actually runs your tool, and compares the set of candidates you picked with the planted answer.
Steps
- Create and run /root/diff/gen_cases.py to create /root/diff/cases.json (240 cases), cases2.json (120 cases), cases3.json (180 cases), and repro.py.
- In /root/diff/table.json, build a difference table that records the failure count and success count for each value of each attribute.
- Create /root/diff/differential.py so that it picks as candidates the values that are in every failure and in no success.
- In /root/diff/candidates.json, write the candidate list and the eliminated candidates along with the grounds.
- Merge in the new samples and run again, and in /root/diff/verdict.json write the one surviving cause, the candidates that dropped out, and the sample numbers that were the grounds.
- Add pair analysis to differential.py and write the pair candidates of cases3.json in /root/diff/pairs.json.
- Use repro.py to change just one attribute and check whether the result flips, and write it in /root/diff/confirm.json.
- Report in /root/diff/summary.json and /root/diff/diff_report.md in four sections.
Notes
- Sample format: each case is
{"case_id": …, "outcome": "ok"|"fail", "region": …, "client_build": …, "encoding": …, "auth_mode": …, "device": …, "endpoint": …, "window": …}. Everything except case_id and outcome is an attribute. - Tool contract:
python3 /root/diff/differential.py --cases <파일> [--cases <파일2> …] [--pairs] --out <json>(the placeholders are the case file paths and the output path) prints one JSON object containing cases, failures, successes, candidates, and eliminated to standard output. If you pass--pairs, it also contains single_candidates and pair_candidates. - Definition of a candidate: an
속성=값string (in the form attribute=value) that is in every failing sample and in no successful sample. The notation joins with an equals sign, as inencoding=utf-8-sig. - Definition of an eliminated candidate: a value that is in every failure but also in at least one success. The grounds for its not being a candidate are in the successful samples.
- Definition of a pair candidate: a case where two values are together in every failure and no success has both. A pair that includes a value that is already a candidate by itself is not new information, so leave it out.
- Reproducer contract:
python3 /root/diff/repro.py --region … --client-build … --encoding … --auth-mode … --device … --endpoint … --window …prints one word, ok or fail. It always gives the same answer for the same combination. - Common mistakes: gathering only failing samples, picking the plausible one when there are two candidates, not recording the eliminated candidates, and changing two things at once in an intervention experiment.
- Assumption of this lab: the attribute list of the samples is fixed as left by the gateway. In the field, deciding this list is itself part of the investigation.
Get the samples and the reproducer in hand
Create and run /root/diff/gen_cases.py to create /root/diff/cases.json (240 cases), cases2.json (120 cases), cases3.json (180 cases), and repro.py.
What you hold in your hands in the field is not a conclusion but samples. Just save this script as it is and run it. Briefly open the generated cases.json and see in what shape the successes and failures are written.
Build the difference table by attribute
In /root/diff/table.json, write failures, successes, and attributes. For each attribute name, attributes holds {"fail": 건수, "ok": 건수} (the failure count and the success count) per value. case_id and outcome are not attributes.
The table comes first. If you skim by eye without a table, you cannot tell 'a value that appears often among failures' from 'a value that appears only among failures.' Read the attribute names from the samples as they are — if you write them by hand, you miss one, and a missed attribute can never become a candidate.
Build the tool that picks candidates
Create /root/diff/differential.py so that it outputs as candidates the 속성=값 (attribute=value) strings that are in every failure and in no success, and as eliminated the values that are in every failure but also in successes. --cases must be allowed to be given more than once.
Two set operations are all it takes. Take the intersection of the failing samples and subtract the union of the successful samples. The values that vanish in the subtraction are eliminated. The reason to use these two lines instead of a person skimming the table by eye is that the procedure is the same even when there are twenty attributes.
Write the candidates and eliminated candidates with grounds
In /root/diff/candidates.json, write candidates (the candidate list) and eliminated (the eliminated candidates). Each item in eliminated is {"attribute": "속성=값", "reason": "…"} (where the attribute value has the attribute=value form), and reason must contain, as a number, the count of successful samples in which that value appeared.
The reason to write down what you eliminated is so that the next person does not walk the same path again. The fact that 'every failure had token authentication' has meaning only when placed beside the fact that token was also present in some number of successes. Take that number from the table.
Knock out the false candidate with new samples
Put cases.json and cases2.json in together and run again, and in /root/diff/verdict.json write cause (the one surviving cause), eliminated_by_new_batch (the candidates that dropped out), and evidence_case_ids (the case numbers from cases2.json that were the grounds).
The reason you could not tell the two apart from the samples was that the data had no combination that separates them. The new samples contain that combination. The cases that serve as grounds are of two kinds — a case that had the attribute to be dropped but succeeded, or a case that lacked that attribute but failed.
When a single attribute does not split it
Add --pairs to differential.py so that it finds pair candidates, run it on cases3.json, and write single_candidates and pair_candidates in /root/diff/pairs.json. Write a pair as a list containing the two values.
There are samples with not a single standalone candidate. That is because they fail only when two values are together. Among the values present in all failures, pair them two at a time and look for pairs where no success has both. A pair that includes a value that is already a candidate by itself is not new information.
Change just one attribute and see whether it flips
In /root/diff/confirm.json, write claim (the cause you assert), base (the combination of seven baseline attributes), flipped_attribute, base_outcome, and flipped_outcome. The base and the flipped combination must differ in exactly one attribute, and the two outcomes must differ from each other.
Samples are observation and the reproducer is experiment. If you hold everything else fixed and change only the one attribute you claim, and the result flips, that attribute is not a correlation but a cause. If you change two things at once, you cannot tell which one flipped it.
Report on one page, including what you eliminated
In /root/diff/summary.json, write week1_cases, week1_candidates, cause, eliminated_by_new_batch, pair_only_tenant, and intervention, and in /root/diff/diff_report.md, report in four sections: ## 무엇이 갈렸나 ## 무엇을 지웠나 ## 상관인가 원인인가 ## 남은 위험 (in order, these mean: what split, what was eliminated, correlation or cause, and the remaining risk).
If you write only the remaining candidates in the report, the next person walks the same path again. Write the eliminated candidates and the grounds for eliminating them together. The two results of the intervention experiment are words rather than numbers, but they are evidence.