Put what works next to what does not
One-line summary
Differential diagnosis is a procedure that, instead of staring only at failing samples, puts successful samples side by side and keeps only the attributes that are present in every failure and absent from every success.
Why this is needed
"Some customers work and some don't." It is the sentence heard most often in the field, and at the same time the best news. That there is a side that works means there is a control group to compare against.
Yet most people throw this news away and dig only into the logs of failed requests. Failure logs contain TLS handshakes, retries, and warnings. They all look suspicious. To know whether that suspiciousness is actually the cause, you must look at whether the same thing also appears in the logs of successful requests. If it is in the successes too, it is not a cause but background.
This procedure has a name. It is a term that comes from medicine, differential diagnosis, and software borrows the same name. The core is one thing — if you look at only one side, everything looks like a candidate cause, and if you look at both sides, most of it erases itself.
How it works
The procedure has four steps.
1. Gather samples. Gathering only twenty failing samples is meaningless. You must gather a similar number of successful samples. The closer the successful samples are to the successes with conditions most like the failures, the better. A success sample from a completely different country and version erases very little.
2. Expand into attributes. Turn each sample into a list of (attribute = value). Region, client version, encoding, authentication method, device, endpoint, time window. Deciding this list is actually the hardest part — an attribute that is not on the list can never become a candidate.
3. Subtract the two sets. From the set of values present in every failure, subtract the set of values present in any success at all. What remains are the candidates. The values that get erased here matter. The fact that "every failure had token authentication" means nothing the moment it is placed beside the fact that a third of the successes were also token.
4. If there are two or more candidates, separate them. This spot is the real difficulty of the procedure.
Correlation is not cause. If two things deployed at the same time always travel together, you cannot tell them apart from the samples alone. That the table points to both equally is not because the table is wrong but because the data has no case that separates the two. There are only two ways to separate them.
- Get more data. If even one combination in which the two separate comes in, one of them drops out. In the new samples, look for a case where A is present but it succeeded, or B is absent but it failed.
- Intervene. Change only one attribute and hold the rest fixed, and see whether the result flips. This is the difference between observation and experiment, and only an experiment can speak of causation.
There are also cases that do not split on a single attribute. If it fails only when the encoding is utf-8-sig and the device is a kiosk, not a single standalone candidate remains. That is because each one is also in the successful samples. In this case you take pairs of values as candidates and do the same subtraction once more. Going up to combinations of three or more makes the number of cases grow quickly, so you usually look up to pairs and then confirm with an intervention.
The numbers make it clear why this procedure is worth it. Say there are 240 samples, seven attributes, and about twenty values. If you take the intersection of the 40 failures, usually two or three values remain, and subtracting the union of the 200 successes leaves one or two. If you skim by eye, all twenty look suspicious, but two subtractions bring the candidates down to one or two. It is not that the person got smarter; the successful samples erased the rest.
One more thing. This procedure gets stronger the more attributes you add. An attribute that was not recorded cannot be a candidate, so deciding "what to record alongside" at the start of an investigation is in effect half of the investigation. Request headers, client version, tenant, path, and authentication method proved useful in most incidents.
What you see in the field
First, gathering only failing samples. The customer sends only the failed ones. That is because nobody keeps the successful ones. So the first request must always be "please send the same number of successful ones too."
Second, the samples are skewed. If you gather all failures from the nightly batch and all successes from the daytime screens, the time window and the path both remain as candidates. The way the data was gathered left a pattern in the data.
Third, stopping after finding just one candidate. If two candidates remain and you pick the plausible one and report it, you are wrong half the time. And that report makes people spend days on the fix. If there are two candidates, it is right to write that there are two and design an experiment that separates them.
Fourth, not writing down what you erased. The fact that "auth_mode is not the cause" keeps the next person from walking the same path again. In the report, write not only the remaining candidates but also the candidates you erased and the grounds for erasing them.
What really matters in practice
- The successful samples are half. An investigation without a control group is a guess.
- If it is not on the attribute list, there is no candidate. Decide what to record first.
- If two candidates are perfectly correlated, get more data or intervene. Do not pick the plausible one.
- If a single attribute does not split it, look at pairs. An interaction disappears entirely in a single-attribute check.
What you will do in the next lab
With 240 payment gateway samples, you build a difference table by attribute and pick out the values that are in every failure and in no success. At the point where two candidates remain, you add 120 new samples to knock one of them out, and use the reproducer to change just one attribute and confirm that the result flips. Finally, from another tenant's samples that do not split on a single attribute, you find pair candidates and report on one page.