Founding as a Developer — Validate Before You Build
A/B Testing — What to Check Before the p-Value
In one line
The conclusion of an A/B test does not come from a single p-value. (1) Is the split as designed (SRM)? (2) Did enough people gather to catch the effect (power)? (3) Did you look once, at a point set in advance (peeking)? (4) How many comparisons did you make (multiple comparisons)? — check these four first, and then state the size of the effect with a confidence interval.
Why this was needed
The smaller the team, the more it hurries experiments. With little traffic, they open the dashboard every day and ship on the day p drops below 0.05, saying "we won". If nothing shows up, they split the results by device and by new versus returning users, and when one comes out significant, they build a feature for that segment. Both are common, and both are the quickest ways to mistake chance for a discovery.
Johari et al.'s Peeking at A/B Tests (KDD 2017) reports that the classical p-value and confidence interval hold only when the sample size was fixed in advance, and that if you keep watching the dashboard and stop the moment it turns significant, the false positive probability inflates greatly. Also, Fabijan et al.'s study of sample ratio mismatch (KDD 2019) treats SRM, where the observed allocation ratio differs from what was expected, as a "symptom" of various data quality problems, and states that the analysis of an experiment with SRM cannot be trusted.
How it works
SRM — first of all. If you split 50:50, the number of users on the two sides should be similar. Measure the gap from the expected value with a chi-square goodness-of-fit test (1 degree of freedom), and this lab treats p < 0.001 as SRM. A bug that drops some users of one side from the records (a failed redirect, a bot filter, an app version) also tilts the conversion rate to one side. When you see SRM, do not interpret the conversion rates; look for the cause first.
The two-proportion z-test and the confidence interval. Difference = B conversion rate − A conversion rate. The test statistic gets its standard error from the pooled proportion under the assumption that "the two sides are the same", and the 95% confidence interval multiplies the unpooled standard error, computed from each side's own proportion, by 1.96 (precisely, NormalDist().inv_cdf(0.975)). The Python standard library's statistics.NormalDist provides the normal distribution's cdf and inv_cdf, so no external package is needed. If the confidence interval does not include 0 and points in the positive direction, then in addition to "B is better" you can say how much better (what the lower bound is).
Power and sample size. To detect an absolute difference of δ from a baseline conversion rate p1 at significance level α (two-sided) and power 1−β, the number of people needed on each side, by normal approximation, is
n = ⌈ ( z(1−α/2)·√(2·p̄(1−p̄)) + z(1−β)·√(p1(1−p1) + p2(1−p2)) )² / δ² ⌉, p2 = p1 + δ, p̄ = (p1+p2)/2
To detect 1 percentage point near a baseline conversion rate of 11% with α=0.05 and 80% power, you need well over ten thousand on each side (an example). An effect that comes out significant in an under-sampled experiment tends to be estimated as larger than it really is — Gelman and Carlin (2014) called this a Type M (magnitude) error. That is why you fix the sample size before you start the experiment.
Peeking. Even in an A/A experiment (both sides show the same screen), if you compute the cumulative p-value every day, some day it can dip below 0.05. If you look only once, on the day you fixed, that day's p is the conclusion. If you want to look every day, use a method designed for that purpose, such as a sequential test.
Multiple comparisons. If you compare five variants against the control one by one, or split into ten segments, it is only natural that one comes out below 0.05 even with no effect. The step-down method of Holm (1979) multiplies the k-th smallest p-value by (m−k+1), compares it with α, and stops at the first place where it is exceeded. It is less conservative than Bonferroni while keeping the overall false positive probability at or below α.
What it looks like in the field
- "B won in 3 days" — it was significant once, on day 3. Two weeks later p was 0.6.
- An experiment where the changed side recorded 2% fewer users. The cause was people who left before the events were sent because of the new screen's slow loading. If you look only at those who remain, the conversion rate looks high.
- A report that says there is no difference overall but it is significant only for "tablet, new users". This is what commonly happens when you look at ten segments.
How to notice when you are wrong
- Look at the allocation ratio first. If one side is a few percent smaller in a 50:50 experiment, do not interpret the conversion rates.
- Check whether the date you drew the conclusion is the same as the date you set before the experiment. If it differs, you peeked.
- Check whether the report states how many comparisons were made (number of variants × number of metrics × number of segments). If it does not, there was likely no correction either.
- If you state the effect size as a single point estimate, also write the lower bound of the confidence interval.
What you will do in the next lab
With four experiments, you build the SRM check, the two-proportion z-test and confidence interval, and the sample size calculation as functions, see the peeking trap in the daily cumulative p of an A/A experiment, and see the difference before and after Holm correction with five variants and ten segments. Finally you produce a ship decision per experiment on one page. The grader calls your functions again with random inputs.