TT Lab
Get started
Learn Learning paths Courses

Founding as a Developer — Validate Before You Build

Judge Four Experiments by the Rules

Continue in TT Lab

Goal

Build functions that calculate SRM, the z-test, confidence intervals, sample size, and Holm correction with the standard library, and judge the results of four experiments by fixed rules.

Why it matters

A small team's A/B tests have small samples, are watched every day, and get split up when there is no result. All three are habits that turn chance into discovery. You must set the rules first and calculate by those rules for an experiment to protect the product decision.

Materials

Definitions

Steps

  1. In /root/founder/ab/ab.py, create srm(n_a, n_b) → {"chi2": x, "p": y}, and in /root/founder/ab/srm.json, write {"a": n, "b": n, "chi2": x, "p": y, "srm": true/false} for each of exp1 and exp2.
  2. In ab.py, create ztest(ca, na, cb, nb) → {"diff", "z", "p", "ci_low", "ci_high"} (ca and cb are conversion counts, na and nb are user counts).
  3. In /root/founder/ab/result.json, write a_users,a_conv,b_users,b_conv,diff,p,ci_low,ci_high,decision for exp1.
  4. In ab.py, create sample_size(p_base, mde), and in /root/founder/ab/power.json, write baseline (exp1's A conversion rate, to four decimal places), mde (0.01), n_per_arm (= sample_size(baseline, 0.01)), exp1_min_arm (the smaller of exp1's two user counts), and powered (exp1_min_arm ≥ n_per_arm).
  5. In /root/founder/ab/peeking.json, write for exp3 daily_p (the 21 p-values computed with the cumulative totals from day 1), first_significant_day (the first day with cumulative p < 0.05, null if none), final_p, and final_decision (the decision made with the last day's cumulative values).
  6. In ab.py, create holm(pvalues) (a name → p dictionary → a sorted list of the names that are rejected), and in /root/founder/ab/multi.json, write for exp4 p (each of v1 through v5 against control), naive (the names with p < 0.05, sorted), and holm.
  7. In /root/founder/ab/segments.json, write p, naive (p < 0.05, sorted), and holm for each of the 10 segments of exp1.
  8. In /root/founder/ab/decision.json, write exp1 (the decision), exp2_trustworthy (true if there is no SRM), exp3_decision (the last day's decision), exp4_ship (the list Holm rejects), and segment_claims (the list of segments Holm rejects).

Notes

Check the sample ratio first

In /root/founder/ab/ab.py, create srm(n_a, n_b), and in /root/founder/ab/srm.json write a, b, chi2, p, and srm for exp1 and exp2.

The expected value is (n_a+n_b)/2. The chi-square tail probability with 1 degree of freedom can be computed as math.erfc(math.sqrt(chi2/2)). The SRM threshold is p < 0.001.

The two-proportion z-test and the 95% confidence interval

In ab.py, create ztest(ca, na, cb, nb). Return diff, z, p (two-sided), ci_low, and ci_high to six decimal places.

The test statistic uses the standard error of the pooled proportion, and the confidence interval uses the standard error of each side's own proportion (unpooled). Φ is statistics.NormalDist().cdf, and the critical value is inv_cdf(0.975).

Judge exp1

In /root/founder/ab/result.json, write a_users, a_conv, b_users, b_conv, diff, p, ci_low, ci_high, and decision for exp1.

Decision rule: ship_b if p < 0.05 and the lower bound > 0, keep_a if p < 0.05 and the upper bound < 0, and inconclusive otherwise.

Was this experiment big enough to catch 1 percentage point

In ab.py, create sample_size(p_base, mde), and in /root/founder/ab/power.json write baseline, mde, n_per_arm, exp1_min_arm, and powered.

p2 = p1 + mde, p̄ = (p1+p2)/2, and the z values are inv_cdf(0.975) and inv_cdf(0.8). Divide the whole parenthesized quantity by mde, square it, and round up.

If you look at an A/A experiment every day

In /root/founder/ab/peeking.json, write daily_p (21 cumulative values), first_significant_day, final_p, and final_decision for exp3.

exp3_daily.csv holds that day's increments. Add them up from day 1 and call ztest with the cumulative totals through that day. Count days from 1.

Five variants — Holm correction

In ab.py, create holm(pvalues), and in /root/founder/ab/multi.json write p (v1 through v5), naive, and holm for exp4.

Run ztest on each variant against control (control is the a side). Holm looks at p·(m−k) ≤ 0.05 for the k-th smallest (from 0), and stops at the first failure.

If you split into ten segments

In /root/founder/ab/segments.json, write p, naive, and holm for the 10 segments of exp1.

Count A and B separately for each segment and call the same ztest, and apply the same holm to the 10 p-values.

A ship decision per experiment

In /root/founder/ab/decision.json, write exp1, exp2_trustworthy, exp3_decision, exp4_ship, and segment_claims.

Gather them from the files of the earlier steps. An experiment with SRM is an experiment you cannot trust. The conclusion of the peeking counterexample is the last day's decision, not the first significant day.