Founding as a Developer — Validate Before You Build
Judge Four Experiments by the Rules
Goal
Build functions that calculate SRM, the z-test, confidence intervals, sample size, and Holm correction with the standard library, and judge the results of four experiments by fixed rules.
Why it matters
A small team's A/B tests have small samples, are watched every day, and get split up when there is no result. All three are habits that turn chance into discovery. You must set the rules first and calculate by those rules for an experiment to protect the product decision.
Materials
/opt/fixtures/founder/ab/exp1.csv—user_id,variant,day,segment,converted(A is the control, B is the change, 14 days, 10 segments)/opt/fixtures/founder/ab/exp2.csv—user_id,variant,converted(a 50:50 design)/opt/fixtures/founder/ab/exp3_daily.csv—day,a_users,a_conv,b_users,b_conv(the daily increments of an A/A experiment)/opt/fixtures/founder/ab/exp4.csv—variant,users,conversions(control and v1 through v5)
Definitions
- SRM: chi-square goodness of fit against the expected 50:50, 1 degree of freedom.
chi2 = Σ(관측−기대)²/기대(the Korean words in the code mean observed and expected),p = math.erfc(math.sqrt(chi2/2)). SRM if p < 0.001. - z-test:
diff = pb − pa. z uses the standard error√(p(1−p)(1/na+1/nb))of the pooled proportionp = (ca+cb)/(na+nb), and the two-sidedp = 2(1−Φ(|z|)). - 95% confidence interval:
diff ± Φ⁻¹(0.975)·√(pa(1−pa)/na + pb(1−pb)/nb)(unpooled). Φ isstatistics.NormalDist(). - Decision: p < 0.05 and ci_low > 0 →
"ship_b", p < 0.05 and ci_high < 0 →"keep_a", otherwise"inconclusive". - Sample size (per side): the formula in the reading, α = 0.05 two-sided, power 0.8,
math.ceil. - Holm: sort p ascending, and at the k=0,1,… th, reject if
p·(m−k) ≤ 0.05, and stop at the first that fails. - Round every p, rate, and difference to the sixth decimal place.
Steps
- In
/root/founder/ab/ab.py, createsrm(n_a, n_b)→{"chi2": x, "p": y}, and in/root/founder/ab/srm.json, write{"a": n, "b": n, "chi2": x, "p": y, "srm": true/false}for each of exp1 and exp2. - In
ab.py, createztest(ca, na, cb, nb)→{"diff", "z", "p", "ci_low", "ci_high"}(ca and cb are conversion counts, na and nb are user counts). - In
/root/founder/ab/result.json, writea_users,a_conv,b_users,b_conv,diff,p,ci_low,ci_high,decisionfor exp1. - In
ab.py, createsample_size(p_base, mde), and in/root/founder/ab/power.json, writebaseline(exp1's A conversion rate, to four decimal places),mde(0.01),n_per_arm(= sample_size(baseline, 0.01)),exp1_min_arm(the smaller of exp1's two user counts), andpowered(exp1_min_arm ≥ n_per_arm). - In
/root/founder/ab/peeking.json, write for exp3daily_p(the 21 p-values computed with the cumulative totals from day 1),first_significant_day(the first day with cumulative p < 0.05, null if none),final_p, andfinal_decision(the decision made with the last day's cumulative values). - In
ab.py, createholm(pvalues)(a name → p dictionary → a sorted list of the names that are rejected), and in/root/founder/ab/multi.json, write for exp4p(each of v1 through v5 against control),naive(the names with p < 0.05, sorted), andholm. - In
/root/founder/ab/segments.json, writep,naive(p < 0.05, sorted), andholmfor each of the 10 segments of exp1. - In
/root/founder/ab/decision.json, writeexp1(the decision),exp2_trustworthy(true if there is no SRM),exp3_decision(the last day's decision),exp4_ship(the list Holm rejects), andsegment_claims(the list of segments Holm rejects).
Notes
from statistics import NormalDist; N = NormalDist(); N.cdf(z); N.inv_cdf(0.975)- Common mistakes: using the pooled standard error for the confidence interval, a one-sided p, computing each day's p without accumulating the increments, reporting segments without correction, and judging an experiment with SRM as it is.
Check the sample ratio first
In /root/founder/ab/ab.py, create srm(n_a, n_b), and in /root/founder/ab/srm.json write a, b, chi2, p, and srm for exp1 and exp2.
The expected value is (n_a+n_b)/2. The chi-square tail probability with 1 degree of freedom can be computed as math.erfc(math.sqrt(chi2/2)). The SRM threshold is p < 0.001.
The two-proportion z-test and the 95% confidence interval
In ab.py, create ztest(ca, na, cb, nb). Return diff, z, p (two-sided), ci_low, and ci_high to six decimal places.
The test statistic uses the standard error of the pooled proportion, and the confidence interval uses the standard error of each side's own proportion (unpooled). Φ is statistics.NormalDist().cdf, and the critical value is inv_cdf(0.975).
Judge exp1
In /root/founder/ab/result.json, write a_users, a_conv, b_users, b_conv, diff, p, ci_low, ci_high, and decision for exp1.
Decision rule: ship_b if p < 0.05 and the lower bound > 0, keep_a if p < 0.05 and the upper bound < 0, and inconclusive otherwise.
Was this experiment big enough to catch 1 percentage point
In ab.py, create sample_size(p_base, mde), and in /root/founder/ab/power.json write baseline, mde, n_per_arm, exp1_min_arm, and powered.
p2 = p1 + mde, p̄ = (p1+p2)/2, and the z values are inv_cdf(0.975) and inv_cdf(0.8). Divide the whole parenthesized quantity by mde, square it, and round up.
If you look at an A/A experiment every day
In /root/founder/ab/peeking.json, write daily_p (21 cumulative values), first_significant_day, final_p, and final_decision for exp3.
exp3_daily.csv holds that day's increments. Add them up from day 1 and call ztest with the cumulative totals through that day. Count days from 1.
Five variants — Holm correction
In ab.py, create holm(pvalues), and in /root/founder/ab/multi.json write p (v1 through v5), naive, and holm for exp4.
Run ztest on each variant against control (control is the a side). Holm looks at p·(m−k) ≤ 0.05 for the k-th smallest (from 0), and stops at the first failure.
If you split into ten segments
In /root/founder/ab/segments.json, write p, naive, and holm for the 10 segments of exp1.
Count A and B separately for each segment and call the same ztest, and apply the same holm to the 10 p-values.
A ship decision per experiment
In /root/founder/ab/decision.json, write exp1, exp2_trustworthy, exp3_decision, exp4_ship, and segment_claims.
Gather them from the files of the earlier steps. An experiment with SRM is an experiment you cannot trust. The conclusion of the peeking counterexample is the last day's decision, not the first significant day.