TT Lab
Get started
Learn Learning paths Courses

Founding as a Developer — Validate Before You Build

Code Interview Notes as Evidence and Pick a First Segment

Continue in TT Lab

Goal

Filter out the bias from question forms in the interview records, count only the evidence of "experienced recently and paid a cost", and choose the first customer segment using an interval that reflects sample size.

Why it matters

People answer politely to someone else's idea. If you read the answers to leading and hypothetical questions, and compliments, as demand, you spend months on a market that does not exist. The habit of encoding records as evidence and accounting even for sample size is the whole of pre-build validation.

Materials — /opt/fixtures/founder/problem/interviews.csv

id,segment,question_style,said_problem,days_since_last,workaround,spent_krw_month,commitment

Definitions

Steps

  1. In /root/founder/problem/counts.json, write all (the number of interviews) and by_style (the count for each of past, hypothetical, and leading) for each segment.
  2. In /root/founder/problem/said.json, write said_all_rate (the said_problem rate over all interviews) and said_usable_rate (over usable interviews only) for each segment.
  3. In /root/founder/problem/problem.py, create evidence(row) (one row of csv.DictReader → True/False).
  4. In /root/founder/problem/evidence.json, write usable, evidence, and rate (evidence ÷ usable) for each segment.
  5. In /root/founder/problem/signals.json, write for each segment median_spend (the median monthly spend of people who are usable, have a recent problem, and spend money; 0 if none), strong (the number of pilot and preorder among usable interviews), and compliments (the number of compliment among all interviews).
  6. In problem.py, add wilson(k, n) → [low, high] (if n is 0, [0.0, 1.0]).
  7. In /root/founder/problem/ranking.json, write by_said_all (the segment with the highest said_all_rate), by_point (the highest evidence rate), by_lower_bound (the highest Wilson lower bound of the evidence), and lower_bounds (segment → lower bound).
  8. In /root/founder/problem/decision.json, write target (by_lower_bound), need_more (segments with lower bound < 0.5 ≤ upper bound, sorted), and drop (segments with upper bound < 0.5, sorted).

Notes

Who was asked which questions

In /root/founder/problem/counts.json, write all and by_style (past, hypothetical, leading) for each segment.

Read with csv.DictReader and count by segment and question_style. See whether the share of question forms differs from segment to segment.

The question form inflates the answer

In /root/founder/problem/said.json, write said_all_rate and said_usable_rate for each segment.

The numerator is the interviews with said_problem == "1"; for the denominator, the former uses the whole segment, and the latter uses the interviews whose question_style is past.

The evidence condition as a function

In /root/founder/problem/problem.py, create evidence(row) (past question · experienced within the last 30 days · money, or a pilot or preorder).

All three conditions must hold for it to return True. Filter out an empty days_since_last before converting it to int. The 30 days is inclusive.

Evidence by segment

In /root/founder/problem/evidence.json, write usable, evidence, and rate for each segment.

The denominator is the number of usable (past) interviews. Do not divide by the total number of interviews.

Money, commitments, and compliments

In /root/founder/problem/signals.json, write median_spend, strong, and compliments for each segment.

The population for median_spend is people who are usable, have a recent problem, and spend more than 0 (0 if none). Count compliments regardless of question form.

A small sample as an interval

In problem.py, add wilson(k, n) → [low, high].

z is NormalDist().inv_cdf(0.975). Transcribe the center and half-width formulas from the instructions as they are. If n=0, [0.0, 1.0].

Three rankings

In /root/founder/problem/ranking.json, write by_said_all, by_point, by_lower_bound, and lower_bounds.

You rank the same records three times — by the "has the problem" rate of all answers, by the evidence rate (point estimate), and by the Wilson lower bound of the evidence.

The first segment and the next interviews

In /root/founder/problem/decision.json, write target, need_more, and drop.

If the interval straddles 0.5, you do not know yet (need_more); if even the upper bound is below 0.5, fold it (drop).