Founding as a Developer — Validate Before You Build
The Problem Lives in Customers' Past Behavior
In one line
In a customer interview, what you can trust is "I did", and what you should not trust is "I would" and "sounds nice". Ask about the other person's recent behavior in concrete terms, and encode the record by when they last experienced the problem, whether they are spending money or time on it now, and whether they committed to a next step. And when you choose a segment, look not at the rate but at an interval that accounts for sample size.
Why this was needed
When a developer starts a company, the first thing they build is the product, and the last thing they do is talk to customers. Even when they do talk, they ask, "Would you use a service like this?" Almost everyone answers, "Sounds nice, I think I'd use it." Polite people do not tear down someone else's idea to their face. Three months later, those people do not show up for the product you launched.
Rob Fitzpatrick's book "The Mom Test" (2013) sums up the rules for avoiding this trap in three — talk about the other person's life, not your idea; ask about specific things in the past, not generalities or opinions about the future; and talk less and listen. The book also advises you to treat as real signals not compliments but commitments that carry a cost, such as time, reputation, or money.
How it works
The form of a question determines the quality of the answer. "When did you last send a quote? What did you make it with?" asks about a fact. "Making quotes is a hassle, isn't it?" is leading, and "Would you use it if it made them for you automatically?" is hypothetical. The answers to leading and hypothetical questions are opinions, so you do not use them as evidence of a problem. If you record the form of each question as you collect the notes, you can filter them out later.
Encode the evidence. This lab sets three conditions for counting an interview as "evidence".
| Condition | Record | Why |
|---|---|---|
| Past-behavior question | question_style = past | A fact, not an opinion |
| Experienced recently | Within 30 days of the last time they experienced it | A hassle from long ago gives little reason to solve it now |
| Paid a cost | Spends money now, or committed to a pilot or preorder | A priority shown by action, not words |
The 30-day threshold is an example for this lab. Set it to fit your product's usage cycle, but once you set it, apply it the same way to every interview.
When you choose a segment — rate and sample. Which is the stronger signal: 8 out of 10 people (80%) or 21 out of 30 people (70%)? Looking only at the point estimate, it is the first, but 10 people are swayed a lot by chance. Among confidence intervals for a binomial proportion, Brown, Cai, and DasGupta (2001, Statistical Science) showed in a comparative study that the Wilson score interval works well even when the sample is small. This lab ranks segments by the lower bound of the 95% Wilson interval — in effect comparing on "at least this much holds". When the interval is wide and straddles the threshold (here 50%), it is a signal to do more interviews rather than a conclusion.
The record format shapes the analysis. If you summarize from memory after the interview ends, only the good remarks survive. So you attach a code, one line at a time, right after the conversation — what form of question it was, whether they mentioned a problem, when they last experienced it, what they solve it with now and how much they spend, and what they committed to as a next step. Once these five columns are collected, even a hundred interviews become one table, and you can recount them in code, as in this lab. If two people attach the codes, code a few interviews together first so that you agree on the criteria.
What it looks like in the field
- "Making quotes is a hassle, isn't it?" is mixed in only into conversations with the segment the founder has in mind, so that segment's "has the problem" rate comes out highest. Keep only past-behavior questions and the ranking flips.
- "Sounds nice, let me know when it's out" has piled up dozens of times, yet there is not a single preorder. The number of compliments is not evidence of demand.
- You chose a segment from ten interviews, and the next ten people say something entirely different. If you had looked at the interval, it would have been "we don't know yet" from the start.
How to notice when you are wrong
- First count the share of question forms in each segment. If leading and hypothetical questions are concentrated in only one segment, that segment's "has the problem" rate is likely a number the questions created.
- Write the number of compliments and the number of commitments side by side. If compliments pile up at several times the number of commitments and pilots and preorders do not grow, the conversations were pleasant but demand has not been confirmed.
- If the interval straddles the threshold, do not write a conclusion. "We don't know yet" is also a conclusion — it becomes the basis for deciding whom to do the next ten interviews with.
- Write down how you gathered the interviewees. If you met only the founder's acquaintances and followers, the sample itself is already friendly.
- Check that you have not counted the same person twice. If you count several employees of one company separately, that company's circumstances look inflated into a signal for the whole segment.
What you will do in the next lab
From about 90 interview records across three segments, you measure how much the form of the question inflates the answer, turn the evidence conditions into a function, and count per segment. You count money, commitments, and compliments side by side, then re-rank with the Wilson interval and choose the first customer segment. The grader also runs against records that shake your function.