Founding as a Developer — Validate Before You Build
Write the Hypothesis as a Number, Then Cut the MVP
In one line
An MVP is not a "small product" but a device that tests the single riskiest hypothesis as cheaply as possible. To do that, write the hypothesis as a numeric sentence containing "who · what · how much", set the criteria for pass, fail, and hold before the test, and read the result by accounting even for who took part in the test. Feature scope comes after that, cut only as far as you need to test the hypothesis.
Why this was needed
"I think people will like it" is not a hypothesis. Whatever comes out, you can insist you were right. To be testable, it has to be able to be wrong. "At least 3% of freelance designer visitors click a prepayment of 9,900 won a month" can be wrong.
The second trap is who you test on. If you send your first landing page first to the founder's newsletter and acquaintances, the conversion rate comes out high. Those people click because they trust the founder, not the product. The number you get when you show the same page to strangers is the market's number.
The third is scope. If you rank feature candidates by score, the things that are easy to build rise to the top. Dark mode is cheap and many people use it, but it does nothing to help test the prepayment hypothesis.
How it works
The hypothesis sentence and the criteria. Write the metric (preorder rate = preorders ÷ visitors), the threshold (for example 3%), and the target (stranger visitors). Judge not by a single rate but by an interval — if the lower bound of the 95% Wilson interval used in the previous module is at or above the threshold, it passes; if the upper bound is below the threshold, it fails; and if the threshold is inside the interval, it is on hold (not yet known). The 3% threshold is an example for this lab; in practice it is better to work it out backward from unit economics (price and CAC) — this is covered in the pricing module.
If on hold, increase the sample. If the current rate stays the same, you can calculate how many visitors you need before the interval moves away from the threshold. This lab uses a simple method: it increases n from 1 and recomputes the interval with k = round(p·n). If the visitors needed are far more than now, that also means the channel is expensive as a means of testing.
Look at the warm audience separately. If you look only at a single number with the channels combined, the acquaintance channel, with its high conversion rate, pulls the average up and it looks like "almost a pass". Make a separate judgment for only the stranger channels, excluding the acquaintance channel, and base the decision on that.
Cut scope with RICE, but hypothesis first. Intercom's RICE article (Sean McBride) explains that the score is (Reach × Impact × Confidence) ÷ Effort, with Impact measured on a scale of 3, 2, 1, 0.5, and 0.25, Confidence as 100%, 80%, or 50%, and Effort in person-months. This lab cuts scope with a greedy method that takes items in descending score order and includes only those that fit the budget (person-months). Then it checks the result against the hypothesis — if a feature that does not test the hypothesis (for example dark mode) has made it into the scope despite a high score, that is a signal that the hypothesis, not the scoring table, should set the scope.
| Decision | Basis |
|---|---|
| Should we build it | Is the stranger channels' verdict a pass |
| Should we test more | Among on-hold channels, the one that needs the fewest additional visitors |
| What should we build | Features that test the hypothesis → in RICE order within the budget |
The ethics of a smoke test. If you take prepayments for a product that does not exist yet, tell people clearly right after payment that "it is not ready yet and you will be fully refunded", and actually refund them. If people feel deceived, you can no longer use that channel for the next test, and above all, trust is almost the only asset an early product has. You can also test with a weaker signal, such as "waitlist + price confirmation" instead of payment, but you must reflect in the decision criteria that the signal is that much weaker.
What it looks like in the field
- You built an investor deck on a 7% preorder rate from a landing page sent to 200 acquaintances, but the people who came in through ads were in the 1% range.
- You started the MVP with login, notifications, and a settings screen, and three months went by without a payment button.
- You read an on-hold result as "it will pass if we collect a little more". When you calculate the sample you need, it is several times the current one.
How to notice when you are wrong
- Look at the per-channel verdicts first, and check that the combined verdict is not being pulled by one channel's number. If the combined result differs from the result without the warm audience, make the decision on the latter.
- Stop if you feel like changing the threshold after seeing the results. Once the criterion starts following the conclusion, it is persuasion, not a test.
- For each feature that went into the MVP scope, ask "Can the hypothesis not be tested without this feature?" If more than half answer no, the scope is following the scoring table, not the hypothesis.
- If the sample you need is several times your current visitors, before waiting on the same channel, first look for a channel where you can collect the sample more cheaply.
What you will do in the next lab
From the smoke-test results of four channels, you make per-channel, overall, and stranger-channels-only verdicts, and calculate how many more visitors the on-hold channels need. You cut scope within the budget using the RICE scores of eight feature candidates, and finally decide on one sheet "should we build, should we test more, what should we build".