TT Lab
Get started
Learn Learning paths Courses

Founding as a Developer — Validate Before You Build

Write the Hypothesis as a Number, Then Cut the MVP

Continue in TT Lab

In one line

An MVP is not a "small product" but a device that tests the single riskiest hypothesis as cheaply as possible. To do that, write the hypothesis as a numeric sentence containing "who · what · how much", set the criteria for pass, fail, and hold before the test, and read the result by accounting even for who took part in the test. Feature scope comes after that, cut only as far as you need to test the hypothesis.

Why this was needed

"I think people will like it" is not a hypothesis. Whatever comes out, you can insist you were right. To be testable, it has to be able to be wrong. "At least 3% of freelance designer visitors click a prepayment of 9,900 won a month" can be wrong.

The second trap is who you test on. If you send your first landing page first to the founder's newsletter and acquaintances, the conversion rate comes out high. Those people click because they trust the founder, not the product. The number you get when you show the same page to strangers is the market's number.

The third is scope. If you rank feature candidates by score, the things that are easy to build rise to the top. Dark mode is cheap and many people use it, but it does nothing to help test the prepayment hypothesis.

How it works

The hypothesis sentence and the criteria. Write the metric (preorder rate = preorders ÷ visitors), the threshold (for example 3%), and the target (stranger visitors). Judge not by a single rate but by an interval — if the lower bound of the 95% Wilson interval used in the previous module is at or above the threshold, it passes; if the upper bound is below the threshold, it fails; and if the threshold is inside the interval, it is on hold (not yet known). The 3% threshold is an example for this lab; in practice it is better to work it out backward from unit economics (price and CAC) — this is covered in the pricing module.

If on hold, increase the sample. If the current rate stays the same, you can calculate how many visitors you need before the interval moves away from the threshold. This lab uses a simple method: it increases n from 1 and recomputes the interval with k = round(p·n). If the visitors needed are far more than now, that also means the channel is expensive as a means of testing.

Look at the warm audience separately. If you look only at a single number with the channels combined, the acquaintance channel, with its high conversion rate, pulls the average up and it looks like "almost a pass". Make a separate judgment for only the stranger channels, excluding the acquaintance channel, and base the decision on that.

Cut scope with RICE, but hypothesis first. Intercom's RICE article (Sean McBride) explains that the score is (Reach × Impact × Confidence) ÷ Effort, with Impact measured on a scale of 3, 2, 1, 0.5, and 0.25, Confidence as 100%, 80%, or 50%, and Effort in person-months. This lab cuts scope with a greedy method that takes items in descending score order and includes only those that fit the budget (person-months). Then it checks the result against the hypothesis — if a feature that does not test the hypothesis (for example dark mode) has made it into the scope despite a high score, that is a signal that the hypothesis, not the scoring table, should set the scope.

Decision Basis
Should we build it Is the stranger channels' verdict a pass
Should we test more Among on-hold channels, the one that needs the fewest additional visitors
What should we build Features that test the hypothesis → in RICE order within the budget

The ethics of a smoke test. If you take prepayments for a product that does not exist yet, tell people clearly right after payment that "it is not ready yet and you will be fully refunded", and actually refund them. If people feel deceived, you can no longer use that channel for the next test, and above all, trust is almost the only asset an early product has. You can also test with a weaker signal, such as "waitlist + price confirmation" instead of payment, but you must reflect in the decision criteria that the signal is that much weaker.

What it looks like in the field

How to notice when you are wrong

What you will do in the next lab

From the smoke-test results of four channels, you make per-channel, overall, and stranger-channels-only verdicts, and calculate how many more visitors the on-hold channels need. You cut scope within the budget using the RICE scores of eight feature candidates, and finally decide on one sheet "should we build, should we test more, what should we build".