Founding as a Developer — Validate Before You Build
Retention Cohorts — Get the Denominator Wrong and Everything Is Wrong
In one line
Retention is "the share of the people who signed up who are still using it N weeks later". Treat the people who signed up in the same week as one bundle (a cohort), make the denominator always the number of signups in that cohort, and leave weeks that have not arrived yet as blank, not 0. Just keeping to these two makes the three most common illusions of the early startup days disappear.
Why this was needed
Monthly active users had been rising every month. But new signups were rising even faster — the bucket had a hole and they were pouring water in faster. The overall active count mixes the water coming in and the water leaking out into one number. Whether the product holds on to people only shows when you look separately at how many of the people who came in at the same time remain as time passes. That is a cohort.
Whether to raise ad spend, build more features, or stop and rethink is usually decided by this one curve. If the curve flattens at some point (stops dropping further), it means there are people who have been held, and if it keeps descending toward 0, it means the product has not yet become anyone's habit. So if the calculation is wrong, the decision is wrong in its entirety.
How it works
Cohorts and week numbers. In this lab a cohort is the Seoul-time week (starting Monday) of the signup time. Retention at week N is "counting the signup week as week 0, there was a session_start in the Nth week". Drawn as a table, it becomes a triangle with cohorts as rows and N as columns — the later a cohort signed up, the more of its right-hand cells are empty.
Trap 1 — the denominator. If you take the denominator of week-N retention as "people active in week 0", people who signed up and never came even once quietly drop out and retention goes up. It amounts to erasing a problem in the signup funnel (a failed first experience) from the retention table. The denominator is the number of signups. If you want to look at the first experience separately, record it as a different metric, the activation rate.
Trap 2 — unobserved cells. Week 6 for the cohort that signed up in week 11 has not arrived yet. If you fill that cell with 0 and average several cohorts, the more recent cohorts there are, the more falsely the curve bends down. When you combine several cohorts into a week-N retention, collect only the cohorts whose observation through week N is complete, and weight by number of people (a simple average of cohort rates gives a 20-person cohort and a 200-person cohort the same weight). In survival analysis, such cells are called "censoring".
Trap 3 — survivorship bias. "We asked the people who are still using it now, and most of them were still using it at week 4" is only natural. If you look at the past by picking only those who remain now, the people who left are not in the sample. A cohort fixes its roster at signup time and follows that roster to the end. You must not reselect the roster using later information (are they active now, are they paying).
Two kinds of retention. "Did they come in week N (bounded)" and "did they come at least once in week N or after (unbounded)" are different questions. The former suits a tool used every week, and the latter suits a tool used occasionally (tax filing, travel booking), matching the nature of the product. If you do not write down which of the two you use, two people look at the same table and state different numbers.
Where the curve flattens. This lab defines the flattening point as "the first N at which the drop in retention from the previous week is under 2 percentage points". The 2-point threshold is an example rule of this lab. Set it to fit the product and sample size, but once set, do not change it.
What activity to count. This lab treats "there was a session_start that week" as activity. As you saw in the metrics module, if you change the definition of activity to the core action (share_doc), the whole cohort curve moves down and the point where the curve flattens changes too. Rather than one definition being right, when you compare curves you must compare only those with the same definition. If the curve improved in the month you changed the definition, it is the definition that changed, not the product.
What it looks like in the field
- In a growth meeting someone said "retention is better than last month", and it turned out to be the month the denominator was switched to week-0 active users.
- The cohort average on the dashboard has kept falling for the last two months. In fact, the empty cells of new cohorts were being filled with 0.
- When you split the cohorts by channel, overall retention is fine but only the cohort that came in through ads is almost 0 at week 4. If you raise ad spend, all that is left is for the overall curve to get worse. We look again in the funnel module at how the people who come in are themselves different depending on the acquisition path.
How to notice when you are wrong
- If the week-0 column of the retention table is all 1.0, it is a sign that you took the denominator as week-0 active users. If you use the number of signups as the denominator, even week 0 can be less than 1.
- If zeros line up in the right-hand cells of recent cohorts, check whether you filled the not-yet-observed cells with 0. Those cells must be empty.
- If the combined curve resembles none of the cohorts' curves, check the weights (number of people).
- If the conditions used to pick the roster include "current" information (currently active, currently paying), it is survivorship bias.
What you will do in the next lab
From the 12 weekly cohorts of "Moanote", you build the retention triangle and write side by side the numbers where you deliberately commit the three traps (denominator, blanks, survivors) and the correct numbers. Then you calculate the weighted retention that collects only the cohorts whose observation is complete, the difference between bounded and unbounded, and the week-4 retention by signup channel, and bundle them into a one-page report. The grader also runs your functions against variant materials.