Founding as a Developer — Validate Before You Build
Build the Retention Triangle and See Three Traps in Numbers
Goal
Build a retention triangle from the weekly signup cohorts of "Moanote", confirm the three traps of denominator, blanks, and survivors with numbers, then judge with the correct weighted retention and retention by signup channel.
Why it matters
The overall active count mixes new entrants and leavers into one number. You can see whether the product holds on to people only when you follow the people who signed up in the same week separately. And that calculation easily looks better or worse just by changing the denominator or filling the blanks with 0 — it is a number ad spend and the roadmap depend on, so it must be recomputed exactly as defined.
Materials and definitions
/opt/fixtures/founder/product/users.csvandevents.csv(the same materials as the metrics module), and the signup channel is the channel of the last touch before signup in/opt/fixtures/founder/product/touches.csv(inuser_id,ts,channel, the line with the latest ts).- Weeks by Seoul time (starting Monday 00:00), week 0 = 2026-03-02, and observation is through week 15. Exclude internal accounts (is_internal=1).
- Cohort c = the week number of the signup time (0–11). Retention at week N = at least one session_start in week c+N. The denominator is the number of cohort signups. A cell with c+N > 15 is not yet observed, so it is
null. - The first argument of every function is the materials directory. Round rates to four decimal places.
Steps
- In
/root/founder/cohort/sizes.json, write the number of signups per cohort as{"0": n, ..., "11": n}. - In
/root/founder/cohort/cohort.py, createretention(fix, max_n)—{"0": [r0, r1, ..., r_max_n], ...}, withNonefor unobserved cells. - In
/root/founder/cohort/denominator.json, write the week-4 retention of cohort 4 with two denominators:rate_cohort_size(the correct denominator) andrate_week0_active(week-0 active users as the denominator — a mistake you commit on purpose). - In
/root/founder/cohort/censor.json, write the week-6 retention by two methods:pooled_rate(only cohorts whose observation is complete, weighted by number of people) andnaive_rate_with_zeros(all cohorts, with unobserved cells as 0), pluseligible_cohorts(the number of cohorts whose observation is complete). - In
/root/founder/cohort/survivor.json, write the week-4 retention with two rosters:all_users_rate(everyone in cohorts 0 to 11, only cohorts whose observation is complete),survivors_only_rate(the same calculation, picking only people active in week 15), andsurvivors(the number of those people). - In
cohort.py, addpooled(fix, n, unbounded=False)— the people-weighted retention over only cohorts whose observation is complete,{"cohorts": k, "users": u, "retained": r, "rate": 0.0000}. If unbounded is true, count "activity in any week after week c+N (including c+N)". - In
/root/founder/cohort/channel.json, write the week-4 weighted retention by signup channel (last touch) for the five channels as{"organic_search": r, ...}. - In
/root/founder/cohort/report.json, writecurve(the nine bounded weighted retentions for n=0..8),week4,week4_unbounded,flatten_week(the first n where the drop from the previous week is under 0.02, at least 1),best_channel, andworst_channel.
Notes
- Week number:
(서울 시각 - datetime(2026, 3, 2, tzinfo=서울)).days // 7(the Korean placeholders are the Seoul time and the Seoul time zone) - If you first build each person's "set of weeks with activity", every step is a single set check.
- Common mistakes: taking the denominator as week-0 active users, averaging with unobserved cells filled with 0, simply averaging the cohort rates, and looking at the past by picking only the people who remain now.
Signups per cohort
In /root/founder/cohort/sizes.json, write the number of external signups per signup week (0–11) with string keys.
Convert signup_at to Seoul time, and the quotient of the number of days since 2026-03-02 divided by 7 is the cohort. Exclude internal accounts.
The retention triangle — unobserved cells are None
In /root/founder/cohort/cohort.py, create retention(fix, max_n). Cohort string key → list of retentions for n=0..max_n, with None for cells where c+n>15.
First build, for each person, the set of weeks that had a session_start. The denominator is the number of cohort signups, and cells whose observation is not complete are None, not 0.
Changing the denominator makes retention look better
In /root/founder/cohort/denominator.json, write the week-4 retention of cohort 4 with the two denominators rate_cohort_size and rate_week0_active.
The numerator is the same (the people of cohort 4 active in week 8). Change only the denominator, to the number of signups and to the number of week-0 (week 4) active users.
Filling unobserved cells with 0 makes the curve falsely bend
In /root/founder/cohort/censor.json, write pooled_rate, naive_rate_with_zeros, and eligible_cohorts for week-6 retention.
The correct one collects only cohorts with c+6 ≤ 15 and divides total retained by total signups. The wrong one puts the signups of all cohorts in the denominator (people not yet observed are counted as retained 0).
Looking only at who remains makes the past look good
In /root/founder/cohort/survivor.json, write the week-4 weighted retention as all_users_rate (everyone) and survivors_only_rate (only those active in week 15), and survivors (the number of those people).
The two calculations share the cohort and week rules and differ only in the roster of people. The principle is to fix the roster at signup time, so the second number is a mistake you commit on purpose.
The weighted retention function — bounded and unbounded
Add pooled(fix, n, unbounded=False) to cohort.py. It returns {cohorts, users, retained, rate} over only cohorts whose observation is complete, weighted by number of people.
It is not a simple average of cohort rates but the sum of retained ÷ the sum of signups. unbounded checks whether there was activity in week c+n or in any week after it.
Week-4 retention by signup channel
In /root/founder/cohort/channel.json, write the week-4 weighted retention of each of the five last-touch channels.
In touches.csv, the channel of each person's line with the latest ts is the signup channel. Narrow only the roster to people of that channel and do the same calculation as pooled.
The curve and the judgment on one page
In /root/founder/cohort/report.json, write curve (n=0..8), week4, week4_unbounded, flatten_week, best_channel, and worst_channel.
flatten_week is the first n, looking from n=1, where curve[n-1] − curve[n] < 0.02. For the channels, use the highest and the lowest from the step 7 result.