TT Lab
开始
学习 学习路径 课程

开发者创业 — 先验证再动手

把访谈记录编码为证据并选出第一批客群

在 TT Lab 中继续学习

目标

从访谈记录中筛除提问形式带来的偏差,只统计“最近遇到过且付出了代价”的证据,并用考虑了样本量的区间来选出第一个客群。

为什么重要

人们会礼貌地回应别人的想法。如果把诱导和假设类问题的回答以及夸赞读成需求,就会在并不存在的市场上花几个月。把记录编码为证据,并连样本的大小也要考虑,这种习惯就是动手做之前的验证的全部。

材料——/opt/fixtures/founder/problem/interviews.csv

id,segment,question_style,said_problem,days_since_last,workaround,spent_krw_month,commitment

定义

步骤

  1. 在 /root/founder/problem/counts.json 中,为每个细分客群写入 all(访谈数)和 by_style(按 past、hypothetical、leading 分别的数量)。
  2. 在 /root/founder/problem/said.json 中,为每个细分客群写入 said_all_rate(所有访谈的 said_problem 比率)和 said_usable_rate(仅可用的访谈)。
  3. 在 /root/founder/problem/problem.py 中创建 evidence(row)(csv.DictReader 的一行 → True/False)。
  4. 在 /root/founder/problem/evidence.json 中,为每个细分客群写入 usable、evidence、rate(evidence ÷ usable)。
  5. 在 /root/founder/problem/signals.json 中,为每个细分客群写入 median_spend(可用、属于近期问题、且在花钱的人的月支出中位数,没有则为 0)、strong(可用访谈中 pilot、preorder 的数量)、compliments(所有访谈中 compliment 的数量)。
  6. 在 problem.py 中加入 wilson(k, n) → [low, high](n 为 0 时为 [0.0, 1.0])。
  7. 在 /root/founder/problem/ranking.json 中写入 by_said_all(said_all_rate 最高的细分客群)、by_point(证据比率最高的)、by_lower_bound(证据的 Wilson 下限最高的)、lower_bounds(细分客群 → 下限)。
  8. 在 /root/founder/problem/decision.json 中写入 target(by_lower_bound)、need_more(下限 < 0.5 ≤ 上限的细分客群,排序)、drop(上限 < 0.5 的细分客群,排序)。

参考

对谁问了什么样的问题

在 /root/founder/problem/counts.json 中,为每个细分客群写入 all 和 by_style(past、hypothetical、leading)。

用 csv.DictReader 读取,按 segment 和 question_style 统计。请看看各细分客群的提问形式所占比重是否不同。

提问形式会放大回答

在 /root/founder/problem/said.json 中,为每个细分客群写入 said_all_rate 和 said_usable_rate。

分子是 said_problem == "1" 的访谈,分母则是:前一个为该细分客群的全部,后一个为 question_style 是 past 的访谈。

把证据条件做成函数

在 /root/founder/problem/problem.py 中创建 evidence(row)(过去行为问题、最近 30 天内遇到过、花钱或者试点、预购)。

三个条件必须全部满足才是 True。空的 days_since_last 要在转成 int 之前先过滤掉。30 天是包含在内的。

各细分客群的证据

在 /root/founder/problem/evidence.json 中,为每个细分客群写入 usable、evidence、rate。

分母是可用的(past)访谈数。不要除以全部访谈数。

金钱、承诺,以及夸赞

在 /root/founder/problem/signals.json 中,为每个细分客群写入 median_spend、strong、compliments。

median_spend 的对象是可用、属于近期问题、并且支出大于 0 的人(没有则为 0)。compliments 与提问形式无关,全部统计。

用区间处理小样本

在 problem.py 中加入 wilson(k, n) → [low, high]。

z 为 NormalDist().inv_cdf(0.975)。把说明中的中心和半宽公式原样搬过来。n=0 时为 [0.0, 1.0]。

三种排名

在 /root/founder/problem/ranking.json 中写入 by_said_all、by_point、by_lower_bound、lower_bounds。

用同一份记录排三次序——所有回答中“有问题”的比率,证据比率(点估计),以及证据的 Wilson 下限。

第一个客群和下一批访谈

在 /root/founder/problem/decision.json 中写入 target、need_more、drop。

区间把 0.5 夹在中间的,是还不知道的(need_more),连上限也低于 0.5 的,是该放弃的(drop)。