把访谈记录编码为证据并选出第一批客群
目标
从访谈记录中筛除提问形式带来的偏差,只统计“最近遇到过且付出了代价”的证据,并用考虑了样本量的区间来选出第一个客群。
为什么重要
人们会礼貌地回应别人的想法。如果把诱导和假设类问题的回答以及夸赞读成需求,就会在并不存在的市场上花几个月。把记录编码为证据,并连样本的大小也要考虑,这种习惯就是动手做之前的验证的全部。
材料——/opt/fixtures/founder/problem/interviews.csv
id,segment,question_style,said_problem,days_since_last,workaround,spent_krw_month,commitment
- segment:freelance_designer、inhouse_marketer、small_agency
- question_style:past、hypothetical、leading
- said_problem:是否说有问题(0/1),days_since_last:上一次遇到距今多少天(没有说的话为空)
- spent_krw_month:现在每个月为这个问题花的钱,commitment:none、compliment、followup、pilot、preorder
定义
- 可用的访谈:question_style == "past"。
- 近期问题:said_problem == 1 并且 days_since_last ≤ 30。
- 代价:spent_krw_month > 0,或者 commitment 为 pilot、preorder。
- 证据:可用的访谈中,属于近期问题并且有代价的。
- Wilson 95% 区间:z =
NormalDist().inv_cdf(0.975),当 p = k/n 时,中心 = (p + z²/2n)/(1 + z²/n),半宽 = z·√(p(1−p)/n + z²/4n²)/(1 + z²/n)。 - 比率保留到小数点后第四位。
步骤
- 在
/root/founder/problem/counts.json中,为每个细分客群写入all(访谈数)和by_style(按 past、hypothetical、leading 分别的数量)。 - 在
/root/founder/problem/said.json中,为每个细分客群写入said_all_rate(所有访谈的 said_problem 比率)和said_usable_rate(仅可用的访谈)。 - 在
/root/founder/problem/problem.py中创建evidence(row)(csv.DictReader 的一行 → True/False)。 - 在
/root/founder/problem/evidence.json中,为每个细分客群写入usable、evidence、rate(evidence ÷ usable)。 - 在
/root/founder/problem/signals.json中,为每个细分客群写入median_spend(可用、属于近期问题、且在花钱的人的月支出中位数,没有则为 0)、strong(可用访谈中 pilot、preorder 的数量)、compliments(所有访谈中 compliment 的数量)。 - 在
problem.py中加入wilson(k, n)→[low, high](n 为 0 时为 [0.0, 1.0])。 - 在
/root/founder/problem/ranking.json中写入by_said_all(said_all_rate 最高的细分客群)、by_point(证据比率最高的)、by_lower_bound(证据的 Wilson 下限最高的)、lower_bounds(细分客群 → 下限)。 - 在
/root/founder/problem/decision.json中写入target(by_lower_bound)、need_more(下限 < 0.5 ≤ 上限的细分客群,排序)、drop(上限 < 0.5 的细分客群,排序)。
参考
csv.DictReader会把值全部作为字符串给出。空的 days_since_last 是""。- 常见错误:把诱导和假设类问题的回答放进证据,把夸赞算作承诺,把边界(30 天)按小于来处理,用点估计去选小样本。
对谁问了什么样的问题
在 /root/founder/problem/counts.json 中,为每个细分客群写入 all 和 by_style(past、hypothetical、leading)。
用 csv.DictReader 读取,按 segment 和 question_style 统计。请看看各细分客群的提问形式所占比重是否不同。
提问形式会放大回答
在 /root/founder/problem/said.json 中,为每个细分客群写入 said_all_rate 和 said_usable_rate。
分子是 said_problem == "1" 的访谈,分母则是:前一个为该细分客群的全部,后一个为 question_style 是 past 的访谈。
把证据条件做成函数
在 /root/founder/problem/problem.py 中创建 evidence(row)(过去行为问题、最近 30 天内遇到过、花钱或者试点、预购)。
三个条件必须全部满足才是 True。空的 days_since_last 要在转成 int 之前先过滤掉。30 天是包含在内的。
各细分客群的证据
在 /root/founder/problem/evidence.json 中,为每个细分客群写入 usable、evidence、rate。
分母是可用的(past)访谈数。不要除以全部访谈数。
金钱、承诺,以及夸赞
在 /root/founder/problem/signals.json 中,为每个细分客群写入 median_spend、strong、compliments。
median_spend 的对象是可用、属于近期问题、并且支出大于 0 的人(没有则为 0)。compliments 与提问形式无关,全部统计。
用区间处理小样本
在 problem.py 中加入 wilson(k, n) → [low, high]。
z 为 NormalDist().inv_cdf(0.975)。把说明中的中心和半宽公式原样搬过来。n=0 时为 [0.0, 1.0]。
三种排名
在 /root/founder/problem/ranking.json 中写入 by_said_all、by_point、by_lower_bound、lower_bounds。
用同一份记录排三次序——所有回答中“有问题”的比率,证据比率(点估计),以及证据的 Wilson 下限。
第一个客群和下一批访谈
在 /root/founder/problem/decision.json 中写入 target、need_more、drop。
区间把 0.5 夹在中间的,是还不知道的(need_more),连上限也低于 0.5 的,是该放弃的(drop)。