TT Lab
开始
学习 学习路径 课程

CRD 与 Operator

日志里有,用户却看不到 - 让 Operator 开口说话

在 TT Lab 中继续学习

目标

用两个事件 API 亲手创建事件,确认必填字段和汇总,看看违反惯例时在工具中会怎样消失,然后制作一张只读取事件就能重建事故的表格。

为什么重要

用户能得知 Operator 做了什么的地方只有三处——控制器日志、资源的 status,以及事件。日志只有集群运维人员会看,status 只说明当前状态,不包含过程。用户实际会读的地方,只有 kubectl describe 下方的 Events。所以开发 Operator 的工作有一半,就是决定把什么作为事件留下来。不过事件不是日志。默认保留一小时,相同原因的重复不是新对象,而是合并为 count 的增加,而且 API 有两个,存储却只有一个。如果不了解这些性质,期待事件提供审计追踪,真正调查事故时就什么都不剩了。

步骤

  1. 在 /root/op-events/pipeline-crd.yaml 中编写 CRD pipelines.ev.labhub.io——group 为 ev.labhub.io,kind 为 Pipeline,复数形式为 pipelines,只有一个版本 v1,schema 中只有 spec.stage(string)。创建命名空间 op-events,通过 /root/op-events/pipeline-build1.yaml 应用 Pipeline build-1(stage 为 build),然后把它的 metadata.uid 用一行保存到 /root/op-events/pipeline-uid.txt。
  2. 在 /root/op-events/event-started.yaml 中编写 v1 Event build-1-started——involvedObject 是 Pipeline build-1(apiVersion、kind、name、namespace、实际 uid),reason 为 ReconcileStarted,type 为 Normal,message 为一句话,firstTimestamp 和 lastTimestamp 为当前时间,count 为 1,source.component 为 pipeline-operator。应用之后,把 kubectl -n op-events events 的输出保存到 /root/op-events/events-list.txt。
  3. 在 /root/op-events/event-bad.yaml 中编写 Event build-1-orphan,但完全不写 involvedObject(reason 为 NoTarget,type 为 Normal,message 为一句话)。尝试应用,并把结果汇总到 /root/op-events/invalid-event.txt——第一行是 apply-rc=<종료 코드>(占位符为退出码),下面原样贴上服务器给出的语句。
  4. 在 /root/op-events/event-weird.yaml 中编写 Event build-1-weird——目标是 Pipeline build-1,reason 为 StageUnknown,并把 type 写成超出惯例的值 Critical(时间字段以与第 2 步相同的格式填写)。应用之后,依次执行 kubectl -n op-events events --types=Critical 和 kubectl -n op-events events --types=Warning,并把结果汇总到 /root/op-events/type-report.txt——必须包含两条命令的输出和各自的退出码(critical-rc=、warning-rc=)。
  5. 假设第 2 步的事件又发生了四次,更新汇总字段——用 kubectl -n op-events patch event build-1-started --type=merge 把 count 改为 5,把 lastTimestamp 改为当前时间。然后把 kubectl -n op-events events 的输出保存到 /root/op-events/count-report.txt。
  6. 在 /root/op-events/event-done.yaml 中编写 events.k8s.io/v1 Event build-1-done——目标是 regarding(Pipeline build-1,包含实际 uid),reason 为 ReconcileSucceeded,note 为一句话,type 为 Normal,eventTime 为当前时间(小数点后 6 位),reportingController 为 pipeline-operator,reportingInstance 为 pipeline-operator-0,action 为 Reconcile。应用之后,用两个 API 各取出一份列表,保存到 /root/op-events/both-apis.txt——以 core: 开头的各行和以 new: 开头的各行,必须分别是 kubectl get events 和 kubectl get events.events.k8s.io 的名称列表。
  7. 在 /root/op-events/hungry-pod.yaml 中编写 Pod hungry——一个容器(app,镜像 busybox:1.36),要求 requests.cpu 为 "64"。应用之后,等待调度器留下事件,然后把 kubectl -n op-events get events --field-selector reason=FailedScheduling -o wide 的输出保存到 /root/op-events/scheduler-event.txt。
  8. 创建 /root/op-events/timeline.sh——把 op-events 中的所有事件,每条以 <type> <reason> <대상종류>/<대상이름> 一行的形式打印(占位符依次为类型、原因、目标类型、目标名称),并排序、去重。不要放入存在时长、时刻、count 这类每次查看都会变化的值。把输出保存到 /root/op-events/timeline.txt,然后仅凭这张表,在 /root/op-events/incident.txt 中用人能读懂的语句写明发生了什么——必须同时提到调度器留下的原因和 Operator 留下的原因,还要指出其中混有超出惯例的 type。

参考

创建用来挂载事件的对象

在 /root/op-events/pipeline-crd.yaml 中编写 CRD pipelines.ev.labhub.io——group 为 ev.labhub.io,kind 为 Pipeline,复数形式为 pipelines,只有一个版本 v1,schema 中只有 spec.stage(string)。创建命名空间 op-events,通过 /root/op-events/pipeline-build1.yaml 应用 Pipeline build-1(stage 为 build),然后把它的 metadata.uid 用一行保存到 /root/op-events/pipeline-uid.txt。

事件总是“关于某个对象的故事”。所以不仅要写明目标的类型、名称、命名空间,还要写明 uid,才不会与以相同名称重新创建的另一个对象的故事混在一起。uid 在下一步中原样使用。

Operator 留下第一句话

在 /root/op-events/event-started.yaml 中编写 v1 Event build-1-started——involvedObject 是 Pipeline build-1(apiVersion、kind、name、namespace、实际 uid),reason 为 ReconcileStarted,type 为 Normal,message 为一句话,firstTimestamp 和 lastTimestamp 为当前时间,count 为 1,source.component 为 pipeline-operator。应用之后,把 kubectl -n op-events events 的输出保存到 /root/op-events/events-list.txt。

reason 不是给人读的句子,而是机器用来统计的键。所以惯例是用简短的 CamelCase 单词,相同的原因重复出现时,会被合并为一个,而不是新事件。详细说明请放在 message 中。即使目标是自定义资源,它也会照样显示在 kubectl describe 的下方。

没有目标的事件无法创建

在 /root/op-events/event-bad.yaml 中编写 Event build-1-orphan,但完全不写 involvedObject(reason 为 NoTarget,type 为 Normal,message 为一句话)。尝试应用,并把结果汇总到 /root/op-events/invalid-event.txt——第一行是 apply-rc=<종료 코드>(占位符为退出码),下面原样贴上服务器给出的语句。

事件只有有了目标才有意义。没有目标,它就无法挂到任何对象的 describe 中,只会在列表里漂着一个原因。所以 API 会直接拒绝,而读一读错误到底指出了哪个字段,就能同时知道事件与目标的命名空间关系。

超出惯例的 type 会在工具中消失

在 /root/op-events/event-weird.yaml 中编写 Event build-1-weird——目标是 Pipeline build-1,reason 为 StageUnknown,并把 type 写成超出惯例的值 Critical(时间字段以与第 2 步相同的格式填写)。应用之后,依次执行 kubectl -n op-events events --types=Critical 和 kubectl -n op-events events --types=Warning,并把结果汇总到 /root/op-events/type-report.txt——必须包含两条命令的输出和各自的退出码(critical-rc=、warning-rc=)。

API 服务器不检查 type 的值。所以对象会照常创建。问题出在之后——读取事件的工具,在设计时就假定只有 Normal 和 Warning 两种。为什么没有强制,却仍应遵守惯例,原因就在这里。

相同原因的重复不是新事件

假设第 2 步的事件又发生了四次,更新汇总字段——用 kubectl -n op-events patch event build-1-started --type=merge 把 count 改为 5,把 lastTimestamp 改为当前时间。然后把 kubectl -n op-events events 的输出保存到 /root/op-events/count-report.txt。

事件记录器再次遇到相同的对象、相同的原因、相同的消息时,不会创建新对象,只会修改这两个字段。请亲自看看列表界面中的 LAST SEEN 一栏此时会怎样变化——括号里的写法就是这一步的答案。这也是为什么调谐循环每秒运行好几次,etcd 也不会被撑爆。

API 有两个,存储只有一个

在 /root/op-events/event-done.yaml 中编写 events.k8s.io/v1 Event build-1-done——目标是 regarding(Pipeline build-1,包含实际 uid),reason 为 ReconcileSucceeded,note 为一句话,type 为 Normal,eventTime 为当前时间(小数点后 6 位),reportingController 为 pipeline-operator,reportingInstance 为 pipeline-operator-0,action 为 Reconcile。应用之后,用两个 API 各取出一份列表,保存到 /root/op-events/both-apis.txt——以 core: 开头的各行和以 new: 开头的各行,必须分别是 kubectl get events 和 kubectl get events.events.k8s.io 的名称列表。

新 API 的字段名称不同——involvedObject 变成了 regarding,message 变成了 note,source 拆成了 reportingController 和 reportingInstance。但存储的位置相同,所以用旧 API 也能原样看到。用旧 API 查看时,确认哪些栏位是空的,两种 schema 的差别就会映入眼帘。

接收真实的控制器留下的事件

在 /root/op-events/hungry-pod.yaml 中编写 Pod hungry——一个容器(app,镜像 busybox:1.36),要求 requests.cpu 为 "64"。应用之后,等待调度器留下事件,然后把 kubectl -n op-events get events --field-selector reason=FailedScheduling -o wide 的输出保存到 /root/op-events/scheduler-event.txt。

这个 Pod 无法放入任何节点。调度器不仅会把这个事实写在 Pod 的 status 中,还会作为事件留下,消息中原样写着“几个里有几个、因为什么不行”。人在查找原因时真正会读的,就是这句话。事件挂上去需要几秒钟,请用条件循环来等待。

只读取事件来重建事故

创建 /root/op-events/timeline.sh——把 op-events 中的所有事件,每条以 <type> <reason> <대상종류>/<대상이름> 一行的形式打印(占位符依次为类型、原因、目标类型、目标名称),并排序、去重。不要放入存在时长、时刻、count 这类每次查看都会变化的值。把输出保存到 /root/op-events/timeline.txt,然后仅凭这张表,在 /root/op-events/incident.txt 中用人能读懂的语句写明发生了什么——必须同时提到调度器留下的原因和 Operator 留下的原因,还要指出其中混有超出惯例的 type。

在事故调查中,事件之所以有价值,是因为“谁针对什么做了什么判断”一条一条地留了下来。但默认保留时间是一小时,所以去得晚了就什么都没有。因此需要长期保存的内容,应该是 status 中的条件,或者发送到外部存储的记录,而不是事件。