TT Lab
开始
学习 学习路径 课程

GitLab CI/CD

一个慢 e2e 让文档发布白等了 8 秒

在 TT Lab 中继续学习

目标

实际运行一条混有慢作业的流水线,测量 stage 和 needs 如何改变作业的开始时刻,并通过运行结果确认产物传递、等待条件作业和允许失败的方式。

为什么重要

stage 方式容易理解,但最慢的作业会拖住所有后续作业的开始。用 needs 只等待必要的作业,流水线会变快,但代价是必须为每个作业准确写出等待什么、接收什么。如果等待可能被 rules 排除的作业,流水线会根本创建不出来,而如果过宽地允许失败,真正的事故就会以绿色通过。这些选择,比起读文档,看到时刻和结果时才会变得清楚。

步骤

  1. 把 /root/glci-dag 创建为 Git 仓库,并在 .gitlab-ci.yml 中放入 stages [build, test, deploy] 和四个作业:compile(build,sleep 3 之后创建 bin/app 文件,并用 artifacts 上传 bin/)、unit(test,test -f bin/app 之后 echo unit-ok)、e2e(test,sleep 8 之后 echo e2e-ok)、publish-docs(deploy,test -f bin/app 之后 echo docs-published)。用 gitlab-ci-local --shell-isolation --no-artifacts-to-source --timestamps 运行,看 publish-docs 要等到慢的 e2e 结束之后才开始。
  2. 给 publish-docs 加上 needs: [compile]。运行后,publish-docs 必须在 e2e 结束之前开始,并且仍要接收 compile 的产物(bin/app)。
  3. 用 needs: [] 添加作业 lint(stage test,sleep 1 之后 echo lint-ok)。运行后,lint 必须在 compile 结束之前开始。
  4. 把作业 audit(stage test)用 needs 的长格式写成 job: compile、artifacts: false,script 为 test ! -e bin/app && echo no-artifact。运行后,audit 必须成功(没有 bin/app),而 unit 仍然要接收 bin/app。
  5. 给作业 integration(stage test)设置 rules,使其只在 $RUN_INTEGRATION == "yes" 时才创建,并执行 sleep 2 之后 echo integration-ok。作业 release(stage deploy)在 needs 中放入 unit 和 job: integration, optional: true,并执行 echo release。不带变量运行时,没有 integration 的 release 会运行,而用 --variable RUN_INTEGRATION=yes 运行时,release 必须在 integration 结束之后才开始。
  6. 再添加三个作业。flaky(test)先 echo flaky-run 然后 exit 1,但有 allow_failure: true。cleanup(deploy)以 when: always 执行 echo cleanup,notify-failure(deploy)以 when: on_failure 执行 echo notify。运行后,flaky 必须以警告结束,流水线成功,cleanup 运行,而 notify-failure 不运行。评分器也会在副本中关掉 flaky 的 allow_failure 再运行一次。
  7. 把作业 check-config 添加到 stage .pre(不写进 stages 列表)。sleep 2 之后用 test ! -e STOP,如果仓库中有 STOP 文件就失败,没有就执行 echo config-ok。运行后,compile 必须在 check-config 结束之后才开始。评分器会在副本中放入 STOP 文件,确认预检失败时 compile 是否根本不运行。

参考

stage 要等待前面所有 stage

把 /root/glci-dag 创建为 Git 仓库,并在 .gitlab-ci.yml 中放入 stages [build, test, deploy] 和四个作业:compile(build,sleep 3 之后创建 bin/app 文件,并用 artifacts 上传 bin/)、unit(test,test -f bin/app 之后 echo unit-ok)、e2e(test,sleep 8 之后 echo e2e-ok)、publish-docs(deploy,test -f bin/app 之后 echo docs-published)。用 gitlab-ci-local --shell-isolation --no-artifacts-to-source --timestamps 运行,看 publish-docs 要等到慢的 e2e 结束之后才开始。

没有 needs 的作业,要等前面 stage 的所有作业结束才开始,并接收前面所有 stage 的产物。加上 --timestamps,每一行前面都会打印时刻,可以比较开始(starting shell)和结束(finished in)。

发布文档只需要等待编译

给 publish-docs 加上 needs: [compile]。运行后,publish-docs 必须在 e2e 结束之前开始,并且仍要接收 compile 的产物(bin/app)。

needs 直接指定要等待的作业。只接收所写作业的产物,stage 顺序不再决定开始时刻。一旦使用了 needs,没有写出的作业既不会被等待,也不会接收其产物。

needs: [] 在流水线一开始就出发

用 needs: [] 添加作业 lint(stage test,sleep 1 之后 echo lint-ok)。运行后,lint 必须在 compile 结束之前开始。

空的 needs 就是“不等任何人”。即使 stage 是 test,也不会等待 build 结束。把只需要看源码的检查这样提前,就能早几分钟知道失败。

等待顺序,但不接收产物

把作业 audit(stage test)用 needs 的长格式写成 job: compile、artifacts: false,script 为 test ! -e bin/app && echo no-artifact。运行后,audit 必须成功(没有 bin/app),而 unit 仍然要接收 bin/app。

needs 的长格式可以按作业关闭是否接收产物。这样可以避免只需要顺序、不需要文件的作业,因为下载大产物而变慢。

等待可能有也可能没有的作业

给作业 integration(stage test)设置 rules,使其只在 $RUN_INTEGRATION == "yes" 时才创建,并执行 sleep 2 之后 echo integration-ok。作业 release(stage deploy)在 needs 中放入 unit 和 job: integration, optional: true,并执行 echo release。不带变量运行时,没有 integration 的 release 会运行,而用 --variable RUN_INTEGRATION=yes 运行时,release 必须在 integration 结束之后才开始。

如果把可能被 rules 排除的作业直接写进 needs,当那个作业不存在时,GitLab 根本不会创建流水线。optional: true 的意思是“有就等,没有就过去”。

允许失败的作业、失败也会运行的作业、失败才会运行的作业

再添加三个作业。flaky(test)先 echo flaky-run 然后 exit 1,但有 allow_failure: true。cleanup(deploy)以 when: always 执行 echo cleanup,notify-failure(deploy)以 when: on_failure 执行 echo notify。运行后,flaky 必须以警告结束,流水线成功,cleanup 运行,而 notify-failure 不运行。评分器也会在副本中关掉 flaky 的 allow_failure 再运行一次。

allow_failure 会把失败变成“警告”,使其不阻挡后面的 stage。when: on_failure 只在前面有作业失败时才运行,always 则不管结果如何都会运行。被允许的失败不会触发 on_failure。

先于所有 stage 运行的预检

把作业 check-config 添加到 stage .pre(不写进 stages 列表)。sleep 2 之后用 test ! -e STOP,如果仓库中有 STOP 文件就失败,没有就执行 echo config-ok。运行后,compile 必须在 check-config 结束之后才开始。评分器会在副本中放入 STOP 文件,确认预检失败时 compile 是否根本不运行。

.pre 是即使不写进 stages 也始终在最前面的保留 stage(最后面是 .post)。用于在不改动 stage 列表的情况下,为整条流水线加上预检。如果前面的 stage 失败,后面 stage 的普通作业只会被创建而不会运行。