TT Lab
Get started
Learn Learning paths Courses

Ansible Fundamentals

It reported 37 hosts succeeded: designing for failure

Continue in TT Lab

Goal

In a three-host inventory where you have made just one machine fail, you turn on the handles for dealing with failure one by one and confirm directly with marker files, each time, what stops and what keeps going.

Why it matters

Failures happen. What you have to design is not getting rid of failures but deciding in advance what will happen when one occurs. Ansible's default is "only the failed host quietly drops out and the rest keep going", which is a good default for fixing servers one by one, independently, and a bad default if 40 machines together make up one service. So this lab takes up three questions in order — is this failure a real failure (failed_when), what to do with the failed host (ignore_errors, block/rescue), and what to do with the other hosts (any_errors_fatal, max_fail_percentage). Add to that assert, which blocks before starting, and unreachable, which is counted in a different column from failure, and you can read a deployment report. In the last step, you build that report yourself.

Steps

  1. In /root/ans/err/hosts.ini, write the web group (web1, web2), the db group (db1), and the ghosts group (ghost1). All four hosts have ansible_host=127.0.0.1 and ansible_user=root, and only ghost1 has the port 2223, while the rest have 2222. Then, with /root/ans/err/p01.yml, create a play that runs on web:db — the first task fails only on web2, and the second task leaves /root/ans/err/a1/reached-<호스트이름> (the placeholder stands for the host name). Save the run output to /root/ans/err/out/run1.txt.
  2. Create /root/ans/err/p02.yml with the same structure as step 1, but attach ignore_errors to the failing task and leave the markers at /root/ans/err/a2/reached-<호스트이름>. Save the output to /root/ans/err/out/run2.txt. This time the marker must remain for all three hosts and the PLAY RECAP must show the ignored column going up.
  3. Create /root/ans/err/sample.conf with the three lines listen_port=8080, env=prod, and workers=4. Then, with /root/ans/err/p03.yml, create a task on web1 that counts how many times nosuchkey appears in that file — since it is only a query, do not count it as a change, and set the criterion so that the exit code produced by not finding it is not a failure. Then leave that exit code as one line in /root/ans/err/a3/grep-rc.txt and save the output to /root/ans/err/out/run3.txt.
  4. With /root/ans/err/p04.yml, create a play that runs on web1. Put in a block a failing task and, after it, a task that creates /root/ans/err/a4/never.txt; have rescue put the name of the failed task into /root/ans/err/a4/rescued.txt, and have always leave /root/ans/err/a4/always.txt. Save the output to /root/ans/err/out/run4.txt. never.txt must not be created, and the PLAY RECAP must show rescued as 1.
  5. Make /root/ans/err/p05.yml a play that runs on localhost, but put in just one assert task. It checks whether deploy_env is one of dev, stage, or prod, and fail_msg must start with 허용되지 않은 배포 환경입니다 (this means "The deployment environment is not allowed") and show the current value together. Save the output of running with -e deploy_env=prod to /root/ans/err/out/assert-ok.txt and the output of running with -e deploy_env=qa to /root/ans/err/out/assert-fail.txt.
  6. Create /root/ans/err/p06.yml with the same structure as step 1, but turn on any_errors_fatal for the play and leave the markers at /root/ans/err/a6/reached-<호스트이름>. Save the output to /root/ans/err/out/run6.txt. This time no host may leave a marker.
  7. With /root/ans/err/p07.yml, create a play that runs on web:db. max_fail_percentage is 50, the default of the variable run_tag is continue, and the default of the variable doomed is just web2. The first task creates the directory /root/ans/err/a7/<run_tag>, the next task fails only on hosts contained in doomed, and the last task leaves /root/ans/err/a7/<run_tag>/reached-<호스트이름>. Run once with the defaults and save it to /root/ans/err/out/run7-continue.txt, and run once more with run_tag set to abort and, in doomed, web2 and db1 put, and save it to /root/ans/err/out/run7-abort.txt.
  8. Make /root/ans/err/p08.yml a play that runs on all — the first task is ping but ignores hosts that cannot be reached, and the second task leaves /root/ans/err/a8/reached-<호스트이름>. Save the output to /root/ans/err/out/run8.txt. Then write a script that takes one run log file as an argument and prints one line, hosts=N ok=N changed=N unreachable=N failed=N skipped=N rescued=N ignored=N, as /root/ans/err/recap.sh, and use it to make /root/ans/err/out/failure-report.json. There are six keys — default_failed (the sum of failed from the step 1 run), ignored (step 2), rescued (step 4), unreachable (step 8), fatal_reached (the number of markers in a6), and maxfail_reached (the number of markers in a7/continue) — and all the values are numbers.

Notes

The default behavior — only the failed host drops out

In /root/ans/err/hosts.ini, write the web group (web1, web2), the db group (db1), and the ghosts group (ghost1). All four hosts have ansible_host=127.0.0.1 and ansible_user=root, and only ghost1 has the port 2223, while the rest have 2222. Then, with /root/ans/err/p01.yml, create a play that runs on web:db — the first task fails only on web2, and the second task leaves /root/ans/err/a1/reached-<호스트이름> (the placeholder stands for the host name). Save the run output to /root/ans/err/out/run1.txt.

All the hosts connect to the same sshd, so you make "only this machine fails" not with a file but with an inventory_hostname condition. The port of ghost1 is a number nobody is listening on, so in step 8 it plays the part of a host that cannot be reached. The playbook ends with a non-zero value, so make sure the script does not stop there when you save the output.

What changes if you only ignore the result

Create /root/ans/err/p02.yml with the same structure as step 1, but attach ignore_errors to the failing task and leave the markers at /root/ans/err/a2/reached-<호스트이름> (the placeholder stands for the host name). Save the output to /root/ans/err/out/run2.txt. This time the marker must remain for all three hosts and the PLAY RECAP must show the ignored column going up.

ignore_errors does not change the verdict; it only ignores the result. So the task is still recorded as a failure, but the host does not drop out. On the PLAY RECAP, see for yourself which column goes up instead of failed.

Decide for yourself what counts as a failure

Create /root/ans/err/sample.conf with the three lines listen_port=8080, env=prod, and workers=4. Then, with /root/ans/err/p03.yml, create a task on web1 that counts how many times nosuchkey appears in that file — since it is only a query, do not count it as a change, and set the criterion so that the exit code produced by not finding it is not a failure. Then leave that exit code as one line in /root/ans/err/a3/grep-rc.txt and save the output to /root/ans/err/out/run3.txt.

If the word you look for is absent, 1 comes out, and if the file is missing or the arguments are wrong, 2 or more comes out. The former is an answer and the latter is an error. Write the criterion using the exit code of the result received with register. It is a different thing from covering it with ignore_errors — this playbook must end cleanly with failed=0.

Build a rollback as a structure with block, rescue, and always

With /root/ans/err/p04.yml, create a play that runs on web1. Put in a block a failing task and, after it, a task that creates /root/ans/err/a4/never.txt; have rescue put the name of the failed task into /root/ans/err/a4/rescued.txt, and have always leave /root/ans/err/a4/always.txt. Save the output to /root/ans/err/out/run4.txt. never.txt must not be created, and the PLAY RECAP must show rescued as 1.

The name of the failed task is held in a variable usable only inside rescue — the name contains failed and task. If rescue succeeds to the end, that host is treated as not having failed and turns into a green light. So leaving evidence in a file that a rollback ran is half of this structure.

If the preconditions are off, stop before touching anything

Make /root/ans/err/p05.yml a play that runs on localhost, but put in just one assert task. It checks whether deploy_env is one of dev, stage, or prod, and fail_msg must start with 허용되지 않은 배포 환경입니다 (this means "The deployment environment is not allowed") and show the current value together. Save the output of running with -e deploy_env=prod to /root/ans/err/out/assert-ok.txt and the output of running with -e deploy_env=qa to /root/ans/err/out/assert-fail.txt.

The default failure message is just one line, so when there are several conditions you cannot tell which one was false. If you embed the current value in the message, a single log line settles it. A command-line variable is stronger than the play's vars, so you can run the same playbook twice, changing only the value. The grader also runs this playbook directly in the same way.

Stop everything if even one machine fails

Create /root/ans/err/p06.yml with the same structure as step 1, but turn on any_errors_fatal for the play and leave the markers at /root/ans/err/a6/reached-<호스트이름> (the placeholder stands for the host name). Save the output to /root/ans/err/out/run6.txt. This time no host may leave a marker.

This handle attaches at the play level (not the task). Hosts that did not fail also do not run the remaining tasks, so the marker directory stays empty. On the PLAY RECAP, see what happens to the ok column of the hosts that did not fail. You must create the marker directory in advance to tell "empty" from "missing".

How many machines can you tolerate — deciding by ratio

With /root/ans/err/p07.yml, create a play that runs on web:db. max_fail_percentage is 50, the default of the variable run_tag is continue, and the default of the variable doomed is just web2. The first task creates the directory /root/ans/err/a7/<run_tag>, the next task fails only on hosts contained in doomed, and the last task leaves /root/ans/err/a7/<run_tag>/reached-<호스트이름> (the placeholder stands for the host name). Run once with the defaults and save it to /root/ans/err/out/run7-continue.txt, and run once more with run_tag set to abort and, in doomed, web2 and db1 put, and save it to /root/ans/err/out/run7-abort.txt.

Out of three machines, one is 33 percent and two are 66 percent. Whether the threshold is exceeded is the fork in the road. When you override a list variable on the command line, it is convenient to pass it whole as JSON. The task that creates the marker directory must come before the failing task so that an empty directory remains even in a stopped run.

Unreachable hosts and six runs on a single sheet

Make /root/ans/err/p08.yml a play that runs on all — the first task is ping but ignores hosts that cannot be reached, and the second task leaves /root/ans/err/a8/reached-<호스트이름> (the placeholder stands for the host name). Save the output to /root/ans/err/out/run8.txt. Then write a script that takes one run log file as an argument and prints one line, hosts=N ok=N changed=N unreachable=N failed=N skipped=N rescued=N ignored=N, as /root/ans/err/recap.sh, and use it to make /root/ans/err/out/failure-report.json. There are six keys — default_failed (the sum of failed from the step 1 run), ignored (step 2), rescued (step 4), unreachable (step 8), fatal_reached (the number of markers in a6), and maxfail_reached (the number of markers in a7/continue) — and all the values are numbers.

Not being reachable is not a task failure but a connection failure, so there is a separate argument to ignore it. That argument applies only to the task it is attached to, and even if you decided to ignore it, the playbook's exit code is not 0. When summing, you must read only below the PLAY RECAP line — similar text is mixed into the task output above it. The grader also runs this script against a log file it made itself, so a script that has the answer written into it will not pass.