It reported 37 hosts succeeded: designing for failure
Goal
In a three-host inventory where you have made just one machine fail, you turn on the handles for dealing with failure one by one and confirm directly with marker files, each time, what stops and what keeps going.
Why it matters
Failures happen. What you have to design is not getting rid of failures but deciding in advance what will happen when one occurs. Ansible's default is "only the failed host quietly drops out and the rest keep going", which is a good default for fixing servers one by one, independently, and a bad default if 40 machines together make up one service. So this lab takes up three questions in order — is this failure a real failure (failed_when), what to do with the failed host (ignore_errors, block/rescue), and what to do with the other hosts (any_errors_fatal, max_fail_percentage). Add to that assert, which blocks before starting, and unreachable, which is counted in a different column from failure, and you can read a deployment report. In the last step, you build that report yourself.
Steps
- In
/root/ans/err/hosts.ini, write thewebgroup (web1,web2), thedbgroup (db1), and theghostsgroup (ghost1). All four hosts haveansible_host=127.0.0.1andansible_user=root, and onlyghost1has the port2223, while the rest have2222. Then, with/root/ans/err/p01.yml, create a play that runs onweb:db— the first task fails only onweb2, and the second task leaves/root/ans/err/a1/reached-<호스트이름>(the placeholder stands for the host name). Save the run output to/root/ans/err/out/run1.txt. - Create
/root/ans/err/p02.ymlwith the same structure as step 1, but attachignore_errorsto the failing task and leave the markers at/root/ans/err/a2/reached-<호스트이름>. Save the output to/root/ans/err/out/run2.txt. This time the marker must remain for all three hosts and thePLAY RECAPmust show theignoredcolumn going up. - Create
/root/ans/err/sample.confwith the three lineslisten_port=8080,env=prod, andworkers=4. Then, with/root/ans/err/p03.yml, create a task onweb1that counts how many timesnosuchkeyappears in that file — since it is only a query, do not count it as a change, and set the criterion so that the exit code produced by not finding it is not a failure. Then leave that exit code as one line in/root/ans/err/a3/grep-rc.txtand save the output to/root/ans/err/out/run3.txt. - With
/root/ans/err/p04.yml, create a play that runs onweb1. Put in ablocka failing task and, after it, a task that creates/root/ans/err/a4/never.txt; haverescueput the name of the failed task into/root/ans/err/a4/rescued.txt, and havealwaysleave/root/ans/err/a4/always.txt. Save the output to/root/ans/err/out/run4.txt.never.txtmust not be created, and thePLAY RECAPmust showrescuedas 1. - Make
/root/ans/err/p05.ymla play that runs onlocalhost, but put in just oneasserttask. It checks whetherdeploy_envis one ofdev,stage, orprod, andfail_msgmust start with허용되지 않은 배포 환경입니다(this means "The deployment environment is not allowed") and show the current value together. Save the output of running with-e deploy_env=prodto/root/ans/err/out/assert-ok.txtand the output of running with-e deploy_env=qato/root/ans/err/out/assert-fail.txt. - Create
/root/ans/err/p06.ymlwith the same structure as step 1, but turn onany_errors_fatalfor the play and leave the markers at/root/ans/err/a6/reached-<호스트이름>. Save the output to/root/ans/err/out/run6.txt. This time no host may leave a marker. - With
/root/ans/err/p07.yml, create a play that runs onweb:db.max_fail_percentageis50, the default of the variablerun_tagiscontinue, and the default of the variabledoomedis justweb2. The first task creates the directory/root/ans/err/a7/<run_tag>, the next task fails only on hosts contained indoomed, and the last task leaves/root/ans/err/a7/<run_tag>/reached-<호스트이름>. Run once with the defaults and save it to/root/ans/err/out/run7-continue.txt, and run once more withrun_tagset toabortand, indoomed,web2anddb1put, and save it to/root/ans/err/out/run7-abort.txt. - Make
/root/ans/err/p08.ymla play that runs onall— the first task ispingbut ignores hosts that cannot be reached, and the second task leaves/root/ans/err/a8/reached-<호스트이름>. Save the output to/root/ans/err/out/run8.txt. Then write a script that takes one run log file as an argument and prints one line,hosts=N ok=N changed=N unreachable=N failed=N skipped=N rescued=N ignored=N, as/root/ans/err/recap.sh, and use it to make/root/ans/err/out/failure-report.json. There are six keys —default_failed(the sum of failed from the step 1 run),ignored(step 2),rescued(step 4),unreachable(step 8),fatal_reached(the number of markers ina6), andmaxfail_reached(the number of markers ina7/continue) — and all the values are numbers.
Notes
- The working directory is
/root/ans/err. The inventory's hosts connect to the sshd at 127.0.0.1:2222 inside this Pod, and onlyghost1looks at 2223, where nobody is listening. - A failing playbook ends with a non-zero value. When you save the output to a file, make sure the shell does not stop there.
- The
PLAY RECAPhas seven columns — ok, changed, unreachable, failed, skipped, rescued, and ignored. - Common mistake: expecting
ignore_errorsto work on hosts that cannot be reached. That is not a task failure but a connection failure. - Common mistake: running without creating the marker directory, so you cannot tell empty from nonexistent.
- Common mistake: seeing a run that
rescuesaved only as a green light and leaving nowhere the fact that a rollback ran. - Error handling in playbooks · Blocks · assert module · fail module · Strategies and forks
The default behavior — only the failed host drops out
In /root/ans/err/hosts.ini, write the web group (web1, web2), the db group (db1), and the ghosts group (ghost1). All four hosts have ansible_host=127.0.0.1 and ansible_user=root, and only ghost1 has the port 2223, while the rest have 2222. Then, with /root/ans/err/p01.yml, create a play that runs on web:db — the first task fails only on web2, and the second task leaves /root/ans/err/a1/reached-<호스트이름> (the placeholder stands for the host name). Save the run output to /root/ans/err/out/run1.txt.
All the hosts connect to the same sshd, so you make "only this machine fails" not with a file but with an inventory_hostname condition. The port of ghost1 is a number nobody is listening on, so in step 8 it plays the part of a host that cannot be reached. The playbook ends with a non-zero value, so make sure the script does not stop there when you save the output.
What changes if you only ignore the result
Create /root/ans/err/p02.yml with the same structure as step 1, but attach ignore_errors to the failing task and leave the markers at /root/ans/err/a2/reached-<호스트이름> (the placeholder stands for the host name). Save the output to /root/ans/err/out/run2.txt. This time the marker must remain for all three hosts and the PLAY RECAP must show the ignored column going up.
ignore_errors does not change the verdict; it only ignores the result. So the task is still recorded as a failure, but the host does not drop out. On the PLAY RECAP, see for yourself which column goes up instead of failed.
Decide for yourself what counts as a failure
Create /root/ans/err/sample.conf with the three lines listen_port=8080, env=prod, and workers=4. Then, with /root/ans/err/p03.yml, create a task on web1 that counts how many times nosuchkey appears in that file — since it is only a query, do not count it as a change, and set the criterion so that the exit code produced by not finding it is not a failure. Then leave that exit code as one line in /root/ans/err/a3/grep-rc.txt and save the output to /root/ans/err/out/run3.txt.
If the word you look for is absent, 1 comes out, and if the file is missing or the arguments are wrong, 2 or more comes out. The former is an answer and the latter is an error. Write the criterion using the exit code of the result received with register. It is a different thing from covering it with ignore_errors — this playbook must end cleanly with failed=0.
Build a rollback as a structure with block, rescue, and always
With /root/ans/err/p04.yml, create a play that runs on web1. Put in a block a failing task and, after it, a task that creates /root/ans/err/a4/never.txt; have rescue put the name of the failed task into /root/ans/err/a4/rescued.txt, and have always leave /root/ans/err/a4/always.txt. Save the output to /root/ans/err/out/run4.txt. never.txt must not be created, and the PLAY RECAP must show rescued as 1.
The name of the failed task is held in a variable usable only inside rescue — the name contains failed and task. If rescue succeeds to the end, that host is treated as not having failed and turns into a green light. So leaving evidence in a file that a rollback ran is half of this structure.
If the preconditions are off, stop before touching anything
Make /root/ans/err/p05.yml a play that runs on localhost, but put in just one assert task. It checks whether deploy_env is one of dev, stage, or prod, and fail_msg must start with 허용되지 않은 배포 환경입니다 (this means "The deployment environment is not allowed") and show the current value together. Save the output of running with -e deploy_env=prod to /root/ans/err/out/assert-ok.txt and the output of running with -e deploy_env=qa to /root/ans/err/out/assert-fail.txt.
The default failure message is just one line, so when there are several conditions you cannot tell which one was false. If you embed the current value in the message, a single log line settles it. A command-line variable is stronger than the play's vars, so you can run the same playbook twice, changing only the value. The grader also runs this playbook directly in the same way.
Stop everything if even one machine fails
Create /root/ans/err/p06.yml with the same structure as step 1, but turn on any_errors_fatal for the play and leave the markers at /root/ans/err/a6/reached-<호스트이름> (the placeholder stands for the host name). Save the output to /root/ans/err/out/run6.txt. This time no host may leave a marker.
This handle attaches at the play level (not the task). Hosts that did not fail also do not run the remaining tasks, so the marker directory stays empty. On the PLAY RECAP, see what happens to the ok column of the hosts that did not fail. You must create the marker directory in advance to tell "empty" from "missing".
How many machines can you tolerate — deciding by ratio
With /root/ans/err/p07.yml, create a play that runs on web:db. max_fail_percentage is 50, the default of the variable run_tag is continue, and the default of the variable doomed is just web2. The first task creates the directory /root/ans/err/a7/<run_tag>, the next task fails only on hosts contained in doomed, and the last task leaves /root/ans/err/a7/<run_tag>/reached-<호스트이름> (the placeholder stands for the host name). Run once with the defaults and save it to /root/ans/err/out/run7-continue.txt, and run once more with run_tag set to abort and, in doomed, web2 and db1 put, and save it to /root/ans/err/out/run7-abort.txt.
Out of three machines, one is 33 percent and two are 66 percent. Whether the threshold is exceeded is the fork in the road. When you override a list variable on the command line, it is convenient to pass it whole as JSON. The task that creates the marker directory must come before the failing task so that an empty directory remains even in a stopped run.
Unreachable hosts and six runs on a single sheet
Make /root/ans/err/p08.yml a play that runs on all — the first task is ping but ignores hosts that cannot be reached, and the second task leaves /root/ans/err/a8/reached-<호스트이름> (the placeholder stands for the host name). Save the output to /root/ans/err/out/run8.txt. Then write a script that takes one run log file as an argument and prints one line, hosts=N ok=N changed=N unreachable=N failed=N skipped=N rescued=N ignored=N, as /root/ans/err/recap.sh, and use it to make /root/ans/err/out/failure-report.json. There are six keys — default_failed (the sum of failed from the step 1 run), ignored (step 2), rescued (step 4), unreachable (step 8), fatal_reached (the number of markers in a6), and maxfail_reached (the number of markers in a7/continue) — and all the values are numbers.
Not being reachable is not a task failure but a connection failure, so there is a separate argument to ignore it. That argument applies only to the task it is attached to, and even if you decided to ignore it, the playbook's exit code is not 0. When summing, you must read only below the PLAY RECAP line — similar text is mixed into the task output above it. The grader also runs this script against a log file it made itself, so a script that has the answer written into it will not pass.