TT Lab
Get started
Learn Learning paths Courses

Ansible Fundamentals

When something fails, what stops and what keeps going

Continue in TT Lab

In one sentence

Ansible's default is "only the failed host quietly drops out and the rest keep going", and designing for failure means asking whether that default fits our situation and changing only the places where it does not, using handles.

Why this was needed

There was a playbook that deployed to 40 servers. One day the disk filled up on three of them and writing the configuration file failed. The deployment printed a few lines in red and kept running, and the last line said 37 servers succeeded. The pipeline ended in failure, but nobody knew which machine was in what state. On three machines, the new code was looking at the old configuration, and in that state they stayed attached to the load balancer as they were.

What went wrong here is not "it failed". It is that nobody had decided what should happen when it fails. Ansible only picked one default, and that default is "one host's failure does not block other hosts' work". It is a good default for fixing servers one by one, independently, but it is a bad default if 40 machines together make up one service.

So designing for failure is not memorizing exception-handling syntax but settling the answers to three questions. Is this failure a real failure? What will you do with the host that failed? What will you do with the other hosts?

How it works

The default behavior — only the failed host drops out

When a task fails on a host, Ansible excludes that host from the targets of this play. The remaining tasks are not executed on that host at all. The other hosts are not affected and go to the end. The exit code of the whole playbook is not 0, but that is a matter for after the run is over.

As a result of this behavior, a partially applied state remains. A host that failed at task 5 stays stopped with tasks 1 through 4 applied. If the playbook is idempotent, you can fix the cause and run it again, but where the middle is dangerous, as in a state where "the configuration was changed but the restart could not be done", this default does not fit.

Is this failure a real failure — failed_when and ignore_errors

command and shell treat a non-zero exit code as a failure. But there are many commands for which a non-zero exit code is normal. grep gives 1 when it finds nothing. diff gives 1 when files differ. systemctl is-active gives 3 when the unit is off. These values are not errors but answers.

- name: 설정에 그 키가 있는지 센다
  ansible.builtin.command: grep -c max_conn /etc/app/app.conf
  register: keycount
  changed_when: false
  failed_when: keycount.rc > 1

failed_when redefines "what counts as a failure" task by task. grep gives 1 when it finds nothing and 2 or more when the file is missing or the arguments are wrong, so the condition above means "not finding it is an answer, and only a real error is a failure". The reason to also put changed_when: false is to keep a query task from being reported as changed every time.

ignore_errors: true is a completely different thing. It does not change the verdict; it only ignores the result. The task is still recorded as a failure and the PLAY RECAP column ignored goes up, but the host does not drop out and goes on to the next task.

The task's verdict Next task RECAP
Default Failure That host drops out failed
failed_when As I decided Continues if not a failure Depends on the condition
ignore_errors Failure as it is Continues ignored

If you confuse the two, a quiet incident happens. Use ignore_errors only when it is "yes this is a failure, but let's move on for now". If it is really not a failure, the right thing is to correct the criterion with failed_when. A playbook that attaches ignore_errors out of habit only swallows failures, so half-applied servers remain wearing a green light. Also, ignore_errors does not apply to hosts that cannot be reached — that is not a task failure but a connection failure, so there is a separate ignore_unreachable.

What to do with the failed host — block, rescue, always

block groups several tasks into one. If you attach rescue and always to that group, it takes the same shape as try, except, and finally in other languages.

- name: 새 판으로 바꿔 보기
  block:
    - name: 새 설정 배치
      ansible.builtin.copy: {dest: /etc/app/app.conf, src: new.conf}
    - name: 헬스 체크
      ansible.builtin.command: /usr/local/bin/healthcheck
  rescue:
    - name: 옛 설정으로 되돌리기
      ansible.builtin.copy: {dest: /etc/app/app.conf, src: old.conf, remote_src: true}
  always:
    - name: 무슨 일이 있었든 기록을 남긴다
      ansible.builtin.lineinfile: {path: /var/log/deploy.log, line: "attempt finished"}

There are a few rules. When a failure occurs inside a block, the block is stopped right there and control moves to rescue. The block tasks after the failure are not executed. If rescue succeeds to the end, that host is treated as not having failed and continues with the play. On the PLAY RECAP, it is not failed but rescued that appears. always runs every time, whether the block succeeded, failed, or was rescued.

Inside rescue, you can see what failed and why with ansible_failed_task and ansible_failed_result. Use them when leaving a log or sending a notification. And if it fails again inside rescue, that is a real failure — piling up structures layer upon layer is usually a sign that the design is wrong.

One thing to watch out for. If rescue succeeds, the pipeline turns green. That means the fact that a rollback ran does not remain in the exit code. So inside rescue you must put a task that leaves evidence that "a rollback happened". Whether a file, a notification, or a metric, it must be in a form a person can count later.

What to do with the other hosts — any_errors_fatal and max_fail_percentage

any_errors_fatal: true stops the whole play right there if even one host fails. Hosts that did not fail also do not run the remaining tasks. Use it for work that must be "all or nothing", such as cluster setup, or for places like a pre-check in front of a deployment where you must not start if even one machine does not meet the condition.

max_fail_percentage is the value in between. If the ratio of failed hosts exceeds this value, the rest stop too.

Out of 3 hosts max_fail_percentage 50 Result
1 failed 33% is 50 or below The other 2 keep going
2 failed 66% exceeds 50 The play stops right there

What matters is that this ratio is judged per batch. In a rolling deployment split with serial, it is calculated each time a batch ends, so it becomes a safety device: "if more than half fail in one batch, do not touch the remaining batches". max_fail_percentage: 0 has almost the same meaning as any_errors_fatal: true, because it stops as soon as the ratio exceeds 0.

Unreachable hosts — unreachable is not failed

When SSH itself does not work, it is counted not as failed but as unreachable. This is because the task did not fail; it could not even start. This distinction shows up as different columns in the PLAY RECAP, and it is not ignored by ignore_errors. If you attach ignore_unreachable: true to that task, the host does not drop out and moves on to the next task.

The reason this distinction matters in practice is that the causes are entirely different. failed is usually a problem in our code or the target's state, and unreachable is usually a network, power, or inventory-typo problem. If you add the two numbers together when you look at a deployment report, you can no longer tell what to fix that day.

Block it before starting — assert

The cheapest failure is a failure that happens before anything is changed. assert takes a list of conditions and fails the task if even one is false.

- name: 배포 전제 확인
  ansible.builtin.assert:
    that:
      - deploy_env in ["dev", "stage", "prod"]
      - app_version is match("^[0-9]+\.[0-9]+\.[0-9]+$")
    fail_msg: "배포 전제가 어긋났습니다 (env={{ deploy_env }}, version={{ app_version }})"
    success_msg: "전제 조건 통과"

There is a reason to recommend always writing fail_msg. The default message is a single line, Assertion failed, so when there are several conditions, you cannot tell which one was false. If you embed the current values in the message, a single log line settles it. And if you put this task at the very front of the play together with any_errors_fatal, then when even one machine's preconditions are off, it stops without touching anything. The fail module is a cousin used together with when, for when you want to write the condition yourself.

How to read the PLAY RECAP

web1 : ok=6  changed=2  unreachable=0  failed=0  skipped=1  rescued=0  ignored=0
web2 : ok=3  changed=1  unreachable=0  failed=1  skipped=0  rescued=0  ignored=0
db1  : ok=5  changed=2  unreachable=0  failed=0  skipped=0  rescued=1  ignored=0
ghost: ok=0  changed=0  unreachable=1  failed=0  skipped=0  rescued=0  ignored=1

What you should read from these four lines is this. web2 failed around the 4th task and did nothing after that. db1 failed, and rescue saved it and it went all the way — a green light, but a rollback ran. ghost could not be reached over SSH and was written to ignore that. A line where ignored is not 0 is always worth a look. Whether what you decided to ignore is still safe to ignore changes over time.

What you see in the field

Case 1 — a playbook where ignore_errors piled up. A role had ignore_errors: true in eleven places. Each had a reason — "this package may already be installed", "some environments don't have this service". But one slipped in among them. It was a task that deployed a certificate. Even when issuing the certificate failed, the deployment was green, and only when the certificate expired four months later did it come out that the server had been running for two months on the old certificate. After that, one rule was set up — always put a comment saying why it is fine to ignore on an ignore_errors. As you write the comment, you find that most of them should have been a failed_when or a when.

Case 2 — the day the pre-check was placed at the back. A deployment playbook pushed out all the configuration and only at the end checked the version format. On a day when a typo made the version 1.2.3-rc, it stopped after the configuration had gone up on all 40 machines. Reverting took two hours. Now the assert set is at the very front of the play together with any_errors_fatal. If the same typo happens, it stops in 3 seconds, without touching anything.

Case 3 — the pipeline where the rollback was a green light. After automatic rollback was attached with block/rescue, the number of deployment failures became 0. It was read as a good sign, but in fact rescue was running a few times every week and nobody was counting it. The day after a task that leaves "rollback occurred" was put inside rescue and that number was put into the weekly report, it came out within an hour that the health check failed only with a particular image tag.

What you will do in the next lab

In a three-host inventory, you make just one machine fail and confirm with marker files that the default behavior really does keep sending the other two on. You measure what changes when you add ignore_errors to it and what changes when you correct the criterion with failed_when. You build a rollback structure with block/rescue/always and see the rescued column go up, and block the preconditions up front with assert. Then you set "what to do with the other hosts" in two ways with any_errors_fatal and max_fail_percentage and check how the results diverge, and add one unreachable host to see how unreachable differs from failed. Finally, you build a tool that reads the PLAY RECAP and sums the numbers, and compile six runs into a single report.

References