TT Lab
Get started
Learn Learning paths Courses

Ansible Fundamentals

Idempotency — The Second Run Should Do Nothing

Continue in TT Lab

In one sentence

The quality of a playbook is judged not by whether the first run succeeds but by whether the second run gives changed=0.

Why this was needed

The real value of a configuration management tool is not "automating installation" but "being able to check at any time whether the current state is the desired state, and to bring it into line". For that, it must be safe to run at any time. Only then can you run it periodically from cron to catch drift, and not be afraid to rerun it in a deployment pipeline after an earlier stage fails.

By contrast, a playbook that is not idempotent restarts the service or appends lines to a file every time it runs. People then come to fear it, and scary automation ends up unused. At that moment the infrastructure goes back to manual work.

How it works

Three mechanisms create idempotency.

First, modules that know the state. Modules such as file, copy, template, and lineinfile read the current state and change only the parts that differ. If there is nothing to change, they report ok.

Second, guards on shell commands. command/shell cannot be idempotent by themselves, so you have to attach conditions.

Mechanism Meaning
creates: /path Do not run if that path already exists
removes: /path Do not run if that path does not exist
changed_when: <조건> Define yourself when to report "changed"
failed_when: <조건> Define the failure decision yourself instead of using the exit code

changed_when: false is especially common. A command that only queries the state (for example, checking a version) is never a change.

Third, handlers. A handler is a task that runs "only when something has actually changed". Use it to make a service reread its configuration only when the configuration file has changed.

tasks:
  - name: 설정 배치
    ansible.builtin.template:
      src: app.conf.j2
      dest: /etc/app/app.conf
    notify: reload app

handlers:
  - name: reload app
    ansible.builtin.command: /usr/bin/app-reload

There are three key rules. (1) notify fires only when that task is changed. (2) Even if several tasks call the same handler, the handler runs only once. (3) By default, handlers run after all the tasks of the play have finished. If you really must run them in the middle, use meta: flush_handlers.

A trap comes from the third rule. If the play fails midway, handlers that have not yet run simply vanish. You end up in a state where the configuration has changed but the service has not reread the new configuration. In such cases, you can make handlers run even after a failure with --force-handlers.

What you see in the field

First, "a deployment that restarts every time". The cause is almost always that the rendered result of a template differs every time. Timestamps, random values, and dictionary iteration with no guaranteed order are the culprits. If the rendered result is the same, the template module reports ok and the handler stays quiet too.

Second, the real way to verify. The standard idempotency test is to run the playbook twice in a row in CI and check that both changed and failed are 0 in the second run. Make it judged by an exit code instead of having a person look at it.

Third, idempotency and drift. An idempotent playbook is itself a drift detector. Run it against a server that was fixed by hand and changed appears, and that number is "how far it has diverged from the code". This property is exactly the same idea as Terraform's plan.

How to keep changed honest

In Ansible, idempotency is not the result being the same but changed being honest. If changed appears when nothing changed, handlers run for nothing, and if ok appears when something changed, the service is left on the old configuration.

command and shell are always changed. This is because Ansible does not know what the command did. So you write down when to run it and what to count as a change yourself.

- name: 스키마 이전
  ansible.builtin.command: /srv/app/migrate --apply
  args:
    creates: /srv/app/.migrated      # 이 파일이 있으면 건너뛴다
  register: migrate
  changed_when: "'applied' in migrate.stdout"
  failed_when: migrate.rc != 0 and 'already up to date' not in migrate.stderr

The first priority is to keep it from running at all with creates/removes, and if it has to run, the second is to correct the judgment with changed_when.

A handler runs only once at the end of the play. Even if you notify the same handler ten times, it runs once. That is an advantage, but it also means that if the play fails midway, the handler does not run at all. The configuration file is changed, but the daemon is left without a restart. On the next run the file is already correct, so ok appears, and the handler never runs. --force-handlers closes this gap, and a better method is to have a separate task that checks for the state that needs a restart.

Validate configuration files before writing them. If you write a bad configuration and then restart, the service dies.

- name: nginx 설정
  ansible.builtin.template:
    src: app.conf.j2
    dest: /etc/nginx/conf.d/app.conf
    validate: nginx -t -c %s      # 통과해야 실제 자리에 놓는다
  notify: reload nginx

Look first with --check and --diff. However, a task that depends on the result of an earlier task either fails or gives a wrong answer in check mode. Attach check_mode: false to such a task so that it really only reads, even in check mode. Once check mode starts to lie, nobody uses it, and from then on people run things straight in production.

What you will do in the next lab

You run the same playbook twice so that the first run has changes and the second gives changed=0. You confirm that the handler fires only in the first run, and put a guard on a shell command so that running it twice leaves only one line in the log. Finally, you build an idempotency verdict report.