TT Lab
Get started
Learn Learning paths Courses

Ansible Fundamentals

What check mode simulates, and what it cannot

Continue in TT Lab

In one sentence

--check is a mode that asks "what would change if I ran it now", and the accuracy of its answer differs by module. Making check mode trustworthy is not a matter of attaching one flag but a matter of deciding, task by task, how each task behaves in check mode.

Why this was needed

The fear on the day you first run a playbook against production servers is concrete. Nobody knows what this playbook will do to twenty machines. Someone may have hand-edited a configuration, and a task put in six months ago by someone else may now do something unexpected. And you cannot go through the machines one at a time by eye either.

--check is the tool placed in that spot. The official documentation calls this mode a dry run and defines what it does in one sentence: it makes no changes to the remote system, and if a change would likely have happened, it reports that task as changed. If you add --diff, it shows what would differ, line by line.

But this tool betrays you twice. Once with false reassurance — it was clean in check mode but fails when really run. And once with false failure — it died in red in check mode but the real run has no problem at all. Both betrayals have the same cause. Check mode is a mode that guesses the result without really executing, and the quality of the guess differs by module.

How it works

Check mode is decided separately for each task. If a module supports check mode, it reports "it will change" instead of making the real change. If it does not support it, that task is skipped (skipping). command and shell are the typical examples — Ansible cannot know what such a command would do, so it does not run it at all.

This is where the first incident happens. If a command task that only queries is skipped, that task's register variable is empty, and the later tasks that use its value collapse one after another. So put check_mode: false on tasks that only read. That task is really executed even in check mode.

- name: 현재 설치된 판을 읽는다
  ansible.builtin.command: myapp --version
  register: current
  changed_when: false
  check_mode: false

There is one criterion for attaching it. Does this task change the target? Attach it only when you are sure it does not. If you attach check_mode: false to a task that changes things, that task is really executed even in the dry run, and at that moment the dry run becomes a lie.

There is also the opposite direction. A task with check_mode: true never changes anything, even in normal runs. Even in the real run, it reports only "it will change". Use it when you want to write down in advance a task that must not be turned on yet and see only its impact, or to hold a dangerous task until a person confirms it.

False failures happen because of ordering. Consider a playbook in which an earlier task creates a file and a later task edits that file. In check mode, the earlier task does not really create the file, so the later task fails trying to edit a file that does not exist. In this lab environment, lineinfile dies with exactly Destination ... does not exist !. It is not that the playbook is wrong; it is a limit of check mode.

There are three ways to fix it, and they differ in value. First, put when: not ansible_check_mode on the later task so that it is skipped in check mode — the most common and honest. In exchange, what that task would change is missing from the plan. Second, put check_mode: false on the earlier task so that it really creates the file — it is no longer a dry run, so it is hard to recommend. Third, change the design so that a single module manages the whole file to begin with — the best but the largest change.

--diff has a fixed format. Under --- before and +++ after, the changing lines appear as - and +. When only attributes change and not contents, the diff of the attribute JSON appears instead of the file contents (for example, a directory's state: absent changing to state: directory). --diff is not check-mode only — if you attach it to a real run, it shows what it changed while changing it. In incident investigation, this side is often more useful.

For files that contain secrets, also apply no_log: true. Otherwise the diff prints those values into the log as they are.

The three tools have different places.

Tool What it does Connects to the target? Expands templates?
--syntax-check Reads the YAML and the playbook structure No No
--list-tasks Only lists what would run for this selection No No
--check Actually runs the tasks in check mode Yes Yes

The difference confirmed by measurement makes this clear. A playbook that uses an undefined variable passes --syntax-check and --list-tasks with exit code 0. That error first shows up in --check with exit code 2. Conversely, a module-name typo or an indentation accident is caught by all three tools with exit code 4. So CI applies the three in order — the cheap ones first.

When you use check mode as an approval procedure, you need an output a person can read. The green and yellow text flowing across the screen does not remain as grounds for approval. If you set the ansible.posix.json callback as the stdout callback, the whole run comes out as a single JSON, and from it you can pull out just "the list of names of tasks that will change" and attach it to a change request.

ANSIBLE_STDOUT_CALLBACK=ansible.posix.json \
  ansible-playbook -i hosts.ini site.yml --check --diff > plan.json

One step further and it becomes a gate. If you put in CI a script that, after convergence is done, runs check mode once more and ends in failure if even one thing would change, a server fixed by hand by someone shows up before the next deployment. This is the spot where check mode turns from a "tool for looking" into a "tool for guarding".

What you see in the field

First, check mode is clean but the real run fails. Usually it is because shell tasks were skipped. If there are many skipped in check mode, the dry run of that playbook has seen that much less. The habit of reading the skipped number on the PLAY RECAP as the reliability of the plan helps.

Second, check mode is red but actually everything is fine. This is where teams adopting it for the first time give up most often. They reach the conclusion "our playbook can't do --check" and stop using dry runs at all. In reality, it is finished by putting guards on a couple of tasks.

Third, the diff leaks secrets into the log. If you deploy a certificate key or a DB password through a template and apply --diff, the values remain in the CI log as they are. It is better to set up a rule of attaching no_log: true from the start.

Fourth, approvals travel as screen captures. A team that pastes dry-run results in as images cannot find "what we approved back then" six months later. With a single JSON plan output, the approval history remains as searchable text.

Fifth, the honest limits of this lab environment. This Pod has no audit log or external approval system. So the "approval procedure" covers only leaving the plan output as a file, and where to attach that file is discussed only in the reading. Tasks such as restarting a service cannot be handled either because there are no capabilities, so the targets of check mode are limited to files and directories.

References

What you will do in the next lab

Right after writing the playbook, you run it with --check first and confirm that nothing was created, converge it once, and then change three values and pull out three kinds of diff — contents, permission, and one line — with --check --diff. You see a read-only command task skipped in check mode and revive it with check_mode: false, and conversely make a task with check_mode: true that never changes anything even in a real run. You deliberately create a false failure in check mode and fix it with an ansible_check_mode guard. You apply --syntax-check, --list-tasks, and --check in turn to a playbook that uses an undefined variable and leave a table of where the exit codes first diverge, build an approval plan output with the JSON callback, and then write a drift gate script that ends with 1 if anything would change.