TT Lab
Get started
Learn Learning paths Courses

Ansible Fundamentals

See what will change before you change it

Continue in TT Lab

Goal

You treat check mode and diff not as "two flags" but as "a behavior set for each task". You deliberately create a false failure and fix it, and you even build a plan output to attach to an approval and a drift gate.

Why it matters

On the day you first run a playbook against production servers, nobody knows what this playbook will do to twenty machines. --check is the tool placed in that spot, but it betrays you twice — check mode is clean but the real run fails, and check mode dies in red but the real run is fine. The cause of both betrayals is the same. Check mode guesses the result without really executing, and the quality of the guess differs by module. So making check mode trustworthy is not attaching a flag but deciding, task by task, "how you behave in check mode". Once that work is done, the dry run goes beyond a picture you look at before deployment and becomes both the grounds for approval and a gate that catches drift.

Steps

  1. Create /root/anschk/hosts.ini — under [web], put web1 (ansible_host=127.0.0.1, ansible_port=2222), and under [all:vars], put ansible_user=root. Then create /root/anschk/site.yml: set the play variables app_env (default lab), app_dir_mode (default "0755"), and motd_owner (default unset), and with four tasks, create the directory /root/anschk/app with the permission app_dir_mode, write to /root/anschk/app/app.conf the two lines env=<app_env> and listen=8080 with permission 0644, write to /root/anschk/app/motd the two lines Welcome=labhub and Owner=unset with permission 0644, and finally, in motd, set the ^Owner= line to Owner=<motd_owner> (turn on create so that it is created if the file does not exist — because in check mode the earlier task does not really create the file). Do not converge yet; run it only with --check --diff and save the output to /root/anschk/out/check1.txt.
  2. Run site.yml for real once with the defaults and --diff, and save the output to /root/anschk/out/converge.txt. Then override three values on the command line and run again with --check --diff — app_env=stage, app_dir_mode=0750, and motd_owner=platform-team. Save that output to /root/anschk/out/diff.txt. The actual files must still be env=lab, permission 0755, and Owner=unset.
  3. Create /root/anschk/probe.yml. Use ansible.builtin.command to run getent passwd root, register it as pw (changed_when: false), and with ansible.builtin.debug print that value and whether it is currently check mode on one line in the form check_mode=<참거짓> pwline=<읽은 값> (a true/false value and the value read). Save the output of running this playbook with --check, including standard error, to /root/anschk/out/probe-check.txt, but the real value must be printed even in check mode. Not a single task may be skipped.
  4. Create /root/anschk/patch.yml — it has two tasks. One writes to /root/anschk/app/fresh.conf the two lines env=lab and listen=9090 with permission 0644, and the other fixes the ^listen= line of that file to listen=9443. First run it as it is with --check and save the failure output, including standard error, to /root/anschk/out/false-failure.txt. Then put on the later task a guard that skips it in check mode, run with --check again, and save that output to /root/anschk/out/false-fixed.txt. Both runs are in check mode, so /root/anschk/app/fresh.conf must never be created.
  5. Create /root/anschk/dryrun.yml. It is a single task that writes to /root/anschk/app/never.conf the one line feature=on with permission 0644, but even when run normally without --check, it must never create the file and must only report that it would change. Put one more task in the same playbook — an ordinary task that writes to /root/anschk/app/dryrun-ran.txt the one line real-run with permission 0644. Run this playbook without any flags and save the output to /root/anschk/out/dryrun.txt. When it finishes, /root/anschk/app/dryrun-ran.txt must exist and /root/anschk/app/never.conf must not.
  6. Create /root/anschk/undef.yml — a single copy task that uses an undefined variable (missing_var) in its content. Also create /root/anschk/badmod.yml — a single task with a typo in the module name, ansible.builtin.coppy. Then create /root/anschk/gates.sh, apply --syntax-check, --list-tasks, and --check in turn to each of the two playbooks, and write only the exit codes, one line each, to /root/anschk/out/gates.txt: <파일이름> syntax-check=<코드> list-tasks=<코드> check=<코드> (the file name, then the exit code of each tool). Run the script to leave the file.
  7. Create /root/anschk/plan.sh. Specify ansible.posix.json as the stdout callback, run site.yml with --check --diff -e app_env=stage, save the entire JSON to /root/anschk/out/plan.raw.json, then pull out only the names of the tasks that will change and save them as a JSON array to /root/anschk/out/plan.json. Run the script to leave the two files. The array must contain exactly one name.
  8. Create /root/anschk/drift-gate.sh. Run site.yml in check mode; if the changed in the summary is 0, print a message whose first line starts with CLEAN and end with 0, and if it is 1 or more, print a message whose first line starts with DRIFT and end with 1. If the check-mode run itself fails, print a message starting with DRIFT-UNKNOWN and end with 1. The arguments given to the script must be passed on to ansible-playbook as they are. Save the output of running with no arguments in the converged state to /root/anschk/out/gate-clean.txt, and the output of running with -e app_env=stage to /root/anschk/out/gate-dirty.txt.

Notes

Look in check mode first, as soon as you create it

Create /root/anschk/hosts.ini — under [web], put web1 (ansible_host=127.0.0.1, ansible_port=2222), and under [all:vars], put ansible_user=root. Then create /root/anschk/site.yml: set the play variables app_env (default lab), app_dir_mode (default "0755"), and motd_owner (default unset), and with four tasks, create the directory /root/anschk/app with the permission app_dir_mode, write to /root/anschk/app/app.conf the two lines env=<app_env> and listen=8080 with permission 0644, write to /root/anschk/app/motd the two lines Welcome=labhub and Owner=unset with permission 0644, and finally, in motd, set the ^Owner= line to Owner=<motd_owner> (turn on create so that it is created if the file does not exist — because in check mode the earlier task does not really create the file). Do not converge yet; run it only with --check --diff and save the output to /root/anschk/out/check1.txt.

Check mode makes no changes to the target, and if a change would likely have happened, it reports that task as changed. So if you run it while nothing exists yet, everything that would be created shows up as changed. If you add --diff, the contents of the files to be created appear as + lines, and the diff for a file that did not exist before prints the earlier range as 0 — that marker remains as evidence that "the file really did not exist at that time". It is safer to capture both standard output and standard error.

Converge once, change three values, and look at three kinds of diff

Run site.yml for real once with the defaults and --diff, and save the output to /root/anschk/out/converge.txt. Then override three values on the command line and run again with --check --diff — app_env=stage, app_dir_mode=0750, and motd_owner=platform-team. Save that output to /root/anschk/out/diff.txt. The actual files must still be env=lab, permission 0755, and Owner=unset.

--diff is not check-mode only — if you attach it to a real run, it shows what it changed while changing it. You see those two uses side by side in one step. If you change the three values, three kinds of diff come out: a diff where the file contents change, a diff where only attributes, not the file contents, change so that JSON is compared, and a diff where only one line changes. A command-line variable takes precedence over a play variable, so you just give -e 이름=값 three times (name and value). Don't forget to check the actual files with your own eyes at the end.

Revive a query task that was skipped in check mode

Create /root/anschk/probe.yml. Use ansible.builtin.command to run getent passwd root, register it as pw (changed_when: false), and with ansible.builtin.debug print that value and whether it is currently check mode on one line in the form check_mode=<참거짓> pwline=<읽은 값> (a true/false value and the value read). Save the output of running this playbook with --check, including standard error, to /root/anschk/out/probe-check.txt, but the real value must be printed even in check mode. Not a single task may be skipped.

Whether check mode is supported is decided by the module. A command module cannot know what the command would do, so it does not support it, and a task of a module that does not support it is simply skipped. When it is skipped, the register variable is empty and the later task collapses — first run it as it is once and see with your own eyes that what follows pwline= is empty. The key that revives it is one line you attach to the task, and there is one criterion for attaching it: are you sure this task does not change the target.

Create a false failure in check mode and fix it with a guard

Create /root/anschk/patch.yml — it has two tasks. One writes to /root/anschk/app/fresh.conf the two lines env=lab and listen=9090 with permission 0644, and the other fixes the ^listen= line of that file to listen=9443. First run it as it is with --check and save the failure output, including standard error, to /root/anschk/out/false-failure.txt. Then put on the later task a guard that skips it in check mode, run with --check again, and save that output to /root/anschk/out/false-fixed.txt. Both runs are in check mode, so /root/anschk/app/fresh.conf must never be created.

In check mode, the earlier task does not really create the file. So a later task that assumes the file exists dies trying to edit a file that does not exist — it is not that the playbook is wrong, it is a limit of check mode. The failure message states that fact as it is, so read it first. A common way to fix it is to put one condition on the later task, and one magic variable tells you whether it is currently check mode. There is also the method of attaching check_mode: false to the earlier task, but at that moment it stops being a dry run, so we don't use it in this step.

Create a task that never changes anything, even in a real run

Create /root/anschk/dryrun.yml. It is a single task that writes to /root/anschk/app/never.conf the one line feature=on with permission 0644, but even when run normally without --check, it must never create the file and must only report that it would change. Put one more task in the same playbook — an ordinary task that writes to /root/anschk/app/dryrun-ran.txt the one line real-run with permission 0644. Run this playbook without any flags and save the output to /root/anschk/out/dryrun.txt. When it finishes, /root/anschk/app/dryrun-ran.txt must exist and /root/anschk/app/never.conf must not.

It is a key with the same name as the one you used in the previous step, but with the opposite value. If you attach that key to a task, only that task is always pinned to a dry run. It is for when you want to write down in advance in the playbook a task that must not be turned on yet and see only its impact, or to hold a dangerous task until a person confirms it. The awkward state of "reported as changed but the file is not there" is exactly what this key does.

Build a table of where the exit codes of the three tools diverge

Create /root/anschk/undef.yml — a single copy task that uses an undefined variable (missing_var) in its content. Also create /root/anschk/badmod.yml — a single task with a typo in the module name, ansible.builtin.coppy. Then create /root/anschk/gates.sh, apply --syntax-check, --list-tasks, and --check in turn to each of the two playbooks, and write only the exit codes, one line each, to /root/anschk/out/gates.txt: <파일이름> syntax-check=<코드> list-tasks=<코드> check=<코드> (the file name, then the exit code of each tool). Run the script to leave the file.

The first two tools neither connect to the target nor expand templates. Variables are expanded only when the task is actually run in check mode — so the results of the two playbooks diverge in different shapes. Which one is caught where first is the answer of this step. Capture the exit code with $? right after the command. Discard the output and keep only the code. When you build CI, this table decides the order — that is why you apply the cheap ones first.

Build a plan output to attach to an approval

Create /root/anschk/plan.sh. Specify ansible.posix.json as the stdout callback, run site.yml with --check --diff -e app_env=stage, save the entire JSON to /root/anschk/out/plan.raw.json, then pull out only the names of the tasks that will change and save them as a JSON array to /root/anschk/out/plan.json. Run the script to leave the two files. The array must contain exactly one name.

The green and yellow text flowing across the screen does not remain as grounds for approval. If you change the callback, the whole run comes out as a single structured JSON — you specify it with the environment variable ANSIBLE_STDOUT_CALLBACK. You can see which callbacks exist with ansible-doc -t callback -l. In the JSON, tasks are under .plays[].tasks[], the task name is .task.name, and per-host results are under .hosts. You changed only app_env, so only one task will change — the rest are already in that state, so they are not changed.

Turn check mode into a gate

Create /root/anschk/drift-gate.sh. Run site.yml in check mode; if the changed in the summary is 0, print a message whose first line starts with CLEAN and end with 0, and if it is 1 or more, print a message whose first line starts with DRIFT and end with 1. If the check-mode run itself fails, print a message starting with DRIFT-UNKNOWN and end with 1. The arguments given to the script must be passed on to ansible-playbook as they are. Save the output of running with no arguments in the converged state to /root/anschk/out/gate-clean.txt, and the output of running with -e app_env=stage to /root/anschk/out/gate-dirty.txt.

If check mode still finds something that would change after convergence is done, it means someone fixed it by hand or the playbook cannot converge by itself — if you put in CI a script that ends in failure under that condition, drift shows up before the next deployment. To pull the number out of the summary line, look at the number after changed=, and you must also handle separately the case where the run fails and there is no summary at all — if you quietly let it pass when you could not read the number, it is decoration, not a gate. The way to pass the script arguments through as they are is "$@". The gate ends in failure, so make sure the shell does not stop there when you save the output.