TT Lab
Get started
Learn Learning paths Courses

Ansible Fundamentals

If you must reach for the shell, what do you have to own yourself?

Continue in TT Lab

Goal

You throw the same command at command and shell and measure by numbers what differs, and practice by hand how to set up your own reporting criterion, failure criterion, and pipe exit code when you have to use a shell. At the end, you build an audit tool that finds tasks that drop out to a shell without a guard.

Why it matters

When you first use Ansible, a playbook easily becomes a shell script executed over SSH. It runs, but it becomes a playbook that can answer nothing about "is it already in that state", "what changed this time", or "did it fail or succeed". Modules were made to answer those three questions, so if a module exists that does the same job, it comes first. That said, you cannot avoid the shell forever — there is always work that has no module. What matters is to know that the moment you drop out to a shell, every judgment Ansible was making for you disappears, and to write those judgments back in by hand. This lab sets up those three judgments one by one.

Steps

  1. Create /root/ansmod/hosts.ini — in the [web] group, put web1 and web2, give both ansible_host=127.0.0.1 and ansible_port=2222, and use [all:vars] to set ansible_user=root. Then save the output of ansible-doc -s ansible.builtin.command to /root/ansmod/out/doc-command.txt and the output of ansible-doc -s ansible.builtin.shell to /root/ansmod/out/doc-shell.txt.
  2. In /root/ansmod/files/, create three empty files a.txt, b.txt, and c.txt. Create /root/ansmod/boundary.yml and run four tasks on web1 — ls /root/ansmod/files/*.txt once with command (letting the run continue even if it fails), ls /root/ansmod/files/*.txt | wc -l once with shell, echo one two three | wc -w once with command, and the same one once with shell. Leave the four results in /root/ansmod/out/boundary.txt as exactly four lines: glob command rc=<값> / glob shell stdout=<값> / pipe command stdout=<값> / pipe shell stdout=<값> (where the placeholder stands for the value).
  3. Create /root/ansmod/report.yml. On web1, use ansible.builtin.command to run id -un and store the result as who with register, then pick only four fields from that return value and save them as JSON to /root/ansmod/out/result.json — rc, stdout, and changed are taken from the return value as they are, and cmd is a string made by joining the return value's argument list with spaces. Do not add changed_when in this step.
  4. Add two more tasks to /root/ansmod/report.yml. One runs cat /etc/hostname with command, registers it as hn, and gets changed_when: false. The other uses shell to run grep -c "^nosuchuser:" /etc/passwd, registers it as hits, and gets changed_when: false together with a failed_when that counts it as a failure only when the exit code is neither 0 nor 1. Then add three more fields to /root/ansmod/out/result.json — hostname_changed (the changed of hn), grep_rc (the rc of hits), and grep_failed (the failed of hits). The playbook must run to the end.
  5. Create /root/ansmod/pipe.yml. Run the same pipeline cat /root/ansmod/missing.txt | wc -l twice — once with plain shell (registered as bare), and once with set -o pipefail in front and executable set to /bin/bash (registered as guarded). Give both ignore_errors: true and changed_when: false. Leave the result in /root/ansmod/out/pipe.json with four fields — bare_rc, bare_failed, guarded_rc, and guarded_failed. Do not create missing.txt.
  6. Create /root/ansmod/modernize.yml. Without using either command or shell even once, do the following — create the directory /root/ansmod/app with permission 0750, in /root/ansmod/app/app.conf, write the single line env=lab with permission 0640, and read the state of the two paths with a module and leave it in /root/ansmod/out/modernize.json with five fields — dir_mode, dir_isdir, conf_mode, conf_size, and conf_checksum_len (the length of the checksum string).
  7. Create /root/ansmod/bootstrap.yml. With ansible.builtin.raw, run command -v python3 || echo NOPYTHON, register it, and save only the path, with trailing whitespace and CR stripped, as one line to /root/ansmod/out/raw.txt (add changed_when: false). Then tidy the five playbooks you have made so far (boundary.yml, report.yml, pipe.yml, modernize.yml, and bootstrap.yml) so that every task uses an FQCN starting with ansible.builtin..
  8. First create /root/ansmod/legacy.yml, which will be the audit target — it has four tasks: one that runs ansible --version with command and has changed_when: false, one that runs mkdir -p /root/ansmod/legacy/logs with shell, one that runs echo seeded > /root/ansmod/legacy/logs/stamp.txt with shell, and one that runs ls /root/ansmod/legacy/logs with command (put no guard on the last three). Then create /root/ansmod/shell-audit.sh <플레이북경로> (where the placeholder stands for the playbook path): it prints, one per line in alphabetical order, only the names of the tasks in that playbook that use command or shell without changed_when (it must recognize both short names and FQCNs). Finally, save the output of ./shell-audit.sh /root/ansmod/legacy.yml to /root/ansmod/out/audit.txt.

Notes

Write down the targets and find modules in the documentation

Create /root/ansmod/hosts.ini — in the [web] group, put web1 and web2, give both ansible_host=127.0.0.1 and ansible_port=2222, and use [all:vars] to set ansible_user=root. Then save the output of ansible-doc -s ansible.builtin.command to /root/ansmod/out/doc-command.txt and the output of ansible-doc -s ansible.builtin.shell to /root/ansmod/out/doc-shell.txt.

Nothing starts without an inventory. The sshd of this Pod runs at 127.0.0.1:2222, and even if you list two hosts, both connect to the same sshd. ansible-doc is not an internet search but a list of what is installed on this machine right now. -s prints a skeleton you can paste into a playbook. Put the two files side by side and look for option names that exist on only one side — that shows exactly the difference in nature between the two modules.

Throw the same command at command and shell and measure the difference

In /root/ansmod/files/, create three empty files a.txt, b.txt, and c.txt. Create /root/ansmod/boundary.yml and run four tasks on web1 — ls /root/ansmod/files/*.txt once with command (letting the run continue even if it fails), ls /root/ansmod/files/*.txt | wc -l once with shell, echo one two three | wc -w once with command, and the same one once with shell. Leave the four results in /root/ansmod/out/boundary.txt as exactly four lines: glob command rc=<값> / glob shell stdout=<값> / pipe command stdout=<값> / pipe shell stdout=<값> (where the placeholder stands for the value).

command splits the string it receives into words and passes them as they are to the executable — with no shell in between, neither globs nor pipes are interpreted as shell syntax. A glob is noticed immediately because the command dies with a non-zero code, but a pipe pretends to succeed. Leaving that difference as numbers is the whole point of this step. To keep the playbook from stopping at a failing task, add ignore_errors: true, and since it only queries, also add changed_when: false.

Receive the JSON a module returns with register and read it

Create /root/ansmod/report.yml. On web1, use ansible.builtin.command to run id -un and store the result as who with register, then pick only four fields from that return value and save them as JSON to /root/ansmod/out/result.json — rc, stdout, and changed are taken from the return value as they are, and cmd is a string made by joining the return value's argument list with spaces. Do not add changed_when in this step.

A module's return value is not a string but JSON with keys, and register stores that JSON whole in a variable. It contains rc, stdout, stdout_lines, stderr, changed, failed, and cmd — cmd is the list of arguments that was actually executed, so you have to join it to make a string. There is a filter that turns a dictionary straight into a JSON string. And check with your own eyes what changed comes out as even though this task changed nothing — it is the starting point of the next step.

Set the reporting criteria yourself with changed_when and failed_when

Add two more tasks to /root/ansmod/report.yml. One runs cat /etc/hostname with command, registers it as hn, and gets changed_when: false. The other uses shell to run grep -c "^nosuchuser:" /etc/passwd, registers it as hits, and gets changed_when: false together with a failed_when that counts it as a failure only when the exit code is neither 0 nor 1. Then add three more fields to /root/ansmod/out/result.json — hostname_changed (the changed of hn), grep_rc (the rc of hits), and grep_failed (the failed of hits). The playbook must run to the end.

grep -c returns exit code 1 when it finds nothing. That is not an error but the answer "0 matches", yet the default judgment treats anything other than 0 as a failure, so the playbook stops right there. failed_when is the place where a person rewrites the definition of failure, and inside the condition you can use the variable you just registered as it is. For the task with changed_when: false and the task without it, compare how the changed value differs in result.json.

Confirm by numbers that a pipe swallows failures, and block it

Create /root/ansmod/pipe.yml. Run the same pipeline cat /root/ansmod/missing.txt | wc -l twice — once with plain shell (registered as bare), and once with set -o pipefail in front and executable set to /bin/bash (registered as guarded). Give both ignore_errors: true and changed_when: false. Leave the result in /root/ansmod/out/pipe.json with four fields — bare_rc, bare_failed, guarded_rc, and guarded_failed. Do not create missing.txt.

In a shell, the exit code of a pipeline is that of the last command. Even if cat dies at the front, if wc ends with 0 the whole thing is 0, and that task passes in green. set -o pipefail makes the whole pipeline fail if any one part of it fails. However, the default shell (/bin/sh) may not know that option, so you must specify the shell. If the two rc values come out different, you have succeeded — that difference is why the risky-shell-pipe rule of ansible-lint exists.

Move three shell lines to the file, copy, and stat modules

Create /root/ansmod/modernize.yml. Without using either command or shell even once, do the following — create the directory /root/ansmod/app with permission 0750, in /root/ansmod/app/app.conf, write the single line env=lab with permission 0640, and read the state of the two paths with a module and leave it in /root/ansmod/out/modernize.json with five fields — dir_mode, dir_isdir, conf_mode, conf_size, and conf_checksum_len (the length of the checksum string).

There is a module that replaces mkdir -p and chmod in one go, a module that replaces echo >, and a module that replaces the ls -l or stat command. If the three names don't come to mind, find them with ansible-doc -l ansible.builtin | grep -i <낱말> (where the placeholder stands for a word) — that is why you saved the documentation in step 1. Under a key called stat, the return value of the module that reads the state holds mode, isdir, size, and checksum. A permission is a string, so if you leave out the quotes, the octal is read as decimal.

Find Python with raw and tidy the playbooks with FQCNs

Create /root/ansmod/bootstrap.yml. With ansible.builtin.raw, run command -v python3 || echo NOPYTHON, register it, and save only the path, with trailing whitespace and CR stripped, as one line to /root/ansmod/out/raw.txt (add changed_when: false). Then tidy the five playbooks you have made so far (boundary.yml, report.yml, pipe.yml, modernize.yml, and bootstrap.yml) so that every task uses an FQCN starting with ansible.builtin..

raw throws the string over SSH as it is and takes what comes out as it is — so it runs even when the target has no Python, but the received string comes with CR and a newline attached. There is a Jinja filter that strips whitespace from both ends. You can check whether CR remains in the file with od -c. Tidy the FQCNs with your eyes, not by hand — if you pull out all the keys the tasks use with yq -r '.[].tasks[] | keys | .[]' <파일> (where the placeholder stands for the file), the short names stand out at a glance.

An audit tool that finds tasks dropping out to a shell without a guard

First create /root/ansmod/legacy.yml, which will be the audit target — it has four tasks: one that runs ansible --version with command and has changed_when: false, one that runs mkdir -p /root/ansmod/legacy/logs with shell, one that runs echo seeded > /root/ansmod/legacy/logs/stamp.txt with shell, and one that runs ls /root/ansmod/legacy/logs with command (put no guard on the last three). Then create /root/ansmod/shell-audit.sh <플레이북경로> (where the placeholder stands for the playbook path): it prints, one per line in alphabetical order, only the names of the tasks in that playbook that use command or shell without changed_when (it must recognize both short names and FQCNs). Finally, save the output of ./shell-audit.sh /root/ansmod/legacy.yml to /root/ansmod/out/audit.txt.

This is the step where, instead of a person asking "is this shell really needed" in every review, a tool asks it. If the tool closes its eyes depending on the notation, it is not an audit — it must treat shell: and ansible.builtin.shell: as the same thing. The yq in the image is the mikefarah version, so you follow the play list with .[].tasks[], and a key that contains a dot is written in brackets like .["ansible.builtin.shell"]. Asking for a missing key gives null, so you can filter with select(... != null).