If you must reach for the shell, what do you have to own yourself?
Goal
You throw the same command at command and shell and measure by numbers what differs, and practice by hand how to set up your own reporting criterion, failure criterion, and pipe exit code when you have to use a shell. At the end, you build an audit tool that finds tasks that drop out to a shell without a guard.
Why it matters
When you first use Ansible, a playbook easily becomes a shell script executed over SSH. It runs, but it becomes a playbook that can answer nothing about "is it already in that state", "what changed this time", or "did it fail or succeed". Modules were made to answer those three questions, so if a module exists that does the same job, it comes first. That said, you cannot avoid the shell forever — there is always work that has no module. What matters is to know that the moment you drop out to a shell, every judgment Ansible was making for you disappears, and to write those judgments back in by hand. This lab sets up those three judgments one by one.
Steps
- Create
/root/ansmod/hosts.ini— in the[web]group, putweb1andweb2, give bothansible_host=127.0.0.1andansible_port=2222, and use[all:vars]to setansible_user=root. Then save the output ofansible-doc -s ansible.builtin.commandto/root/ansmod/out/doc-command.txtand the output ofansible-doc -s ansible.builtin.shellto/root/ansmod/out/doc-shell.txt. - In
/root/ansmod/files/, create three empty filesa.txt,b.txt, andc.txt. Create/root/ansmod/boundary.ymland run four tasks onweb1—ls /root/ansmod/files/*.txtonce withcommand(letting the run continue even if it fails),ls /root/ansmod/files/*.txt | wc -lonce withshell,echo one two three | wc -wonce withcommand, and the same one once withshell. Leave the four results in/root/ansmod/out/boundary.txtas exactly four lines:glob command rc=<값>/glob shell stdout=<값>/pipe command stdout=<값>/pipe shell stdout=<값>(where the placeholder stands for the value). - Create
/root/ansmod/report.yml. Onweb1, useansible.builtin.commandto runid -unand store the result aswhowithregister, then pick only four fields from that return value and save them as JSON to/root/ansmod/out/result.json—rc,stdout, andchangedare taken from the return value as they are, andcmdis a string made by joining the return value's argument list with spaces. Do not addchanged_whenin this step. - Add two more tasks to
/root/ansmod/report.yml. One runscat /etc/hostnamewithcommand, registers it ashn, and getschanged_when: false. The other usesshellto rungrep -c "^nosuchuser:" /etc/passwd, registers it ashits, and getschanged_when: falsetogether with afailed_whenthat counts it as a failure only when the exit code is neither 0 nor 1. Then add three more fields to/root/ansmod/out/result.json—hostname_changed(the changed of hn),grep_rc(the rc of hits), andgrep_failed(the failed of hits). The playbook must run to the end. - Create
/root/ansmod/pipe.yml. Run the same pipelinecat /root/ansmod/missing.txt | wc -ltwice — once with plainshell(registered asbare), and once withset -o pipefailin front andexecutableset to/bin/bash(registered asguarded). Give bothignore_errors: trueandchanged_when: false. Leave the result in/root/ansmod/out/pipe.jsonwith four fields —bare_rc,bare_failed,guarded_rc, andguarded_failed. Do not createmissing.txt. - Create
/root/ansmod/modernize.yml. Without using eithercommandorshelleven once, do the following — create the directory/root/ansmod/appwith permission0750, in/root/ansmod/app/app.conf, write the single lineenv=labwith permission0640, and read the state of the two paths with a module and leave it in/root/ansmod/out/modernize.jsonwith five fields —dir_mode,dir_isdir,conf_mode,conf_size, andconf_checksum_len(the length of the checksum string). - Create
/root/ansmod/bootstrap.yml. Withansible.builtin.raw, runcommand -v python3 || echo NOPYTHON, register it, and save only the path, with trailing whitespace and CR stripped, as one line to/root/ansmod/out/raw.txt(addchanged_when: false). Then tidy the five playbooks you have made so far (boundary.yml,report.yml,pipe.yml,modernize.yml, andbootstrap.yml) so that every task uses an FQCN starting withansible.builtin.. - First create
/root/ansmod/legacy.yml, which will be the audit target — it has four tasks: one that runsansible --versionwithcommandand haschanged_when: false, one that runsmkdir -p /root/ansmod/legacy/logswithshell, one that runsecho seeded > /root/ansmod/legacy/logs/stamp.txtwithshell, and one that runsls /root/ansmod/legacy/logswithcommand(put no guard on the last three). Then create/root/ansmod/shell-audit.sh <플레이북경로>(where the placeholder stands for the playbook path): it prints, one per line in alphabetical order, only the names of the tasks in that playbook that usecommandorshellwithoutchanged_when(it must recognize both short names and FQCNs). Finally, save the output of./shell-audit.sh /root/ansmod/legacy.ymlto/root/ansmod/out/audit.txt.
Notes
- First create the inventory in step 1. The sshd of this Pod runs at 127.0.0.1:2222 and key authentication already works.
- Command hint: find modules with
ansible-doc -l ansible.builtin | grep -i <낱말>(where the placeholder stands for a word), see the option skeleton withansible-doc -s <모듈>(the module name), and see the parsed inventory withansible-inventory -i hosts.ini --graph. - Command hint:
yq -r '.[].tasks[] | keys | .[]' <플레이북>(the playbook file) prints every key the tasks use. Check withjq . <파일>(the file) that the JSON you made is real JSON. - Common mistake: a
commandtask that only queries has nochanged_when: false, so every run piles up as changed. - Common mistake: passing a pipe to
shellwithoutset -o pipefail, so the task passes as a success even when the earlier command dies. - Common mistake: writing a permission without quotes, like
mode: 0640, so that the octal is read as decimal. - This Pod has no capabilities, so
systemctl,mount, andsysctl -wdo not work. So this lab covers only files, directories, and query commands — the principle is the same in service management. - command module · shell module · raw module · ansible-doc · Error handling
Write down the targets and find modules in the documentation
Create /root/ansmod/hosts.ini — in the [web] group, put web1 and web2, give both ansible_host=127.0.0.1 and ansible_port=2222, and use [all:vars] to set ansible_user=root. Then save the output of ansible-doc -s ansible.builtin.command to /root/ansmod/out/doc-command.txt and the output of ansible-doc -s ansible.builtin.shell to /root/ansmod/out/doc-shell.txt.
Nothing starts without an inventory. The sshd of this Pod runs at 127.0.0.1:2222, and even if you list two hosts, both connect to the same sshd. ansible-doc is not an internet search but a list of what is installed on this machine right now. -s prints a skeleton you can paste into a playbook. Put the two files side by side and look for option names that exist on only one side — that shows exactly the difference in nature between the two modules.
Throw the same command at command and shell and measure the difference
In /root/ansmod/files/, create three empty files a.txt, b.txt, and c.txt. Create /root/ansmod/boundary.yml and run four tasks on web1 — ls /root/ansmod/files/*.txt once with command (letting the run continue even if it fails), ls /root/ansmod/files/*.txt | wc -l once with shell, echo one two three | wc -w once with command, and the same one once with shell. Leave the four results in /root/ansmod/out/boundary.txt as exactly four lines: glob command rc=<값> / glob shell stdout=<값> / pipe command stdout=<값> / pipe shell stdout=<값> (where the placeholder stands for the value).
command splits the string it receives into words and passes them as they are to the executable — with no shell in between, neither globs nor pipes are interpreted as shell syntax. A glob is noticed immediately because the command dies with a non-zero code, but a pipe pretends to succeed. Leaving that difference as numbers is the whole point of this step. To keep the playbook from stopping at a failing task, add ignore_errors: true, and since it only queries, also add changed_when: false.
Receive the JSON a module returns with register and read it
Create /root/ansmod/report.yml. On web1, use ansible.builtin.command to run id -un and store the result as who with register, then pick only four fields from that return value and save them as JSON to /root/ansmod/out/result.json — rc, stdout, and changed are taken from the return value as they are, and cmd is a string made by joining the return value's argument list with spaces. Do not add changed_when in this step.
A module's return value is not a string but JSON with keys, and register stores that JSON whole in a variable. It contains rc, stdout, stdout_lines, stderr, changed, failed, and cmd — cmd is the list of arguments that was actually executed, so you have to join it to make a string. There is a filter that turns a dictionary straight into a JSON string. And check with your own eyes what changed comes out as even though this task changed nothing — it is the starting point of the next step.
Set the reporting criteria yourself with changed_when and failed_when
Add two more tasks to /root/ansmod/report.yml. One runs cat /etc/hostname with command, registers it as hn, and gets changed_when: false. The other uses shell to run grep -c "^nosuchuser:" /etc/passwd, registers it as hits, and gets changed_when: false together with a failed_when that counts it as a failure only when the exit code is neither 0 nor 1. Then add three more fields to /root/ansmod/out/result.json — hostname_changed (the changed of hn), grep_rc (the rc of hits), and grep_failed (the failed of hits). The playbook must run to the end.
grep -c returns exit code 1 when it finds nothing. That is not an error but the answer "0 matches", yet the default judgment treats anything other than 0 as a failure, so the playbook stops right there. failed_when is the place where a person rewrites the definition of failure, and inside the condition you can use the variable you just registered as it is. For the task with changed_when: false and the task without it, compare how the changed value differs in result.json.
Confirm by numbers that a pipe swallows failures, and block it
Create /root/ansmod/pipe.yml. Run the same pipeline cat /root/ansmod/missing.txt | wc -l twice — once with plain shell (registered as bare), and once with set -o pipefail in front and executable set to /bin/bash (registered as guarded). Give both ignore_errors: true and changed_when: false. Leave the result in /root/ansmod/out/pipe.json with four fields — bare_rc, bare_failed, guarded_rc, and guarded_failed. Do not create missing.txt.
In a shell, the exit code of a pipeline is that of the last command. Even if cat dies at the front, if wc ends with 0 the whole thing is 0, and that task passes in green. set -o pipefail makes the whole pipeline fail if any one part of it fails. However, the default shell (/bin/sh) may not know that option, so you must specify the shell. If the two rc values come out different, you have succeeded — that difference is why the risky-shell-pipe rule of ansible-lint exists.
Move three shell lines to the file, copy, and stat modules
Create /root/ansmod/modernize.yml. Without using either command or shell even once, do the following — create the directory /root/ansmod/app with permission 0750, in /root/ansmod/app/app.conf, write the single line env=lab with permission 0640, and read the state of the two paths with a module and leave it in /root/ansmod/out/modernize.json with five fields — dir_mode, dir_isdir, conf_mode, conf_size, and conf_checksum_len (the length of the checksum string).
There is a module that replaces mkdir -p and chmod in one go, a module that replaces echo >, and a module that replaces the ls -l or stat command. If the three names don't come to mind, find them with ansible-doc -l ansible.builtin | grep -i <낱말> (where the placeholder stands for a word) — that is why you saved the documentation in step 1. Under a key called stat, the return value of the module that reads the state holds mode, isdir, size, and checksum. A permission is a string, so if you leave out the quotes, the octal is read as decimal.
Find Python with raw and tidy the playbooks with FQCNs
Create /root/ansmod/bootstrap.yml. With ansible.builtin.raw, run command -v python3 || echo NOPYTHON, register it, and save only the path, with trailing whitespace and CR stripped, as one line to /root/ansmod/out/raw.txt (add changed_when: false). Then tidy the five playbooks you have made so far (boundary.yml, report.yml, pipe.yml, modernize.yml, and bootstrap.yml) so that every task uses an FQCN starting with ansible.builtin..
raw throws the string over SSH as it is and takes what comes out as it is — so it runs even when the target has no Python, but the received string comes with CR and a newline attached. There is a Jinja filter that strips whitespace from both ends. You can check whether CR remains in the file with od -c. Tidy the FQCNs with your eyes, not by hand — if you pull out all the keys the tasks use with yq -r '.[].tasks[] | keys | .[]' <파일> (where the placeholder stands for the file), the short names stand out at a glance.
An audit tool that finds tasks dropping out to a shell without a guard
First create /root/ansmod/legacy.yml, which will be the audit target — it has four tasks: one that runs ansible --version with command and has changed_when: false, one that runs mkdir -p /root/ansmod/legacy/logs with shell, one that runs echo seeded > /root/ansmod/legacy/logs/stamp.txt with shell, and one that runs ls /root/ansmod/legacy/logs with command (put no guard on the last three). Then create /root/ansmod/shell-audit.sh <플레이북경로> (where the placeholder stands for the playbook path): it prints, one per line in alphabetical order, only the names of the tasks in that playbook that use command or shell without changed_when (it must recognize both short names and FQCNs). Finally, save the output of ./shell-audit.sh /root/ansmod/legacy.yml to /root/ansmod/out/audit.txt.
This is the step where, instead of a person asking "is this shell really needed" in every review, a tool asks it. If the tool closes its eyes depending on the notation, it is not an audit — it must treat shell: and ansible.builtin.shell: as the same thing. The yq in the image is the mikefarah version, so you follow the play list with .[].tasks[], and a key that contains a dot is written in brackets like .["ansible.builtin.shell"]. Asking for a missing key gives null, so you can filter with select(... != null).