TT Lab
Get started
Learn Learning paths Courses

Ansible Fundamentals

Why modules come first, and how to use the shell when you must

Continue in TT Lab

In one sentence

ansible.builtin.command does not go through a shell, and ansible.builtin.shell does. That one-line difference decides whether redirects, pipes, and globs work, and the moment you drop out to a shell, idempotency, reporting, and failure judgment all become your own responsibility.

Why this was needed

The playbook of someone using Ansible for the first time almost always looks the same. Because they carry over what they used to do as it is, every task is a shell:.

- name: 설정 배포
  ansible.builtin.shell: |
    mkdir -p /etc/myapp
    echo "env=prod" > /etc/myapp/app.conf
    chmod 640 /etc/myapp/app.conf

It runs. The problem is that this task cannot answer any question you ask it. Is it already in that state? It doesn't know. What changed in this run? It doesn't know. Can it show what would change before changing? It can't. Did it fail or succeed? It only says it succeeded if the last command's exit code is 0. So this task is reported as changed every time it runs, and even if one machine out of twenty quietly gives a different result, nobody knows.

Written with modules, the same job looks like this.

- name: 설정 자리를 만든다
  ansible.builtin.file:
    path: /etc/myapp
    state: directory
    mode: "0755"

- name: 설정을 쓴다
  ansible.builtin.copy:
    dest: /etc/myapp/app.conf
    content: "env=prod\n"
    mode: "0640"

The number of lines is similar, but the nature is different. A module first reads the target's current state, and if it is already in that state, it does nothing and returns changed: false. In check mode, it reports what it would change instead of changing it. The result is not a string but JSON with keys, so a later task can take .mode and .checksum straight out. These three — state judgment, advance notice, and structured return — are the whole reason for "modules first".

How it works

The real difference between command and shell is just the presence of a shell. command splits the string it receives into words and passes them as they are to the executable. There is no /bin/sh in between. shell passes the whole string to a shell. So everything the shell used to do now differs.

What you pass command shell
ls /srv/app/*.conf The glob is not expanded, so it looks for a file literally named *.conf The shell expands it
echo a b c 뒤에 파이프와 wc -w Everything from the pipe symbol onward becomes arguments to echo A pipe is actually set up
echo x > /tmp/f The greater-than sign is also an argument. The file is not created The redirect works
; 와 && Arguments Shell operators

There is one spot here that many people find confusing. Environment variables are expanded even in command. Since ansible-core 2.16, the command module has an expand_argument_vars option that defaults to true, so the module itself expands $HOME without a shell. If you want to turn it off, give expand_argument_vars: false. So the accurate sentence is not "if it doesn't go through a shell, nothing gets expanded" but "shell syntax doesn't work".

It is also worth knowing how command goes wrong quietly. When a glob isn't expanded, the command fails with exit code 2 and is noticed immediately. Pipes are different. If you pass echo one two three | wc -w to command, echo prints one two three | wc -w as it is and succeeds with exit code 0. The task is green and only the result is wrong. A success that is worse than a failure.

If you must use a shell, you set three things yourself.

First, changed_when. command and shell have no way to know what they changed, so they unconditionally report changed. A task that only queries needs changed_when: false. Without it, a playbook that changes nothing piles up changed every day, and the moment that number loses its meaning, real changes go unnoticed too.

Second, failed_when. The default judgment is "a non-zero exit code means failure". But grep returns 1 when it finds nothing, and that is an answer, not an error. In such cases you write the definition of failure yourself, like failed_when: result.rc not in [0, 1].

Third, the exit code of a pipe. In a shell, the exit code of a pipeline is that of the last command. In cat 없는파일 | wc -l, even if cat dies, wc ends with 0, so the whole thing is 0. The task is a success and the result is 0. To prevent this, put bash's set -o pipefail in front and also give executable: /bin/bash (the default shell may not know pipefail). ansible-lint has a rule named risky-shell-pipe that catches exactly this spot.

- name: 로그에서 오류 줄을 센다
  ansible.builtin.shell:
    cmd: set -o pipefail; grep ERROR /var/log/app.log | wc -l
    executable: /bin/bash
  register: errors
  changed_when: false
  failed_when: errors.rc not in [0, 1]

What a module returns is JSON. If you receive it with register, it contains rc, stdout, stdout_lines, stderr, changed, failed, and cmd. stdout_lines is already a list of lines, so there is no reason to do split('\n') yourself, and cmd keeps the list of arguments that was actually executed, which is used in incident investigation. The keys returned differ by module, and that list is written in the RETURN section of ansible-doc.

ansible-doc is not a search engine but a list of what is installed. ansible-doc -l prints every module this machine can actually use right now. With -s you get a skeleton you can paste into a playbook. With -t you can pick the plugin type (callback, filter, lookup, connection). Look here first, not on the internet, for "is there a module that does this job".

raw is a tool for exceptions. Both command and shell need Python on the target to run — because the module code is Python. For machines that do not have Python yet (a freshly installed server, network equipment, a container with Python removed), use raw. raw throws the string over SSH as it is and takes what comes out as it is. It is cheap but does nothing for you — no idempotency, no return structure, no line-ending cleanup, so it is common for the received string to come with a CR attached. Use it only for bootstrapping, and stop using it once Python is installed.

Why use FQCNs. Writing it short, like copy:, works today. But on a machine with several collections installed, there can be more than one module with the same name, and which one gets picked is decided by the search path. If you write ansible.builtin.copy all the way, that ambiguity disappears. A reader also sees at a glance that "this one is from core". ansible-lint has a rule named fqcn that requires it.

What you see in the field

First, a playbook that starts with shell goes back to being a shell script. If one task is shell, the next task easily becomes shell too. In half a year the playbook becomes a shell script executed over SSH, and no reason remains to use Ansible. So the first question asked in review is "is there really no module to replace this shell".

Second, in incident investigation changed lies. A team that did not put changed_when: false on query tasks gets dozens of changed every day. Even if a real change is mixed in among them, nobody can find it. Conversely, a team that keeps changed honest can say "nothing changed yesterday" as evidence.

Third, pipe exit code incidents are quiet. A backup verification task was tar -tzf backup.tar.gz | wc -l, and on the day the file was corrupted, tar died but wc returned 0 and the task succeeded. You learn the backup is broken on the day you need to restore. One line of set -o pipefail stands in between.

Fourth, the honest limits of this lab environment. Lab Pods have no capabilities, so systemctl, mount, and sysctl -w do not work. So examples such as "move the service restart that used to be done with the shell to ansible.builtin.service" cannot be judged in this environment and were left out of the lab. Instead, we handle only what really runs in this Pod, such as files, directories, and query commands. The principle is the same.

References

What you will do in the next lab

You throw the same command at command and shell and measure for yourself what differs for globs and pipes, and leave those numbers in a file. Using register, you take rc, stdout, changed, and cmd out of the received JSON, put changed_when: false on a query task and failed_when on a task for which exit code 1 is normal, and correct the criteria. You confirm by numbers that a pipe without pipefail swallows a failure and then fix it, and move three shell lines to the file, copy, and stat modules so that only structured return values remain. Finally, you look for Python with raw, tidy the whole playbook with FQCNs, and write an audit script that finds tasks that drop out to a shell without a guard.