TT Lab
Get started
Learn Learning paths Courses

Ansible in Practice

Rules, profiles, exceptions, and the gate that uses them

Continue in TT Lab

Summary in one line

--syntax-check only looks as far as whether the YAML makes sense, and ansible-lint looks at quality. And the secret to keeping a linter alive for a long time is not turning on many rules but placing exceptions in the narrowest possible scope.

Why this is needed

Ansible is dangerous because even a badly written playbook finishes green. A task with no name, a shell command containing a pipe, a file write that sets no permissions all finish as ok or changed. The trouble comes six months later.

A linter pulls those six months forward to right before the commit. But the moment you turn it on, the team immediately meets the next problem. If you apply a linter to an old repository, hundreds of findings come out, and a few of them really do need exceptions. How you place exceptions then decides whether the linter is still alive half a year later.

How it works

The layers differ

Tool What it sees What it cannot see
--syntax-check YAML structure, the shape of plays and tasks, whether module names exist Quality, idempotency, whether values make sense
ansible-lint Conventions such as names, idempotency, permissions, FQCN Values at run time
assert task Whether values make sense (port range, environment name) Static conventions
molecule Bringing a role up for real to run convergence, idempotency, and verification scenarios It does not replace the three above

The further down, the more expensive. So the gates run the cheap ones first.

Rule ids and profiles

A finding from ansible-lint always carries a rule id. Sometimes the detail is split with square brackets, as in name[play]. The following are ones that actually appear in the version (6.17.2) in this course's lab image.

Rule id What it catches
name[play] A play with no name
name[missing] A task with no name
name[casing] A name that starts with a lowercase letter
no-free-form A call written on one line, like copy: src=a dest=b
no-changed-when A command that may change state has no reporting criterion
risky-shell-pipe Uses a pipe without setting pipefail
risky-file-permissions Creates a file without setting mode
command-instead-of-module Calls via the shell a command for which a dedicated module exists
fqcn[action-core] A short module name

Rules are bundled into profiles. In the order min · basic · moderate · safety · shared · production, each one up is stricter, and a higher profile includes all the rules of the lower ones. So when you bring a linter into an old repository, you do not jump to production in one go. You apply basic to make CI green, and raise it to moderate the following month. The reason profiles exist is this migration path.

If the version differs, rule names differ too. So always look at your own version's list first with ansible-lint -L.

Two places to put exceptions

- name: Pack the release bundle  # noqa: command-instead-of-module
  ansible.builtin.command: tar -czf /tmp/rel.tgz -C /srv app.conf
  changed_when: false

# noqa: <규칙id> (with the rule id in the placeholder) removes only that task from that rule. Because it stays right next to the code, a reviewer can ask "why was it removed?", and it can be deleted when the reason disappears.

# .ansible-lint
profile: production
exclude_paths:
  - legacy/
skip_list:
  - name[casing]

skip_list turns off that rule in the whole repository. Use it only for conventions the team has agreed on (for example, allowing task names that start with a lowercase product name). If you dump many rules in here because there are many findings, all that remains is the illusion that "the linter is on."

exclude_paths is different again. It does not turn rules off; it means not scanning that path at all. You put code made by others or bad examples left in on purpose for teaching. However, if you give a file name directly as an argument, this list is ignored — exclusion is a rule for "when scanning."

The real reason to have a configuration file is not convenience but making people's hands and CI run with the same rules. If only CI gets --profile production, developers think they passed and push, and fail in CI.

Where the linter cannot see — assert

A linter looks at static conventions. But incidents also come from values. 80 comes in for a port, an environment name has a typo, the replica count becomes 0. These can be known only at run time, so you make the playbook ask itself with ansible.builtin.assert.

- name: Assert that the port is usable
  ansible.builtin.assert:
    that:
      - app_port is integer
      - app_port >= 1024
    fail_msg: "app_port must be an integer of 1024 or above, got {{ app_port }}"

Two things are important. First, ask before changing anything. If you stop after deploying about half of it, undoing it is much more expensive. Second, always write fail_msg. Without it, the failure message comes out as the raw condition expression and the recipient does not know what to fix. And a value passed with -e is a string unless otherwise specified — this is where the reason is integer is false comes from.

What molecule adds

Molecule is a framework for testing roles. For each scenario it brings up a target (Docker, Podman, cloud), applies the role, applies it once more to see whether changed=0 (idempotency), checks the result with a verification playbook, and cleans up. It automatically checks "does it really converge?", which the linter cannot see.

This lab image does not include molecule. The lab Pod has no internet so it cannot be installed, and launching containers is also blocked. So the lab for this module leaves molecule out and sets up only the stage before it (syntax, lint, preconditions). That front stage is still needed even in teams that use molecule — because molecule takes minutes and a syntax check takes less than 1 second.

What it looks like in the field

First, adoption always starts not with "fix everything" but with "don't get worse." Apply a low profile to make CI green, apply the higher standard only to new code, put old paths in exclude_paths, and take them out when you touch them.

Second, half of the no-changed-when findings are real defects. Attaching changed_when: false to a lookup command is not about fitting a format but about making the report honest. A playbook that is always changed hides the day something really did change.

Third, always verify that a gate script actually blocks. Scripts that only check and always exit with 0 are actually common. Such a gate is worse than none — it creates the illusion that checking is happening. The first test of that script is, on the day you make it, to feed it bad input on purpose and see whether it exits with a non-zero value.

Fourth, exceptions have a deadline. When you add a # noqa, if you write why you added it in one line on the same line or right above, you can remove it half a year later when the reason has gone. An exception with no reason written stays forever.

References

What you will do in the next lab

You deliberately write a bad playbook that passes the syntax check, pull out the rule ids the linter catches, and leave them as a list. You raise a clean playbook that does the same job up to the basic profile, and then add FQCN and mode to raise it up to production. You put an exception on a single line with # noqa for one command that has no module, write the profile, excluded paths, and rules to skip in .ansible-lint, and pass the whole repository in one go. You block preconditions with assert, and finally build a gate script that looks at syntax, lint, and preconditions in one go, and check whether it really blocks when fed a bad directory.