TT Lab
Get started
Learn Learning paths Courses

Ansible in Practice

Fix without measuring and you fix the part you already knew

Continue in TT Lab

Summary in one line

A playbook is slow for one of about five reasons, and each is fixed differently — so measure first with profile_tasks and only then pick a knob. And every knob has a cost.

Why this is needed

Every team gets a report that "the deployment takes 40 minutes." But those 40 minutes are a mix of times of completely different natures.

If you start fixing without measuring, people fix what they know. Usually they raise forks and call it done. But if the bottleneck was fact gathering, raising forks barely speeds things up, and only increases the controller's memory and the load on the targets. Measurement is not a warm-up to performance work; it is the first step of the work itself.

How it works

Measure first — profile_tasks

An Ansible callback is a plugin that receives what happens during execution and produces output. If you turn on ansible.posix.profile_tasks, the time taken by each task is printed, and at the end of the run a summary in order of longest time appears once more. There is nothing to install; one line of configuration is enough.

[defaults]
callbacks_enabled = ansible.posix.profile_tasks

The summary is per task. So it tells you "which task" but not "what in that task." The next step is to try the knobs below one at a time.

Knob one — how many facts to gather

gather_facts: true is like secretly slipping a setup task in at the very front of the play. This task scrapes the target's CPU, memory, disks, network interfaces, and mount list. More than a hundred facts come back, but what you actually use is usually just ansible_distribution.

There are two ways to choose.

Knob two — the fact cache, and false facts

If you turn on fact_caching, facts gathered once are kept in a file or Redis and reused on the next run. The jsonfile cache leaves one JSON file per host, so you can open and look at it with your own eyes, which also makes it good for learning. Together with gathering = smart, it becomes "if it is in the cache, don't ask again."

Here lies the most important cost in this module. Not asking again means Ansible does not know when the target changes. It trusts and decides on the disk capacity, kernel version, and in-house local facts that were put in the cache yesterday, as they are today. If a conditional looks at facts, that condition branches on yesterday's facts. There are three defenses used in practice — keep the cache lifetime shorter than the deployment cycle, empty it with --flush-cache in the first step of the deployment pipeline, and do not make risky decisions on cached facts.

Knob three — SSH round trips and pipelining

With pipelining off, Ansible copies the module file to the target for every task, runs it, and deletes it. That is several SSH round trips. With it on, the module is streamed into the remote Python's standard input, reducing the round trips to one. It is usually the cheapest and biggest gain.

There is a reason the default is off. If requiretty is enabled in the target's sudoers, pipelining does not work. So Ansible left it as "turn it on where it works." This setting goes in the [ssh_connection] section, and to see whether it really took effect you must use ansible-config dump --only-changed -t all — without -t all, the settings of connection plugins are not shown.

Knob four — forks and strategy

The two are on different axes.

What it decides Default
forks Number of targets held at the same time 5
strategy Whether hosts wait for each other at every task linear

With forks at 5 and 200 targets, Ansible runs in 40 groups. Raising it makes things faster, but the controller's CPU and memory and the number of open SSH connections rise together.

With strategy: linear, all hosts must finish one task before moving to the next. With free, each host runs to the end at its own speed. In a situation where one slow host holds up everyone, free wins by a lot. But free breaks the ordering promises between hosts. A playbook like "take it out of all web servers and then deploy" loses its meaning under free. You must be especially careful when mixing it with serial or run_once.

Knob five — asynchronous

async and poll are a pair. If poll is not 0, Ansible waits right there, and if it is 0, it receives only the job number and goes straight on to the next task. You collect the result later with async_status.

- name: Start the long job
  ansible.builtin.command: /opt/bin/reindex
  async: 1800
  poll: 0
  register: job

The async value is the maximum time allowed for that job. If you set it short, a perfectly healthy job is cut off midway. This approach fits "jobs that take long but are independent of each other," and using it for a job whose result is needed right away only complicates the playbook.

What it looks like in the field

First, the biggest gain usually comes from facts. The more targets there are, the more fact gathering grows linearly, and most of it goes unused. It is common for the single line gather_facts: false to win by much more than doubling forks.

Second, always check that the setting took effect. ansible.cfg can exist in several places, and only one is used. It is frequent to edit one, run from a different directory, find nothing changed, and say "no effect." The CONFIG_FILE line of ansible-config dump --only-changed tells you which one.

Third, separate where you measure time from where you judge. On emulation or shared runners, the execution time of the same playbook can swing by up to twofold. So when you prove that "it got faster," leave not the time of a single run but evidence that the structure changed — things like a config dump, a cache file, or a measurement file left by a callback.

Fourth, turn the habit of reading measurement results into a tool. Skimming the summary by eye is different from extracting the top few and leaving them as a record. Once records accumulate, you can say "which tasks got slower this week," and from then on performance is something that is managed.

References

What you will do in the next lab

You first leave a baseline with profile_tasks, then turn the knobs one at a time. You save the result of gathering all facts and the result of gathering only the minimal set side by side and count what is missing, turn on the jsonfile cache, change a local fact on the target, and see for yourself that the cache returns the old value as it is, then release it with --flush-cache. You turn on pipelining and use a config dump to check that it really took effect, apply forks and strategy: free, and throw a long job out with poll: 0 and collect it later with async_status. At the end you build a small tool yourself that picks out the slowest tasks from the measurement results.