TT Lab
Get started
Learn Learning paths Courses

Ansible in Practice

Measure the playbook first, then turn the dials

Continue in TT Lab

Goal

You first measure why a playbook is slow using a callback, then turn five knobs in turn: fact gathering, caching, pipelining, concurrent execution, and asynchronous execution. At the end you build a small tool yourself that picks out where to fix from the measurement results.

Why it matters

Reports that "the playbook is slow" always come in, but they mix causes of different natures. Being slow because all facts are gathered on a hundred targets, being slow because every task goes back and forth over SSH three times, and being slow because ninety-nine hosts stand waiting while one host does a ten-minute job are fixed in completely different ways. That is why order matters — measure first, then fix. If you fix without measuring, you fix what you know, and what you know is usually not the bottleneck. And every knob has a cost. A cache is fast but creates false facts, and the free strategy is fast but breaks the ordering promises between hosts. This lab has you turn each knob and look at its cost too.

Steps

  1. In /root/ansperf/ansible.cfg, set the default inventory to ./inventory/hosts.ini and turn on the ansible.posix.profile_tasks callback. In /root/ansperf/inventory/hosts.ini, list web1 (work_sec 1) and web2 (2) in the web group and db1 (3) in the db group (all three have ansible_host=127.0.0.1 ansible_port=2222, and ansible_user in [all:vars] is root). Make /root/ansperf/slow.yml with facts all gathered, a set of three tasks that sleep for 3 seconds, 2 seconds, and 1 second respectively, and one task that writes /root/ansperf/out/baseline.marker. Save the execution result to /root/ansperf/out/profile_baseline.txt.
  2. Create /root/ansperf/subsets.yml — at the play level, turn off automatic gathering with gather_facts: false, and call ansible.builtin.setup on web1 twice. Once with no arguments to gather everything, and once with gather_subset: [min] to gather only the minimum. Capture each with register, save the facts that task returned as JSON to /root/ansperf/out/facts_all.json and /root/ansperf/out/facts_min.json, and run it.
  3. Add gathering = smart, fact_caching = ansible.builtin.jsonfile, fact_caching_connection = ./factcache, and fact_caching_timeout = 3600 to /root/ansperf/ansible.cfg. Then create a local fact that the target reports about itself — write the [app] section and release=1.0.0 in /etc/ansible/facts.d/lab.fact. /root/ansperf/release.yml is a playbook that writes ansible_local.lab.app.release of web1 to the file the variable outfile points to. Empty the cache directory, run this playbook, and leave /root/ansperf/out/rel1.txt.
  4. With the previous step done, the cache now holds 1.0.0. Raise only the release in /etc/ansible/facts.d/lab.fact to 2.0.0. Run the playbook again as is and leave /root/ansperf/out/rel2_cached.txt, then add the option that empties the cache, run it again, and leave /root/ansperf/out/rel3_fresh.txt. Check for yourself how the values of the three files differ.
  5. Add a [ssh_connection] section to /root/ansperf/ansible.cfg and turn on pipelining = True. Then save the output that contains all settings that differ from the defaults (ansible-config dump --only-changed -t all) to /root/ansperf/out/config.txt, and check that an ad-hoc ping to web1 works with pipelining on.
  6. Add forks = 10 to [defaults] of /root/ansperf/ansible.cfg. Create /root/ansperf/free.yml — it uses strategy: free, does not gather facts, and consists of two tasks. The first sleeps for that host's work_sec, and the second writes one line <호스트이름> <work_sec> (host name and work_sec) to out/free-<호스트이름>.txt (with the host name in place of the placeholder). Save the execution output to /root/ansperf/out/free.txt.
  7. Create /root/ansperf/async.yml — on web1, throw out an 8-second job with async: 120 and poll: 0 and capture it with register. Then, in the next task, write /root/ansperf/out/meanwhile.txt (meaning there is something to do in the meantime), wait until it finishes with ansible.builtin.async_status, save the result as JSON to /root/ansperf/out/async.json, and run it.
  8. Create /root/ansperf/slowest.sh — it reads the callback summary from the measurement output file given as the first argument and prints <초>s <태스크이름> (seconds and task name) one per line in order of longest time. The number of lines to print is taken from the second argument, with a default of 3. If the file does not exist, it reports to standard error and exits with a non-zero value. Run this tool on /root/ansperf/out/profile_baseline.txt from step 1 and save the result to /root/ansperf/out/slowest.txt.

Notes

Measure before you fix

In /root/ansperf/ansible.cfg, set the default inventory to ./inventory/hosts.ini and turn on the ansible.posix.profile_tasks callback. In /root/ansperf/inventory/hosts.ini, list web1 (work_sec 1) and web2 (2) in the web group and db1 (3) in the db group (all three have ansible_host=127.0.0.1 ansible_port=2222, and ansible_user in [all:vars] is root). Make /root/ansperf/slow.yml with facts all gathered, a set of three tasks that sleep for 3 seconds, 2 seconds, and 1 second respectively, and one task that writes /root/ansperf/out/baseline.marker. Save the execution result to /root/ansperf/out/profile_baseline.txt.

A callback is configuration, not installation — it is turned on by writing its name in callbacks_enabled. The ansible.posix collection is already in this image. This callback prints the time taken by each task and, at the end of the run, prints a summary in order of longest time once more. The value of this step is not optimization but the baseline. If you fix without measuring, nobody can say what got better.

Count what fact gathering brings back

Create /root/ansperf/subsets.yml — at the play level, turn off automatic gathering with gather_facts: false, and call ansible.builtin.setup on web1 twice. Once with no arguments to gather everything, and once with gather_subset: [min] to gather only the minimum. Capture each with register, save the facts that task returned as JSON to /root/ansperf/out/facts_all.json and /root/ansperf/out/facts_min.json, and run it.

gather_facts: true is like secretly slipping a setup task in at the very front of the play. If you turn it off and call it yourself where needed, you decide when and how much to gather. If you capture a setup task with register, only the facts that run returned are held in <이름>.ansible_facts (where the placeholder stands for the registered variable name) — they are not mixed with facts already attached to the host, which makes comparison easy. The filter for saving pretty JSON is to_nice_json. To see what is missing, use jq 'keys'.

Cache facts in a file and skip gathering on the second run

Add gathering = smart, fact_caching = ansible.builtin.jsonfile, fact_caching_connection = ./factcache, and fact_caching_timeout = 3600 to /root/ansperf/ansible.cfg. Then create a local fact that the target reports about itself — write the [app] section and release=1.0.0 in /etc/ansible/facts.d/lab.fact. /root/ansperf/release.yml is a playbook that writes ansible_local.lab.app.release of web1 to the file the variable outfile points to. Empty the cache directory, run this playbook, and leave /root/ansperf/out/rel1.txt.

When gathering is smart, it means "if it is in the cache, don't gather again." It sits between implicit, which has no cache, and explicit, which does not use the cache. The jsonfile cache leaves one JSON file per host, so you can open it and see what is in it. The setup module reads /etc/ansible/facts.d/*.fact on its own and puts it under ansible_local.<파일이름> (where the placeholder stands for the file name) — if it is in INI form, the section name becomes one level as it is.

False facts that a cache creates

With the previous step done, the cache now holds 1.0.0. Raise only the release in /etc/ansible/facts.d/lab.fact to 2.0.0. Run the playbook again as is and leave /root/ansperf/out/rel2_cached.txt, then add the option that empties the cache, run it again, and leave /root/ansperf/out/rel3_fresh.txt. Check for yourself how the values of the three files differ.

A cache means "it does not ask again," and so Ansible does not know when the target changes. This is the risk that lives alongside turning on a cache. ansible-playbook has an option that throws the cache away on the spot and gathers again (look for "cache" in --help). In practice, you either keep the cache lifetime shorter than the deployment cycle or empty the cache in the first step of the deployment pipeline.

Reduce SSH round trips and leave the changed settings in a file

Add a [ssh_connection] section to /root/ansperf/ansible.cfg and turn on pipelining = True. Then save the output that contains all settings that differ from the defaults (ansible-config dump --only-changed -t all) to /root/ansperf/out/config.txt, and check that an ad-hoc ping to web1 works with pipelining on.

With pipelining off, Ansible copies the module file to the target for every task, runs it, and deletes it — several SSH round trips. With it on, the module is streamed into the remote Python's standard input, reducing the round trips to one. It is usually the cheapest and biggest gain. But it cannot be turned on if requiretty is enabled in the target's sudoers — that is why the default is off. To check whether this setting really took effect, use ansible-config dump --only-changed -t all. Without -t all, the settings of connection plugins are not shown.

How many to hold at once, and whether to make them wait for each other

Add forks = 10 to [defaults] of /root/ansperf/ansible.cfg. Create /root/ansperf/free.yml — it uses strategy: free, does not gather facts, and consists of two tasks. The first sleeps for that host's work_sec, and the second writes one line <호스트이름> <work_sec> (host name and work_sec) to out/free-<호스트이름>.txt (with the host name in place of the placeholder). Save the execution output to /root/ansperf/out/free.txt.

forks is the number of targets Ansible holds at the same time. The default of 5 immediately becomes a bottleneck for an inventory of dozens of servers. strategy is a different axis — the default linear makes all hosts finish one task before moving to the next, and free lets each host run to the end at its own speed. free is not always good. In a playbook where there are ordering promises between hosts (take out first, put back later), that promise is broken. So you must be especially careful when using it with serial or run_once.

Throw a long job out and collect it later

Create /root/ansperf/async.yml — on web1, throw out an 8-second job with async: 120 and poll: 0 and capture it with register. Then, in the next task, write /root/ansperf/out/meanwhile.txt (meaning there is something to do in the meantime), wait until it finishes with ansible.builtin.async_status, save the result as JSON to /root/ansperf/out/async.json, and run it.

If poll is not 0, Ansible waits right there. If you set it to 0, it receives only the job number and goes straight on to the next task — this is "throwing it out." The async value is the maximum time allowed for that job. If you set it short, a perfectly healthy job is cut off midway. What collects it is async_status, and you pass it the ansible_job_id you received. It may not have finished in one go, so ask a few more times with until. This approach fits "jobs that take long but are independent of each other." Using it for a job whose result is needed right away only makes things complicated.

A tool that picks out where to fix from the measurement results

Create /root/ansperf/slowest.sh — it reads the callback summary from the measurement output file given as the first argument and prints <초>s <태스크이름> (seconds and task name) one per line in order of longest time. The number of lines to print is taken from the second argument, with a default of 3. If the file does not exist, it reports to standard error and exits with a non-zero value. Run this tool on /root/ansperf/out/profile_baseline.txt from step 1 and save the result to /root/ansperf/out/slowest.txt.

The summary starts after a separator line made of equals signs, and each line has the form <이름> ----- <초>s (name and seconds). Before the separator, lines with the time printed for each task are mixed in, so you must read only from after you meet the separator. The value in seconds is always printed to two decimal places, so you can carry it over as is. Sorting is sort -rn, and since names contain spaces, separating the seconds from the name with a tab makes it easier to handle later. What this tool does is small, but its place matters — the mistake people make most often in optimization is fixing what they know without measuring.