Pipes Join Small Programs Into a Sentence
In one sentence
grep selects, sed edits, and awk counts. Connect these three tools with pipes, and five lines are enough to extract an answer from hundreds of thousands of log lines.
Why this matters
You get an incident report: "Responses have been slow since 3 p.m." The log has 400,000 lines, and just opening it in an editor takes a few seconds. What you need here is not a log viewer but the ability to turn a question into a command.
- How many requests failed? -> select, then count
- Which endpoint is the problem? -> extract the field, group and count, then sort
- Is a particular IP unusual? -> do the same with a different field
All three questions have the same shape. So you only need to memorize one idiom.
... | sort | uniq -c | sort -rn | head
uniq -c counts only adjacent duplicates, so sort in front is essential, and the sort -rn at the end sorts by count in descending order. These four pieces are the standard answer to the question "what is the most frequent?"
How it works
The criteria for choosing a tool are simple. If you only select lines, use grep; if you edit lines, use sed; if you work with fields or calculate, use awk. Using a tool stronger than you need makes things slower and harder to read.
You should also know what a pipe really is. Each command connected by pipes runs in a separate process at the same time, and the stdout of the earlier one is connected to the stdin of the later one. Two properties follow from this.
- Even a large file is processed as a stream without loading all of it into memory. That is why 400,000 lines are fast.
- A variable changed inside a pipeline does not remain outside it, because it runs in a subshell.
And the exit code of a pipeline is, by default, that of the last command. In curl ... | jq ... | wc -l, even if curl fails, if wc prints 0 the whole thing looks like a success. The mechanism that prevents this trap in scripts is set -o pipefail, which we cover in detail in the next course.
It is a good idea to fix a few performance anti-patterns now.
| Instead of this | Write this |
|---|---|
| `cat f | grep p` |
| `sort | uniq` |
| `cat f | wc -l` |
| `echo "$s" | cut -d. -f1` |
What you see in the field
First, aggregating status codes. One line, awk '$9 ~ /^5[0-9][0-9]$/ {print $7}' access.log | sort | uniq -c | sort -rn | head -10, gives you the ten paths that returned the most 5xx. Type this before you open a dashboard.
Second, do ratio calculations with awk. Counts alone are often not enough to judge. If you calculate everything at once in the END block, as in awk '{t++; if ($5 != 200) e++} END {printf "%.1f\n", e*100/t}', you do not need to read the file twice.
Third, do not read the log from the beginning. The first move is to cut out only the five minutes before and after the time of the incident. The habit of scanning all 400,000 lines not only eats time; on a disk whose I/O is already saturated, the diagnosis itself makes the incident worse.
Habits with a large file
When logs reach tens of gigabytes, reducing how much you read matters far more than which tool you choose. Just a few habits can turn a command that takes minutes into one that takes seconds.
Cut first. Narrowing the range by time is the first move. If the log is sorted by time, extract only the interval with something like sed -n '/03:00/,/03:10/p', or pick only the file for that time range in the first place. And if you are sure what you are looking for, stop at the first match with grep -m 1. There is no reason to read all 400,000 lines to the end.
Take out only the columns you need. If you narrow the fields first, as in awk '{print $7}', the data the following sort handles becomes much smaller. sort is the most expensive spot in a pipeline, so reducing data before it has the biggest effect.
Check whether sorting is really needed. If you only need to count, you can count in one pass with an awk associative array. awk '{c[$7]++} END {for (k in c) print c[k], k}' | sort -rn | head does not sort everything and sorts only the result at the end, so it is much faster when there are many lines and few distinct values.
Turn off the locale. Putting LC_ALL=C in front makes sort and grep switch their character comparison rules to simple byte comparison, which is noticeably faster. However, the sort order changes, so be careful with results you show to people.
And read compressed logs as they are, without decompressing them. With zgrep and zcat you do not need to create temporary files on disk, and you avoid the accident of filling up the remaining space while diagnosing on a disk that is already full. Making sure the diagnosis does not make the incident worse is basic courtesy when working on a large system.
What you will do in the next lab
Using /opt/lab/data/app.log as your material, you extract the error count, the IP with the most requests, counts by status code, the top paths, and the error rate in turn, and finally bundle those results into a three-line summary report.