Pulling the Answer Out of a Log
Goal
You practice turning questions into commands on a single log file. /opt/lab/data/app.log is in the format IP 날짜 메서드 경로 상태코드 [ERROR] (IP, date, method, path, status code, [ERROR]) and has 480 lines.
Why it matters
What you need in an incident is not a log viewer but the ability to turn a question into a pipeline. The question "what is the most frequent?" can almost always be answered with one idiom, sort | uniq -c | sort -rn | head. Two points explain why this idiom is fast: uniq -c counts only adjacent duplicates, so the sort in front is essential, and the commands connected by pipes each run at the same time in separate processes, so even a huge file is processed without being loaded into memory.
Steps
- Write only the number of lines that contain
ERRORto/root/error_count.txt. - Write the single IP that appears most often to
/root/top_ip.txt. - Save only the lines whose status code is
500, unchanged, to/root/five_hundreds.txt. - Save the list of unique paths (4th field) that appear to
/root/paths.txt, one per line. - Save the count for each status code, in the format
코드 개수(code, then count), as three lines in ascending code order, to/root/status_counts.txt. - Save the top three paths by request count, in the format
경로 개수(path, then count), as three lines in descending count order, to/root/top3_paths.txt. - Write the share of responses that are not 200, to one decimal place and without a % sign, to
/root/error_rate.txt. - Create the following three lines in this order in
/root/report.txt.total=<전체 줄 수>(total number of lines)errors=<200 이 아닌 줄 수>(number of lines that are not 200)top_path=<가장 많이 요청된 경로>(most requested path)
Notes
- With
awk '{print $4}' 파일you can extract only the 4th field. - Calculate the ratio in the END block, as in
awk '{t++; if ($5 != 200) e++} END {printf "%.1f\n", e*100/t}'. - Common mistake 1: if you use
uniq -cwithoutsortin front of it, it counts only adjacent lines and the result is wrong. - Common mistake 2: in step 3, using
grep 500also catches lines where 500 appears in an IP or another field. Specify the field.
Count the ERROR lines
Write only the number of lines that contain ERROR to /root/error_count.txt.
Connect a selecting tool and a counting tool with a pipe. grep also has an option that only counts.
The IP with the most requests
Write the single IP that appears most often to /root/top_ip.txt.
Extract the first field and chain sort | uniq -c | sort -rn. uniq -c counts only adjacent duplicates.
Extract only the 500 responses
Save only the lines whose status code is 500, unchanged, to /root/five_hundreds.txt.
The status code is the 5th field. If you grep for the string, it may also catch 500 in other fields.
A list of unique paths
Save the list of unique paths (4th field) that appear to /root/paths.txt, one per line.
The path is the 4th field. sort already has an option to remove duplicates.
Count by status code
Save the count for each status code, in the format 코드 개수 (code, then count), as three lines in ascending code order, to /root/status_counts.txt.
It is the format 'code count', three lines in ascending code order. Use an awk associative array and an END block, or combine sort|uniq -c.
Top 3 most requested paths
Save the top three paths by request count, in the format 경로 개수 (path, then count), as three lines in descending count order, to /root/top3_paths.txt.
The format is 'path count'. The output of uniq -c is in the order 'count value', so you need to swap the order at the end.
Calculate the error rate
Write the share of responses that are not 200, to one decimal place and without a % sign, to /root/error_rate.txt.
You need both the total number of lines and the number of lines that are not 200. If you calculate them at once in awk's END block, you do not need to read the file twice.
Create a summary report
Create the following three lines in this order in /root/report.txt.
total=<전체 줄 수>(total number of lines)errors=<200 이 아닌 줄 수>(number of lines that are not 200)top_path=<가장 많이 요청된 경로>(most requested path)
Combine the values you found in the earlier steps into three lines. The format and order must be exact.