TT Lab
Get started
Learn Learning paths Courses

LFCS — Linux Foundation System Administrator

Shaping and Aggregating Logs

Continue in TT Lab

Goal

You process a log you made yourself with grep, sed, awk, cut, tr, sort, and uniq, and pull values out of structured data with jq and yq.

Why it matters

The text-processing problems on the LFCS do not test "do you know the one-line answer" but do you think in terms of fields. If you sweep the whole line with a regular expression, even a 500 in the byte-count column gets caught. So you first grasp the structure that the status code is the fourth field, and put the tool on top of that. Aggregation is the same — sort | uniq -c | sort -rn is an idiom, but if you do not know why you have to sort first (uniq removes only adjacent duplicates), it collapses with even a slight change in the situation. The last two steps are today's operational reality. Configuration is mostly JSON or YAML, and you need hands that finish with jq and yq instead of starting Python to pull out a single value.

Steps

  1. Create /root/lfcs-text/access.log exactly as the 12 lines below. Fields are separated by a single space and are <IP> <메서드> <경로> <상태코드> <바이트> (IP, method, path, status code, bytes).
10.0.0.117 GET /api/users 200 512
10.0.0.118 POST /api/users 201 128
10.0.0.117 GET /api/orders 500 64
10.0.0.120 GET /healthz 200 12
10.0.0.117 GET /api/users 200 512
10.0.0.116 DELETE /api/users 403 96
10.0.0.118 GET /api/orders 500 64
10.0.0.117 POST /api/orders 201 256
10.0.0.120 GET /healthz 200 12
10.0.0.117 GET /api/users 404 32
10.0.0.115 GET /api/orders 200 1024
10.0.0.118 GET /healthz 200 12
  1. Put only the lines whose status code starts with 5 into /root/lfcs-text/errors.txt, in the original order.
  2. Write the byte total per IP to /root/lfcs-text/bytes_by_ip.txt in the format <IP> <합계> (IP and total), in descending order of total.
  3. Write the IP that appears most often to /root/lfcs-text/top_ip.txt as one line, <횟수> <IP> (count and IP).
  4. Extract the second field (the HTTP method), convert it to lowercase, remove duplicates, sort it in dictionary order, and write it to /root/lfcs-text/methods.txt.
  5. Copy access.log to /root/lfcs-text/access-v2.log, and then in that copy replace /api/ in the paths with /v2/api/. The content from before the edit must remain as /root/lfcs-text/access-v2.log.bak, and the original access.log must be unchanged.
  6. Create /root/lfcs-text/services.json with the content below, and write the sum of replicas of the services whose tier is backend to /root/lfcs-text/jq_out.txt as a single number.
{
  "cluster": "homelab",
  "services": [
    {"name": "api", "replicas": 3, "tier": "backend"},
    {"name": "web", "replicas": 5, "tier": "frontend"},
    {"name": "worker", "replicas": 2, "tier": "backend"}
  ]
}
  1. Create /root/lfcs-text/deploy.yaml with the content below, and write two lines, replicas=<값> and image=<값> (each followed by the value), to /root/lfcs-text/yq_out.txt.
apiVersion: apps/v1
kind: Deployment
metadata:
  name: lfcs-web
  namespace: lfcs
spec:
  replicas: 4
  template:
    spec:
      containers:
        - name: web
          image: nginx:1.27

Notes

Create the log file to analyze

Create /root/lfcs-text/access.log exactly as the 12 lines below. Fields are separated by a single space and are <IP> <메서드> <경로> <상태코드> <바이트> (IP, method, path, status code, bytes).

Just create the 12 lines from the instructions as they are. The field separator is a single space and the line order must stay the same for the later results to come out right.

Picking out only the 5xx responses

Put only the lines whose status code starts with 5 into /root/lfcs-text/errors.txt, in the original order.

The status code is the fourth field. If you look for 500 in the whole line, lines with 500 in the byte count can get caught too, so specify the field when you decide.

Byte total per IP

Write the byte total per IP to /root/lfcs-text/bytes_by_ip.txt in the format <IP> <합계> (IP and total), in descending order of total.

Use the IP as the key in an awk associative array and keep adding the fifth field. At the end, iterate over the array to print, and sort numerically in descending order.

The IP that appears most often

Write the IP that appears most often to /root/lfcs-text/top_ip.txt as one line, <횟수> <IP> (count and IP).

To count duplicates you have to sort first. Make use of the fact that the output of the tool that attaches counts is in 'count value' order.

Normalizing the method list

Extract the second field (the HTTP method), convert it to lowercase, remove duplicates, sort it in dictionary order, and write it to /root/lfcs-text/methods.txt.

Chain together a tool that cuts out fields and a tool that changes characters. Removing duplicates can also be done with an option of the sorting tool.

Substituting while protecting the original

Copy access.log to /root/lfcs-text/access-v2.log, and then in that copy replace /api/ in the paths with /v2/api/. The content from before the edit must remain as /root/lfcs-text/access-v2.log.bak, and the original access.log must be unchanged.

If you attach a suffix to the in-place editing option, the content from before the edit remains under that name. You may change the substitution delimiter to a character other than a slash.

JSON aggregation

Create /root/lfcs-text/services.json with the content below, and write the sum of replicas of the services whose tier is backend to /root/lfcs-text/jq_out.txt as a single number.

Filter the array by a condition, gather only the field you want into an array, and then use the function that adds up that array, and it is done in one line.

Extracting YAML values

Create /root/lfcs-text/deploy.yaml with the content below, and write two lines, replicas=<값> and image=<값> (each followed by the value), to /root/lfcs-text/yq_out.txt.

Path expressions are almost the same as jq. Array elements are accessed by index, and depending on the implementation the output may come with quotes attached, so check the result by eye.