TT Lab
Get started
Learn Learning paths Courses

It failed, but the exit code was 0

Reading 5,000 log lines without a spreadsheet

Continue in TT Lab

Summary

You do not need pandas to pull "how many, what percentage, the top 5, p95, the worst minute" out of thousands of log lines and a hundred-odd CSV rows. csv, collections.Counter, statistics, and datetime are enough, and a tool built that way runs on any server with no dependencies.

Why this matters

In an incident meeting someone asked, "at what time did the 5xx errors cluster?", and somebody copied the logs to a laptop and pasted them into Excel. It worked because there were 5,000 lines, but the next month there were 500,000. The production server has no pandas and no internet. Yet the calculations you need are only counting, grouping, quantiles, and bucketing by time range. The standard library has all four, and a tool built from it is a single file to deploy.

How it works

Reading. Split each log line with a regular expression. The nginx combined format has a timestamp like [10/Sep/2026:03:14:07 +0900], a request like "GET /api/orders HTTP/1.1", a status code, and the processing time appended at the end. Read the file one line at a time (for line in f). Because the whole file is never loaded into memory, the same code works for 500,000 lines. Count the lines that fail to parse, but do not stop. Logs always contain broken lines.

Read CSV with DictReader from the csv module. It treats the first row as column names and gives you a dictionary per row. If you handle values containing commas and quoting by hand with split(","), you will certainly get it wrong, which is why the module exists. Every value is a string, so you do the int() and float() conversions yourself.

Counting. Counter from collections is a dictionary that counts hashable values. With Counter(status // 100 for ...) you get status classes grouped as 2, 4, and 5, and most_common(5) returns the top 5. defaultdict(list) is for collecting values per service.

Quantiles. quantiles(data, n=100) from statistics returns 99 cut points that divide the data into 100 groups of equal probability. So p50 is [49], p95 is [94], and p99 is [98]. A cut point is a value linearly interpolated between the two nearest data points, and the default method is exclusive. Even for the same data the values differ slightly by method, so a tool should record which method it used so that others can reproduce the result. Use median() when you need just the median.

Time. strptime(s, "%d/%b/%Y:%H:%M:%S %z") from datetime turns an nginx timestamp into an object with a time zone attached (aware). %z reads +0900. A naive object (no time zone) and an aware object cannot be compared, so inputs such as --since should also be accepted with a time zone, as in fromisoformat("2026-09-10T03:00:00+09:00"). For "one-minute buckets", count by the key dt.replace(second=0, microsecond=0).

import re, statistics
from collections import Counter
from datetime import datetime

LINE = re.compile(r'\[(?P<ts>[^\]]+)\] "(?P<method>\S+) (?P<path>\S+) [^"]*" (?P<status>\d{3}) \d+ "[^"]*" "[^"]*" (?P<rt>[\d.]+)$')

def parse(line):
    m = LINE.search(line)
    if not m:
        return None
    return {"ts": datetime.strptime(m["ts"], "%d/%b/%Y:%H:%M:%S %z"),
            "path": m["path"], "status": int(m["status"]), "rt": float(m["rt"])}

by_class = Counter(); times = []
for line in open("/opt/fixtures/pyops/access.log"):
    r = parse(line)
    if r:
        by_class[f"{r['status'] // 100}xx"] += 1
        times.append(r["rt"])
q = statistics.quantiles(times, n=100)
print(by_class["5xx"], round(q[94], 3))   # 5xx 건수, p95

Exporting. Produce the result in two forms: one line for people and JSON for machines (json). A datetime cannot go into JSON directly, so convert it to an isoformat() string. Fix the number of decimal places of floats with round(x, 3) so that two runs look the same.

What it looks like in the field

The most common mistake is the average. It is typical to see a report of an average response time of 0.08 seconds while p99 is 4 seconds: the slowest 1% is buried in the average. That is why the habit of reporting quantiles is needed. The second is time zones. The log says +0900, but --since is given in UTC or as a naive value, so everything is off by an hour and you conclude "there was no problem at that time". The third is silently discarding parse failures. From the day the format changes slightly, the tool reports 0 records; if it also reports the skipped count, as in parsed=0 skipped=52000, you notice that very day.

What you will do in the next lab

You analyze /opt/fixtures/pyops/access.log (about 5,000 lines, with an outage window between 03:12 and 03:17) and deploys.csv using only the standard library, by building logstat.py and deploys.py. The tools report the line count and the number of successful parses, counts per status class, the top paths, p50, p95, and p99, the minute with the most 5xx responses, the failure rate and median per service, a --since/--until filter, and finally a JSON report that holds everything.