TT Lab
Get started
Learn Learning paths Courses

It failed, but the exit code was 0

Turn what the log says into numbers

Continue in TT Lab

Goal

Using only the standard library, build a tool that extracts operational metrics (counts, status distribution, top paths, quantiles, the worst minute, and failure rate per service) from an access log and a deployment CSV.

Why it matters

Production servers have neither pandas nor internet, but they do have csv, Counter, statistics, and datetime. Once a line-by-line parser, quantiles, and comparison of time-zone-aware timestamps come naturally to you, the same tool answers whether the log has 5,000 lines or 500,000. The materials are /opt/fixtures/pyops/access.log (nginx combined with the processing time in seconds as the last column, from 02:00 to 05:00 KST on 2026-09-10, with the outage window from 03:12 to 03:17) and /opt/fixtures/pyops/deploys.csv (service,version,deployed_at,duration_s,status).

Steps

  1. Create /root/pyops/data/logstat.py. logstat.py count <로그> prints the total number of lines and the number of lines parsed successfully by the regular expression as one line, lines=<n> parsed=<m>. Parsing extracts the timestamp, path, status code, and processing time. Here the placeholder stands for the log file.
  2. logstat.py status <로그> prints the counts per status code class as one line, 2xx=<n> 3xx=<n> 4xx=<n> 5xx=<n> (Counter).
  3. logstat.py top <로그> --n 5 prints the top n paths by request count, one per line in the format <건수> <경로>, most frequent first (most_common). The placeholders are the count and the path.
  4. logstat.py latency <로그> prints the p50, p95, and p99 of the processing time as p50=<x> p95=<x> p99=<x>. Use [49], [94], and [98] of statistics.quantiles(times, n=100) (default method), rounded to three decimal places.
  5. logstat.py worst <로그> prints the one-minute window with the most 5xx responses as worst_minute=<YYYY-MM-DDTHH:MM> count=<n>. Build an aware datetime with strptime and group by setting the seconds to 0. Write the time in the log's own time zone (+0900).
  6. Create /root/pyops/data/deploys.py. deploys.py summary <csv> prints, one line per service in order of service name, service=<이름> total=<n> failed=<m> fail_rate=<0.000> median_s=<x> (DictReader and statistics.median; fail_rate to three decimal places). The placeholder in that line is the service name.
  7. Add a time filter with logstat.py status <로그> --since <ISO> --until <ISO>. The two values are ISO 8601 with a time zone, such as 2026-09-10T03:00:00+09:00, and only lines with since ≤ ts < until are counted.
  8. logstat.py report <로그> --json prints to standard output one JSON object that holds all of the metrics above. The keys are lines, parsed, status (an object per class), top (5 entries of [[경로, 건수], ...]), latency (p50, p95, p99), and worst_minute ({"minute": ..., "count": ...}). Save it to /root/pyops/data/report.json.

Notes

Read line by line and count the successful parses

Create /root/pyops/data/logstat.py. logstat.py count <로그> prints lines=<n> parsed=<m>. Parse with a regular expression that extracts the timestamp, path, status, and processing time. The placeholder stands for the log file.

Read the file line by line with for line in f, and if the search of the compiled pattern returns None, skip the line but still count it in lines. You may use the regular expression from the Notes section as it is.

Counts per status code class

logstat.py status <로그> prints one line, 2xx=<n> 3xx=<n> 4xx=<n> 5xx=<n>. The placeholder stands for the log file.

Put f"{status // 100}xx" into a Counter as the key. A class that does not appear must be printed as 0 (dict.get(k, 0)).

The top n paths by request count

logstat.py top <로그> --n 5 prints <건수> <경로> one per line, most frequent first. The placeholders are the count and the path.

Counter.most_common(n) returns a list of (value, count) pairs, most frequent first. The output order is the count, then the path.

Quantiles, not the average

logstat.py latency <로그> prints p50=<x> p95=<x> p99=<x>. Round [49], [94], and [98] of statistics.quantiles(times, n=100) to three decimal places.

With n=100 there are 99 cut points, so indexes 49, 94, and 98 are p50, p95, and p99. Leave method at its default (exclusive).

The minute with the most 5xx responses

logstat.py worst <로그> prints worst_minute=<YYYY-MM-DDTHH:MM> count=<n>. Keep the log's own time zone and group by one-minute windows by setting the seconds to 0. The placeholder stands for the log file.

Using %z with strptime gives an aware datetime. Use ts.replace(second=0, microsecond=0) as the Counter key and look at most_common(1). Format the output with strftime("%Y-%m-%dT%H:%M").

Failure rate per service from the deployment CSV

Create /root/pyops/data/deploys.py. deploys.py summary <csv> prints, one line per service in order of service name, service=<이름> total=<n> failed=<m> fail_rate=<0.000> median_s=<x>. The placeholder in that line is the service name.

csv.DictReader gives a dictionary per row, and every value is a string. Collect duration_s into a defaultdict(list) and compute statistics.median. fail_rate is failed / total to three decimal places.

Filter by time-zone-aware timestamps

logstat.py status <로그> --since <ISO> --until <ISO> counts only lines with since ≤ ts < until. The two values are ISO 8601 with a time zone (for example 2026-09-10T03:00:00+09:00).

If you accept them with type=datetime.fromisoformat in argparse, they become aware datetimes. Comparing with a naive value raises TypeError, so check that the input includes a time zone.

Everything in one JSON report

logstat.py report <로그> --json prints one JSON object holding lines, parsed, status, top (5 entries, [path, count]), latency, and worst_minute ({minute, count}), and saves it to /root/pyops/data/report.json.

Putting a datetime into json.dumps fails, so convert it to a string first. Save with logstat.py report ... --json > /root/pyops/data/report.json.