The Logs Arrived in Three Different Formats
Goal
You normalize three log files in different formats into NDJSON of one line per record, group multi-line exceptions into one record, quarantine lines that could not be parsed, and then merge the three into a common schema.
Why it matters
The first log bundle you receive in the field is not tidy. If you run grep -c with the formats still mixed, one exception is counted as ten errors, and that number goes straight into the report. The time notation also differs from file to file, so sorting as strings does not give time order. That is why, before analysis, you must first decide "what is one record" and "what string to fix the time into." This decision is a contract, not code, and if it is not left in a document, the next person produces different numbers from the same file.
Steps
- Create and run
/root/norm/gen_norm.pyto create app.log, gateway.log, and sys.log under/root/norm/raw/. - In
/root/norm/counts.json, write the physical line count and the logical record count separately. /root/norm/app.ndjson— group the application log into one line per record./root/norm/gw.ndjson— structure the Combined access log./root/norm/sys.ndjson— structure the RFC 5424 syslog./root/norm/rejects.ndjson— quarantine the lines that could not be parsed, with the original text and line number./root/norm/all.ndjson— merge the three into a common schema and line them up in time order./root/norm/norm_report.md— leave the normalization rules and the dropped lines as a report.
Notes
- Convert every
tsto UTC and write it in the form2026-03-05T05:22:31.118Z— an uppercaseTbetween date and time, three millisecond digits, and an uppercaseZat the end. - Python's
datetime.strptimereads both+09:00and+0900with%z. The time in a Combined log can be read with%d/%b/%Y:%H:%M:%S %z, andMaris interpreted only if the locale is C. - The
<134>in syslog is the PRI. The facility is the quotient of dividing by 8, and the severity is the remainder. - Common mistakes: counting stack trace lines as independent records, leaving the response time as a floating-point number in seconds, quietly skipping lines that could not be parsed, and sorting times as strings without converting them to UTC.
- Collect all outputs of this lab under
/root/norm/. They disappear when the session ends, so keep the important ones on screen.
Reproduce the customer's log bundle
Create and run /root/norm/gen_norm.py to create app.log (91 lines), gateway.log (60 lines), and sys.log (30 lines) under /root/norm/raw/.
The three files were written by different programs, so their formats differ. app.log has multi-line exceptions, so its line count is greater than its record count, and sys.log has a few lines truncated in transit. First create /root/norm/raw and write the three files in it.
Count lines and records separately
In /root/norm/counts.json, write app_lines, app_records, gateway_records, syslog_lines, syslog_records, rejected_lines, and total_records. total_records is the sum of the records that survived from the three files.
In app.log, only the lines that start with a time are the head of a record. sys.log has 30 lines, but lines that lack an RFC 5424 header are mixed in, so the record count is smaller. That the two numbers differ is itself the answer to this step.
Group a multi-line exception into one record
In /root/norm/app.ndjson, write one line per record containing ts (UTC, Z), level, logger, msg, and stack_lines. stack_lines is the number of additional lines attached to that record.
If a line starts with a time, it is a new record; otherwise it is a continuation of the previous record. This one rule makes the at … frames and Caused by: attach to the previous record automatically. +09:00 is read by %z, and the UTC conversion is astimezone(timezone.utc).
Structure the access log
In /root/norm/gw.ndjson, write one line per record containing ts (UTC, Z), method, path, status (integer), bytes (integer), and rt_ms (integer milliseconds).
The time in a Combined log is inside square brackets in the form 05/Mar/2026:14:22:31 +0900. It is read by %d/%b/%Y:%H:%M:%S %z, and %b interprets Mar in the C locale. The response time in the last column is a floating-point number in seconds, so multiply by 1000 and round.
Split the PRI into facility and severity
In /root/norm/sys.ndjson, write one line per record containing ts (UTC, Z), host, app, facility (integer), severity (integer), msgid, and msg. Do not put lines that could not be parsed here.
One RFC 5424 line is <PRI>1 TIMESTAMP HOSTNAME APP-NAME PROCID MSGID STRUCTURED-DATA MSG. A single number inside the PRI carries both — the quotient and the remainder of dividing by 8. Skip lines that lack a header for now, and collect them separately in the next step.
Quarantine unparsed lines instead of discarding them
In /root/norm/rejects.ndjson, leave the lines that could not be parsed as src_file, line_no (from 1), raw (the original text as it is), and reason.
If you skip a line the parser could not read with continue, the loss is recorded nowhere. A parse failure is not an exceptional situation but one kind of result. Count the line number from 1 in the original file, and raw must be the untouched original text so you can trace back later.
Merge the three originals into a common schema
In /root/norm/all.ndjson, write ts, source (app|gateway|syslog), severity (ERROR|WARN|INFO), and message in ascending ts order. For the access log, 5xx is ERROR, 4xx is WARN, and the rest is INFO, and for syslog, a severity number of 3 or less is ERROR, 4 is WARN, and the rest is INFO.
If you read and write the three NDJSON files made in the earlier steps, you do not have to parse again. Folding severity into three values is the heart of this step — each original has a different grading system, so if you leave them as they are, you cannot compare them. You have already fixed the times into the same notation, so string sorting is enough for the sort.
Leave the normalization rules in a document
In /root/norm/norm_report.md, write four sections: ## 무엇이 섞여 있었나 ## 어떻게 한 줄 한 레코드로 만들었나 ## 버린 줄과 그 이유 ## 다음에 받을 때의 요구사항 (in order, these mean: what was mixed, how it was made one record per line, the dropped lines and why, and the requirements for the next delivery). It must include, as numbers, the total record count, the physical line count of app.log, and the number of quarantined lines.
Normalization rules are a contract, not code. If you do not write down which lines you treated as the start of a record, how you fixed the time, and where you put the unreadable lines, the next person produces different numbers from the same file. Take the numbers from counts.json.