One Line Is Not One Event
One-line summary
The first step of log analysis is not counting but deciding what one record is. In a bundle of mixed formats, if you mistake the line count for the record count, one exception becomes ten errors and goes straight into the report.
Why this is needed
The first bundle you receive from a customer is usually not tidy. The application prints in a format of its own taste, the front web server writes a Combined access log, and the firewall and proxy send syslog. If you put the three files in one directory and run grep -c ERROR, one number comes out, and that number means nothing.
There are two reasons. First, Java and Python exceptions span several lines. After one header line come five or six at … frames, and then a Caused by: is attached as well. Counted by lines, one exception becomes ten. Second, each line writes the time differently. One carries +09:00, one is 05/Mar/2026:14:22:31 +0900, and one is UTC with a Z. If you sort the times as strings, these three line up in a jumble.
So one step is needed before analysis. It is turning the three formats into a structured form of one line, one record, aligning the times to one standard, and unifying the field names. This work is called normalization.
How it works
Normalization consists of four decisions.
First, what to treat as the start of a record. The sturdiest way is "if a line starts with a time, it is a new record; otherwise it is a continuation of the previous record." Exception frames start with a tab or spaces and Caused by: has no time, so they attach to the previous record automatically. This one rule solves multi-line exceptions.
Second, what string to fix the time into. RFC 3339 sets the notation used on the internet. Section 5.1 of this document also writes down why to use this notation — if the offset notation is the same and the number of decimal places is the same, string sorting is time-order sorting. So if you fix every time as UTC, three millisecond digits, and Z, as in 2026-03-05T05:22:31.118Z, then in all later work a single sort is enough.
Third, where to put the structured result. The practical default is JSON Lines (NDJSON). Because it is a format with one JSON object per line, line-oriented tools such as head, tail, and split work as is, and even a 100 GB file can be read by streaming. A single JSON array is not suited to large logs, because you have to read the file to the end before you can use the first record.
Fourth, what to unify the field names to. There is no reason to reinvent the wheel here. The OpenTelemetry log data model defines a log record with fields such as Timestamp, ObservedTimestamp, SeverityText, SeverityNumber, Body, and Attributes, and writes as a design requirement that "existing log formats must be able to be mapped to this model unambiguously." The Elastic Common Schema is another dictionary aiming at the same spot. Whichever you use, it matters to settle on one within the team.
If you deal with syslog, RFC 5424 is worth reading once. The <134> at the head is called the PRI, and a single number in it carries both facility and severity — facility is the quotient and severity is the remainder (you divide by 8). Section 6.2.3 tightens further: TIMESTAMP is a subset of RFC 3339, and T and Z must be uppercase.
What you see in the field
Dropped lines quietly appear. If the parser skips a line it could not read with continue, that loss is recorded nowhere. RFC 5424 writes in section 6.1 that a transport receiver may truncate or discard a message that exceeds its supported size (at least 480 octets, and up to 2048 octets should be supported). That is, truncated lines occur normally. So you treat a parse failure not as an exceptional situation but as one kind of result, and keep it separately with the original text, file name, and line number. Later, when you ask "why was the record empty at this time," that file becomes the answer.
For the same reason, RFC 5424 section 8.3 recommends putting important information at the front of the message. That is because even if the back gets cut off, the front remains.
Normalization rules are a contract, not code. If you do not write down which lines are treated as the start of a record, how the time is fixed, and where unreadable lines are put, the next person produces different numbers from the same file. This document is also the basis for demanding "from now on, please send it as one JSON line" in the next conversation with the customer.
What you will do in the next lab
You build three files that reproduce a customer bundle as it is, and count lines and records separately. Then you group a multi-line exception into one record, structure the access log and the syslog each, and quarantine the lines that could not be parsed. Finally, you merge the three into a common schema, line them up in time order, and leave the normalization rules and the dropped lines as a report. The grader re-parses the originals itself and compares them with your results.