Without Knowing the Normal You Cannot See the Spike
One-line summary
When you pull a single number out of a log, you can judge whether it is a bad value only if you know the usual value.
Why this is needed
"There are 103 errors" is not information. If the usual was 5, it is serious, and if the usual was 120, things actually got better. The fact alone that the response time is now 200 milliseconds does not tell you whether it is good or bad. If the usual was 40 milliseconds, it is serious, and if it was 220 milliseconds, it is nothing.
So the first step of log analysis is not finding the cause but building a baseline. And at an unfamiliar customer, the baseline does not exist as a document, so you must build it from within the same log. The interval remaining after you exclude the time of the incident is the normal value.
How it works
When you first open an access log, the order is this.
1. Measure the overall size. How many lines, and what interval it covers. If you are looking for an event from yesterday afternoon in a 6-hour log, that log does not hold the answer in the first place.
2. Look at the status code distribution. How many 200s, how many 4xx, how many 5xx. Here you must always look at 4xx and 5xx separately. 4xx means the client sent something wrong and 5xx means the server broke, so whether the two numbers move together or separately is the direction of the cause. If only 4xx increased after a deployment, the API contract changed, and if only 5xx increased, something inside broke.
3. Slice along the time axis. Count 5xx by minute. If they are spread evenly, it is a chronic problem, and if they are concentrated at one point, it is an event. This distinction changes the response completely. If they are concentrated, the investigation can end by asking what was there at that time.
4. Slice by path. Whether errors are concentrated on a particular endpoint or spread across everything. If concentrated, it is a problem of that feature, and if spread, it is a problem of a common dependency (database, authentication, network).
5. Subtract the baseline. Count the errors in the remainder, excluding the incident interval. Only with this number can you write the sentence "it was 33 as usual and then 70 in one minute," and that sentence becomes the first line of the report.
What you see in the field
Here is one practical sense to add. Do not be fooled by the average.
An average response time of 1380 milliseconds does not mean users waited 1.4 seconds on average. The average comes out to that value even in a situation where 240 of 340 requests are at or under 260 milliseconds and 70 exceed 3 seconds. The average describes a user who does not exist.
And this distortion works in only one direction. The average barely detects the tail getting worse. Even if 100 slow requests get twice as bad, from 3 seconds to 6 seconds, the overall average rises by only about 30 milliseconds, and that is not enough to cross any alert threshold.
So it is no contradiction when the customer says "it sometimes freezes for a few seconds" while the dashboard average is normal. Both are true, and they are simply looking at different things.
When the log does not hold the answer
In the course of an investigation you meet cases where the log itself cannot be trusted. If you do not know this and keep digging, you spend hours looking for an answer that does not exist, so there are things you should check first.
Loss. If the structure is that logs gather in a buffer and are sent later, then when the process is forcibly terminated, the last few seconds vanish entirely. Yet the decisive moment of an outage is exactly those few seconds. "There is no log right before it died" does not mean there is no clue; it is a strong clue that the exit was abnormal. That is because if the exit had been normal, closing logs would have been left.
Truncation. It is common for a collector to limit the length of a line. A long stack trace or a JSON body gets cut off in the middle, and the lines that follow get picked up as separate entries. On the parsing side they look like malformed lines and are quietly discarded. This is one of the reasons the error count you tallied comes out lower than the real one.
Time. When you merge the logs of several machines and their clocks are off, causality looks reversed. If the result seems to have been recorded before the cause, you should suspect the clock rather than set up a strange hypothesis. And if time zones get mixed, a nine-hour illusion arises. Keeping log times in UTC and converting only when displaying is the only way to eliminate this problem.
Sampling. A high-traffic service sometimes records only a portion of its logs instead of all of them. In that case, "that user's request is not in the log" does not mean there was no request. If you count without knowing the sampling ratio, the number means nothing.
To sum up, before opening a log, first check four things. What interval does it cover, how far is it intact, is the clock right, and is it all or part. The few minutes that check takes prevent the hours spent looking for an answer that does not exist. And if, after the check, the log does not hold the answer, writing exactly that in the report is worth far more than writing a guess. That is because what to record additionally next time comes out of that sentence.
What you will do in the next lab
From a web access log of 1269 lines, you find the total error volume, the time of the incident, and the causal path, and finally compute the normal value separately to turn the size of the surge into a number.