TT Lab
Get started
Learn Learning paths Courses

Finding the Cause in Logs

The Absence of Logs Is Also Information

Continue in TT Lab

One-line summary

No logs does not mean the investigation is over. Your job includes narrowing the time with other traces, and then making sure logs are left for the next incident.

Why this is needed

You get a report that "a few failed yesterday afternoon" and go looking for the logs. But retention is 24 hours so they are already deleted, or the log level is WARN so the needed information was not printed, or there was no log statement in that interval to begin with.

If you close it at this point with "cannot confirm because there are no logs," the same thing repeats.

Traces that are not logs

Trace What it tells you How to check
File mtime When something changed find /etc -mmin -1440 -type f
DB timestamps When the data was created/changed The distribution of created_at, updated_at
Process start time When it was restarted ps -eo pid,lstart,cmd
System boot time When the server came up uptime -s
File size changes When it surged Comparison between backup snapshots
Package install time What came in and when /var/log/dpkg.log, rpm -qa --last
Shell history What people did ~/.bash_history (if timestamps are set)
Certificate validity period The time of an outage caused by expiry openssl x509 -noout -dates

The data itself is a log. If you aggregate the created_at distribution of failed orders per minute, you can pinpoint the minute the incident started even without logs.

SELECT date_trunc('minute', created_at) AS m, count(*)
FROM orders WHERE status='FAILED' AND created_at >= now() - interval '2 days'
GROUP BY 1 ORDER BY 1;

The point where the value spikes in this result is the time of the incident. And if you cross-check that time against the deployment records, configuration changes, and boot times, the candidates shrink to one or two.

Absence is evidence too

An empty interval in a log is itself information.

If you do not translate "it was not printed" into "we don't know," and instead ask "what was not printed," the scope narrows.

Leave something for next time

If you found the cause at the end of the investigation, the last task is to make it possible to know faster than this time if the same thing happens again.

  1. Add log statements — on the failure path, what failed and why. Along with the values that were needed in the post-incident investigation (request ID, target, time taken).
  2. Adjust the retention period — 24 hours is not enough for most investigations. At the very least, keep the error logs longer.
  3. Correlation ID — attach an ID to each request so it can be traced across systems. Without it, you have to align the logs of several services by time alone.
  4. One metric — make what was the problem this time a counter or a gauge. Next time it will be visible on a graph.

If you write these four under "preventing recurrence" in the report, it is not a formal sentence; it actually shortens the next investigation by hours.

Post-hoc instrumentation — when you need it right now

If the problem is in progress right now and there are no logs, attach observation while it reproduces.

It is rough, but overwhelmingly better than nothing. And this file becomes the only basis at the next meeting.

When you first attach observation to someone else's system

At a site you enter as an FDE, there are usually no logs, or there are but you cannot see them. With no permissions and no freedom to deploy as you like, there is an order of seeing the most with the fewest changes.

First, find what already exists. Before attaching anything new, count what that system is already producing. Web server access logs, the database's slow query log, load balancer statistics, cloud flow logs. They are usually turned on and nobody is looking at them.

Second, measure from the outside. Even if you cannot fix the code, you can knock from the outside. If you call a critical path at intervals of a few minutes and record the response time and status code, when the outage started can be known from that alone. Even if you do not know the cause, knowing the time is half the investigation.

Third, attach at the boundary. Even if you cannot fix the application, in many cases you can change the proxy or sidecar in front of it. If you make it record time and status per request, you get per-request observation without changing a single line of code.

Fourth, only then fix the code. By the time you get here, you already know where to instrument. If you touch the code from the start, you end up instrumenting the wrong place, and you also pay the cost of reverting that change.

Decide at the start what to keep and what not to keep. The more it is someone else's system, the easier it is to turn on logging without knowing what counts as personal information. For a setting that keeps the whole request body, look at what is in it before turning it on.

When reporting, write a number together with that number's source. If you write not "it's slow" but "this path's p95 is 4.2 seconds, and the median over the past 7 days at the load balancer is 0.4 seconds," then the next action is decided on the spot. Without a source, you end up arguing about that number first.

What you see in the field

What you will do in the following lab

You receive a site where the error logs of the incident interval are gone. Using only the data's created_at, the empty interval of the access log, and the modification times of configuration files, you narrow the time of the incident down to the minute and reduce the candidates to one.

At the end, you build a post-hoc instrumentation tool yourself. The grader actually runs it to check that times and observed values are left periodically.