TT Lab
Get started
Learn Learning paths Courses

Loki — A Log Store That Does Not Index Logs

When a parser silently extracts nothing

Continue in TT Lab

In one line

Loki does not index the body, so to filter by a value inside the body you have to extract it with a parser every time you query. A parser attaches an __error__ label when it fails, but there are parsers that raise no error even when they extract nothing.

Why this was needed

A team was measuring the error rate of a payment service with Loki. The query was {app="pay"} | logfmt | status="500", and the dashboard stayed calm for months. But when they counted the actual 5xx requests for the same period in an incident meeting, it was three times the dashboard figure.

The cause was simple. Some paths of that service were logging in JSON. When the logfmt parser meets a JSON line, it raises no error. It just fails to extract any label and moves on. Without the label, status="500" is false, and the line silently drops out. The dashboard shows neither "no data" nor "parse error." The number is simply smaller.

How it works

LogQL has four parsers. All of them attach after the selector and line filters, and each time you query they reread those lines to create labels.

Parser Format it reads When it fails
logfmt Lines of space-separated 키=값 pairs (the placeholders are the key and the value) Extracts nothing, with no error
json One line holding a JSON object Attaches __error__="JSONParserErr"
pattern A fixed shape written with <이름> placeholders (the placeholder is the field name) Extracts nothing if the shape differs
regexp RE2 with named capture groups Extracts nothing if it does not match

So an investigation must always go two ways. First count how many lines are broken with | json | __error__!="", and then count how many lines had no error but produced no value with | logfmt | status="". Only when both numbers are 0 can you trust the parser for that stream.

If you want to remove lines with parse errors, add | __error__="". But do not add it out of habit — because those lines are exactly the signal that "logs of a different format are mixed in." First count, then learn the cause, and only then remove.

Extracted labels live only inside that query. They are neither stored nor indexed. So when you change the parser, past data is reinterpreted along with it — the opposite of metrics. With metrics, if you instrument them wrongly the past is lost forever, but with logs, as long as the body remains you can reread it later with different eyes. This is the biggest benefit of the choice "do not index."

Values are always strings. To compare like | dur_ms > 500, LogQL converts them to numbers, but if units are mixed (milliseconds and seconds in one field), a wrong answer comes out silently. For fields that hold numbers, it is safer to put the unit in the name.

What it looks like in the field

The most common incident is "two formats in one stream." The application logs in JSON, but a proxy or runtime in front of it mixes plain-text warnings into the same stdout. The stream labels are the same, so it is one stream, and a query can use only one parser.

The fix is in the placement, not the query. Send logs of a different format to a different stream with different labels. Once you split them at the collector, the query gets simpler and the parser errors disappear. If it is a legacy system where that is not possible, at least build a query that counts separately per parser and adds them up, and write that fact next to the dashboard.

One more thing. The pattern parser is faster and easier to read than regexp, so for lines with fixed positions such as access logs it is almost always better. Use regular expressions only for lines where the positions vary.

What you will do in the next lab

You put one stream with three mixed formats and one stream with clean formats into the real Loki in the Pod, and attach the four parsers in turn. You count the lines where json raises an error and the lines logfmt passed over silently, extract values from the legacy lines with pattern and regexp, and finally work out the true number of 5xx requests, which comes out only when you count separately by format and add them up.