TT Lab
Get started
Learn Learning paths Courses

Loki — A Log Store That Does Not Index Logs

Not a single error, yet the 5xx count was a third of the real one

Continue in TT Lab

Goal

In the real Loki inside the Pod, you attach the four LogQL parsers yourself, count the lines with parse errors and the lines that silently failed to yield a value without any error, and work out the true number of 5xx requests.

Why it matters

Loki does not index the body. So to filter by a value inside the body you have to extract it with a parser every time you query, and if you choose the wrong parser the numbers quietly shrink. json reports failure with the __error__ label, but logfmt does not — when it meets a line of a different format it creates no label and moves on, and without the label the label filter that follows silently drops that line. The dashboard shows neither an error nor a warning. That is why, before querying a new stream, you must always measure two numbers first — the number of lines where parsing failed, and the number of lines that did not fail but yielded no value.

Steps

  1. In /root/lk-parsers, start Loki, write date +%s in /root/lk-parsers/anchor.txt, and then load the data with python3 /opt/lab/d5/gen.py parsers "$(cat anchor.txt)". Then count how many lines {app="mixed"} and {app="orders"} each have over the past hour from the reference time, and write them in /root/lk-parsers/01-boot.txt on two lines as mixed=<정수> and orders=<정수> (the placeholders are the integer counts).
  2. In the orders stream, count with the logfmt parser how many lines have status equal to 500. Write the query in /root/lk-parsers/02-logfmt.logql and the answer in /root/lk-parsers/02-logfmt.txt as lines=<정수> (the placeholder is the integer count).
  3. Attach the json parser to the mixed stream and count the lines that parsed successfully and the lines that received the __error__ label. Write the answer in /root/lk-parsers/03-json.txt on two lines as ok=<정수> and err=<정수> (the placeholders are the integer counts), and write one actual value of __error__ on one line in /root/lk-parsers/03-json-name.txt.
  4. Count how many lines, when you attach the logfmt parser to the mixed stream, raised no error but did not get a status label created. Write the query in /root/lk-parsers/04-silent.logql and the answer in /root/lk-parsers/04-silent.txt as lines=<정수> (the placeholder is the integer count).
  5. Select only the lines that contain legacy in the body, extract the status code with the pattern parser, and count the lines where that value is 500. Write the query in /root/lk-parsers/05-pattern.logql and the answer in /root/lk-parsers/05-pattern.txt as lines=<정수> (the placeholder is the integer count). The query must include the pattern parser.
  6. From the same legacy lines, extract the elapsed time (in milliseconds) with the regexp parser and count the lines greater than 1000. Write the query in /root/lk-parsers/06-regexp.logql and the answer in /root/lk-parsers/06-regexp.txt as lines=<정수> (the placeholder is the integer count). The query must include the regexp parser.
  7. Split the lines of the mixed stream into three formats, count them, and create /root/lk-parsers/tally.tsv. It has three lines with no header, and each line has two tab-separated columns, <형식><탭><줄수> (the placeholders are the format, a tab, and the line count). The format names are, in order, json, legacy, and logfmt, and the sum of the three lines must equal the total number of mixed lines counted in step 1.
  8. Work out the true total number of lines in the mixed stream whose status code is 500, write it in /root/lk-parsers/08-total.txt as total=<정수> (the placeholder is the integer count), and on the next line write one sentence starting with reason= (at least 40 characters excluding spaces). Explain in your own words why a single parser cannot count this number.

Notes

Get hold of a stream with mixed formats

In /root/lk-parsers, start Loki, write date +%s in /root/lk-parsers/anchor.txt, and then load the data with python3 /opt/lab/d5/gen.py parsers "$(cat anchor.txt)". Then count how many lines {app="mixed"} and {app="orders"} each have over the past hour from the reference time, and write them in /root/lk-parsers/01-boot.txt on two lines as mixed=<정수> and orders=<정수> (the placeholders are the integer counts).

It takes about 20 seconds until /ready returns ready. Give start and end in nanoseconds based on the value in anchor.txt. You can count with the selector alone, without a parser.

When the format is clean, one logfmt line is enough

In the orders stream, count with the logfmt parser how many lines have status equal to 500. Write the query in /root/lk-parsers/02-logfmt.logql and the answer in /root/lk-parsers/02-logfmt.txt as lines=<정수> (the placeholder is the integer count).

You attach a parser after the selector with a pipe. A label created by the parser can be used after it as a label filter. This is different from looking for a string with a line filter like |= "status=500" — this time, use the value extracted by the parser.

The json parser reports failure

Attach the json parser to the mixed stream and count the lines that parsed successfully and the lines that received the __error__ label. Write the answer in /root/lk-parsers/03-json.txt on two lines as ok=<정수> and err=<정수> (the placeholders are the integer counts), and write one actual value of __error__ on one line in /root/lk-parsers/03-json-name.txt.

When a parser fails, the __error__ label is attached. You just select the lines where that label is empty and the lines where it is not empty. You can see what the label value is by looking directly at the stream in the result — __error_details__ is attached along with it.

Count the lines from which nothing was extracted, without any error

Count how many lines, when you attach the logfmt parser to the mixed stream, raised no error but did not get a status label created. Write the query in /root/lk-parsers/04-silent.logql and the answer in /root/lk-parsers/04-silent.txt as lines=<정수> (the placeholder is the integer count).

In LogQL, a "missing label" is compared as equal to an empty string. Using that property, you can select "lines where the parser could not create a value." This is the scariest number in this lab — it does not show up anywhere on the dashboard, yet it eats away at the total.

For lines with fixed positions, pattern is the right fit

Select only the lines that contain legacy in the body, extract the status code with the pattern parser, and count the lines where that value is 500. Write the query in /root/lk-parsers/05-pattern.logql and the answer in /root/lk-parsers/05-pattern.txt as lines=<정수> (the placeholder is the integer count). The query must include the pattern parser.

A legacy line has the shape legacy handler finished code 500 in 123ms. With pattern you write a name in angle brackets where you want to extract and write the rest literally. It is important to put a line filter before the parser so that only the legacy lines are passed on — lines of other formats do not have this shape.

Extract a number with the regular expression parser and compare it

From the same legacy lines, extract the elapsed time (in milliseconds) with the regexp parser and count the lines greater than 1000. Write the query in /root/lk-parsers/06-regexp.logql and the answer in /root/lk-parsers/06-regexp.txt as lines=<정수> (the placeholder is the integer count). The query must include the regexp parser.

regexp turns only named capture groups into labels — unnamed parentheses are discarded. The extracted value is a string, but if you compare it with an inequality, LogQL converts it to a number. You can get the same answer with pattern too — compare which one is easier to read.

Applied ① — build a table of how many lines each format has

Split the lines of the mixed stream into three formats, count them, and create /root/lk-parsers/tally.tsv. It has three lines with no header, and each line has two tab-separated columns, <형식><탭><줄수> (the placeholders are the format, a tab, and the line count). The format names are, in order, json, legacy, and logfmt, and the sum of the three lines must equal the total number of mixed lines counted in step 1.

A JSON line is a line where the json parser raises no error. A legacy line has legacy in its body. What remains are the logfmt lines, and for these lines logfmt creates status. Be sure to check that the sum of the three numbers matches the total — if it does not, you counted something twice or missed something somewhere.

Applied ② — record the true 5xx count and the fact behind it

Work out the true total number of lines in the mixed stream whose status code is 500, write it in /root/lk-parsers/08-total.txt as total=<정수> (the placeholder is the integer count), and on the next line write one sentence starting with reason= (at least 40 characters excluding spaces). Explain in your own words why a single parser cannot count this number.

Even if you chain two parsers in one query, it cannot read both formats together. Count separately for each format and add them up. If you think about where the "silently dropped lines" counted in step 4 went, you can see how many branches to split into. When you write the answer, also check what each of the three branch numbers was.