Loki — A Log Store That Does Not Index Logs
Not a single error, yet the 5xx count was a third of the real one
Goal
In the real Loki inside the Pod, you attach the four LogQL parsers yourself, count the lines with parse errors and the lines that silently failed to yield a value without any error, and work out the true number of 5xx requests.
Why it matters
Loki does not index the body. So to filter by a value inside the body you have to extract it with a parser every time you query, and if you choose the wrong parser the numbers quietly shrink. json reports failure with the __error__ label, but logfmt does not — when it meets a line of a different format it creates no label and moves on, and without the label the label filter that follows silently drops that line. The dashboard shows neither an error nor a warning. That is why, before querying a new stream, you must always measure two numbers first — the number of lines where parsing failed, and the number of lines that did not fail but yielded no value.
Steps
- In
/root/lk-parsers, start Loki, writedate +%sin/root/lk-parsers/anchor.txt, and then load the data withpython3 /opt/lab/d5/gen.py parsers "$(cat anchor.txt)". Then count how many lines{app="mixed"}and{app="orders"}each have over the past hour from the reference time, and write them in/root/lk-parsers/01-boot.txton two lines asmixed=<정수>andorders=<정수>(the placeholders are the integer counts). - In the
ordersstream, count with thelogfmtparser how many lines havestatusequal to 500. Write the query in/root/lk-parsers/02-logfmt.logqland the answer in/root/lk-parsers/02-logfmt.txtaslines=<정수>(the placeholder is the integer count). - Attach the
jsonparser to themixedstream and count the lines that parsed successfully and the lines that received the__error__label. Write the answer in/root/lk-parsers/03-json.txton two lines asok=<정수>anderr=<정수>(the placeholders are the integer counts), and write one actual value of__error__on one line in/root/lk-parsers/03-json-name.txt. - Count how many lines, when you attach the
logfmtparser to themixedstream, raised no error but did not get astatuslabel created. Write the query in/root/lk-parsers/04-silent.logqland the answer in/root/lk-parsers/04-silent.txtaslines=<정수>(the placeholder is the integer count). - Select only the lines that contain
legacyin the body, extract the status code with thepatternparser, and count the lines where that value is 500. Write the query in/root/lk-parsers/05-pattern.logqland the answer in/root/lk-parsers/05-pattern.txtaslines=<정수>(the placeholder is the integer count). The query must include thepatternparser. - From the same legacy lines, extract the elapsed time (in milliseconds) with the
regexpparser and count the lines greater than 1000. Write the query in/root/lk-parsers/06-regexp.logqland the answer in/root/lk-parsers/06-regexp.txtaslines=<정수>(the placeholder is the integer count). The query must include theregexpparser. - Split the lines of the
mixedstream into three formats, count them, and create/root/lk-parsers/tally.tsv. It has three lines with no header, and each line has two tab-separated columns,<형식><탭><줄수>(the placeholders are the format, a tab, and the line count). The format names are, in order,json,legacy, andlogfmt, and the sum of the three lines must equal the total number ofmixedlines counted in step 1. - Work out the true total number of lines in the
mixedstream whose status code is 500, write it in/root/lk-parsers/08-total.txtastotal=<정수>(the placeholder is the integer count), and on the next line write one sentence starting withreason=(at least 40 characters excluding spaces). Explain in your own words why a single parser cannot count this number.
Notes
- The working directory is
/root/lk-parsers. You start Loki yourself in step 1. - The data generator is
/opt/lab/d5/gen.pyand it uses theparsersdata. The grader does not read this file. - It is safer to wrap queries in backticks — inside double quotes the backslash is interpreted as an escape.
- Common mistake: measuring with
since=1h. As time passes the answer changes and you fail the re-grading. Givestartandendbased on the reference time inanchor.txt. - Common mistake: adding
| __error__=""out of habit. If you remove the broken lines before counting them, the very fact that formats are mixed becomes invisible. - Log queries and parsers · LogQL overview · Labels · HTTP API
Get hold of a stream with mixed formats
In /root/lk-parsers, start Loki, write date +%s in /root/lk-parsers/anchor.txt, and then load the data with python3 /opt/lab/d5/gen.py parsers "$(cat anchor.txt)". Then count how many lines {app="mixed"} and {app="orders"} each have over the past hour from the reference time, and write them in /root/lk-parsers/01-boot.txt on two lines as mixed=<정수> and orders=<정수> (the placeholders are the integer counts).
It takes about 20 seconds until /ready returns ready. Give start and end in nanoseconds based on the value in anchor.txt. You can count with the selector alone, without a parser.
When the format is clean, one logfmt line is enough
In the orders stream, count with the logfmt parser how many lines have status equal to 500. Write the query in /root/lk-parsers/02-logfmt.logql and the answer in /root/lk-parsers/02-logfmt.txt as lines=<정수> (the placeholder is the integer count).
You attach a parser after the selector with a pipe. A label created by the parser can be used after it as a label filter. This is different from looking for a string with a line filter like |= "status=500" — this time, use the value extracted by the parser.
The json parser reports failure
Attach the json parser to the mixed stream and count the lines that parsed successfully and the lines that received the __error__ label. Write the answer in /root/lk-parsers/03-json.txt on two lines as ok=<정수> and err=<정수> (the placeholders are the integer counts), and write one actual value of __error__ on one line in /root/lk-parsers/03-json-name.txt.
When a parser fails, the __error__ label is attached. You just select the lines where that label is empty and the lines where it is not empty. You can see what the label value is by looking directly at the stream in the result — __error_details__ is attached along with it.
Count the lines from which nothing was extracted, without any error
Count how many lines, when you attach the logfmt parser to the mixed stream, raised no error but did not get a status label created. Write the query in /root/lk-parsers/04-silent.logql and the answer in /root/lk-parsers/04-silent.txt as lines=<정수> (the placeholder is the integer count).
In LogQL, a "missing label" is compared as equal to an empty string. Using that property, you can select "lines where the parser could not create a value." This is the scariest number in this lab — it does not show up anywhere on the dashboard, yet it eats away at the total.
For lines with fixed positions, pattern is the right fit
Select only the lines that contain legacy in the body, extract the status code with the pattern parser, and count the lines where that value is 500. Write the query in /root/lk-parsers/05-pattern.logql and the answer in /root/lk-parsers/05-pattern.txt as lines=<정수> (the placeholder is the integer count). The query must include the pattern parser.
A legacy line has the shape legacy handler finished code 500 in 123ms. With pattern you write a name in angle brackets where you want to extract and write the rest literally. It is important to put a line filter before the parser so that only the legacy lines are passed on — lines of other formats do not have this shape.
Extract a number with the regular expression parser and compare it
From the same legacy lines, extract the elapsed time (in milliseconds) with the regexp parser and count the lines greater than 1000. Write the query in /root/lk-parsers/06-regexp.logql and the answer in /root/lk-parsers/06-regexp.txt as lines=<정수> (the placeholder is the integer count). The query must include the regexp parser.
regexp turns only named capture groups into labels — unnamed parentheses are discarded. The extracted value is a string, but if you compare it with an inequality, LogQL converts it to a number. You can get the same answer with pattern too — compare which one is easier to read.
Applied ① — build a table of how many lines each format has
Split the lines of the mixed stream into three formats, count them, and create /root/lk-parsers/tally.tsv. It has three lines with no header, and each line has two tab-separated columns, <형식><탭><줄수> (the placeholders are the format, a tab, and the line count). The format names are, in order, json, legacy, and logfmt, and the sum of the three lines must equal the total number of mixed lines counted in step 1.
A JSON line is a line where the json parser raises no error. A legacy line has legacy in its body. What remains are the logfmt lines, and for these lines logfmt creates status. Be sure to check that the sum of the three numbers matches the total — if it does not, you counted something twice or missed something somewhere.
Applied ② — record the true 5xx count and the fact behind it
Work out the true total number of lines in the mixed stream whose status code is 500, write it in /root/lk-parsers/08-total.txt as total=<정수> (the placeholder is the integer count), and on the next line write one sentence starting with reason= (at least 40 characters excluding spaces). Explain in your own words why a single parser cannot count this number.
Even if you chain two parsers in one query, it cannot read both formats together. Count separately for each format and add them up. If you think about where the "silently dropped lines" counted in step 4 went, you can see how many branches to split into. When you write the answer, also check what each of the three branch numbers was.