Loki — A Log Store That Does Not Index Logs
All the logs were there; I just did not know what to ask
Goal
You throw LogQL directly at the real Loki inside the Pod, pull out only the lines you want with a stream selector and line filters, and learn to read even the shape of the JSON that comes back.
Why it matters
Asking Loki a question splits into two layers. First the selector decides whose chunks of which streams to read, and then the line filter checks the lines that were read one by one. If you do not know this order, you will never be able to explain "why this query is fast and that one is slow." The body is not indexed, so a line filter reads everything the selector picked — that is why the first thing you do in an incident investigation is narrow the selector and the time range. limit and direction also change the result of an incident investigation. If you set limit without knowing the direction, you see only the last error instead of the most important first error.
Steps
- In
/root/lk-logql, start Loki, write the value ofdate +%son one line in/root/lk-logql/anchor.txt, and then load the data withpython3 /opt/lab/d5/gen.py logql "$(cat anchor.txt)". Then count how many streams Loki has now and write it on one line in/root/lk-logql/01-boot.txtasstreams=<정수>(the placeholder is the integer count). - Find how many lines over the past hour from the reference time have
appequal towebandlevelequal toerror. Write the query you used in/root/lk-logql/02-selector.logqland the line count in/root/lk-logql/02-selector.txton one line aslines=<정수>(the placeholder is the integer count). - Find how many lines over the same hour, across all three services (
web,api, andworker), containtimeoutin the body. Write the query in/root/lk-logql/03-line.logqland the answer in/root/lk-logql/03-line.txtaslines=<정수>(the placeholder is the integer count). - Across all three services, count the lines whose body contains
refusedand does not containrpc. Write the query in/root/lk-logql/04-chain.logqland the answer in/root/lk-logql/04-chain.txtaslines=<정수>(the placeholder is the integer count). - Across all three services, count the lines whose body contains the three-digit number 500 or 503, using a single regular expression line filter. Write the query in
/root/lk-logql/05-regex.logqland the answer in/root/lk-logql/05-regex.txtaslines=<정수>(the placeholder is the integer count). The query must include the regular expression line filter operator. - Query
{app="web",level="error"}twice over the same one-hour range withlimit=5. One query usesdirection=forwardand the otherdirection=backward. Write the results in/root/lk-logql/06-order.txton three lines —forward_first=<나노초 타임스탬프>,backward_first=<나노초 타임스탬프>, andtotal=<구간 전체 줄 수>(the placeholders are the nanosecond timestamp and the total number of lines in the range). - Run any log query once, look at the raw JSON, and write three lines in
/root/lk-logql/07-shape.txt. Afterresult_type=write the value ofdata.resultType, afterts_unit=write which time unit the first cell ofvaluesuses (one ofns,us,ms, ors), and afterentry_len=write as an integer how many cells long one element ofvaluesis. - Among the
webandapiservices, select only the lines whose body containstimeoutorrefused, and find the nanosecond timestamp of the earliest of those lines and the total number of lines. Write the query in/root/lk-logql/08-triage.logqland the answer in/root/lk-logql/08-triage.txton two lines aslines=<정수>andfirst_ts=<나노초>(the placeholders are the integer count and the nanosecond timestamp). The query must include a regular expression line filter, and it must not include worker.
Notes
- The working directory is
/root/lk-logql. Loki does not start automatically when the Pod comes up — you start it yourself in step 1. - The data generator is
/opt/lab/d5/gen.py. Run it with thelogqldata and the reference time argument. The grader does not read it, so there is no need to edit its contents. - You can also send queries with
logcli:LOKI_ADDR=http://localhost:3100 logcli query --limit=5 '{app="web"}'. However, the answers in this lab must be measured over an absolute range, so calling thequery_rangeAPI directly is more accurate. - Common mistake: measuring with
since=1h. As time passes the answer changes and you fail the re-grading. Givestartandendbased on the reference time inanchor.txt. - Common mistake: wrapping the line filter string in double quotes. In LogQL, backticks are safer — the backslash is not interpreted as an escape.
- LogQL overview · Log queries · HTTP API · Labels
Start Loki and load one hour of data instead of a whole day
In /root/lk-logql, start Loki, write the value of date +%s on one line in /root/lk-logql/anchor.txt, and then load the data with python3 /opt/lab/d5/gen.py logql "$(cat anchor.txt)". Then count how many streams Loki has now and write it on one line in /root/lk-logql/01-boot.txt as streams=<정수> (the placeholder is the integer count).
Copy /opt/lab/loki/loki.yaml and use it as the configuration. It takes about 20 seconds until /ready returns ready, so use a loop that checks the condition instead of a fixed sleep. /opt/lab/loki/streams.sh counts the streams for you — the number of streams is the number of distinct label combinations.
Decide where to look with the selector alone
Find how many lines over the past hour from the reference time have app equal to web and level equal to error. Write the query you used in /root/lk-logql/02-selector.logql and the line count in /root/lk-logql/02-selector.txt on one line as lines=<정수> (the placeholder is the integer count).
A stream selector is the label matchers inside the curly braces. If you join conditions with a comma, it selects only streams that satisfy both. Give the range as start and end in nanoseconds — multiply the value in anchor.txt by 1000000000 to get nanoseconds.
Dig through the body with a line filter
Find how many lines over the same hour, across all three services (web, api, and worker), contain timeout in the body. Write the query in /root/lk-logql/03-line.logql and the answer in /root/lk-logql/03-line.txt as lines=<정수> (the placeholder is the integer count).
A line filter comes after the selector. The operator that looks for a string as it is differs from the one that looks for a regular expression. The body is not indexed, so a line filter reads every line of the streams the selector picked and checks them one by one.
Chain filters to screen lines out
Across all three services, count the lines whose body contains refused and does not contain rpc. Write the query in /root/lk-logql/04-chain.logql and the answer in /root/lk-logql/04-chain.txt as lines=<정수> (the placeholder is the integer count).
You can chain several line filters, and they are applied in order from the left. There is a separate operator that means "does not contain." Do not try to merge the two conditions into one regular expression — chaining is easier to read for negation.
Regular expression line filter
Across all three services, count the lines whose body contains the three-digit number 500 or 503, using a single regular expression line filter. Write the query in /root/lk-logql/05-regex.logql and the answer in /root/lk-logql/05-regex.txt as lines=<정수> (the placeholder is the integer count). The query must include the regular expression line filter operator.
A regular expression line filter uses RE2 syntax. You write "one character can be any of several values" with square brackets. If you chain two |= filters, it becomes "lines that contain both" and the answer changes.
limit and direction — what gets cut off
Query {app="web",level="error"} twice over the same one-hour range with limit=5. One query uses direction=forward and the other direction=backward. Write the results in /root/lk-logql/06-order.txt on three lines — forward_first=<나노초 타임스탬프>, backward_first=<나노초 타임스탬프>, and total=<구간 전체 줄 수> (the placeholders are the nanosecond timestamp and the total number of lines in the range).
limit means "how many lines in the sort direction," not "how many lines from the start." If you change the direction, the same limit returns exactly the opposite lines. If you do not know this during an incident investigation, you miss the most important first error and see only the last one. To count the total number of lines, give a generous limit.
Read the JSON that comes back
Run any log query once, look at the raw JSON, and write three lines in /root/lk-logql/07-shape.txt. After result_type= write the value of data.resultType, after ts_unit= write which time unit the first cell of values uses (one of ns, us, ms, or s), and after entry_len= write as an integer how many cells long one element of values is.
Look at it as it is with curl ... | jq .. Log queries and metric queries have different resultType values. You can tell the unit by counting the digits of the timestamp — an epoch in seconds has ten digits. You need to know this shape to write automation scripts.
Applied — narrow the incident window with one query
Among the web and api services, select only the lines whose body contains timeout or refused, and find the nanosecond timestamp of the earliest of those lines and the total number of lines. Write the query in /root/lk-logql/08-triage.logql and the answer in /root/lk-logql/08-triage.txt on two lines as lines=<정수> and first_ts=<나노초> (the placeholders are the integer count and the nanosecond timestamp). The query must include a regular expression line filter, and it must not include worker.
Use together a matcher that picks only the two services in the selector and a regular expression that looks for either one of the two strings. To find the earliest line, put the sort direction forward, or sort the results you received yourself. This is the first move of an incident investigation — narrowing the range and pinning down the start time.