Loki — A Log Store That Does Not Index Logs
I measured p95 from logs and got the same number as p50
Goal
You throw the two kinds of LogQL metric queries yourself to turn logs into numbers, reproduce the phenomenon where extracted labels split the series and make quantiles meaningless, and then fix it.
Why it matters
When you urgently need the latency or error rate of a service that has no metrics, the logs are already there. LogQL metric queries turn those logs into numbers, but there is one trap — the labels that decide the identity of a series include even the labels the parser created in that query. If even one field with a value that differs on every line is extracted, the series split into as many pieces as there are lines, and since each series has only one sample, whichever quantile you ask for gives the same value. There is no error and no warning, and only the graph draws fine, so it goes unnoticed for months. And log-based metrics read the raw data every time you query, so you need the habit of measuring the amount read once before putting one on a dashboard.
Steps
- In
/root/lk-metrics, start Loki, writedate +%sin/root/lk-metrics/anchor.txt, and then load the data withpython3 /opt/lab/d5/gen.py metrics "$(cat anchor.txt)". Then give the reference time astime, throwsum by (app) (count_over_time({app=~"checkout|search"}[1h])), and write the line counts of the two services in/root/lk-metrics/01-boot.txton two lines ascheckout=<정수>andsearch=<정수>(the placeholders are the integer counts). - Find the per-second rate of lines with
statusequal to 500 in thecheckoutservice over a one-hour range. Write the query in/root/lk-metrics/02-rate.logqland the value in/root/lk-metrics/02-rate.txton one line asrate=<소수 여섯째 자리>(the placeholder is the value to six decimal places). Wrap the outside insum(...)so that the result comes out as one series. - Find the one-hour error ratio of
checkout(lines equal to 500 ÷ total lines) and write it in/root/lk-metrics/03-ratio.txton one line asratio=<소수 여섯째 자리>(the placeholder is the value to six decimal places). Write the query in/root/lk-metrics/03-ratio.logql. The series line up only if you wrap the numerator and the denominator each insum(...)before dividing. - Apply
quantile_over_time(0.50, ...)andquantile_over_time(0.95, ...)to{app="checkout"} | logfmt | unwrap dur_ms [1h]and throw each one. Write the number of series the two queries returned and the value of the first series in/root/lk-metrics/04-trap.txton four lines —series=<정수>,p50_first=<숫자>,p95_first=<숫자>, andlines=<그 구간의 전체 줄 수>(the placeholders are the integer count, the numbers, and the total number of lines in that range). - Fix the same two quantiles so that they come out as one series and throw them again. Write the queries in
/root/lk-metrics/05-p50.logqland/root/lk-metrics/05-p95.logqlrespectively, and write the values in/root/lk-metrics/05-fix.txton three lines asp50=<숫자>,p95=<숫자>, andseries=<정수>(the placeholders are the numbers and the integer count). The two values must clearly separate. - Create
/root/lk-metrics/compare.tsv. It has two lines with no header, and each line has four tab-separated columns,<서비스><탭><줄수><탭><오류비율><탭><p95>(the placeholders are the service, a tab, the line count, a tab, the error ratio, a tab, and the p95). The services are, in order,checkoutandsearch; write the error ratio to six decimal places and p95 as an integer rounded with no decimals. - Throw the step 5 p95 query over a one-hour range with
query_range, measure the bytes read in the response statistics, and write three lines in/root/lk-metrics/07-cost.txt—bytes_per_query=<정수>,refresh_sec=30, andbytes_per_day=<정수>(the placeholders are the integer counts). Compute the day's total asbytes_per_query × 2880, assuming it runs once every 30 seconds. - Write three lines in
/root/lk-metrics/08-decide.txt. Afterchoice=write eitherlogormetric, afterevidence=write one line containing two or more numbers you measured in earlier steps, and afterreason=write why you chose it, in at least 60 characters excluding spaces. There is more than one correct answer, but the evidence must be the measured values from earlier steps.
Notes
- The working directory is
/root/lk-metrics. You start Loki yourself in step 1. - The data generator is
/opt/lab/d5/gen.pyand it uses themetricsdata. The grader does not read this file. - You throw metric queries to
/loki/api/v1/querywithtime, and if you need values per interval, you givestepto/loki/api/v1/query_range.mq.sh, which the answer key creates, is a helper for convenience. - Common mistake: measuring with
since=1hor the current time. Give the reference time inanchor.txtastime. - Common mistake: the label sets of the numerator and denominator differ, so the division gives an empty result. If you wrap both sides in
sum(...), all the labels are erased and they line up. - Metric queries · Log queries · LogQL overview · HTTP API
Load the logs of two services and count lines first
In /root/lk-metrics, start Loki, write date +%s in /root/lk-metrics/anchor.txt, and then load the data with python3 /opt/lab/d5/gen.py metrics "$(cat anchor.txt)". Then give the reference time as time, throw sum by (app) (count_over_time({app=~"checkout|search"}[1h])), and write the line counts of the two services in /root/lk-metrics/01-boot.txt on two lines as checkout=<정수> and search=<정수> (the placeholders are the integer counts).
Throw metric queries not to query_range but to /loki/api/v1/query, and give time in nanoseconds. Also look once at how the response's resultType differs from a log query. sum by (app) groups the series by service.
A query that counts lines — errors per second
Find the per-second rate of lines with status equal to 500 in the checkout service over a one-hour range. Write the query in /root/lk-metrics/02-rate.logql and the value in /root/lk-metrics/02-rate.txt on one line as rate=<소수 여섯째 자리> (the placeholder is the value to six decimal places). Wrap the outside in sum(...) so that the result comes out as one series.
rate divides the number of lines in the range by the seconds of the range. To count only 500, you have to extract the status code with a parser and apply a label filter. It is normal for the value to come out very small — it is a few requests per hour.
A ratio divides two metric queries
Find the one-hour error ratio of checkout (lines equal to 500 ÷ total lines) and write it in /root/lk-metrics/03-ratio.txt on one line as ratio=<소수 여섯째 자리> (the placeholder is the value to six decimal places). Write the query in /root/lk-metrics/03-ratio.logql. The series line up only if you wrap the numerator and the denominator each in sum(...) before dividing.
The denominator needs no parser — just count all the lines. If the label sets of the numerator and denominator differ, the division gives an empty result, so the simplest way is to wrap both sides in sum(...) and erase all the labels.
p50 and p95 come out identical
Apply quantile_over_time(0.50, ...) and quantile_over_time(0.95, ...) to {app="checkout"} | logfmt | unwrap dur_ms [1h] and throw each one. Write the number of series the two queries returned and the value of the first series in /root/lk-metrics/04-trap.txt on four lines — series=<정수>, p50_first=<숫자>, p95_first=<숫자>, and lines=<그 구간의 전체 줄 수> (the placeholders are the integer count, the numbers, and the total number of lines in that range).
The number of series is the length of the response's data.result array. Compare that number with the line count from step 1. Why the two quantiles give the same value is visible right in that comparison. Also check in the result's metric which labels logfmt created from this log.
Tidy the labels and the quantiles separate
Fix the same two quantiles so that they come out as one series and throw them again. Write the queries in /root/lk-metrics/05-p50.logql and /root/lk-metrics/05-p95.logql respectively, and write the values in /root/lk-metrics/05-fix.txt on three lines as p50=<숫자>, p95=<숫자>, and series=<정수> (the placeholders are the numbers and the integer count). The two values must clearly separate.
You can strip off the labels whose values differ on every line, keep only the labels you will use, or wrap the outside in an aggregation. Any of the three methods is fine — but the series must become one. You can tell which label is the culprit by looking at the metric in the step 4 result.
Compare the two services in one table
Create /root/lk-metrics/compare.tsv. It has two lines with no header, and each line has four tab-separated columns, <서비스><탭><줄수><탭><오류비율><탭><p95> (the placeholders are the service, a tab, the line count, a tab, the error ratio, a tab, and the p95). The services are, in order, checkout and search; write the error ratio to six decimal places and p95 as an integer rounded with no decimals.
Just change the service name in the queries from steps 3 and 5. Confirm in the table that, even on the same data, the tails of the two services have different shapes — it is a difference you see only by looking at p95, not the average.
Applied ① — how much does this query read if you hang it on a panel?
Throw the step 5 p95 query over a one-hour range with query_range, measure the bytes read in the response statistics, and write three lines in /root/lk-metrics/07-cost.txt — bytes_per_query=<정수>, refresh_sec=30, and bytes_per_day=<정수> (the placeholders are the integer counts). Compute the day's total as bytes_per_query × 2880, assuming it runs once every 30 seconds.
You can also throw metric queries with query_range, and then you give step. The statistics are in the same place as for log queries (data.stats.summary). 2880 a day is the number of runs per day at a 30-second interval — it means a single dashboard panel queries far more often than a person.
Applied ② — keep measuring from logs, or export as a metric?
Write three lines in /root/lk-metrics/08-decide.txt. After choice= write either log or metric, after evidence= write one line containing two or more numbers you measured in earlier steps, and after reason= write why you chose it, in at least 60 characters excluding spaces. There is more than one correct answer, but the evidence must be the measured values from earlier steps.
The two paths have different values. The log-based path needs no instrumentation fix and lets you look back at the past, but reads the raw data on every query. The metric-based path is cheap and fast, but must be exported in advance and has no past. Using the amount read per day from step 7 and the trap you met in steps 4 and 5 as your evidence will make your case persuasive.