Loki — A Log Store That Does Not Index Logs
Two queries with the same answer: one took four minutes, the other two seconds
Goal
You read the stats.summary that comes with the Loki response to measure how many lines and bytes a query actually read, and confirm for yourself what reduces that number and what does not.
Why it matters
Loki's index holds only label combinations and the time ranges of chunks. The body is not indexed. So a query always runs in the same order — the selector picks the chunks to open, the time range keeps only those that overlap, and the rest are read in full and the filters are applied. Only the first two steps set the amount read; the later steps just throw away what has already been read. If you do not know this order, you end up putting a line filter first or lowering limit and waiting for the query to get faster. In a store where cost is charged by bytes, this difference is your bill.
Steps
- In
/root/lk-filters, start Loki, writedate +%sin/root/lk-filters/anchor.txt, and then load the data withpython3 /opt/lab/d5/gen.py filters "$(cat anchor.txt)". Then query{app=~"edge|cart|ship|batch"}over the past hour from the reference time, and write the number of lines and bytes read, taken from the response'sstats.summary, in/root/lk-filters/01-base.txton two lines asprocessed=<정수>andbytes=<정수>(the placeholders are the integer counts). - Over the same one-hour range, write in
/root/lk-filters/02-marker.logqla query that finds lines whose body containsSETTLE-LATEacross all four streams, and write the number of result lines and the number of lines read in/root/lk-filters/02-marker.txton two lines aslines=<정수>andprocessed=<정수>(the placeholders are the integer counts). - Write in
/root/lk-filters/03-stream.logqla query that gives the same answer (the same lines) while narrowing only the stream selector, and write the number of result lines and the number of lines read in/root/lk-filters/03-stream.txtaslines=<정수>andprocessed=<정수>(the placeholders are the integer counts). Then write the ratio of the lines read to step 2 in/root/lk-filters/03-ratio.txton one line asratio=<소수 둘째 자리>(the placeholder is the ratio to two decimal places). - Throw the same query as step 3 again, changing only the range to 30 minutes (1800 seconds back from the reference time). Write the number of result lines and the number of lines read in
/root/lk-filters/04-window.txtaslines=<정수>andprocessed=<정수>(the placeholders are the integer counts), and in/root/lk-filters/04-note.txtwrite one sentence starting withtradeoff=(at least 40 characters excluding spaces) describing in your own words what you gain and what you lose by narrowing the range. - Over all four streams and the one-hour range, throw two queries and compare the number of lines read. One is
{app=~"edge|cart|ship|batch"}with no filter at all, and the other adds two line filters to it. Write the results in/root/lk-filters/05-linefilter.txton four lines —plain_processed=,plain_lines=,filtered_processed=, andfiltered_lines=. - Throw a query over all four streams and the one-hour range with
| logfmt | status="500"attached, and write three numbers in/root/lk-filters/06-parser.txt—processed=<읽은 줄 수>,post_filter=<필터를 통과해 남은 줄 수>, andlines=<결과 줄 수>(the placeholders are the lines read, the lines remaining after passing the filter, and the result lines). Compare the three numbers with the step 1 baseline and see what stayed the same and what changed. - Organize what you have measured so far into
/root/lk-filters/cost.tsv. It has four lines with no header, and each line has three tab-separated columns,<방법><탭><읽은줄수><탭><결과줄수>(the placeholders are the method, a tab, the lines read, a tab, and the result lines). The method names are, in order,all(step 2),stream(step 3),window(step 4), andparser(step 6). - Write one query in
/root/lk-filters/08-budget.logql. There are three conditions — (1) when thrown over the one-hour range back from the reference time, it must return exactly the same lines as step 2, (2) the number of lines read must be at most one third of the step 1 baseline, and (3) do not rely onlimit. Then write the number of lines that query read in/root/lk-filters/08-budget.txton one line asprocessed=<정수>(the placeholder is the integer count).
Notes
- The working directory is
/root/lk-filters. You start Loki yourself in step 1. - The data generator is
/opt/lab/d5/gen.pyand it uses thefiltersdata. The grader does not read this file. - You view the statistics with
curl ... | jq '.data.stats.summary'. The fields this lab looks at aretotalLinesProcessed,totalBytesProcessed, andtotalPostFilterLines.execTimediffers every time you run even the same query, so do not use it as a basis for judgment. cost.sh, which the answer key creates, is a helper for convenience. You may also throw the queries directly withcurl.- Common mistake: measuring with
since=1h. As time passes the answer changes and you fail the re-grading. - Common mistake: missing that a cheaper query also changed the answer. Every time you cut the cost, check the number of result lines as well.
- LogQL overview · Log queries · Labels · Architecture · HTTP API
Load four streams and measure the baseline
In /root/lk-filters, start Loki, write date +%s in /root/lk-filters/anchor.txt, and then load the data with python3 /opt/lab/d5/gen.py filters "$(cat anchor.txt)". Then query {app=~"edge|cart|ship|batch"} over the past hour from the reference time, and write the number of lines and bytes read, taken from the response's stats.summary, in /root/lk-filters/01-base.txt on two lines as processed=<정수> and bytes=<정수> (the placeholders are the integer counts).
The statistics are in data.stats.summary of the query_range response — look at it whole once with jq '.data.stats.summary'. The two fields are totalLinesProcessed and totalBytesProcessed. This number is the baseline you compare against throughout this lab.
How many lines did we read to find five?
Over the same one-hour range, write in /root/lk-filters/02-marker.logql a query that finds lines whose body contains SETTLE-LATE across all four streams, and write the number of result lines and the number of lines read in /root/lk-filters/02-marker.txt on two lines as lines=<정수> and processed=<정수> (the placeholders are the integer counts).
One line filter is enough. What matters is not the answer but the amount read to get the answer — put the two numbers side by side. Also check whether the number of lines read equals the step 1 baseline.
Narrowing the selector reduces the amount read
Write in /root/lk-filters/03-stream.logql a query that gives the same answer (the same lines) while narrowing only the stream selector, and write the number of result lines and the number of lines read in /root/lk-filters/03-stream.txt as lines=<정수> and processed=<정수> (the placeholders are the integer counts). Then write the ratio of the lines read to step 2 in /root/lk-filters/03-ratio.txt on one line as ratio=<소수 둘째 자리> (the placeholder is the ratio to two decimal places).
You can tell which service the marker comes from by looking at the stream label in the step 2 result. Narrow the selector to that one. The number of result lines must stay the same — the point of this step is to reduce only the cost without changing the answer. The ratio is the step 3 lines read divided by the step 2 value.
Narrowing the range reduces the amount read, and the answer shrinks too
Throw the same query as step 3 again, changing only the range to 30 minutes (1800 seconds back from the reference time). Write the number of result lines and the number of lines read in /root/lk-filters/04-window.txt as lines=<정수> and processed=<정수> (the placeholders are the integer counts), and in /root/lk-filters/04-note.txt write one sentence starting with tradeoff= (at least 40 characters excluding spaces) describing in your own words what you gain and what you lose by narrowing the range.
Leave the query as it is and change only start. The number of lines read drops to about half, but the number of result lines drops with it — answers outside the range are not visible at all. Think about what is risky about narrowing the range first without knowing the incident time.
Adding line filters leaves the amount read unchanged
Over all four streams and the one-hour range, throw two queries and compare the number of lines read. One is {app=~"edge|cart|ship|batch"} with no filter at all, and the other adds two line filters to it. Write the results in /root/lk-filters/05-linefilter.txt on four lines — plain_processed=, plain_lines=, filtered_processed=, and filtered_lines=.
Any line filters will do (for example |= "level=error" and != "route=/a"). What you need to look at is the relationship among the four numbers — which pairs are equal and which differ. The equal side is "the amount read," and the differing side is "the amount remaining."
Attaching a parser leaves the amount read unchanged
Throw a query over all four streams and the one-hour range with | logfmt | status="500" attached, and write three numbers in /root/lk-filters/06-parser.txt — processed=<읽은 줄 수>, post_filter=<필터를 통과해 남은 줄 수>, and lines=<결과 줄 수> (the placeholders are the lines read, the lines remaining after passing the filter, and the result lines). Compare the three numbers with the step 1 baseline and see what stayed the same and what changed.
totalPostFilterLines is the third field of the statistics. cost.sh does not output this field, so look at it directly with curl ... | jq '.data.stats.summary' or edit the helper. A parser does more expensive work than a line filter, but by the standard of the amount read, the two are in the same place.
Applied ① — a cost table of four methods
Organize what you have measured so far into /root/lk-filters/cost.tsv. It has four lines with no header, and each line has three tab-separated columns, <방법><탭><읽은줄수><탭><결과줄수> (the placeholders are the method, a tab, the lines read, a tab, and the result lines). The method names are, in order, all (step 2), stream (step 3), window (step 4), and parser (step 6).
Just pull the values out of the files you made in the earlier steps and collect them. Once the table is built, it is visible at a glance — how many rows actually reduced the number of lines read.
Applied ② — a query that gives the same answer within a budget
Write one query in /root/lk-filters/08-budget.logql. There are three conditions — (1) when thrown over the one-hour range back from the reference time, it must return exactly the same lines as step 2, (2) the number of lines read must be at most one third of the step 1 baseline, and (3) do not rely on limit. Then write the number of lines that query read in /root/lk-filters/08-budget.txt on one line as processed=<정수> (the placeholder is the integer count).
There are only two levers that reduce the amount read. You cannot change the range in this step, so you must solve it with the remaining one. The stream label in the step 2 result tells you the answer. After changing the query, be sure to measure the number of result lines again — if it got cheaper but the answer changed, that is not tuning but an incident.