Loki — A Log Store That Does Not Index Logs
One Label Makes It Thirty Times More Expensive
Goal
Loki does not index log content. It indexes only labels. That is why it is cheap, and why it blows up if you choose labels badly.
In this lab you blow it up yourself instead of just hearing about it.
Getting started
cp -r /opt/lab/loki/* . && chmod +x *.sh
setsid nohup loki -config.file=loki.yaml > loki.log 2>&1 </dev/null &
curl -s localhost:3100/ready # "ready" 가 될 때까지 20초쯤
export LOKI_ADDR=http://localhost:3100
Add setsid nohup. If you start it with just &, it dies along with the shell when the shell changes.
Tools
| File | What it does |
|---|---|
push.sh |
Push one log line — the first argument is the label JSON |
streams.sh |
The current number of streams (with -v, the combinations too) |
logcli |
LogQL queries |
LogQL starts with label selection
{app="web", level="error"} |= "timeout" | logfmt | status="500"
└── 색인으로 좁힌다 ──┘ └── 여기부터는 훑어 읽는다 ──────────┘
The curly braces at the front set the range to read, and the rest reads and filters within that range. If you do not narrow the range, it reads everything.
Steps
- Start it and push one line →
01-boot.txt - Label combination = stream →
02-streams.txt - Content is not indexed →
03-notindexed.md - Cardinality explosion →
04-explode.txt - The same information in the content →
05-fix.txt - Extract with a parser →
06-parser.txt - The criterion for labels →
07-rule.md - Wrap-up →
08-notes.md
Notes
The difference in numbers between steps 4 and 5 is the whole point of this lab. Depending on where you put the same information, the number of streams rises by 30 or by 1.
Start Loki and push one line
Copy /opt/lab/loki/ to start Loki, push one log line, then retrieve that line again and save it in 01-boot.txt.
cp -r /opt/lab/loki/* . && chmod +x *.sh
setsid nohup loki -config.file=loki.yaml > loki.log 2>&1 </dev/null &
curl -s localhost:3100/ready # ready 가 될 때까지 20초쯤 걸린다
./push.sh '{"app":"web"}' 'GET /health 200'
export LOKI_ADDR=http://localhost:3100
logcli query --limit=5 --since=1h '{app="web"}'
As in {app="web"}, what is inside the curly braces is the label selection. LogQL starts here.
One label combination is one stream
Push several log lines with different label combinations, and save in 02-streams.txt that the number of streams rises by the number of combinations.
You can see the current combinations with ./streams.sh -v. If you push 2 kinds of app × 2 kinds of level, you get 4 streams.
This number decides most of the cost of Loki. Every stream gets its own chunks and index entries.
Content is not indexed
Find logs by a string that is not in the labels (|=), and write in 03-notindexed.md that this is not an index lookup but a scan.
logcli query --since=1h '{app="web"} |= "500"'. It works — but it finds the lines by reading everything inside the range narrowed down by {app="web"}.
That is why LogQL must always start with label selection. {app=~".+"} |= "500" scans everything. This is why Loki is cheaper than Elasticsearch, and why label design matters.
Blow up the cardinality on purpose
Push 30 log lines with request_id as a label, and save in 04-explode.txt what happens to the number of streams.
for i in $(seq 30); do ./push.sh "{\"app\":\"bad\",\"request_id\":\"r$i\"}" "req $i"; done
./streams.sh
Record the counts both before and after pushing. Thirty lines add 30 streams — one stream per line.
Put the same information in the content
Push the same 30 request_id values again, this time in the content, and save in 05-fix.txt how many streams are added this time.
Keep a single label, {"app":"good"}, and write the line content like request_id=r1 status=200 dur=12ms. Only 1 stream is added.
One line of difference makes a factor of 30. And you lose nothing — in the next step you extract it with a parser as it is.
Extract values from the content
Use the | logfmt parser to filter on status inside the content, find only the lines you want, and save the result in 06-parser.txt.
logcli query --since=1h '{app="good"} | logfmt | status="200"'. The parser runs at query time, so it does not add to the index.
This is the key point — even without putting it in a label, your filtering ability stays the same. All you lose is "narrowing instantly by index," and what you gain is not suffering a stream explosion.
Decide what may go into a label
In 07-rule.md, write 3 things that are fine to use as labels and 3 things that must not be used, with reasons.
There is one criterion — does the number of distinct values stay flat over time. app, env, and level are bounded. user_id, request_id, trace_id, ip, and url grow without limit.
The mistake made most often in practice is pod or container_id — every deployment creates new values, so it blows up slowly.
Summarize the three points
Write at least three lines in 08-notes.md: what a stream is, how a content filter differs from an index, and the criterion for labels.
The text must include 스트림, 색인, and 가짓수 (the Korean words for stream, index, and number of distinct values).