TT Lab
Get started
Learn Learning paths Courses

Loki — A Log Store That Does Not Index Logs

LogQL Has Three Parts

Continue in TT Lab

In one line

A LogQL query has three parts: stream selector → log filter → aggregation. Performance is decided almost entirely by the first part.

Why this was needed

Most reports that a Loki query is slow are not about query syntax but about not narrowing the range. If the curly braces at the front select 100,000 streams, nothing you do afterward will make it fast.

Three parts

sum by (pod) (
  rate({namespace="labhub-prod", app="labhub"} |= "error" | json | status >= 500 [5m])
)
 ^--집계--^        ^------- 스트림 셀렉터 -------^ ^------ 로그 필터 ------^

1) Stream selector {...} — picks candidates by indexed labels. There must be at least one, and an equality matcher (=) is much cheaper than a regular expression (=~).

2) Log filters |= != |~ !~ — scan the bodies of the selected streams. Order matters.

# 나쁨 — 파싱부터 하고 거른다
{app="labhub"} | json | level = "error"

# 좋음 — 문자열로 먼저 줄이고 파싱한다
{app="labhub"} |= "error" | json | level = "error"

|= is a simple substring check and very cheap. | json parses every line and is expensive. If you put the cheap one first, the number of lines to parse drops sharply.

3) Aggregation — rate, count_over_time, sum by, and so on. This is where logs become metrics.

# 초당 오류 수를 파드별로
sum by (pod) (rate({app="labhub"} |= "error" [5m]))

# 오류 메시지 상위 10개
topk(10, sum by (msg) (count_over_time({app="labhub"} | json [1h])))

Three parsers

Parser When to use
` json`
` logfmt`
` pattern`
` regexp`

Fields created by parsing are not indexed. So | status >= 500 is a scan, not an index lookup. If you use a value often, consider promoting it to a label at collection time, but calculate the cardinality first.

Common misconceptions

"Adding labels makes queries faster" — adding one label multiplies the number of streams by the number of values of that label. Putting status in as a label multiplies the streams by 5 (2xx/3xx/4xx/5xx/other), and putting user_id in multiplies them by the number of users. When there are many streams, every query gets slower.

"The |~ regular expression is convenient" — |= "error" or |= "warn" is much cheaper than |~ "error|warn". A regular expression runs the engine on every line.

Knowing the order a query runs in shows the cost

LogQL is a pipeline, and reducing early is always the cheapest.

{namespace="labhub-prod", app="api"}   ← 1. 색인으로 스트림을 고른다 (거의 공짜)
  |= "timeout"                          ← 2. 원문 문자열 필터 (싸다)
  | json                                ← 3. 파싱 (비싸다)
  | status >= 500                       ← 4. 파싱된 필드로 필터
  | line_format "{{.msg}}"              ← 5. 출력 성형

Changing the order changes the cost. If you do | json first and |= "timeout" afterward, you parse every line and then throw it away. Pulling string filters as far forward as possible is the first rule of LogQL tuning.

Turning logs into metrics — draw a graph from logs

Counting logs gives you a metric. This is especially useful in systems with no instrumentation.

# 5분간 5xx 비율
sum(rate({app="api"} |= "HTTP/1.1" 5" [5m]))
  /
sum(rate({app="api"} [5m]))

# 파싱한 필드로 p95 지연
quantile_over_time(0.95,
  {app="api"} | json | unwrap duration_ms [5m]) by (route)

But do not leave this on a permanent dashboard. A log scan is far more expensive than a time series lookup. Once the value is confirmed, put a real metric into the application, and use log-based queries only when investigating. A metric built from logs is a scouting tool for finding out whether "it is worth adding a metric."

Decide retention and compression in advance

Loki bundles logs into chunks and puts them in object storage. Two things must be decided here.

Keep queries used for alerts on a short range (5–15 minutes), and separate them from long-range queries for investigation. If you run the same query every minute over a 24-hour range, the cost quietly grows.

What really matters in practice

When a query is slow, the first thing to measure is the number of selected streams.

count(count by (pod, container) ({namespace="labhub-prod"}))

Grafana's query stats (Query inspector) also show Total bytes processed. If that value is in GB, the range is too wide. Shortening the time range or adding one more label has a bigger effect than query tuning.