TT Lab
Get started
Learn Learning paths Courses

Designing a Log Pipeline

What to index and what to leave in the body

Continue in TT Lab

Goal

The reason a log store gets slow or expensive is almost always that you chose wrongly what to index. That decision is very hard to undo after the data has piled up.

Here you count real logs and decide with numbers.

Materials

/opt/lab/logidx/events.jsonl   로그 2만 줄 (약 4MB)
mkdir -p /root/logidx && cp /opt/lab/logidx/* /root/logidx/ && cd /root/logidx
head -1 events.jsonl | python3 -m json.tool

What to leave behind

01-cardinality.txt  필드마다 값이 몇 가지인가
02-labels.txt       라벨과 본문을 가른 결과
03-explode.txt      라벨 하나를 더하면 스트림이 몇 개가 되나
04-shards.txt       샤드 수
05-tiers.txt        보관 단계와 기간
06-cost.txt         색인과 검색의 비용
07-notes.md         왜 그런지

Count first

Count how many distinct values each field (service level env pod route request_id user_id) has, write it in 01-cardinality.txt, and mark which are high and which are low.

python3 - <<'PY'
import json
rows = [json.loads(l) for l in open('events.jsonl')]
for k in ('service','level','env','pod','route','request_id','user_id'):
    print(k, len({r[k] for r in rows}))
PY

You will find a mix of two-digit numbers and numbers in the tens of thousands. That difference becomes the difference in cost.

Separate labels from the body

Separate what to keep as labels from what to keep in the body, and write it in 02-labels.txt. Also calculate the number of combinations of the ones you chose as labels.

A label is a stream (or a time series). The number of combinations becomes the count.

With the three service · level · env, it is 5 × 3 × 2, but if you count only the combinations that actually appear, it may be smaller. Count with the real data.

Keep high-cardinality fields in the body. Then you narrow down by label first and search within that, and that is the design intent of this kind of store.

Add just one

To the three labels, add pod · route · request_id · user_id one at a time, write in 03-explode.txt how many combinations you get each time, and write how many times larger it becomes with pod alone.

Calculate four times and compare. Each service has six pod values, so it multiplies.

"It would be convenient to have just this one label" is where incidents start. The convenience belongs to one person, and the cost belongs to the whole cluster.

How many shards to split into

Assuming these logs arrive at 500 lines per second, calculate a day's storage and the number of shards in 04-shards.txt. Decide a target shard size, and also write what to do if a day's data is smaller than that.

You get the size of one line as file size ÷ number of lines.

Aim for 10–50GB per shard. If there are thousands of small shards, they consume that much memory and file handles and the cluster slows down.

If a day's data is smaller than the target, leave it to weekly indexes or a rollover condition (by size) instead of daily ones.

When to throw what away

Decide the period of each of the four stages, hot, warm, cold, and delete, and calculate in 05-tiers.txt how much accumulates in each stage. Also write what changes in warm or cold.

In hot, both writes and searches happen, and warm is read-only, so you can merge segments and reduce replicas. Cold moves to slower storage or object storage.

If you don't decide, one day the disk fills up, and then you delete in a hurry and end up deleting what you need too.

Where the cost comes from

Separate the indexing-side cost from the search-side cost and write it in 06-cost.txt, and include how to reduce the indexing cost and how to handle fields you don't search.

Documents per second is CPU. If you send debug-level logs to production as they are, you pay that cost every day.

For a field that is only viewed and never searched, setting index: false reduces indexing cost and storage. Conversely, a number used only for aggregation needs only doc_values.

Another way is to move toward using traces for normal requests and keeping logs only for unusual cases.

For the next person to read this

Pick four or more of the things you saw here and organize them in 07-notes.md. Write why it is so, not what you did.

Imagine it is read by you, when you are wondering whether to put a new field into the labels. "I counted the cardinality" does not help, but "if you put pod in the labels, the streams become six times as many — the convenience belongs to one person and the cost belongs to the whole cluster" does.