Loki — A Log Store That Does Not Index Logs
It vanished from the screen immediately, yet storage stayed flat for a month
Goal
You write and verify a configuration that turns retention on and applies a different period per stream, start a second Loki with that configuration, submit a deletion request, and confirm in numbers the time gap between acceptance and actual deletion.
Why it matters
Loki starts out with retention turned off. If nobody turns it on, the logs you put in stay forever, and you learn of it only after the bill has quadrupled. Deletion requests are two-layered in the same way — in the default mode, queries filter out those lines the moment the request is accepted, but the actual removal from stored objects is the compactor's job after the cancellation waiting period has passed. If you read "not visible" as "deleted," your regulatory compliance report will be wrong and your capacity plan will be off. Deciding the retention period must start from measurement, not estimation — the number of bytes that comes with the query response is that starting point.
Steps
- In
/root/lk-retention, start Loki, writedate +%sin/root/lk-retention/anchor.txt, and then load the data withpython3 /opt/lab/d5/gen.py retention "$(cat anchor.txt)". Then query the two streamsauditanddebugeach over a one-hour range, and write the bytes read from the response statistics in/root/lk-retention/01-bytes.txton two lines asaudit_bytes=<정수>anddebug_bytes=<정수>(the placeholders are the integer counts). - From the one-hour bytes in step 1, calculate the storage per stream and create
/root/lk-retention/budget.tsv. It has two lines with no header, and each line has four tab-separated columns,<스트림><탭><한시간바이트><탭><보존일수><탭><보존기간총바이트>(the placeholders are the stream, a tab, the one-hour bytes, a tab, the retention days, a tab, and the total bytes over the retention period). The streams are, in order,auditanddebug, and the retention days are taken as365and1respectively. The total bytes are한시간바이트 × 24 × 보존일수(the one-hour bytes times 24 times the retention days). - Write
/root/lk-retention/loki-ret.yaml. Base it on theloki.yamlyou are using now and add the following —server.http_listen_portset to3200,server.grpc_listen_portset to9096, the storage path under/tmp/lokiret,compactor.retention_enabledset totrue,compactor.delete_request_storeset tofilesystem,compactor.working_directoryspecified, andlimits_config.retention_periodset to720h. Andloki -config.file=/root/lk-retention/loki-ret.yaml -verify-configmust finish without any error. - Add a
retention_streamlist tolimits_configin/root/lk-retention/loki-ret.yamlso that{app="debug"}is kept for24hand{app="audit"}for8760h. State the priorities explicitly as1and2respectively. Then write two lines in/root/lk-retention/04-rules.txt—rules=<규칙 개수>andverify=ok(when-verify-configpassed) (the placeholder is the number of rules). - Start a second Loki with
/root/lk-retention/loki-ret.yamland wait untilhttp://localhost:3200/readyreturnsready. Then read three values from that server's/configand write them in/root/lk-retention/05-run.txt—retention_enabled=<값>,retention_period=<값>, andcompaction_interval=<값>(the placeholders are the values). - Push a few lines into the Loki on port 3200, and then request that a range of one of those streams be deleted. You make the request by giving
query,start, andendtoPOST /loki/api/v1/delete. Then write three lines in/root/lk-retention/06-delete.txt—post_code=<HTTP 상태 코드>,requests=<목록에 있는 요청 수>, andstatus=<그 요청의 status 값>(the placeholders are the HTTP status code, the number of requests in the list, and the status value of that request). - Query again the range you asked to delete and count how many lines come out now, and find and confirm in the server configuration when the actual deletion happens. Write four lines in
/root/lk-retention/07-async.txt—still_visible=<정수>,mode=<deletion_mode 값>,param=<실제 삭제까지 기다리는 시간을 정하는 설정 이름>, andvalue=<서버가 찍은 값 그대로>(the placeholders are the integer count, the deletion_mode value, the name of the setting that decides the wait before actual deletion, and the value exactly as the server printed it). - Write five lines in
/root/lk-retention/policy.txt—audit_period=anddebug_period=(both exactly the values you wrote in the configuration),audit_bytes_total=anddebug_bytes_total=(the last column of the step 2 table), andnote=(at least 60 characters excluding spaces, stating together why you set the two periods differently and the fact that a deletion request is not reflected immediately).
Notes
- The working directory is
/root/lk-retention. You start the first Loki in step 1 and the second in step 5. - The data generator is
/opt/lab/d5/gen.pyand it uses theretentiondata. The grader does not read this file. - In this lab you cannot see retention deletion actually happen. The compactor runs on a default 10-minute interval and deletion requests wait 24 hours by default. What this lab confirms is whether the configuration is valid, whether the server came up with those values, whether the request is accepted, and the fact that all of this is asynchronous.
- The second Loki dies if you change only
http_listen_port— you must changegrpc_listen_porttoo. The storage paths (path_prefixandstorage_config) must also not overlap with the first. - The
startandendof a deletion request are integers in seconds. This differs from the nanoseconds of the query API. - Retention (Compactor) · Log deletion requests · Configuration docs · HTTP API
Load two kinds of logs and measure the real bytes
In /root/lk-retention, start Loki, write date +%s in /root/lk-retention/anchor.txt, and then load the data with python3 /opt/lab/d5/gen.py retention "$(cat anchor.txt)". Then query the two streams audit and debug each over a one-hour range, and write the bytes read from the response statistics in /root/lk-retention/01-bytes.txt on two lines as audit_bytes=<정수> and debug_bytes=<정수> (the placeholders are the integer counts).
The statistic is data.stats.summary.totalBytesProcessed of the query_range response. The two streams differ in character — one is an audit record that accumulates rarely, and the other is a debug log that pours out several lines per second. Capacity planning starts from this difference.
Calculate the retention budget from measurements
From the one-hour bytes in step 1, calculate the storage per stream and create /root/lk-retention/budget.tsv. It has two lines with no header, and each line has four tab-separated columns, <스트림><탭><한시간바이트><탭><보존일수><탭><보존기간총바이트> (the placeholders are the stream, a tab, the one-hour bytes, a tab, the retention days, a tab, and the total bytes over the retention period). The streams are, in order, audit and debug, and the retention days are taken as 365 and 1 respectively. The total bytes are 한시간바이트 × 24 × 보존일수 (the one-hour bytes times 24 times the retention days).
The point of this calculation is to put the last columns of the two lines side by side. If the side with ten times as many lines has a short retention period, the totals can flip — that reversal is the reason to split the retention policy per stream.
Write and verify the configuration that turns retention on
Write /root/lk-retention/loki-ret.yaml. Base it on the loki.yaml you are using now and add the following — server.http_listen_port set to 3200, server.grpc_listen_port set to 9096, the storage path under /tmp/lokiret, compactor.retention_enabled set to true, compactor.delete_request_store set to filesystem, compactor.working_directory specified, and limits_config.retention_period set to 720h. And loki -config.file=/root/lk-retention/loki-ret.yaml -verify-config must finish without any error.
-verify-config checks not only the syntax but also whether the combination of settings makes sense — if it passes, there is no output. The reason you change both ports will become clear in the next steps. If you do not change the storage path, two processes will use the same directory as the Loki that is running now.
A different retention period per stream
Add a retention_stream list to limits_config in /root/lk-retention/loki-ret.yaml so that {app="debug"} is kept for 24h and {app="audit"} for 8760h. State the priorities explicitly as 1 and 2 respectively. Then write two lines in /root/lk-retention/04-rules.txt — rules=<규칙 개수> and verify=ok (when -verify-config passed) (the placeholder is the number of rules).
retention_stream is a list of entries with three fields, selector, priority, and period. The selector is exactly the LogQL stream selector syntax, put inside quotes. If you do not write the priority, it becomes hard to predict which side wins among overlapping rules.
Start the second Loki — there are two ports
Start a second Loki with /root/lk-retention/loki-ret.yaml and wait until http://localhost:3200/ready returns ready. Then read three values from that server's /config and write them in /root/lk-retention/05-run.txt — retention_enabled=<값>, retention_period=<값>, and compaction_interval=<값> (the placeholders are the values).
Keep the log in a file (> ret.log 2>&1). If it dies right away, look at the last line of that file — the error that appears when you change only one port is written there as it is. /ready takes about 20 seconds, so use a loop that checks the condition instead of a fixed sleep.
Submit a deletion request and look at the list
Push a few lines into the Loki on port 3200, and then request that a range of one of those streams be deleted. You make the request by giving query, start, and end to POST /loki/api/v1/delete. Then write three lines in /root/lk-retention/06-delete.txt — post_code=<HTTP 상태 코드>, requests=<목록에 있는 요청 수>, and status=<그 요청의 status 값> (the placeholders are the HTTP status code, the number of requests in the list, and the status value of that request).
Give the range as integers in seconds (not nanoseconds). You ask for the list with a GET on the same path. If the request is rejected with 400, check whether delete_request_store is set and whether the range is backwards.
Applied ① — not visible and deleted are different things
Query again the range you asked to delete and count how many lines come out now, and find and confirm in the server configuration when the actual deletion happens. Write four lines in /root/lk-retention/07-async.txt — still_visible=<정수>, mode=<deletion_mode 값>, param=<실제 삭제까지 기다리는 시간을 정하는 설정 이름>, and value=<서버가 찍은 값 그대로> (the placeholders are the integer count, the deletion_mode value, the name of the setting that decides the wait before actual deletion, and the value exactly as the server printed it).
A query result of 0 does not mean it was deleted from storage. In /config, find the setting that means the deletion mode, and the "period during which a request can be canceled" in the compactor: block. If you put the two values side by side, you get the answer to "why isn't the size shrinking."
Applied ② — pin down the retention policy in a document
Write five lines in /root/lk-retention/policy.txt — audit_period= and debug_period= (both exactly the values you wrote in the configuration), audit_bytes_total= and debug_bytes_total= (the last column of the step 2 table), and note= (at least 60 characters excluding spaces, stating together why you set the two periods differently and the fact that a deletion request is not reflected immediately).
If the values in a policy document disagree with the configuration file, that document is worse than none. Take them from the configuration as they are. The last line is a sentence that the next person to take over this cluster will read — if you write down the basis of the numbers and the time delay together, you will not be asked the same question twice.