TT Lab
Get started
Learn Learning paths Courses

Loki — A Log Store That Does Not Index Logs

If you never turn it on, logs stay forever

Continue in TT Lab

In one line

Loki starts out with retention turned off. If you do not turn it on, the logs you put in stay forever, and even when you turn it on, deletion is an asynchronous job the compactor does on its own schedule, so there is a time gap between "deleted" and "no longer visible."

Why this was needed

There was a team whose storage bill quadrupled in three months. When they opened an old range on the dashboard, logs from a year ago still came up. Nobody had configured retention, and nobody knew that — because the default is "do not delete."

There is an incident in the opposite direction too. A team that discovered personal information had been mixed in and filed a deletion request saw the lines vanish from queries right after the request and reported "deleted." A month later the storage size was the same. The default deletion mode (filter-and-delete) filters the lines out at query time the moment the request is accepted, but the actual removal from the objects is done by the compactor after the cancellation waiting period (default 24 hours) has passed. Not visible and deleted are different things.

How it works

Retention is done by the compactor. You have to look at three settings together.

Setting Default Meaning
compactor.retention_enabled false If off, nothing is deleted
compactor.compaction_interval 10m The interval at which compaction and retention run
compactor.retention_delete_delay 2h The time to wait between marking something for deletion and actual removal

The retention period itself is set with limits_config.retention_period, and if you want to apply it differently per stream, you write the selector, priority, and period in the retention_stream list. When several rules overlap, the one with the higher priority wins. Audit logs for a year and debug logs for a day — splitting them like this in the same cluster is the biggest lever for cutting cost.

Deleting specific lines is a different road from retention. If you give a selector and a range to POST /loki/api/v1/delete, a deletion request is accepted and its status becomes received. In the default deletion_mode, filter-and-delete, queries filter out those lines from that moment, so the query looks empty immediately. But the request can be canceled during delete_request_cancel_period (default 24 hours), and the actual removal from stored objects is the compactor's job after that time has passed.

What often goes wrong here is expectation versus guarantee. "It vanished from queries" is an indication, not proof that "it was deleted." If you have a regulatory deadline, you have to calculate that deadline together with the cancellation waiting period and the compactor interval, and all the more so if you are waiting for the storage size to shrink.

Capacity planning is actually simple. Measure the real bytes for one hour with totalBytesProcessed in the query response, multiply up to a day and a month, then multiply by each stream's retention period and add them up. Correct the compression ratio by comparing with the size of the chunk files actually stored. If you start from measurement rather than estimation, the discussion about shortening the retention period proceeds with numbers.

What it looks like in the field

The most frequent mistake is turning retention on and not running the compactor. In a microservices deployment the compactor is a separate component, and exactly one must be running in the cluster. If you write retention_enabled: true in the configuration and there is no compactor, nothing happens, yet from the configuration file alone it looks as if it were on.

The second is leaving out the priority of a stream rule. If two rules with the same priority apply to the same stream, it becomes hard to predict which one wins. The habit of stating the priority every time you add a rule is safe.

The third is "I deleted it but the size did not shrink." There is a waiting time between the deletion mark and the actual object deletion, and there may be yet another layer of lifecycle policy on the object storage side.

What you will do in the next lab

You write a configuration file that turns retention on and applies a different period per stream, and use loki -verify-config to check that the configuration itself is valid, not just its syntax. With that configuration you start a second Loki on a different port (you must change the gRPC port too), check with /config that the server actually came up holding those values, then submit a deletion request and record together that it is accepted and that it does not vanish immediately. Finally, you calculate a retention budget from the measured bytes and make a policy document.