TT Lab
Get started
Learn Learning paths Courses

Loki — A Log Store That Does Not Index Logs

The push returns 204 but the query comes back empty

Continue in TT Lab

In one line

A line you put into Loki first piles up in the ingester's memory and later goes down to storage as a chunk. A query asks the ingester only about "the last few hours," so a line put in with a timestamp older than that is accepted but invisible.

Why this was needed

A team was in the middle of incident recovery. The collector had been stopped for two hours, and they recovered the files that had piled up in the meantime and pushed them into Loki. Every push returned 204. But when they opened that time period in Grafana, there was nothing.

People thought "the push is lying." In reality the lines were in the ingester just fine, and the query simply did not ask the ingester about that range. When they forced a flush to bring the chunks down to storage, the same query returned everything.

How it works

The write path goes like this.

  1. When a push arrives, the ingester appends the line to the stream's open chunk. It also writes the same content to the WAL so it can be revived even if the ingester dies.
  2. A chunk is closed when it meets one of three conditions — it reaches the target size (chunk_target_size), no new lines arrive for a set time (chunk_idle_period), or it has been open too long (max_chunk_age).
  3. A closed chunk is uploaded to object storage, and the index records that stream and time range.

The read path has two branches. A query asks storage, and if the range is recent, it also asks the ingester for the lines that have not yet gone down to storage. What decides "recent" here is querier.query_ingesters_within, and the default is 3 hours. If the end of the range is further in the past than that, the query skips the ingester.

So a line put in with a past timestamp falls into the gap between these two paths. It is not yet in storage, and it is in the ingester but nobody asks. As time passes and the chunk naturally closes and is uploaded, it starts to appear — that is where the story "it showed up by itself a long time later" comes from.

Loki's write path going from a push through the ingester's open chunk to object storage, and the read path that always asks storage but asks the ingester only when the range is within the last 3 hours, so a line put in with a timestamp from five hours ago falls into the gap between the two paths.

This design is not a mistake. Asking the ingester is expensive, and asking about old ranges every time would slow down every query. But in recovery work that pushes data in late, this assumption breaks.

There are three things to remember in operations. First, if you pushed in past data for recovery, wait for the flush or force one. Second, data that is too old may be rejected in the first place — reject_old_samples and reject_old_samples_max_age are the levers for that (this lab's configuration deliberately turns them off). Third, the Loki in this Pod is in single-binary mode, with all components in one process. If the ingester, querier, and compactor run separately in production, the same principle applies as it is, but the place where you call the flush and the metrics to watch differ.

What it looks like in the field

The symptom you see most often is "the dashboard's last 15 minutes look fine but yesterday's range is empty." The recent range is answered by the ingester and the past range by storage, and if the road up to storage is blocked, it takes exactly this shape. It always appears like this in incidents where the object storage credentials expired.

The second is the empty range right after an ingester restart. Until the WAL replay finishes, that range looks empty, and once the replay finishes it fills in. So if a "logs disappeared" report comes in during a deployment that cycles the ingesters, waiting a few minutes first is the right move.

The third is chunk size tuning. If you set the target size too small, objects multiply endlessly and the index and queries slow down, and if you set it too large, the ingester's memory grows and there is more to lose on restart.

What you will do in the next lab

You put recent data and data from five hours ago into the Pod's Loki, and see for yourself that the second is accepted but empty in queries. You find and confirm the setting that decides that boundary in the server's /config, and record as numbers how the number of chunk files in storage and the query result change before and after a forced flush. Finally, you go as far as putting in a line from six hours ago yourself and making it visible.