TT Lab
Get started
Learn Learning paths Courses

Loki — A Log Store That Does Not Index Logs

Fix a 400, wait out a 429

Continue in TT Lab

In one line

There are two kinds of rejection on the write side. A 400 must be fixed and a 429 must be waited out. A collector that handles the two with the same code mishandles both.

Why this was needed

On the day a new service was hooked up, rejections poured into the collector's log. The person in charge turned on retry logic and went home. The next morning the rejections were unchanged, and instead the collector's buffer was full and pushing back the logs of other services as well.

The rejection was a 400. The label name contained a hyphen, and that line is rejected forever no matter how many times it is resent. Retrying did not fix the situation and only filled the queue. Conversely, the 429 that came from another service the same week was a case where retrying was the right answer, but that side was just discarding without retrying.

How it works

There are roughly four reasons Loki rejects an incoming request.

Situation Code What to do
The label name does not fit the syntax 400 Fix the sender
The number of labels exceeds the limit (max_label_names_per_series) 400 Reduce the labels
The line is too long (max_line_size) 400 Truncate or split it
The per-stream rate is exceeded (per_stream_rate_limit) 429 Retry after a backoff

Label names follow the same rules as Prometheus — start with a letter or an underscore and use only letters, digits, and underscores. Hyphens are not allowed, and they cannot start with a digit. Many people hit this wall trying to carry Kubernetes labels over as they are, so it is better to put in advance a setting on the collector side that renames them.

Rate limits are per stream. So even for the same service, if there are several label combinations, each combination gets its own limit separately. Conversely, if you reduce labels and merge streams, traffic concentrates on one stream and you get 429s. Lowering cardinality and avoiding rate limits pull in opposite directions — if you do not know this tension, you fix one side and blow up the other.

Bursts add another layer to this picture. per_stream_rate_limit_burst is the amount that may momentarily exceed the limit. So traffic that is quiet and then surges passes for the first while and receives 429s after that. This is usually what lies behind the symptom "it worked at first and then suddenly didn't."

A correct client, when it gets a 429, has three things — exponential backoff, jitter, and a retry cap. Without jitter, many collectors come back in the same rhythm and hit the same wall together. Without a cap, if you mistake a 400 for a 429, you retry forever.

What it looks like in the field

The most common incident is the "retry on 400" above. The fix is to state the retry conditions by status code in the collector configuration and send 400-class responses to a quarantine queue where a person looks at them.

The second is long lines. If a single stack trace exceeds the limit, the whole request is rejected and even the healthy lines in the same batch fall with it. So truncate at the collector first, or keep batches small to reduce the blast radius.

The third is the number of labels. If you move all the Kubernetes metadata over as labels, you easily exceed the limit. If you keep only the five or six you need and leave the rest in the body, the rejections disappear and the number of streams shrinks too.

What you will do in the next lab

You read the limits the server holds directly from /config and make a table, then actually send a bad label and an overly long line and receive the 400 and its body. Next you start a second Loki with a low rate limit, receive a real 429, and explain in numbers why the first few passed because of the burst. Finally you add backoff retries and go as far as storing everything without losing a single line.