TT Lab
Get started
Learn Learning paths Courses

Loki — A Log Store That Does Not Index Logs

We reported the incident start forty minutes too late

Continue in TT Lab

Goal

You exceed Loki's query limits yourself to tell apart what gets silently cut from what is rejected with a 400, and build a way to scan a wide range without missing anything within the limits.

Why it matters

Three levers decide the number of results of a log query. The request's limit cuts silently, while the server's max_entries_limit_per_query and max_query_length reject with a 400. A rejection is visible right away, but a silently cut answer looks like "this is everything." On top of that, direction decides what gets cut off, so if you look for the start of an incident with the default, backward, exactly what you want to find is cut off. The answer when you have to scan a wide range is not to raise the limit but to cut the range, ask several times, and combine — the fact that each piece is smaller than the limit is what makes the total trustworthy.

Steps

  1. In /root/lk-query-limits, start Loki, write date +%s in /root/lk-query-limits/anchor.txt, and then load the data with python3 /opt/lab/d5/gen.py querylimits "$(cat anchor.txt)". Then run a log query for {app="gateway"} over two hours back from the reference time and write the total number of lines in /root/lk-query-limits/01-total.txt on one line as total=<정수> (the placeholder is the integer count). Give limit as 5000, and be sure to check that the number you get back is smaller than that value.
  2. Run a log query over the same two-hour range with limit=100 and direction=backward, and write the number of lines you got back and the nanosecond timestamp of the earliest line among them in /root/lk-query-limits/02-cut.txt on two lines — returned=<정수> and oldest_ts=<나노초> (the placeholders are the integer count and the nanosecond timestamp). Confirm for yourself that there is no "truncated" indication anywhere in the response.
  3. Throw the same query with limit=6000. Write the response's status code, the limit value written in the body, and the value of that setting read from the server's /config in /root/lk-query-limits/03-entries.txt on three lines — code=<정수>, limit_in_body=<정수>, and limit_in_config=<정수> (the placeholders are the integer values).
  4. Throw the same query over a 60-day range. Write the status code and the range length limit from the server's /config in /root/lk-query-limits/04-length.txt on three lines — code=<정수>, param=<설정 이름>, and value=<서버가 찍은 값 그대로> (the placeholders are the integer code, the setting name, and the value exactly as the server printed it).
  5. Throw the same query over a two-hour range and a 30-minute range, measure splits in the response statistics for each, also read the server's piece interval setting, and write three lines in /root/lk-query-limits/05-splits.txt — splits_2h=<정수>, splits_30m=<정수>, and interval=<서버가 찍은 값 그대로> (the placeholders are the integer counts and the value exactly as the server printed it).
  6. Throw the same query three times with limit set to 50, 500, and 5000, and write the results in /root/lk-query-limits/detect.tsv on three lines with no header, each with three tab-separated columns, <limit><탭><returned><탭><truncated> (the placeholders are the limit, a tab, the returned count, a tab, and the truncated flag). truncated is yes or no, and the criterion for deciding is the true total you found in step 1.
  7. Create /root/lk-query-limits/scan.sh, split the two-hour range into four 30-minute pieces, query each, and compute the total. Throw each piece with limit=5000, and if a piece's number of results equals the limit, it must also print INCOMPLETE. Write the result of running the script in /root/lk-query-limits/07-scan.txt on two lines as sum=<정수> and chunks=4 (the placeholder is the integer total). The total must equal the step 1 total.
  8. Design a query that fetches the most recent 50 lines with status=500 in the two-hour range, and write three lines in /root/lk-query-limits/08-latest.txt — returned=<정수>, newest_ts=<나노초>, and oldest_ts=<나노초> (the placeholders are the integer count and the nanosecond timestamps). If there are fewer than 50, only as many as exist need to come out.

Notes

Load two hours of data and find the true total

In /root/lk-query-limits, start Loki, write date +%s in /root/lk-query-limits/anchor.txt, and then load the data with python3 /opt/lab/d5/gen.py querylimits "$(cat anchor.txt)". Then run a log query for {app="gateway"} over two hours back from the reference time and write the total number of lines in /root/lk-query-limits/01-total.txt on one line as total=<정수> (the placeholder is the integer count). Give limit as 5000, and be sure to check that the number you get back is smaller than that value.

If the number you get back equals limit, it was cut off, so you cannot use it as the total. If it is smaller, you received that whole range. This number becomes the criterion for judging "was it cut?" throughout the later steps. You can also count with a metric query (sum(count_over_time(...[2h]))), but a range aggregation does not include the left end of the range, so it may come out about one line different.

limit cuts without saying a word

Run a log query over the same two-hour range with limit=100 and direction=backward, and write the number of lines you got back and the nanosecond timestamp of the earliest line among them in /root/lk-query-limits/02-cut.txt on two lines — returned=<정수> and oldest_ts=<나노초> (the placeholders are the integer count and the nanosecond timestamp). Confirm for yourself that there is no "truncated" indication anywhere in the response.

backward fills from the most recent. So the earliest time of the 100 lines you got back is not the start of the range but much later. Think about what happens if you report this value as the incident start time.

Server limits reject outright

Throw the same query with limit=6000. Write the response's status code, the limit value written in the body, and the value of that setting read from the server's /config in /root/lk-query-limits/03-entries.txt on three lines — code=<정수>, limit_in_body=<정수>, and limit_in_config=<정수> (the placeholders are the integer values).

The body of the 400 is helpful — it even writes the numbers, which limit you exceeded and by how much. Find the same name in limits_config of /config and check that the two values are the same. If automation throws this body away and keeps only the status code, you will not be able to find the cause later.

A range that is too long is rejected too

Throw the same query over a 60-day range. Write the status code and the range length limit from the server's /config in /root/lk-query-limits/04-length.txt on three lines — code=<정수>, param=<설정 이름>, and value=<서버가 찍은 값 그대로> (the placeholders are the integer code, the setting name, and the value exactly as the server printed it).

Just give start as 60 days ago. The body writes the query length and the limit side by side. Find the same name in limits_config of /config — the notation of the value may differ from the body.

A query is split into time pieces and run

Throw the same query over a two-hour range and a 30-minute range, measure splits in the response statistics for each, also read the server's piece interval setting, and write three lines in /root/lk-query-limits/05-splits.txt — splits_2h=<정수>, splits_30m=<정수>, and interval=<서버가 찍은 값 그대로> (the placeholders are the integer counts and the value exactly as the server printed it).

splits is in data.stats.summary. Work out how the number of pieces comes from the range length and the setting value — getting 0 for a short range has a meaning too.

How to notice that it was cut

Throw the same query three times with limit set to 50, 500, and 5000, and write the results in /root/lk-query-limits/detect.tsv on three lines with no header, each with three tab-separated columns, <limit><탭><returned><탭><truncated> (the placeholders are the limit, a tab, the returned count, a tab, and the truncated flag). truncated is yes or no, and the criterion for deciding is the true total you found in step 1.

If the number of lines you got back equals limit, it was almost certainly cut, and if it is smaller than the total, it was certainly cut. Also think about whether the two criteria can disagree — if you received a number equal to the total, it was not cut even if it equals limit.

Applied ① — cut the range and count without missing anything

Create /root/lk-query-limits/scan.sh, split the two-hour range into four 30-minute pieces, query each, and compute the total. Throw each piece with limit=5000, and if a piece's number of results equals the limit, it must also print INCOMPLETE. Write the result of running the script in /root/lk-query-limits/07-scan.txt on two lines as sum=<정수> and chunks=4 (the placeholder is the integer total). The total must equal the step 1 total.

The key is to set the start and end of each piece so they do not overlap — if two pieces both count the boundary second, the total gets bigger, and if you miss it, it gets smaller. Setting the left end of the range inclusive and the right end exclusive is clean.

Applied ② — fetch the latest 50 errors exactly

Design a query that fetches the most recent 50 lines with status=500 in the two-hour range, and write three lines in /root/lk-query-limits/08-latest.txt — returned=<정수>, newest_ts=<나노초>, and oldest_ts=<나노초> (the placeholders are the integer count and the nanosecond timestamps). If there are fewer than 50, only as many as exist need to come out.

You have to decide the direction and the limit together. The point of this lab is that the direction you want when you want "the most recent" is the opposite of the direction you needed in step 2 to find the start of the incident. Also check whether the number you got back equals the limit.