Keep every log line and you will find nothing
In one line
Logs are not priced per line. Whether it is kept or lost per request decides investigability, and level and sampling rate are cost knobs layered on top of that.
Why this matters
The first step of investigation, as described in the troubleshooting chapter of the SRE Book, is to find the request that left the symptom, and this is what happens during an outage investigation. You found out the request_id of a failed request. You put that ID into the log store and two lines come out — "request started" and "request failed". What happened in between is not there. That is because DEBUG was turned off after last month's cost meeting.
The opposite situation is just as bad. The team that turned DEBUG back on saw its storage bill quadruple a month later, and search became so slow it was unusable for investigation. The next easy choice was "let's keep only 10%". But when they sampled per line, only half of the same request's lines remained, and no request could be read to the end. Costs went down and investigability went to 0.
All three times, the same mistake is being made. It treats the unit of a log as a line. The unit a person reads when investigating is not a line but the bundle of lines one request left behind.
How it works
So sampling is done per request. If you hash the request_id and keep only requests whose remainder after dividing by N is 0, a request's lines are either kept whole or lost whole. This is called consistent head sampling. For the hash, hashlib in the Python standard library is enough. The hash does one job — giving the same answer for the same ID no matter when or where it is computed. So if several services use the same rule, the flow of one request continues across service boundaries.
One more rule is layered on top. Keep every request that ended in an error, regardless of the sampling rate. The requests worth investigating are usually failed requests, and such requests are only a few percent of the total, so keeping them hardly increases cost. If you use a 10% sampling rate together with error retention, the number of lines goes up a little from 9.6% to 14.3%, but the retention rate of failed requests goes from 0% to 100%. Few knobs have a bigger effect per cost.
The remaining two knobs deal with repetition. Code caught in a retry loop prints the same line hundreds of times per second. If you collapse consecutive identical lines into one and attach repeated=<k>, you reduce only the volume without losing information. If there is still too much left, you apply a per-second cap — but the cap must be applied only to DEBUG. If ERROR disappears because of the cap, costs go down and investigation becomes impossible.
| Knob | What it reduces | What it loses |
|---|---|---|
| Lowering the level | Most of the volume | The whole flow of the failed request |
| Per-request sampling | In proportion to volume | Every request not caught in the sample |
| Error retention | Increases (a little) | Nothing |
| Collapsing repeats | The volume of retry loops | Nothing (the count is kept) |
| Per-second cap | The runaway period | The lines over the cap |
Along with these knobs, the policy document writes down the retention period per level. There is almost no reason to keep DEBUG for 90 days, and if you keep ERROR for only 3 days, you cannot find anything in the quarterly retrospective. If you tie the levels to a single retention period, you end up with one of two things — losing ERROR to match DEBUG, or paying several times more for DEBUG to match ERROR.
For all these knobs to work there is one prerequisite. The line must have a request_id. This is why OpenTelemetry's log data model makes trace_id and span_id first-class fields of a log record. Logs without an ID can neither be grouped per request nor sampled per request, so in the end there is no choice but to cut per line. Half of reducing cost is already decided at the instrumentation stage.
What it looks like in the field
One team lowered the sampling rate to 1% and reported "we cut costs by 99%". Three months later, in a payment failure investigation, there was not a single related request in the logs. It was because there was no error retention rule. When the rule was added, cost rose to 1.3% and all failed requests were kept. What the difference between 1.3% and 1% bought was 'being able to investigate'.
Another team applied a per-second cap without distinguishing levels. There was no problem in normal times, but at the moment a real outage occurred and thousands of lines per second poured out, the cap kicked in and ERROR lines were cut off. Only the logs of the time they were most needed were missing. It was a design that forgot that a cap operates not in normal times but at the worst moment.
The third case is about collapsing. A batch job caught in a retry loop printed 400 million lines a day, and those lines did not differ by a single character. With one consecutive-collapse step, those 400 million lines shrank to 20,000, and thanks to the repeated= number, the question "how many times did it retry?" could still be answered. Discarding and summarizing are different things.
What you will do in the next lab
Using a fixed log file preloaded in the Pod, you build a cost table by level and count for yourself how many of the nine lines of the failed request remain when you turn off DEBUG. Then you build a filter that draws a per-request sample by request_id hash, add a rule that keeps requests that ended in an error, and make a table of the trade-off between sampling rate and investigability. After building one more filter that collapses repeats and applies a per-second cap, you chain the two filters into a pipeline and confirm that while cutting 88%, all nine of those lines remain, and finally write a policy file containing the levels, sampling rates, retention periods, and estimated cost.