Slowing Down Beats Cutting Off
In one line
A request passes through a filter chain before it reaches a route. The chain has an order, and at the end there must always be a terminal filter (router) that actually sends the request out. Depending on what you put in front of it, fault injection, rate limiting, authentication and header manipulation attach without changing application code.
Why this was needed
There are two ways to answer the question "what happens to the order service if the payment service gets slow". One is to read the code and reason about it, and the other is to actually make it slow and see. The first way confidently produces wrong answers — because where the timeouts are set, how many connection pools there are and how many retries there are are scattered across several places in the code.
If you put a filter in the proxy, the second way becomes possible. The application does not even know it is the subject of an experiment, and no deployment is needed. You only change the configuration and then change it back.
Rate limiting is here for the same reason. If every service implements its own rate limiting, each language uses a different library and the behavior differs subtly. If you do it once in the proxy, everyone gets the same rule.
How it works
Order carries meaning. Filters see the request in the order they are written. So the authentication filter usually goes before the rate limit (there is no reason to spend a token on a request that has not even been authenticated), and router is always last. router is the terminal filter that sends the request to the upstream and brings back the response, so anything you put after it never runs. Envoy does not leave that state until run time; it rejects it when it reads the configuration.
Fault injection (fault) can do two things.
abort— cuts immediately with a set status codedelay— holds for a set time and then sends
Both can be conditioned with percentage for a ratio and headers for a condition. Setting a condition is the key. If you turn it on without a condition, every user becomes a subject of the experiment. And the ratio is a probability drawn independently for each request, not "exactly ten out of twenty".
One thing often misunderstood here — a slowing-down experiment matters more than a cutting experiment. When it is cut you know immediately, but when it slows, connections pile up, threads get tied up, and only then does it become an outage. Whether the timeouts and circuit breakers are set properly shows itself only with delay.
Local rate limiting (local_ratelimit) is a token bucket. It holds up to max_tokens and is refilled by tokens_per_fill every fill_interval. When there are no tokens, a 429 goes out. The important property is in the word "local" in the name — it counts only inside that proxy. With ten proxies, the overall limit is effectively ten times as large. In exchange, it does not ask anything outside, so there is no latency, and there is no effect even if that service dies.
The same filter, differently in each place. A filter attaches to a listener, but its configuration can be overridden per route or per virtual host (typed_per_filter_config). A requirement such as tight on the login path and loose on the read path is solved this way. A setup with the global configuration turned off and turned on only in routes is also common.
What it looks like in the field
Leaving fault injection on and forgetting it. An abort left on at 5% with no condition comes back weeks later as "I get a 500 sometimes". Always put a condition on fault injection, leave a record somewhere that you turned it on, and delete it when it is over.
When the rate limit does not match reality. If you do not factor the number of proxies into the calculation, the limit becomes several times what you intended. Conversely, if you divide by the number of proxies, traffic gets unfairly blocked when it concentrates on one side. Adjust it by looking at the ratio of enabled and rate_limited in the statistics.
When you turn on rate limiting and nobody gets blocked. Most of the time it is confusing two knobs. filter_enabled is "whether to make this request a target of this filter", and filter_enforced is "whether to actually block a request that became a target". If you set the latter to 0, tokens are counted but nothing is blocked — it becomes an observation mode that measures with real traffic before you set the limit. The statistics keep that number separately.
When performance got worse after changing the filter order. If you put a heavy filter (such as authentication that asks something outside) at the front of the chain, you pay that cost even for requests that will be blocked later anyway. The rule is to put cheap filters that filter out a lot at the front.
Official documentation: Fault Injection · Local rate limit · HTTP filters
What you will do in the next lab
First you see that putting the terminal filter anywhere other than last gets the configuration rejected, and you put in a conditional abort and a fixed delay yourself. Next you cut by ratio only and count how many of twenty get caught, make a 429 with a token bucket, and then configure the same filter differently for each route so that only one side is tight. At the end you read the statistics the filters left behind.