We Added Proxies and the Limit Grew
In one line
Global rate limiting puts the place that counts in one spot outside the proxy. For each request Envoy builds a descriptor and asks the rate limit service over gRPC, and the service raises a counter in Redis and answers allowed or exceeded. So however many proxies there are, there is one limit.
Why this was needed
Local rate limiting (a token bucket) is fast and has no dependencies. In exchange, the bucket is inside a single Envoy. The official documentation also states that the bucket of a local limit applies by default to "one Envoy process". The moment you grow the ingress proxies from two to six, a promise of "1,000 times per hour per customer" becomes 6,000 times. Once autoscaling is attached, the limit grows by itself along with traffic — the more of the surge you were trying to stop arrives, the wider the door opens.
If you have a promise that applies to the whole, as with selling an API, there has to be one place to count. That is the global rate limit service, and on the Envoy side there is an HTTP filter (envoy.filters.http.ratelimit) that asks that service. The reference implementation on the service side is envoyproxy/ratelimit.
How it works
1. The route builds the descriptor. The filter does not know by itself what to count. The rate_limits of the route (or virtual host) extracts values from the request and builds the descriptor.
rate_limits:
- actions:
- request_headers: { header_name: x-plan, descriptor_key: plan }
# x-plan: free 인 요청 → 설명자 [("plan", "free")]
If the header is absent, no descriptor is built, and that request is not limited. The rate_limits written on a virtual host applies to all the routes of that host, and it is replaced only when a route has its own rate_limits.
2. The service matches against its rules. The service configuration is a list of descriptor rules under one domain.
| Setting | Meaning |
|---|---|
domain: edge |
It must be the same as the domain of the Envoy filter for this rule to be used |
key: plan, value: free |
When the descriptor is exactly this pair |
unit: hour, requests_per_unit: 3 |
3 times per time window |
3. Redis counts. For each window, the service raises one key (edge_plan_free_<창 시작 시각>, where the placeholder is the window start time) with INCRBY and sets a TTL of the window length. The service itself has no state, so you can run several — the limit is gathered in one Redis key.
4. The result shows in the response. When exceeded, Envoy returns a 429. If you turn on enable_x_ratelimit_headers, the headers x-ratelimit-limit, x-ratelimit-remaining and x-ratelimit-reset are attached, so the client can tell how much is left.
When the service dies, failure_mode_deny decides. The default, false, is let through. Rate limiting is a protective device, and this is the choice not to stop the service just because it broke. In exchange, that moment is left only in the statistics cluster.<업스트림>.ratelimit.error and failure_mode_allowed (the placeholder is the upstream name).
Using the two together. A global limit costs one more network round trip per request, and that service becomes a new bottleneck. So you put the local limit in front to filter out the surge first, and use the global limit to keep the promise. This is also the combination the official documentation recommends.
What it looks like in the field
"I added proxies and the customer limit went up." This is the case where only the local limit was used. Writing the limit divided by the number of proxies collapses in front of autoscaling.
"I changed the limit configuration but it is not applied." If the descriptor differs from the rule by even one character, the rule does not apply, and a request the rule did not apply to passes with no limit. Check separately whether the rule was read through /rlconfig on the service, and which descriptor Envoy sends.
"Nobody was blocked during the Redis outage." Because the default is let through. If that is intended, set an alert on the statistics, and if it is not intended, change to failure_mode_deny: true, but then you have to accept together that Redis holds the availability of every request.
Official documentation: Global rate limiting · Rate limit filter · Local rate limit filter · envoyproxy/ratelimit
What you will do in the next lab
You start Redis and the reference-implementation rate limit service inside a Pod, and set up two Envoys that call the same service. You count side by side that requests sent alternately to the two share one limit, and that the same requests sent through a local token bucket are counted separately per proxy. At the end you kill the service and confirm from the responses and statistics that the default is let through.