TT Lab
Get started
Learn Learning paths Courses

Envoy Internals

Two Envoys, One Counter

Continue in TT Lab

Goal

Start a real rate limit service (envoyproxy/ratelimit) and Redis, and count side by side, together with the local limit, how two Envoys calling the same service share one limit.

Why it matters

If you use only the local limit, the limit grows every time you add proxies. In front of autoscaling, the more of the surge you were trying to stop arrives, the wider the door opens. The global limit solves that problem but makes every request go through one more service, and if you do not know what it does when that service dies (the default is let through), you cannot find out in an outage why nobody was blocked. If you count the two things yourself, why you use them together stays with you.

Steps

  1. Start Redis as a daemon on 127.0.0.1:6379 so that it does not write to disk (--save '', --appendonly no). Then save the output of redis-cli ping and the output of redis-cli config get save, in that order, to /root/envd-rl/01-redis.txt.
  2. In /root/envd-rl/runtime/config/edge.yaml, write two rules for the domain edge — the descriptor plan=free is 3 per hour and plan=pro is 100 per hour. Then start the ratelimit (envoyproxy/ratelimit) that is in the image. Give the configuration through environment variables: Redis 127.0.0.1:6379, RUNTIME_ROOT=/root/envd-rl/runtime, RUNTIME_SUBDIRECTORY=., RUNTIME_WATCH_ROOT=false, USE_STATSD=false, and the ports HTTP 8180, gRPC 8181 and debug 6070 (all three on 127.0.0.1).
  3. Start the upstream python3 /opt/lab/envoy/echo.py 8092, and start an Envoy with /root/envd-rl/envoy-a.yaml with admin port 9983 and listener 127.0.0.1:10083 (--base-id 21). The HTTP filter order is local_ratelimit (stat_prefix local_rl) → ratelimit (domain edge, gRPC to the cluster ratelimit, failure_mode_deny: false, enable_x_ratelimit_headers: DRAFT_VERSION_03) → router. There are two routes — /local puts a local token bucket (2 tokens, 2 more every 60 seconds) on that route only, and / makes the request header x-plan into the descriptor key plan through rate_limits on that route.
  4. In /root/envd-rl/envoy-b.yaml, write a configuration the same as the first Envoy but with admin port 9984 and listener 127.0.0.1:10084, and start it with --base-id 22. Then send one x-plan: pro request to each of the two Envoys, see whether x-ratelimit-remaining in the responses keeps decreasing, and write the two values to /root/envd-rl/04-remaining.txt as two lines, a= and b= (send a first).
  5. Send eight x-plan: free requests to / alternately to A (10083) and B (10084) (starting with A), and write them in order to /root/envd-rl/05-global.txt as eight lines, a1=코드, b1=코드, a2=코드 … b4=코드 (in each line, the Korean word stands for the response code).
  6. In the same way, send eight requests, this time to /local, and write them to /root/envd-rl/06-local.txt as eight lines, a1=코드 … b4=코드 (in each line, the Korean word stands for the response code).
  7. With the ratelimit process stopped, send one x-plan: free request to A and write its status code to /root/envd-rl/07-fail.txt as down_code=. Then append A's statistics cluster.echo.ratelimit.error and cluster.echo.ratelimit.failure_mode_allowed as two lines, error= and allowed=, and start the service again.
  8. In /root/envd-rl/08-report.md, write four lines — global_ok_total= (the number of 200s in step 5), local_ok_total= (the number of 200s in step 6), fail_open_code= (the down_code of step 7) and counted_in= (where the global limit actually counted: envoy or redis) — and below them write what you learned in at least four lines.

Notes

Set up the place that counts first — Redis

Start Redis as a daemon on 127.0.0.1:6379 so that it does not write to disk (--save '', --appendonly no). Then save the output of redis-cli ping and the output of redis-cli config get save, in that order, to /root/envd-rl/01-redis.txt.

The global rate limit service has no state. How many times requests came is all kept in the storage behind it (Redis), and the service only raises one key per window. That is why the limit stays one even if you run several services. The counter is a value to be discarded once the window passes, so there is no reason to leave it on disk. Start it like redis-server --port 6379 --bind 127.0.0.1 --save '' --appendonly no --daemonize yes.

Start the rate limit service

In /root/envd-rl/runtime/config/edge.yaml, write two rules for the domain edge — the descriptor plan=free is 3 per hour and plan=pro is 100 per hour. Then start the ratelimit (envoyproxy/ratelimit) that is in the image. Give the configuration through environment variables: Redis 127.0.0.1:6379, RUNTIME_ROOT=/root/envd-rl/runtime, RUNTIME_SUBDIRECTORY=., RUNTIME_WATCH_ROOT=false, USE_STATSD=false, and the ports HTTP 8180, gRPC 8181 and debug 6070 (all three on 127.0.0.1).

This service reads the YAML under RUNTIME_ROOT/RUNTIME_SUBDIRECTORY/config/. One file is one domain, and a rule applies only when the descriptor's key and value match what the request sent exactly. If no rule matches, it does not limit. Whether it was read properly is shown by curl localhost:6070/rlconfig on the debug port — you should see a line like edge.plan_free: unit=HOUR requests_per_unit=3. When you start it, write it in the shape env 변수=값 … setsid --fork nohup ratelimit > 로그 2>&1 </dev/null (the placeholders are the variable, the value and the log file).

The first Envoy builds a descriptor for each request and asks

Start the upstream python3 /opt/lab/envoy/echo.py 8092, and start an Envoy with /root/envd-rl/envoy-a.yaml with admin port 9983 and listener 127.0.0.1:10083 (--base-id 21). The HTTP filter order is local_ratelimit (stat_prefix local_rl) → ratelimit (domain edge, gRPC to the cluster ratelimit, failure_mode_deny: false, enable_x_ratelimit_headers: DRAFT_VERSION_03) → router. There are two routes — /local puts a local token bucket (2 tokens, 2 more every 60 seconds) on that route only, and / makes the request header x-plan into the descriptor key plan through rate_limits on that route.

The global limit filter does not count by itself. The rate_limits of the route (or virtual host) builds the descriptor from the request, and the filter only sends it to the service. If you put rate_limits on the virtual host, it applies to all the routes of that host, and writing an empty list (rate_limits: []) on a route does not turn it off either (measured) — so put it on the / route. The cluster that asks over gRPC needs HTTP/2 turned on. A request without the x-plan header gets no descriptor built and passes with no limit.

The second Envoy also calls the same service

In /root/envd-rl/envoy-b.yaml, write a configuration the same as the first Envoy but with admin port 9984 and listener 127.0.0.1:10084, and start it with --base-id 22. Then send one x-plan: pro request to each of the two Envoys, see whether x-ratelimit-remaining in the responses keeps decreasing, and write the two values to /root/envd-rl/04-remaining.txt as two lines, a= and b= (send a first).

If the two Envoys count separately, the remaining values of the two responses start from the same number. If they raise the same service and the same Redis key, the second is one less than the first. Header names are case-insensitive, so get only the headers with curl -s -D - -o /dev/null and look for it. If you start the two Envoys with the same --base-id, the shared-memory names collide and the second one does not start.

Even sent alternately, the limit is one

Send eight x-plan: free requests to / alternately to A (10083) and B (10084) (starting with A), and write them in order to /root/envd-rl/05-global.txt as eight lines, a1=코드, b1=코드, a2=코드 … b4=코드 (in each line, the Korean word stands for the response code).

Free is 3 per hour. If the two Envoys share the limit, only three of the eight should be 200 and the rest 429 — regardless of which Envoy received them. If you have already sent free in this time window, there may be fewer 200s, and that too is evidence that it counts globally. If you look at the keys in Redis with redis-cli --scan --pattern 'edge_plan_free_*' and get them, the requests of both Envoys are gathered in one key.

If you send the same eight to the local limit

In the same way, send eight requests, this time to /local, and write them to /root/envd-rl/06-local.txt as eight lines, a1=코드 … b4=코드 (in each line, the Korean word stands for the response code).

The /local route has no global limit, only a local token bucket (2 tokens). The token bucket is inside one Envoy, so A gives 2 and B gives 2. If you put it side by side with the global result, what changes when you add proxies shows up in numbers. The bucket refills every 60 seconds, so wait a minute when you try again.

If the limit service dies, requests are let through

With the ratelimit process stopped, send one x-plan: free request to A and write its status code to /root/envd-rl/07-fail.txt as down_code=. Then append A's statistics cluster.echo.ratelimit.error and cluster.echo.ratelimit.failure_mode_allowed as two lines, error= and allowed=, and start the service again.

Free has already used up its limit, so if the service were alive, it would be 429. When the service is dead, failure_mode_deny: false (the default) lets the request through. Rate limiting is a protective device, and this is the choice not to block the whole service just because it broke. In exchange, that moment is left only in the statistics. The cluster.echo in the statistic name is the name of the upstream cluster the request was heading to.

Write down what happens to the limit when you increase the number of proxies

In /root/envd-rl/08-report.md, write four lines — global_ok_total= (the number of 200s in step 5), local_ok_total= (the number of 200s in step 6), fail_open_code= (the down_code of step 7) and counted_in= (where the global limit actually counted: envoy or redis) — and below them write what you learned in at least four lines.

Copy the values from the files of the earlier steps. In the explanation lines, write "if you increase to three proxies, what the real limit becomes for the local limit and the global limit respectively" and "what is good about using the two together (the local one is the first wall that protects the limit service)".