Create an Outage Without Touching Code
Goal
Put fault injection and rate limiting into the HTTP filter chain, and confirm for yourself how order, conditions and ratios change the result.
Why it matters
A filter is a place to do something in front of the application without touching it. So you can answer a question like "what happens to us if payment gets slow" by measuring instead of reasoning. However, used wrongly, this tool becomes an outage itself — fault injection turned on without a condition, a rate limit that left out the number of proxies, and a chain with its order reversed are all shapes that really caused incidents. If you build them once by hand, you can recognize all three from the configuration alone.
Steps
- Start an upstream on
8095asok. In/root/envd-filter/f-ok.yaml, put a configuration whose only filter isrouter(admin 9976, listener127.0.0.1:10076, clusterorigin), and in/root/envd-filter/f-badorder.yaml, put a configuration with afaultfilter afterrouter. Check each of the two files withenvoy --mode validateand in/root/envd-filter/01-order.txtwrite two lines,ok_rc=andbad_rc=. - Create
/root/envd-filter/f-abort.yaml— put afaultfilter beforerouter, with anabortof 418, and the condition is the request headerx-fault: yes. After you start it, send a request with the header and one without, and in/root/envd-filter/02-abort.txtwrite two lines,with=andwithout=(the HTTP code of each). - Create
/root/envd-filter/f-delay.yaml— setdelay.fixed_delayof thefaultfilter to 2 seconds, with the condition headerx-slow: yes. After you start it, measure the elapsed time of a request with the header and one without, and in/root/envd-filter/03-delay.txtwrite three lines,slow_ms=,fast_ms=andgap_ms=(integer milliseconds,gap_msis the difference of the two values). - Create
/root/envd-filter/f-half.yaml— give a 503 withabortwith no condition, but setpercentageto 50%. After you start it, request 20 times and in/root/envd-filter/04-percentage.txtwrite three lines:total=20,aborted=(the number of times you received 503) andpassed=(the number of times you received 200). - Create
/root/envd-filter/f-lrl.yaml— put alocal_ratelimitfilter beforerouterinstead offault, withstat_prefixlrl, tokensmax_tokens3,tokens_per_fill3 andfill_interval300s, and both enabled and enforced at 100%. After you start it, request 6 times and in/root/envd-filter/05-ratelimit.txtwrite three lines,ok=,limited=andcodes=(codesis the codes received, separated by spaces). - Create
/root/envd-filter/f-perroute.yaml— leave the globallocal_ratelimitturned off by settingfilter_enabledto 0%, and give the route/tight1 token (stat_prefix: tight) and/loose5 tokens (stat_prefix: loose) withtyped_per_filter_config. After you start it, request/tight3 times and/loose3 times and in/root/envd-filter/06-perroute.txtwrite four lines:tight_ok=,tight_limited=,loose_ok=andloose_limited=. - From
/statsof the Envoy that is running now, pull out the rate-limit statistics and in/root/envd-filter/07-stats.txtwrite three lines:tight_rate_limited=,loose_rate_limited=andenabled_total=(the first two are therate_limitedvalue of eachstat_prefix, and the last is the sum of theenabledvalues of the twostat_prefix). - In
/root/envd-filter/08-report.md, write four lines —router_last=(yes if the terminal filter must be last),abort_status=(the code you received in step 2),delay_ms=(thegap_msof step 3) andtight_allowed=(the number of times/tightlet through in step 6) — and below them write what you learned in at least four lines.
Notes
- When you start Envoy, use
setsid --fork nohup envoy -c <파일> --log-level warn --concurrency 1 > <로그> 2>&1 </dev/null(the placeholders are the file and the log), and before you start it again, clean up withpkill -x envoy. - Wait for startup not with a fixed
sleepbut with a loop that runs until/readyreturns LIVE. - To get only the status code, use
curl -s -o /dev/null -w '%{http_code}', and for the elapsed time use-w '%{time_total}'. The time is a decimal in seconds, so you have to multiply to convert it to milliseconds. - Start the upstream with
python3 /opt/lab/envoy/upstream.py <포트> ok(the placeholder is the port). - Common mistake — putting
routerat the very front or in the middle of the chain. A filter after it never runs, so Envoy rejects the whole configuration. - Common mistake — setting
fill_intervalshort. Tokens refill during the lab and the results wobble.
The end of the chain is fixed
Start an upstream on 8095 as ok. In /root/envd-filter/f-ok.yaml, put a configuration whose only filter is router (admin 9976, listener 127.0.0.1:10076, cluster origin), and in /root/envd-filter/f-badorder.yaml, put a configuration with a fault filter after router. Check each of the two files with envoy --mode validate and in /root/envd-filter/01-order.txt write two lines, ok_rc= and bad_rc=.
The last filter in a filter chain has the role of actually sending the request out to the upstream. That is router, and such a filter is called a terminal filter. If you put anything after a terminal filter, that filter never runs, so Envoy does not leave that state until run time and rejects it when it reads the configuration. The reason appears as it is in the rejection message, so read it.
Cut only the requests that meet the condition
Create /root/envd-filter/f-abort.yaml — put a fault filter before router, with an abort of 418, and the condition is the request header x-fault: yes. After you start it, send a request with the header and one without, and in /root/envd-filter/02-abort.txt write two lines, with= and without= (the HTTP code of each).
Fault injection is a tool for checking "what happens to our service if this service dies" while in operation. That is why setting a condition is the key — if you turn it on without a condition, every user becomes a target. If you use a header condition, only the test tool attaches that header and makes only its own requests experience the fault. 418 is a code not used in real services, so it stands out for experiments.
Make it slow instead of cutting it
Create /root/envd-filter/f-delay.yaml — set delay.fixed_delay of the fault filter to 2 seconds, with the condition header x-slow: yes. After you start it, measure the elapsed time of a request with the header and one without, and in /root/envd-filter/03-delay.txt write three lines, slow_ms=, fast_ms= and gap_ms= (integer milliseconds, gap_ms is the difference of the two values).
Getting slow is harder to deal with than getting cut. When it is cut you know immediately, but when it slows, connections pile up, threads get tied up, and only then does it become an outage. So to check whether timeouts and circuit breakers are set properly, you need a slowing-down experiment, not a cutting experiment. Measure the elapsed time with curl -w '%{time_total}' — it is a decimal in seconds, so you have to multiply to convert it to milliseconds.
Cut only some of them
Create /root/envd-filter/f-half.yaml — give a 503 with abort with no condition, but set percentage to 50%. After you start it, request 20 times and in /root/envd-filter/04-percentage.txt write three lines: total=20, aborted= (the number of times you received 503) and passed= (the number of times you received 200).
A ratio is a probability drawn independently for each request, not "exactly ten out of twenty". So if you count again with the same configuration, the numbers differ. It is the same reason "the ratio does not match" reports come up in canary verification. This property is also why, when you use fault injection in production, you start with a very low ratio — even one in a hundred becomes a sufficient sample if there are many requests.
Count tokens inside the proxy
Create /root/envd-filter/f-lrl.yaml — put a local_ratelimit filter before router instead of fault, with stat_prefix lrl, tokens max_tokens 3, tokens_per_fill 3 and fill_interval 300s, and both enabled and enforced at 100%. After you start it, request 6 times and in /root/envd-filter/05-ratelimit.txt write three lines, ok=, limited= and codes= (codes is the codes received, separated by spaces).
Local rate limiting counts tokens only inside that proxy. With ten proxies the limit is effectively ten times as large, so to keep the overall limit you have to set it divided by the number of proxies or use an external rate limit service. In exchange, it does not ask anything outside, so there is no latency, and there is no effect even if that service dies. A request over the limit gets a 429. If you set fill_interval long, tokens do not refill during the lab and the results stay stable.
Configure the same filter differently for each route
Create /root/envd-filter/f-perroute.yaml — leave the global local_ratelimit turned off by setting filter_enabled to 0%, and give the route /tight 1 token (stat_prefix: tight) and /loose 5 tokens (stat_prefix: loose) with typed_per_filter_config. After you start it, request /tight 3 times and /loose 3 times and in /root/envd-filter/06-perroute.txt write four lines: tight_ok=, tight_limited=, loose_ok= and loose_limited=.
A filter attaches per listener, but its configuration can be overridden per route or per virtual host — that is typed_per_filter_config. A requirement such as tight on the login path and loose on the read path is solved this way. Turning the global configuration off and turning it on only in routes is also a common shape. The key is the name of the filter, and inside the value you write that filter's configuration again in full. If you give each route a different stat_prefix, you can also tell them apart in the statistics.
Read the numbers the filter leaves behind
From /stats of the Envoy that is running now, pull out the rate-limit statistics and in /root/envd-filter/07-stats.txt write three lines: tight_rate_limited=, loose_rate_limited= and enabled_total= (the first two are the rate_limited value of each stat_prefix, and the last is the sum of the enabled values of the two stat_prefix).
A filter leaves statistics under its own name. The name of local rate limiting has the shape <그 자리의 stat_prefix>.http_local_rate_limit.<항목> (the placeholders are the stat_prefix at that spot and the item), and the items include enabled (the number of requests that became targets), rate_limited (the number actually blocked) and ok (the number let through). What matters in operations is not the blocked number itself but the ratio of the blocked number to the target number — if that ratio stays high, the limit does not match reality. If you do not know the exact name, first look for it with curl -s localhost:<admin>/stats | grep rate_limit.
Leave a filter design memo
In /root/envd-filter/08-report.md, write four lines — router_last= (yes if the terminal filter must be last), abort_status= (the code you received in step 2), delay_ms= (the gap_ms of step 3) and tight_allowed= (the number of times /tight let through in step 6) — and below them write what you learned in at least four lines.
In the explanation lines, write when to use and when not to use each mechanism. For example, fault injection is used with a condition, and local rate limiting multiplies the limit by the number of proxies. Take the values from the files of the earlier steps.