Add a Broken Upstream and Watch the Cluster
Goal
Set up the way endpoints are discovered and the way health is judged separately, and build yourself the two paths Envoy takes when too few healthy ones are left.
Why it matters
The screen you actually look at often in proxy operations is not the route table but the cluster list — how many servers are where, and how many of them are healthy. If you can read those numbers, half of the 503 reports end on the spot. Panic mode in particular, if you do not see it with your own eyes once, leaves you stuck at "they are all dead, so why does it keep sending" the first time you meet it in production. The priority step-over is the same; you cannot get a feel for when the lower tier takes traffic just by reading the configuration.
Steps
- Start four upstreams —
8084and8086areok,8085and8087arefail. In/root/envd-cluster/cluster.yaml, put aSTATICclusterpool(endpoints8084,8085and8086) and a/poolroute and start it (admin 9941, listener127.0.0.1:10041). From/clusters, pick only the lines that start withpool::and save them to/root/envd-cluster/01-pool.txt. - Add a
STRICT_DNSclusterbydns— the address islocalhost, the port is8084anddns_lookup_familyisV4_ONLY. Put a/bydnsroute with it as well. From/config_dump?resource=static_clusters, pull out one line per cluster as이름 타입(the placeholders are the name and the type) and save it to/root/envd-cluster/02-types.txt. - Add an active health check to
pool—interval1 second,timeout1 second, both thresholds 1, and the path ofhttp_health_checkis/healthz. After a moment, from/clusterspick only thehealth_flagslines of thepoolendpoints and save them to/root/envd-cluster/03-health.txt. - Request
/pool12 times, count which upstream received each one, and write it to/root/envd-cluster/04-spread.txtas four lines,p8084=,p8085=,p8086=andtotal=(the values are the counts received). - Add a cluster
panic— its endpoints are8085and8087, bothfail, and it has the same active health check. Put a/panicroute too. After you request 6 times, write three lines to/root/envd-cluster/05-panic.txt:codes=(the HTTP codes received, separated by spaces),lb_healthy_panic=(the value of the statistic with that name) andmembership_healthy=. - Add a cluster
tiered— put8085(fail) inpriority: 0and8086(ok) inpriority: 1, with the same active health check. Put a/tieredroute too, request 6 times, and write two lines to/root/envd-cluster/06-priority.txt:served_by=(the port that responded) andcount=(the number of times that port received). - In
/root/envd-cluster/07-inventory.txt, write one line per cluster with four values separated by spaces,이름 타입 엔드포인트수 성한수(the placeholders are the name, the type, the number of endpoints and the number of healthy ones), sorted by name. Do not count the values by eye; pull them out of/config_dumpand/stats. - In
/root/envd-cluster/08-report.md, write four lines —healthy=(the number of healthy endpoints of pool),total=(the total number of endpoints of pool),panic_requests_sent=(yes if requests went out in panic mode) andfailover_port=(the port that actually responded in step 6) — and below them write what you learned in at least four lines.
Notes
- Start the upstreams with
python3 /opt/lab/envoy/upstream.py <포트> ok|fail(the placeholder is the port).failreturns 503 whatever path you ask for, so the health check fails along with it. - When you start Envoy, use
setsid --fork nohup envoy -c <파일> --log-level warn --concurrency 1 > <로그> 2>&1 </dev/null(the placeholders are the file and the log), and before you start it again, clean up withpkill -x envoy. - The health check interval is 1 second, so right after you load the configuration there is no verdict yet. Instead of a fixed
sleep, use a loop that runs until the state you want shows up in/clustersor/stats. /clustersproduces more than twenty lines per endpoint. Pick out only the lines you need, as ingrep 'health_flags'.- Common mistake — treating outlier detection and the active health check as the same thing. This lab is about the active kind.
Hang three addresses on one name
Start four upstreams — 8084 and 8086 are ok, 8085 and 8087 are fail. In /root/envd-cluster/cluster.yaml, put a STATIC cluster pool (endpoints 8084, 8085 and 8086) and a /pool route and start it (admin 9941, listener 127.0.0.1:10041). From /clusters, pick only the lines that start with pool:: and save them to /root/envd-cluster/01-pool.txt.
A cluster is a group of "servers you can call by this name". STATIC is the simplest way, writing the addresses in the configuration as they are, so if you do not change the configuration, the list does not change. Start the upstreams with python3 /opt/lab/envoy/upstream.py <포트> ok|fail (the placeholder is the port) — fail returns 503 whatever path you ask for. /clusters produces several lines per endpoint, so filter it with grep before saving.
Write a name instead of addresses
Add a STRICT_DNS cluster bydns — the address is localhost, the port is 8084 and dns_lookup_family is V4_ONLY. Put a /bydns route with it as well. From /config_dump?resource=static_clusters, pull out one line per cluster as 이름 타입 (the placeholders are the name and the type) and save it to /root/envd-cluster/02-types.txt.
How endpoints are discovered is the cluster's type. STATIC uses the addresses written in the configuration, STRICT_DNS resolves the name again periodically and makes all the addresses in the response endpoints, and LOGICAL_DNS holds on to only one of them. EDS is the way the control plane pushes in the list. You can pull out the two values with jq -r '.configs[]?.cluster | "\(.name) \(.type)"'.
Probe periodically, separately from requests
Add an active health check to pool — interval 1 second, timeout 1 second, both thresholds 1, and the path of http_health_check is /healthz. After a moment, from /clusters pick only the health_flags lines of the pool endpoints and save them to /root/envd-cluster/03-health.txt.
Outlier detection takes endpoints out by watching actual requests fail, but an active health check probes separately, regardless of requests. So it knows the state even in hours with no traffic, and the first request does not go to a broken server. In exchange, each server gets one more periodic load. health_flags in /clusters is healthy when it is healthy and /failed_active_hc when it fails the active check. The check interval is 1 second, so right after you load the configuration there may be no verdict yet — wait until the flag appears.
Do not send to places that are not healthy
Request /pool 12 times, count which upstream received each one, and write it to /root/envd-cluster/04-spread.txt as four lines, p8084=, p8085=, p8086= and total= (the values are the counts received).
Load balancing runs only among healthy endpoints. So the one that failed the active check should not receive a single request. The response body contains the port, so count with that. If you started it with --concurrency 1, the two healthy ones will get exactly half each.
What to do when there is no healthy one at all
Add a cluster panic — its endpoints are 8085 and 8087, both fail, and it has the same active health check. Put a /panic route too. After you request 6 times, write three lines to /root/envd-cluster/05-panic.txt: codes= (the HTTP codes received, separated by spaces), lb_healthy_panic= (the value of the statistic with that name) and membership_healthy=.
When the healthy ratio falls below the threshold (50% by default), Envoy ignores health information and sends to all of them. This is called panic mode. It looks strange, but the reasoning is simple — if the checks may be wrong and you send nowhere, that is a certain outage, and if you send, at least some requests may survive. You see whether it happened in the cluster.<이름>.lb_healthy_panic statistic (the placeholder is the cluster name). Pull out just the value with curl -s localhost:<admin>/stats | grep ... (the placeholder is the admin port).
When the upper tier is all dead, it steps down to the lower one
Add a cluster tiered — put 8085 (fail) in priority: 0 and 8086 (ok) in priority: 1, with the same active health check. Put a /tiered route too, request 6 times, and write two lines to /root/envd-cluster/06-priority.txt: served_by= (the port that responded) and count= (the number of times that port received).
You can give each endpoint group a priority. Normally only priority 0 is used. If the healthy ratio of priority 0 drops, priority 1 takes the shortfall, and if priority 0 is wiped out, everything goes to priority 1. A setup that leaves standby servers in another region idle normally and uses them only during an outage is built this way. The /clusters output has a priority line for each endpoint, so you can confirm which tier it is.
Build a list of the four clusters
In /root/envd-cluster/07-inventory.txt, write one line per cluster with four values separated by spaces, 이름 타입 엔드포인트수 성한수 (the placeholders are the name, the type, the number of endpoints and the number of healthy ones), sorted by name. Do not count the values by eye; pull them out of /config_dump and /stats.
This is the table you look at first in production — how many servers each cluster has and how many of them are healthy. The two statistics cluster.<이름>.membership_total and membership_healthy (the placeholder is the cluster name) are those two numbers. The type comes from /config_dump?resource=static_clusters. Join the values pulled from the two places by cluster name.
Leave a cluster operations memo
In /root/envd-cluster/08-report.md, write four lines — healthy= (the number of healthy endpoints of pool), total= (the total number of endpoints of pool), panic_requests_sent= (yes if requests went out in panic mode) and failover_port= (the port that actually responded in step 6) — and below them write what you learned in at least four lines.
Take the values from the files you made in the earlier steps — if you write them from memory, they will not match. In the explanation lines, write sentences you will use the next time you design, such as "the difference between an active health check and outlier detection".