TT Lab
Get started
Learn Learning paths Courses

Envoy Internals

Run Envoy Yourself and Break It

Continue in TT Lab

Goal

You start Envoy yourself and build and break routing, timeouts, retries and outlier detection one at a time. Istio's data plane is exactly this Envoy, so what you learn here applies directly to responding to mesh outages.

Environment

envoy --version
curl -s localhost:9901/stats     # admin (띄운 뒤)

The starting configuration is at /opt/lab/envoy/minimal.yaml. Copy it and edit it.

Creating upstreams

python3 -m http.server 8081 &          # 정상

If you need a failing or slow upstream, use /opt/lab/envoy/upstream.py.

python3 /opt/lab/envoy/upstream.py 8082 fail &   # 항상 503
python3 /opt/lab/envoy/upstream.py 8083 slow &   # 3초 걸림

Starting it and starting it again

setsid --fork nohup envoy -c e.yaml --log-level warn > envoy.log 2>&1 </dev/null
# 고친 뒤에는
pkill -f 'envoy -c'
setsid --fork nohup envoy -c e.yaml --log-level warn > envoy.log 2>&1 </dev/null

Be sure to add setsid --fork. If you just start it with &, it dies along with the shell when the shell changes. If you start it with setsid nohup … & without --fork, it survives in an interactive shell, but on paths where a shell attaches and detaches briefly, such as grading or step preparation, it dies when that shell ends. --fork splits off once more and attaches to PID 1, so it stays alive either way. Grading looks at the admin of a live Envoy, so if it is dead, you fail from step 1.

If the configuration is wrong, Envoy does not start; it dies. The reason is on the first line of envoy.log.

Only one Envoy can run at a time

Even if you give it different ports, the second one dies like this.

unable to bind domain socket with base_id=0, errno=98 (see --base-id option)

It is not a port problem; it is the shared-memory domain sockets colliding. When you compare two configurations, stop the first with pkill -f 'envoy -c' and start again (or use --base-id 1).

When you cannot see the access log

Two things are common.

Reading order

If you read Envoy configuration in this order, you will not get confused.

listener → filter chain → http_connection_manager → route_config → cluster

Steps

  1. Start with a minimal configuration → 01-boot.txt
  2. Look inside admin → 02-admin.txt
  3. Route order → 03-route.txt
  4. Timeout → 04-timeout.txt
  5. Retries → 05-retry.txt
  6. Outlier detection → 06-outlier.txt
  7. Response flags → 07-flags.txt
  8. Wrap-up → 08-notes.md

Start with a minimal configuration

Start an Envoy with one listener and one cluster, confirm that a request reaches the upstream, and record it in 01-boot.txt.

Copy /opt/lab/envoy/minimal.yaml and edit it. When you start it, be sure to add setsid --fork — if you just start it with &, it dies along with the shell when the shell changes and you fail grading. With --fork it splits off once more and attaches to PID 1, so it survives even after the shell ends.

setsid --fork nohup envoy -c e.yaml --log-level warn > envoy.log 2>&1 </dev/null

To check: curl -s localhost:10000/. If you read Envoy configuration in the order listener → filter chain → route → cluster, you will not get confused.

Look inside with admin

From the admin interface (9901), pull out the cluster list and the config_dump section names and record them in 02-admin.txt.

curl -s localhost:9901/clusters, curl -s localhost:9901/config_dump | jq -r '.configs[]."@type"'. In production, config_dump is the only way to confirm "did Envoy really receive my configuration" — what you wrote in a file and what Envoy holds can differ.

The route written first wins

Send the two routes /api and / to different clusters, and show that changing the order changes the result for the same request. In 03-route.txt, record the response to the same request on two lines, before= and after=.

Envoy scans routes from the top and stops at the first match. If you put prefix: "/" on top, everything below it is dead. Send the same /api/x request under each of the two orders and record it like this (along with the configuration you wrote):

before=api:8084 /api/x
after=ok:8081 /api/x

In practice, most "I added a route but it does not take effect" cases are this.

Cut off a slow upstream

Put a 1-second route timeout on the upstream that takes 3 seconds so that it returns 504, and record it in 04-timeout.txt together with the response flag in the access log.

Put timeout: 1s on the route. If you put %RESPONSE_FLAGS% in the access log format, UT (Upstream Timeout) is printed. Check whether the elapsed time is exactly 1000ms — that is the proof that Envoy cut it off.

If you do not see the log right away, wait about 10 seconds. Envoy batches file access logs before writing them — if you grep right after a request, it is not there yet.

How many retries actually went out

Put num_retries: 3 retries on the upstream that always returns 503, confirm the number of retries with statistics, and record it in 05-retry.txt.

retry_policy: {retry_on: "5xx", num_retries: 3}. To check, curl -s localhost:9901/stats | grep upstream_rq_retry. The access log shows URX (retries exhausted). Retries are not free — if the upstream is already dying, retries make the load four times larger.

Pull out the broken endpoint

In a cluster with 2 endpoints, make only one return 503, show with statistics that outlier detection removes that one, and record it in 06-outlier.txt.

outlier_detection: {consecutive_5xx: 2, interval: 1s, base_ejection_time: 30s}. After about 20 requests, run curl -s localhost:9901/stats | grep -E 'ejections_active|membership_healthy'. If it works properly, about 18 of the 20 succeed — because after working out from the first 2 which side is broken, it never sends there again.

Read the response flags

Collect the access logs of the failures you have made so far and record them in 07-flags.txt. There must be at least two different kinds of flags.

UT upstream timeout, URX retries exhausted, UF connection failure, NR no matching route, UH no healthy upstream.

Watch out for two things when you collect them — the log is written in batches, so you have to wait about 10 seconds before it shows up, and if you start Envoy again, the log file is emptied. It is safer to append as you go each time you create a failure.

Being able to read these flags is the longest-lasting skill in responding to Envoy and Istio outages — when you see a 5xx, look at the flags before you go digging through the app.

Wrap up three things

At least three lines in 08-notes.md: the first thing to look at when a route does not take effect, the case where retries become dangerous, and the difference between UT and URX.

The body must contain 순서, 재시도 and 플래그 (the Korean words for "order", "retry" and "flag").