TT Lab
Get started
Learn Learning paths Courses

Envoy Internals

Screen Before Deploy, Then Run a Second Proxy

Continue in TT Lab

Goal

Write a bootstrap by hand, check only the configuration before deployment, and start a second Envoy on one machine.

Why it matters

A proxy with a wrong configuration shows it the moment it restarts. By then the old process is already down, so there is no time to roll back. --mode validate reads only the configuration without binding ports, which moves this incident to before deployment. If you can point to the layered structure of the bootstrap, you can read someone else's configuration in a few seconds, and if you do not go through the dangers of --base-id and the admin port once, you will certainly meet them for the first time in production.

Steps

  1. Write a bootstrap to /root/envd-boot/boot.yaml — admin port 9911, a listener named edge at 127.0.0.1:10011, stat_prefix set to edge, and a cluster named origin at 127.0.0.1:8081. Run envoy --mode validate -c /root/envd-boot/boot.yaml and save its output and exit code to /root/envd-boot/01-validate.txt. The last line must be rc=0.
  2. Copy /root/envd-boot/boot.yaml to /root/envd-boot/boot-typo.yaml, then delete one character so HttpConnectionManager becomes HttpConnectionManger. Check it with envoy --mode validate and save the output and exit code to /root/envd-boot/02-typo.txt (the last line is rc=1).
  3. Add node to /root/envd-boot/boot.yaml — id is edge-1 and cluster is envd-edge. Start an upstream on 8081 and start Envoy, then from /config_dump on the admin port pull out only the node in the bootstrap section and save it to /root/envd-boot/03-node.json.
  4. Read /root/envd-boot/boot.yaml with yq and write five values to /root/envd-boot/04-layers.txt as five lines of 열쇠=값 (the placeholders are the key and the value) — listener= (the listener name), filter= (the network filter name), stat_prefix=, route_cluster= (the cluster the route points to), cluster_endpoint= (the endpoint of that cluster, in the format 주소:포트, where the placeholders are the address and the port).
  5. Create /root/envd-boot/boot-second.yaml, a copy of /root/envd-boot/boot.yaml with only the ports changed (admin 9912, listener 10012). First start it with no options and confirm the failure, then add --base-id 7 and make it succeed. In /root/envd-boot/05-baseid.txt, write without_base_id_rc= (a nonzero value), a line containing the failure reason, and with_base_id_ready= (the /ready response of the second Envoy).
  6. Start Envoy again with --concurrency 1 --log-level info, and from /server_info on the admin port pull out only command_line_options and save it to /root/envd-boot/06-cli.json. concurrency must be 1 and log_level must be info.
  7. POST to /quitquitquit on the admin port to end Envoy, and in /root/envd-boot/07-quit.txt write after_quit= (the HTTP code of /ready afterward). Then start it again and append after_restart= (the response string of /ready). At the end, Envoy must be alive on 9911.
  8. In /root/envd-boot/08-report.md, write four lines validate_rc=, typo_rc=, second_envoy= and admin_bind= (respectively the exit codes of step 1 and step 2, the name of the option the second Envoy needed, and the address the admin port should be bound to), and below them write what you learned in at least four lines.

Notes

Check only the configuration before you deploy

Write a bootstrap to /root/envd-boot/boot.yaml — admin port 9911, a listener named edge at 127.0.0.1:10011, stat_prefix set to edge, and a cluster named origin at 127.0.0.1:8081. Run envoy --mode validate -c /root/envd-boot/boot.yaml and save its output and exit code to /root/envd-boot/01-validate.txt. The last line must be rc=0.

--mode validate reads the configuration, checks even the schema, and finishes without binding a single port. So you can run it even on a machine where Envoy is already running, and in CI. The reason to write the exit code separately is that people misjudge by looking at the output alone — if you go through a pipe, $? becomes that of the last command, so capture it without a pipe.

Get one character of a name wrong and the process does not start at all

Copy /root/envd-boot/boot.yaml to /root/envd-boot/boot-typo.yaml, then delete one character so HttpConnectionManager becomes HttpConnectionManger. Check it with envoy --mode validate and save the output and exit code to /root/envd-boot/02-typo.txt (the last line is rc=1).

@type is not a comment; it is the key that chooses which protocol buffer message to interpret it as. If the name is not in the list, Envoy has no way to interpret that filter and rejects the whole configuration. In production this typo shows itself the moment you restart, and by then the old process is already down.

Write into the configuration who this proxy is

Add node to /root/envd-boot/boot.yaml — id is edge-1 and cluster is envd-edge. Start an upstream on 8081 and start Envoy, then from /config_dump on the admin port pull out only the node in the bootstrap section and save it to /root/envd-boot/03-node.json.

node is the label by which this proxy introduces itself to the control plane. It starts without it when you use only static configuration, but the moment you attach xDS, the server uses this value to choose who gets which configuration. The first section of /config_dump is BootstrapConfigDump, and bootstrap.node is inside it. Pull out only that spot with jq.

Point to the five layers by path

Read /root/envd-boot/boot.yaml with yq and write five values to /root/envd-boot/04-layers.txt as five lines of 열쇠=값 (the placeholders are the key and the value) — listener= (the listener name), filter= (the network filter name), stat_prefix=, route_cluster= (the cluster the route points to), cluster_endpoint= (the endpoint of that cluster, in the format 주소:포트, where the placeholders are the address and the port).

The order for reading Envoy configuration is always the same — listener → filter_chain → http_connection_manager → route_config → cluster. If you can point to these five places with your finger, you will not get lost in a configuration you are seeing for the first time. Go down one layer at a time, as in yq '.static_resources.listeners[0].name'. Pull the values out with the tool instead of copying them by eye, and you will not get them wrong.

The second Envoy does not start

Create /root/envd-boot/boot-second.yaml, a copy of /root/envd-boot/boot.yaml with only the ports changed (admin 9912, listener 10012). First start it with no options and confirm the failure, then add --base-id 7 and make it succeed. In /root/envd-boot/05-baseid.txt, write without_base_id_rc= (a nonzero value), a line containing the failure reason, and with_base_id_ready= (the /ready response of the second Envoy).

When Envoy starts, it grabs one shared-memory region (the place used to hand over statistics during a hot restart). The name of that region is set by the base id and its default is 0, so a second process on the same machine collides at that spot even if you have moved all the ports out of the way. The first line of the failure log says so directly. The side that fails ends immediately, so you can just run it without setsid and take the exit code.

Check where the command-line options remain

Start Envoy again with --concurrency 1 --log-level info, and from /server_info on the admin port pull out only command_line_options and save it to /root/envd-boot/06-cli.json. concurrency must be 1 and log_level must be info.

The command line is exactly the place where things that are not in the configuration file change the behavior. --concurrency is the number of worker threads, and its default is the number of cores, so the load-balancing state is separate for each worker — the most common cause of experiments that count the distribution wobbling. Do not guess what was actually applied; read it from /server_info. Pull out only that spot with jq '.command_line_options'.

The admin port even holds a button that kills it

POST to /quitquitquit on the admin port to end Envoy, and in /root/envd-boot/07-quit.txt write after_quit= (the HTTP code of /ready afterward). Then start it again and append after_restart= (the response string of /ready). At the end, Envoy must be alive on 9911.

The admin port has more than statistics — /quitquitquit ends the process and /drain_listeners cuts traffic. There is no authentication, so anyone who can reach this port can take the proxy down. That is why you must not set admin.address to 0.0.0.0. If you curl a dead server, the connection itself fails, so the HTTP code comes out as 000.

Summarize it as a pre-deployment checklist

In /root/envd-boot/08-report.md, write four lines validate_rc=, typo_rc=, second_envoy= and admin_bind= (respectively the exit codes of step 1 and step 2, the name of the option the second Envoy needed, and the address the admin port should be bound to), and below them write what you learned in at least four lines.

A checklist is something other people read. When you write a value, write it so that where the value came from is clear, and in the explanation lines, write "so from now on I will do this" instead of "I did this". For second_envoy=, write the option name as it is.