TT Lab
Get started
Learn Learning paths Courses

Envoy Internals

Broken Config Shows Up Mid-Deploy

Continue in TT Lab

In one line

Envoy becomes a process by reading one configuration file. That one file has four chunks — admin port, identity, static resources, dynamic resources — and everything else is a layer inside them. If you can point to this structure with your finger, you will not get lost in a configuration you are seeing for the first time.

Why this was needed

When a proxy configuration is wrong, usually one of two things happens. One is that the process does not start at all, and the other is that it does start but requests go to the wrong place. The scary one in production is the first. A restart usually happens during a deployment, and by then the old process is already down. One character of @type or one space of indentation becomes an outage as it is.

Envoy knows about this problem, so it gives you a separate mode that only checks the configuration.

envoy --mode validate -c boot.yaml

This command reads the configuration through to the end, checks even the schema, and then finishes, leaving only an exit code, without binding a single port. So you can run it even on a machine where Envoy is already running, and you can put it in a CI step that builds the container image. If you changed the configuration and did not run this command, nobody yet knows whether that configuration will start.

How it works

For the top-level keys of the bootstrap, you only need to remember four.

Key What If missing
admin Admin port. Statistics, configuration dump, changing the log level No window to look inside with
node This proxy's identifying label (id, cluster) Fine if you use only static configuration. Required once you attach xDS
static_resources Listeners and clusters written in the file It listens to nothing
dynamic_resources Configuration fetched from outside (xDS) The configuration is fixed

Inside static_resources there are again five layers. listener → filter_chain → http_connection_manager → route_config → cluster. This is exactly the order in which a request comes in. Starting from the address and port, it picks which filter chain receives it, interprets it as HTTP, gets the destination cluster name from the route table, and the cluster of that name sends it to an actual endpoint. If you keep to this order when reading a configuration, you can quickly see which layer is empty.

You do not confirm by eye whether what you wrote in the bootstrap was actually applied. The first section of /config_dump on the admin port is BootstrapConfigDump, and the values Envoy read are in it as they are. The habit of putting the file and the dump side by side removes "I fixed it, so why did nothing change".

What it looks like in the field

The moment you start a second Envoy on one machine. Even though you moved all the ports out of the way, the second process dies with unable to bind domain socket with base_id=0. When Envoy starts, it grabs one shared-memory region (used to hand over statistics during a hot restart), and its name is set by the base id, whose default is 0. If you separate them with --base-id, both start. You will definitely run into this in a setup with two proxies on the same node (one for ingress, one for egress).

An incident where the admin port was left open to the outside. Some people set admin.address to 0.0.0.0 for convenience. But this port has no authentication, /quitquitquit ends the process, and /drain_listeners cuts traffic. A hole opened to scrape statistics becomes a button that takes the service down. Bind the admin port to loopback, and if necessary, put a separate proxy in front of it that exposes only the read path.

Experimenting without knowing about --concurrency. The default is the number of cores. Because the load-balancing state is separate for each worker thread, if you send eight requests and count whether they were spread evenly over three endpoints, the numbers come out different every time. An experiment that counts the distribution is deterministic only with --concurrency 1.

"I fixed it but it does not change." What you write in static_resources is read once when the process starts. Even if you edit the file, nothing happens until you start it again. That is why in production the habit of putting the file and /config_dump side by side matters — if it is not in the dump, it has not been applied yet, and if it is in the dump but the behavior is different, from then on it is time to suspect a different layer, not the configuration. The way to fetch configuration from outside and change it without starting again is dynamic_resources (xDS), and that is covered in the last module of this course.

Official documentation: Bootstrap configuration · Command line options · Administration interface

What you will do in the next lab

You write a bootstrap by hand and pass it through --mode validate, then get one character wrong and watch it get rejected. Next you start a second Envoy and run into the moment --base-id is needed, and at the end you end the process through the admin port and start it again.