TT Lab
Get started
Learn Learning paths Courses

Envoy Internals

503 Is the Cluster Talking, Not the Route

Continue in TT Lab

In one line

A route says only as far as the cluster name. How many addresses stand behind that name, how they are discovered, which of them are healthy, and what to do when too few healthy ones are left are all decided by the cluster.

Why this was needed

A report saying "the route is right but I get a 503" is not a routing problem. It means the cluster the route points to has nowhere suitable to send to. So the screen you actually look at often when operating a proxy is not the route table but the cluster list — how many servers each cluster has and how many of them are healthy.

You need to tell two things apart here. How endpoints are discovered and how you judge which of them are healthy are separate axes. The first is the cluster's type, and the second is the health check. If you mix the two, you get lost on questions like "it was removed from DNS, so why is it still being sent to".

How it works

Four ways to discover endpoints.

type How When
STATIC Write the addresses in the configuration as they are When the addresses are fixed
STRICT_DNS Resolve the name periodically and make all the addresses in the response endpoints When several addresses come back, as with a headless Service
LOGICAL_DNS Hold on to only one of the resolved results A large peer with long-lived connections (for example an external API)
EDS The control plane pushes in the list Service mesh. Updated every time a Pod starts or stops

Two ways to judge which are healthy. The two work in opposite directions.

The two are not exclusive. Used together, it becomes "filter out ahead of time by probing, and filter out by results what still leaks through". health_flags in /clusters shows a different mark for each — an active check failure is /failed_active_hc.

When too few healthy ones are left. Envoy has two mechanisms.

One is panic mode. When the healthy ratio falls below a threshold (50% by default), it ignores health information and sends to all of them. It looks strange at first, but the reasoning is simple — if the checks may be wrong and you send nowhere, that is a certain outage, and if you send, at least some requests may survive. You see whether it happened in the cluster.<이름>.lb_healthy_panic statistic (the placeholder is the cluster name). If this number is rising, load balancing has already lost its meaning.

The other is priority. If you give each endpoint group a priority, normally only priority 0 is used, and priority 1 takes the amount by which the healthy ratio of priority 0 has dropped. If priority 0 is wiped out, everything goes to priority 1. A setup that leaves standby servers in another region idle normally and uses them only during an outage is built this way.

What it looks like in the field

"I removed it from DNS, but it still goes to that server." STRICT_DNS resolves again periodically, but until that period passes, it uses the old list. With LOGICAL_DNS, it may hold on even longer. You have to plan the order of your work on the assumption that there is time between removing and the change taking effect.

"I turned on the health check and the backend CPU went up." If twenty proxies probe every second, from the server's point of view that is twenty extra requests per second. Calculate the interval and the number of proxies together, and keep the health check path light (a path that does not touch the database).

Misunderstanding panic mode. The question "they are all dead, so why does it keep sending" comes up. The answer is in the statistics. What you fix at this point is not Envoy but the health check threshold or the backend.

Official documentation: Service discovery · Health checking · Cluster configuration

What you will do in the next lab

You create a cluster with three endpoints, set up one more cluster that finds endpoints by name, and watch an active health check take out the one broken endpoint. Next you confirm panic mode with a cluster where everything is broken, and the step-over to the standby tier with a cluster given priorities, and at the end you build a list table of the four clusters from the statistics.