What Actually Stops When the Control Plane Dies
In one line
xDS is the channel through which a proxy receives its configuration from outside. When it gives listeners it is called LDS, when it gives clusters CDS, and when it gives only endpoints EDS. There are two key properties — it swaps configuration in without a restart, and it rejects a bad update and keeps using the old configuration.
Why this was needed
Every configuration so far was written in a file and the process was started again. That is bearable with a few servers. But in Kubernetes, Pods start and stop hundreds of times a day. If you restart the proxy every time the endpoint list changes, what happens to the requests during the restart, and how do you prevent a situation where one restart causes another?
So the direction was turned around. Leave the proxy alone and push in only the configuration. When it starts, the proxy introduces itself ("this is the kind of proxy I am", node) and subscribes to the resources it needs. The control plane watches the state of the cluster and sends what changed to that proxy. This is the seam where the data plane and control plane of a service mesh split, and it is also the whole of what an ingress controller does.
How it works
The name splits by what it gives.
| Name | What it gives | When it changes |
|---|---|---|
| LDS | Listeners | When a port or a filter chain changes |
| CDS | Clusters | When services are added or removed |
| EDS | Endpoints only | When Pods start and stop (most often) |
| RDS | Route tables | When routing rules change |
| SDS | Certificates | When certificates are renewed |
The reason endpoints are separate is that they change most often. Sending the whole cluster definition again every time one Pod starts is wasteful.
There are three ways to receive it. Opening a stream over gRPC is what real meshes use, there is also a way of asking periodically over REST, and there is also a way of subscribing to a file. The last one is used for experiments and tests, and the core of the protocol (subscription, update, rejection, version) is the same in all three.
Updates are atomic, and if they fail, it rolls back. This is the most important property in xDS. If the configuration received does not match the schema or cannot be interpreted, the proxy rejects that whole update and keeps using the last configuration that succeeded. Traffic is not cut. So the saying "if the control plane dies, the mesh dies" is not accurate — it simply cannot receive new configuration, and keeps running on the current configuration.
In exchange, nobody tells you. That is why the numbers to watch are fixed.
cluster_manager.cds.update_success 갱신이 도착해 반영된 횟수
cluster_manager.cds.update_rejected 도착했지만 거절된 횟수
listener_manager.lds.update_success 리스너 쪽 같은 숫자
If update_success is stuck, configuration is not arriving, and if update_rejected is rising, it arrives but cannot be accepted. The two are entirely different problems and the place to fix them differs.
There is a separate place to check whether it arrived. /config_dump shows static resources and dynamic resources separately (static_clusters and dynamic_active_clusters). When you check "did the configuration I sent arrive", look at the dynamic side.
What it looks like in the field
"I changed the configuration but it is not taking effect." It splits into two branches. If it is not in the dynamic side of the dump, it has not arrived (a subscription or connection problem), and if it is there but the behavior is different, the configuration itself differs from what was intended. If you separate these two first, the search range is cut in half.
"I restarted the control plane and nothing happened." That is normal. The proxy keeps running holding the last configuration. The problem is that if that state lasts long, newly started Pods cannot join the mesh. The fact that it is fine now and the fact that it will soon become a problem are both true together.
When rejections quietly pile up. Even if something goes wrong with the configuration the control plane produced, traffic still flows, so nobody knows. A few weeks later it shows up as "only this service goes through the old routing". The answer is to set an alert on update_rejected.
Official documentation: xDS REST and gRPC protocol · Bootstrap configuration · Administration interface
What you will do in the next lab
First you confirm that static configuration does not change even when you edit the file, and then you switch the same proxy to the file-subscription method. You remove an endpoint from the cluster resource and count that it takes effect without a restart, swap the listener resource in with no downtime too, and finally send a deliberately broken configuration and see for yourself that it is rejected yet traffic keeps flowing. Then you confirm which number records that event.
This lab environment has no real gRPC control plane, so file subscription stands in for it. The shape of what is sent and the way the proxy processes it are the same, but behavior of the transport layer, such as resubscribing when the connection drops, cannot be seen here.