Swap Config Without a Restart
Goal
After confirming the limits of static configuration, switch the same proxy to the file-subscription method, swap clusters and listeners in without a restart, and also see a bad update get rejected.
Why it matters
This is the core of what service meshes and ingress controllers do. But if you only read it in documentation, wrong sentences like "if the control plane dies, the mesh dies" stay in your head. Once you cause a rejection yourself, you keep both facts together: traffic keeps flowing after a rejection, and therefore weeks can pass with nobody knowing. Only someone who knows both sets an alert on update_rejected.
Steps
- Start two upstreams as
ok(8096and8097). In/root/envd-xds/static.yaml, put a static configuration in which/returnsstatic=v1and start it (admin 9977, listener127.0.0.1:10077), then while it is running, editv1tov2in that file. Request before and after the edit and in/root/envd-xds/01-static.txtwrite two lines,before=andafter_edit=(do not start it again). - Create
/root/envd-xds/dyn.yaml— with nostatic_resources, onlynode(idenvd-xds-1, clusterenvd-xds) anddynamic_resources, wherelds_configsubscribes to/root/envd-xds/xds/lds.yamlandcds_configsubscribes to/root/envd-xds/xds/cds.yamlthroughpath_config_source. Create the two resource files as well — the listener returnslds=v1at/markerand sends everything else to the clusterpool, and the clusterpoolhas two endpoints,8096and8097. After you start it, request/markerand/and in/root/envd-xds/02-subscribe.txtwrite two lines,marker=andupstream_ok=(yes if the/request is 200). - Remove the endpoint
8097from/root/envd-xds/xds/cds.yaml(leave just one endpoint). Leaving the proxy as it is, wait until it takes effect, then request eight times and in/root/envd-xds/03-reload.txtwrite three lines,p8096=,p8097=andtotal=. - Pull out only the cluster section from
/config_dumpand save it to/root/envd-xds/04-dump.json, and in/root/envd-xds/04-dump.txtwrite three lines:static_clusters=(the number of static clusters),dynamic_clusters=(the number of dynamic clusters) anddynamic_name=(the name of the dynamic cluster). - Change the
/markerresponse in/root/envd-xds/xds/lds.yamlfromlds=v1tolds=v2. Leaving the proxy as it is, wait until it takes effect, then in/root/envd-xds/05-lds.txtwrite two lines:marker_after=(the/markerresponse) andupstream_still_ok=(yes if the/request is still 200). - Delete the cluster's
nameline in/root/envd-xds/xds/cds.yamlto send a deliberately bad update (the schema requires the name). After a moment, in/root/envd-xds/06-rejected.txtwrite three lines:still_serving=(yes if the/request is still 200),update_rejected=(the value of thecluster_manager.cds.update_rejectedstatistic) andendpoints=(the number ofpoolendpoints remaining in/clusters). - In
/root/envd-xds/07-stats.txt, write four lines,cds_success=,cds_rejected=,lds_success=andcds_reload=— these are the values of the statisticscluster_manager.cds.update_success,cluster_manager.cds.update_rejected,listener_manager.lds.update_successandcluster_manager.cds.config_reloadrespectively. - In
/root/envd-xds/08-report.md, write four lines —static_needs_restart=(yes if the result of step 1 was "it did not change"),reload_count=(thecds_reloadof step 7),rejected_kept_old=(yes if traffic kept flowing after the rejection in step 6) andendpoints_after=(the number of endpoints remaining after step 3) — and below them write what you learned in at least four lines.
Notes
- This environment has no real gRPC control plane, so file subscription (
path_config_source) stands in for it. The shape of what is sent and the way the proxy processes it are the same, but transport-layer behavior, such as resubscribing when the connection drops, cannot be seen here. - After you edit the file, it does not take effect right away. Wait with a loop that runs until the result you want appears instead of a fixed
sleep(up to about forty iterations). - Do not overwrite a resource file in place; write a new file and swap it in with
mv. File subscription wakes up on a move — overwriting in place does not bring an update. It is the same reason Kubernetes applies a ConfigMap by swapping a symbolic link. - When you start Envoy, use
setsid --fork nohup envoy -c <파일> --log-level warn --concurrency 1 > <로그> 2>&1 </dev/null(the placeholders are the file and the log), and before you start it again, clean up withpkill -x envoy. - The shape of a resource file is a list under
resources:that states"@type". For a listener the type is Listener, and for a cluster it is Cluster. - Common mistake — starting the proxy again in step 3 and step 5. The point of this lab is to see things change without starting it again.
Static configuration does not change even if you edit the file
Start two upstreams as ok (8096 and 8097). In /root/envd-xds/static.yaml, put a static configuration in which / returns static=v1 and start it (admin 9977, listener 127.0.0.1:10077), then while it is running, edit v1 to v2 in that file. Request before and after the edit and in /root/envd-xds/01-static.txt write two lines, before= and after_edit= (do not start it again).
static_resources is read once when the process starts. So even if you edit the file, nothing happens until you start it again. This is what is behind "changing proxy configuration needs a restart", and it is also the problem a service mesh set out to solve. sed -i is handy for the edit. Be careful not to start it again — the point of this step is to show "it does not change".
Switch it to receive the configuration from outside
Create /root/envd-xds/dyn.yaml — with no static_resources, only node (id envd-xds-1, cluster envd-xds) and dynamic_resources, where lds_config subscribes to /root/envd-xds/xds/lds.yaml and cds_config subscribes to /root/envd-xds/xds/cds.yaml through path_config_source. Create the two resource files as well — the listener returns lds=v1 at /marker and sends everything else to the cluster pool, and the cluster pool has two endpoints, 8096 and 8097. After you start it, request /marker and / and in /root/envd-xds/02-subscribe.txt write two lines, marker= and upstream_ok= (yes if the / request is 200).
The bootstrap does not need static_resources at all. In that case the proxy subscribes as soon as it starts, and only starts to actually listen once the resources arrive. The shape of a resource file is a list under resources: that states "@type" — for a listener the type is Listener, and for a cluster it is Cluster. A real control plane sends this over gRPC, but the shape of what is sent is the same whether it is a file or gRPC.
Edit the file and it takes effect without starting again
Remove the endpoint 8097 from /root/envd-xds/xds/cds.yaml (leave just one endpoint). Leaving the proxy as it is, wait until it takes effect, then request eight times and in /root/envd-xds/03-reload.txt write three lines, p8096=, p8097= and total=.
The file-subscription method notices that the file changed and reads it again. It is not immediate and can take a few seconds, so use a loop that runs until the result you want appears instead of a fixed sleep. This is what happens in a service mesh every time a Pod starts or stops — it does not restart the proxy; it only swaps in the endpoint list. You can also check whether it took effect with /clusters. Note: if you overwrite the file in place, no update arrives. Write a new file and swap it in with mv.
Configuration received dynamically is in a different place in the dump
Pull out only the cluster section from /config_dump and save it to /root/envd-xds/04-dump.json, and in /root/envd-xds/04-dump.txt write three lines: static_clusters= (the number of static clusters), dynamic_clusters= (the number of dynamic clusters) and dynamic_name= (the name of the dynamic cluster).
The cluster section of /config_dump has static_clusters and dynamic_active_clusters separately, because what was written in the bootstrap and what was received from outside have to be told apart. In production, the place you look when you check "did the configuration I sent arrive" is this dynamic side. If the static side is empty and the dynamic side has it, the subscription is running properly. You can also pick a section with ?resource=, but in this step it is easier to fetch everything and count both places together with jq.
Swap the listener in with no downtime too
Change the /marker response in /root/envd-xds/xds/lds.yaml from lds=v1 to lds=v2. Leaving the proxy as it is, wait until it takes effect, then in /root/envd-xds/05-lds.txt write two lines: marker_after= (the /marker response) and upstream_still_ok= (yes if the / request is still 200).
Swapping a listener is heavier than a cluster — because it means opening sockets again. Envoy prepares the new listener and then swaps it in while draining the old listener, so requests are not cut in between. That is why you do not have to deploy just to fix one routing rule. You did not touch the cluster configuration, so the path to the upstream should stay the same — check that too.
A bad update is rejected and the old configuration remains
Delete the cluster's name line in /root/envd-xds/xds/cds.yaml to send a deliberately bad update (the schema requires the name). After a moment, in /root/envd-xds/06-rejected.txt write three lines: still_serving= (yes if the / request is still 200), update_rejected= (the value of the cluster_manager.cds.update_rejected statistic) and endpoints= (the number of pool endpoints remaining in /clusters).
This is the most important property in xDS. If an update is bad, the proxy rejects it and keeps using the last configuration that succeeded. Even if the control plane sends broken configuration, traffic keeps flowing. So the saying "if the control plane dies, the mesh dies" is not accurate — it simply cannot receive new configuration, and keeps running on the current configuration. In exchange, nobody tells you, so you have to watch the rejection statistics and the version. update_rejected is in the statistic name. One thing to watch out for — not every mistake is rejected. There are also things where the default validator only warns and moves on, such as an unknown enum value. Only what the schema cannot accept, such as a missing required field, is caught as a rejection.
Check whether an update arrived by the numbers
In /root/envd-xds/07-stats.txt, write four lines, cds_success=, cds_rejected=, lds_success= and cds_reload= — these are the values of the statistics cluster_manager.cds.update_success, cluster_manager.cds.update_rejected, listener_manager.lds.update_success and cluster_manager.cds.config_reload respectively.
These are the numbers you look at first in an environment that has a control plane. If update_success is stuck, configuration is not arriving, and if update_rejected is rising, it arrives but the proxy cannot accept it. The two are entirely different problems, so the place to fix them differs — the former is a connection or subscription problem, and the latter is a problem with the configuration the sending side produced. In this lab you deliberately created a rejection in an earlier step, so you will see that number.
Leave a dynamic configuration operations memo
In /root/envd-xds/08-report.md, write four lines — static_needs_restart= (yes if the result of step 1 was "it did not change"), reload_count= (the cds_reload of step 7), rejected_kept_old= (yes if traffic kept flowing after the rejection in step 6) and endpoints_after= (the number of endpoints remaining after step 3) — and below them write what you learned in at least four lines.
In the explanation lines, write in your own words "when the control plane dies, what stops and what continues". This one sentence is used most often in responding to mesh outages. Take the values from the files of the earlier steps.