One Snapshot, One Stream
Goal
Write yourself the smallest ADS server that reads a snapshot file and pushes CDS and LDS, and with an Envoy that receives all its configuration from that server, observe ACK, NACK, reconnection and hot restart on the stream.
Why it matters
This is what happens between istiod and the sidecars. But in a mesh that conversation is invisible, so when you meet a symptom like "I changed the configuration but only some behave the old way", you do not know where to start looking. If you read nonce, ACK, NACK and version_info once in the log of a server you wrote yourself, you learn what SYNCED in proxy-status and the rejection log in istiod point to.
Steps
- Start three upstreams —
python3 /opt/lab/envoy/upstream.py 8096 ok,… 8097 okand… 8098 slow. Then write a snapshot of version"1"to/root/envd-ads/snapshot.json: inclusters,pool(two endpoints, 8096 and 8097,connect_timeout1s) andslow(8098), and inlisteners,edgeat127.0.0.1:10088— the path/markerreturns the bodysnapshot=1directly,/slowgoes toslow, and the rest goes topool. Write the resources in Envoy's JSON shape as it is. - In
/root/envd-ads/ads_server.py, implementStreamAggregatedResourcesofenvoy.service.discovery.v3.AggregatedDiscoveryServiceand start it on127.0.0.1:18000(run it with/opt/xds/bin/python3). There are four requirements — (1) read the snapshot file and send all the resources of the requested type (CDS, LDS) together withversion_infoandnonce, (2) when the version in the file changes, push a new response to open streams without waiting for a request, (3) for every request received, leave one JSON line in/root/envd-ads/ads.log—event(requestif the nonce is empty,nackif there is an error_detail, otherwiseack),stream,node,type,version_info,response_nonceanderror, (4) when sending a response too, leaveevent: sentwith the version and nonce. - In
/root/envd-ads/bootstrap.yaml, write the node idenvd-ads-1(clusterenvd-ads), admin port9988, theads_configofdynamic_resources(gRPC, V3, clusterxds_cluster), andcds_configandlds_configasads: {}, and put only one static resource,xds_cluster(127.0.0.1:18000, HTTP/2), for reaching the control plane. Start it withenvoy -c /root/envd-ads/bootstrap.yaml --base-id 50 --restart-epoch 0 --concurrency 1 --drain-time-s 5 --parent-shutdown-time-s 10. - Change the snapshot to version
"2"— remove 8097 frompool, and make the/markerbodysnapshot=2. Write the file anew and swap it in withmv. After Envoy accepts it, in/root/envd-ads/04-push.txtwrite two lines:marker=(the/markerresponse body) andendpoints=(the number of pool endpoints remaining in/clusters). - Change the snapshot to version
"3", but setconnect_timeoutofpoolto"0s"and the/markerbody tosnapshot=3. Then write four lines to/root/envd-ads/05-nack.txt— theversion_infoof the CDS NACK you found in ads.log asnack_version_info=, the phrase starting withConnectTimeoutin the rejection reason asreason=, and Envoy's currentcluster_manager.cds.version_textandlistener_manager.lds.version_textascds_version=andlds_version=. - Stop the ADS server and in
/root/envd-ads/06-down.txtwrite two lines:traffic=(the status code of a/request) andconnected=(control_plane.connected_state). Then fix the snapshot to version"4"(revert connect_timeout from 3 to1sand make/markersnapshot=4), start the server again, and wait until Envoy reconnects and accepts version 4. Append theversion_infothat the first CDS request of the reconnected stream carried to the same file asresume_cds_version=. - Send one request to
/slow(it takes 3 seconds) in the background so that the result goes to/root/envd-ads/inflight.txt, and while that request is running, start an Envoy with epoch 1 with the same bootstrap (--base-id 50 --restart-epoch 1 --concurrency 1 --drain-time-s 5 --parent-shutdown-time-s 10). After the new process is LIVE and the old process has stepped back, in/root/envd-ads/07-restart.txtwrite three lines:epoch=(restart_epoch from/server_infoon the admin port),inflight=(the status code of the background request) andnew_stream_version=(the version_info of the first CDS request on the stream the new process opened —emptyif it is empty). - In
/root/envd-ads/08-report.md, write five lines —first_request_version=(the version_info of the very first CDS request Envoy sent,emptyif it is empty),nack_version_info=,resume_cds_version=(recorded in steps 5 and 6),epoch_after_restart=andinflight_code=(recorded in step 7) — and below them write what you learned in at least four lines.
Notes
- Run the ADS server with
/opt/xds/bin/python3(a virtual environment that has grpcio and the Envoy API protobuf). - Write the snapshot to a new file and swap it in with
mv. If the server reads a half-written file, it skips it with a JSON error (snapshot_error is left in the server log). - Start Envoy with
--restart-epoch 0in step 3 and do not start it again until step 7. If you start it again, the epoch of step 7 goes off. To stop it, usecurl -X POST localhost:9988/quitquitquit. - ads.log is one JSON line at a time. Pick lines out like
jq -c 'select(.event=="nack")' ads.log. - Common mistake —
pkill -f ads_server.py. The shell that contains that string dies too. Write it as'[a]ds_server.py'.
Write the configuration to push as a snapshot
Start three upstreams — python3 /opt/lab/envoy/upstream.py 8096 ok, … 8097 ok and … 8098 slow. Then write a snapshot of version "1" to /root/envd-ads/snapshot.json: in clusters, pool (two endpoints, 8096 and 8097, connect_timeout 1s) and slow (8098), and in listeners, edge at 127.0.0.1:10088 — the path /marker returns the body snapshot=1 directly, /slow goes to slow, and the rest goes to pool. Write the resources in Envoy's JSON shape as it is.
The job of a control plane is, in the end, to bundle "all the resources this proxy should have right now" into one version. That bundle is called a snapshot. You can write the resources in the JSON representation of the Cluster and Listener protobufs as it is, and typed_config inside the listener needs @type. The grader reads this JSON as real protobuf with xds-protos — even one wrong field name gets caught there.
Write the smallest ADS server
In /root/envd-ads/ads_server.py, implement StreamAggregatedResources of envoy.service.discovery.v3.AggregatedDiscoveryService and start it on 127.0.0.1:18000 (run it with /opt/xds/bin/python3). There are four requirements — (1) read the snapshot file and send all the resources of the requested type (CDS, LDS) together with version_info and nonce, (2) when the version in the file changes, push a new response to open streams without waiting for a request, (3) for every request received, leave one JSON line in /root/envd-ads/ads.log — event (request if the nonce is empty, nack if there is an error_detail, otherwise ack), stream, node, type, version_info, response_nonce and error, (4) when sending a response too, leave event: sent with the version and nonce.
A bidirectional gRPC stream, in Python, is "a function that takes a request iterator and yields responses". But since it has to push snapshot changes even while waiting for a request, it is convenient to have another thread read requests and put them in a queue, while the main loop waits briefly on the queue and checks the snapshot. Pack resources into google.protobuf.any_pb2.Any with Pack, and use json_format.ParseDict when converting JSON to protobuf (you have to import the HCM and router types inside the listener first for it to resolve). The grader opens a stream directly to this server and requests CDS.
Receive everything from the control plane with no static configuration
In /root/envd-ads/bootstrap.yaml, write the node id envd-ads-1 (cluster envd-ads), admin port 9988, the ads_config of dynamic_resources (gRPC, V3, cluster xds_cluster), and cds_config and lds_config as ads: {}, and put only one static resource, xds_cluster (127.0.0.1:18000, HTTP/2), for reaching the control plane. Start it with envoy -c /root/envd-ads/bootstrap.yaml --base-id 50 --restart-epoch 0 --concurrency 1 --drain-time-s 5 --parent-shutdown-time-s 10.
All that remains in the bootstrap is "the way to find the control plane". Listeners and business clusters are all received through ADS. ads: {} means "do not subscribe to CDS and LDS separately; receive them over one ADS stream", so the two types go back and forth in order on one stream. --base-id and --restart-epoch are attached from now on for the hot restart in step 7. You check that it connected well with the statistic control_plane.connected_state and the dynamic section of curl localhost:9988/config_dump.
Remove an endpoint without a restart
Change the snapshot to version "2" — remove 8097 from pool, and make the /marker body snapshot=2. Write the file anew and swap it in with mv. After Envoy accepts it, in /root/envd-ads/04-push.txt write two lines: marker= (the /marker response body) and endpoints= (the number of pool endpoints remaining in /clusters).
The server notices by itself that the file changed and sends a new response to the open stream. You confirm whether Envoy accepted it from the ack in ads.log (version_info is 2) and the statistic cluster_manager.cds.version_text. Write to a new file and move it into place so that the server does not read a half-written file.
How does a configuration Envoy rejected come back
Change the snapshot to version "3", but set connect_timeout of pool to "0s" and the /marker body to snapshot=3. Then write four lines to /root/envd-ads/05-nack.txt — the version_info of the CDS NACK you found in ads.log as nack_version_info=, the phrase starting with ConnectTimeout in the rejection reason as reason=, and Envoy's current cluster_manager.cds.version_text and listener_manager.lds.version_text as cds_version= and lds_version=.
A 0-second time limit is perfectly fine as protobuf, so it passes the server's JSON conversion. What catches it is Envoy's validation rules, and Envoy does not apply that response, but asks again with the same nonce, putting the reason in error_detail, and sends it back. See what the version_info is at that time. And CDS and LDS are applied separately even if they came from the same snapshot — the listener has no problem, so it is accepted. Also check what changed through /marker.
The control plane dies and comes back
Stop the ADS server and in /root/envd-ads/06-down.txt write two lines: traffic= (the status code of a / request) and connected= (control_plane.connected_state). Then fix the snapshot to version "4" (revert connect_timeout from 3 to 1s and make /marker snapshot=4), start the server again, and wait until Envoy reconnects and accepts version 4. Append the version_info that the first CDS request of the reconnected stream carried to the same file as resume_cds_version=.
An Envoy that has already received its configuration keeps running on it even without the control plane — the only thing that stops is receiving new configuration. When it reconnects, Envoy does not start empty-handed; it puts the last version it accepted for each type in its first request. When you start the server again, server_started is printed anew in the log, so look at the first Cluster request after that. Wrap the first letter of the pkill -f pattern in brackets, as in '[a]ds_server.py'.
Replace even the proxy with a new process without cutting it
Send one request to /slow (it takes 3 seconds) in the background so that the result goes to /root/envd-ads/inflight.txt, and while that request is running, start an Envoy with epoch 1 with the same bootstrap (--base-id 50 --restart-epoch 1 --concurrency 1 --drain-time-s 5 --parent-shutdown-time-s 10). After the new process is LIVE and the old process has stepped back, in /root/envd-ads/07-restart.txt write three lines: epoch= (restart_epoch from /server_info on the admin port), inflight= (the status code of the background request) and new_stream_version= (the version_info of the first CDS request on the stream the new process opened — empty if it is empty).
Hot restart is the method in which the new process takes over the listener sockets from the old process. So the port is never closed for a moment, and the old process finishes the requests it was handling during the drain time and then steps back. The two processes find each other through --base-id, so it must be the same, and the epoch goes up by one. The new process does not inherit the old process's xDS state — it connects to the control plane anew and receives from the start. You check whether the old process has stepped back with pgrep -af 'restart-epoch 0'.
Summarize what you saw on the stream
In /root/envd-ads/08-report.md, write five lines — first_request_version= (the version_info of the very first CDS request Envoy sent, empty if it is empty), nack_version_info=, resume_cds_version= (recorded in steps 5 and 6), epoch_after_restart= and inflight_code= (recorded in step 7) — and below them write what you learned in at least four lines.
Copy the values from ads.log and the evidence files. In the explanation lines, write in your own words "why the version_info of a NACK is not the rejected version" and "what a control plane outage and a proxy restart each make you lose".