TT Lab
Get started
Learn Learning paths Courses

Envoy Internals

Why a NACK's version_info Is Not the Rejected Version

Continue in TT Lab

In one line

ADS opens one gRPC stream between Envoy and the control plane and sends and receives configuration of every type (CDS, LDS and so on) over it. Responses carry a version and a nonce, and Envoy reports the result of trying to apply them in the next request with the same nonce — if the reason is empty it is an ACK, and if it is filled in it is a NACK.

Why this was needed

File subscription also changes configuration with no downtime, but there is no way to distribute files to thousands of proxies, and no way to know whether the proxy accepted that configuration. That is exactly what the control plane needs to know first — was the configuration I sent applied or rejected, and if rejected, why.

So xDS has a streaming mode. The proxy opens the connection first (it is an outgoing connection from the proxy side, so even a proxy behind a firewall can connect), and the control plane pushes over that stream whenever it is needed. You can also open a separate stream for each type, but then you cannot guarantee an order such as "receive the cluster first and the listener that uses it later". ADS (Aggregated Discovery Service) gathers all types into one stream and lets the control plane decide the order. What Istio's istiod and the sidecars use is this mode.

How it works

One exchange — based on the State of the World (SotW) mode.

Envoy → 요청  type=Cluster  version_info=""   response_nonce=""     (처음)
서버  → 응답  type=Cluster  version_info="1"  nonce="1"  resources=[pool, slow]
Envoy → 요청  type=Cluster  version_info="1"  response_nonce="1"    ← ACK
          … 파일이 바뀌어 서버가 밀어 넣는다 …
서버  → 응답  type=Cluster  version_info="3"  nonce="5"  resources=[…]
Envoy → 요청  type=Cluster  version_info="2"  response_nonce="5"
              error_detail="…ConnectTimeout: value must be greater than 0s"   ← NACK

There are three things to read.

Field Meaning
response_nonce Which response this is an answer to. If it is empty, it is a new subscription request
error_detail If filled in, it is a NACK. The reason comes as a string
version_info The last version accepted. Even in a NACK it is not the version that was rejected

The last row is the most confusing. The version_info of a NACK request is not the rejected 3 but the 2 that is still in use. The rejected response is pointed to by the nonce.

In the State of the World mode, a response carries all the resources of that type. Even if you change one, everything is sent again. In a mesh with many resources this is a burden, so there is a separate incremental (Delta) mode that sends only what changed.

It is applied separately for each type. If in the same snapshot version 3 the listener is fine and only the cluster is wrong, LDS accepts 3 and CDS stays at 2. That is the moment the statistics cluster_manager.cds.version_text and listener_manager.lds.version_text show different values.

If the control plane disappears, the proxy keeps running on the current configuration (control_plane.connected_state is 0). When it reconnects, Envoy does not start empty-handed; it puts in its first request the last version it accepted for each type — the server can look at that and decide what to send again.

Hot restart — this is the mechanism used when you change the proxy itself (replacing the binary, changing the bootstrap). The new process (epoch N+1) finds the old process with the same --base-id and takes over the listener sockets, and the old process finishes the requests it was handling during --drain-time-s and then steps back at --parent-shutdown-time-s. The port is never closed for a moment. In exchange, the new process does not inherit the xDS state; it connects to the control plane anew and receives from the start.

What it looks like in the field

"I changed the configuration but only some proxies behave the old way." A NACK is the most common cause. In Istio, istioctl proxy-status shows for each proxy whether CDS and LDS are SYNCED, and if there is a rejection, the reason is left in the istiod log. Traffic keeps flowing, so if nobody is looking, weeks go by.

"I restarted istiod and all the sidecars reconnected at once." When they reconnect, each states its last version, so the control plane has to answer those thousands of first requests. That is why you run the control plane as several instances.

"When I restart Envoy, connections get cut." In a place with many long-lived connections, such as a gateway, you need a hot restart or a sequential replacement that goes through drain. A Kubernetes sidecar is replaced Pod by Pod, so it does not use hot restart — that is also why Istio starts it with --disable-hot-restart.

Official documentation: xDS REST and gRPC protocol · Aggregated Discovery Service · Hot restart · Command line options

What you will do in the next lab

You write in Python an ADS server that reads one snapshot file and pushes CDS and LDS, and start an Envoy that receives everything from that server with no static configuration. You remove an endpoint and see it take effect with no restart, send a deliberately wrong cluster and see a NACK and its reason come back, and read from the log the version Envoy states when the server dies and comes back. Finally you hot restart that Envoy and see that requests in progress are not cut.