TT Lab
Get started
Learn Learning paths Courses

ICA — Istio Certified Associate

503s Everywhere, Yet Every Config Says Applied

Continue in TT Lab

Goal

On a real Istio 1.31 mesh (k3s inside a VM), you narrow the four symptoms produced by the handed-over configuration — header rules being ignored, a 503 on /api, a 503 on the canary header, and a connection drop from a client without a sidecar — from evidence to cause and fix them. Finally, you take istiod down for a moment and see for yourself what continues and what is blocked when there is no control plane.

Why it matters

A 503 in a mesh does not have a single cause. The cluster the configuration points to may not exist at all (NC), the cluster may exist but there may be no Pod to select (UH), or the request may not even have been interpreted as HTTP (NR). That kubectl apply succeeded means only that the API server accepted it, and how Envoy uses that rule must be checked separately. So divide troubleshooting into three layers.

Steps

  1. Apply the handed-over configuration and record the uid and sidecar status of the four Pods.
  2. Find in the listener why the header rules are ignored (tcp_proxy), and correct the protocol with the service port name.
  3. Turn on the namespace access log with the Telemetry API.
  4. Confirm the 503 on /api with the access log (NC) and analyze, and fix the reference to a subset that does not exist.
  5. Confirm the 503 on the canary header with the access log (UH) and the empty endpoints, and fix the subset label.
  6. Confirm the drop of the sidecar-less legacy from the server log (NR) and put it into the mesh while keeping STRICT.
  7. Take istiod down and bring it back up, and record what happens to existing traffic, configuration writes, and new Pods respectively.
  8. Write a report transcribed from the evidence files.

Notes

Writing down the handed-over mesh's Pods and proxy list

Apply the handed-over configuration /opt/fixtures/ica-debug/incident.yaml as is and wait until the four Pods (web-v1, web-v2, client, legacy) in the namespace ica-debug are Ready. Save the list at that point to /root/ica-debug/inventory.json as a JSON array. Each element has three keys: name, uid, and sidecar (true if the Pod has istio-proxy). Do not fix the configuration yet.

You must first write down what has come into the mesh to interpret the later symptoms. On this VM, Istio 1.31 puts istio-proxy not in spec.containers but in spec.initContainers with restartPolicy: Always (a Kubernetes native sidecar). So you must look at both lists to judge sidecar correctly. And also look in istioctl proxy-status at which Pods are attached to istiod. The legacy Pod refused injection through a label. The uid is a value the cluster gave, so if you make it up, you fail.

I wrote a header rule, but traffic just goes 50/50

If you call http://web/ several times from client with the header x-canary: yes, v1 and v2 come out mixed (by the rule, only one side should come out). Before fixing, save the client sidecar's port 80 listener to /root/ica-debug/listener-before.json with istioctl proxy-config listeners client.ica-debug --port 80 -o json. Then change the port name of the Service web to http-web so that Istio handles this port as HTTP (the port number and targetPort stay the same).

A VirtualService's http rules are used only when Envoy interprets that port as HTTP. Istio decides the protocol by the service port's name prefix (<프로토콜>-이름, where the placeholder is the protocol) or appProtocol. In the saved listener JSON, look for whether the filter name of the Service ClusterIP listener is tcp_proxy or http_connection_manager. You can change the port name with a json patch of kubectl patch svc.

Making the proxy say what happened

Turn on the Envoy access log for the entire ica-debug namespace. Write a Telemetry resource into /root/ica-debug/telemetry.yaml and apply it — the provider name in spec.accessLogging is envoy. Once it is on, requests sent from client must leave one line each in kubectl logs client -c istio-proxy.

The minimal profile does not specify an access log file in meshConfig, so by default no log is left. Instead of fixing the mesh-wide configuration, you can turn it on at namespace scope with the Telemetry API (apiVersion telemetry.istio.io/v1). If you attach an x-request-id header to the request yourself, it is easy to find that line in the log.

Only /api is 503, yet the server Pods are fine

If you call http://web/api/orders from client, you get a 503. Save the one line of the client sidecar's access log that corresponds to that request as is to /root/ica-debug/nc.log, and then fix the cause: change the subset that the api rule of the VirtualService web points to into v1, which actually exists in the DestinationRule. After the fix, /api/orders must return 200 and the body v1, and there must be no IST0101 in istioctl analyze -n ica-debug.

The field right after the response code in an Envoy access log is the response flag. Check the flag and the detailed reason after it against the response flags table in the Envoy documentation. The fact that this request did not arrive at the server sidecar is also a clue. istioctl analyze finds the same cause just from the configuration.

With the canary header, no healthy upstream

If you call http://web/ from client with the header x-canary: yes, this time you get no healthy upstream and a 503. Save the one line of the client sidecar's access log for that request as is to /root/ica-debug/uh.log, and save the output of istioctl proxy-config endpoints client.ica-debug --cluster 'outbound|80|v2|web.ica-debug.svc.cluster.local' -o json at the same point to /root/ica-debug/endpoints-before.json. Then fix subset v2 of the DestinationRule web so that it selects the actual Pod label (version: v2).

Unlike NC, this time the cluster (outbound|80|v2|…) exists. The problem is the endpoints that went into that cluster. A subset filters the service endpoints once more by Pod labels, so even a one-character difference in the label value produces an empty list. Put kubectl get pod --show-labels and the DestinationRule's labels side by side. In 1.31, analyze tells you about this case as IST0173.

Only the old client without a sidecar gets its connection dropped

If you run curl http://web/ from the legacy Pod, the connection drops with no response (curl exit code 56). In the server-side (web-v1 or web-v2) sidecar log, find the line where the legacy Pod IP is recorded as the source and save it as is to /root/ica-debug/nr.log. Then fix it by putting legacy into the mesh while leaving STRICT as it is — write in /root/ica-debug/legacy.yaml a Pod definition with the label that refused injection removed, and recreate the Pod. After the fix, http://web/ from legacy must be 200.

A server sidecar under PeerAuthentication STRICT accepts only connections arriving over mTLS. A client without a sidecar sends plaintext, and the server sidecar cannot find a filter chain that fits that connection — it is at a stage before being interpreted as an HTTP request, so the method and path fields in the log are empty. Lowering to PERMISSIVE also makes the connection work, but it is not the answer to this step. A Pod's labels and annotations are not injected even if you change them while it is running, so you must delete and recreate it.

What stops while istiod is down

Take istiod down with kubectl -n istio-system scale deploy istiod --replicas=0 and confirm that the Pod is gone, then try three things and write the results in three lines in /root/ica-debug/istiod-down.txt — after traffic=, the body from calling http://web/ from client; after config-write=, the output of kubectl -n ica-debug annotate virtualservice web lab.example/probe=1 --overwrite; and after new-pod=, the output of kubectl -n ica-debug run probe --image=curlimages/curl:8.10.1 --restart=Never --command -- sleep 60. After recording them, return istiod to replicas 1 and wait until the four proxies (client, web-v1, web-v2, legacy) appear in istioctl proxy-status again.

The control plane distributes configuration, issues certificates, and receives the injection and validation webhooks. An Envoy that has already received its configuration keeps working with the last configuration even without istiod. On the other hand, requests that write Istio resources and creating new Pods require the API server to call a webhook, so check with kubectl get validatingwebhookconfiguration,mutatingwebhookconfiguration -o yaml what the webhook's failurePolicy is. If you forget to bring it back, every later configuration change is blocked.

A report that classifies the four 503s and drops by evidence

Write a single JSON object in /root/ica-debug/report.json. The meaning of the keys and values — header_rules_ignored_filter: the last fragment of the network filter name seen in the listener before the fix (for example, xxx_proxy); api_503_flag, canary_503_flag, legacy_reset_flag: the response flags recorded in nc.log, uh.log, and nr.log; istiod_down_traffic: the body received while istiod was gone; istiod_down_config_write: accepted if the configuration write went through at that time, rejected if it was rejected. The values must match the evidence files you left, and all four symptoms must be fixed by now.

You transcribe the report from evidence, not from memory. Read the field after the response code in the log lines, and find the filter name of the Service ClusterIP listener in the listener JSON. The grader reads the same files and cross-checks them, and sends the header, /api, and legacy requests again now to check.