TT Lab
Get started
Learn Learning paths Courses

ICA — Istio Certified Associate

Sizing Resilience Settings

Continue in TT Lab

Goal

While writing retries, timeouts, circuit breakers, fault injection, and mirroring yourself, you understand what constraints each value places on the others.

Why it matters

Resilience settings that are copied over usually do not work or, worse, amplify failures. If you put retries on every layer of a call chain, then in a 3-hop chain, in the worst case, 9 times the requests, and with 4 hops 27 times, pile onto the innermost service. It is a structure in which retries become the final blow to an already dying service.

Timeouts are the same. timeout is the overall budget including retries and backoff, and perTryTimeout is the cap for each attempt. For a fast failure, retries can run even when the overall budget is smaller than the sum of the per-try caps. Conversely, if one attempt uses up the entire budget, the next attempt may not start. The 3 seconds / 1 second / 2 additional retries in this lab are values for writing the configuration, not a guarantee that three attempts each necessarily run for 1 second. The actual count must be observed together with the failure condition, the response time, and the backoff. Even when the outer caller finishes waiting, there is no guarantee that the inner business processing is canceled, so design duplicate prevention and cancellation propagation separately so that an already processed write is not executed again.

A circuit breaker whose values are too large is effectively the same as off, and one whose values are too small produces 503s even in normal times. So you size it based on measured concurrency and put in a safety valve with maxEjectionPercent. This lab is training in carrying that sense into manifests.

You actually apply the basic Kubernetes resources and write the Istio CRDs as files under /root/ica-resilience/. Because this is a kwok-based design lab, do not regard real Envoy HTTP responses, retry counts, or business-level duplicate processing as verified here.

Steps

  1. Create the namespace ica-resilience and attach the label istio-injection=enabled.
  2. In the namespace ica-resilience, deploy a Deployment inventory with 3 replicas. The Pod label is app=inventory and the image is nginx:1.27-alpine.
  3. In the same namespace, create two Services. They are inventory and inventory-canary, both with port 8080 and port name http. The selector of inventory is app=inventory.
  4. In /root/ica-resilience/vs-inventory-retry.yaml, write a VirtualService. spec.hosts[0] is inventory.ica-resilience.svc.cluster.local; in the http rule, include timeout: 3s, retries.attempts: 2, retries.perTryTimeout: 1s, and in retries.retryOn, include 5xx and connect-failure.
  5. In /root/ica-resilience/vs-inventory-fault.yaml, write a VirtualService. There are two rules. The first rule applies only to requests whose header x-chaos-test is true and has fault.delay.fixedDelay: 3s / fault.delay.percentage.value: 50 / fault.abort.httpStatus: 503 / fault.abort.percentage.value: 10. The second rule just routes, with neither a match nor a fault.
  6. In /root/ica-resilience/dr-inventory.yaml, write a DestinationRule. spec.host is inventory.ica-resilience.svc.cluster.local, and under trafficPolicy put outlierDetection (consecutive5xxErrors 5, interval 10s, baseEjectionTime 30s, maxEjectionPercent 50) and connectionPool (tcp.maxConnections 100, http.http1MaxPendingRequests 50).
  7. In /root/ica-resilience/vs-inventory-mirror.yaml, write a VirtualService. Route 100% to inventory.ica-resilience.svc.cluster.local while mirroring 20% to inventory-canary.ica-resilience.svc.cluster.local, and set the overall timeout to 5s.

Notes

Preparing the lab namespace

Create the namespace ica-resilience and attach the label istio-injection=enabled.

It is the same way as the previous lab. Do not forget the automatic injection label.

Deploying a Deployment with 3 instances

In the namespace ica-resilience, deploy a Deployment inventory with 3 replicas. The Pod label is app=inventory and the image is nginx:1.27-alpine.

To get a feel for maxEjectionPercent of outlierDetection, you need multiple instances. Wait until the Pods are Ready before grading.

Creating the main Service and the canary Service

In the same namespace, create two Services. They are inventory and inventory-canary, both with port 8080 and port name http. The selector of inventory is app=inventory.

It is safer to keep the mirroring target as a separate Service. Name the ports of both Services so that the protocol is evident.

Matching the relationship between timeout and retries

In /root/ica-resilience/vs-inventory-retry.yaml, write a VirtualService. spec.hosts[0] is inventory.ica-resilience.svc.cluster.local; in the http rule, include timeout: 3s, retries.attempts: 2, retries.perTryTimeout: 1s, and in retries.retryOn, include 5xx and connect-failure.

timeout is the overall deadline including retries. If perTryTimeout x (attempts + 1) exceeds timeout, retries cannot be executed. Include in retryOn the cases where the request failed before it even arrived.

Fault injection gated by a header

In /root/ica-resilience/vs-inventory-fault.yaml, write a VirtualService. There are two rules. The first rule applies only to requests whose header x-chaos-test is true and has fault.delay.fixedDelay: 3s / fault.delay.percentage.value: 50 / fault.abort.httpStatus: 503 / fault.abort.percentage.value: 10. The second rule just routes, with neither a match nor a fault.

Applying a fault to all traffic on a production cluster is not chaos testing but an outage. Split the matching rule from the rule for ordinary traffic, and put the fault only on the matched side.

Setting circuit breaker limits

In /root/ica-resilience/dr-inventory.yaml, write a DestinationRule. spec.host is inventory.ica-resilience.svc.cluster.local, and under trafficPolicy put outlierDetection (consecutive5xxErrors 5, interval 10s, baseEjectionTime 30s, maxEjectionPercent 50) and connectionPool (tcp.maxConnections 100, http.http1MaxPendingRequests 50).

connectionPool and outlierDetection go side by side under the trafficPolicy of the same DestinationRule. Think about what would happen if you set the simultaneous ejection percentage to 100 when there are 3 instances.

Sending shadow traffic

In /root/ica-resilience/vs-inventory-mirror.yaml, write a VirtualService. Route 100% to inventory.ica-resilience.svc.cluster.local while mirroring 20% to inventory-canary.ica-resilience.svc.cluster.local, and set the overall timeout to 5s.

Mirroring discards the response, so put only one real route. The mirror target is the canary Service you created in step 3, and the percentage is specified in a separate field next to mirror.