ICA — Istio Certified Associate
Response failure and business failure are different events
In one line
A retry is an attempt to get a response again, not a command that cancels the earlier business operation. On a real Istio mesh, this module counts the number of HTTP receipts and the SQLite business records separately, and builds the boundary that does not execute the same business operation twice even if a response is lost or a process dies.
Why this was needed
Suppose a user pressed the pay button once, yet the screen showed a timeout. A customer support agent says it failed, so try again. A server operator explains that the proxy automatically retries two more times. But the business owner says there are three payment records for the same order. All three observations can be true.
One request on the screen, several send attempts by the proxy, and a business write in the database are different events. If you count only HTTP 200s, duplicate business operations are invisible, and if you count only 504s, you mistake a business operation that was already committed for one that was canceled. So you must leave the business key, the transport request ID, the response code, and the actual stored business ID together. Before writing "it ran three times" in an incident report, you must state what was counted.
LabHub's earlier synthetic experiment also observed this difference. Among concurrent requests with the same business key, some clients received a 504. One of those requests had 200 as the last state in the server's receipt record. Only one business record was added. You must not memorize this ratio as a law of every environment. Instead of a coincidental ratio depending on load, this lab uses a comparison path that delays the response by 2 seconds after the business commit.
How it works
Three ledgers
When you run curl once from the client Pod, it attaches one request ID. If Istio resends the same request, the same request ID appears several times in the upstream's receipts ledger. The business key is a separate header. Even if you start the transmission anew, if the request is meant to retrieve the same order, it keeps the same business key. Conversely, even with identical content, a separate order must have a different key.
The server has a deliberately wrong /naive path and a /flaky path that uses duplicate-prevention records. On both, the first two responses are real upstream 503s and the next response is 200. They are not fake errors produced only by the proxy. /naive adds a business record on every receipt, and /flaky reuses the existing response for the business key. Even with the same HTTP response sequence, the business effect can differ.
The count setting is a cap, not a guarantee
The VirtualService's attempts is the cap on retries added to the initial send. If attempts is 2, it is at most three including the initial attempt. The overall timeout is the budget these attempts share. perTryTimeout is the cap on how long each attempt can wait, and a request that failed fast is not counted as having used all of that time.
retryOn chooses which kinds of failure to retry. Specifying HTTP 503 does not mean every connection error and read timeout is bundled under the same condition. This lab configures things so that a previous host can be chosen again, in order to compare a single upstream. Do not regard this as having solved retry storms, backoff, circuit breaking, and capacity planning in a real service.
The business operation that remains even after a 504
/lost first commits the business operation and the duplicate-prevention record and delays only the response by 2 seconds. You set the proxy budget for this path to 500ms and turn off automatic retries. You look up the client's 504 and the server's business record with the same key. Then you send the same key to the fast /charge path to retrieve the existing business response. You are not trying to create a new order, so you must not issue a new key here.
This is not an external payment provider's API. It is a server that writes synthetic business operations to a local SQLite. The 200 in the response record is the response the server meant to produce, and it is not proof that the client received those bytes. That is why you distinguish the receipt record from the client's raw text.
Atomicity within the same transaction
If you commit the business operation first and write the duplicate-prevention key later, a gap opens between the two commits. If the process dies in that gap, the business operation remains but the key does not. A retry with the same key is mistaken for a new business operation. You must bind the business write and the storing of the key, the content fingerprint, and the response in one transaction.
SQLite allows one write transaction at a time. Even if you start a write transaction with BEGIN IMMEDIATE, you cannot necessarily obtain the lock forever, and when there is contention you must handle BUSY or a timeout. The lab uses only a single local DB file and short synthetic jobs. It does not implement a distributed transaction that atomically wraps several DBs or an external payment call.
What it looks like in the field
Even if a service restarts, an in-memory dictionary does not come back. In the lab, you actually delete and recreate the Pod and check that the UID and the process instance changed. You read the existing response from the SQLite file left on the same VM disk. You must not extend the fact that it survived a Pod replacement to VM deletion, disk corruption, or power-failure recovery.
A content fingerprint is also needed. If the same key comes in with a changed amount or customer, instead of unconditionally returning the existing response, it must be rejected as a conflict. On the other hand, if only the order of JSON properties changed, it must be treated as the same content. Merging different orders with the same amount looking only at the payload fingerprint is also an error.
At the end, you fix the split-commit defect in the provided transaction.py. The checker runs the student implementation on a new temporary DB and sends SIGKILL to a real child process. At two points before the commit, the business operation and the key must both be absent, and after the commit, both must remain. Then it checks whether there is one business operation when retried over a new connection. Throwing an exception alone does not substitute for a forced kill.
What you will do in the next lab
You observe the Pod inventory, the retry route, the comparison of duplicate business operations, the timeout, the content conflict, and the Pod replacement. Next you fix the transaction code and also check concurrent requests and different business keys. The observation tool leaves the actual raw text in files and does not run again if the same observation file exists. Grading cross-checks those files against the receipt records left on the server.
The business DB, test files, and namespaces exist only inside this learning VM. Do not put in private keys or real payment data. Everything is reclaimed when the lab ends. The counterpart API's idempotency key, outbox, reconciliation, and compensation processing that you need when combining with an external system are later design tasks, and this lab does not claim to have implemented them.
Check further in the official documentation
- Istio HTTPRetry: the cap on additional attempts, the per-try time, and the retry conditions.
- Envoy Router: the relationship between retry conditions and the overall request time.
- SQLite Transaction: concurrent writes, transaction start and end, and BUSY.
- SQLite Atomic Commit: the atomicity of the rollback journal and the assumptions about the storage device.