Building an EAI Middleware Layer
Retransmission Is the Default
In one line
When a timeout occurs, the channel resends with the same GUID. If core banking is not idempotent, preventing duplicates is the relay layer's job, and its tool is a state ledger with the GUID as the primary key. The ledger must record not only "processed" but also distinguish "sent but unknown" from "could not send," and reprocessing is safe only on top of that distinction.
Why it was needed
The relay of module 4 calls as it receives. There is no problem if the network is fine, but if one response arrives late between the channel and the hub, the channel takes it as a failure and resends — and the hub passes the second request to core banking too. The core banking fixture in this course was deliberately made non-idempotent. If it receives the same transfer twice, it deducts twice. Real ledger systems often cannot judge "is this the same request" by themselves, and even when they can, someone has to carry the basis for that (the unique transaction number) all the way.
Think of messaging delivery guarantees, and duplication is not an exception but the default. To deliver a request reliably, you must resend when there is no response (at least once), and the moment you resend, the receiving side can receive the same thing twice. Enterprise Integration Patterns sums up the way to solve this problem on the receiving side as the Idempotent Receiver — making the result the same even if the same message is received several times, as if it had been received once. In this module, the relay layer takes on that role in front of core banking.
How it works
One ledger row = one GUID. If you make guid the primary key, the very state of "two rows with the same GUID" becomes impossible. The ledger holds state, response code, the original request, the original response, the number of attempts and the update time. The reason to store the original request is reprocessing, and the reason to store the original response is to return the same answer as the first one to a resend.
There are five states.
| State | Meaning | If the same GUID comes again |
|---|---|---|
| SENT | Being sent to core banking (claimed) | E903 — being processed; answer immediately without waiting |
| DONE | There is a confirmed answer (0000, B2xx, E500) | The stored response as is |
| UNKNOWN | Sent but no answer received (E901) | E901 — do not call core banking again |
| FAILED | Certain that it could not be sent (E902) | It may be sent again |
| RECEIVED | Received but not yet sent | Treated the same as SENT, depending on the implementation |
Write before sending. The key is the order. If you write to the ledger after calling core banking, then when the hub dies in between, the ledger has no trace, and the channel's resend is processed as a new transaction. So first claim a row as SENT, then call, then record the result. The claim must be atomic. When two threads receive the same GUID at the same time, if both judge "not there, I'll process it," the ledger is useless even if it exists. SQLite's INSERT ... ON CONFLICT DO NOTHING (UPSERT documentation) inserts nothing on a primary key conflict and tells you, by the number of affected rows, "did I claim it." You judge with a single insert, not with the two steps of look up and then insert.
UNKNOWN is resolved only by lookup. Resending a transaction whose result is unknown is a gamble. If core banking processed it, it becomes a double transfer. Instead, ask core banking's lookup API whether that GUID was processed. If it was processed, change to DONE with that result; if there is no record, change to FAILED (confirmed unprocessed). There is one trap here. A transaction that just timed out may still be being processed by core banking. If you look it up now, get "none," change it to FAILED and reprocess, core banking finishes the first request a second later and deducts twice. So lookup is done only for UNKNOWN rows older than the target's maximum processing time (--min-age).
The DLQ is not a place to discard but a place to wait. A transaction that could not be sent (E902) can simply be sent again after the cause is cleared. The Dead Letter Channel is a channel that collects messages that cannot be processed. In this lab, the FAILED/E902 rows are that place. There are three principles for reprocessing — ① send with the same GUID (if you make a new number, the ledger cannot see the duplicate) ② put a limit on the number of attempts (so that a transaction that fails forever does not hammer core banking every cycle) ③ UNKNOWN is never a target for reprocessing.
Clean up the ledger too. A ledger cannot grow forever. But the rule for deleting is itself the limit of duplicate prevention. Deleting DONE after 7 days means "a resend arriving after 8 days cannot be blocked," so the retention period must be longer than the longest period over which the channel can resend. And transactions that are not finalized (UNKNOWN, FAILED) are kept regardless of the period — the moment you delete them, the basis for investigating them is gone.
What it looks like in the field
The most common incident is a reprocessing batch that resends "everything that failed." The failure list has timeout cases mixed in, and some of them were already processed by core banking. The next morning, double-withdrawal inquiries pour into the call center. The second is making a new GUID when reprocessing — the way for the ledger to recognize it as the same transaction is gone. The third is keeping the ledger in memory (a dictionary). The moment you restart the hub, it forgets "what was sent." The fourth is the race between the ledger lookup and insert. It never shows up in a low-load development environment, and double processing occurs only in an outage situation where the channel resends at short intervals.
What we do in the next lab
You write the ledger schema, copy the fixture's relay (relay_base.py, with no duplicate prevention) and attach the ledger — a resend of a finished transaction, a resend while processing (E903), a timeout (UNKNOWN) and a connection failure (FAILED), in turn. Then you build resolve.py, which confirms UNKNOWN by lookup, reprocess.py, which sends again only confirmed-unsent transactions with the same GUID, and purge.py, which cleans up by retention period. The grader catches double calls with a temporary ledger and core banking's call statistics.