Building an EAI Middleware Layer
The Middleware Reads Only the Header
In one line
An intermediate layer such as an EAI decides the route by looking only at the common header of a message. So the header is fixed to a few fields unrelated to the business — length, transaction code, global ID, sending and receiving institutions, request/response indicator and response code — and if even one system cuts these fields one byte differently, every stretch after it is wrong.
Why it was needed
A single bank has dozens each of channels (internet banking, mobile, branch counters), core banking (the ledger), information systems, cards and external systems. If you connect every system directly to every other, the lines grow to n×(n-1)/2, and each time one system's format changes, everything attached to it has to be fixed. An intermediate layer (EAI; on the channel side often called MCI, and on the external side FEP) gathers these lines into a single hub. Each system makes an agreement only with the hub, and the hub decides on their behalf where to send a message it received, what format to convert it to, and what to return if it fails.
But for the hub to decide, there has to be a place it can read without knowing the business content. A funds transfer message and an insurance claim message have completely different shapes in the business part (the body). If the hub had to know the body layout of each transaction to choose a route, it would have to be modified every time a new transaction appears. So a common header of the same shape is attached in front of every message, and the hub reads only the header. The body is interpreted by the destination. This separation is the starting point of this whole course.
How it works
This course uses the fictional LabHub bank standard (LH-STD). The header is 80 bytes and entirely ASCII. The specifications of real institutions are often not public and their field order and lengths vary, but the roles of the fields they contain are almost the same.
| Field | Length | Role |
|---|---|---|
| Message length (MSG_LEN) | 4 | The number of bytes excluding this field. The only basis for finding message boundaries in TCP |
| Transaction code (TX_CODE) | 8 | What transaction this is. The key to routing (module 2) |
| Global ID (GUID) | 32 | A number that follows this one transaction to the end. The key to tracing (module 7) and duplicate prevention (module 8) |
| Sending and receiving institutions | 3+3 | Who to whom. If it is an external institution, it is sent to the external system |
| Request/response indicator | 1 | Q request / R response |
| Response code | 4 | Blank for a request; 0000 or an error code for a response |
| Transmission time | 14 | The time this stretch sent it |
Length is in bytes. The body is EUC-KR. EUC-KR is an encoding that holds the KS X 1001 character set, so one Hangul character is 2 bytes, and the same character is 3 bytes in UTF-8. If someone opens a message file in an editor and saves it as UTF-8, the characters look the same but the byte count grows, while the message length field keeps its old value. The receiving side reads as many bytes as the length says and mistakes the rest for the beginning of the next message. Also, KS X 1001 contains only 2,350 precomposed Hangul characters, so a customer whose name contains a character such as the syllable U+B620 has no 2-byte precomposed code. What happens then differs by implementation — the glibc iconv in the lab image refuses the conversion, while Python's euc_kr codec converts it without error into an 8-byte combining sequence (measured in the course of this work). If it quietly becomes 8 bytes, the fixed-length field calculations are thrown off entirely. How to treat such customers is a business rule that must be decided at the moment you choose the encoding (we handle it directly in module 3).
TCP does not know messages. TCP is a stream that delivers bytes in order, without gaps (RFC 9293). Just because the sender called send twice does not mean the receiver's recv arrives in two parts. A single recv may bring half a message, or two and a half. So the receiving side always first reads exactly 4 bytes of length, and then reads exactly that length again. This is why the message length field is at the very start of the header. Code that breaks this rule works almost always in the development environment (same machine, short messages) and breaks only with long messages and a slow network in production.
A GUID is made once and carried to the end. The system that first creates the transaction (the channel) issues it, and the hub, core banking and external systems pass on the received value as is. If each stretch creates a new one, one transaction scatters into four numbers and there is no way to connect them during an outage. LH-STD fixed the GUID to the same shape as the trace-id of W3C Trace Context — 32 lowercase hexadecimal characters, all zeros invalid. Then in stretches that pass over HTTP, it can be carried as is in the traceparent header. Issue it with unpredictable random numbers (secrets, uuid4). If you make it from the time and a sequence number, two servers produce the same value in the same millisecond.
A response is made by turning the request around. The transaction code and GUID stay as they are, the sending and receiving institutions are swapped, the indicator is R, the response code is filled in, the transmission time is now, and the length is recalculated. Code that puts the length in by hand will surely get it wrong someday — calculate it when building.
What it looks like in the field
The most common incident is a misunderstanding of the length criterion. Some institutions include the length field itself in the length, and others exclude it. Because of this difference, written on a single line in the specification, the integration with one external institution stops entirely on opening day (in module 6 we handle the differences by institution directly). The second is a parser that quietly passes over format errors. If a message whose request/response indicator is X is guessed to be a "request" and processed, the ledger changes without anyone knowing where and why that message broke. A message with a wrong format should be rejected, and the reason for the rejection left as a code (E102). The third is a system that creates a new GUID in each stretch. When someone says in an outage meeting, "I can't see that transaction number on our side," this is usually it.
What we do in the next lab
You read the specification (SPEC.md), write the header layout down to the offsets, and then compare eight messages in the inbox byte by byte. Then you build in turn a parser (hdr.py), a response generator (reply.py), a GUID issuer (guid.py) and a stream splitter (split.py), and produce a verdict report for the whole inbox. From module 2, you use the in-house common library (lhstd.py) that does this work — there should be only one copy of the header.