Building an EAI Middleware Layer
One Number Joins Four Teams' Logs
In one line
The first question in an outage meeting is always the same — "How far did that transaction get?" To answer this question within minutes, every stretch must leave the same number (GUID) in its log, and those logs must be joinable in one format and one time zone. If the number differs per stretch or the times are all over the place, the answer becomes hours of guesswork.
Why it was needed
One transfer passes through the channel (MCI) → hub (EAI) → core banking (CORE) → and, if needed, the external gateway (FEP). When a customer calls saying "the money was deducted but it showed the transfer failed," four teams dig through their own logs. MCI is key=value, the hub is pipe-delimited, core banking is JSON, and the external gateway is fixed columns. Core banking writes in UTC and the rest in Korean time. And in the worst case — when the hub calls core banking, it makes up a new number. If you search core banking's logs with the number from the channel's log, nothing comes out. "I can't see that transaction on our side" repeats four times, and in the end they look for similar lines by time and amount and join them by hand.
This is why module 1 decided that a GUID is "issued once by the system that first creates it, and every stretch carries it as is." This module covers what you can do with logs in which that promise is kept, and how to fix a relay that breaks the promise.
How it works
Normalization. Convert logs of different formats into the same columns (guid, hop, time, event, response code). Rather than asking four teams to change their original log formats, it is faster to align them once on the reading side. But a newly built system writes structured logs from the start (one line of JSON, fixed keys) — it can be read without regular expressions.
Time zones. A log time must always carry its time zone. The Z in 2026-09-23T00:02:08.268Z means UTC, and in Korean time (UTC+9) it is 09:02:08.268. A time without a time zone (2026-09-23 09:02:08) is a time that "only someone who knows which time zone can read." If you join them without realizing that only one system is in UTC, it looks as if core banking processed 9 hours before the hub.
Joining in time order. If you line up the rows of one GUID in time order, the path of the transaction emerges. If you attach the elapsed time from the first event, you can see in which stretch the time was spent. But the clocks of different servers drift slightly (several milliseconds even when each is set by NTP), so the millisecond-level order between stretches is taken only as a reference. The difference between two events inside the same server (core banking RECV → APPLY) is reliable.
Vanished transactions. If the hub wrote "sent (OUT)" but core banking has no "received (RECV)," that transaction vanished between the two. It could be a network disconnection, or it could have been dropped in front of core banking. The hub would have answered such a transaction with E901 (module 4), and it must be confirmed by the lookup of module 8. Only by joining on the GUID can this list be pulled out mechanically.
Decomposing slow transactions. Break transactions whose total time exceeded 3 seconds into stretches — the time spent inside core banking (RECV→APPLY) and the time spent at the external institution (REQ→RSP). You must decompose individual transactions, not averages, to separate "a day when core banking was slow" from "a day when a particular institution was slow."
Propagation rules. A relay carries the received GUID to the next stretch as is. In message stretches it is the GUID position in the header, and in HTTP stretches it is a header (this course uses X-GUID) and the body. And there is also a standard. W3C Trace Context defines the traceparent header for passing trace context over HTTP. Its shape is 버전-trace-id-parent-id-flags (version-trace-id-parent-id-flags), and in version 00, trace-id is 16 bytes (32 lowercase hexadecimal characters) and parent-id is 8 bytes (16 characters), and both are invalid if all zeros. The trace-id is one per transaction, and the parent-id is made new for each call — so multiple calls within one transaction can be drawn as parent and child. Because the LH-STD GUID was set to the same shape as a trace-id (module 1), if you carry the GUID as is in the trace-id, tracing is not broken between the message stretch and the HTTP stretch and stays connected.
Required log columns. A single stretch-event line must have at least the time (with time zone), GUID, stretch name, event and response code. Personal information such as amounts and account numbers is either not included or masked — one number is enough for tracing.
What it looks like in the field
The most common is the incident where a relay makes up a new GUID. Someone, to keep "our system's transaction number rules," discarded the received number and attached their own. The intention was good, but tracing breaks at that stretch. If you really need your own number, write it additionally and pass on the received GUID as is. The second is logs without a time zone. From the day one server's time zone setting changed, the logs drift by 9 hours and nobody notices. The third is error paths that do not leave the GUID in the log. The GUID is printed in the normal flow, but one log line in the exception handling block lacks it — and that line is the one you actually need.
What we do in the next lab
You normalize the logs of four stretches (MCI, EAI, CORE, FEP; four formats, two time zones) for one business day into one format and align them to Korean time. You build trace.py, which draws the path of one GUID, a list of transactions that did not reach core banking, and a stretch breakdown of slow transactions. Finally, you fix the relay that makes up a new GUID (relay_buggy.py) so that it carries the received GUID as is and leaves structured logs, and carry traceparent in the HTTP stretch.