TT Lab
Get started
Learn Learning paths Courses

Building an EAI Middleware Layer

Convert Only at the Boundary, Reject What You Cannot Convert

Continue in TT Lab

In one line

A transformation adapter changes the format in just one place, at the boundary between the standard message (fixed-length, EUC-KR, external codes) and internal JSON (UTF-8, internal codes). The rules are kept as data — a layout and a code mapping table — and unknown codes, overflowing lengths and characters that cannot be represented are rejected, not fixed up and passed along.

Why it was needed

Two worlds live together inside a bank. Core banking and external institutions use decades-old fixed-length messages, while new channels and internal APIs use JSON. If you let the two worlds connect directly, every team building a new channel has to learn EUC-KR and byte padding all over again, and each carries its own copy of the correspondence table between external and internal codes. When there are several copies of a table, one copy will surely go stale someday.

"Enterprise Integration Patterns" puts a Message Translator in this place. It inserts a filter that changes the format between applications that use different data formats, and it explains this as moving the Adapter pattern of object orientation into messaging. If you gather transformation in the hub, the inner systems know only JSON and internal codes, and even when the outside standard changes, the only place to fix is one adapter.

How it works

The layout is data. You keep a table with one line per field giving the name, length, format, JSON key, mapping domain and number of decimal places (scale), and the converter only interprets that table. When the specification changes, you fix the table. If you write separate conversion code per transaction, with a hundred transactions you end up with a hundred sets of padding rules.

Length is in bytes. There are three formats. N (numeric) is right-aligned and padded with zeros at the front, while AN (alphanumeric) and H (mixed Hangul) are left-aligned and padded with spaces at the back. The length of an H field is in EUC-KR bytes, so 10 Hangul characters completely fill a 20-byte field. If you count by characters, a name of 11 Hangul characters passes a "20 characters or fewer" check and becomes 22 bytes, and all the following fields shift by 2 bytes.

If it overflows, don't cut. If you cut by bytes to fit the length, a Hangul character is split in half. The receiving side cannot decode that field, or joins the remaining first byte with the first byte of the next field and reads a completely different character. Even cutting by characters leaves the problem — a transfer whose payee name was truncated is a business incident. The converter rejects, and whether to shorten is decided by the channel and the business.

Convert codes by table, and reject unknown codes. The bank code 201 in a message is GRM internally. The correspondence is kept in one mapping table and used in both directions. If you let a code not in the table flow through as it is, the internal system interprets that value in its own code system — if the same value happens to exist with a different meaning, money goes to the wrong place. Codes discontinued because of a merger are also removed from the table so that they are caught the moment they come in.

Round trip is the converter's specification. If you convert a body made to the standard into JSON and then back into a body, it must not differ from the original bytes by even a single byte. Several rules come out of this property by themselves. For AN and H you must strip only trailing spaces — if you strip leading spaces too, they are gone when you convert back. N can be converted to an integer because the leading zeros are only padding, not part of the value. If you run a round trip over the whole sample set every time you fix the converter, any change that loses something shows up on the spot.

Amounts and interest rates do not go through floating point. The interest rate 0032500 in a message has an implicit decimal point at the fourth place, so it is 3.2500. The Python decimal documentation explains that in binary floating point 1.1 + 2.2 shows as 3.3000000000000003 and 0.1 + 0.1 + 0.1 - 0.3 is not 0, so you cannot trust equality checks, and therefore says decimal suits accounting that needs strict equality. The same document also shows that Decimal(0.1) and Decimal('0.1') differ — once a value has become a float, converting it to decimal is already too late. So even in JSON, values are exchanged not as numbers but as the string "3.2500". The two trailing zeros are information conveying the number of digits, and decimal keeps this notation by leaving 1.30 + 1.20 as 2.50. In fact float("0.29") * 100 is 28.999999999999996, so converting it to an integer comes up short by 1.

Don't trust the name "EUC-KR." The Hangul that EUC-KR holds is the 2,350 KS X 1001 precomposed characters. CP949 (UHC), which is common on Windows, adds the remaining Hangul to these, and counting with Python, it holds all 11,172 Unicode Hangul syllables in 2 bytes. The 2,350 characters have the same bytes in both encodings, so normally there is no problem, and it shows up only with characters such as the syllable U+B620. The harder part is that it differs by implementation. The euc_kr codec in the Python standard encodings table does not reject the syllable U+B620 as an error. As the comment in the CPython source says, it converts it into a KS X 1001:1998 combining sequence and outputs 8 bytes. glibc's iconv -t EUC-KR rejects the same character. A peer that knows only 2-byte precomposed codes reads those 8 bytes as four jamo letters. So the converter itself decides, for each character, "is it a 2-byte precomposed character?", and rejects it if not.

What it looks like in the field

Half a Hangul character. If there is a channel that cut a summary by bytes to fit 30 bytes, the external institution sends the whole message back as a format error. The error appears on the external side, but the cause is our channel's cutting. A body mixed with CP949. If a channel saves the file as CP949 and sends it, everything is fine for months, and it breaks on the day the first customer arrives whose name has a character outside the precomposed set. If the converter reads strictly as EUC-KR and rejects it, the cause is visible that very day. Numbers padded with spaces. Some systems send amount fields padded with spaces instead of 0. It can be read, but the bytes differ on a round trip. A round-trip test finds not only the converter's bugs but also such peers that violate the standard.

What we do in the next lab

You transcribe the specification into a layout file and the code definition into a mapping table, and then build two converters (f2j.py and j2f.py) yourself. You run the round-trip test over all the samples, handle the implicit decimal point with decimal, fix it to reject characters outside the precomposed set, and then convert the whole inbox. The common library lhconv.py does the same work but is not used in this lab — it is used from module 4. The grader runs your converters with a new layout and new values (including Hangul) each time.