TT Lab
Get started
Learn Learning paths Courses

Building an EAI Middleware Layer

The Hub Decides Where to Send Messages from a Table

Continue in TT Lab

In one line

The hub decides the destination by looking only at the header's transaction code and receiving institution. That decision is kept not in code but as data (a routing table), who wins when rules overlap is nailed down in a document, and a message that cannot be sent anywhere is not swallowed but immediately returned as an error response.

Why it was needed

As we saw in module 1, with point-to-point connections linking systems directly, the lines grow quadratically as the number of systems increases. A hub-and-spoke structure gathers all the lines into one hub, and in exchange gives the hub one new job — who to give this message to. Hohpe and Woolf's "Enterprise Integration Patterns" calls this role the Content-Based Router. It is a component that chooses the channel to send on by looking at the data in the message (whether a field exists, what a particular field's value is). Our hub looks at just two fields of the header.

The same text warns of something. A router easily becomes a frequent maintenance point that has to be modified every time a destination changes, and if the rules are complex, it says you can choose a form in which the destination is computed from a configurable rule set. At a bank hub this warning is reality as it stands. Transaction codes are newly created every month, and when a system is migrated, the destinations of dozens of codes change at once. If you write routing as code like if tx_code == "BKTR0001", you have to redeploy the hub every time you add a transaction, and a hub deployment briefly shakes all the transactions of that bank. So the routing rules are taken out into a table (data), and the hub code does only the job of "read the table and choose by priority." The table changes often and the way the rules are interpreted changes rarely — you separate what changes often from what changes rarely.

How it works

The source is the business side's interface list. Which transaction goes to which system is often managed in Excel by the business owner, not the developer. This list mixes retired interfaces, interfaces under development and old versions of the same transaction. The routing table the hub uses is the refined result of this. It loads only rows whose status is in operation, and if a transaction code appears several times, it keeps only the one with the largest version. A common trap here is version comparison. The version in a CSV exported from Excel is text, so if you compare it as a string, "9" is greater than "10". You have to convert it to an integer and compare.

The table carries not only "where to" but also "how." One row of the routing table has, along with the destination, the sync/async distinction and the timeout. There is a reason timeouts differ by transaction. A balance inquiry makes sense only if it answers within a few seconds, but submitting an insurance claim document taking tens of seconds is normal. If you lump all transactions under one value, setting it short makes claims all look like failures, and setting it long makes the hub's connections pile up waiting for a stalled core system. A timeout is a property of the transaction, so it goes in the same row as the transaction code.

There are four rules, and they have a ranking. This course's fictional standard (ROUTING.md) set it like this.

Rank Rule Condition Why this position
1 ORG The receiving institution is not our own bank → FEP A message going to another institution must go out only through one exit where external lines and per-institution formats are gathered
2 EXACT Exact match on the transaction code Write exceptions as the most specific rule
3 PREFIX The longest prefix match Group a whole business family with CD*, but a sub-group such as CDLN* wins
4 NONE Nothing matches → E101 Do not guess and send

The idea that "the specific one wins" is the same as IP routing. RFC 1812 states that when a router chooses a route it picks a specific host route first, then network prefixes in order of longest prefix. If both CD* and CDLN* match, the longer one is the rule that knows more.

The priority must not be the row order of the table. If you implement it as "the first matching row from the top," the row order becomes a hidden rule. Even if someone just sorts the Excel file by transaction code and exports it again, CD* moves above CDLN*, and card loan messages quietly go to the card system. No error occurs at all. So the priority is computed explicitly by the code — find the exact match first, and sort prefixes by length.

Unroutable is returned as a response. If a hub that received a transaction code not in the table only logs it and drops the message, the sending channel learns of the failure only when it waits for a response and times out. In the meantime the customer screen keeps spinning, and if the channel even resends, the same wasted trip happens twice. The hub immediately builds a response message carrying response code E101 by turning the received request around (exactly the response building of module 1 — keep the transaction code and GUID, swap the institutions). On the other hand, this response cannot be built for a message with a wrong format (E102). Since the header itself cannot be trusted, neither can the GUID.

Leave an audit log for every decision. To answer the question "why did this transaction go to the information system," you need not only the destination but which rule matched (EXACT, PREFIX, ORG, NONE) and the GUID. The log is append-only. An audit log that is overwritten is not an audit log.

What it looks like in the field

The most common incident is a broad prefix swallowing a new transaction. With a CD* rule that goes to the card system in place, a new transaction CDLN0010 that should go to the core banking system appears, and if you forget to put an exact-match row in the table, that transaction goes to the card system without an error. The card system rejects it as an unknown transaction, and the outage meeting suspects the card system first. If PREFIX is recorded in the audit log, the cause shows in one line. The very fact that a transaction expected to match exactly went out by prefix is a signal.

The second is retired interfaces remaining in the table. If you migrate a system and do not delete the old rows, it is quiet for a while because nobody sends to that code. One day someone reuses that code and traffic goes to the retired system. This is why you keep a status column and mechanically keep to "load only in operation." The third is, like this module's version comparison trap, a trivial mistake in the script that refines the list. After building the routing table, you should run a test that checks it against the original list once.

What we do in the next lab

You refine the business side's messy interface list into the routing table routes.csv and grow the router router.py one step at a time — exact match, prefix (the longer wins), institution-based rule, routing by the message header, the E101 response, the audit log. The grader builds a random table with new transaction codes and destination names each time and runs your router, so if you hard-code the rules, you cannot pass. At the end you dispatch 30 messages from the inbox.