TT Lab
Get started
Learn Learning paths Courses

Building an EAI Middleware Layer

A Timeout Means Unknown, Not Failed

Continue in TT Lab

In one line

Synchronous relay means receiving a message, calling the target system and converting its result into a standard response code. The most important distinction is "it failed" versus "I don't know." If you could not even connect, it is certain that nothing was sent (E902), but if you sent and got no answer, nobody knows whether the target processed it (E901). The moment you lump these two into the same code, double transfers begin.

Why it was needed

If a channel (internet banking) calls core banking directly, the channel has to know all of core banking's address, HTTP conventions and error format. Whether core banking gives an error as 422 {"result":"INSUFFICIENT_FUNDS"} or as 200 {"code":"E-17"}, each channel ends up with its own interpretation code, and when the target system changes, ten channels have to be fixed. The relay layer takes on this interpretation in one place and returns a single standard response code to the channel — insufficient funds is B201 no matter which target it came from.

But with a relay in between, there is one more point of failure. It can stop anywhere on the channel → hub → core banking stretch, and depending on where it stopped, the meaning is completely different. The core banking fixture in this module is built exactly like reality. Even if slow, it carries the processing through to the end. Even if the hub gives up after 2 seconds, core banking deducts the balance at the 4-second mark. If the hub then answers the channel "failed," the channel sends again, and the money is deducted twice.

How it works

The path of one request. ① Read exactly 4 bytes of length over TCP and then read that much again (module 1) ② Parse the header, and if the format is wrong, E102 ③ If the transaction code is not a relay target, E101 (routing, module 2) ④ Convert the body to JSON according to the layout (the rules of module 3, now the common library lhconv) ⑤ Call core banking's POST /v1/transfers — carry the GUID in the body and in the X-GUID header ⑥ Convert the result into a standard code and build the response message (a 45-byte response body only on success).

Divide results into four branches. In HTTP semantics (RFC 9110), 2xx means the request was processed successfully, 4xx means a problem on the request side, and 5xx means the server could not process it. Core banking gives a business rejection as 422 (RFC 9110 15.5.21, content that the server understood the format of but cannot process) and distinguishes the reason by result in the body. So the conversion table looks at the status code and the business code together.

Situation What you can know Standard code
200 Processed 0000
422 + INSUFFICIENT_FUNDS etc. Rejected (money did not move) B201–B203
500 The target said it could not process E500
Timeout while waiting for a response Sent. Whether it was processed is unknown E901
Connection refused, connect timeout Could not send (confirmed not sent) E902

There are two kinds of timeouts. A timeout while establishing the connection and a timeout while waiting for a response after sending the request have opposite meanings. With the former, not a single byte of the request went out, so sending again is safe. With the latter, the request has already arrived at the target. Python 3.12's urllib.request.urlopen raises these as different exceptions (measured in the course of this work) — a problem at the connection stage is wrapped in URLError (with reason being ConnectionRefusedError or TimeoutError), and a timeout that occurs while waiting for a response comes up as an unwrapped TimeoutError. If you catch one exception wholesale and convert it to the same code, this distinction disappears.

The timeout budget must get shorter going inward. If the channel waits 10 seconds and the hub waits 15 seconds for core banking, the channel has already given up and resent while the hub is still holding the first request. Nobody receives the hub's result. So the budget shrinks in the order channel > hub > target, and the hub always answers something within its own budget — if it does not know, it says it does not know (E901).

Handle each connection separately. A server that does not accept the next connection while processing one request lines up all the transfers behind a single slow transfer. The Python standard library's socketserver.ThreadingTCPServer starts one thread per connection. Since threads can grow without limit, module 11 adds a per-target concurrency limit (bulkhead).

What it looks like in the field

The most expensive incident is a channel that treated E901 as a failure. With just one line of screen text, "Response timed out — please try again," the customer presses the button once more, and two transfers are left in core banking. That is why no resend and result lookup are attached as a pair to a response whose result is undetermined (module 8). The second is a hub that passes the target system's error messages straight through to the channel. Core banking's internal error text (stack traces, table names) appears on the customer screen, and the channel piles up code to interpret the different wording of each target. The third is a relay that always worked locally but loses half of each message in production — code that ended with a single recv instead of reading the whole length.

What we do in the next lab

You start the core banking fixture and call it directly, then read the specification and write the response code conversion table. Then you grow the relay server relay.py step by step — a skeleton that reads fragmented messages to the end, relaying a normal transfer, converting business errors, a read timeout (E901), a connection failure (E902), and concurrent handling. The grader starts your relay.py itself and even checks how many times it called using the core banking fixture's call statistics.