Building an EAI Middleware Layer
Rejecting Fast Is the Kind Thing to Do
In one line
When one slow target drags down the entire hub, it is called a cascading failure. There are two ways to stop it, and you layer them — for a dead target, do not call it and reject immediately (a circuit breaker), and for a target that is alive but slow, limit, per target, the amount you can send concurrently (a bulkhead). Both are the same principle, "reject quickly rather than make them wait" — that is, backpressure.
Why it was needed
Say core banking has stopped. The hub connects to core banking for every request and waits for the response. If module 4's timeout is 3 seconds, while 100 requests per second come in, 300 waiting threads pile up in the hub. Each thread holds memory and a socket, so the hub chokes first, and even insurance claim transactions that have nothing to do with core banking line up in front of the hub. The failure of one core banking system goes through the hub and becomes a failure of all transactions.
A timeout alone is not enough. A timeout answers only "how long do we wait for one request," and does not answer "do we keep calling a target we plainly know is dead." And continuing to pour requests onto a dying target also robs it of the chance to recover.
How it works
Circuit breaker. As summarized in Martin Fowler's article, an object that wraps the call to be protected counts failures and opens the circuit when it reaches a threshold. While open, it does not call and returns an error immediately. Fowler writes that this pattern became widely known through Michael Nygard's book "Release It!". There are three states.
| State | Calls | Transition |
|---|---|---|
| CLOSED | Yes | When consecutive failures reach the threshold → OPEN |
| OPEN | No (immediate E904) | When the retry time passes → HALF_OPEN |
| HALF_OPEN | Only one trial call | Success → CLOSED, failure → OPEN (time restarts from the beginning) |
Half-open (HALF_OPEN) is the key. If you leave the circuit open forever, you don't notice when the target recovers, and if you close it entirely just because time has passed, the backlog of requests pours onto the just-revived target and kills it again. You test with one call and close if that one succeeds. Even if several threads ask at the same time, the trial call must be only one, so the state change is done inside a lock.
What to count as a failure. Insufficient funds (B201) is core banking answering healthily. If you count this as a failure, then just because insufficient-funds cases pile up at month end, you cut off the way to a perfectly fine core banking yourself. A failure is when the target said it could not process (E500), when there was no answer (E901), or when it could not even connect (E902) — count only system errors.
Bulkhead. Like the bulkheads of a ship, it is a design in which even if one compartment fills with water, the other compartments are left dry. You set a concurrency limit (a semaphore) separately for each target. If core banking's share is 3, only 3 requests to core banking go in at once, and from the fourth on, it does not wait and rejects with overload (E905). Claims have a separate share, so even if core banking is slow, claim transactions pass quickly on their own share. If you make requests over the limit line up and wait, threads eventually pile up and you return to the original problem — if you reject quickly, the caller (the channel) can choose among retrying, an alternate path and a message to the customer.
Circuit and retry. Retries are used to get past transient failures, and the circuit is used to stop calling on persistent failures. If the retrying side retries even the circuit's rejection (E904), there is no point in the circuit opening. So you give the rejection code separately, and callers do not retry E904 and E905 immediately.
How to decide the numbers. If the consecutive-failure threshold is too low, the circuit opens at momentary wobbles, and if too high, you keep calling a dead target for a long time. If the retry time is shorter than the time a target usually takes to recover, the half-open trial just keeps failing, and if too long, it keeps rejecting for a long time even after recovery. Set the concurrency limit so that it does not exceed the concurrency the target can handle (such as its connection pool size). There is no right answer; you decide by looking at the target's normal metrics and then adjust by looking at the transition records.
Record the state transitions. The fact that a circuit opened is an outage signal. If you leave the transitions (CLOSED→OPEN, OPEN→HALF_OPEN, HALF_OPEN→CLOSED) together with time and target, "core banking opened at 14:02 and closed at 14:05" appears in one line. In production you connect these transitions to alerts.
What it looks like in the field
The most common mistake is having only one circuit. If you put one breaker over the whole hub, a core banking outage makes even claim and card transactions get E904 — that is not isolation but a total block. The second is applying the concurrency limit with a single thread pool only. If the whole pool is locked up by one target, the result is the same as having no bulkhead. The third is a circuit that counts business errors as failures. The fourth is one built as "close automatically after 30 seconds" without half-open — the moment it closes, the backlog of requests pours out all at once. Finally, without a record that the circuit opened, operators learn only when they receive an inquiry: "why is E904 coming out?"
What we do in the next lab
You build a state machine in breaker.py, and the grader turns time with a fake clock to confirm the transitions. Then you attach it to the skeleton that relays to two targets (core banking and claims) (relay_base.py) — immediate E904 for a dead target, E905 when the concurrency limit is exceeded, isolating by splitting per target, recording state transitions, and not counting business errors as failures. The grader checks with the core banking fixture's call count and its maximum concurrent processing.