Building an EAI Middleware Layer
Don't Let One Slow Target Drag Down the Hub
Goal
With a circuit breaker and per-target bulkheads, keep one slow or dead target from dragging down the entire hub.
Why it matters
A timeout only decides how long to wait for one request. If you keep calling a dead target, waiting threads pile up and the hub chokes first, and unrelated transactions stop along with it. You need a device that rejects quickly, separately for each target, for an outage not to spread.
Steps
- In
/root/eaimw/breaker/breaker.py, createCircuitBreaker(failure_threshold, reset_timeout, clock=time.monotonic). It hasstate(CLOSED,OPEN,HALF_OPEN),allow(),record_success()andrecord_failure(). When consecutive failures reach the threshold, it is OPEN, and when OPEN,allow()is False. You must always read the time withclock()(the grader injects a fake clock). - In OPEN, when
reset_timeouthas passed,allow()gives True only once and it becomes HALF_OPEN. If the trial call succeeds, CLOSED; if it fails, OPEN again (time restarts from the beginning). Even if several threads ask at the same time, allow only one (a lock). - Start with
cp /opt/lab/fixtures/eaimw/breaker/relay_base.py /root/eaimw/breaker/relay.py, add the--fail-threshold(default 5) and--reset-timeout(default 10) arguments, and wrap the target call with the breaker. If open, do not call and answer E904 immediately. At this step, count any result other than 0000 as a failure. - Put a concurrency limit with
--max-inflight(default 10). If the limit is full, do not wait and answer E905 immediately. - Put the breaker and the concurrency limit separately for each target (CORE, CLAIM). A breakdown or saturation of core banking must not spread to claim transactions.
- In
--breaker-log(default/root/eaimw/breaker/breaker.log), leave one line<시각>|<대상>|<FROM>-><TO>(time, target, from-state and to-state) each time the state changes. - Narrow what counts as a failure to system errors (E500, E901, E902). Count a business rejection (B2xx) as a success, since the target answered healthily.
Notes
- Concurrency limit: with
threading.BoundedSemaphore(n), ifacquire(blocking=False)is False, the limit is full. When finished, alwaysrelease()(try/finally). - Recording state transitions: if you compare the before and after of
allow()and ofrecord_*()in terms ofstate, you can tell the moment it changed. - Core banking fixture switches:
curl -XPOST localhost:9201/_ctl -d '{"mode":"fail"}'(500),"slow"(delay),"normal". In the/_statsstatistics,callsandmax_inflight. - Common mistakes: one breaker for the whole hub, making requests over the limit line up and wait, allowing several calls in HALF_OPEN, and calling
time.time()directly and ignoring the fake clock.
Open the circuit when consecutive failures pile up
The CircuitBreaker in /root/eaimw/breaker/breaker.py becomes OPEN at the consecutive failure threshold, and when OPEN, allow() is False.
On success, reset the failure count to 0 to count "consecutive." Record the opened time with self.clock(), not time.time() — the grader holds the clock.
Test with one call and close
After reset_timeout, in HALF_OPEN allow only one trial call, and go to CLOSED if it succeeds and back to OPEN if it fails.
Inside allow(), if it is OPEN and the time has passed, switch to HALF_OPEN and keep a mark of whether the trial call has already gone out. Decide inside the lock so that even if two threads ask at once, only one gets True.
Don't call when open
Copy relay_base.py, wrap the target call with the breaker, and when open answer E904 immediately (--fail-threshold, --reset-timeout).
handle_tx is the place to wrap. If allow() is False, return without calling call_target, and after calling, call record_success/record_failure according to the result.
Don't wait when over the limit
Put a concurrency limit with --max-inflight and answer E905 immediately when it is full.
If BoundedSemaphore's acquire(blocking=False) is False, the limit is full. After calling, use try/finally so that release happens even if an exception occurs.
Divide the bulkheads per target
Put the breaker and the concurrency limit separately for each of the CORE and CLAIM targets, so that core banking's breakdown or saturation does not spread to claims.
Change what you had one of into a dictionary keyed by target name. With a single bulkhead, the water of one compartment fills the whole ship.
Record the state transitions
In --breaker-log, leave one line 'time|target|FROM->TO' each time the state changes.
Compare state before and after allow() and before and after record, and write one line if they differ. Several threads write, so take a lock when writing.
A business rejection is not a failure
Narrow what counts as a failure to E500, E901 and E902, and count B2xx as a success.
Insufficient funds is core banking answering healthily. If you count it as a failure, business rejections at month end alone cut off the way to a perfectly fine core banking.