TT Lab
Get started
Learn Learning paths Courses

Microservice Architecture

The Circuit Breaker's State Machine

Continue in TT Lab

Summary

A circuit breaker does not fix failures. When a failure is certain, it fails fast to protect the caller's resources from drying up.

Why this was needed

Retries alone are not enough. When the downstream is really dead, retrying only makes the situation worse. And from the caller's point of view too, waiting 3 seconds for a call that will fail anyway is a waste of connections and threads.

When a circuit breaker judges "judging from recent results, this service is not working now", it does not attempt the call at all and returns a failure immediately. Because it fails in microseconds, the caller's resources do not dry up, and the downstream gets room to recover.

How it works

There are three states.

The three states and four transitions of a circuit breaker — CLOSED lets everything through and records results in a window; when the failure rate exceeds the threshold it goes to OPEN and rejects immediately; after waitDurationInOpenState passes it goes to HALF_OPEN and allows only trial calls; and depending on the result it returns to CLOSED or becomes OPEN again. 4xx is not counted as a failure

CLOSED is normal. All calls pass through, and the results are recorded in a sliding window. When the failure rate in the window exceeds the threshold, it goes to OPEN.

OPEN is blocking. All calls are rejected immediately. When waitDurationInOpenState passes, it goes to HALF_OPEN.

HALF_OPEN is a trial. It allows only a set number of trial calls, and depending on the result it returns to CLOSED or goes back to OPEN.

There is no right answer for the settings, but there is a starting point. These are the values the author uses in operation. slidingWindowSize: 10, minimumNumberOfCalls: 5, failureRateThreshold: 50, slowCallDurationThreshold: 3s, waitDurationInOpenState: 30s, permittedNumberOfCallsInHalfOpenState: 3. You adjust the threshold according to the nature of the service — 30–40% where failure is fatal, such as payment, and 60–70% where leniency is acceptable, such as notifications.

Here is the single most important design decision. What to count as a failure. A 4xx is not a failure. The client sent something wrong, so it is not a matter for the circuit to step in. There is a real incident. The inventory service began returning 400 for a particular product, that 400 was counted in the failure rate, the circuit went OPEN, and even lookups of normal products were all blocked. The circuit must react only to 5xx, timeouts, and connection failures.

What you meet in the field

The second most common incident is HALF_OPEN flapping. If you set permittedNumberOfCallsInHalfOpenState to 1, every time one trial call happens to fail, it drops back to OPEN. In a situation where the downstream responds only intermittently, it can never return to CLOSED. Use a minimum of 3, usually 5–10.

Third. Many implementations have a default of 100 for minimumNumberOfCalls. For a service with a low call frequency, the outage ends before the window fills, so the circuit never opens even once. Half of "I added a circuit but it doesn't open" is for this reason.

And always use a circuit with a fallback. If the fallback is return null, the circuit has merely turned the outage into a NullPointerException. It must be one of a cached value, a reduced response, or a clear notice.

A circuit breaker alone is not enough

A circuit is a device that reacts after things have already gone bad. There are things that must be placed before and after it, and only when the four come together do they form a single line of defense.

Timeout. The most basic and the one most often omitted. Without a timeout, the caller's thread is tied up indefinitely without the circuit even having a chance to count failures. Decide the value by looking at the other side's latency distribution, but it must get shorter as you go deeper in the call chain. If the outside is 3 seconds and the inside 5 seconds, the outside has already given up before the inside answers, and the inside's work is wasted entirely.

Bulkhead. You separately limit the number of concurrent calls that can be used for one counterpart. Without this, a single slowed service occupies the entire shared thread pool or connection pool, and even features unrelated to that service stop along with it. Most of the paths along which failures spread are these shared resources.

Retry. It works in the opposite direction to the circuit, so rules are needed when using them together. Only on idempotent calls, with a cap, and with randomness mixed into the interval. And if several layers of the chain each retry, it multiplies. If three layers each retry three times, the worst case is 27 attempts. The rule is that only one layer of the chain handles retries.

Fallback. It is what to return when the circuit is open. As said earlier, a fallback that throws the exception again is nothing. A good fallback is usually one of three: a slightly stale cached value, a reduced response with just that part removed, or notifying, in a form the screen can understand, that "this information could not be fetched now".

Finally, these four must be tested. A circuit that has never been opened in normal times works for the first time in a real outage, and if the fallback has a defect at that time, the line of defense actually creates the outage. This is why teams regularly run drills in which they deliberately kill the counterpart.

What you will do in the next lab

You implement a circuit breaker yourself. You see the state transitions with your own eyes, prove with a counter that calls really do not go to the downstream in OPEN, make 4xx not counted as failures, and finally leave the whole transition history in a file.