Articles /

Circuit Breaker: isolating failures before they spread

How Circuit Breakers help contain failures, protect resources and make applications more resilient.

7 min read#architecture #go #java #fault-tolerance

How to keep a fire from spreading while it is still growing

The art of isolation

I have spent the last few weeks thinking about how to avoid widespread failures in different kinds of architecture, from monoliths to microservices, and mostly about how to keep a piece, a domain, a service or an external system that is failing from taking down the rest of the system and the user’s journey along with it.

With that in mind, I like to frame the circuit breaker as an analogy to a forest on fire: when part of the forest catches fire, what we do is isolate the burning areas so the fire does not reach the rest and turn into a much larger disaster.

The Circuit Breaker idea starts from the same principle. While everything is fine, traffic reaches every area it needs to, normally. But when some area degrades or becomes unavailable, the breaker steps in and stops traffic from reaching the affected piece, so it does not make the damage worse or drain resources. And above all, it keeps the application from attempting an action that is doomed to fail, or from causing a broad degradation when the process sits in the middle of the user’s journey.

Circuit Breaker states

Closed

The request is routed to the service (or external system) normally, and the proxy keeps a count of how many times the request failed within a time window. If requests fail, the circuit breaker moves to the Open state and starts a timeout. When that timeout expires, the state changes to Half-open.

Open

Every request sent to the service fails fast, so the caller does not sit waiting and can move on to the next request.

Half-open

In this state a limited number of requests is allowed through to the application, and they are monitored, again by the proxy. If those requests succeed, the circuit breaker goes back to the Closed state. If any of the requests that got through fails, the circuit breaker takes it as the degradation still being there and the state goes back to Open, restarting the timer and the whole flow.

Implementations in Go and Java

Even though the concept is the same, a Circuit Breaker can be implemented in different ways. It can live inside the application itself, wrapping calls to external dependencies, or in an infrastructure layer such as a proxy or a service mesh.

No matter where it is implemented, the logic we are after is pretty much the same:

  1. Run the call normally while the circuit is closed.
  2. Monitor the failures that happen during those calls.
  3. Open the circuit when the number or ratio of failures crosses an acceptable limit.
  4. Fail fast while the circuit is open.
  5. After a given period, allow a few test calls.
  6. Close the circuit again if the dependency has recovered.

Go

A simple way to think about a Go implementation is to wrap the call we want to protect in a function controlled by the Circuit Breaker.

result, err := breaker.Execute(func() (interface{}, error) {
    return callExternalService()
})

What matters here is not exactly the library, but the separation of responsibilities.

The callExternalService function should not need to know there is a circuit breaker guarding its execution. As far as it is concerned, its job is still to make the call.

It is the breaker that decides whether the call can happen at that moment.

That also means that, while the circuit is open, we can return immediately without even consuming an HTTP connection, opening a connection to another service or waiting on a timeout we know is probably coming.

Java

In Java the idea is pretty much the same.

We can wrap an external call with a Circuit Breaker and let it monitor the outcome of each execution.

Supplier<Response> decoratedSupplier =
    circuitBreaker.decorateSupplier(() -> externalService.call());

Response response = decoratedSupplier.get();

Frameworks like Spring also let this kind of behavior be applied more declaratively, but under the hood the decision stays the same.

On tuning

Maybe one of the trickiest parts of using Circuit Breakers is configuring their triggers. There is no magic number that works for every system.

Imagine we configure our circuit to open after two consecutive failures. In an application taking thousands of requests per second, two failures may be nothing but normal noise. On the other hand, if we set the circuit to wait for hundreds of failures before reacting, by the time it finally opens the damage may already be done.

So a few things need to be considered.

Request window

We need to define which set of calls will be used to decide whether a dependency is healthy.

For example:

Out of the last 100 requests, how many failed?

Or:

In the last 30 seconds, what percentage of requests ended in an error?

These two questions look similar, but they can produce quite different behavior depending on the application’s volume.

Failure threshold

Next we need to decide how many failures are enough to open the circuit.

For example, we can consider a dependency degraded when 50% of the calls inside the observed window end in an error.

But again, that depends on the domain.

A 5% error rate may be completely unacceptable for a payment operation, while a less critical process can live with it for a while.

Open state timeout

When the circuit opens, we also need to decide how long to wait before testing the dependency again.

If we retry too early, we may keep pushing on a system that is still trying to recover.

If we wait too long, we may keep a feature unavailable even after the dependency is back to normal, which is exactly what we want to avoid.

That balance is precisely why the Half-open state exists.

Which errors actually count?

Not every error should contribute to opening a Circuit Breaker.

A 400 Bad Request, for instance, usually means something is wrong with the request the caller sent, and trying again five seconds from now probably will not change anything.

Errors like timeouts, connection failures or certain 5xx responses, on the other hand, can mean the dependency really is degraded.

Trade-offs

Since not everything is roses, and the only certainty, besides the dishes waiting in the sink right now, is that everything has pros and cons, there are costs that come with this pattern.

The first one is complexity.

The Circuit Breaker adds one more state to the system and, with it, more behaviors that need to be understood and tested.

We no longer think only about:

Is the dependency working?

Now we also think about:

Is the circuit closed, open or half-open?

Why did it open?

How long until it tries again?

How many requests are being used as a test?

There is also a much stronger need for observability.

If a Circuit Breaker opened in production, we probably want to know about it.

We want to know which dependency caused it to open, what the error rate was at that moment, how long the circuit has been open and how many times it flipped between its states.

Otherwise we can end up in a curious situation where the Circuit Breaker is doing exactly what it was built for, protecting our application, while a dependency stays quietly broken behind it.

Another point is that failing fast is still failing.

The Circuit Breaker does not magically make a dependency available. Quite the opposite: it only keeps a known failure from consuming resources and spreading through the rest of the system.

That is why we usually need to pair the pattern with some degradation strategy.

If a recommendations API is unavailable, maybe we can simply hide the recommendations.

If a shipping cost service is unavailable, maybe the purchase cannot be completed at all.

How we react to an open circuit depends directly on how important that dependency is to the journey being run.

A Circuit Breaker does not work alone

A Circuit Breaker gets more interesting when used together with other resilience mechanisms.

Timeouts bound how long we are willing to wait on a dependency.

Retries allow repeating operations that may have failed for transient reasons.

Circuit Breakers stop us from trying again when we have enough evidence that the dependency is degraded.

These mechanisms complement each other, but they can also work against each other when configured badly.