# Failure Is Normal in Distributed Systems — Rate Limiting, Circuit Breakers & Resilience Patterns

Source: https://www.skillbyai.com/en/resilience-patterns/f-why

> Explain how small failures cascade and what resilience patterns aim to achieve.

## Design for the day something breaks

Every service that calls another over a network inherits that network's problems: lost packets, slow responses, overloaded dependencies, deploys in progress, expired certificates. At scale, **something is always failing**. Without protection, a single slow dependency can take down a whole system: callers wait, their threads and connection pools fill up, they stop answering their own callers, and the failure **cascades** upstream. Retrying everything makes it worse by multiplying load on the struggling service. **Resilience patterns** are well-tested techniques that keep a system useful when parts of it misbehave. Their goals: **fail fast** instead of hanging, **contain** a failure so it does not spread, **protect** overloaded components so they can recover, **recover automatically** when the problem passes, and **degrade gracefully**, giving users a reduced service rather than an error page.

## A failure spreading upstream

One slow dependency backs up its callers, then their callers, unless something stops it.

![A chain of four boxes from right to left, the rightmost dark and slow, with queues of small dots piling up in front of each box towards the left.](assets/figures/resilience-patterns/section-1-map.svg) — Figure 1.1 — A cascading failure moving from a slow dependency to the edge.

## How a slow dependency exhausts a caller

Little's law shows why waiting is so dangerous.

```text
checkout service: 200 worker threads, 400 requests/second
normal payment latency: 100 ms  -> in flight = 400 x 0.1 = 40 threads busy
payment slows to 5 s            -> in flight = 400 x 5   = 2,000 threads needed

only 200 threads exist -> the pool fills in half a second,
every other endpoint of checkout (cart, coupons) now waits too,
and the load balancer starts seeing checkout as down

fix: a 300 ms timeout + circuit breaker + bulkhead for payment calls
```

## Slow is worse than down

A dependency that refuses connections fails in milliseconds. One that accepts connections and never answers holds your resources hostage. Most cascading failures start with slowness, not crashes.

**Quiz:** Why can retries make an outage worse?

- [x] They multiply load on a dependency that is already struggling
- [ ] Retries are always slower than the first attempt
- [ ] They disable timeouts
- [ ] They delete cached data

*Answer:* They multiply load on a dependency that is already struggling. Uncontrolled retries amplify traffic exactly when the dependency has the least capacity.
