पाठ 1 / 25

Failure Is Normal in Distributed Systems

Explain how small failures cascade and what resilience patterns aim to achieve.

Design for the day something breaks

Every service that calls another over a network inherits that network's problems: lost packets, slow responses, overloaded dependencies, deploys in progress, expired certificates. At scale, something is always failing. Without protection, a single slow dependency can take down a whole system: callers wait, their threads and connection pools fill up, they stop answering their own callers, and the failure cascades upstream. Retrying everything makes it worse by multiplying load on the struggling service. Resilience patterns are well-tested techniques that keep a system useful when parts of it misbehave. Their goals: fail fast instead of hanging, contain a failure so it does not spread, protect overloaded components so they can recover, recover automatically when the problem passes, and degrade gracefully, giving users a reduced service rather than an error page.

A failure spreading upstream

One slow dependency backs up its callers, then their callers, unless something stops it.

A chain of four boxes from right to left, the rightmost dark and slow, with queues of small dots piling up in front of each box towards the left.
Figure 1.1 — A cascading failure moving from a slow dependency to the edge.

How a slow dependency exhausts a caller

Little's law shows why waiting is so dangerous.

checkout service: 200 worker threads, 400 requests/second
normal payment latency: 100 ms  -> in flight = 400 x 0.1 = 40 threads busy
payment slows to 5 s            -> in flight = 400 x 5   = 2,000 threads needed

only 200 threads exist -> the pool fills in half a second,
every other endpoint of checkout (cart, coupons) now waits too,
and the load balancer starts seeing checkout as down

fix: a 300 ms timeout + circuit breaker + bulkhead for payment calls

Slow is worse than down

A dependency that refuses connections fails in milliseconds. One that accepts connections and never answers holds your resources hostage. Most cascading failures start with slowness, not crashes.

त्वरित जाँच: Why can retries make an outage worse?

  • They multiply load on a dependency that is already struggling
  • Retries are always slower than the first attempt
  • They disable timeouts
  • They delete cached data
Answer

They multiply load on a dependency that is already struggling — Uncontrolled retries amplify traffic exactly when the dependency has the least capacity.