पाठ 1 / 25
Failure Is Normal in Distributed Systems
Explain how small failures cascade and what resilience patterns aim to achieve.
Design for the day something breaks
Every service that calls another over a network inherits that network's problems: lost packets, slow responses, overloaded dependencies, deploys in progress, expired certificates. At scale, something is always failing. Without protection, a single slow dependency can take down a whole system: callers wait, their threads and connection pools fill up, they stop answering their own callers, and the failure cascades upstream. Retrying everything makes it worse by multiplying load on the struggling service. Resilience patterns are well-tested techniques that keep a system useful when parts of it misbehave. Their goals: fail fast instead of hanging, contain a failure so it does not spread, protect overloaded components so they can recover, recover automatically when the problem passes, and degrade gracefully, giving users a reduced service rather than an error page.
A failure spreading upstream
One slow dependency backs up its callers, then their callers, unless something stops it.
How a slow dependency exhausts a caller
Little's law shows why waiting is so dangerous.
checkout service: 200 worker threads, 400 requests/second
normal payment latency: 100 ms -> in flight = 400 x 0.1 = 40 threads busy
payment slows to 5 s -> in flight = 400 x 5 = 2,000 threads needed
only 200 threads exist -> the pool fills in half a second,
every other endpoint of checkout (cart, coupons) now waits too,
and the load balancer starts seeing checkout as down
fix: a 300 ms timeout + circuit breaker + bulkhead for payment callsSlow is worse than down
A dependency that refuses connections fails in milliseconds. One that accepts connections and never answers holds your resources hostage. Most cascading failures start with slowness, not crashes.
त्वरित जाँच: Why can retries make an outage worse?
- They multiply load on a dependency that is already struggling
- Retries are always slower than the first attempt
- They disable timeouts
- They delete cached data
Answer
They multiply load on a dependency that is already struggling — Uncontrolled retries amplify traffic exactly when the dependency has the least capacity.