पाठ 13 / 25
Redundancy and Failover
Remove single points of failure with redundancy and well-tested failover.
Nothing important should exist only once
A single point of failure (SPOF) is any component whose failure takes the service down: one database server, one load balancer, one DNS provider, one engineer who knows the deploy script. Redundancy removes SPOFs by having spares. N+1 means enough capacity to serve peak load plus one spare; N+2 allows maintenance and a failure at the same time. In active-active setups every copy serves traffic and survivors absorb the load when one fails; this requires capacity headroom and data that can be served from several places. In active-passive setups a standby waits to take over: simpler for stateful systems like databases, but the standby may be cold, and failover takes time and can fail itself. Failover needs reliable failure detection (health checks and heartbeats, tuned to avoid false positives) and protection against split brain, where two nodes both believe they are the leader; consensus systems and fencing solve that.
Active-active and active-passive
In active-active both sides serve; in active-passive the standby waits to take over.
Hunting for single points of failure
Walk a request path and ask "what if this one disappears?" at every hop.
request path redundant? notes
-------------------------------- ----------- -------------------------------------
DNS provider no -> fix add secondary provider or anycast DNS
CDN / edge yes multiple PoPs
load balancer yes managed, multi-zone
app instances yes 6 pods across 3 zones, N+1 capacity
Redis cache no -> fix single node; add replica + failover
PostgreSQL partial HA standby in 2nd zone; failover tested?
payment gateway no external; add timeout + fallback message
TLS certificate renewal no -> fix automate; alert 21 days before expiry
the one person who can deploy no -> fix document runbook, train two more peopleAn untested failover is a hope, not a plan
Failovers that are never exercised tend to fail when needed: expired credentials, missing config on the standby, DNS TTLs longer than expected. Practise failover on a schedule.
त्वरित जाँच: What is split brain in an active-passive database cluster?
- Both nodes believing they are the primary and accepting writes
- Two replicas with different schemas
- A slow replica
- A load balancer with two algorithms
Answer
Both nodes believing they are the primary and accepting writes — Split brain causes divergent writes; quorum-based leader election and fencing prevent it.