पाठ 13 / 25

Redundancy and Failover

Remove single points of failure with redundancy and well-tested failover.

Nothing important should exist only once

A single point of failure (SPOF) is any component whose failure takes the service down: one database server, one load balancer, one DNS provider, one engineer who knows the deploy script. Redundancy removes SPOFs by having spares. N+1 means enough capacity to serve peak load plus one spare; N+2 allows maintenance and a failure at the same time. In active-active setups every copy serves traffic and survivors absorb the load when one fails; this requires capacity headroom and data that can be served from several places. In active-passive setups a standby waits to take over: simpler for stateful systems like databases, but the standby may be cold, and failover takes time and can fail itself. Failover needs reliable failure detection (health checks and heartbeats, tuned to avoid false positives) and protection against split brain, where two nodes both believe they are the leader; consensus systems and fencing solve that.

Active-active and active-passive

In active-active both sides serve; in active-passive the standby waits to take over.

Left: two equal boxes both receiving traffic arrows. Right: one bright box receiving traffic and a faded box beside it connected by a dashed heartbeat line.
Figure 5.1 — Two redundancy styles.

Hunting for single points of failure

Walk a request path and ask "what if this one disappears?" at every hop.

request path                      redundant?   notes
--------------------------------  -----------  -------------------------------------
DNS provider                      no -> fix    add secondary provider or anycast DNS
CDN / edge                        yes          multiple PoPs
load balancer                     yes          managed, multi-zone
app instances                     yes          6 pods across 3 zones, N+1 capacity
Redis cache                       no -> fix    single node; add replica + failover
PostgreSQL                        partial      HA standby in 2nd zone; failover tested?
payment gateway                   no           external; add timeout + fallback message
TLS certificate renewal           no -> fix    automate; alert 21 days before expiry
the one person who can deploy     no -> fix    document runbook, train two more people

An untested failover is a hope, not a plan

Failovers that are never exercised tend to fail when needed: expired credentials, missing config on the standby, DNS TTLs longer than expected. Practise failover on a schedule.

त्वरित जाँच: What is split brain in an active-passive database cluster?

  • Both nodes believing they are the primary and accepting writes
  • Two replicas with different schemas
  • A slow replica
  • A load balancer with two algorithms
Answer

Both nodes believing they are the primary and accepting writes — Split brain causes divergent writes; quorum-based leader election and fencing prevent it.