# Redundancy and Failover — Scalability, Availability & Reliability

Source: https://www.skillbyai.com/en/scalability/a-redundancy

> Remove single points of failure with redundancy and well-tested failover.

## Nothing important should exist only once

A **single point of failure (SPOF)** is any component whose failure takes the service down: one database server, one load balancer, one DNS provider, one engineer who knows the deploy script. **Redundancy** removes SPOFs by having spares. **N+1** means enough capacity to serve peak load plus one spare; **N+2** allows maintenance and a failure at the same time. In **active-active** setups every copy serves traffic and survivors absorb the load when one fails; this requires capacity headroom and data that can be served from several places. In **active-passive** setups a standby waits to take over: simpler for stateful systems like databases, but the standby may be cold, and **failover** takes time and can fail itself. Failover needs reliable **failure detection** (health checks and heartbeats, tuned to avoid false positives) and protection against **split brain**, where two nodes both believe they are the leader; consensus systems and fencing solve that.

## Active-active and active-passive

In active-active both sides serve; in active-passive the standby waits to take over.

![Left: two equal boxes both receiving traffic arrows. Right: one bright box receiving traffic and a faded box beside it connected by a dashed heartbeat line.](assets/figures/scalability/section-5-map.svg) — Figure 5.1 — Two redundancy styles.

## Hunting for single points of failure

Walk a request path and ask "what if this one disappears?" at every hop.

```text
request path                      redundant?   notes
--------------------------------  -----------  -------------------------------------
DNS provider                      no -> fix    add secondary provider or anycast DNS
CDN / edge                        yes          multiple PoPs
load balancer                     yes          managed, multi-zone
app instances                     yes          6 pods across 3 zones, N+1 capacity
Redis cache                       no -> fix    single node; add replica + failover
PostgreSQL                        partial      HA standby in 2nd zone; failover tested?
payment gateway                   no           external; add timeout + fallback message
TLS certificate renewal           no -> fix    automate; alert 21 days before expiry
the one person who can deploy     no -> fix    document runbook, train two more people
```

## An untested failover is a hope, not a plan

Failovers that are never exercised tend to fail when needed: expired credentials, missing config on the standby, DNS TTLs longer than expected. Practise failover on a schedule.

**Quiz:** What is split brain in an active-passive database cluster?

- [x] Both nodes believing they are the primary and accepting writes
- [ ] Two replicas with different schemas
- [ ] A slow replica
- [ ] A load balancer with two algorithms

*Answer:* Both nodes believing they are the primary and accepting writes. Split brain causes divergent writes; quorum-based leader election and fencing prevent it.
