# Replication — Scalability, Availability & Reliability

Source: https://www.skillbyai.com/en/scalability/d-replication

> Use leader-follower replication to scale reads and survive failures, and handle replica lag.

## Copies of the data

**Replication** keeps copies of data on several machines, for availability (another copy can take over) and read scaling (reads spread across copies). In **leader-follower** (primary-replica) replication, all writes go to the leader, which streams changes to followers that serve reads. **Synchronous** replication waits for a follower to confirm before acknowledging a write: no data loss on failover, but higher write latency and writes stall if the follower is slow. **Asynchronous** replication acknowledges immediately: fast, but followers lag and a failover can lose the last writes. Many systems use **semi-synchronous** replication, waiting for one follower. **Replica lag** causes anomalies: a user updates a profile and then reads a stale copy. Fixes include **read-your-writes** (read from the leader for a short time after the user writes), monotonic reads (keep a user on the same replica) and lag-aware routing. **Multi-leader** and **leaderless** designs accept writes in many places, at the cost of conflict resolution.

## One leader, many followers

Writes go to the leader; changes stream to followers that serve reads.

![A central cylinder receiving a write arrow, with streaming arrows to three smaller cylinders, each receiving read arrows from clients.](assets/figures/scalability/section-3-map.svg) — Figure 3.1 — Leader-follower replication for read scaling and failover.

## Read-your-writes routing

After a write, this user's reads go to the leader for a few seconds.

```python
import time

RECENT_WRITE_WINDOW = 5  # seconds, longer than typical replica lag

def db_for_read(user_id, session):
    last_write = session.get("last_write_at", 0)
    if time.time() - last_write < RECENT_WRITE_WINDOW:
        return leader          # guarantees the user sees their own change
    return pick_replica()      # spread other reads across followers

def update_profile(user_id, changes, session):
    leader.execute("UPDATE profiles SET ... WHERE user_id = %s", (user_id,))
    session["last_write_at"] = time.time()
```

## Replicas are not backups

A mistaken `DELETE` replicates to every follower within milliseconds. Keep real backups with point-in-time recovery, and test restoring them.

**Quiz:** A user edits their name and the next page still shows the old name. What is the likely cause?

- [x] The read was served by an asynchronous replica that had not caught up
- [ ] The leader rejected the write
- [ ] Synchronous replication
- [ ] The load balancer cached the name forever

*Answer:* The read was served by an asynchronous replica that had not caught up. Replica lag causes stale reads right after a write; read-your-writes routing fixes it.
