Lesson 7 / 25
Replication
Use leader-follower replication to scale reads and survive failures, and handle replica lag.
Copies of the data
Replication keeps copies of data on several machines, for availability (another copy can take over) and read scaling (reads spread across copies). In leader-follower (primary-replica) replication, all writes go to the leader, which streams changes to followers that serve reads. Synchronous replication waits for a follower to confirm before acknowledging a write: no data loss on failover, but higher write latency and writes stall if the follower is slow. Asynchronous replication acknowledges immediately: fast, but followers lag and a failover can lose the last writes. Many systems use semi-synchronous replication, waiting for one follower. Replica lag causes anomalies: a user updates a profile and then reads a stale copy. Fixes include read-your-writes (read from the leader for a short time after the user writes), monotonic reads (keep a user on the same replica) and lag-aware routing. Multi-leader and leaderless designs accept writes in many places, at the cost of conflict resolution.
One leader, many followers
Writes go to the leader; changes stream to followers that serve reads.
Read-your-writes routing
After a write, this user's reads go to the leader for a few seconds.
import time
RECENT_WRITE_WINDOW = 5 # seconds, longer than typical replica lag
def db_for_read(user_id, session):
last_write = session.get("last_write_at", 0)
if time.time() - last_write < RECENT_WRITE_WINDOW:
return leader # guarantees the user sees their own change
return pick_replica() # spread other reads across followers
def update_profile(user_id, changes, session):
leader.execute("UPDATE profiles SET ... WHERE user_id = %s", (user_id,))
session["last_write_at"] = time.time()Replicas are not backups
A mistaken DELETE replicates to every follower within milliseconds. Keep real backups with point-in-time recovery, and test restoring them.
Quick check: A user edits their name and the next page still shows the old name. What is the likely cause?
- The read was served by an asynchronous replica that had not caught up
- The leader rejected the write
- Synchronous replication
- The load balancer cached the name forever
Answer
The read was served by an asynchronous replica that had not caught up — Replica lag causes stale reads right after a write; read-your-writes routing fixes it.