Lesson 19 / 25
Persistence, Replication and Data Loss
Understand what survives restarts and failovers when Redis is your message bus.
In-memory first, durable by configuration
Redis keeps data in memory. For streams to survive restarts, enable persistence: RDB snapshots periodically write the whole dataset (fast restarts, but you can lose the changes since the last snapshot), and the AOF (append-only file) logs every write, with appendfsync everysec (the common default, losing at most about a second of writes on a crash) or always (safest, slowest). Many deployments use both. Replication is asynchronous: a primary acknowledges XADD before replicas have it, so if the primary fails and a replica is promoted, the most recent writes can be lost. The WAIT numreplicas timeout command lets a client block until a number of replicas have acknowledged its writes, reducing (not eliminating) that window. Some managed offerings, such as Amazon MemoryDB, add a durable multi-zone transaction log for stronger guarantees. Pub/Sub messages are never persisted at all. Match these properties to your requirements and document them.
Where messages can be lost
Writes live in memory, then the AOF and replicas; each gap is a window for loss on failure.
Durability settings for a stream workload
AOF every second plus a replica acknowledgement for important writes.
# redis.conf
appendonly yes
appendfsync everysec # at most ~1 s of writes lost on a crash
save 900 1 300 10 # RDB snapshots as well
maxmemory-policy noeviction # never evict stream data to make room
# client side, for an important event:
XADD events:payments '*' eventId 77 type PaymentCaptured
WAIT 1 100 # wait up to 100 ms for at least 1 replica to acknowledgeNever let eviction delete streams
With an eviction policy such as allkeys-lru, Redis may evict a whole stream key under memory pressure. Run message workloads on an instance with noeviction (or volatile-* policies with TTLs only on cache keys), separate from caches.
Quick check: Why can a Redis failover lose recently added stream entries?
- Replication is asynchronous, so the promoted replica may not have the latest writes
- Streams are never replicated
- XADD deletes old entries
- Consumer groups are not replicated
Answer
Replication is asynchronous, so the promoted replica may not have the latest writes — Acknowledged writes may not have reached replicas when the primary fails.