SkillByAIOpen interactive version →

Lesson 19 / 25

Persistence, Replication and Data Loss

Understand what survives restarts and failovers when Redis is your message bus.

In-memory first, durable by configuration

Redis keeps data in memory. For streams to survive restarts, enable persistence: RDB snapshots periodically write the whole dataset (fast restarts, but you can lose the changes since the last snapshot), and the AOF (append-only file) logs every write, with appendfsync everysec (the common default, losing at most about a second of writes on a crash) or always (safest, slowest). Many deployments use both. Replication is asynchronous: a primary acknowledges XADD before replicas have it, so if the primary fails and a replica is promoted, the most recent writes can be lost. The WAIT numreplicas timeout command lets a client block until a number of replicas have acknowledged its writes, reducing (not eliminating) that window. Some managed offerings, such as Amazon MemoryDB, add a durable multi-zone transaction log for stronger guarantees. Pub/Sub messages are never persisted at all. Match these properties to your requirements and document them.

Where messages can be lost

Writes live in memory, then the AOF and replicas; each gap is a window for loss on failure.

Figure 7.1 — Memory, persistence and asynchronous replication.

Durability settings for a stream workload

AOF every second plus a replica acknowledgement for important writes.

# redis.conf
appendonly yes
appendfsync everysec          # at most ~1 s of writes lost on a crash
save 900 1 300 10             # RDB snapshots as well
maxmemory-policy noeviction   # never evict stream data to make room

# client side, for an important event:
XADD events:payments '*' eventId 77 type PaymentCaptured
WAIT 1 100                    # wait up to 100 ms for at least 1 replica to acknowledge

Never let eviction delete streams

With an eviction policy such as allkeys-lru, Redis may evict a whole stream key under memory pressure. Run message workloads on an instance with noeviction (or volatile-* policies with TTLs only on cache keys), separate from caches.

Quick check: Why can a Redis failover lose recently added stream entries?

  • Replication is asynchronous, so the promoted replica may not have the latest writes
  • Streams are never replicated
  • XADD deletes old entries
  • Consumer groups are not replicated
Answer

Replication is asynchronous, so the promoted replica may not have the latest writes — Acknowledged writes may not have reached replicas when the primary fails.