Lesson 16 / 25
Clustering and High Availability
Run a RabbitMQ cluster and understand what is replicated.
Several nodes, one logical broker
A RabbitMQ cluster is a group of nodes that share metadata (users, vhosts, exchanges, bindings, queue definitions, policies), so clients can connect to any node. Message data is replicated only for quorum queues and streams; classic queues still live on one node. Use an odd number of nodes, typically 3 (or 5 for larger deployments), so quorum-based queues and metadata keep a majority if one node fails. Recent releases use Khepri, a Raft-based metadata store, in place of the older Mnesia database, improving behaviour during network partitions. Spread nodes across availability zones with low latency between them; clusters are not designed to span distant regions, so for multi-region setups use separate clusters connected by the Shovel or Federation plugins. Put a load balancer or DNS name in front of nodes, enable automatic connection recovery in clients, and upgrade with rolling upgrades following the documented version path.
A three-node cluster across zones
Metadata is shared everywhere; quorum queues keep replicas on several nodes.
Cluster health checks
Use these in readiness probes and runbooks.
rabbitmq-diagnostics cluster_status
rabbitmq-diagnostics check_running
rabbitmq-diagnostics check_local_alarms # memory or disk alarms on this node
rabbitmq-queues check_if_node_is_quorum_critical # unsafe to stop this node now?
# on Kubernetes, the Cluster Operator manages this declaratively:
# kind: RabbitmqCluster, spec.replicas: 3Check before stopping a node
Stopping a node that holds the only in-sync majority member of quorum queues can make them unavailable. check_if_node_is_quorum_critical tells you whether it is safe to take the node down for maintenance.
Quick check: Why are RabbitMQ clusters usually built with an odd number of nodes such as three?
- Licensing rules
- Exchanges require three nodes
- So quorum queues and metadata keep a majority if one node fails
- Odd numbers make TLS faster
Answer
So quorum queues and metadata keep a majority if one node fails — Majority-based replication needs a quorum; three nodes tolerate one failure.