# Clustering and High Availability — RabbitMQ

Source: https://www.skillbyai.com/en/rabbitmq/o-cluster

> Run a RabbitMQ cluster and understand what is replicated.

## Several nodes, one logical broker

A RabbitMQ **cluster** is a group of nodes that share **metadata** (users, vhosts, exchanges, bindings, queue definitions, policies), so clients can connect to any node. Message data is replicated only for **quorum queues** and **streams**; classic queues still live on one node. Use an **odd number** of nodes, typically **3** (or 5 for larger deployments), so quorum-based queues and metadata keep a majority if one node fails. Recent releases use **Khepri**, a Raft-based metadata store, in place of the older Mnesia database, improving behaviour during network partitions. Spread nodes across **availability zones** with low latency between them; clusters are not designed to span distant regions, so for multi-region setups use separate clusters connected by the **Shovel** or **Federation** plugins. Put a load balancer or DNS name in front of nodes, enable **automatic connection recovery** in clients, and upgrade with **rolling upgrades** following the documented version path.

## A three-node cluster across zones

Metadata is shared everywhere; quorum queues keep replicas on several nodes.

![Three server boxes in three zone columns linked in a triangle, each holding small queue replicas with one highlighted as leader.](assets/figures/rabbitmq/section-6-map.svg) — Figure 6.1 — Nodes, replicas and leaders in a cluster.

## Cluster health checks

Use these in readiness probes and runbooks.

```bash
rabbitmq-diagnostics cluster_status
rabbitmq-diagnostics check_running
rabbitmq-diagnostics check_local_alarms          # memory or disk alarms on this node
rabbitmq-queues check_if_node_is_quorum_critical   # unsafe to stop this node now?

# on Kubernetes, the Cluster Operator manages this declaratively:
# kind: RabbitmqCluster, spec.replicas: 3
```

## Check before stopping a node

Stopping a node that holds the only in-sync majority member of quorum queues can make them unavailable. `check_if_node_is_quorum_critical` tells you whether it is safe to take the node down for maintenance.

**Quiz:** Why are RabbitMQ clusters usually built with an odd number of nodes such as three?

- [ ] Licensing rules
- [ ] Exchanges require three nodes
- [x] So quorum queues and metadata keep a majority if one node fails
- [ ] Odd numbers make TLS faster

*Answer:* So quorum queues and metadata keep a majority if one node fails. Majority-based replication needs a quorum; three nodes tolerate one failure.
