SkillByAIOpen interactive version →

Lesson 23 / 25

Performance and Capacity

Tune throughput and avoid the patterns that make RabbitMQ slow.

Keep queues short and messages small

RabbitMQ performs best when queues are short: messages flow through rather than piling up. Long backlogs consume disk and memory, slow down operations and increase recovery time after restarts, so scale consumers to keep up and use length limits as a safety net. Keep messages small (kilobytes); for large payloads store the data in object storage and send a reference. Use publisher confirms asynchronously or in batches rather than waiting for each confirm. Set consumer prefetch appropriately. Spread load across several queues when one queue is a bottleneck, because each queue is served by a single Erlang process (and for quorum queues, a single leader), so one very hot queue is limited by one CPU core's work; the consistent hash exchange or super streams help. Avoid huge numbers of queues or bindings created dynamically without cleanup. Measure with PerfTest (the official load-testing tool) on hardware similar to production before committing to throughput numbers.

Load testing with PerfTest

Simulate producers and consumers to find throughput and latency on your hardware.

# 2 producers, 4 consumers, quorum queue, 1 KB messages, confirms in flight up to 100
docker run -it --rm pivotalrabbitmq/perf-test:latest \
  --uri amqp://app:secret@rabbit.internal:5672/shop \
  --producers 2 --consumers 4 \
  --quorum-queue --queue perf-test-q \
  --size 1000 --confirm 100 --qos 50 \
  --time 120

# watch publish/consume rates and latency percentiles in the output,
# and node CPU, memory and disk in Grafana while it runs

A conveyor belt, not a warehouse

RabbitMQ is happiest as a conveyor belt with parcels moving steadily. When parcels pile up into a warehouse, everything gets slower: finding space, counting stock, reopening after a power cut.

Quick check: A single queue receives so many messages that its node shows one CPU core saturated. What usually helps?

  • Making messages larger
  • Spreading the work across several queues (for example with a consistent hash exchange or super streams)
  • Disabling acknowledgements
  • Adding more bindings to the same queue
Answer

Spreading the work across several queues (for example with a consistent hash exchange or super streams) — Each queue has one process (or leader), so partitioning across queues spreads the load.