Lesson 14 / 25
Message Queues and Worker Sizing
Decouple slow work.
Queues, Little's law and backoff
A message queue (Kafka, SQS, RabbitMQ) lets a service accept work quickly and process it asynchronously with workers, smoothing spikes and isolating failures. Most queues deliver at least once, so consumers must handle duplicates. Little's law (items in system = arrival rate times time in system) sizes worker pools. Failed calls should retry with exponential backoff and jitter to avoid synchronised retry storms, and give up into a dead-letter queue.
Decouple, retry safely, protect
Queues absorb spikes, idempotency makes retries safe, and rate limits protect shared resources.
Sizing workers and spreading retries, run
I ran this with Python 3.12.3 using only the standard library; inputs are fixed or seeded, so the output is reproducible. At 400 jobs per second taking 0.25 s each, 100 jobs are in service on average, so about 143 workers keep utilisation near 70%. Full jitter picks a random sleep up to the exponential limit.
# Sizing workers with Little's law, and retry backoff with jitter
import math, random
arrival_rate = 400 # jobs per second at peak
service_time = 0.25 # seconds per job per worker
utilisation_target = 0.7
busy = arrival_rate * service_time # L = lambda * W (jobs in service)
workers = math.ceil(busy / utilisation_target)
print(f"jobs in service on average: {busy:.0f}")
print(f"workers for 70% utilisation: {workers}")
rng = random.Random(3)
base, cap = 0.1, 10.0
for attempt in range(6):
exp = min(cap, base * 2 ** attempt)
print(f"attempt {attempt}: max {exp:5.1f}s full-jitter sleep {rng.uniform(0, exp):5.2f}s")
Output:
jobs in service on average: 100 workers for 70% utilisation: 143 attempt 0: max 0.1s full-jitter sleep 0.02s attempt 1: max 0.2s full-jitter sleep 0.11s attempt 2: max 0.4s full-jitter sleep 0.15s attempt 3: max 0.8s full-jitter sleep 0.48s attempt 4: max 1.6s full-jitter sleep 1.00s attempt 5: max 3.2s full-jitter sleep 0.21s
Monitor queue depth and age
A growing backlog or oldest-message age is the earliest sign that workers cannot keep up.
Quick check: Why add jitter to retry backoff?
- To reduce message size
- To make retries faster
- To guarantee exactly-once delivery
- To stop many clients retrying at the same moment
Answer
To stop many clients retrying at the same moment — Spread retries out in time.