SkillByAIOpen interactive version →

Lesson 14 / 25

Message Queues and Worker Sizing

Decouple slow work.

Queues, Little's law and backoff

A message queue (Kafka, SQS, RabbitMQ) lets a service accept work quickly and process it asynchronously with workers, smoothing spikes and isolating failures. Most queues deliver at least once, so consumers must handle duplicates. Little's law (items in system = arrival rate times time in system) sizes worker pools. Failed calls should retry with exponential backoff and jitter to avoid synchronised retry storms, and give up into a dead-letter queue.

Decouple, retry safely, protect

Queues absorb spikes, idempotency makes retries safe, and rate limits protect shared resources.

Figure 5.1 — Queues, idempotency and rate limiting.

Sizing workers and spreading retries, run

I ran this with Python 3.12.3 using only the standard library; inputs are fixed or seeded, so the output is reproducible. At 400 jobs per second taking 0.25 s each, 100 jobs are in service on average, so about 143 workers keep utilisation near 70%. Full jitter picks a random sleep up to the exponential limit.

# Sizing workers with Little's law, and retry backoff with jitter
import math, random

arrival_rate = 400        # jobs per second at peak
service_time = 0.25       # seconds per job per worker
utilisation_target = 0.7

busy = arrival_rate * service_time            # L = lambda * W (jobs in service)
workers = math.ceil(busy / utilisation_target)
print(f"jobs in service on average: {busy:.0f}")
print(f"workers for 70% utilisation: {workers}")

rng = random.Random(3)
base, cap = 0.1, 10.0
for attempt in range(6):
    exp = min(cap, base * 2 ** attempt)
    print(f"attempt {attempt}: max {exp:5.1f}s  full-jitter sleep {rng.uniform(0, exp):5.2f}s")

Output:

jobs in service on average: 100
workers for 70% utilisation: 143
attempt 0: max   0.1s  full-jitter sleep  0.02s
attempt 1: max   0.2s  full-jitter sleep  0.11s
attempt 2: max   0.4s  full-jitter sleep  0.15s
attempt 3: max   0.8s  full-jitter sleep  0.48s
attempt 4: max   1.6s  full-jitter sleep  1.00s
attempt 5: max   3.2s  full-jitter sleep  0.21s

Monitor queue depth and age

A growing backlog or oldest-message age is the earliest sign that workers cannot keep up.

Quick check: Why add jitter to retry backoff?

  • To reduce message size
  • To make retries faster
  • To guarantee exactly-once delivery
  • To stop many clients retrying at the same moment
Answer

To stop many clients retrying at the same moment — Spread retries out in time.