# Message Queues and Worker Sizing — System Design Interview Prep

Source: https://www.skillbyai.com/en/system-design-interview/r-queue

> Decouple slow work.

## Queues, Little's law and backoff

A **message queue** (Kafka, SQS, RabbitMQ) lets a service accept work quickly and process it asynchronously with workers, smoothing spikes and isolating failures. Most queues deliver **at least once**, so consumers must handle duplicates. **Little's law** (items in system = arrival rate times time in system) sizes worker pools. Failed calls should retry with **exponential backoff and jitter** to avoid synchronised retry storms, and give up into a dead-letter queue.

## Decouple, retry safely, protect

Queues absorb spikes, idempotency makes retries safe, and rate limits protect shared resources.

![Three ideas: queues and workers, idempotency, rate limiting.](assets/figures/system-design-interview/section-5-map.svg) — Figure 5.1 — Queues, idempotency and rate limiting.

## Sizing workers and spreading retries, run

I ran this with Python 3.12.3 using only the standard library; inputs are fixed or seeded, so the output is reproducible. At 400 jobs per second taking 0.25 s each, 100 jobs are in service on average, so about 143 workers keep utilisation near 70%. Full jitter picks a random sleep up to the exponential limit.

```python
# Sizing workers with Little's law, and retry backoff with jitter
import math, random

arrival_rate = 400        # jobs per second at peak
service_time = 0.25       # seconds per job per worker
utilisation_target = 0.7

busy = arrival_rate * service_time            # L = lambda * W (jobs in service)
workers = math.ceil(busy / utilisation_target)
print(f"jobs in service on average: {busy:.0f}")
print(f"workers for 70% utilisation: {workers}")

rng = random.Random(3)
base, cap = 0.1, 10.0
for attempt in range(6):
    exp = min(cap, base * 2 ** attempt)
    print(f"attempt {attempt}: max {exp:5.1f}s  full-jitter sleep {rng.uniform(0, exp):5.2f}s")
```

Output:

```
jobs in service on average: 100
workers for 70% utilisation: 143
attempt 0: max   0.1s  full-jitter sleep  0.02s
attempt 1: max   0.2s  full-jitter sleep  0.11s
attempt 2: max   0.4s  full-jitter sleep  0.15s
attempt 3: max   0.8s  full-jitter sleep  0.48s
attempt 4: max   1.6s  full-jitter sleep  1.00s
attempt 5: max   3.2s  full-jitter sleep  0.21s
```

## Monitor queue depth and age

A growing backlog or oldest-message age is the earliest sign that workers cannot keep up.

**Quiz:** Why add jitter to retry backoff?

- [ ] To reduce message size
- [ ] To make retries faster
- [ ] To guarantee exactly-once delivery
- [x] To stop many clients retrying at the same moment

*Answer:* To stop many clients retrying at the same moment. Spread retries out in time.
