Lesson 5 / 25

Retries, Backoff and Jitter

Retry only transient, safe failures with exponential backoff and jitter.

Retry the right things, the right way

Many failures are transient: a connection reset, a timeout during a brief network blip, an HTTP 503 during a deploy. Retrying them often succeeds. But retry only when it is safe and useful. Safe: the operation is idempotent (GET, PUT, DELETE by design, or a POST with an idempotency key), so doing it twice has the same effect as once. Useful: the error is transient (timeouts, 502/503/504, connection errors, 429 after waiting), not permanent (400 Bad Request, 401, 403, 404, validation errors). Space retries with exponential backoff (for example 100 ms, 200 ms, 400 ms) so a struggling service gets breathing room, add jitter (randomness) so thousands of clients do not retry in synchronised waves, cap the number of attempts (two or three is typical) and the maximum delay, and honour Retry-After headers on 429 and 503 responses.

Retry policy with full jitter

Only transient errors on idempotent requests are retried.

import random, time

RETRYABLE_STATUS = {502, 503, 504}

def with_retries(send, attempts=3, base=0.1, cap=2.0, idempotent=True):
    for attempt in range(attempts):
        try:
            resp = send()
        except (TimeoutError, ConnectionError):
            resp = None
        if resp is not None and resp.status not in RETRYABLE_STATUS and resp.status != 429:
            return resp                               # success or a permanent error
        if not idempotent or attempt == attempts - 1:
            if resp is None:
                raise TimeoutError("dependency unavailable")
            return resp
        retry_after = resp.headers.get("Retry-After") if resp is not None else None
        delay = float(retry_after) if retry_after else random.uniform(0, min(cap, base * 2 ** attempt))
        time.sleep(delay)

Calling a busy helpline

If the line is busy, you do not redial every second, and you would not want every caller in the city to redial at exactly the same moment. You wait a little longer each time, at a slightly random moment.

Quick check: Which failure should normally NOT be retried?

  • A connection reset
  • HTTP 400 Bad Request caused by invalid input
  • HTTP 503 Service Unavailable
  • A timeout on an idempotent GET
Answer

HTTP 400 Bad Request caused by invalid input — A 400 is a permanent error; repeating the same request will fail the same way.