# Retry Amplification, Retry Budgets and Hedging — Rate Limiting, Circuit Breakers & Resilience Patterns

Source: https://www.skillbyai.com/en/resilience-patterns/t-budgets

> Prevent retry storms across layers and use hedged requests carefully.

## Retries multiply across layers

If a mobile app retries 3 times, the gateway retries 3 times and a service retries 3 times, one user action can turn into **27** calls to the deepest dependency, exactly when it is failing. This **retry amplification** turns brief blips into outages. Defences: retry at **one layer only**, usually the one closest to the failing dependency, and make other layers fail fast; use a **retry budget**, which limits retries to a percentage of normal traffic (for example, retries may add at most 10% extra load) so they stop automatically during widespread failure; and combine retries with **circuit breakers**, which stop retries once a dependency is clearly down. **Hedged requests** are a related tool for tail latency: if a response has not arrived by roughly the p95 time, send a second request to another replica and use whichever answers first. Hedge only idempotent reads, and cap the extra load.

## A simple retry budget

Retries are allowed only while they stay below 10% of recent requests.

```python
import time
from collections import deque

class RetryBudget:
    def __init__(self, ratio=0.1, window_s=10, min_per_sec=10):
        self.ratio, self.window, self.min_per_sec = ratio, window_s, min_per_sec
        self.requests, self.retries = deque(), deque()

    def _trim(self, q, now):
        while q and now - q[0] > self.window:
            q.popleft()

    def record_request(self):
        self.requests.append(time.monotonic())

    def can_retry(self):
        now = time.monotonic()
        self._trim(self.requests, now); self._trim(self.retries, now)
        allowed = self.ratio * len(self.requests) + self.min_per_sec * self.window
        if len(self.retries) < allowed:
            self.retries.append(now)
            return True
        return False
```

## Count total attempts in traces

Add the attempt number to spans and logs. During an incident, a graph of attempts per request quickly shows whether retries are amplifying load.

**Quiz:** Three layers each retry 3 times. How many calls can one user request cause at the deepest dependency in the worst case?

- [x] 27
- [ ] 3
- [ ] 9
- [ ] 81

*Answer:* 27. 3 × 3 × 3 = 27 attempts, which is why retries should happen at one layer with a budget.
