# Concurrency Limits and Back-Pressure — Rate Limiting, Circuit Breakers & Resilience Patterns

Source: https://www.skillbyai.com/en/resilience-patterns/o-concurrency

> Use adaptive concurrency limits and back-pressure to keep queues short.

## Limit work in flight, adaptively

A **concurrency limit** caps how many requests a server (or a client towards a dependency) processes at once; anything beyond is rejected or briefly queued. Fixed limits are hard to tune because the right number changes with hardware, request mix and dependency latency. **Adaptive concurrency limits**, inspired by TCP congestion control, adjust the limit automatically: when latency stays near its minimum, raise the limit; when latency rises (a sign of queueing), lower it. Netflix's open-source **concurrency-limits** library implements algorithms such as Vegas and Gradient, and Envoy offers an adaptive concurrency filter. **Back-pressure** is the general principle of letting overload signals flow **upstream**: bounded queues that block or reject, HTTP 429/503 with `Retry-After`, flow control in gRPC and HTTP/2, and reactive streams where consumers request only as many items as they can handle. Without back-pressure, overload just moves into ever-growing queues and memory.

## A simple adaptive concurrency limit (gradient idea)

The limit shrinks when latency rises above the best observed latency and grows when it does not.

```python
class AdaptiveLimit:
    def __init__(self, initial=50, min_limit=5, max_limit=1000):
        self.limit, self.min, self.max = initial, min_limit, max_limit
        self.best_rtt = None

    def on_sample(self, rtt_s, in_flight):
        self.best_rtt = rtt_s if self.best_rtt is None else min(self.best_rtt, rtt_s)
        gradient = max(0.5, min(1.0, self.best_rtt / rtt_s))   # 1.0 = no queueing
        headroom = in_flight ** 0.5                              # allow growth when healthy
        new_limit = self.limit * gradient + headroom
        self.limit = int(max(self.min, min(self.max, 0.8 * self.limit + 0.2 * new_limit)))

# requests beyond self.limit in flight are rejected with 503 / Retry-After
```

## A metro station gate

At rush hour, staff close some gates so the platform does not overflow. Passengers wait in the concourse (back-pressure) instead of being crushed on the platform, and the gates open again as trains clear the crowd.

**Quiz:** What signal do adaptive concurrency limiters mainly use to reduce the limit?

- [ ] Number of deploys
- [x] Rising latency, which indicates queueing
- [ ] Disk usage on the client
- [ ] HTTP 200 responses

*Answer:* Rising latency, which indicates queueing. Latency above the baseline indicates requests are queueing, so the limit is lowered.
