# Distributed Rate Limiting — Rate Limiting, Circuit Breakers & Resilience Patterns

Source: https://www.skillbyai.com/en/resilience-patterns/r-distributed

> Enforce limits across many instances with Redis or a gateway.

## One limit, many servers

With ten API instances behind a load balancer, a limit counted in each instance's memory allows ten times the intended rate. **Distributed rate limiting** keeps counters in a shared, fast store, typically **Redis**, and updates them **atomically**, usually with a Lua script or atomic commands so concurrent requests cannot both pass on the last token. Trade-offs: every request adds a network round trip to the limiter; the limiter becomes a dependency, so decide whether to **fail open** (allow requests if Redis is unavailable, favouring availability) or **fail closed** (reject, favouring protection); and very hot keys can overload one Redis shard. Alternatives: let the **API gateway** or **service mesh** enforce limits (Kong, Envoy's global rate limit service, cloud API gateways), or use **approximate local limits** per instance (global limit divided by instance count), which is cheap and often good enough when traffic is balanced.

## Atomic token bucket in Redis with Lua

The script refills and consumes in one atomic step on the Redis server.

```lua
-- KEYS[1] = bucket key, ARGV = rate_per_sec, capacity, now_ms, cost
local rate     = tonumber(ARGV[1])
local capacity = tonumber(ARGV[2])
local now      = tonumber(ARGV[3])
local cost     = tonumber(ARGV[4])

local state  = redis.call('HMGET', KEYS[1], 'tokens', 'ts')
local tokens = tonumber(state[1]) or capacity
local ts     = tonumber(state[2]) or now

tokens = math.min(capacity, tokens + (now - ts) / 1000 * rate)
local allowed = 0
if tokens >= cost then
  tokens = tokens - cost
  allowed = 1
end

redis.call('HSET', KEYS[1], 'tokens', tokens, 'ts', now)
redis.call('PEXPIRE', KEYS[1], math.ceil(capacity / rate * 1000) + 1000)
return { allowed, math.floor(tokens) }
```

## Decide fail-open or fail-closed explicitly

For a public API's fairness limits, failing open during a Redis outage is usually right. For login brute-force protection, failing closed (or falling back to strict local limits) may be safer. Write the decision down.

**Quiz:** Why must a distributed rate limiter update its counter atomically?

- [ ] Redis requires it for all keys
- [ ] To make responses smaller
- [x] Otherwise concurrent requests on different instances could all read the same count and exceed the limit
- [ ] To avoid using TTLs

*Answer:* Otherwise concurrent requests on different instances could all read the same count and exceed the limit. Atomic read-modify-write prevents race conditions between instances.
