# Load Shedding — Rate Limiting, Circuit Breakers & Resilience Patterns

Source: https://www.skillbyai.com/en/resilience-patterns/o-shedding

> Reject excess work early and by priority so the server keeps serving what it accepts.

## Say no early to keep saying yes

When demand exceeds capacity, a server cannot serve everyone. If it accepts everything, queues grow, every request gets slow, most time out, and **goodput** (useful completed work) collapses even though the server is busy. **Load shedding** deliberately rejects some requests **early and cheaply** (HTTP 503 or 429 with `Retry-After`) so the requests it does accept complete in time. Signals for shedding: **in-flight request count** above a limit, **queue wait time** above a threshold (a request that waited 2 seconds in a queue will probably time out anyway), CPU or memory pressure. Shed by **priority**: health checks and critical user actions (checkout, login) are kept; background jobs, prefetching and analytics are dropped first. Techniques from large-scale systems include **CoDel-style** queue management (drop requests that have waited too long) and **adaptive LIFO** under overload, serving the newest requests first because older ones are likely already abandoned.

## Goodput with and without shedding

Without shedding, useful throughput collapses under overload; with shedding it stays near capacity.

![Two curves rising together; beyond a capacity line one curve falls sharply while the other flattens out at a plateau.](assets/figures/resilience-patterns/section-6-map.svg) — Figure 6.1 — Load shedding preserves goodput beyond capacity.

## Shedding by queue time and priority

Requests that already waited too long, or low-priority work under pressure, are rejected immediately.

```python
import time

MAX_QUEUE_WAIT = 0.5          # seconds
MAX_IN_FLIGHT = 400
in_flight = 0

def admit(request):
    waited = time.monotonic() - request.enqueued_at
    if waited > MAX_QUEUE_WAIT:
        return reject(503, retry_after=2, reason="queue timeout")
    if in_flight > MAX_IN_FLIGHT * 0.8 and request.priority == "low":
        return reject(503, retry_after=5, reason="shedding low priority")
    if in_flight > MAX_IN_FLIGHT:
        return reject(503, retry_after=1, reason="overloaded")
    return None   # accepted
```

## Rejections must be cheap

Shedding only helps if rejecting costs far less than serving. Check limits before parsing large bodies, authenticating with remote calls or touching the database.

**Quiz:** Why can accepting every request under overload reduce the number of successful responses?

- [ ] Servers refuse to queue requests
- [ ] Rejected requests use more CPU than accepted ones
- [x] Queues grow until most requests time out, so busy work produces few useful results
- [ ] Load balancers stop routing traffic

*Answer:* Queues grow until most requests time out, so busy work produces few useful results. Overloaded queues make almost every request too slow; shedding keeps accepted requests within their deadlines.
