# Why and Where to Rate Limit — Rate Limiting, Circuit Breakers & Resilience Patterns

Source: https://www.skillbyai.com/en/resilience-patterns/r-why

> Explain the goals of rate limiting and choose keys and placement.

## Fair share for everyone

A **rate limiter** caps how many requests a client may make in a period. Goals: **protect capacity** so a burst from one client cannot degrade service for everyone; **fairness** between tenants or users; **cost control** for expensive operations or paid third-party APIs; **abuse prevention** against scraping, credential stuffing and brute-force login attempts; and enforcing **commercial plans** (free tier 100 requests per minute, paid tier 1,000). Choose the **key** that identifies a client: API key or tenant ID for authenticated APIs, user ID for per-user limits, IP address for anonymous traffic (careful with shared IPs such as offices and mobile carriers). Place limits at several layers: at the **edge** (CDN or WAF) for abusive traffic, at the **API gateway** for per-client quotas, and inside **services** for expensive operations. When a client exceeds its limit, respond with **HTTP 429 Too Many Requests** and a **`Retry-After`** header.

## Limits at several layers

Coarse limits at the edge, per-client quotas at the gateway, fine-grained limits in services.

![Three nested gates from left to right, each narrower than the previous, with many arrows entering the first and fewer passing each gate.](assets/figures/resilience-patterns/section-5-map.svg) — Figure 5.1 — Rate limits at the edge, gateway and service.

## A rate-limited response

Clients learn the limit, what is left and when to retry.

```http
HTTP/1.1 429 Too Many Requests
Content-Type: application/problem+json
Retry-After: 12
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1790000012

{
  "type": "https://api.example.com/problems/rate-limited",
  "title": "Too many requests",
  "detail": "Limit of 100 requests per minute exceeded for API key ak_live_...9f"
}

# an IETF draft standardises RateLimit-Policy / RateLimit fields;
# many APIs still use the X-RateLimit-* names shown here
```

## Limit login attempts by account and by IP

Per-IP limits alone are bypassed by botnets; per-account limits alone let one attacker lock out many users. Combine them, and add progressive delays or CAPTCHA for repeated failures.

**Quiz:** Which status code and header should a rate-limited API return?

- [x] 429 Too Many Requests with Retry-After
- [ ] 500 with Location
- [ ] 404 with ETag
- [ ] 301 with Cache-Control

*Answer:* 429 Too Many Requests with Retry-After. 429 signals rate limiting, and Retry-After tells clients when to try again.
