पाठ 16 / 25
Load Shedding
Reject excess work early and by priority so the server keeps serving what it accepts.
Say no early to keep saying yes
When demand exceeds capacity, a server cannot serve everyone. If it accepts everything, queues grow, every request gets slow, most time out, and goodput (useful completed work) collapses even though the server is busy. Load shedding deliberately rejects some requests early and cheaply (HTTP 503 or 429 with Retry-After) so the requests it does accept complete in time. Signals for shedding: in-flight request count above a limit, queue wait time above a threshold (a request that waited 2 seconds in a queue will probably time out anyway), CPU or memory pressure. Shed by priority: health checks and critical user actions (checkout, login) are kept; background jobs, prefetching and analytics are dropped first. Techniques from large-scale systems include CoDel-style queue management (drop requests that have waited too long) and adaptive LIFO under overload, serving the newest requests first because older ones are likely already abandoned.
Goodput with and without shedding
Without shedding, useful throughput collapses under overload; with shedding it stays near capacity.
Shedding by queue time and priority
Requests that already waited too long, or low-priority work under pressure, are rejected immediately.
import time
MAX_QUEUE_WAIT = 0.5 # seconds
MAX_IN_FLIGHT = 400
in_flight = 0
def admit(request):
waited = time.monotonic() - request.enqueued_at
if waited > MAX_QUEUE_WAIT:
return reject(503, retry_after=2, reason="queue timeout")
if in_flight > MAX_IN_FLIGHT * 0.8 and request.priority == "low":
return reject(503, retry_after=5, reason="shedding low priority")
if in_flight > MAX_IN_FLIGHT:
return reject(503, retry_after=1, reason="overloaded")
return None # acceptedRejections must be cheap
Shedding only helps if rejecting costs far less than serving. Check limits before parsing large bodies, authenticating with remote calls or touching the database.
त्वरित जाँच: Why can accepting every request under overload reduce the number of successful responses?
- Servers refuse to queue requests
- Rejected requests use more CPU than accepted ones
- Queues grow until most requests time out, so busy work produces few useful results
- Load balancers stop routing traffic
Answer
Queues grow until most requests time out, so busy work produces few useful results — Overloaded queues make almost every request too slow; shedding keeps accepted requests within their deadlines.