पाठ 17 / 25
Concurrency Limits and Back-Pressure
Use adaptive concurrency limits and back-pressure to keep queues short.
Limit work in flight, adaptively
A concurrency limit caps how many requests a server (or a client towards a dependency) processes at once; anything beyond is rejected or briefly queued. Fixed limits are hard to tune because the right number changes with hardware, request mix and dependency latency. Adaptive concurrency limits, inspired by TCP congestion control, adjust the limit automatically: when latency stays near its minimum, raise the limit; when latency rises (a sign of queueing), lower it. Netflix's open-source concurrency-limits library implements algorithms such as Vegas and Gradient, and Envoy offers an adaptive concurrency filter. Back-pressure is the general principle of letting overload signals flow upstream: bounded queues that block or reject, HTTP 429/503 with Retry-After, flow control in gRPC and HTTP/2, and reactive streams where consumers request only as many items as they can handle. Without back-pressure, overload just moves into ever-growing queues and memory.
A simple adaptive concurrency limit (gradient idea)
The limit shrinks when latency rises above the best observed latency and grows when it does not.
class AdaptiveLimit:
def __init__(self, initial=50, min_limit=5, max_limit=1000):
self.limit, self.min, self.max = initial, min_limit, max_limit
self.best_rtt = None
def on_sample(self, rtt_s, in_flight):
self.best_rtt = rtt_s if self.best_rtt is None else min(self.best_rtt, rtt_s)
gradient = max(0.5, min(1.0, self.best_rtt / rtt_s)) # 1.0 = no queueing
headroom = in_flight ** 0.5 # allow growth when healthy
new_limit = self.limit * gradient + headroom
self.limit = int(max(self.min, min(self.max, 0.8 * self.limit + 0.2 * new_limit)))
# requests beyond self.limit in flight are rejected with 503 / Retry-AfterA metro station gate
At rush hour, staff close some gates so the platform does not overflow. Passengers wait in the concourse (back-pressure) instead of being crushed on the platform, and the gates open again as trains clear the crowd.
त्वरित जाँच: What signal do adaptive concurrency limiters mainly use to reduce the limit?
- Number of deploys
- Rising latency, which indicates queueing
- Disk usage on the client
- HTTP 200 responses
Answer
Rising latency, which indicates queueing — Latency above the baseline indicates requests are queueing, so the limit is lowered.