पाठ 5 / 25

Load Balancing

Pick load-balancing layers and algorithms, and configure health checks correctly.

Spreading requests across instances

A load balancer distributes requests across healthy instances. Layer 4 balancers route TCP/UDP connections by IP and port: fast and protocol-agnostic. Layer 7 balancers understand HTTP: they can route by host or path, terminate TLS, add headers and retry. Common algorithms: round robin (simple, assumes equal instances and requests), weighted round robin, least connections or least outstanding requests (better when request costs vary), random with two choices (pick two instances, use the less loaded; nearly as good as least-loaded and cheap), and consistent hashing on a key (useful when the same user should reach the same cache). Health checks remove failing instances: a readiness check says "send me traffic", a liveness check says "restart me". Load balancers themselves must be redundant, and DNS or anycast spreads traffic across load balancers and regions.

An NGINX upstream with least connections and passive health checks

Instances that fail repeatedly are taken out of rotation for a short time.

upstream api_backend {
    least_conn;
    server 10.0.1.11:8080 max_fails=3 fail_timeout=10s;
    server 10.0.1.12:8080 max_fails=3 fail_timeout=10s;
    server 10.0.1.13:8080 max_fails=3 fail_timeout=10s;
    keepalive 64;
}

server {
    listen 443 ssl;
    location /api/ {
        proxy_pass http://api_backend;
        proxy_next_upstream error timeout http_502 http_503;
        proxy_next_upstream_tries 2;
        proxy_connect_timeout 1s;
        proxy_read_timeout 5s;
    }
}

Billing counters at a supermarket

Round robin sends each shopper to the next counter in turn, even if that counter has a customer with three full trolleys. Least connections sends you to the shortest queue, which is what a sensible shopper does.

त्वरित जाँच: Request costs vary widely: some take 5 ms, others 2 seconds. Which algorithm usually balances better than round robin?

  • Round robin
  • Least outstanding requests (or least connections)
  • Alphabetical by hostname
  • Always the first server
Answer

Least outstanding requests (or least connections) — Load-aware algorithms avoid piling requests onto an instance busy with slow work.