# Load Balancing — Scalability, Availability & Reliability

Source: https://www.skillbyai.com/en/scalability/s-lb

> Pick load-balancing layers and algorithms, and configure health checks correctly.

## Spreading requests across instances

A **load balancer** distributes requests across healthy instances. **Layer 4** balancers route TCP/UDP connections by IP and port: fast and protocol-agnostic. **Layer 7** balancers understand HTTP: they can route by host or path, terminate TLS, add headers and retry. Common algorithms: **round robin** (simple, assumes equal instances and requests), **weighted round robin**, **least connections** or **least outstanding requests** (better when request costs vary), **random with two choices** (pick two instances, use the less loaded; nearly as good as least-loaded and cheap), and **consistent hashing** on a key (useful when the same user should reach the same cache). **Health checks** remove failing instances: a **readiness** check says "send me traffic", a **liveness** check says "restart me". Load balancers themselves must be redundant, and DNS or anycast spreads traffic across load balancers and regions.

## An NGINX upstream with least connections and passive health checks

Instances that fail repeatedly are taken out of rotation for a short time.

```nginx
upstream api_backend {
    least_conn;
    server 10.0.1.11:8080 max_fails=3 fail_timeout=10s;
    server 10.0.1.12:8080 max_fails=3 fail_timeout=10s;
    server 10.0.1.13:8080 max_fails=3 fail_timeout=10s;
    keepalive 64;
}

server {
    listen 443 ssl;
    location /api/ {
        proxy_pass http://api_backend;
        proxy_next_upstream error timeout http_502 http_503;
        proxy_next_upstream_tries 2;
        proxy_connect_timeout 1s;
        proxy_read_timeout 5s;
    }
}
```

## Billing counters at a supermarket

Round robin sends each shopper to the next counter in turn, even if that counter has a customer with three full trolleys. Least connections sends you to the shortest queue, which is what a sensible shopper does.

**Quiz:** Request costs vary widely: some take 5 ms, others 2 seconds. Which algorithm usually balances better than round robin?

- [ ] Round robin
- [x] Least outstanding requests (or least connections)
- [ ] Alphabetical by hostname
- [ ] Always the first server

*Answer:* Least outstanding requests (or least connections). Load-aware algorithms avoid piling requests onto an instance busy with slow work.
