# Failures, Retries and Keepalive to Upstreams — Nginx

Source: https://www.skillbyai.com/en/nginx/l-health

> Handle failing upstream servers and reuse connections efficiently.

## Passive health checks and safe retries

NGINX Open Source performs **passive health checks**: if a server fails **`max_fails`** times (default 1) within **`fail_timeout`** (default 10 s), it is considered unavailable for the next `fail_timeout` period, after which NGINX tries it again with real traffic. What counts as a failure is set by **`proxy_next_upstream`**: by default connection errors and timeouts, and you can add `http_502`, `http_503` and `http_504`. When a request fails on one server, NGINX can retry it on the next server, limited by **`proxy_next_upstream_tries`** and **`proxy_next_upstream_timeout`**. Be careful with **non-idempotent** requests: NGINX does not retry POST and other non-idempotent methods after a request has been sent unless you add `non_idempotent`, which risks duplicate orders or payments. **Active health checks**, which probe a health endpoint independently of traffic, are an NGINX Plus feature; with open source, rely on passive checks plus your orchestrator's health checks. **`keepalive N`** in the upstream keeps idle connections to backends open for reuse, which needs `proxy_http_version 1.1` and an empty `Connection` header.

## Tuned failure handling

Fail fast, retry once elsewhere for safe requests, and reuse connections.

```nginx
upstream api_backend {
    server 10.0.1.11:8080 max_fails=3 fail_timeout=15s;
    server 10.0.1.12:8080 max_fails=3 fail_timeout=15s;
    keepalive 64;
    keepalive_timeout 60s;
}

server {
    location /api/ {
        proxy_pass http://api_backend;
        proxy_http_version 1.1;
        proxy_set_header Connection "";
        proxy_connect_timeout 1s;
        proxy_read_timeout 10s;
        proxy_next_upstream error timeout http_502 http_503;
        proxy_next_upstream_tries 2;
        proxy_next_upstream_timeout 5s;
    }
}
```

## Skipping a closed counter

If a counter refuses three customers in a row, the queue manager stops sending people there for fifteen minutes, then sends one person to check. That is passive health checking.

**Quiz:** By default, will NGINX retry a POST request on another upstream after it was already sent and the upstream timed out?

- [ ] Yes, always
- [x] No, not unless non_idempotent is added to proxy_next_upstream
- [ ] Only on Sundays
- [ ] Only with ip_hash

*Answer:* No, not unless non_idempotent is added to proxy_next_upstream. NGINX avoids retrying non-idempotent requests to prevent duplicate side effects.
