SkillByAIOpen interactive version →

Lesson 11 / 25

Failures, Retries and Keepalive to Upstreams

Handle failing upstream servers and reuse connections efficiently.

Passive health checks and safe retries

NGINX Open Source performs passive health checks: if a server fails max_fails times (default 1) within fail_timeout (default 10 s), it is considered unavailable for the next fail_timeout period, after which NGINX tries it again with real traffic. What counts as a failure is set by proxy_next_upstream: by default connection errors and timeouts, and you can add http_502, http_503 and http_504. When a request fails on one server, NGINX can retry it on the next server, limited by proxy_next_upstream_tries and proxy_next_upstream_timeout. Be careful with non-idempotent requests: NGINX does not retry POST and other non-idempotent methods after a request has been sent unless you add non_idempotent, which risks duplicate orders or payments. Active health checks, which probe a health endpoint independently of traffic, are an NGINX Plus feature; with open source, rely on passive checks plus your orchestrator's health checks. keepalive N in the upstream keeps idle connections to backends open for reuse, which needs proxy_http_version 1.1 and an empty Connection header.

Tuned failure handling

Fail fast, retry once elsewhere for safe requests, and reuse connections.

upstream api_backend {
    server 10.0.1.11:8080 max_fails=3 fail_timeout=15s;
    server 10.0.1.12:8080 max_fails=3 fail_timeout=15s;
    keepalive 64;
    keepalive_timeout 60s;
}

server {
    location /api/ {
        proxy_pass http://api_backend;
        proxy_http_version 1.1;
        proxy_set_header Connection "";
        proxy_connect_timeout 1s;
        proxy_read_timeout 10s;
        proxy_next_upstream error timeout http_502 http_503;
        proxy_next_upstream_tries 2;
        proxy_next_upstream_timeout 5s;
    }
}

Skipping a closed counter

If a counter refuses three customers in a row, the queue manager stops sending people there for fifteen minutes, then sends one person to check. That is passive health checking.

Quick check: By default, will NGINX retry a POST request on another upstream after it was already sent and the upstream timed out?

  • Yes, always
  • No, not unless non_idempotent is added to proxy_next_upstream
  • Only on Sundays
  • Only with ip_hash
Answer

No, not unless non_idempotent is added to proxy_next_upstream — NGINX avoids retrying non-idempotent requests to prevent duplicate side effects.