Lesson 4 / 25
Choosing Timeouts
Set connect and request timeouts from measured latency.
Every network call needs a timeout
Many client libraries default to no timeout or a very long one (minutes). That is the single most common resilience bug. Set at least two: a connect timeout (how long to wait to establish a TCP or TLS connection; usually short, such as 100 ms to 1 s inside a datacentre) and a request or read timeout (how long to wait for the response). Some libraries also offer a total timeout covering retries. Choose request timeouts from measured latency: a common starting point is a little above the dependency's p99 or p99.9 under normal load, so you only cut off genuinely abnormal requests. Too short, and you fail requests that would have succeeded, causing retries and extra load; too long, and resources are held during an incident. Review timeouts when latency profiles change, and remember the caller's own deadline caps whatever you choose.
Timeout from the latency distribution
Set the timeout just beyond the normal tail, so only abnormal requests are cut off.
Explicit timeouts in common clients
Never rely on library defaults.
import httpx
# Python httpx: separate connect / read / write / pool timeouts
client = httpx.Client(timeout=httpx.Timeout(connect=0.5, read=1.0, write=1.0, pool=0.5))
# Java HttpClient:
# HttpClient.newBuilder().connectTimeout(Duration.ofMillis(500)).build();
# HttpRequest.newBuilder(uri).timeout(Duration.ofSeconds(1)).build();
# Node.js fetch:
# await fetch(url, { signal: AbortSignal.timeout(1000) });
# database drivers: set statement / query timeouts too, e.g. PostgreSQL
# SET statement_timeout = '2s';Time out the connection pool wait too
When a pool is exhausted, callers queue for a connection. Without a pool acquisition timeout, they can wait far longer than the request timeout, invisibly.
Quick check: What is a sensible starting point for a request timeout to a dependency?
- The mean latency
- One hour
- Slightly above the dependency's normal p99 (or p99.9) latency, within the caller's deadline
- No timeout, to avoid failing requests
Answer
Slightly above the dependency's normal p99 (or p99.9) latency, within the caller's deadline — Basing timeouts on the tail of normal latency cuts off only abnormal requests.