# Tail Latency and Fan-Out — Scalability, Availability & Reliability

Source: https://www.skillbyai.com/en/scalability/p-tail

> Measure percentiles and reduce tail latency in fan-out architectures.

## The slowest requests matter most

Averages hide pain. Report latency as **percentiles**: p50 (median), p95, **p99** and p99.9. At scale, tail latency dominates user experience because one page often calls many services: if a request fans out to 100 backends and each has a 1% chance of being slow, about 63% of requests (1 − 0.99^100) hit at least one slow backend. Tail latency comes from garbage-collection pauses, noisy neighbours, cold caches, retries and queueing. Techniques to reduce it: keep fan-out small; set tight **timeouts** with partial results (show the page without recommendations); use **hedged requests**, sending a second copy to another replica if the first has not answered by the p95 time, and using whichever replies first; reduce variability through capacity headroom and fewer synchronous dependencies. Measure percentiles from histograms; averaging percentiles across servers gives wrong numbers.

## Fan-out amplifies the tail

Probability that a page waits on at least one slow call.

```python
for fan_out in (1, 10, 50, 100):
    p_slow_page = 1 - 0.99 ** fan_out      # each call slow with probability 1%
    print(fan_out, f"{p_slow_page:.0%}")

# hedged request: after the p95 deadline, ask a second replica
async def hedged_get(key, replicas, hedge_after=0.05):
    first = asyncio.create_task(replicas[0].get(key))
    done, _ = await asyncio.wait({first}, timeout=hedge_after)
    if done:
        return first.result()
    second = asyncio.create_task(replicas[1].get(key))
    done, pending = await asyncio.wait({first, second}, return_when=asyncio.FIRST_COMPLETED)
    for t in pending:
        t.cancel()
    return done.pop().result()
```

## Hedge only idempotent reads

A hedged request sends duplicate work. Use it for safe reads, cap the extra load (for example only after the p95 deadline), and never for operations like payments.

**Quiz:** Why should latency SLOs use percentiles such as p99 rather than averages?

- [x] Averages hide the slow requests that many users experience, especially with fan-out
- [ ] Averages are harder to compute
- [ ] p99 is always lower than the average
- [ ] Percentiles ignore errors

*Answer:* Averages hide the slow requests that many users experience, especially with fan-out. A good average can coexist with a painful tail that users regularly hit.
