# Hit Ratio, Effective Latency and the Working Set — Caching Strategies & CDN Design

Source: https://www.skillbyai.com/en/caching-strategies/f-metrics

> Quantify a cache with hit ratio and effective latency, and size it from the working set.

## The numbers that describe a cache

The **hit ratio** is hits divided by total lookups. Small changes matter a lot: the **effective latency** is `hit_ratio × hit_time + (1 − hit_ratio) × miss_time`, and the **load on the origin** is proportional to the **miss ratio**. Going from 90% to 99% hits cuts origin traffic **tenfold**. Real access patterns are usually skewed (often close to a **Zipf** distribution): a small set of popular items receives most requests, so a cache holding the **working set**, the items accessed in a typical period, achieves a high hit ratio even if it holds a small fraction of all data. Measure hit ratio per key prefix or endpoint, not just globally, because a great global number can hide a feature with a 20% hit ratio. Also watch **eviction rate** (memory too small), **memory use** and **latency of the cache itself**.

## Hit ratio maths

Raising the hit ratio from 90% to 99% cuts database queries by ten times.

```python
def effective_latency(hit_ratio, hit_ms, miss_ms):
    return hit_ratio * hit_ms + (1 - hit_ratio) * miss_ms

rps = 20_000
for h in (0.80, 0.90, 0.99):
    db_qps = rps * (1 - h)
    print(f"hit {h:.0%}: {effective_latency(h, 0.5, 20):.2f} ms avg, {db_qps:,.0f} DB queries/s")

# by hand: 80% -> 0.8*0.5 + 0.2*20 = 4.4 ms and 4,000 DB queries/s
#          90% -> 2.45 ms and 2,000/s;  99% -> about 0.7 ms and 200/s
```

## Plan capacity for the miss rate

Your database must handle the miss traffic at peak, plus a margin for when the hit ratio drops after a deploy or cache restart. If it only survives at 99% hits, a dip to 90% is a tenfold load spike.

**Quiz:** A cache serving 10,000 requests per second improves from a 95% to a 99% hit ratio. What happens to origin traffic?

- [ ] It halves
- [ ] It stays the same
- [x] It drops from 500 to 100 requests per second
- [ ] It increases

*Answer:* It drops from 500 to 100 requests per second. Misses fall from 5% to 1% of 10,000, so origin load drops from 500 to 100 per second.
