SkillByAIOpen interactive version →

Lesson 3 / 25

Hit Ratio, Effective Latency and the Working Set

Quantify a cache with hit ratio and effective latency, and size it from the working set.

The numbers that describe a cache

The hit ratio is hits divided by total lookups. Small changes matter a lot: the effective latency is hit_ratio × hit_time + (1 − hit_ratio) × miss_time, and the load on the origin is proportional to the miss ratio. Going from 90% to 99% hits cuts origin traffic tenfold. Real access patterns are usually skewed (often close to a Zipf distribution): a small set of popular items receives most requests, so a cache holding the working set, the items accessed in a typical period, achieves a high hit ratio even if it holds a small fraction of all data. Measure hit ratio per key prefix or endpoint, not just globally, because a great global number can hide a feature with a 20% hit ratio. Also watch eviction rate (memory too small), memory use and latency of the cache itself.

Hit ratio maths

Raising the hit ratio from 90% to 99% cuts database queries by ten times.

def effective_latency(hit_ratio, hit_ms, miss_ms):
    return hit_ratio * hit_ms + (1 - hit_ratio) * miss_ms

rps = 20_000
for h in (0.80, 0.90, 0.99):
    db_qps = rps * (1 - h)
    print(f"hit {h:.0%}: {effective_latency(h, 0.5, 20):.2f} ms avg, {db_qps:,.0f} DB queries/s")

# by hand: 80% -> 0.8*0.5 + 0.2*20 = 4.4 ms and 4,000 DB queries/s
#          90% -> 2.45 ms and 2,000/s;  99% -> about 0.7 ms and 200/s

Plan capacity for the miss rate

Your database must handle the miss traffic at peak, plus a margin for when the hit ratio drops after a deploy or cache restart. If it only survives at 99% hits, a dip to 90% is a tenfold load spike.

Quick check: A cache serving 10,000 requests per second improves from a 95% to a 99% hit ratio. What happens to origin traffic?

  • It halves
  • It stays the same
  • It drops from 500 to 100 requests per second
  • It increases
Answer

It drops from 500 to 100 requests per second — Misses fall from 5% to 1% of 10,000, so origin load drops from 500 to 100 per second.