पाठ 10 / 25

Caching as a Scaling Tool

Use caches to cut load and latency, and avoid the classic cache failure modes.

Not doing the work at all

The cheapest request is the one you never compute. Caches exist at every layer: the browser and CDN cache static assets and even API responses at the edge; a reverse proxy caches whole pages; an application cache (Redis, Memcached) stores computed objects and query results; databases cache pages in memory. The common cache-aside pattern reads from the cache, falls back to the database on a miss and populates the cache with a TTL. Caches raise throughput dramatically but add failure modes: stale data (decide what staleness is acceptable), cache stampede (many requests miss at once when a hot key expires and all hit the database), and cold caches after a restart or deploy that suddenly push full load to the database. Mitigations include request coalescing, early refresh, TTL jitter and warming. A separate course in this catalogue covers caching and CDN design in depth.

Layers of cache between user and database

Each layer answers what it can and passes only misses inward.

Concentric rings around a central database cylinder, with arrows from the outside that mostly stop at the outer rings and only a few reach the centre.
Figure 4.1 — Browser, CDN, application and database caches.

Cache-aside with stampede protection

Only one caller rebuilds a missing value; TTL jitter spreads expiries.

import json, random

def get_product(product_id):
    key = f"product:{product_id}"
    cached = redis.get(key)
    if cached:
        return json.loads(cached)

    # only one process rebuilds; others wait briefly or serve stale
    if redis.set(f"lock:{key}", "1", nx=True, ex=10):
        try:
            product = db.fetch_product(product_id)
            ttl = 300 + random.randint(0, 60)      # jitter avoids synchronised expiry
            redis.set(key, json.dumps(product), ex=ttl)
            return product
        finally:
            redis.delete(f"lock:{key}")
    return wait_and_retry_cache(key)

Plan for the cache being empty

If the database can only survive because the cache absorbs 95% of reads, a cache flush becomes an outage. Know your database's capacity with a cold cache, and warm caches before shifting traffic.

त्वरित जाँच: Many requests miss the cache at the same moment when a popular key expires and all hit the database. What is this called?

  • Cache warming
  • Write-through caching
  • Cache stampede (thundering herd)
  • Replica lag
Answer

Cache stampede (thundering herd) — Simultaneous misses on a hot key overload the backing store; coalescing and jitter help.