# Caching as a Scaling Tool — Scalability, Availability & Reliability

Source: https://www.skillbyai.com/en/scalability/p-caching

> Use caches to cut load and latency, and avoid the classic cache failure modes.

## Not doing the work at all

The cheapest request is the one you never compute. Caches exist at every layer: the **browser** and **CDN** cache static assets and even API responses at the edge; a **reverse proxy** caches whole pages; an **application cache** (Redis, Memcached) stores computed objects and query results; **databases** cache pages in memory. The common **cache-aside** pattern reads from the cache, falls back to the database on a miss and populates the cache with a **TTL**. Caches raise throughput dramatically but add failure modes: **stale data** (decide what staleness is acceptable), **cache stampede** (many requests miss at once when a hot key expires and all hit the database), and **cold caches** after a restart or deploy that suddenly push full load to the database. Mitigations include request coalescing, early refresh, TTL jitter and warming. A separate course in this catalogue covers caching and CDN design in depth.

## Layers of cache between user and database

Each layer answers what it can and passes only misses inward.

![Concentric rings around a central database cylinder, with arrows from the outside that mostly stop at the outer rings and only a few reach the centre.](assets/figures/scalability/section-4-map.svg) — Figure 4.1 — Browser, CDN, application and database caches.

## Cache-aside with stampede protection

Only one caller rebuilds a missing value; TTL jitter spreads expiries.

```python
import json, random

def get_product(product_id):
    key = f"product:{product_id}"
    cached = redis.get(key)
    if cached:
        return json.loads(cached)

    # only one process rebuilds; others wait briefly or serve stale
    if redis.set(f"lock:{key}", "1", nx=True, ex=10):
        try:
            product = db.fetch_product(product_id)
            ttl = 300 + random.randint(0, 60)      # jitter avoids synchronised expiry
            redis.set(key, json.dumps(product), ex=ttl)
            return product
        finally:
            redis.delete(f"lock:{key}")
    return wait_and_retry_cache(key)
```

## Plan for the cache being empty

If the database can only survive because the cache absorbs 95% of reads, a cache flush becomes an outage. Know your database's capacity with a cold cache, and warm caches before shifting traffic.

**Quiz:** Many requests miss the cache at the same moment when a popular key expires and all hit the database. What is this called?

- [ ] Cache warming
- [ ] Write-through caching
- [x] Cache stampede (thundering herd)
- [ ] Replica lag

*Answer:* Cache stampede (thundering herd). Simultaneous misses on a hot key overload the backing store; coalescing and jitter help.
