# Monitoring and Testing Caches — Caching Strategies & CDN Design

Source: https://www.skillbyai.com/en/caching-strategies/p-measure

> Monitor cache health and include caches in load tests and failure drills.

## Know how your caches behave

Treat caches as production systems with their own dashboards. For application caches, track **hit ratio** (overall and per key prefix), **latency** of cache calls, **evictions**, **memory use**, **connection counts**, **errors and timeouts**, and the **load on the origin** behind them. For Redis, `INFO stats` and `INFO memory` expose `keyspace_hits`, `keyspace_misses`, `evicted_keys` and memory figures, and the slow log shows expensive commands. For CDNs, track **edge hit ratio** (by requests and by bytes), origin requests and errors, and cache status by path. In **load tests**, run both **warm-cache** and **cold-cache** scenarios, because the cold case shows what happens after a deploy or cache failure. Run **failure drills**: restart a cache node, block the cache entirely, and confirm the application degrades gracefully (timeouts on cache calls, falling back to the origin with limits) rather than failing.

## A cache health dashboard

Hit ratio, latency, evictions and origin load viewed together tell the real story.

![Four small chart panels in a grid: a high flat line, a low flat line, a small bar chart and a line with a dip.](assets/figures/caching-strategies/section-8-map.svg) — Figure 8.1 — Key cache metrics side by side.

## Reading Redis cache statistics

Hit ratio from server counters; evictions suggest the cache is too small.

```bash
redis-cli INFO stats | grep -E 'keyspace_hits|keyspace_misses|evicted_keys|expired_keys'
redis-cli INFO memory | grep -E 'used_memory_human|maxmemory_human|mem_fragmentation_ratio'
redis-cli SLOWLOG GET 10            # recent slow commands
redis-cli --bigkeys                 # sample the keyspace for the largest keys

# hit ratio = keyspace_hits / (keyspace_hits + keyspace_misses)
```

## Cache calls need timeouts too

A slow or unreachable cache can hang every request if calls have no timeout. Give cache clients short timeouts (tens of milliseconds) and treat a cache error as a miss.

**Quiz:** Why should load tests include a cold-cache scenario?

- [ ] Cold caches are faster
- [ ] It removes the need for warm-cache tests
- [x] It shows how the origin copes after a deploy or cache failure when the hit ratio is near zero
- [ ] Caches cannot be tested warm

*Answer:* It shows how the origin copes after a deploy or cache failure when the hit ratio is near zero. Cold-cache tests reveal whether the origin survives when caching suddenly stops helping.
