पाठ 7 / 25

Why Invalidation Is Hard

Explain the sources of stale data and the trade-off between TTLs and explicit invalidation.

Two hard things

Phil Karlton's famous quip says there are only two hard things in computer science: cache invalidation and naming things. Invalidation is hard because the cache and the source are separate systems without a shared transaction, so updates can happen in different orders, partially fail or be missed. Data reaches caches through many paths (several services, admin tools, batch jobs, direct SQL fixes) and any one of them can forget to invalidate. One record may also appear in many cached objects: a product name lives in the product page, search results, category listings and the cart. There are two basic approaches. Time-based expiry (TTL) is simple and self-healing, but data is stale for up to the TTL. Explicit invalidation on change keeps data fresh but must catch every write path and every derived key. Most systems use both: explicit invalidation for freshness, TTL as a safety net.

One change, many cached copies

A single source record feeds several cached objects, all of which may need invalidating.

A central record icon with arrows to five different cached tiles, some marked fresh and some faded as stale.
Figure 3.1 — A change must reach every cached copy derived from it.

Mapping a change to the keys it affects

List derived keys explicitly so none are forgotten.

def keys_affected_by_product_change(product):
    return [
        f"product:v2:{product.id}",
        f"category-page:v1:{product.category_id}:*",     # pattern, see tags below
        f"search:v1:brand:{product.brand_id}",
        f"home:v1:featured" if product.featured else None,
    ]

# better: tag cached entries with what they depend on, then invalidate by tag
# product page -> tags {product:42, category:7}
# category page -> tags {category:7}
# invalidate(tag="product:42") clears every entry built from product 42

Every write path must invalidate

A bulk price update run directly in SQL bypasses application code and its invalidation logic. Either route all writes through code that invalidates, or invalidate from the database's change stream so no path is missed.

त्वरित जाँच: Why do most systems combine explicit invalidation with a TTL?

  • The TTL acts as a safety net that eventually fixes entries whose invalidation was missed
  • TTLs make caches faster
  • Explicit invalidation cannot work without a TTL
  • Redis requires both
Answer

The TTL acts as a safety net that eventually fixes entries whose invalidation was missed — Explicit invalidation keeps data fresh; the TTL bounds how long a missed invalidation can cause staleness.