SkillByAIOpen interactive version →

Lesson 15 / 25

In-Process and Two-Level Caches

Use local caches safely and keep them consistent across instances.

The fastest cache is in your own memory

An in-process cache lives inside the application's memory: no network, no serialisation, nanosecond lookups. Libraries such as Caffeine (Java), cachetools or functools.lru_cache (Python), MemoryCache (.NET) and lru-cache (Node.js) provide size limits, TTLs and eviction. The catches: memory is limited and competes with the application (and the garbage collector), every instance has its own copy that warms separately, and instances can disagree after an update. A two-level cache puts a small local L1 with a short TTL in front of a shared L2 such as Redis: most reads are served locally, and the short TTL bounds inconsistency. For faster coherence, broadcast invalidations to all instances through pub/sub (for example Redis pub/sub or a message topic) so each instance drops the affected keys. Keep local caches for data that is safe to be briefly inconsistent: configuration, reference data, feature flags, rendered fragments.

A two-level cache with pub/sub invalidation (Java, Caffeine)

L1 serves most reads; an invalidation message clears the key on every instance.

Cache<String, Product> l1 = Caffeine.newBuilder()
        .maximumSize(50_000)
        .expireAfterWrite(Duration.ofSeconds(30))
        .build();

Product getProduct(String id) {
    return l1.get(id, key -> {
        Product p = redisGet("product:v2:" + key);            // L2
        if (p == null) {
            p = repository.findById(key);                    // source of truth
            redisSet("product:v2:" + key, p, Duration.ofMinutes(10));
        }
        return p;
    });
}

// every instance subscribes to the invalidation channel
void onInvalidate(String productId) {
    l1.invalidate(productId);
}

Watch memory and GC

A large on-heap cache increases garbage-collection work and can cause long pauses. Bound caches by size, measure their memory, and consider off-heap or shared caches for very large data sets.

Quick check: What is the main drawback of an in-process cache in a service with ten instances?

  • It is slower than Redis
  • It cannot have a TTL
  • Each instance holds its own copy, so instances can briefly serve different values
  • It requires a database
Answer

Each instance holds its own copy, so instances can briefly serve different values — Per-instance copies can diverge after updates until TTLs expire or invalidations arrive.