Caching
The first tool you reach for at scale: cache-aside vs. write-through vs. write-behind, eviction and TTLs, the hard problem of invalidation, and where a CDN fits.
On this page
When an estimate reveals a read-heavy system - and most systems are - the first and highest-leverage tool you reach for is a cache. Caching stores a copy of data somewhere faster or closer than its source, so repeated reads skip the expensive work. It's how you turn a database that would melt under 100,000 reads/sec into one that handles the writes while a cache absorbs the reads. But caching introduces the two genuinely hard problems in computer science - so it demands respect, not just enthusiasm.
Why caching wins: the latency hierarchy
Recall the latency numbers: RAM (~100 ns) is roughly 1000× faster than an SSD read, and both crush a cross-network database query. A cache exploits this by keeping hot data in memory (often a dedicated in-memory store like Redis or Memcached) so a read that would hit the database is served from RAM instead. Add the 80/20 rule - a small fraction of data gets most of the requests - and a modestly-sized cache can absorb the overwhelming majority of reads. That's the leverage: cache the hot 20% and your database load drops by most of the read traffic.
Where caches live
Caching happens at many layers, and real systems stack them:
- Client / browser cache - the response is stored on the user's device.
- CDN (Content Delivery Network) - caches static assets (images, JS, video) at edge locations physically near users, killing that ~150 ms cross-continent round trip. Essential for global, static-heavy content.
- Application / distributed cache - a shared in-memory store (Redis/Memcached) in front of the database, holding hot query results, sessions, computed values.
- Database cache - the DB's own buffer pool caching pages in memory.
The closer to the user a cache sits, the bigger the latency win - but the harder it is to keep fresh.
Caching strategies (write patterns)
The subtlety is how the cache stays in sync with the database. The main patterns:
- Cache-aside (lazy loading) - the most common. The app checks the cache; on a miss, it reads the DB, populates the cache, and returns. Writes go to the DB and invalidate (or update) the cache entry. Simple, and only-requested data is cached - but the first read is always a miss, and there's a staleness window.
- Read-through - the cache itself loads from the DB on a miss (the app only talks to the cache). Like cache-aside but the loading logic lives in the cache layer.
- Write-through - every write goes to the cache and the DB synchronously. The cache is always fresh, but writes are slower (two hops).
- Write-behind (write-back) - writes go to the cache and are flushed to the DB asynchronously later. Fast writes, but you risk losing data if the cache dies before the flush.
Cache-aside is the sensible default; write-through when reads must never be stale; write-behind only when write throughput matters more than durability.
The two hard problems: eviction and invalidation
"There are only two hard things in Computer Science: cache invalidation and naming things." - Phil Karlton
A cache has finite memory, so it must evict entries. Common policies: LRU (least recently used - evict what hasn't been touched longest, the usual default), LFU (least frequently used), and TTL (time-to-live - entries expire after a set time, a simple way to bound staleness).
But the truly hard problem is invalidation: keeping the cache consistent with the source of truth. When the underlying data changes, stale cache entries must be updated or removed - and getting this wrong serves users old data. Approaches: TTL (accept staleness up to the expiry - simplest, often good enough), explicit invalidation on write (delete/update the entry when the DB changes - precise but you must catch every write path), or event-based invalidation (a change event evicts the entry). There's no perfect answer; you trade freshness against complexity, which is exactly why it's famously hard.
Caching trades freshness for speed - decide how much staleness is OK
Every cache introduces the possibility of serving stale data. The engineering question is never 'should we cache?' but 'how stale can this data be?' A user's account balance may need to be fresh to the second (short/no TTL or write-through); a product's review count can be minutes stale (a long TTL is fine); a static image effectively never changes (cache forever at the CDN). Match the strategy and TTL to the data's tolerance for staleness - there's no one setting, and over-caching mutable data is how you ship confusing bugs.
A cache is a chef's mise en place - the small tray of pre-chopped, ready-to-grab ingredients at their station. Instead of walking to the walk-in fridge (the database) for every order, the chef keeps the hot, frequently-used items right at hand (in fast-access memory), serving most dishes in seconds. The tray is small, so they keep only what's used most (eviction - LRU: yesterday's unused garnish gets cleared off), and the hard part is keeping it fresh: if the supplier changes the recipe or an ingredient spoils, the chef must remember to refresh the tray (invalidation), or they'll plate the old version. A brilliant mise en place makes the kitchen fly; a stale one serves yesterday's dish. Speed is easy; keeping the tray in sync with the fridge is the craft.
An e-commerce product page shows: (a) the product's name and description (changes rarely), (b) its current price (changes occasionally, must not be wrong at checkout), and (c) its live stock count (changes constantly). The page gets enormous read traffic. Propose a caching approach for each piece of data, choosing strategy and rough TTL, and justify based on staleness tolerance.
What are the two genuinely hard problems that caching introduces?
Key takeaways
- Caching stores a copy of data somewhere faster/closer (often RAM via Redis/Memcached) so repeated reads skip expensive work - the first tool for read-heavy systems.
- It exploits the latency hierarchy (RAM ~1000x faster than SSD) and the 80/20 rule: caching the hot fraction absorbs most reads and offloads the database.
- Caches stack across layers: browser, CDN (static assets at edge locations near users), application/distributed cache, and the DB's own buffer pool.
- Write strategies trade freshness vs speed: cache-aside (default), read-through, write-through (always fresh, slower writes), write-behind (fast writes, durability risk).
- The two hard problems are eviction (LRU/LFU/TTL) and, hardest, invalidation - match strategy and TTL to how much staleness each piece of data can tolerate.