Caching in System Design: A Walkthrough With Diagram

By Aniruddha · September 28, 2026 · 8 min read

Caching comes up in almost every system design interview, usually as the answer to "the database is too slow". The idea is simple: keep hot data in memory. The interesting part is everything around it: when to fill the cache, how to keep it correct, and what to throw out when it is full.

The 50 second animation below follows one key, product:7, through a miss, a hit, a price change and an eviction. I drew it in CanvasKeep. The sections after it cover the write strategies, the invalidation race, and the failure modes that did not fit in the video.

1. Why cache

A database read involves a network hop, query planning, and often disk reads and joins. A simple indexed lookup can take a millisecond or two, but real queries under load often take tens of milliseconds. Reading the same value from an in-memory store like Redis or Memcached usually takes well under a millisecond inside the same data center.

Most apps are read heavy, and reads repeat: the same product page, the same user profile, the same feed, over and over. Without a cache, every one of those is a database query, and as traffic grows the database becomes the bottleneck.

2. The diagram

The app server sits between the client and two stores. It asks the cache first and only reads the database on a miss.

Caching system design diagram: a client calls an app server, which checks a Redis cache first and reads the database only on a miss
Click to open full size.

Get the editable diagram

An .excalidraw file. Open it on excalidraw.com, or in a CanvasKeep diagram with Open from the menu.

Download

3. Cache-aside, step by step

This is the most common pattern, also called lazy loading. The application owns the logic:

v = cache.get("product:7")
if v: return v                     # hit
v = db.query("... WHERE id = 7")   # miss
cache.set("product:7", v, ttl=60)
return v

On a miss

The first request for product:7 finds nothing in the cache, reads the database (about 50 ms in the video), and stores the result. That request is slow, once.

On a hit

Every request after that is served from memory, and the database is never asked. The cache only holds data someone actually requested, and if the cache goes down the app still works, just slower.

4. Hit ratio is the number that matters

Hit ratio is hits divided by all reads. At 90%, the database sees 1 in 10 reads. At 99%, 1 in 100. Going from 90% to 99% cuts database load by 10x, which is why it is worth watching.

Hit ratio depends on how skewed your traffic is (a few very popular keys cache well), how much memory the cache has, and how long entries live.

5. Write strategies

  • Write-around (pairs with cache-aside). Writes go straight to the database and skip the cache. Any cached copy is deleted, and the next read reloads it.
  • Write-through. Writes go to the cache and the database together. Reads right after a write are fast, but you cache data that may never be read.
  • Write-behind (write-back). Writes go to the cache and are flushed to the database later. Very fast writes, but a cache crash can lose data.
  • Read-through. Like cache-aside, but the cache library loads from the database itself on a miss, so the app only talks to the cache.

6. The catch: stale data

When the price of product:7 changes from $10 to $12 in the database, the cache still says $10. Every hit returns the wrong price until something fixes it. Two tools, used together:

  • Delete on write. After updating the database, delete the key. The next read misses and loads $12. Deleting is safer than updating the cache, because two concurrent updates can land in the wrong order and leave the old value behind.
  • A TTL. Every entry expires after, say, 60 seconds, so anything invalidation misses is wrong for at most that long.

There is still a small race: a slow read can load $10 from the database, then the write deletes the key, then the slow read puts $10 back. The TTL bounds how long that lasts. Facebook's Memcache paper, linked below, solves it with leases: the cache hands out a token on a miss and rejects a set whose token was invalidated by a delete.

7. Eviction: memory is small

A cache holds far less than the database. When it is full, adding a key means evicting one. The usual policies:

  • LRU (least recently used). Evict whatever was read longest ago. Recently read keys survive. This is the common default. Redis approximates it by sampling a few keys rather than tracking exact order, which works almost as well and saves memory.
  • LFU (least frequently used). Evict what is read least often. Better when a few keys are always hot.
  • TTL-based. Evict keys closest to expiring.

In Redis, set maxmemory (on 64-bit systems there is no limit by default) and choose a policy such as allkeys-lru or allkeys-lfu. The default policy, noeviction, returns errors on writes once the limit is reached, which is rarely what a cache wants.

8. What the 60 second version leaves out

  • Cache stampede. A hot key expires and thousands of requests miss at once. Let one request rebuild it (a lock or request coalescing) and add jitter to TTLs.
  • Cache penetration. Requests for keys that do not exist always miss and always hit the database. Cache the "not found" result briefly.
  • Cold start. A fresh cache has a 0% hit ratio. Warm it with the most popular keys before sending it full traffic.
  • Hot keys. One extremely popular key can overload a single cache node. Replicate it or keep a small local cache in each app server.
  • Layers. Caching also happens in the browser, at the CDN, and inside each app process. The same trade-offs apply at every layer.

Interview checklist

  • Say what you cache and why: read-heavy, rarely changing data.
  • Use cache-aside and walk through a miss and a hit.
  • Mention hit ratio and what it does to database load.
  • Handle writes: delete on write plus a TTL.
  • Pick an eviction policy (LRU is a fine default).
  • Cover stampedes, penetration, cold start, and hot keys.

To see caching inside a full design, the URL shortener walkthrough uses exactly this pattern for its redirects.

FAQ

What is the cache-aside pattern?

The application checks the cache first. On a hit it returns the cached value. On a miss it reads the database, writes the result into the cache, and returns it. The cache only ever holds data that something actually asked for.

Should I update or delete the cache when data changes?

Delete it. Updating the cache from two concurrent writes can leave the older value in the cache for good. Deleting means the next read misses and loads the current value from the database. Keep a TTL as a safety net for anything invalidation misses.

What is a good cache hit ratio?

It depends on the workload. For read-heavy data like product pages or profiles, teams often aim for 90% or higher. A 90% hit ratio means the database sees only 1 in 10 reads, so small improvements at the top end cut database load a lot. In Redis, compute it from keyspace_hits and keyspace_misses in INFO stats.

What is a cache stampede?

When a popular key expires, every request for it misses at once and they all hit the database together. Prevent it by letting only one request rebuild the value while the others wait or get the old value, and by adding random jitter to TTLs so keys do not expire together.

Sources and further reading

More system design in 60 seconds

The diagram was drawn in CanvasKeep, which keeps Excalidraw diagrams in workspaces and folders with version history.

Keep your system design diagrams in one place

The Excalidraw canvas, with workspaces, nested folders, and named versions. Free to start.

Start drawing