← All posts
ConceptsJuly 18, 202610 min read

Caching Strategies Explained: Cache-Aside, Write-Through & More

Caching is the highest-leverage tool in system design, and "I'll add a cache" is where a lot of interview answers stop. That sentence commits you to nothing. The strategy you pick underneath it decides how stale your reads can be, whether a crash loses writes, and what happens to your database at the exact moment a popular key expires.

This post covers why caching works at all (with the arithmetic), the read and write patterns and what each costs, eviction and sizing, why you should delete cache entries rather than update them, and the four failure modes interviewers reliably probe.

Why caching works, with numbers

Caches work because real access patterns are heavily skewed. A small fraction of keys takes a large fraction of requests — Zipf-distributed, roughly, in almost every human-facing system. That means a cache holding a few percent of your data can serve most of your traffic.

The payoff is worth being able to compute on a whiteboard. Effective latency is just a weighted average:

effective = (hit_rate x cache_latency) + (miss_rate x db_latency)

cache 1ms, database 50ms:

  90% hit rate  ->  0.90(1) + 0.10(50)  =  5.9ms
  95% hit rate  ->  0.95(1) + 0.05(50)  =  3.45ms
  99% hit rate  ->  0.99(1) + 0.01(50)  =  1.49ms

Two things fall out of that, and both are worth saying aloud.

First, the last few percent of hit rate matter enormously. Going from 90% to 95% cuts effective latency by 42%. Going from 95% to 99% more than halves it again. The misses dominate the average long after the hit rate looks respectable, which is why chasing hit rate has such outsized returns.

Second, the same arithmetic applies to database load. At a 95% hit rate your database sees 5% of read traffic — a 20× reduction. At 99% it sees 1%, a 100× reduction. That is usually the real reason a cache is in the design: not user-facing latency, but keeping the database alive.

Reading through a cache

HIT — one hop app cache ~1ms get value The database is never touched. This is 95%+ of requests. MISS — four steps, and the app does the work app cache miss database ~50ms 1 get 2 read 3 row 4 populate Only requested keys are ever cached, so memory stays lean.
Cache-aside, the default for most systems. The application owns the logic, which is why it is also called lazy loading. The variant worth naming is read-through, where the cache itself knows how to fetch from the database — identical behaviour, cleaner application code, but it requires cache support for a loader.

The consequences of cache-aside are worth stating because interviewers ask about both: the first request for any key is always a miss, so a cold cache after a restart can briefly hammer the database; and the cache can drift from the database whenever something else writes, which is what invalidation is for.

Writing through a cache

Read strategy is where most candidates stop. Write strategy is where the interesting failures live.

WRITE-THROUGH app cache database sync Always fresh. Pays both latencies. WRITE-BACK app cache database later Fastest writes. Crash here = data gone. WRITE-AROUND app skipped database No cache churn from write-once data.
Same three boxes, three different write paths. Write-back is the one to be careful about in an interview — it is genuinely the fastest, and it will lose acknowledged writes if the cache dies before flushing. Volunteering that trade-off is the difference between sounding fast and sounding trustworthy.

The combination most systems actually run is cache-aside for reads plus write-around for writes, with explicit invalidation on update. Write-through is worth it when reads immediately follow writes. Write-back belongs in workloads where throughput genuinely dominates and the cache is itself replicated and persistent — metrics ingestion, view counters — and it should be named as a deliberate durability trade, never as a default.

Delete the key, don't update it

This is a small detail that reliably impresses, because it shows you have thought about concurrency rather than just data flow.

The tempting thing on a write is to update both the database and the cached value. Consider two concurrent writers:

Writer A: sets x = 1        Writer B: sets x = 2

A writes x=1 to database
                            B writes x=2 to database
                            B writes x=2 to cache
A writes x=1 to cache       <- A is slower, lands last

database: x = 2    cache: x = 1    ...and it never self-corrects

The cache is now permanently wrong, with no TTL short enough to make that acceptable and no error anywhere. If instead both writers delete the key, the next read repopulates from the database and gets 2. Deletion is idempotent and order-independent; updating is neither.

The general rule: on a write, invalidate rather than refresh, unless you have a specific reason and a version number to enforce ordering.

Eviction and sizing

A cache is bounded, so something has to go.

LRU evicts whatever has gone untouched longest. It is the right default because it matches how skewed workloads behave — recently used is likely to be used again. LFU evicts the least-accessed, which handles a stable hot set better but adapts poorly when popularity shifts, and can let yesterday's hot keys squat forever unless it decays counts. TTL expires entries on a clock regardless of use, which is really a staleness control rather than a memory control, and is usually combined with LRU rather than used instead of it.

Sizing is a back-of-the-envelope exercise worth doing rather than guessing:

20M items, 500 bytes each, cache the hottest 20%

  20M x 0.20 x 500B  =  2 GB of values
  + ~100 bytes/entry overhead x 4M  =  0.4 GB
  ~= 2.5 GB, comfortably one node with headroom

The reason to run that calculation in an interview is that it frequently shows the working set fits entirely in memory — at which point eviction policy stops mattering and the conversation moves to more interesting problems. The same estimation habit is covered in back-of-the-envelope estimation.

The four ways caches fail

STAMPEDE — the hot key's TTL expires 10,000 req/s expired database 10,000 identical One expiry, and the database takes the full uncached load. COALESCED — one recomputes, the rest wait 10,000 req/s lock held by one database 1 query The other 9,999 either block briefly or are served the stale value while the refresh runs.
Stampede, and the standard fix. Serving the stale value while one worker refreshes — stale-while-revalidate — is usually better than blocking, because a value that is two seconds out of date is almost always preferable to a request that waits. Randomising TTLs by a few percent also stops thousands of keys expiring on the same tick.

The other three are worth naming:

Hot key. One key is so popular that the single cache node holding it saturates, no matter how well the rest of the keyspace is distributed. Sharding does not help, because the key hashes to one place. Fix it with a small in-process cache in front of the shared one, or by replicating that key under several suffixed names and picking one at random. This is the follow-up that catches people who only memorised consistent hashing — even distribution of keys says nothing about distribution of traffic.

Cache penetration. Requests for keys that do not exist miss the cache every time, by definition, and go straight to the database. An attacker can turn that into a denial of service by requesting random nonexistent ids. Fix it by caching the negative result with a short TTL, or with a Bloom filter in front that can say "definitely not present" cheaply.

The inconsistency window. Any TTL or async scheme means some reads see stale data for a bounded period. This is not a bug to be fixed, it is a parameter to be chosen — and the right value depends entirely on the data. Seconds of staleness on a follower count is invisible; seconds on an account balance is a support ticket.

A worked example: caching one product end to end

Rules are easier to apply when you have seen them applied, so here is the whole toolkit pointed at one system — a news feed, the kind of thing you might be asked to design from scratch. Four different pieces of data, four different answers.

The user's profile — name, avatar, bio. Read constantly, changed maybe twice a year. The working set is every active user, which at ten million actives and 400 bytes each is about 4GB — comfortably cacheable in full.

Cache-aside, TTL of an hour, explicit delete on profile update. The long TTL is safe precisely because writes are rare and invalidation is explicit; the TTL is a backstop against a missed invalidation, not the primary mechanism.

The rendered feed page for a user. Expensive to build — it fans out across followed accounts, ranks, and hydrates. Changes whenever anyone the user follows posts, which for an active user is constantly.

Cache-aside with a short TTL, 30 seconds, and stale-while-revalidate. Explicit invalidation is hopeless here: a single post by a popular account would need to invalidate millions of cached feeds. This is the case where bounded staleness is not a compromise but the only sane design, and where being able to say "I would not invalidate this, I would let it expire" shows judgement.

The post itself — text, media URLs, author id. Immutable once written, except for deletes.

Write-through, and cache indefinitely. Immutable data is the easy case and worth calling out as such: there is no invalidation problem when nothing changes. Deletes are handled by removing the key, which is rare enough that the cost does not matter.

The like counter. Written constantly, read constantly, and nobody can tell the difference between 12,481 and 12,490 likes.

Write-back, batched. This is the one place the durability risk is acceptable, because the consequence of losing a few seconds of increments is a slightly wrong number that self-corrects on the next reconciliation. Contrast that with the same strategy applied to an account balance, where it would be indefensible. The strategy is not good or bad in itself; it is good or bad for a specific piece of data.

Notice what happened across those four: one used explicit invalidation, one deliberately refused to, one needed no invalidation at all, and one accepted write loss on purpose. "I'll add a cache" collapses all four into one sentence and hides every decision that mattered.

Warming, and what a deploy does to you

One operational detail worth carrying into an interview, because it catches teams in production and almost never appears in prep material.

An empty cache is not merely slower — it can be fatal. If your database is provisioned for 5% of read traffic because the cache absorbs the other 95%, then a cache that restarts cold presents it with 20× its normal load, all at once, from a standing start. The system was sized for the steady state and the steady state assumed a warm cache.

This is a real risk at three moments: deploying a change that alters cache keys (which silently invalidates everything), restarting or failing over the cache cluster, and scaling the cache out, which redistributes keys and misses everything that moved.

The defences are straightforward once you have thought about it: roll cache nodes one at a time rather than all at once, so only a fraction of the keyspace goes cold; keep the cache persistent where the technology supports it, so a restart reloads rather than starts empty; pre-warm the hottest keys before taking traffic; and change cache key formats behind a gradual rollout rather than a flag flip. Mentioning even one of these unprompted signals operational experience rather than textbook knowledge.

What caching does not solve

It does not fix a slow query, it hides one. At a 95% hit rate you have made a 2-second query happen 20× less often. The remaining 5% are still 2 seconds, and if the cache ever goes cold — a restart, a deploy, an eviction storm — every request becomes that query at once. A cache in front of an unindexed table is a loaded gun.

It does not help write-heavy workloads. Caches accelerate reads. If your bottleneck is write throughput, a cache in front of it changes nothing.

It adds a failure mode you did not have. Now there are two systems, and the interesting question is what happens when the cache is unavailable. Fail open — go straight to the database — and a cache outage may take the database with it. Fail closed and a cache outage is a full outage. Pick deliberately, and say which you picked.

It makes reasoning about correctness harder. Every cached value is a claim about the past. Debugging "user says they updated it but it still shows the old value" is materially harder with a cache in the path, and that cost is real even when everything is working.

Answering it in an interview

Be specific about all four dimensions, briefly, when you introduce a cache:

  1. Which layer — browser, CDN, application cache, or database buffer pool. "Add Redis" skips this.
  2. Read strategy — cache-aside unless there's a reason.
  3. Write strategy and invalidation — "write-around, and I delete the key on update" is a complete answer in eight words.
  4. Staleness tolerance — say the number. "Up to 30 seconds stale is fine here" invites the right follow-up questions.

Then volunteer one failure mode before being asked. "The risk is a stampede when a hot key expires, so I'd coalesce requests and jitter the TTLs" is the sentence that separates a design from a buzzword.

If you want to be pushed on exactly those trade-offs out loud — "what happens when that key expires?" is the one interviewers always reach for — you can run a full system design mock by voice on Whitepad.

Practice this out loud

Reading is the easy part. Sit across from a senior AI interviewer that talks, watches your whiteboard, and scores you like the real thing — your first mock is free.

Start a free mock →