Four cache layers with their hit rates, where the query cache hits three per cent of the time and therefore costs more than it saves

Caching Strategy

Back to Performance Tuning and Capacity Planning · Kernel and Filesystem Parameters · Connection Pool Sizing · Before-and-After Measurement · Service Offerings

Application, query, and CDN caching applied where it actually reduces load, not layered on everywhere by default. The default assumption is that a cache can only help, and it is wrong in a specific and measurable way.

1. A Cache That Misses Is Not Free

Every miss costs twice: the lookup that found nothing, and the write that stores a result nobody will ask for again. At a 3 per cent hit rate — the bottom row of the illustration above — ninety-seven requests in a hundred pay both and gain nothing.

So the question for each layer is not “could this be cached” but “how often is the same thing asked for twice”. Where the answer is rarely, the cache is a tax. Measure the hit rate of what you already have before adding another layer, because a layer with a low rate is usually easier to remove than to improve.

2. The Layer Closest to the User Is Worth the Most

The same hit avoids more work the further out it happens. A CDN hit skips your network, your server, your application and your database; a query cache hit skips only the query.

Layer What a hit avoids What makes it hard
BrowserThe request entirelyYou cannot clear it; a wrong long expiry is wrong for months
CDNEverything from your network inwardsAnything personalised fragments it into uselessness
ApplicationThe computation and the queries behind itKnowing when the answer stopped being true
QueryOne queryLow hit rates, because exact repeats are rarer than they seem

Effort therefore belongs at the top of that table, and the common mistake is the reverse: the query cache is the easiest to switch on and the least valuable when it works.

3. Invalidation, Which Is the Actual Work

Storing is trivial. Knowing when a stored answer stopped being true is the engineering, and it is why a cache is a permanent obligation rather than a one-off change.

4. The Stampede

A popular entry expires, and every request that wanted it arrives at the database at once. The cache was absorbing a hundred requests a second; now all hundred land together, on a system sized for the cached load.

5. What Not to Cache

That last one is worth dwelling on. A cache over a bad query hides it: the query stays bad, the first request after every expiry still waits, and the problem reappears the moment the hit rate drops.

6. Measure the Layer, Not the Idea

Every cache should report hit rate, miss cost and eviction rate, and those three numbers decide whether it stays.

How We Approach It

  1. Measure what the existing layers achieve — hit rate, eviction rate, and what a miss actually costs.
  2. Find the repeats. Caching only pays where the same answer is wanted more than once, and the request log says where that is.
  3. Work from the outside in, since a hit further out avoids more.
  4. Decide invalidation before adding anything, including what the key must contain to keep one user's data away from another.
  5. Protect against the stampede with stale-while-revalidate, a recomputation lock, and jittered expiry.
  6. Remove the layers that do not earn their keep, which is the part that normally never happens.

What You Get

The uncomfortable question worth asking of any existing cache: what is its hit rate? If nobody knows, it is as likely to be costing you as saving you.