Tech Behind ThingsHow the ordinary machinery actually works

Software

Almost everything on your screen was fetched before you asked for it

Caching is a bet that you will want the same thing twice. Every layer of a computer makes that bet, and the losses are invisible.

A laptop screen shows a coding application with a calculator design in a tech office setting.
Photograph by Eduardo Rosas via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

The theory of caching is well covered elsewhere. This is about the version you meet in practice.

What holds up in practice

  • A cache trades memory for time and is useless without repetition.
  • Eviction policy decides what survives when space runs out.
  • Knowing when a cached copy is stale is the genuinely hard part.

A cache is a bet about repetition

Caching keeps a copy of something expensive to obtain, on the assumption that it will be wanted again before long. The bet pays whenever the copy is reused and costs memory whenever it is not, so the value depends entirely on access patterns.

Real workloads repeat heavily, because programs revisit the same data and people revisit the same pages, which is why caching works so well. A workload that touches everything once and never returns gains nothing and pays the full cost of maintaining the cache. This is why streaming a huge file can push useful data out of a cache and slow down everything else afterwards.

Layers all the way down

The processor keeps small, extremely fast caches beside its cores, because reading main memory takes an eternity in processor terms. The operating system holds recently read file contents in memory, so a second read of the same file never touches storage.

The short version: storage devices themselves cache writes, acknowledging them before the data has physically reached the medium. Applications cache their own computed results, and browsers cache downloaded resources on disk between sessions. Each layer is unaware of the others, which is why a single request can be answered from any of five different places.

Deciding what to throw away

A cache is always smaller than the thing it caches, so it must constantly choose what to discard to make room. The common approach discards whatever has gone longest without being used, on the assumption that the past predicts the future.

In the datasheet, that assumption fails during a large sequential scan, which floods the cache with data that will never be requested again. Better schemes track how often something is used as well as how recently, which resists that flooding effect. Choosing badly is not merely a missed opportunity, because evicting hot data forces expensive re-fetches immediately afterwards.

Invalidation is the hard problem

A cache holds a copy, and nothing automatically informs it when the original has changed somewhere else. Systems handle this with expiry times, with version identifiers, or by having the source actively announce changes.

In practice, expiry is simple and always wrong in one direction: either data is held too long or it is discarded too soon. Version identifiers embedded in a resource name sidestep the problem, because a changed resource has a different name entirely.

Bugs here are unusually confusing, since the system behaves correctly for some users and incorrectly for others simultaneously.

Why clearing the cache fixes things

Clearing forces every layer to fetch fresh copies, which resolves any problem caused by something stale being held somewhere. It works so often because staleness is a common failure and the fix is indiscriminate rather than targeted. The cost is a slow session afterwards, since everything must be fetched and stored again from the beginning.

At the protocol level, it also discards data that was perfectly correct, which is why it is a blunt instrument rather than a maintenance routine. Regularly clearing caches as a habit generally makes a device slower without improving anything measurable.

This is the general case; a specific device may behave differently by design.

When caching actively hurts

A cold cache after a restart makes a system feel sluggish until the useful data has been read once and stored. Caching authenticated content is a security failure rather than a performance issue, because one person's data can reach another. Caches consume memory that other work could use, and an oversized cache can starve the very programs it was helping.

Some systems thrash, spending more effort managing entries than they save by reusing them, which is worse than no cache. Measuring the hit rate rather than assuming one is the only way to tell which side of that line a system sits on.

The takeaway

A cache is a guess about the future, and its worth is entirely in the hit rate.

The constraint is almost always physical, and marketing rarely mentions which one.

Questions readers ask

Why does an app get faster the second time I open it?

Its code and data were read from storage into memory and are still there. The second launch skips the slowest step entirely.

Do bigger caches always help?

No. Beyond the size of the working set, extra capacity holds data nobody wants while taking memory from something that did.

Softwareperformancesoftwarememoryweb
Mikkel Aas
Editor, Tech Behind Things

Mikkel edits Tech Behind Things and has taken apart more devices than he has successfully reassembled.

Also by Mikkel Aas