Tech Behind ThingsHow the ordinary machinery actually works

Software

Why Clearing A Cache Fixes Things It Should Not

Caches store copies to avoid repeating work, and every one of them must guess how long a copy stays valid, which is where the bugs live.

Vibrant and engaging code displayed on a computer screen, showcasing programming concepts.
Photograph by Seraphfim Gallery via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Clearing a cache is the standard remedy for a problem nobody can explain. It works often enough to be a reflex, and the reason is a genuinely hard problem underneath.

A cache is a bet about the future

Storing a copy of something is only useful if it will be needed again and if it has not changed in the meantime. Both are predictions.

The first prediction is usually safe, since access patterns repeat heavily. The second is the difficult one, because the cache rarely controls the original.

Every caching layer therefore needs a rule for when a copy stops being trustworthy. Getting that rule wrong produces stale data rather than an error.

Expiry and validation are the two approaches

The simplest rule is a lifetime. The copy is used freely until a stated time, after which it is discarded and fetched again.

The alternative is to ask the source whether anything changed, using a short identifier for the current version. If the answer is no, the existing copy is kept.

Validation is accurate and costs a round trip. Expiry is free and can be wrong for the whole duration, which is why lifetimes on frequently changing content are kept short.

Layers stack without knowing about each other

A single request may pass through a cache in the application, one in the browser, one at an intermediate network, and one at the origin.

Each applies its own rules, and a copy considered fresh by one layer can be stale by the standard of another. No layer sees the whole chain.

Clearing works because it collapses the ambiguity. It forces every layer to start again from a known state rather than reasoning about which copy is correct.

Invalidation is the part that fails

When content changes, every stored copy must be told. Systems that push such notifications can be precise, but reaching every layer is often impossible.

The practical workaround is to change the name. A file published under a new identifier cannot collide with any cached copy of the old one.

This is why asset filenames on websites contain long strings of characters. The name encodes the content, so a change of content is automatically a change of address.

The failure is silent by design

A stale cache produces plausible output. Nothing errors, nothing logs, and the user sees a page or a setting that was correct at some point in the past.

Because the symptom resembles an application bug, effort goes into investigating code that is behaving perfectly on data that is old.

Clearing the cache is the fastest way to distinguish the two. It is a diagnostic step being used as a repair, which is why it works and explains nothing.

Questions readers ask

Why does a copied folder show a different size?

Block allocation, compression and metadata differ between filesystems. The contents are identical while the space consumed is not.

Is defragmenting a solid state drive useful?

No. There is no seek penalty to remove, and rewriting every block consumes write endurance for no measurable benefit.

Softwarestoragesoftwareoperating systemsdata
Junko Ishida
Contributing writer, Tech Behind Things

Junko covers batteries, charging and energy density, and is unimpressed by most battery claims.

Also by Junko Ishida