Tech Behind ThingsHow the ordinary machinery actually works

Software

Why A Database Refuses To Lose Half A Change

Systems that must never be partially updated write their intentions down before acting, so an interrupted operation can be finished or discarded on the next start.

Vibrant and engaging code displayed on a computer screen, showcasing programming concepts.
Photograph by Seraphfim Gallery via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Moving money between two accounts involves two changes that must both happen or neither. Power failing between them would be a serious problem, and the solution is older than the machines running it.

Storage completes writes in an order you did not choose

A program issuing several writes has little control over the sequence in which they reach the physical medium. Caching at every level reorders for efficiency.

An interruption can therefore leave the later change applied and the earlier one missing, which is worse than losing both.

Nothing in the storage layer understands that two writes belong together. That relationship exists only in the software above it.

The log is written before the data

The standard technique is to record the intended change in a sequential log and force that record to durable storage before touching anything else.

Only once the log entry is safely stored does the system modify the actual data. A crash at any point leaves the log as the authoritative account of what should have happened.

On restart, the system reads the log. Completed transactions are reapplied if necessary, and incomplete ones are rolled back so their partial effects disappear.

Sequential writing is cheaper than it sounds

Writing everything twice appears wasteful, but the log is appended sequentially while data updates are scattered across the storage.

Sequential writes are dramatically faster on nearly every medium, so the log can be forced to durable storage quickly while data updates are batched later at leisure.

The system can therefore confirm a change as permanent before it has been written to its final location, which is why databases are both durable and fast.

Durability depends on hardware telling the truth

The guarantee rests on a flush instruction genuinely committing data rather than acknowledging it while holding it in a volatile cache.

Drives and controllers have historically reported completion early to appear faster, which silently converts a durable system into one that merely usually works.

Devices intended for this use include capacitors or batteries sufficient to complete their pending writes after power is lost, which is the difference the specification is describing.

Filesystems and applications use the same idea

Journalling filesystems apply the technique to their own structures, which is why a modern system boots after a crash instead of spending an hour checking itself.

Applications reproduce it in miniature by writing to a temporary file, forcing it to storage, then renaming it over the original in a single indivisible step.

Skipping that pattern is the common cause of configuration files that end up empty. The write began, the machine stopped, and nothing recorded the intention first.

Questions readers ask

Why does a copied folder show a different size?

Block allocation, compression and metadata differ between filesystems. The contents are identical while the space consumed is not.

Is defragmenting a solid state drive useful?

No. There is no seek penalty to remove, and rewriting every block consumes write endurance for no measurable benefit.

Softwarestoragesoftwareoperating systemsdata
Junko Ishida
Contributing writer, Tech Behind Things

Junko covers batteries, charging and energy density, and is unimpressed by most battery claims.

Also by Junko Ishida