Tech Behind ThingsHow the ordinary machinery actually works

Software

Why Sorting Names Is Harder Than It Looks

Alphabetical order is not a property of text but a set of language-specific rules, which is why the same list sorts differently on two machines.

Vibrant and engaging code displayed on a computer screen, showcasing programming concepts.
Photograph by Seraphfim Gallery via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Sorting a list of words seems like a solved problem. It is solved separately for every language, and the answers disagree with each other.

Byte order is not alphabetical order

Text is stored as numbers, and the simplest sort compares those numbers. For plain unaccented English this coincidentally resembles alphabetical order.

It breaks immediately elsewhere. Capital letters sort before lowercase ones, so a list ordered this way puts every capitalised word ahead of every uncapitalised one.

Accented characters sit far away from their unaccented counterparts in the numbering, so they land at the end of the list rather than beside their base letters.

Languages disagree about the alphabet

Some languages treat a pair of letters as a single unit that sorts as its own entry rather than between the letters it resembles.

Others place accented forms at the end of the alphabet as distinct letters, while their neighbours treat the same marks as decoration to be ignored in the first pass.

There is no neutral answer. A correct sort requires knowing which language's conventions apply, and that is a property of the reader rather than the data.

Comparison happens in several passes

Proper collation compares base letters first, ignoring case and accents entirely. Only where those are equal does it consider accents, and only then case.

This produces the ordering people expect, where a word and its accented variant sit adjacent rather than at opposite ends.

The rules are maintained as data tables rather than code, because they change and because every language needs its own adjustments to a shared default.

The same characters can be stored differently

An accented letter may be a single character or a base letter followed by a combining mark. Both display identically and compare as different byte sequences.

Text must therefore be normalised to a consistent form before comparison, or two visually identical strings will be treated as unequal.

This is a frequent source of bugs where a search fails, a login is rejected or a duplicate slips through, and the two values look the same in every log.

Numbers inside text confound everything

Sorting filenames alphabetically places item ten before item two, because comparison proceeds character by character and one precedes two.

Natural sorting detects digit sequences and compares them as numbers, which matches expectation and adds ambiguity around leading zeros and version strings.

Because both behaviours are defensible, different applications choose differently, and a list can be sorted correctly by two programs and still not match.

Questions readers ask

Why does a copied folder show a different size?

Block allocation, compression and metadata differ between filesystems. The contents are identical while the space consumed is not.

Is defragmenting a solid state drive useful?

No. There is no seek penalty to remove, and rewriting every block consumes write endurance for no measurable benefit.

Softwarestoragesoftwareoperating systemsdata
Junko Ishida
Contributing writer, Tech Behind Things

Junko covers batteries, charging and energy density, and is unimpressed by most battery claims.

Also by Junko Ishida