Software
Why text sometimes turns into question marks and boxes
Characters, encodings and fonts are three separate layers, and mangled text tells you which one broke.

Most explanations of character encoding stop at the point where it starts to matter. This one carries on.
The short version
- A character set assigns numbers to characters; an encoding stores those numbers as bytes.
- Reading bytes with the wrong assumed encoding produces predictable garbage.
- Empty boxes mean the font lacks a glyph, not that the text is broken.
Three layers that get confused
A character set assigns a number to each character, an encoding decides how those numbers become bytes, and a font supplies a shape to draw. Each layer can fail independently, and the visible symptom differs in each case. Replacement characters and strings of accented nonsense indicate an encoding failure; empty rectangles indicate a missing glyph.
Knowing which one you are looking at usually identifies the fix immediately.
The universal character set solved the first problem
Historically each language region had its own incompatible byte-to-character mapping, so a document was only readable with the right table. A single universal set assigns a unique number to every character in every script, which removed the ambiguity at that layer.
In the datasheet, it also includes combining marks, so some visible characters are built from several code points rather than one. This is why counting characters, reversing a string and truncating text are all harder than they look.
Variable-length encoding is why bytes are not characters
The dominant encoding uses one byte for the basic Latin range and up to four for everything else. That makes existing plain English text valid without change while supporting every script, which is why it won. It also means the number of bytes and the number of characters are different, and truncating by bytes can cut a character in half.
The short version: the resulting fragment is invalid and is displayed as a replacement character.
Mojibake has a signature
Text stored in one encoding and read as another produces characteristic patterns, such as pairs of accented Latin letters where an accented character should be. That pattern identifies the mismatch and often allows the original to be recovered by re-decoding correctly. Databases, web pages and file transfers each declare an encoding, and a mismatch anywhere in the chain corrupts the result.
Declaring the encoding explicitly at every layer is the fix, and guessing is the cause.
Fonts are a separate failure
A font contains glyphs for a subset of characters, and no font covers the whole universal set. When a glyph is missing the system substitutes another font if it can, and shows an empty box when it cannot. Some systems display the character number inside the box, which is genuinely helpful for diagnosis.
In practice, installing a font covering the script in question fixes it; changing the encoding does not.
Implementations differ, and vendors are not obliged to document the differences.
Emoji made everything harder
Skin tone and profession variants are built by joining several code points with an invisible joiner character. Systems that do not understand a particular sequence display the components separately, which is why a single emoji sometimes appears as two. New characters are added periodically and older systems simply lack them, which is a font and platform version issue rather than a bug.
In practice, the same mechanism underlies flags and family sequences, which is why they render inconsistently across platforms.
The takeaway
Garbled letters are an encoding fault. Empty boxes are a font fault. They need different fixes.
Once you know what it is trading away, the design stops looking arbitrary.
Questions readers ask
Why does my exported spreadsheet show strange accented characters?
The file was written in one encoding and opened as another. Most spreadsheet tools let you specify the encoding on import, and choosing the correct one recovers the text.
Why do I see empty boxes instead of some characters?
The text is fine and the font has no shape for those characters. Install a font covering the relevant script or emoji set.
Also by Junko Ishida
- USB-C solved the connector and not the confusionPower & Batteries
- File formats decide whether your files outlive the softwareSoftware
- Where the electricity goes when everything is switched offPower & Batteries
- The memory effect was real, and it has nothing to do with your phonePower & Batteries





