Mojibake (Character Corruption)
A phenomenon where text displays as garbled symbols or incorrect characters due to a mismatch between the encoding used to write the data and the encoding used to read it.
Mojibake is a Japanese term that has been adopted internationally to describe the garbled text that appears when a text file is decoded using a different character encoding than the one it was encoded with. For example, opening a UTF-8 encoded Japanese file as Shift_JIS produces strings of meaningless characters, while opening a Shift_JIS file as UTF-8 floods the screen with replacement characters (U+FFFD) or completely wrong glyphs.
- 1. The writer saves the text as UTF-8 The five Japanese characters こんにちは take three bytes each in UTF-8, so the file holds a sequence of 15 bytes.
- the promise "this file is UTF-8" never reaches the reader
- 2. The reader splits those bytes as Shift_JIS The same 15 bytes are regrouped into one- and two-byte units under Shift_JIS rules, so characters are assembled at different boundaries than the original five.
- a different code table decides which glyphs to show
- 3. The screen shows 縺薙s縺ォ縺。縺ッ Only the display changed. The original five characters were not replaced by different data; the reading side is simply applying the wrong encoding.
| Encoding used to save | Encoding used to read | What appears on screen |
|---|---|---|
| UTF-8 | UTF-8 (match) | こんにちは |
| UTF-8 | Shift_JIS | 縺薙s縺ォ縺。縺ッ |
| UTF-8 | Latin-1 | Unreadable symbols, and the five characters are counted as 15 |
| Shift_JIS | UTF-8 | A flood of replacement characters �, or completely wrong glyphs |
Only one of the four rows reads correctly: the row where the encoding used to save matches the encoding used to read. Confirm that pairing before you trust any character count.
Three typical scenarios produce mojibake. First, a file is saved in one encoding but opened in another. Second, a database connection's character set does not match the table's character set. Third, the Content-Type HTTP response header specifies a charset that differs from the actual encoding of the HTML file. All three share the same root cause: the writer and reader disagree on the encoding contract.
Historically, mojibake has been closely tied to Japanese computing culture. From the 1980s through the 1990s, three major Japanese encodings coexisted: JIS, Shift_JIS, and EUC-JP. This fragmentation caused rampant character corruption in email and web pages. Email was especially problematic because ISO-2022-JP was the nominal standard, yet different mail clients would silently use other encodings, producing garbled messages on the receiving end.
Preventing mojibake in practice comes down to a few clear rules. Save files as UTF-8 (without BOM). Set database character sets to utf8mb4. Include Content-Type: text/html; charset=UTF-8 in HTTP responses. When generating CSV files for Excel, output BOM-prefixed UTF-8. Applied consistently, these conventions keep the writer and the reader on the same encoding, which is the condition that prevents mojibake.
From a character counting perspective, mojibake-affected text shows a large discrepancy between the visible character count and the actual byte count. For instance, when UTF-8 Japanese text is misinterpreted as Latin-1, each character appears to expand into roughly three characters, inflating the character count to nearly triple its true value. Accurate character counting depends on the text's encoding being correctly identified in the first place.