Character Encoding
A system of rules that maps characters to sequences of bits. It consists of two layers: a character set (which characters are included) and an encoding scheme (how those characters are converted to byte sequences).
Character encoding is the foundational technology that allows computers to handle text. It defines the rules for converting human-readable characters like "A" or "あ" into numeric values that machines can process. Without these conversion rules, storing, transmitting, and displaying text would be impossible.
Character encoding is best understood as two distinct layers. The first layer is the character set, which defines which characters are available. ASCII covers 128 characters, JIS X 0208 covers 6,879 characters (6,355 kanji and 524 non-kanji), and Unicode covers over 150,000 characters (159,801 in version 17.0, released in September 2025). The second layer is the encoding scheme, which determines how each character in the set is represented as a byte sequence. This is why multiple encoding schemes (UTF-8, UTF-16, UTF-32) can exist for the same Unicode character set.
Japanese character encoding has a particularly complex history. JIS C 6226 (later renamed JIS X 0208) was established in 1978, and from it emerged three competing encoding schemes: Shift_JIS (developed for PCs and also known as the MS Kanji code), EUC-JP (adopted on UNIX systems), and ISO-2022-JP (used for email). Because all three represented the same character set with different byte sequences, converting between them became a source of mojibake (character corruption).
Today, the industry has largely converged on Unicode + UTF-8. UTF-8 maintains full backward compatibility with ASCII, representing English text at 1 byte per character and Japanese text at 3 bytes per character. This variable-length design lets it handle every writing system in the world while remaining compatible with ASCII-based infrastructure. In W3Techs' usage survey, 99.0% of websites with a known character encoding used UTF-8 as of August 2026, and for new systems the reasons to choose anything else are mostly limited to exchanging data with existing systems that require Shift_JIS or EUC-JP.
A practical concern when working with legacy systems is that converting from Unicode to older encodings like Shift_JIS can be lossy. Characters that exist in Unicode but not in Shift_JIS (certain emoji, CJK Unified Ideographs Extension B and beyond) are replaced with "?" or "〓" during conversion. Character encoding conversion is a potentially irreversible operation that can result in permanent data loss.