Machine Translation
Technology that automatically translates text from one language to another using a computer. The advent of neural machine translation (NMT) has dramatically improved quality, enabling cross-language conversion that involves changes in character count.
Machine translation (MT) is the technology that converts text from one language to another without human intervention. Services such as Google Translate, DeepL, and Microsoft Translator offer it to the public, and it is used daily for translating web pages, drafting business documents, and enabling real-time conversation across languages.
The history of machine translation spans three generations. First-generation rule-based translation grew out of research that began in the 1950s and converted text with hand-written grammar rules and bilingual dictionaries. Second-generation statistical machine translation built on research from the 1990s and reached the major translation services in the 2000s, learning word choice and word order probabilistically from large parallel corpora. Third-generation neural machine translation (NMT) came into practical use in the mid-2010s; because it treats the whole sentence as one sequence instead of replacing words and phrases, its output reads in a more natural order. The boundaries between generations are not sharp calendar years, and several approaches coexisted in the same period.
Machine translation and character count are closely linked. Expressing the same content in different languages produces significant variation in character count. The Japanese word "情報" (2 characters) becomes "information" (11 characters) in English. As a rule of thumb, translating from Japanese to English increases the character count by a factor of 1.5 to 2, while English to Japanese reduces it to 0.5 to 0.7 times the original. This expansion ratio directly affects the sizing of buttons and labels in UI localization.
Post-translation character limits are a major practical challenge. Take a post on X (formerly Twitter): the limit is stated as 280, but it is a weighted count in which each Japanese, Chinese, or Korean character counts as 2, so a post written only in Japanese is capped at 140 characters; Latin letters count as 1 each, so an English translation effectively gets twice the room, yet 280 characters may still be too few to say everything a 140-character Japanese post contains. For meta descriptions, ad copy, and UI labels with character constraints, simple translation is not enough; paraphrasing or summarizing to fit within the limit is necessary.
BLEU is one of the standard metrics for evaluating machine translation quality. BLEU compares the machine output against a human reference translation using N-gram match rates, producing a score from 0 to 100. Note that the same output can receive different scores depending on the tokenization procedure and the number of reference translations, so figures published by different papers or services cannot be lined up and compared directly; for language pairs such as Japanese-English, whose word order and segmentation differ greatly, N-gram matches are harder to obtain even when the translation reads well, so a low score does not by itself mean low quality.