Compression Ratio
In data compression, the ratio of the compressed size to the original size. Text data is highly redundant and can typically achieve compression ratios of 60% to 80%.
Compression ratio is a metric that measures the efficiency of data compression. If the original data size is D and the compressed size is C, English usage defines the compression ratio as D/C, while the percentage (1 - C/D) × 100% is called space savings; in Japanese, however, the word for compression ratio is also used for this percentage. A 100 KB text file compressed to 25 KB therefore has a compression ratio of 4:1, or a space saving of 75%. A higher compression ratio means the data can be stored and transmitted using less storage and network bandwidth.
Text data tends to achieve higher compression ratios than images or video. Natural language text contains many forms of redundancy: skewed character frequency distributions (in English, "e" is the most common letter), repeated words, and formulaic phrases. What decides the result is the content rather than the language: how far the same phrases and fixed formats repeat is reflected almost directly in the outcome. Japanese text in UTF-8 starts from a larger byte count because each kanji or kana takes 3 bytes, but the same 3-byte sequences recur and compress well, so its compressed size is not necessarily far behind English. When a figure is needed, measuring the actual data you deliver is the only reliable method.
Text compression algorithms fall into two broad categories. Huffman coding assigns variable-length bit sequences based on character frequency, representing common characters with shorter sequences. LZ77/LZ78-family algorithms detect repeated patterns in the text and replace them with references to earlier occurrences (position and length). gzip uses the DEFLATE algorithm, which combines both approaches.
The relationship between character count and compression ratio has interesting properties. Two texts with the same character count can have vastly different compression ratios depending on their content. A string of the same character repeated (such as "aaaaaaaaaa") compresses extremely well, while a random character string is nearly incompressible. This connects directly to the concept of entropy in information theory: the more redundant the text, the higher the compression ratio.
In practice, compression ratio translates directly into storage cost and bandwidth optimization. Log files and chat histories repeat the same format over and over, which makes them a particularly compressible kind of text and a natural target for reducing storage costs. Compressing API responses shortens response times on mobile networks and improves the user experience. The longer the text, the greater the benefit of compression, so for long-form content delivery, whether compression is enabled can make a noticeable difference in performance.