Last updated:
Steganography - The Art of Hiding Secret Messages in Text and Character Count
There's a secret message hidden in this paragraph - if someone told you that, where would you look? The first letter of each sentence? The spacing between certain characters? Or perhaps invisible characters embedded within? Steganography is the art of hiding the very existence of a message. If encryption makes a message "unreadable," steganography makes it "undetectable." And this technique can sometimes be detected through the simple act of character counting.
The Ancient Art of Hiding
The history of steganography dates back to 5th century BC Greece. Two episodes that the historian Herodotus set down in his Histories are regarded as the oldest surviving records. In the first, Histiaeus shaved the head of his most trusted servant, marked a message into the scalp, waited for the hair to grow back, and sent the man off to Aristagoras. The messenger himself could not read what he was carrying, and the recipient recovered the text from a single instruction: shave his head when he arrives. In the second, Demaratus wanted to warn of a planned attack on Greece, so he wrote directly on the wooden base of a wax writing tablet and then coated the wood with fresh wax. Tablets of that era were reusable writing tools whose wax was melted and smoothed over again, so as long as the wax surface was blank, nobody thought to look underneath.
The two methods hide in opposite ways. The tattooed scalp cannot be dispatched until the hair has grown, but an inspection along the route turns up nothing. The wax tablet can be sent at once, yet scraping the wax away exposes everything in one stroke. That trade-off between being hard to detect and being quick to deliver runs through the modern digital techniques described below in exactly the same shape.
Later centuries added more ways of hiding things physically: invisible inks (lemon juice, milk, urine and the like), and knitting Morse code into a length of yarn worn by the carrier. In the 20th century came the microdot, which shrank a document photographically. Reduced to a circular image roughly one millimetre across, it was indistinguishable from a printed period or the dot over a lowercase i, and could travel through the ordinary mail as it was. The method was first used in Germany between the world wars, and other countries later adopted it as a way of getting material past censored postal routes.
Text-Based Steganography Techniques
Digital-age text steganography employs several representative techniques.
Acrostics - Messages Hidden in Initial Letters
An acrostic is a technique where connecting the first letters of each line or sentence reveals a secret message. It's the most classical form of text steganography, used in poetry and lyrics since ancient times.
One case that genuinely caused trouble was the veto message California Governor Arnold Schwarzenegger issued in October 2009 on a bill introduced by state legislator Tom Ammiano. Read downward, the first letters of lines three through nine of the message spelled out an insult. The governor's office said the arrangement was a coincidence, but mathematicians argued back that coincidence was statistically hard to credit. What the episode really shows is that the awkward part of an acrostic is not how cleverly it is embedded, but that nobody except the author can settle whether it was accident or intent.
An acrostic embeds a message without increasing the character count, but the amount of information it can carry is tied to the number of lines. Ten lines of prose hide only ten letters, and every line has to open with a prescribed letter, which makes the writing read awkwardly. Add the fact that anyone who thinks to read down the initials will find it, and the secrecy on offer is low compared with the modern techniques.
Whitespace Manipulation
This technique embeds bit information by manipulating the number of spaces between words. One space represents "0" and two spaces represent "1," encoding binary data. The subtle difference in spacing is hard for humans to notice, but a character counting tool can detect that "there are too many spaces relative to the visible word count."
Zero-Width Character Steganography - The World of Invisible Characters
The most powerful modern text steganography technique uses zero-width invisible characters. Unicode defines several "zero-width characters" that don't display on screen but exist as character data.
| Unicode Code Point | Name | Original Purpose | Steganographic Role |
|---|---|---|---|
| U+200B | Zero Width Space | Specifying line break opportunities | Represents bit "0" |
| U+200C | Zero Width Non-Joiner | Suppressing ligatures | Represents bit "1" |
| U+200D | Zero Width Joiner | Promoting ligatures | Additional bit value |
| U+FEFF | Zero Width No-Break Space (BOM) | Byte order mark | Delimiter character |
Using two types of zero-width characters, U+200B and U+200C, you can represent 1 bit with 2 values (0 and 1). Eight zero-width characters make 1 byte, meaning 1 ASCII character. Hiding the 5-character message "Hello" requires 40 zero-width characters.
Distributing these 40 zero-width characters between words in normal text makes the appearance completely unchanged. However, comparing "visible character count" with "actual code point count" using a character counting tool reveals an unnatural discrepancy. Each zero-width character occupies 3 bytes in UTF-8, so the same 40 characters also surface as 120 extra bytes. Understanding Unicode fundamentals helps identify that zero-width characters cause this discrepancy.
Zero-Width Character Steganography Implementation
Let's look at a concrete embedding process. Consider hiding the secret message "Hi" in the normal text "Good morning."
"H" has ASCII code 72, binary 01001000. "i" is 105, binary 01101001. Converting 0 to U+200B (zero-width space) and 1 to U+200C (zero-width non-joiner) generates a string of 16 zero-width characters.
Inserting these 16 zero-width characters between "Good" and "morning" leaves the appearance as "Good morning," but the actual data contains 16 invisible characters. A text editor counts 12 characters, but programmatically counting Unicode code points yields 28. The difference of 16 characters is the hidden message.
More advanced implementations use 3 or more types of zero-width characters for ternary or higher encoding, representing the same message with fewer zero-width characters. Using U+200B, U+200C, and U+200D carries log₂3 ≒ 1.58 bits per character, so 8 bits of information fits into 5.05 characters in theory. An implementation, however, has to round that up. Five characters drawn from three symbols yield only 3 to the 5th power = 243 combinations, which cannot cover the 256 values of a byte, so real encoders use 6 characters (3 to the 6th power = 729 combinations). That is 6 characters instead of the 8 the binary scheme needs, a reduction of 25 percent.
Homoglyph Attacks - Different Characters That Look Identical
Homoglyphs are characters that look nearly identical but have different Unicode code points. For example, Latin "a" (U+0061) and Cyrillic "а" (U+0430) appear completely identical in many fonts.
| Latin Character | Code Point | Cyrillic Character | Code Point | Visual Difference |
|---|---|---|---|---|
| a | U+0061 | а | U+0430 | Nearly identical |
| e | U+0065 | е | U+0435 | Nearly identical |
| o | U+006F | о | U+043E | Nearly identical |
| p | U+0070 | р | U+0440 | Nearly identical |
| c | U+0063 | с | U+0441 | Nearly identical |
Homoglyph attacks exploit this property. Replacing the "a" in "apple.com" with Cyrillic "а" in a phishing URL looks identical but redirects to a completely different domain. In steganography, replacing specific characters with homoglyphs embeds bit information.
Detecting homoglyphs requires checking the Unicode code point of each character. As discussed in password length and security, cases where appearance is identical but byte sequences differ pose serious security risks.
As a countermeasure, major browsers restrict IDN (Internationalized Domain Name) display. When domain names mix multiple scripts (Latin and Cyrillic, etc.), browsers display the domain in Punycode (encoded format starting with xn--) to warn users of fake sites. Firefox has applied this kind of check since version 22 and Chrome since version 51, while Safari renders problematic character sets in Punycode.
Text Watermarking Technology
As an application of steganography, text digital watermarking technology exists. While image and video watermarks are widely known, techniques for embedding watermarks in text also exist.
| Watermark Method | Principle | Detection Method | Resilience |
|---|---|---|---|
| Zero-width character embedding | Stores bit information in invisible characters | Character counting | May be lost on copy-paste |
| Synonym substitution | Substitutes synonyms like "big" → "large" | Comparison with original | Resilient to text editing |
| Syntactic transformation | Transforms active → passive voice | Comparison with original | Resilient to text editing |
| Whitespace manipulation | Manipulates space and tab counts | Statistical analysis of whitespace | Lost on format changes |
Synonym substitution watermarking embeds bit information without changing text meaning. For example, substituting "big" with "large" represents 1 bit of information. This method may change character count but is resilient to text editing and copy-paste.
Encryption vs. Steganography
Encryption and steganography are often confused but are fundamentally different technologies.
| Property | Encryption | Steganography |
|---|---|---|
| Purpose | Make message content unreadable | Hide message existence |
| Detectability | Ciphertext existence is obvious | Message existence itself is unknown |
| Effect on character count | Similar to original text | Cover text character count may increase |
| Key requirement | Key needed for decryption | May be extractable with technique knowledge |
| Combination | Can be used alone | Combined with encryption, content stays protected even after discovery |
The safest approach is to encrypt a message and then hide it with steganography. Even if the steganography is broken and the message's existence is discovered, the content remains unreadable if encrypted.
Detecting Steganography Through Character Counting
The simplest method for detecting text-based steganography is character counting. The following unnatural discrepancies serve as detection hints.
Mismatch between visible character count and actual character count (code point count). When zero-width characters are embedded, the character count visible in a text editor is less than the programmatic count. For example, if text that appears to be 100 characters actually contains 180 characters of data, 80 zero-width characters may be embedded.
Unnatural character encoding sizes also provide clues. Pure ASCII text (alphanumeric only) should be 1 character = 1 byte in UTF-8. However, if Cyrillic homoglyphs are mixed in, some characters that look ASCII become 2 bytes. If the total byte count exceeds the character count, homoglyph presence should be suspected.
Character Counting and Zero-Width Characters on Twitter (now X)
"twitter-text," the character counting library used by Twitter (now X), assigns a weight to every character and decides whether a post is permitted by whether the weights add up to no more than the limit of 280. In its configuration the default weight is 2, and only a few ranges, such as U+0000 to U+10FF and U+2000 to U+200D, are defined with a weight of 1. Zero-width space (U+200B), zero-width non-joiner (U+200C) and zero-width joiner (U+200D) all fall inside that weight-1 range, so each one consumes a character of the allowance even though nothing appears on screen. U+FEFF sits outside the range and is counted at the default weight of 2.
A post that hides information in zero-width characters is therefore always "longer" than it looks. The binary scheme needs 8 zero-width characters per character of the secret message, so pouring the entire allowance of 280 into it carries only 35 characters, and subtracting the cover text leaves less again. The character limit itself is what caps the capacity of zero-width steganography.
Steganography Detection Tools and Techniques
Several specialized tools and techniques exist for detecting text-based steganography.
| Detection Method | Target | Principle | Limitation |
|---|---|---|---|
| Character count vs. byte count comparison | Zero-width characters | Mismatch between visible count and actual bytes | Difficult to distinguish from legitimate zero-width chars |
| Unicode category analysis | Homoglyphs | Verify Unicode block consistency of characters | Many false positives in multilingual text |
| Statistical analysis | Whitespace manipulation | Verify space distribution matches natural language statistics | Low accuracy for short texts |
| Entropy analysis | General | Verify text information entropy is within natural language range | Difficult against advanced techniques |
The simplest and most effective detection method is to copy-paste text to plain text and compare byte counts with the original. If zero-width characters or homoglyphs are present, byte counts will differ. If a character counting tool displays both "visible character count" and "Unicode code point count," confirming the discrepancy is a matter of reading the two numbers side by side.
The Practical Risks Zero-Width Characters Introduce
Zero-width characters cause trouble as accidental contamination long before they cause trouble as deliberate secret communication. Text copied from a web page carries a stray zero-width character, and then a search term fails to match, an identifier fails a comparison, or a form field rejects the input for exceeding its character limit. Since the text looks exactly the same, no amount of staring at it leads to the cause. What decides the outcome is whether you notice that the "visible character count" and the code point count disagree.
The same property can be turned to tracking. If each distributed copy of a document is given a different sequence of zero-width characters, the copy carries a mark identifying its recipient without any change to how it looks. That is the idea behind document fingerprinting: read the sequence out of a leaked document and you can narrow down which copy it came from. Zero-width characters can be stripped along copy-paste routes or by processing that normalizes text, though, so how far the mark survives depends on the route it travels.
Steganography itself leans neither one way nor the other. It can hide the existence of a communication in order to protect the sender, and it can erase the traces left by someone carrying information out. Viewed from either side, the plain check of setting the visible character count against the actual code point and byte counts is the first handhold on whatever has been hidden.