Last updated:

Steganography - The Art of Hiding Secret Messages in Text and Character Count

10 min read

There's a secret message hidden in this paragraph - if someone told you that, where would you look? The first letter of each sentence? The spacing between certain characters? Or perhaps invisible characters embedded within? Steganography is the art of hiding the very existence of a message. If encryption makes a message "unreadable," steganography makes it "undetectable." And this technique can sometimes be detected through the simple act of character counting.

The Ancient Art of Hiding

The history of steganography dates back to 5th century BC Greece. Two episodes that the historian Herodotus set down in his Histories are regarded as the oldest surviving records. In the first, Histiaeus shaved the head of his most trusted servant, marked a message into the scalp, waited for the hair to grow back, and sent the man off to Aristagoras. The messenger himself could not read what he was carrying, and the recipient recovered the text from a single instruction: shave his head when he arrives. In the second, Demaratus wanted to warn of a planned attack on Greece, so he wrote directly on the wooden base of a wax writing tablet and then coated the wood with fresh wax. Tablets of that era were reusable writing tools whose wax was melted and smoothed over again, so as long as the wax surface was blank, nobody thought to look underneath.

The two methods hide in opposite ways. The tattooed scalp cannot be dispatched until the hair has grown, but an inspection along the route turns up nothing. The wax tablet can be sent at once, yet scraping the wax away exposes everything in one stroke. That trade-off between being hard to detect and being quick to deliver runs through the modern digital techniques described below in exactly the same shape.

Later centuries added more ways of hiding things physically: invisible inks (lemon juice, milk, urine and the like), and knitting Morse code into a length of yarn worn by the carrier. In the 20th century came the microdot, which shrank a document photographically. Reduced to a circular image roughly one millimetre across, it was indistinguishable from a printed period or the dot over a lowercase i, and could travel through the ordinary mail as it was. The method was first used in Germany between the world wars, and other countries later adopted it as a way of getting material past censored postal routes.

Text-Based Steganography Techniques

Digital-age text steganography employs several representative techniques.

Acrostics - Messages Hidden in Initial Letters

An acrostic is a technique where connecting the first letters of each line or sentence reveals a secret message. It's the most classical form of text steganography, used in poetry and lyrics since ancient times.

One case that genuinely caused trouble was the veto message California Governor Arnold Schwarzenegger issued in October 2009 on a bill introduced by state legislator Tom Ammiano. Read downward, the first letters of lines three through nine of the message spelled out an insult. The governor's office said the arrangement was a coincidence, but mathematicians argued back that coincidence was statistically hard to credit. What the episode really shows is that the awkward part of an acrostic is not how cleverly it is embedded, but that nobody except the author can settle whether it was accident or intent.

An acrostic embeds a message without increasing the character count, but the amount of information it can carry is tied to the number of lines. Ten lines of prose hide only ten letters, and every line has to open with a prescribed letter, which makes the writing read awkwardly. Add the fact that anyone who thinks to read down the initials will find it, and the secrecy on offer is low compared with the modern techniques.

Whitespace Manipulation

This technique embeds bit information by manipulating the number of spaces between words. One space represents "0" and two spaces represent "1," encoding binary data. The subtle difference in spacing is hard for humans to notice, but a character counting tool can detect that "there are too many spaces relative to the visible word count."

Zero-Width Character Steganography - The World of Invisible Characters

The most powerful modern text steganography technique uses zero-width invisible characters. Unicode defines several "zero-width characters" that don't display on screen but exist as character data.

Unicode Code PointNameOriginal PurposeSteganographic Role
U+200BZero Width SpaceSpecifying line break opportunitiesRepresents bit "0"
U+200CZero Width Non-JoinerSuppressing ligaturesRepresents bit "1"
U+200DZero Width JoinerPromoting ligaturesAdditional bit value
U+FEFFZero Width No-Break Space (BOM)Byte order markDelimiter character

Using two types of zero-width characters, U+200B and U+200C, you can represent 1 bit with 2 values (0 and 1). Eight zero-width characters make 1 byte, meaning 1 ASCII character. Hiding the 5-character message "Hello" requires 40 zero-width characters.

Distributing these 40 zero-width characters between words in normal text makes the appearance completely unchanged. However, comparing "visible character count" with "actual code point count" using a character counting tool reveals an unnatural discrepancy. Each zero-width character occupies 3 bytes in UTF-8, so the same 40 characters also surface as 120 extra bytes. Understanding Unicode fundamentals helps identify that zero-width characters cause this discrepancy.

Zero-Width Character Steganography Implementation

Let's look at a concrete embedding process. Consider hiding the secret message "Hi" in the normal text "Good morning."

"H" has ASCII code 72, binary 01001000. "i" is 105, binary 01101001. Converting 0 to U+200B (zero-width space) and 1 to U+200C (zero-width non-joiner) generates a string of 16 zero-width characters.

Inserting these 16 zero-width characters between "Good" and "morning" leaves the appearance as "Good morning," but the actual data contains 16 invisible characters. A text editor counts 12 characters, but programmatically counting Unicode code points yields 28. The difference of 16 characters is the hidden message.

More advanced implementations use 3 or more types of zero-width characters for ternary or higher encoding, representing the same message with fewer zero-width characters. Using U+200B, U+200C, and U+200D carries log₂3 ≒ 1.58 bits per character, so 8 bits of information fits into 5.05 characters in theory. An implementation, however, has to round that up. Five characters drawn from three symbols yield only 3 to the 5th power = 243 combinations, which cannot cover the 256 values of a byte, so real encoders use 6 characters (3 to the 6th power = 729 combinations). That is 6 characters instead of the 8 the binary scheme needs, a reduction of 25 percent.

Homoglyph Attacks - Different Characters That Look Identical

Homoglyphs are characters that look nearly identical but have different Unicode code points. For example, Latin "a" (U+0061) and Cyrillic "а" (U+0430) appear completely identical in many fonts.

Latin CharacterCode PointCyrillic CharacterCode PointVisual Difference
aU+0061аU+0430Nearly identical
eU+0065еU+0435Nearly identical
oU+006FоU+043ENearly identical
pU+0070рU+0440Nearly identical
cU+0063сU+0441Nearly identical

Homoglyph attacks exploit this property. Replacing the "a" in "apple.com" with Cyrillic "а" in a phishing URL looks identical but redirects to a completely different domain. In steganography, replacing specific characters with homoglyphs embeds bit information.

Detecting homoglyphs requires checking the Unicode code point of each character. As discussed in password length and security, cases where appearance is identical but byte sequences differ pose serious security risks.

As a countermeasure, major browsers restrict IDN (Internationalized Domain Name) display. When domain names mix multiple scripts (Latin and Cyrillic, etc.), browsers display the domain in Punycode (encoded format starting with xn--) to warn users of fake sites. Firefox has applied this kind of check since version 22 and Chrome since version 51, while Safari renders problematic character sets in Punycode.

Text Watermarking Technology

As an application of steganography, text digital watermarking technology exists. While image and video watermarks are widely known, techniques for embedding watermarks in text also exist.

Watermark MethodPrincipleDetection MethodResilience
Zero-width character embeddingStores bit information in invisible charactersCharacter countingMay be lost on copy-paste
Synonym substitutionSubstitutes synonyms like "big" → "large"Comparison with originalResilient to text editing
Syntactic transformationTransforms active → passive voiceComparison with originalResilient to text editing
Whitespace manipulationManipulates space and tab countsStatistical analysis of whitespaceLost on format changes

Synonym substitution watermarking embeds bit information without changing text meaning. For example, substituting "big" with "large" represents 1 bit of information. This method may change character count but is resilient to text editing and copy-paste.

Encryption vs. Steganography

Encryption and steganography are often confused but are fundamentally different technologies.

PropertyEncryptionSteganography
PurposeMake message content unreadableHide message existence
DetectabilityCiphertext existence is obviousMessage existence itself is unknown
Effect on character countSimilar to original textCover text character count may increase
Key requirementKey needed for decryptionMay be extractable with technique knowledge
CombinationCan be used aloneCombined with encryption, content stays protected even after discovery

The safest approach is to encrypt a message and then hide it with steganography. Even if the steganography is broken and the message's existence is discovered, the content remains unreadable if encrypted.

Detecting Steganography Through Character Counting

The simplest method for detecting text-based steganography is character counting. The following unnatural discrepancies serve as detection hints.

Mismatch between visible character count and actual character count (code point count). When zero-width characters are embedded, the character count visible in a text editor is less than the programmatic count. For example, if text that appears to be 100 characters actually contains 180 characters of data, 80 zero-width characters may be embedded.

Unnatural character encoding sizes also provide clues. Pure ASCII text (alphanumeric only) should be 1 character = 1 byte in UTF-8. However, if Cyrillic homoglyphs are mixed in, some characters that look ASCII become 2 bytes. If the total byte count exceeds the character count, homoglyph presence should be suspected.

Character Counting and Zero-Width Characters on Twitter (now X)

"twitter-text," the character counting library used by Twitter (now X), assigns a weight to every character and decides whether a post is permitted by whether the weights add up to no more than the limit of 280. In its configuration the default weight is 2, and only a few ranges, such as U+0000 to U+10FF and U+2000 to U+200D, are defined with a weight of 1. Zero-width space (U+200B), zero-width non-joiner (U+200C) and zero-width joiner (U+200D) all fall inside that weight-1 range, so each one consumes a character of the allowance even though nothing appears on screen. U+FEFF sits outside the range and is counted at the default weight of 2.

A post that hides information in zero-width characters is therefore always "longer" than it looks. The binary scheme needs 8 zero-width characters per character of the secret message, so pouring the entire allowance of 280 into it carries only 35 characters, and subtracting the cover text leaves less again. The character limit itself is what caps the capacity of zero-width steganography.

Steganography Detection Tools and Techniques

Several specialized tools and techniques exist for detecting text-based steganography.

Detection MethodTargetPrincipleLimitation
Character count vs. byte count comparisonZero-width charactersMismatch between visible count and actual bytesDifficult to distinguish from legitimate zero-width chars
Unicode category analysisHomoglyphsVerify Unicode block consistency of charactersMany false positives in multilingual text
Statistical analysisWhitespace manipulationVerify space distribution matches natural language statisticsLow accuracy for short texts
Entropy analysisGeneralVerify text information entropy is within natural language rangeDifficult against advanced techniques

The simplest and most effective detection method is to copy-paste text to plain text and compare byte counts with the original. If zero-width characters or homoglyphs are present, byte counts will differ. If a character counting tool displays both "visible character count" and "Unicode code point count," confirming the discrepancy is a matter of reading the two numbers side by side.

The Practical Risks Zero-Width Characters Introduce

Zero-width characters cause trouble as accidental contamination long before they cause trouble as deliberate secret communication. Text copied from a web page carries a stray zero-width character, and then a search term fails to match, an identifier fails a comparison, or a form field rejects the input for exceeding its character limit. Since the text looks exactly the same, no amount of staring at it leads to the cause. What decides the outcome is whether you notice that the "visible character count" and the code point count disagree.

The same property can be turned to tracking. If each distributed copy of a document is given a different sequence of zero-width characters, the copy carries a mark identifying its recipient without any change to how it looks. That is the idea behind document fingerprinting: read the sequence out of a leaked document and you can narrow down which copy it came from. Zero-width characters can be stripped along copy-paste routes or by processing that normalizes text, though, so how far the mark survives depends on the route it travels.

Steganography itself leans neither one way nor the other. It can hide the existence of a communication in order to protect the sender, and it can erase the traces left by someone carrying information out. Viewed from either side, the plain check of setting the visible character count against the actual code point and byte counts is the first handhold on whatever has been hidden.

Share this article