Last updated:

The World of Invisible Characters - Troubles Caused by Zero-Width and Invisible Characters

12 min read

Your string should be 10 characters, but the system insists it's 12. No matter how hard you look, you can't see any extra characters. The culprit is "zero-width characters" - invisible characters that don't appear on screen at all, yet undeniably exist as data. This article explains the types and purposes of invisible characters defined in Unicode, their impact on character counting, and the typical patterns of what breaks once they get in, together with how to deal with them.

Invisible Character Catalog - Characters That Exist Without Being Seen

Unicode defines multiple characters that are not displayed on screen (or have zero width). These are not "bugs" - they exist for legitimate reasons in text processing.

Character NameCode PointPurposeCharacter CountDisplay Width
Zero Width Space (ZWSP)U+200BSpecifying line break opportunitiesCounted as 1 character0
Zero Width Joiner (ZWJ)U+200DJoining characters (emoji composition)Counted as 1 character0
Zero Width Non-Joiner (ZWNJ)U+200CPreventing character joiningCounted as 1 character0
Left-to-Right Mark (LRM)U+200EText direction controlCounted as 1 character0
Right-to-Left Mark (RLM)U+200FText direction controlCounted as 1 character0
Byte Order Mark (BOM)U+FEFFEncoding identificationCounted as 1 character (disappears in runtimes that strip it on read)0
Soft Hyphen (SHY)U+00ADSpecifying hyphenation pointsCounted as 1 characterUsually 0 (shown only at line breaks)
Word Joiner (WJ)U+2060Specifying no-break positionsCounted as 1 character0

All of these characters serve legitimate roles in text processing. The problem is that when they unintentionally infiltrate text, they silently throw off character counts.

Zero Width Space (U+200B) - The Most Troublesome Invisible Character

The Zero Width Space (ZWSP) is a character that embeds "you may break the line here" information into text. It's used in languages like Thai and Khmer that don't use spaces between words, allowing browsers to break lines at appropriate positions.

However, ZWSP easily infiltrates text when copying and pasting from web pages, causing troubles like:

Password infiltration is particularly serious. When ZWSP sneaks into a password copied from a website, you get a situation where the password looks correct but login fails. When considering password length and security, the existence of invisible characters cannot be ignored.

Zero Width Joiner (U+200D) - The Magic Character That Composes Emoji

The Zero Width Joiner (ZWJ) plays the most positive role among invisible characters. As explained in detail in emoji character counting, ZWJ combines multiple emoji to create new ones.

Displayed EmojiComponentsCode Point CountCharacter Count (JavaScript)
👨‍👩‍👧‍👦 (Family)👨 + ZWJ + 👩 + ZWJ + 👧 + ZWJ + 👦711 (including surrogate pairs)
👩‍💻 (Woman Technologist)👩 + ZWJ + 💻35
🏳️‍🌈 (Rainbow Flag)🏳️ + ZWJ + 🌈46
👨‍🍳 (Man Cook)👨 + ZWJ + 🍳35

The family emoji 👨‍👩‍👧‍👦 looks like a single emoji, but internally consists of 4 emoji and 3 ZWJs. JavaScript's .length property returns 11. On social media with character limits, a single emoji like this can consume a large number of characters.

Direction Control Characters - Mechanisms for Right-to-Left Languages

Arabic and Hebrew are languages written right-to-left (RTL). In text where these languages coexist with English (left-to-right, LTR), invisible characters that control text direction are necessary.

U+200E (Left-to-Right Mark) and U+200F (Right-to-Left Mark) are characters for explicitly specifying text direction. When these unintentionally infiltrate text, they can disrupt display order or throw off character counts.

In 2021, a security vulnerability called "Trojan Source" was reported that exploits direction control characters. By embedding direction control characters in source code, code that looks normal to human eyes is interpreted as different logic by the compiler. This vulnerability demonstrated that invisible characters can also pose security risks.

BOM (U+FEFF) - The Invisible Character Lurking at File Beginnings

The Byte Order Mark (BOM) is a character added at the beginning of text files to identify encoding. The UTF-8 BOM is 3 bytes (EF BB BF) and is sometimes added by Windows Notepad when saving files.

BOM is ignored by many programs, but causes problems in these cases:

Steganography Using Zero-Width Characters (Watermarking Technology)

Steganography (digital watermarking) is a technology that turns the "invisible" property of invisible characters on its head. By embedding patterns of zero-width characters in text, hidden information can be embedded without changing the appearance.

There is really only one technique at work here, and the characters it uses are limited to zero-width characters such as U+200B, U+200C, U+200D and U+FEFF. The only thing that differs from one use to the next is what the embedded bit sequence is read as.

PurposeEmbedded ContentOperation Needed to Read It
Passing a hidden messageAn arbitrary bit sequenceThe receiver decodes it with the same mapping table
Identifying a leak source (a watermark per recipient)A bit sequence that identifies the recipientMatch it against the pattern recorded at distribution time
Detecting unauthorized copyingA bit sequence that indicates the originInspect the copied text code point by code point

For example, by treating 4 types of zero-width characters as 2-bit information (U+200B = 00, U+200C = 01, U+200D = 10, U+FEFF = 11) and inserting zero-width characters between each word in text, binary data can be hidden within it.

This technology is sometimes used by companies to identify the source of confidential document leaks. By embedding different zero-width character patterns for each recipient, when a document leaks externally, the source can be identified.

Detecting and Removing Invisible Characters

To correctly process text infiltrated by invisible characters, you need to know detection and removal methods.

MethodTargetCode Example
JavaScript regexMajor zero-width charactersstr.replace(/[\u200B-\u200F\u2028-\u202F\uFEFF]/g, '')
Python regexSame as abovere.sub(r'[\u200b-\u200f\u2028-\u202f\ufeff]', '', text)
Text editorInvisible and confusable charactersVS Code: editor.unicodeHighlight.invisibleCharacters
Command lineInvisible characters in filescat -v filename or xxd filename
PHPMajor zero-width characterspreg_replace('/[\x{200B}-\x{200F}\x{FEFF}]/u', '', $str)

What the JavaScript regex /[\u200B-\u200F\u2028-\u202F\uFEFF]/g covers is limited to the "most common zero-width characters." Run text through it and ZWSP (U+200B) and the BOM (U+FEFF) do disappear, but the word joiner (U+2060) and the soft hyphen (U+00AD) from the table at the top of this article, along with the isolation-type direction control characters covered later (U+2066 to U+2069), fall outside the range and stay where they are. Cleaning form input on the server side is a sound measure, but you have to count up the characters you actually want to remove and write them into the character class yourself.

However, unconditionally removing all invisible characters is dangerous. ZWJ is necessary for emoji composition, and removing it will decompose emoji. ZWNJ is essential for correct rendering in Persian and Hindi. Invisible character removal must be done carefully with understanding of purpose and context.

Invisible Character Handling by Programming Language

What happens when a ZWSP finds its way into source code depends on the language, and on the version of the toolchain. Some languages stop with an error; others quietly accept the character as part of an identifier, and that second group is far more troublesome. The table below shows representative behavior as of August 2026. Because it can change between compilers and versions, the only conclusive check is to try it on your own toolchain.

LanguageZWSP inside an IdentifierZWSP in String LiteralsDetection Clue
JavaScript (Node.js)Syntax error (not a character allowed in identifiers)Retained as part of stringESLint's no-irregular-whitespace
PythonSyntaxError (rejected as a non-printable character)Retained as part of stringThe interpreter itself refuses to run the file
RustCompile error (cannot be tokenized)Retained as part of stringuncommon_codepoints warning for ZWNJ and ZWJ
C / C++ (clang)Accepted as an identifier, warning onlyRetained as part of stringclang's -Wunicode-zero-width

The easiest thing to get backwards here is the difference between ZWSP and ZWNJ or ZWJ. The set of characters JavaScript allows inside an identifier includes ZWNJ (U+200C) and ZWJ (U+200D), and does not include ZWSP (U+200B). So var he\u200Bllo stops with a syntax error, while var he\u200Cllo passes without complaint and declares a variable that is distinct from the identical-looking hello.

In practice the two cases mean opposite things. A ZWSP is invisible in the editor but brings the program down the moment it runs, so you always find out that it is there. The dangerous one is the pair that passes: the declaration and its references both succeed, and a later hello typed by hand becomes an undefined variable. The character stays invisible, and all that is left to see is a name that does not match.

When considering variable and function name length guidelines, the risk of invisible character infiltration should be kept in mind. Since they cannot be detected visually in code review, it's important to set up automatic detection through linters and editor settings.

Typical Failure Patterns Caused by Invisible Characters

The ways these characters get in, and the ways things break afterwards, follow a small number of patterns. What they share is that the cause is not visible, so narrowing it down takes time.

Character Count Tools and Invisible Characters

How character count tools handle invisible characters varies by tool. Some ignore invisible characters when counting, while others count them as-is. Without understanding Unicode basics, you can't identify why character counts differ between tools.

If you want to accurately count text characters, we recommend first checking for invisible characters and removing them if necessary before counting. Simply knowing that "invisible characters" exist can prevent many character count-related troubles.

Invisible Characters and Security - Unseen Threats

Trojan Source, mentioned earlier, exploits the nine invisible characters used to control bidirectional text (U+202A to U+202E for embedding and override, and U+2066 to U+2069 for isolation) to create a gap between how source code looks and how it actually executes. Other ways in which invisible characters turn into a security problem are known as well.

Attack MethodInvisible Characters UsedImpactCountermeasure
Trojan SourceDirection control chars (U+202A-U+202E, U+2066-U+2069)Malicious logic undetectable in code reviewEnable compiler warnings
Homograph attackVisually identical different chars (U+0430 vs U+0061)Phishing URL spoofingCheck Punycode display
ZWSP injectionU+200BBypassing input validationServer-side invisible character removal
BOM injectionU+FEFFFile parser malfunctionAutomatic BOM removal processing

The mechanism that makes the attack work is simple. Once direction control characters are planted inside a comment or a string literal, editors and web diff views reorder the characters for display exactly as instructed, while the compiler reads the byte sequence in its original order. That is enough to produce a gap where code that reads on screen as "check the permission, then run the operation" actually runs the operation without checking. What is being abused is not an invalid character but the legitimate mechanism that exists to display bidirectional text correctly.

This attack is particularly dangerous because it neutralizes code review - the human-eye verification process. Countermeasures include enabling compiler and linter settings that warn about direction control character usage, and incorporating invisible character detection steps into CI/CD pipelines.

As mentioned in the Git commit message writing article, utilizing linters is essential for code quality management. Detecting invisible characters is one of the important roles of linters.

Share this article