Last updated:
The World of Invisible Characters - Troubles Caused by Zero-Width and Invisible Characters
Your string should be 10 characters, but the system insists it's 12. No matter how hard you look, you can't see any extra characters. The culprit is "zero-width characters" - invisible characters that don't appear on screen at all, yet undeniably exist as data. This article explains the types and purposes of invisible characters defined in Unicode, their impact on character counting, and the typical patterns of what breaks once they get in, together with how to deal with them.
Invisible Character Catalog - Characters That Exist Without Being Seen
Unicode defines multiple characters that are not displayed on screen (or have zero width). These are not "bugs" - they exist for legitimate reasons in text processing.
| Character Name | Code Point | Purpose | Character Count | Display Width |
|---|---|---|---|---|
| Zero Width Space (ZWSP) | U+200B | Specifying line break opportunities | Counted as 1 character | 0 |
| Zero Width Joiner (ZWJ) | U+200D | Joining characters (emoji composition) | Counted as 1 character | 0 |
| Zero Width Non-Joiner (ZWNJ) | U+200C | Preventing character joining | Counted as 1 character | 0 |
| Left-to-Right Mark (LRM) | U+200E | Text direction control | Counted as 1 character | 0 |
| Right-to-Left Mark (RLM) | U+200F | Text direction control | Counted as 1 character | 0 |
| Byte Order Mark (BOM) | U+FEFF | Encoding identification | Counted as 1 character (disappears in runtimes that strip it on read) | 0 |
| Soft Hyphen (SHY) | U+00AD | Specifying hyphenation points | Counted as 1 character | Usually 0 (shown only at line breaks) |
| Word Joiner (WJ) | U+2060 | Specifying no-break positions | Counted as 1 character | 0 |
All of these characters serve legitimate roles in text processing. The problem is that when they unintentionally infiltrate text, they silently throw off character counts.
Zero Width Space (U+200B) - The Most Troublesome Invisible Character
The Zero Width Space (ZWSP) is a character that embeds "you may break the line here" information into text. It's used in languages like Thai and Khmer that don't use spaces between words, allowing browsers to break lines at appropriate positions.
However, ZWSP easily infiltrates text when copying and pasting from web pages, causing troubles like:
- Form input judged as "exceeding character limit" (looks within limit visually)
- Password copy-paste failures (ZWSP infiltrates making it a different string)
- Search mismatches (identical-looking strings don't match in search)
- CSV file data not parsing correctly
- Program source code infiltration causing compile errors
Password infiltration is particularly serious. When ZWSP sneaks into a password copied from a website, you get a situation where the password looks correct but login fails. When considering password length and security, the existence of invisible characters cannot be ignored.
Zero Width Joiner (U+200D) - The Magic Character That Composes Emoji
The Zero Width Joiner (ZWJ) plays the most positive role among invisible characters. As explained in detail in emoji character counting, ZWJ combines multiple emoji to create new ones.
| Displayed Emoji | Components | Code Point Count | Character Count (JavaScript) |
|---|---|---|---|
| 👨👩👧👦 (Family) | 👨 + ZWJ + 👩 + ZWJ + 👧 + ZWJ + 👦 | 7 | 11 (including surrogate pairs) |
| 👩💻 (Woman Technologist) | 👩 + ZWJ + 💻 | 3 | 5 |
| 🏳️🌈 (Rainbow Flag) | 🏳️ + ZWJ + 🌈 | 4 | 6 |
| 👨🍳 (Man Cook) | 👨 + ZWJ + 🍳 | 3 | 5 |
The family emoji 👨👩👧👦 looks like a single emoji, but internally consists of 4 emoji and 3 ZWJs. JavaScript's .length property returns 11. On social media with character limits, a single emoji like this can consume a large number of characters.
Direction Control Characters - Mechanisms for Right-to-Left Languages
Arabic and Hebrew are languages written right-to-left (RTL). In text where these languages coexist with English (left-to-right, LTR), invisible characters that control text direction are necessary.
U+200E (Left-to-Right Mark) and U+200F (Right-to-Left Mark) are characters for explicitly specifying text direction. When these unintentionally infiltrate text, they can disrupt display order or throw off character counts.
In 2021, a security vulnerability called "Trojan Source" was reported that exploits direction control characters. By embedding direction control characters in source code, code that looks normal to human eyes is interpreted as different logic by the compiler. This vulnerability demonstrated that invisible characters can also pose security risks.
BOM (U+FEFF) - The Invisible Character Lurking at File Beginnings
The Byte Order Mark (BOM) is a character added at the beginning of text files to identify encoding. The UTF-8 BOM is 3 bytes (EF BB BF) and is sometimes added by Windows Notepad when saving files.
BOM is ignored by many programs, but causes problems in these cases:
- BOM at the beginning of PHP files prevents
header()from working (output is judged to have already started) - BOM at the beginning of CSV files prevents the first column name from being recognized correctly
- BOM in JSON files may cause parser errors
- BOM at the beginning of shell scripts prevents the shebang (
#!/bin/bash) from being recognized
Steganography Using Zero-Width Characters (Watermarking Technology)
Steganography (digital watermarking) is a technology that turns the "invisible" property of invisible characters on its head. By embedding patterns of zero-width characters in text, hidden information can be embedded without changing the appearance.
There is really only one technique at work here, and the characters it uses are limited to zero-width characters such as U+200B, U+200C, U+200D and U+FEFF. The only thing that differs from one use to the next is what the embedded bit sequence is read as.
| Purpose | Embedded Content | Operation Needed to Read It |
|---|---|---|
| Passing a hidden message | An arbitrary bit sequence | The receiver decodes it with the same mapping table |
| Identifying a leak source (a watermark per recipient) | A bit sequence that identifies the recipient | Match it against the pattern recorded at distribution time |
| Detecting unauthorized copying | A bit sequence that indicates the origin | Inspect the copied text code point by code point |
For example, by treating 4 types of zero-width characters as 2-bit information (U+200B = 00, U+200C = 01, U+200D = 10, U+FEFF = 11) and inserting zero-width characters between each word in text, binary data can be hidden within it.
This technology is sometimes used by companies to identify the source of confidential document leaks. By embedding different zero-width character patterns for each recipient, when a document leaks externally, the source can be identified.
Detecting and Removing Invisible Characters
To correctly process text infiltrated by invisible characters, you need to know detection and removal methods.
| Method | Target | Code Example |
|---|---|---|
| JavaScript regex | Major zero-width characters | str.replace(/[\u200B-\u200F\u2028-\u202F\uFEFF]/g, '') |
| Python regex | Same as above | re.sub(r'[\u200b-\u200f\u2028-\u202f\ufeff]', '', text) |
| Text editor | Invisible and confusable characters | VS Code: editor.unicodeHighlight.invisibleCharacters |
| Command line | Invisible characters in files | cat -v filename or xxd filename |
| PHP | Major zero-width characters | preg_replace('/[\x{200B}-\x{200F}\x{FEFF}]/u', '', $str) |
What the JavaScript regex /[\u200B-\u200F\u2028-\u202F\uFEFF]/g covers is limited to the "most common zero-width characters." Run text through it and ZWSP (U+200B) and the BOM (U+FEFF) do disappear, but the word joiner (U+2060) and the soft hyphen (U+00AD) from the table at the top of this article, along with the isolation-type direction control characters covered later (U+2066 to U+2069), fall outside the range and stay where they are. Cleaning form input on the server side is a sound measure, but you have to count up the characters you actually want to remove and write them into the character class yourself.
However, unconditionally removing all invisible characters is dangerous. ZWJ is necessary for emoji composition, and removing it will decompose emoji. ZWNJ is essential for correct rendering in Persian and Hindi. Invisible character removal must be done carefully with understanding of purpose and context.
Invisible Character Handling by Programming Language
What happens when a ZWSP finds its way into source code depends on the language, and on the version of the toolchain. Some languages stop with an error; others quietly accept the character as part of an identifier, and that second group is far more troublesome. The table below shows representative behavior as of August 2026. Because it can change between compilers and versions, the only conclusive check is to try it on your own toolchain.
| Language | ZWSP inside an Identifier | ZWSP in String Literals | Detection Clue |
|---|---|---|---|
| JavaScript (Node.js) | Syntax error (not a character allowed in identifiers) | Retained as part of string | ESLint's no-irregular-whitespace |
| Python | SyntaxError (rejected as a non-printable character) | Retained as part of string | The interpreter itself refuses to run the file |
| Rust | Compile error (cannot be tokenized) | Retained as part of string | uncommon_codepoints warning for ZWNJ and ZWJ |
| C / C++ (clang) | Accepted as an identifier, warning only | Retained as part of string | clang's -Wunicode-zero-width |
The easiest thing to get backwards here is the difference between ZWSP and ZWNJ or ZWJ. The set of characters JavaScript allows inside an identifier includes ZWNJ (U+200C) and ZWJ (U+200D), and does not include ZWSP (U+200B). So var he\u200Bllo stops with a syntax error, while var he\u200Cllo passes without complaint and declares a variable that is distinct from the identical-looking hello.
In practice the two cases mean opposite things. A ZWSP is invisible in the editor but brings the program down the moment it runs, so you always find out that it is there. The dangerous one is the pair that passes: the declaration and its references both succeed, and a later hello typed by hand becomes an undefined variable. The character stays invisible, and all that is left to see is a name that does not match.
When considering variable and function name length guidelines, the risk of invisible character infiltration should be kept in mind. Since they cannot be detected visually in code review, it's important to set up automatic detection through linters and editor settings.
Typical Failure Patterns Caused by Invisible Characters
The ways these characters get in, and the ways things break afterwards, follow a small number of patterns. What they share is that the cause is not visible, so narrowing it down takes time.
- String comparison fails silently: Two strings that look the same are not equal. It usually surfaces as test data typed by hand passing while only the real data, created by copying, fails to match
- Search returns no hits: When a product name or article title contains a zero-width character, typing exactly what is shown on screen no longer matches
- Unique constraints and duplicate checks are bypassed: With a ZWSP in the middle of an email address or an ID, the system treats it as a different value, so records that look identical end up registered side by side
- Copy-paste from PDFs and web pages carries them along: Soft hyphens (U+00AD) and zero-width spaces embedded for layout purposes get pasted in as-is, causing form input validation to judge the character count exceeded
- Diffs do not show them: Review screens and diff output do not draw zero-width characters, so they slip through both code review and manuscript proofreading
Character Count Tools and Invisible Characters
How character count tools handle invisible characters varies by tool. Some ignore invisible characters when counting, while others count them as-is. Without understanding Unicode basics, you can't identify why character counts differ between tools.
If you want to accurately count text characters, we recommend first checking for invisible characters and removing them if necessary before counting. Simply knowing that "invisible characters" exist can prevent many character count-related troubles.
Invisible Characters and Security - Unseen Threats
Trojan Source, mentioned earlier, exploits the nine invisible characters used to control bidirectional text (U+202A to U+202E for embedding and override, and U+2066 to U+2069 for isolation) to create a gap between how source code looks and how it actually executes. Other ways in which invisible characters turn into a security problem are known as well.
| Attack Method | Invisible Characters Used | Impact | Countermeasure |
|---|---|---|---|
| Trojan Source | Direction control chars (U+202A-U+202E, U+2066-U+2069) | Malicious logic undetectable in code review | Enable compiler warnings |
| Homograph attack | Visually identical different chars (U+0430 vs U+0061) | Phishing URL spoofing | Check Punycode display |
| ZWSP injection | U+200B | Bypassing input validation | Server-side invisible character removal |
| BOM injection | U+FEFF | File parser malfunction | Automatic BOM removal processing |
The mechanism that makes the attack work is simple. Once direction control characters are planted inside a comment or a string literal, editors and web diff views reorder the characters for display exactly as instructed, while the compiler reads the byte sequence in its original order. That is enough to produce a gap where code that reads on screen as "check the permission, then run the operation" actually runs the operation without checking. What is being abused is not an invalid character but the legitimate mechanism that exists to display bidirectional text correctly.
This attack is particularly dangerous because it neutralizes code review - the human-eye verification process. Countermeasures include enabling compiler and linter settings that warn about direction control character usage, and incorporating invisible character detection steps into CI/CD pipelines.
As mentioned in the Git commit message writing article, utilizing linters is essential for code quality management. Detecting invisible characters is one of the important roles of linters.