Last updated:

Emoji Character Counting: Why One Emoji Can Count as Multiple Characters

9 min read

Emoji History and Growth by Unicode Version

The starting point for emoji spreading worldwide was 1999, when Shigetaka Kurita at NTT DoCoMo designed a set of 176 icons for the i-mode mobile platform, rendered as 12×12 pixel art. Initially, each Japanese mobile carrier (DoCoMo, au, SoftBank) implemented emoji independently, causing frequent garbled text when messages were sent between carriers. To resolve this compatibility issue, Google and Apple proposed emoji standardization to the Unicode Consortium, and 722 emoji were officially adopted in Unicode 6.0 in 2010.

Emoji kept being added in later Unicode releases, but the size of each batch varies enormously from version to version. The figures below are the ones that can be traced back to published release material.

Unicode VersionRelease YearEmoji Added
6.02010722
7.02014about 250
8.0201541
16.020248

One thing to keep in mind is that the "total number of emoji" moves with the definition you count by. Counting only standalone code points gives a very different answer from counting skin-tone variants and ZWJ sequences as separate entries, even at the same moment in time, which is why published totals disagree so widely. The Unicode Consortium's own tally, in the Emoji 17.0 edition, comes to 3,953 including derived forms.

The number of new additions has been declining in recent years. The Unicode Consortium now rigorously evaluates whether a proposed emoji can be represented by combining existing emoji via ZWJ sequences, and declines to allocate new code points when composition is feasible. Proposals must also carry evidence of expected usage frequency, drawn from sources such as Google Search, Google Trends, and the Google Books Ngram Viewer, together with a case that the new design is clearly distinguishable from existing emoji.

Byte Size by Encoding: Measured Data

EmojiVisualCode PointsUTF-8 BytesUTF-16 BytesUTF-32 Bytes
😀 (U+1F600)1 char1444
👍🏻 (U+1F44D U+1F3FB)1 char2888
👨‍👩‍👧‍👦1 char7252228
🇺🇸 (U+1F1FA U+1F1F8)1 char2888
1️⃣ (U+0031 U+FE0F U+20E3)1 char37612
🏳️‍🌈1 char4141216

In UTF-8, code points outside the Basic Multilingual Plane (BMP) consume 4 bytes each. UTF-16 represents them as surrogate pairs (two 16-bit units), also totaling 4 bytes. UTF-32 is fixed-width at 4 bytes per code point, making calculations simple but memory efficiency the worst of the three. The family emoji 👨‍👩‍👧‍👦 consuming 25 bytes in UTF-8 is a critical consideration when designing database column sizes.

How ZWJ, Variation Selectors, and Surrogate Pairs Work

Unicode uses ZWJ (Zero Width Joiner, U+200D) to combine multiple code points into a single visual emoji. The family emoji 👨‍👩‍👧‍👦 is composed of "man (U+1F468) + ZWJ + woman (U+1F469) + ZWJ + girl (U+1F467) + ZWJ + boy (U+1F466)" - 7 code points total.

This design was adopted because registering every combination individually would multiply out very quickly. Give each of the four members of a family emoji one of the 5 skin tone levels and you already have 5 to the power of 4, or 625, variants of that single icon, and multiplying by the different family compositions pushes the figure into the thousands. ZWJ composition allows diversity through combining basic building blocks, preventing code point exhaustion.

Inside the family emoji 👨‍👩‍👧‍👦 - how one visible character becomes 7 code points and 25 bytes
On-screen appearance 👨‍👩‍👧‍👦 1
Code points 👨 U+1F468 ZWJ U+200D 👩 U+1F469 ZWJ U+200D 👧 U+1F467 ZWJ U+200D 👦 U+1F466 7
UTF-16 code units 👨 2 ZWJ 1 👩 2 ZWJ 1 👧 2 ZWJ 1 👦 2 11
UTF-8 bytes 👨 4 ZWJ 3 👩 4 ZWJ 3 👧 4 ZWJ 3 👦 4 25

The dashed cells are the ZWJ code points, which never appear on screen. The four person emoji live outside the BMP, so each one takes a surrogate pair of 2 UTF-16 units and 4 UTF-8 bytes, while each ZWJ takes 1 unit and 3 bytes. The very same icon returns 1 (Swift .count), 7 (Python 3 len()), or 11 (JavaScript .length) depending only on which layer you count. Because it reaches 25 bytes in UTF-8, MySQL's utf8 charset - capped at 3 bytes per character - cannot store it at all.

Beyond ZWJ, several invisible code points control emoji display:

JavaScript's internal string representation uses UTF-16, so emoji outside the BMP (U+10000 and above) are represented as surrogate pairs - two 16-bit code units. This is the fundamental reason "😀".length returns 2 instead of 1. Splitting a surrogate pair (high surrogate U+D800–U+DBFF, low surrogate U+DC00–U+DFFF) produces an invalid string, making string truncation particularly dangerous.

Platform-Specific Emoji Counting and Rendering Differences

The same emoji can also look dramatically different across Apple, Google, Samsung, and Microsoft platforms. For example, 🔫 (pistol) was changed to a toy water gun design by Apple, and other vendors followed - but at different times. When emoji appearance matters in marketing or UI design, preview your content across major platforms before publishing.

Character Count Differences Across Programming Languages

The same emoji produces different length values depending on the programming language, because each language uses a different internal string representation.

Language / Method"😀""👍🏻""👨‍👩‍👧‍👦"Count Unit
JavaScript .length2411UTF-16 code units
JavaScript [...str].length127Unicode code points
Python 3 len()127Unicode code points
Rust .len()4825UTF-8 bytes
Rust .chars().count()127Unicode code points
Swift .count111Grapheme clusters
Go len()4825UTF-8 bytes
Java .length()2411UTF-16 code units

Only Swift counts by grapheme clusters, returning 1 for any emoji regardless of internal complexity. JavaScript and Java use UTF-16 internally, so emoji outside the BMP are counted as 2 (surrogate pairs). Rust and Go return byte counts, making them unsuitable for character counting without additional processing. Developers must understand exactly what their language's length returns.

Common Mistakes and How to Avoid Them

Developer Guide: Counting Emoji Accurately

  1. Grapheme cluster segmentation with Intl.Segmenter: Standardized in the internationalization API (ECMA-402), Intl.Segmenter splits strings by grapheme clusters - the units users perceive as single characters. [...new Intl.Segmenter().segment(str)].length gives the accurate visual character count for any emoji. As of August 2026 it is available in Chrome 87+, Safari 14.1+, Node.js 16+, and Firefox 125+.
  2. Emoji detection and removal with regex: The Unicode property escape /\p{Emoji_Presentation}/u detects emoji in strings. Note that \p{Emoji} also matches digits (0-9) and #, so use \p{Emoji_Presentation} or \p{Extended_Pictographic} when targeting only pictorial emoji.
  3. Comprehensive emoji test set: When testing, cover these 5 categories at minimum: (1) basic emoji (😀), (2) skin tone modified (👍🏻), (3) ZWJ sequences (👨‍👩‍👧‍👦), (4) flags (🇺🇸), and (5) keycap sequences (1️⃣).
  4. Database design best practices: Choose the N in VARCHAR(N) knowing that it counts characters and that one grapheme cluster can swallow 7 of them. When you then estimate row size and index length, convert at up to 4 bytes per character. In MySQL, use utf8mb4, and for chat applications with heavy emoji use, consider using the TEXT type instead.

Conclusion

Emoji counting is more complex than it appears. Different platforms and programming languages count emoji differently. Use Character Counter to get accurate character counts that account for emoji complexity.

Share this article