Last updated:

The Curious Relationship Between Kanji Stroke Count and Character Count - A Single Character Can Have 84 Strokes

11 min read

The kanji "靐," which stacks three copies of "雷" (thunder), has 39 strokes. "𪚥," four copies of "龍" (dragon), reaches 64 strokes. And "taito," considered the most stroke-heavy kanji in Japan, hits a staggering 84 strokes. Yet every one of these counts as just "1 character" in a character counter. The single-stroke "一" and the 84-stroke "taito" are both one character. This article explores the relationship between kanji stroke counts and character counting, covering how many kanji Unicode contains and how variant character systems work.

Ranking the Most Stroke-Heavy Kanji

There is no theoretical upper limit to kanji stroke counts. Because existing kanji can be combined to form new ones, stroke counts can increase without bound. However, limiting ourselves to kanji with documented historical usage, the ranking looks like this.

Many high-stroke kanji share a structure called "rigiji" - characters built by stacking identical components. "森" (tree × 3 = 12 strokes), "轟" (vehicle × 3 = 21 strokes), "靐" (thunder × 3 = 39 strokes) - doubling or tripling the same radical multiplies the stroke count. These compound characters typically convey intensified meanings like "many" or "intense," and this method of character creation has existed since ancient times.

KanjiStrokesReadingCompositionUnicode Status
𱁬84taito, daito, otodo3 × cloud + 3 × dragonIncluded (U+3106C, added in Unicode 13.0)
𪚥64tetsu, techi龍 × 4Included (U+2A6A5)
𰻞58byanUsed to write the name of the Shaanxi noodle dish "biangbiang mian"Included (U+30EDE, added in Unicode 13.0)
靐39hyō雷 × 3Included (U+9750)
鬱29utsuHighest stroke count among jōyō kanjiIncluded (U+9B31)

The 84-stroke character is often said to have been used as a surname, but that story has no confirmable basis in dictionaries or public records. As a matter of character encoding, it was added in Unicode 13.0 (2020) as part of Extension G at U+3106C, so it can be displayed as a single character wherever a compatible font is available. The same is true of the 64-stroke "𪚥." A kanji that is not in Unicode, by contrast, cannot be typed or stored as text at all - it can only be shown as an image.

Among the 2,136 jōyō kanji (characters designated for everyday use), "鬱" holds the record at 29 strokes. It was added to the jōyō list in the 2010 revision (Cabinet Notification No. 2 of 2010), and among the public comments submitted on the draft revision, it was the character most frequently proposed for removal. The list is not chosen by counting strokes, however; the test is how necessary a character is in everyday writing. That is why a 29-stroke character earns a place on it.

Stroke Count Distribution of Jōyō Kanji

Examining the stroke count distribution of all 2,136 jōyō kanji reveals the "standard complexity" of characters used in daily Japanese.

Stroke RangeCountPercentageRepresentative Kanji
1-4 strokes1175.5%一, 二, 人, 大, 中, 日
5-8 strokes57526.9%生, 出, 本, 学, 国, 物
9-12 strokes84339.5%食, 海, 時, 動, 朝, 森
13-16 strokes47822.4%話, 歴, 様, 線, 機, 橋
17-20 strokes1135.3%題, 類, 議, 識, 観, 競
21+ strokes100.5%欄, 露, 鶴, 籠, 驚, 鬱

The 9-12 stroke range is the most populated, accounting for about 39.5% of all jōyō kanji. The average is roughly 10.4 strokes. In other words, the kanji used in everyday Japanese text average about 10 strokes in complexity. Only 10 characters reach 21 strokes or more: "欄," "躍," "露," "顧," "鶴" (21 strokes), "籠," "襲," "鑑," "驚" (22 strokes), and "鬱" (29 strokes).

Note that dictionaries sometimes differ by one stroke on the same character. The figures here cover the 2,136 characters of the jōyō list (2010 Cabinet notification), counted by the total stroke values given in Unicode's character properties for kanji.

The peak at 9-12 strokes follows from how kanji were built. The largest group by far is the phono-semantic compounds, which pair a part carrying the meaning with a part carrying the sound. The radical side typically runs 3 to 4 strokes and the sound side 6 to 8, so the combined character lands squarely in this range. Characters with too few strokes blur into each other ("未" and "末," "己" and "已"), while very dense ones are a burden to write and to read. Everyday kanji cluster in this band as a result of both the mechanics of character formation and plain usability.

High Stroke Counts, Same Byte Size - Unicode's Equality

This is the most important point from a character counting perspective. No matter how many strokes a kanji has, the relationship between character count and byte size remains unchanged.

KanjiStrokesUnicode Code PointUTF-8 BytesUTF-16 Bytes
一1U+4E003 bytes2 bytes
鬱29U+9B313 bytes2 bytes
靐39U+97503 bytes2 bytes
𪚥64U+2A6A54 bytes4 bytes (surrogate pair)

All kanji in the CJK Unified Ideographs basic block (U+4E00-U+9FFF) take 3 bytes in UTF-8 and 2 bytes in UTF-16, regardless of stroke count. The 64-stroke "𪚥" resides in Extension B (U+20000-U+2A6DF), so it requires 4 bytes in UTF-8 and a surrogate pair (4 bytes) in UTF-16.

In short, stroke count has no effect on byte size. What matters is which Unicode block the character belongs to. Basic block kanji take 3 bytes; extension block kanji take 4. This difference is determined by when the character was added to Unicode, not by its stroke count.

Growth of CJK Unified Ideographs in Unicode

The number of kanji encoded in Unicode has grown steadily with each version. As explained in Unicode basics, Unicode is a standard for handling all the world's writing systems uniformly, and kanji encoding is an especially large-scale undertaking.

Unicode VersionRelease YearCJK Unified Ideographs (Cumulative)Main Additions
1.0.1199220,902The base block (the starting point of Han unification)
3.0199927,484Extension A (6,582)
3.1200170,195Extension B (42,711)
5.2200974,382Extension C (4,149)
8.0201580,376Extension E (5,762)
10.0201787,870Extension F (7,473)
13.0202092,844Extension G (4,939)
15.1202397,668Extension H (4,192, in 15.0) + Extension I (622)
17.02025101,984Extension J (4,298)

The cumulative column is the total of the base block and the extension blocks (Extensions A through J). It grows by slightly more than each row's additions because the base block itself also gains a handful of characters in most versions. One further wrinkle: twelve characters sitting in the compatibility area are treated as unified ideographs, and counts that include them come out twelve higher on every row. That is why the same version is sometimes quoted as "97,668" and sometimes as "97,680."

From about 20,000 characters in Unicode 1.0.1 (1992), CJK Unified Ideographs reached roughly 100,000 in Unicode 17.0 (2025) - a fivefold increase over three decades. However, the kanji used in daily life number only about 3,000 in Japanese and about 3,500 in Simplified Chinese. The remaining 90,000+ are rare characters found in classical texts, dialects, and historical variant forms.

Ideographic Variation Sequences (IVS) - One Character, Multiple Glyphs

What makes kanji character counting even more complex is the Ideographic Variation Sequence (IVS) system. IVS distinguishes different glyphs (variant forms) of the same kanji by appending a variation selector (U+E0100-U+E01EF) after the base character.

The easy point of confusion here is the difference from older character forms. The traditional forms of "辺," namely "邊" and "邉," each hold their own separate code point in Unicode, so no IVS is needed for them. What IVS handles is something else: differences in glyph shape that were unified into a single character at a single code point. Whether a small stroke is present, whether the final stroke flicks upward or stops - details of that kind are specified by placing a selector after the base character. Family registry and property registration work has to record such details exactly, which is where IVS comes in.

When IVS is used, what appears as one character on screen actually consumes two code points - the base character plus the variation selector. This is the same structure seen in emoji character counting, where a single emoji can consist of multiple code points.

Base CharacterVariation SelectorDisplayed GlyphCode PointsUsage
辺 (U+8FBA)VS17 (U+E0100)Variant 1 of 辺2Family registry names
辺 (U+8FBA)VS18 (U+E0101)Variant 2 of 辺2Family registry names
葛 (U+845B)VS17 (U+E0100)Variant of 葛2Place names (Katsushika vs. Katsuragi)
祇 (U+7947)VS17 (U+E0100)Variant of 祇2Exact glyph for "Gion"

Some character counting tools count IVS-enhanced kanji as "2 characters." It looks like one character to the human eye, but the program sees two. This mismatch causes real problems in name input forms and address database systems.

Legal Restrictions on Kanji in Names

In Japan, the characters allowed in children's names are legally restricted. Article 50 of the Family Register Act requires that "the name of a child shall use characters that are in common use and plain," and the concrete range is set out in Article 60 of the Act's Enforcement Regulations: the kanji listed in the jōyō kanji table, the kanji in Appended Table 2 of the same regulations (the name-use kanji, or jinmeiyō kanji), katakana, and hiragana. A kanji outside that range cannot be used in a name no matter how few strokes it has.

There is no legal limit on the number of characters in a name itself, but practical constraints exist in registry systems. The range of characters each municipal registry system can handle varies, and the treatment of variant forms and old-style characters differs by municipality.

The information on record is expanding rather than shrinking. The revised Family Register Act that took effect on May 26, 2025 added the reading of a person's name, in kana, to the items recorded in the family register (Article 13 of the Act). Systems that handle personal names now carry the reading alongside the kanji, which means one more field to store and validate.

Stroke Count and Education - Design Philosophy of Grade-Level Kanji

The 1,026 educational kanji taught in elementary school are allocated by grade. This allocation considers not only stroke count but also usage frequency and conceptual difficulty, though the correlation with stroke count is clear.

GradeCharactersCumulative
1st grade8080
2nd grade160240
3rd grade200440
4th grade202642
5th grade193835
6th grade1911,026

Looking at what each grade contains, the first-year list is dominated by low-stroke characters such as "一," "日," and "山," while characters of more than ten strokes like "臓" and "警" sit in the upper grades. The steady rise in the number of component parts per character matches children's developing fine motor skills and their widening vocabulary. Not teaching "鬱" (29 strokes) in first grade is the obvious call once writing ability is taken into account.

Stroke Count and Handwriting Input - A Digital-Age Challenge

On smartphones and tablets, stroke count directly affects handwriting recognition accuracy. The more strokes a kanji has, the harder it is to write it accurately on a small screen. Fitting many strokes into a narrow box squeezes the gaps between them, which blurs the question of where one stroke ended and the next began.

Few people can accurately write "鬱" (29 strokes) using smartphone handwriting input. On-screen handwriting does have an advantage over recognizing characters printed on paper: it can also use information from the act of writing itself, such as the order in which the lines were drawn and the direction each one took. The burden on the person writing, however, still scales with the stroke count. What any particular recognition engine does internally is not published, so accuracy cannot be discussed in numbers, but the practical conclusion is simple. For a high-stroke character, do not persist with handwriting; switching to radical search or phonetic input gets you there faster.

This issue connects to form input validation design. When a name input form accepts handwriting, UI design that accounts for misrecognition of high-stroke kanji - such as displaying candidate lists or offering phonetic conversion options - is essential. For government online services where the exact registered glyph must be entered, handwriting recognition accuracy directly impacts user experience.

For high-stroke kanji, typing the reading and converting it takes fewer actions than writing the character by hand. The effort of handwriting grows in step with the stroke count, whereas the number of keystrokes for phonetic input has nothing to do with strokes at all - it depends only on how many kana the reading takes. "鬱" runs to 29 strokes, but its reading is the two kana of "utsu." That asymmetry is why the more complex the character, the more it pays to lean on conversion.

Practical Lessons from Stroke Count and Character Counting

The key practical lesson from the relationship between kanji strokes and character count is that "visual complexity and data size are separate things." Just as with the difference between fullwidth and halfwidth characters, a character's appearance and its data representation do not necessarily match.

When designing character limits for web forms, treating "1 kanji = 1 character" is standard practice, but accounting for IVS-enhanced kanji and surrogate pairs makes the implementation less straightforward. For name input forms in particular, whether variant characters are handled correctly directly affects user experience.

An 84-stroke kanji and a 1-stroke kanji are equally "1 character" in a character counter. This equality is a core design principle of Unicode - the foundation for handling all the world's writing systems uniformly. Beyond the physical complexity of stroke count, every character is treated equally in the digital world. That is both the beauty and the occasional headache of Unicode.

Stroke count is an important metric in calligraphy and education, but it is completely ignored in digital character counting. The 1-stroke "一" and the 29-stroke "鬱" both count as one character in Slack messages and one character in LINE messages. This "democratization of stroke count" is one of the significant changes that digital communication has brought to the kanji-using world.

Share this article