Last updated:

Full-Width vs Half-Width Characters | Impact on Character Counting

13 min read

When working with text that includes East Asian characters, understanding the difference between "full-width" and "half-width" characters is essential. This distinction affects character counting results, form input limits, database storage sizes, and even URL encoding. Whether you are a developer, writer, or general user, this concept is unavoidable. This article systematically covers everything from basic definitions and Unicode technical specifications to byte size comparisons across encodings and real-world edge cases.

Full-Width / Half-Width Converter

Paste your text and convert alphanumerics, katakana, symbols, and spaces between full-width and half-width forms instantly. Everything runs in your browser; nothing you type is sent to a server.

Direction
Convert

Hiragana, kanji, and emoji are never touched. The wave dash 〜 (U+301C) is also preserved as-is; only the fullwidth tilde ~ (U+FF5E) is converted.

How to Use the Converter

Three steps: (1) paste your text into the input box, (2) pick the direction (full-width to half-width, or the reverse), and (3) narrow down what gets converted with the category checkboxes (alphanumerics, katakana, symbols, spaces). The result updates as you type, so there is no convert button to press. Hit "Copy result" and the converted text is ready for form fields or CSV cleanup. To see how the conversion changes your counts, paste the result into Character Counter, which breaks totals down by full-width and half-width.

What Gets Converted, and What Never Does

Instead of a blanket normalization such as NFKC, the tool only converts characters whose full-width and half-width forms map to each other unambiguously, using an explicit conversion table.

Hiragana, kanji, emoji (including ZWJ sequences), and combining characters are left untouched in both directions. The wave dash 〜 (U+301C), covered in the gray-zone section below, is deliberately excluded because it is a different character from the fullwidth tilde ~ (U+FF5E); converting it would recreate the classic wave dash corruption problem. Katakana with no half-width form (ヰ, ヱ, ヵ, ヶ, and similar) also pass through unchanged.

Full-Width vs Half-Width at a Glance

The two terms describe how much horizontal space a character occupies in a fixed-width font: a full-width character takes the width of two half-width characters. For counting and storage, the quick rules are below.

AspectHalf-WidthFull-Width
ExamplesA, 1, !, ア (half-width katakana)あ, 漢, A, 1, 。
Display width1 unit2 units
Character count (Unicode)1 character each1 character each
Byte size in UTF-81 byte (ASCII); 3 bytes for half-width katakana3 bytes (most); 4 bytes for rare extension kanji
Byte size in Shift_JIS1 byte2 bytes

The key thing most people get wrong: full-width and half-width characters both count as one character in standard Unicode counting, even though their byte sizes differ. The old "full-width = 2 bytes, half-width = 1 byte" rule only held under Shift_JIS, not under today's UTF-8. The sections below cover the counting and byte rules in detail.

The Technical Reality Behind "Full-Width" and "Half-Width" - Unicode East Asian Width Property

While the terms "full-width" (全角) and "half-width" (半角) originated in Japanese computing, Unicode formally defines character widths in UAX #11 (Unicode Standard Annex #11: East Asian Width). Each code point is assigned one of six width properties:

  • F (Fullwidth): Fullwidth forms of characters. ASCII fullwidth variants (A, 1, etc., U+FF01–U+FF60)
  • H (Halfwidth): Halfwidth forms. Halfwidth katakana (ア, イ, etc., U+FF61–U+FF9F)
  • W (Wide): Characters that are wide in East Asian contexts. CJK Unified Ideographs, Hiragana, Katakana, etc.
  • Na (Narrow): Characters that are narrow in East Asian contexts. Basic Latin letters (A–Z), etc.
  • A (Ambiguous): Characters whose width varies by context. Some Greek letters, Cyrillic characters, etc.
  • N (Neutral): Characters not used in East Asian contexts

What people commonly call "full-width" includes both F and W categories, while "half-width" includes both H and Na. The A (Ambiguous) category requires special attention - depending on terminal or editor settings, these characters may render as either single-width or double-width. For example, "α" (Greek small letter alpha) may display as full-width in Windows Command Prompt but half-width in macOS Terminal.

Full-Width Characters

Full-width characters occupy twice the display width of half-width characters in fixed-width font environments. In Unicode's East Asian Width property, they are classified as W (Wide) or F (Fullwidth). Most native Japanese characters are full-width:

Half-Width Characters

Half-width characters occupy roughly half the display width of full-width characters. In Unicode, they are classified as Na (Narrow) or H (Halfwidth). Standard ASCII characters fall into this category:

Half-width katakana is discouraged because it originates from the JIS X 0201 standard. Established in 1969, this standard defined dakuten (゙) and handakuten (゚) as separate characters to fit katakana into a limited 7-bit/8-bit code space. As a result, "ガ" becomes "ガ" - counting as 2 characters. Even Unicode NFC normalization does not combine half-width katakana dakuten, making character count discrepancies likely. Unless there is a specific reason, full-width katakana should always be used.

JIS X 0201 and JIS X 0208 - The Historical Origins of Full-Width and Half-Width

The full-width/half-width distinction is closely tied to the evolution of Japanese character encoding standards. JIS X 0201, established in 1969, included ASCII-compatible 7-bit codes plus 63 half-width katakana characters in the 8-bit range. This was a world of 1 character = 1 byte.

JIS X 0208, established in 1978, defined a large character set including 6,349 kanji. Since 1 byte can only represent 256 values, a 2-byte code space was required. This physical size difference between "1-byte characters" and "2-byte characters" was visualized as the "half-width" and "full-width" display width difference in fixed-width font environments.

In other words, "full-width = 2 bytes" was factually correct in Shift_JIS and EUC-JP encodings, but it no longer holds in today's UTF-8 world. The persistence of this equation is due to the many systems built in Japan's IT industry during the 1990s–2000s that assumed Shift_JIS encoding.

Byte Size Comparison Across Encodings

The same character can have vastly different byte sizes depending on the encoding. The following table compares byte sizes for representative characters:

CharacterUTF-8UTF-16Shift_JISEUC-JP
A (half-width letter)1 byte2 bytes1 byte1 byte
あ (hiragana)3 bytes2 bytes2 bytes2 bytes
漢 (kanji)3 bytes2 bytes2 bytes2 bytes
A (full-width letter)3 bytes2 bytes2 bytes2 bytes
ア (half-width katakana)3 bytes2 bytes1 byte2 bytes
€ (euro sign)3 bytes2 bytesN/AN/A
𠮷 (CJK Extension B)4 bytes4 bytes (surrogate pair)N/AN/A

A key takeaway: in UTF-8, half-width katakana "ア" consumes 3 bytes. While it was 1 byte in Shift_JIS, it becomes the same 3 bytes as full-width hiragana in UTF-8. The intuition that "half-width means smaller data size" does not necessarily hold in UTF-8 environments.

Impact on Character Counting - Platform Differences

Most character counting tools count both full-width and half-width characters as "1 character" each. However, counting methods vary by platform, and the same text can produce different results.

Counting Method"Hello 世界" Result
Unicode character count (standard)7 characters
Byte count (Shift_JIS)9 bytes (5+4)
Byte count (UTF-8)11 bytes (5+6)
Byte count (UTF-16)14 bytes (all chars × 2)

Understanding how major platforms handle full-width/half-width counting is also useful in practice:

PlatformCounting MethodFull-Width Handling
X (formerly Twitter)Weighted counting1 Japanese char = 2 units (140 chars out of 280)
LINEUnicode character countFull/half-width both count as 1
SMSEncoding-dependentJapanese: max 70 chars per message (UCS-2)
MySQL VARCHAR(n)Character count (utf8mb4)Full/half-width both count as 1 (byte limit applies)
Oracle VARCHAR2(n BYTE)Byte count1 full-width char = 3 bytes in UTF-8
Google Ads (headlines / descriptions)Weighted counting1 full-width char counts as 2 half-width chars (as of August 2026)

The awkward part is that "one character" does not mean the same thing across these rows. Some platforms count code points, so a full-width character and a half-width character each cost 1; others charge a full-width character double, as though it were two half-width ones. The same visible text therefore fits one field and overflows another, and copy written right up against a limit tends to come out short when it is finally measured by the stricter method, forcing a rewrite at submission time. Checking which method a field uses before writing to the limit avoids that.

Character Counter displays full-width and half-width character counts separately, so you can work with either counting method.

Common Problems from Full-Width/Half-Width Confusion

Full-Width Characters in Programming - A Hidden Trap

Full-width space infiltration (U+3000) in programming is particularly serious. Because full-width and half-width spaces (U+0020) look nearly identical, developers often cannot identify the cause even when reading the error message.

LanguageError Message
PythonSyntaxError: invalid character '\u3000'
Javaillegal character: '\u3000'
JavaScriptSyntaxError: Invalid or unexpected token
C/C++error: stray '\343' in program (UTF-8 lead byte)
RubySyntaxError: invalid multibyte char (UTF-8)

Beyond full-width spaces, accidentally using full-width colons ":" (U+FF1A) instead of half-width colons ":" (U+003A), or mixing in full-width semicolons ";" (U+FF1B), are also common mistakes. In structured data formats like JSON and YAML, a full-width colon causes a syntax error.

In e-commerce search, systems that treat "Tシャツ" (full-width T) and "Tシャツ" (half-width T) as different queries can return vastly different results. The pitfall that catches teams out is partial normalization: if you normalize the incoming query but not the stored data, or the other way round, searches that used to succeed start failing, because the two sides now disagree in a new way. The same rules have to be applied to the indexed values and to the query, otherwise normalization lowers recall instead of raising it.

CSV/TSV and Full-Width Character Pitfalls

In CSV (Comma-Separated Values) files widely used for data exchange, mixing full-width commas "," (U+FF0C) with half-width commas "," (U+002C) causes serious problems. Most CSV parsers only recognize half-width commas as delimiters, so fields containing full-width commas are not split, causing column misalignment.

Similarly, in TSV (Tab-Separated Values) files, full-width spaces used in place of tab characters prevent correct column separation. When opening a CSV in Excel results in garbled text or misaligned columns, full-width character contamination should be suspected.

URL Encoding and Full-Width Characters

When full-width characters appear in URLs, percent-encoding (RFC 3986) converts each byte to %XX format. A Japanese character that is 3 bytes in UTF-8 expands to 9 characters like %E3%81%82.

For example, "東京都" (3 characters) becomes %E6%9D%B1%E4%BA%AC%E9%83%BD (27 characters) in a URL. Considering URL length limits (typically 2,048 characters), URLs containing many full-width characters can quickly reach the limit. When using Japanese in file names or directory names, this expansion must be factored into the design.

Professional Management Techniques

  1. Enable "show invisible characters" in your text editor. In VS Code, set editor.renderWhitespace: "all" to visually distinguish full-width spaces. Additionally, enabling editor.unicodeHighlight.ambiguousCharacters: true highlights Ambiguous-category characters.
  2. Use regex to detect full-width alphanumerics. The pattern [A-Za-z0-9] finds full-width alphanumerics for batch conversion.
  3. Implement server-side normalization for form inputs. Automatically convert full-width input to half-width to prevent errors.
  4. Use IME shortcuts for quick conversion. On Windows, F10 converts to half-width alphanumerics. On macOS, the Japanese input method's shortcut for this converts the text being typed to Roman letters (half-width alphanumerics) by default; to get half-width katakana instead, half-width katakana has to be enabled in the input source settings first (as of August 2026).
  5. Set up Git pre-commit hooks to detect full-width spaces. Running grep -rn $'\xe3\x80\x80' catches full-width spaces across the repository before they are committed.

Web Form Auto-Conversion Implementation Patterns

In Japanese web services, automatic full-width to half-width conversion is widely implemented for phone numbers, postal codes, and email address fields. Here is a common implementation pattern.

The basic JavaScript logic for converting full-width alphanumerics to half-width leverages Unicode code point offsets. Full-width alphanumerics (U+FF01–U+FF5E) differ from their half-width ASCII counterparts (U+0021–U+007E) by exactly 0xFEE0.

function toHalfWidth(str) {
  return str.replace(/[\uFF01-\uFF5E]/g, ch =>
    String.fromCharCode(ch.charCodeAt(0) - 0xFEE0)
  ).replace(/\u3000/g, ' ');
}

This function converts full-width alphanumerics and symbols to half-width, and also converts full-width spaces to half-width spaces. However, full-width katakana to half-width katakana conversion involves complex dakuten/handakuten handling, so using a dedicated library is recommended.

For HTML input elements, instead of the deprecated CSS ime-mode property, the inputmode attribute can control input mode. Setting inputmode="numeric" displays a numeric keyboard on mobile devices, reducing the risk of full-width input.

Regex-Based Full-Width/Half-Width Detection in Practice

Regular expressions are a practical way to detect full-width and half-width characters, but East Asian Width is not among the properties you can reach with \p{...} property escapes: those expose Unicode scripts and general categories, not the width property. Detection therefore has to spell out code point ranges, as in the patterns below.

// Detect full-width characters (Wide + Fullwidth)
const fullwidthPattern = /[\u3000-\u303F\u3040-\u309F\u30A0-\u30FF\u4E00-\u9FFF\uFF01-\uFF60]/;

// Detect half-width katakana
const halfwidthKatakana = /[\uFF61-\uFF9F]/;

// Detect full-width alphanumerics only (useful for conversion targeting)
const fullwidthAlphaNum = /[\uFF10-\uFF19\uFF21-\uFF3A\uFF41-\uFF5A]/;

For normalizing full-width/half-width before database storage, NFKC (Normalization Form Compatibility Composition) is effective. In JavaScript, "A".normalize("NFKC") converts full-width "A" to half-width "A". However, NFKC also expands characters like "㍻" into "平成", so the scope of application must be carefully considered.

Gray-Zone Characters

Two very different problems get lumped together as the "gray zone", and separating them clears up most of the confusion. The first is genuine width ambiguity: characters in the A (Ambiguous) category of the East Asian Width property, whose rendered width really does change with the environment. Greek letters such as α, Cyrillic letters, and circled numbers like ① sit here, and a terminal may draw them one column wide or two depending on its settings. The second is a mapping problem: the width is fixed by the standard, but two look-alike code points coexist, so text can silently change identity as it moves between encodings.

The wave dash (〜, U+301C) versus the fullwidth tilde (~, U+FF5E) is the best-known case of the second kind, not the first. Both sit on the wide side of the East Asian Width property (U+301C is W, U+FF5E is F), so their width is not what is in dispute. The trouble is that Windows' Shift_JIS implementation mapped the wave dash to the fullwidth tilde, so moving a file between operating systems could turn one into the other and produce garbled text. This "wave dash problem" stems from differing interpretations of the wave dash glyph in the JIS X 0208 character code table, and it is a mapping conflict rather than a width conflict.

The yen sign (¥, U+00A5) and the backslash (\, U+005C) are a mapping problem as well. Both are Na (narrow) in the East Asian Width property, and neither is the same character as the full-width yen sign ¥ (U+FFE5), which is F. The pairing goes back to JIS X 0201 placing the yen sign at 0x5C, the ASCII position of the backslash; Windows Japanese fonts still draw the backslash with a yen glyph, which is why C:¥Users and C:\Users both circulate as ways of writing the same path.

The converter at the top of this article follows the same reasoning: it never converts the wave dash (U+301C) and only maps the fullwidth tilde ~ (U+FF5E) to ~ (U+007E). Likewise, it treats the yen sign pair ¥ (U+FFE5) ⇔ ¥ (U+00A5) and the backslash pair \ (U+FF3C) ⇔ \ (U+005C) as separate mappings, never mixing up characters that merely look alike.

Database Best Practices for Full-Width/Half-Width Normalization

Normalizing full-width/half-width text before database storage directly improves search accuracy and data quality.

  1. Normalize at input time: Apply NFKC normalization in the application layer before INSERT. This automatically converts full-width alphanumerics to half-width.
  2. Normalize at search time: Apply the same normalization to search queries to absorb notation variations between stored data and search conditions. In MySQL, using COLLATE utf8mb4_unicode_ci enables case-insensitive and width-insensitive collation.
  3. Column design: Clarify whether VARCHAR length is character-based (MySQL) or byte-based (Oracle), and set byte limits accounting for 1 full-width character = 3 bytes in UTF-8.
  4. Index design: When width-insensitive search is needed, create a separate column storing normalized values and index that column for efficient lookups.

Usage Rules for Full-Width and Half-Width

Knowing when to use full-width versus half-width characters is essential for producing polished Japanese text. While conventions vary by medium and style guide, the following rules are widely accepted.

  1. Use half-width for alphanumeric characters in horizontal text (e.g., 2024年, 100円)
  2. Use full-width brackets for Japanese quotations (e.g., 「こんにちは」)
  3. Always use half-width for URLs and email addresses
  4. Follow the specified format (full-width or half-width) when filling in forms

In web content, the standard practice is to use half-width for all alphanumeric characters and half-width spaces, while keeping Japanese punctuation marks (。and 、) in full-width. Avoid full-width spaces entirely - they are a common source of invisible formatting issues in HTML and code.

Conclusion

The full-width/half-width distinction is not merely cosmetic - it directly impacts character counting, byte calculations, database design, URL design, and programming correctness. At its foundation lie the historical legacy of JIS X 0201/0208 and the technical specification of Unicode's East Asian Width property. By accurately understanding byte size differences across encodings and applying practical techniques like NFKC normalization and regex-based detection, you can prevent full-width/half-width issues before they occur. Use Character Counter to check full-width and half-width breakdowns for accurate character management.

Share this article