Unicode Analyzer

Inspect every character — code points, UTF-8 bytes, categories, and normalization

0 code points
🔍

Start typing to analyze Unicode characters

Supports all Unicode scripts, emoji, and special characters

Character-Level Inspection

Code Points | Every character in Unicode is assigned a unique number (code point) expressed as U+XXXX. This tool reveals the code point for each character in your text.

UTF-8 Encoding | See the exact byte sequence used to encode each character in UTF-8 — the most common encoding on the web. ASCII characters use 1 byte; emoji and CJK use 3–4.

Invisible Characters | Detect zero-width spaces, non-breaking spaces, directional overrides, and other invisible characters that can cause unexpected bugs or security issues.

Unicode Normalization

NFC (Canonical Composition) | The recommended form for web and storage. Combines base characters and combining marks into precomposed forms (e.g., é as one code point).

NFD (Canonical Decomposition) | Splits precomposed characters into base + combining marks. Useful for processing individual components of accented letters.

NFKC / NFKD | Compatibility normalization replaces stylistic variants (fi → fi, ① → 1) and fullwidth forms. Essential for text comparison and search normalization.