Unicode Analyzer
Inspect every character — code points, UTF-8 bytes, categories, and normalization
Start typing to analyze Unicode characters
Supports all Unicode scripts, emoji, and special characters
Character-Level Inspection
Code Points | Every character in Unicode is assigned a unique number (code point) expressed as U+XXXX. This tool reveals the code point for each character in your text.
UTF-8 Encoding | See the exact byte sequence used to encode each character in UTF-8 — the most common encoding on the web. ASCII characters use 1 byte; emoji and CJK use 3–4.
Invisible Characters | Detect zero-width spaces, non-breaking spaces, directional overrides, and other invisible characters that can cause unexpected bugs or security issues.
Unicode Normalization
NFC (Canonical Composition) | The recommended form for web and storage. Combines base characters and combining marks into precomposed forms (e.g., é as one code point).
NFD (Canonical Decomposition) | Splits precomposed characters into base + combining marks. Useful for processing individual components of accented letters.
NFKC / NFKD | Compatibility normalization replaces stylistic variants (fi → fi, ① → 1) and fullwidth forms. Essential for text comparison and search normalization.