Unicode String Forensic Inspector
Dissect any string: analyze individual Unicode codepoints, UTF-8/16/32 raw byte sizes, UAX #29 grapheme clusters, scripts, and invisible zero-width characters in real time.
Interactive Character Anatomy Heatmap38 Glyphs
Color-coded by Unicode classification. Click any tile to inspect byte structures and homoglyph traits.
This text contains invisible zero-width or formatting control characters. These are frequently used in homoglyph phishing, zero-width steganography watermarks, or security filter evasions.
Character-by-Character Forensic Breakdown (38 Characters)
Unicode 16.0 Architecture| # | Glyph | Codepoint (Hex / Dec) | UTF-8 Raw Bytes | UTF-16 Units | HTML Entity | ES6 Escape | Script & Category | Action |
|---|---|---|---|---|---|---|---|---|
| 1 | R | U+0052(82) | 0x52 | 0x0052 | R | \u0052 | Basic Latin (ASCII)Letter, Other (Lo/Lu/Ll) | Lookup |
| 2 | o | U+006F(111) | 0x6F | 0x006F | o | \u006F | Basic Latin (ASCII)Letter, Other (Lo/Lu/Ll) | Lookup |
| 3 | c | U+0063(99) | 0x63 | 0x0063 | c | \u0063 | Basic Latin (ASCII)Letter, Other (Lo/Lu/Ll) | Lookup |
| 4 | k | U+006B(107) | 0x6B | 0x006B | k | \u006B | Basic Latin (ASCII)Letter, Other (Lo/Lu/Ll) | Lookup |
| 5 | e | U+0065(101) | 0x65 | 0x0065 | e | \u0065 | Basic Latin (ASCII)Letter, Other (Lo/Lu/Ll) | Lookup |
| 6 | t | U+0074(116) | 0x74 | 0x0074 | t | \u0074 | Basic Latin (ASCII)Letter, Other (Lo/Lu/Ll) | Lookup |
| 7 | ␣ | U+0020(32) | 0x20 | 0x0020 |   | \u0020 | Symbol / OtherSeparator, Space (Zs) | Lookup |
| 8 | 🚀 | U+1F680(128640) | 0xF0 0x9F 0x9A 0x80 | 0xD83D | 🚀 | \u{1F680} | Emojis & SymbolsSymbol, Other (So) | Lookup |
| 9 | , | U+002C(44) | 0x2C | 0x002C | , | \u002C | Symbol / OtherOther | Lookup |
| 10 | ␣ | U+0020(32) | 0x20 | 0x0020 |   | \u0020 | Symbol / OtherSeparator, Space (Zs) | Lookup |
| 11 | A | U+0041(65) | 0x41 | 0x0041 | A | \u0041 | Basic Latin (ASCII)Letter, Other (Lo/Lu/Ll) | Lookup |
| 12 | s | U+0073(115) | 0x73 | 0x0073 | s | \u0073 | Basic Latin (ASCII)Letter, Other (Lo/Lu/Ll) | Lookup |
| 13 | t | U+0074(116) | 0x74 | 0x0074 | t | \u0074 | Basic Latin (ASCII)Letter, Other (Lo/Lu/Ll) | Lookup |
| 14 | r | U+0072(114) | 0x72 | 0x0072 | r | \u0072 | Basic Latin (ASCII)Letter, Other (Lo/Lu/Ll) | Lookup |
| 15 | o | U+006F(111) | 0x6F | 0x006F | o | \u006F | Basic Latin (ASCII)Letter, Other (Lo/Lu/Ll) | Lookup |
| 16 | n | U+006E(110) | 0x6E | 0x006E | n | \u006E | Basic Latin (ASCII)Letter, Other (Lo/Lu/Ll) | Lookup |
| 17 | a | U+0061(97) | 0x61 | 0x0061 | a | \u0061 | Basic Latin (ASCII)Letter, Other (Lo/Lu/Ll) | Lookup |
| 18 | u | U+0075(117) | 0x75 | 0x0075 | u | \u0075 | Basic Latin (ASCII)Letter, Other (Lo/Lu/Ll) | Lookup |
| 19 | t | U+0074(116) | 0x74 | 0x0074 | t | \u0074 | Basic Latin (ASCII)Letter, Other (Lo/Lu/Ll) | Lookup |
| 20 | ␣ | U+0020(32) | 0x20 | 0x0020 |   | \u0020 | Symbol / OtherSeparator, Space (Zs) | Lookup |
| 21 | 👩 | U+1F469(128105) | 0xF0 0x9F 0x91 0xA9 | 0xD83D | 👩 | \u{1F469} | Emojis & SymbolsSymbol, Other (So) | Lookup |
| 22 | 🏽 | U+1F3FD(127997) | 0xF0 0x9F 0x8F 0xBD | 0xD83C | 🏽 | \u{1F3FD} | Emojis & SymbolsSymbol, Other (So) | Lookup |
| 23 | ◌ | U+200D(8205) | 0xE2 0x80 0x8D | 0x200D | ‍ | \u200D | General PunctuationOther, Format (Cf) | Lookup |
| 24 | 🚀 | U+1F680(128640) | 0xF0 0x9F 0x9A 0x80 | 0xD83D | 🚀 | \u{1F680} | Emojis & SymbolsSymbol, Other (So) | Lookup |
| 25 | ␣ | U+0020(32) | 0x20 | 0x0020 |   | \u0020 | Symbol / OtherSeparator, Space (Zs) | Lookup |
| 26 | & | U+0026(38) | 0x26 | 0x0026 | & | \u0026 | Symbol / OtherOther | Lookup |
| 27 | ␣ | U+0020(32) | 0x20 | 0x0020 |   | \u0020 | Symbol / OtherSeparator, Space (Zs) | Lookup |
| 28 | S | U+0053(83) | 0x53 | 0x0053 | S | \u0053 | Basic Latin (ASCII)Letter, Other (Lo/Lu/Ll) | Lookup |
| 29 | p | U+0070(112) | 0x70 | 0x0070 | p | \u0070 | Basic Latin (ASCII)Letter, Other (Lo/Lu/Ll) | Lookup |
| 30 | a | U+0061(97) | 0x61 | 0x0061 | a | \u0061 | Basic Latin (ASCII)Letter, Other (Lo/Lu/Ll) | Lookup |
| 31 | r | U+0072(114) | 0x72 | 0x0072 | r | \u0072 | Basic Latin (ASCII)Letter, Other (Lo/Lu/Ll) | Lookup |
| 32 | k | U+006B(107) | 0x6B | 0x006B | k | \u006B | Basic Latin (ASCII)Letter, Other (Lo/Lu/Ll) | Lookup |
| 33 | l | U+006C(108) | 0x6C | 0x006C | l | \u006C | Basic Latin (ASCII)Letter, Other (Lo/Lu/Ll) | Lookup |
| 34 | e | U+0065(101) | 0x65 | 0x0065 | e | \u0065 | Basic Latin (ASCII)Letter, Other (Lo/Lu/Ll) | Lookup |
| 35 | s | U+0073(115) | 0x73 | 0x0073 | s | \u0073 | Basic Latin (ASCII)Letter, Other (Lo/Lu/Ll) | Lookup |
| 36 | ␣ | U+0020(32) | 0x20 | 0x0020 |   | \u0020 | Symbol / OtherSeparator, Space (Zs) | Lookup |
| 37 | ✨ | U+2728(10024) | 0xE2 0x9C 0xA8 | 0x2728 | ✨ | \u2728 | Symbol / OtherOther | Lookup |
| 38 | 💖 | U+1F496(128150) | 0xF0 0x9F 0x92 0x96 | 0xD83D | 💖 | \u{1F496} | Emojis & SymbolsSymbol, Other (So) | Lookup |
The Architecture of Unicode Codepoints, Code Units & Grapheme Clusters
In computer programming, a fundamental confusion exists between Code Units, Codepoints, and User-Perceived Characters (Grapheme Clusters):
"🚀".length === 2.U+200D) that render as a single visual glyph.Encoding Efficiency Across UTF-8, UTF-16 & UTF-32
- •UTF-8 (Web Standard): 1 byte for ASCII, 2 bytes for Arabic/Hebrew, 3 bytes for Devanagari/CJK, 4 bytes for Emojis.
- •UTF-16 (Java & Windows): 2 bytes for BMP, 4 bytes for Astral planes.
- •UTF-32 (Fixed Width): Exactly 4 bytes per character. Easiest for array indexing but consumes maximum memory.
- •100% In-Browser Privacy: Inspect secret keys, passwords, and sensitive text locally in your browser.
Frequently Asked Questions (FAQs)
How does this tool detect zero-width invisible characters?+
Our inspector scans each codepoint against the Unicode database for invisible characters (like U+200B Zero-Width Space, U+200C ZWNJ, U+200D ZWJ, and U+FEFF BOM) and displays them with a red dashed indicator box.
What is the difference between a Codepoint and a Grapheme Cluster?+
A codepoint is an individual numeric code assigned by Unicode (e.g. U+0065 for "e" and U+0301 for acute accent). A grapheme cluster is the user-perceived visual character combining the base letter and accent (e.g. "é").
Can I jump directly to the detailed character database page for any codepoint?+
Yes! Click the "Lookup" button on any row in the forensic table to open the full character page (/characters/[codepoint]) with complete Unicode block properties and bi-directional classification.