UTF-8 byte inspector
Enter or paste a short string to see its Unicode code points and UTF-8 byte sequence. Compare visible text, programming-language string lengths, and the bytes stored in a UTF-8 file.
Inspect text
| Text unit | Code point | UTF-8 bytes (hex) | Byte count |
|---|
How to read the result
- UTF-16 code units are what JavaScript’s
string.lengthcounts. A supplementary character such as😀takes two units. - Code points are Unicode numbers displayed as
U+..... A combining accent or zero-width joiner is its own code point. - UTF-8 bytes are the encoded bytes for the text in this browser. A code point takes one to four bytes.
- Grapheme clusters approximate user-perceived text units. A joined emoji can contain several code points but appear as one cluster. This count is shown when the browser provides
Intl.Segmenter.
This tool encodes the text value supplied by the browser as UTF-8. It does not inspect original file bytes, identify a file’s source encoding, or determine Unicode character names. For those tasks, use the converter with a known source encoding and preserve the original file.
See the UTF-8 specification and the Unicode FAQ on encoding forms for the complete rules.