UTF-8 explained: bytes, characters, and web pages
UTF-8 is the byte encoding used to store and transmit Unicode text. This guide shows what that means in practice, how to read a byte sequence, and where to look when accented letters or emoji appear incorrectly.
Use the UTF-8 byte inspector to see code points and bytes for text you enter, or use the converter to encode text, decode a hexadecimal byte sequence, or convert a small text file in your browser.
Three layers that are easy to confuse
- Unicode assigns numbers called code points to characters. For example, ü is U+00FC.
- UTF-8 turns each code point into one to four bytes. The bytes for ü are
C3 BC. - Decoding and rendering turn those bytes back into text and draw the result with a font. If bytes are decoded with the wrong encoding, the stored data may be fine while the displayed text is garbled.
Unicode is the character standard; UTF-8 is one way to encode Unicode text as bytes. They are related, but they are not interchangeable terms.
What UTF-8 bytes look like
| Text | Unicode code point | UTF-8 bytes (hex) | Bytes |
|---|---|---|---|
| A | U+0041 | 41 | 1 |
| ü | U+00FC | C3 BC | 2 |
| € | U+20AC | E2 82 AC | 3 |
| 😀 | U+1F600 | F0 9F 98 80 | 4 |
UTF-8 preserves ASCII: basic English letters, digits, and punctuation have the same byte values as ASCII. A non-ASCII code point can take more than one byte, so “character count” and “file size in bytes” are different measurements. RFC 3629 defines the valid one-to-four-byte sequences.
A visible character is not always one code point
The word “character” can mean different things in programming. The text é can be stored as U+00E9 or as e followed by a combining acute accent, U+0065 U+0301. Both can look the same. A family emoji may use several code points joined together. A string’s byte count, code-point count, and number of user-perceived characters can therefore all differ.
Unicode normalization can make canonically equivalent sequences consistent for comparison, but it changes the code-point sequence. Choose a normalization form only when your application’s comparison or storage rules call for it; do not use normalization to repair wrongly decoded bytes. See the Unicode normalization FAQ.
Set the encoding consistently on a web page
Save the HTML file as UTF-8, then declare that encoding near the start of the document’s <head>:
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<title>Example</title>
</head>
The web server should send a matching response header, for example Content-Type: text/html; charset=utf-8. A declaration describes how bytes should be interpreted; it does not convert a file saved in a different encoding. Check both the actual bytes and HTTP header when a page is displayed incorrectly. The W3C guide to HTML encoding declarations explains the browser side.
Where encoding problems usually enter
- Source files: an editor saved a file in Windows-1252 or another legacy encoding while the page declares UTF-8.
- HTTP response: the server labels the response with a charset that does not match the file bytes.
- Database connection: the table, connection, or application layer uses a different encoding. In MySQL, use
utf8mb4for the full Unicode range and make sure the connection uses it too. - Repeated conversion: text was decoded with the wrong charset, then saved again. Converting it repeatedly can make recovery harder.
Keep an untouched copy before converting files. If you know the original encoding, convert once from that source encoding to UTF-8 and check representative text before replacing production data. A byte sequence can be valid UTF-8 and still display as nonsense if it was originally misinterpreted.
Choose the next step
- Inspect UTF-8 bytes and Unicode code points
- Convert text, hexadecimal bytes, or a local text file
- Follow a web-page and database troubleshooting checklist
- Compare UTF-8 with UTF-16
Technical references: RFC 3629, the Unicode UTF FAQ, and the W3C encoding declaration guide.