October 6, 2026
How to Validate UTF-8 Encoding
Learn how UTF-8 encoding works, what invalid sequences look like, and how to validate text before it breaks your pipeline.
UTF-8 is the default text encoding on the web and in most modern systems. Validating UTF-8 means checking that a byte sequence follows the encoding rules—so decoders, databases, and UIs do not choke on corrupt input.
What UTF-8 actually encodes
UTF-8 maps Unicode code points to one–four bytes:
| Code points | Bytes | Pattern (binary) |
|---|---|---|
| U+0000–U+007F | 1 | 0xxxxxxx |
| U+0080–U+07FF | 2 | 110xxxxx 10xxxxxx |
| U+0800–U+FFFF | 3 | 1110xxxx 10xxxxxx 10xxxxxx |
| U+10000–U+10FFFF | 4 | 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx |
ASCII is valid UTF-8 by design. Multi-byte characters use a leading byte plus continuation bytes (10xxxxxx).
Common ways UTF-8 goes wrong
Invalid UTF-8 usually comes from:
- Truncated multi-byte sequences — cutting a string mid-character
- Unexpected continuation bytes — a
10xxxxxxbyte with no valid lead - Overlong encodings — encoding ASCII with extra bytes (forbidden)
- Surrogate halves — UTF-16 leftovers treated as Unicode scalars
- Wrong encoding labeled as UTF-8 — Latin-1/Windows-1252 bytes mislabeled
Symptoms include replacement characters (�), mojibake (é instead of é), failed JSON parses, or database rejection.
How to validate UTF-8
1. Treat input as bytes
Validation is about bytes, not characters you already decoded. If a language has already turned bytes into a string with a lossy decoder, you may have lost the evidence.
2. Decode strictly
Use a strict UTF-8 decoder that errors on illegal sequences instead of inserting �. In pipelines, fail closed when encoding must be trusted (security-sensitive parsers, signatures, filenames).
3. Inspect suspects in hex
When validation fails, convert the bytes to hex to see the broken sequence. The Convert UTF-8 to Hex and Convert Hex to UTF-8 tools help you inspect and round-trip samples.
4. Validate in the browser
Paste suspicious text or decoded output into Validate UTF8 to check whether the content is well-formed and to surface encoding issues early.
Validation checklist
- Confirm the producer’s declared charset (headers, file metadata, DB collation).
- Decode with a strict UTF-8 checker.
- Reject or quarantine invalid byte sequences.
- Normalize only after validation (NFC/NFKC is a separate step).
- Re-encode explicitly when crossing system boundaries.
Related encoding tools
- Validate UTF8 — check whether text is valid UTF-8
- Convert Unicode to UTF-8 — move from code points to UTF-8 bytes
- Convert UTF-8 to Hex / Convert Hex to UTF-8 — debug byte-level issues
- Browse more in the text encoding tools collection
Summary
UTF-8 validation protects parsers and users from corrupt or mislabeled text. Decode strictly, inspect failures in hex, and fix the producer when possible. Use Validate UTF8 whenever you need a quick, local check before data enters a stricter system.
Related tools
Try these browser tools connected to this article.
Validate UTF8
Check if text is valid UTF8 encoded
Convert UTF8 to Hex
Convert UTF-8 text to hexadecimal representation for debugging and analysis
Convert Hex to UTF8
Convert hexadecimal representation back to UTF-8 text
Convert Unicode to UTF8
Convert Unicode code points to UTF8 encoded bytes