Glossary

UTF-8

UTF-8 (Unicode Transformation Format, 8-bit) is a variable-length byte encoding for Unicode code points, using 1 to 4 bytes per character. It's specified in RFC 3629 and is by far the dominant text encoding on the web — the WHATWG and W3C both mandate it as the default for HTML and JSON.

How it works

UTF-8 encodes each Unicode code point using the smallest number of bytes it needs:

  • 1 byte (0xxxxxxx) — code points U+0000–U+007F, exactly the original 7-bit ASCII range, byte-for-byte identical to ASCII.
  • 2 bytes — U+0080–U+07FF (most Latin-script accented letters, Greek, Cyrillic, Hebrew, Arabic).
  • 3 bytes — U+0800–U+FFFF (most CJK characters, and most of the Basic Multilingual Plane).
  • 4 bytes — U+10000–U+10FFFF (emoji and rarer historic scripts).

Every continuation byte in a multi-byte sequence starts with the bit pattern 10xxxxxx, which is what lets a decoder resynchronize mid-stream even if it starts reading from a random byte offset. Because the ASCII range is untouched, any valid ASCII text is already valid UTF-8 — this backward compatibility is a major reason UTF-8 won out over UTF-16 for web content.

Example:

'A'  (U+0041) → 0x41                      (1 byte, same as ASCII)
'é'  (U+00E9) → 0xC3 0xA9                 (2 bytes)
'😀' (U+1F600)→ 0xF0 0x9F 0x98 0x80       (4 bytes)

Common pitfalls

  • Truncating a UTF-8 string at a fixed byte length can slice a multi-byte character in half, producing invalid bytes at the cut point — truncate on code point or grapheme boundaries instead.
  • String.length in JavaScript counts UTF-16 code units, not UTF-8 bytes or even Unicode code points — emoji outside the Basic Multilingual Plane count as 2.
  • Reading a file with the wrong declared encoding (e.g. Latin-1 bytes interpreted as UTF-8) produces mojibake — garbled but "successfully decoded" text, not an error.
  • An optional UTF-8 byte-order mark (EF BB BF) at the start of a file is unnecessary (UTF-8 has no byte-order ambiguity) but some tools add it anyway, which can break strict parsers expecting the file to start with, say, {.

Related terms

  • Unicode — the character standard UTF-8 encodes; UTF-8 is one of several ways to turn Unicode code points into bytes.
  • Base64 — commonly applied after converting text to UTF-8 bytes, since Base64 itself only knows about bytes, not characters.
  • URL encoding — percent-encodes the individual bytes UTF-8 produces for non-ASCII characters.

See also