UTF-8 (Unicode Transformation Format, 8-bit) is a variable-length byte encoding for Unicode code points, using 1 to 4 bytes per character. It's specified in RFC 3629 and is by far the dominant text encoding on the web — the WHATWG and W3C both mandate it as the default for HTML and JSON.
UTF-8 encodes each Unicode code point using the smallest number of bytes it needs:
0xxxxxxx) — code points U+0000–U+007F, exactly the original 7-bit ASCII range, byte-for-byte identical to ASCII.U+0080–U+07FF (most Latin-script accented letters, Greek, Cyrillic, Hebrew, Arabic).U+0800–U+FFFF (most CJK characters, and most of the Basic Multilingual Plane).U+10000–U+10FFFF (emoji and rarer historic scripts).Every continuation byte in a multi-byte sequence starts with the bit pattern 10xxxxxx, which is what lets a decoder resynchronize mid-stream even if it starts reading from a random byte offset. Because the ASCII range is untouched, any valid ASCII text is already valid UTF-8 — this backward compatibility is a major reason UTF-8 won out over UTF-16 for web content.
Example:
'A' (U+0041) → 0x41 (1 byte, same as ASCII)
'é' (U+00E9) → 0xC3 0xA9 (2 bytes)
'😀' (U+1F600)→ 0xF0 0x9F 0x98 0x80 (4 bytes)
String.length in JavaScript counts UTF-16 code units, not UTF-8 bytes or even Unicode code points — emoji outside the Basic Multilingual Plane count as 2.EF BB BF) at the start of a file is unnecessary (UTF-8 has no byte-order ambiguity) but some tools add it anyway, which can break strict parsers expecting the file to start with, say, {.