A character encoding is the rule that turns characters into bytes and back. Use this sheet to pick an encoding, declare it in HTTP and HTML, and recognize which wrong decoding produced a garbled string. The main confusion it clears up: text is only readable when the decoder uses the same encoding the writer used, and the bytes themselves never say which that was.
| Encoding | Bytes per character | Covers | Notes |
|---|---|---|---|
| ASCII | 1 | 128 characters | Subset of all encodings below except UTF-16 and UTF-32 |
| ISO-8859-1 (Latin-1) | 1 | 256 characters | Western European letters |
| Windows-1252 | 1 | 256 byte values; Python's codec leaves 5 undefined | Adds the euro sign, curly quotes and dashes at 0x80 to 0x9F; undefined in Python: 0x81, 0x8D, 0x8F, 0x90, 0x9D |
| UTF-8 | 1 to 4 | All of Unicode | Default for the web |
| UTF-16 | 2 or 4 | All of Unicode | Used by JavaScript strings and Windows APIs |
| UTF-32 | 4 | All of Unicode | Fixed width, rarely used for storage |
| Character | Code point | UTF-8 | UTF-16 BE | Latin-1 | Windows-1252 |
|---|---|---|---|---|---|
| A | U+0041 | 41 | 00 41 | 41 | 41 |
| é | U+00E9 | c3 a9 | 00 e9 | e9 | e9 |
| € | U+20AC | e2 82 ac | 20 ac | not encodable | 80 |
| 中 | U+4E2D | e4 b8 ad | 4e 2d | not encodable | not encodable |
| 𝄞 | U+1D11E | f0 9d 84 9e | d8 34 dd 1e | not encodable | not encodable |
| Place | Syntax | Rule |
|---|---|---|
| HTTP header | `Content-Type: text/html; charset=utf-8` | Charset names are case-insensitive (RFC 9110, section 8.3.2) |
| HTML | `<meta charset="utf-8">` | Must appear within the first 1024 bytes of the document |
| XML | `<?xml version="1.0" encoding="UTF-8"?>` | Optional for UTF-8 and UTF-16 |
| Byte order mark | EF BB BF for UTF-8, FF FE or FE FF for UTF-16 | A signature at the start of a file |
| Order | Source | Confidence |
|---|---|---|
| 1 | Byte order mark at the start of the file | Certain |
| 2 | User override in the browser | Certain |
| 3 | Charset in the HTTP Content-Type header | Certain |
| 4 | Meta tag found by a prescan of the first 1024 bytes | Tentative |
So a header saying charset=iso-8859-1 beats a meta tag saying UTF-8, and a BOM beats both.
| Encoding | BOM bytes |
|---|---|
| UTF-8 | EF BB BF |
| UTF-16 little endian | FF FE |
| UTF-16 big endian | FE FF |
| UTF-32 little endian | FF FE 00 00 |
| UTF-32 big endian | 00 00 FE FF |
| You see | Should be | Cause |
|---|---|---|
| `é` | é | UTF-8 bytes read as Windows-1252 or Latin-1 |
| `’` | ’ | Curly apostrophe, UTF-8 read as Windows-1252 |
| `€` | € | Euro sign, UTF-8 read as Windows-1252 |
| `caf\xe9` then a decode error | café | Latin-1 bytes read as UTF-8 |
| `�` | any | The decoder replaced bytes it could not read |
| Situation | Use | Reason |
|---|---|---|
| New web page, API or database text | UTF-8 | The WHATWG Encoding Standard says new formats should use UTF-8 exclusively |
| Reading an old Western European file | Windows-1252 | It is what browsers use for the Latin-1 labels |
| In-memory strings in JavaScript or Java | UTF-16 | Fixed by the language, not a storage choice |
| Unknown bytes | Detect, then verify by decoding strictly | A guess that decodes without errors is not proof |
import tempfile, os
d = tempfile.mkdtemp(); p = os.path.join(d, "a.txt")
open(p, "w", encoding="latin-1").write("café")
print(open(p, "rb").read())
try: open(p, encoding="utf-8").read()
except Exception as e: print(type(e).__name__, e)
print(open(p, encoding="latin-1").read())
b'caf\xe9'
UnicodeDecodeError 'utf-8' codec can't decode byte 0xe9 in position 3: unexpected end of data
café
Always pass encoding= to open. The default depends on the platform and locale.
print("café".encode("utf-8").decode("latin-1").encode("latin-1").decode("utf-8"))
print(repr("’".encode("utf-8").decode("cp1252")))
café
'’'
Reverse the mistake: encode with the wrong codec to get the original bytes back, then decode as UTF-8. This only works if no bytes were lost along the way.
try { new TextDecoder("utf-8", { fatal: true }).decode(new Uint8Array([0x63, 0xe9])); }
catch (e) { console.log(e.name + ": " + e.message); }
TypeError: The encoded data was not valid for encoding utf-8
Without fatal: true, TextDecoder silently inserts U+FFFD for each bad sequence.
console.log(Buffer.byteLength("é"), Buffer.byteLength("€"), Buffer.byteLength("\u{1D11E}"), "\u{1D11E}".length, [..."\u{1D11E}"].length);
2 3 4 2 1
The musical G clef U+1D11E is 4 UTF-8 bytes, 2 JavaScript string units and 1 code point.
printf 'caf\xc3\xa9\n' | iconv -f UTF-8 -t ISO-8859-1 | od -An -tx1
printf 'caf\xe9\n' | iconv -f ISO-8859-1 -t UTF-8
63 61 66 e9 0a
café
iconv fails with a non-zero exit code on unconvertible input, such as iconv: illegal input sequence at position 1, instead of corrupting the data.
printf 'caf\xc3\xa9' > t.txt && file -i t.txt
printf 'caf\xe9' > t2.txt && file -i t2.txt
t.txt: text/plain; charset=utf-8
t2.txt: text/plain; charset=iso-8859-1
file -i guesses from the bytes. The second file could equally be Windows-1252, so treat the result as a hint.
echo bin2hex(mb_convert_encoding("é", "ISO-8859-1", "UTF-8"));
e9
The arguments are target encoding first, then the source encoding.
import tempfile, os
d = tempfile.mkdtemp(); p = os.path.join(d, "a.txt")
open(p, "w", encoding="utf-8-sig").write("hi")
print(open(p, "rb").read(), repr(open(p, encoding="utf-8").read()), repr(open(p, encoding="utf-8-sig").read()))
b'\xef\xbb\xbfhi' '\ufeffhi' 'hi'
Read with utf-8-sig to drop the BOM. Plain utf-8 keeps it as an invisible first character.
b"\x80".decode("cp1252") gives €, while the Latin-1 codec gives the control character U+0080.latin1, iso-8859-1 and us-ascii to windows-1252. A charset=iso-8859-1 page therefore shows a euro sign for byte 0x80.latin1 truncates silently: Buffer.from("€", "latin1") prints <Buffer ac>, which keeps only the low byte of U+20AC. Check that text fits in one byte before using it.é for every é. Find the step that decodes with the wrong charset instead of patching the output.#! line and can make JSON parsers fail. RFC 3629 notes the UTF-8 BOM is always the three bytes EF BB BF and is useless as a byte-order hint.strlen("é") returns 2 and mb_strlen("é") returns 1. Use the multibyte functions for user-visible length.'charmap' codec can't decode byte 0x81 in position 0: character maps to <undefined>. Browsers map those five bytes to control characters instead, so a page that works in Chrome can crash your script.b"caf\xe9".decode("utf-8", "replace") returns caf followed by U+FFFD, and the original byte is gone for good. Log the failure or keep the raw bytes instead.<meta charset="utf-8"> first in head.