Cheatsheet

Character Encoding Cheatsheet

A character encoding is the rule that turns characters into bytes and back. Use this sheet to pick an encoding, declare it in HTTP and HTML, and recognize which wrong decoding produced a garbled string. The main confusion it clears up: text is only readable when the decoder uses the same encoding the writer used, and the bytes themselves never say which that was.

Quick reference

Encodings compared

Encoding Bytes per character Covers Notes
ASCII 1 128 characters Subset of all encodings below except UTF-16 and UTF-32
ISO-8859-1 (Latin-1) 1 256 characters Western European letters
Windows-1252 1 256 byte values; Python's codec leaves 5 undefined Adds the euro sign, curly quotes and dashes at 0x80 to 0x9F; undefined in Python: 0x81, 0x8D, 0x8F, 0x90, 0x9D
UTF-8 1 to 4 All of Unicode Default for the web
UTF-16 2 or 4 All of Unicode Used by JavaScript strings and Windows APIs
UTF-32 4 All of Unicode Fixed width, rarely used for storage

The same character in each encoding

Character Code point UTF-8 UTF-16 BE Latin-1 Windows-1252
A U+0041 41 00 41 41 41
é U+00E9 c3 a9 00 e9 e9 e9
€ U+20AC e2 82 ac 20 ac not encodable 80
中 U+4E2D e4 b8 ad 4e 2d not encodable not encodable
𝄞 U+1D11E f0 9d 84 9e d8 34 dd 1e not encodable not encodable

Where the encoding is declared

Place Syntax Rule
HTTP header `Content-Type: text/html; charset=utf-8` Charset names are case-insensitive (RFC 9110, section 8.3.2)
HTML `<meta charset="utf-8">` Must appear within the first 1024 bytes of the document
XML `<?xml version="1.0" encoding="UTF-8"?>` Optional for UTF-8 and UTF-16
Byte order mark EF BB BF for UTF-8, FF FE or FE FF for UTF-16 A signature at the start of a file

Which declaration a browser trusts first

Order Source Confidence
1 Byte order mark at the start of the file Certain
2 User override in the browser Certain
3 Charset in the HTTP Content-Type header Certain
4 Meta tag found by a prescan of the first 1024 bytes Tentative

So a header saying charset=iso-8859-1 beats a meta tag saying UTF-8, and a BOM beats both.

Byte order marks

Encoding BOM bytes
UTF-8 EF BB BF
UTF-16 little endian FF FE
UTF-16 big endian FE FF
UTF-32 little endian FF FE 00 00
UTF-32 big endian 00 00 FE FF

Garbled text and its cause

You see Should be Cause
`é` é UTF-8 bytes read as Windows-1252 or Latin-1
`’` ’ Curly apostrophe, UTF-8 read as Windows-1252
`€` € Euro sign, UTF-8 read as Windows-1252
`caf\xe9` then a decode error café Latin-1 bytes read as UTF-8
`�` any The decoder replaced bytes it could not read

Which encoding should I use

Situation Use Reason
New web page, API or database text UTF-8 The WHATWG Encoding Standard says new formats should use UTF-8 exclusively
Reading an old Western European file Windows-1252 It is what browsers use for the Latin-1 labels
In-memory strings in JavaScript or Java UTF-16 Fixed by the language, not a storage choice
Unknown bytes Detect, then verify by decoding strictly A guess that decodes without errors is not proof

Common patterns

Read and write a file with an explicit encoding

import tempfile, os
d = tempfile.mkdtemp(); p = os.path.join(d, "a.txt")
open(p, "w", encoding="latin-1").write("café")
print(open(p, "rb").read())
try: open(p, encoding="utf-8").read()
except Exception as e: print(type(e).__name__, e)
print(open(p, encoding="latin-1").read())
b'caf\xe9'
UnicodeDecodeError 'utf-8' codec can't decode byte 0xe9 in position 3: unexpected end of data
café

Always pass encoding= to open. The default depends on the platform and locale.

Repair text that was decoded with the wrong encoding

print("café".encode("utf-8").decode("latin-1").encode("latin-1").decode("utf-8"))
print(repr("’".encode("utf-8").decode("cp1252")))
café
'’'

Reverse the mistake: encode with the wrong codec to get the original bytes back, then decode as UTF-8. This only works if no bytes were lost along the way.

Reject invalid bytes instead of guessing

try { new TextDecoder("utf-8", { fatal: true }).decode(new Uint8Array([0x63, 0xe9])); }
catch (e) { console.log(e.name + ": " + e.message); }
TypeError: The encoded data was not valid for encoding utf-8

Without fatal: true, TextDecoder silently inserts U+FFFD for each bad sequence.

Count bytes, UTF-16 units and code points

console.log(Buffer.byteLength("é"), Buffer.byteLength("€"), Buffer.byteLength("\u{1D11E}"), "\u{1D11E}".length, [..."\u{1D11E}"].length);
2 3 4 2 1

The musical G clef U+1D11E is 4 UTF-8 bytes, 2 JavaScript string units and 1 code point.

Convert a file between encodings with iconv

printf 'caf\xc3\xa9\n' | iconv -f UTF-8 -t ISO-8859-1 | od -An -tx1
printf 'caf\xe9\n' | iconv -f ISO-8859-1 -t UTF-8
 63 61 66 e9 0a
café

iconv fails with a non-zero exit code on unconvertible input, such as iconv: illegal input sequence at position 1, instead of corrupting the data.

Detect the charset of a file

printf 'caf\xc3\xa9' > t.txt && file -i t.txt
printf 'caf\xe9' > t2.txt && file -i t2.txt
t.txt: text/plain; charset=utf-8
t2.txt: text/plain; charset=iso-8859-1

file -i guesses from the bytes. The second file could equally be Windows-1252, so treat the result as a hint.

Convert a string between encodings in PHP

echo bin2hex(mb_convert_encoding("é", "ISO-8859-1", "UTF-8"));
e9

The arguments are target encoding first, then the source encoding.

Handle a UTF-8 BOM

import tempfile, os
d = tempfile.mkdtemp(); p = os.path.join(d, "a.txt")
open(p, "w", encoding="utf-8-sig").write("hi")
print(open(p, "rb").read(), repr(open(p, encoding="utf-8").read()), repr(open(p, encoding="utf-8-sig").read()))
b'\xef\xbb\xbfhi' '\ufeffhi' 'hi'

Read with utf-8-sig to drop the BOM. Plain utf-8 keeps it as an invisible first character.

Pitfalls

  • Latin-1 and Windows-1252 are not the same: bytes 0x80 to 0x9F are control codes in ISO-8859-1 but printable symbols in Windows-1252. Python's b"\x80".decode("cp1252") gives €, while the Latin-1 codec gives the control character U+0080.
  • Browsers treat the Latin-1 label as Windows-1252: the WHATWG Encoding Standard maps latin1, iso-8859-1 and us-ascii to windows-1252. A charset=iso-8859-1 page therefore shows a euro sign for byte 0x80.
  • Node.js latin1 truncates silently: Buffer.from("€", "latin1") prints <Buffer ac>, which keeps only the low byte of U+20AC. Check that text fits in one byte before using it.
  • Encoding twice: UTF-8 bytes that are decoded as Latin-1 and then encoded again as UTF-8 give é for every é. Find the step that decodes with the wrong charset instead of patching the output.
  • A BOM in JSON or shell scripts: the BOM breaks a #! line and can make JSON parsers fail. RFC 3629 notes the UTF-8 BOM is always the three bytes EF BB BF and is useless as a byte-order hint.
  • Counting characters with byte length: PHP's strlen("é") returns 2 and mb_strlen("é") returns 1. Use the multibyte functions for user-visible length.
  • Cutting a UTF-8 string in the middle of a character: slicing bytes at a fixed length can split a 2 to 4 byte sequence and produce U+FFFD or a decode error. Slice by characters, or back up to a byte that does not start with the bits 10.
  • Python's cp1252 is stricter than the browser: decoding byte 0x81 raises 'charmap' codec can't decode byte 0x81 in position 0: character maps to <undefined>. Browsers map those five bytes to control characters instead, so a page that works in Chrome can crash your script.
  • Hiding errors with replace: b"caf\xe9".decode("utf-8", "replace") returns caf followed by U+FFFD, and the original byte is gone for good. Log the failure or keep the raw bytes instead.
  • No declaration at all: without a charset in the header or a meta tag, the browser guesses from the document and the locale. Put <meta charset="utf-8"> first in head.

Related ZipKit tools

Related cheatsheets