Cheatsheet

Unicode Cheatsheet

Unicode assigns every character a number called a code point, written like U+20AC. This sheet is for developers who need the ranges, the byte layouts, the escape syntax of each language and the string-length traps. The confusion it clears up: a code point, a UTF-16 unit, a byte and what a user sees as one character are four different things.

Quick reference

Code space

Fact Value
Code point range U+0000 to U+10FFFF
Total code points 1,114,112
Planes 17, numbered 0 to 16, each holding 65,536 code points
Characters assigned in Unicode 17.0 159,801
Surrogate range U+D800 to U+DFFF (2,048 code points, never valid characters on their own)
Replacement character U+FFFD

Planes

Plane Range Name Holds
0 U+0000 to U+FFFF Basic Multilingual Plane (BMP) Most modern scripts, common symbols
1 U+10000 to U+1FFFF Supplementary Multilingual Plane (SMP) Historic scripts, musical symbols, emoji
2 U+20000 to U+2FFFF Supplementary Ideographic Plane (SIP) Rare CJK ideographs
3 U+30000 to U+3FFFF Tertiary Ideographic Plane (TIP) More CJK ideographs
14 U+E0000 to U+EFFFF Supplementary Special-purpose Plane Tag characters and variation selectors
15 and 16 U+F0000 to U+10FFFF Private Use planes Defined by agreement between parties

UTF-8 byte layout

Code point range Bytes Bit pattern
U+0000 to U+007F 1 `0xxxxxxx`
U+0080 to U+07FF 2 `110xxxxx 10xxxxxx`
U+0800 to U+FFFF 3 `1110xxxx 10xxxxxx 10xxxxxx`
U+10000 to U+10FFFF 4 `11110xxx 10xxxxxx 10xxxxxx 10xxxxxx`

UTF-16 rule

Code point Encoded as
Below U+10000 One 16-bit unit, equal to the code point
U+10000 and above Subtract 0x10000, then add the top 10 bits to 0xD800 and the low 10 bits to 0xDC00

Writing a code point in source code

Language Syntax for U+1D11E Notes
JavaScript `"\u{1D11E}"` Braces allow up to six hex digits
Python `"\U0001D11E"` or `"\N{MUSICAL SYMBOL G CLEF}"` `\U` needs exactly eight hex digits
PHP 7+ `"\u{1D11E}"` Braces, as in JavaScript
HTML `𝄞` or `𝄞` Hex or decimal reference
CSS `\1D11E` Up to six hex digits, then a space if more hex follows

Invisible and special code points to know

Code point Name (Python 3.11 data) Category Why you meet it
U+FEFF ZERO WIDTH NO-BREAK SPACE Cf Doubles as the byte order mark at the start of a file
U+FFFD REPLACEMENT CHARACTER So Shown where a decoder found invalid bytes
U+00A0 NO-BREAK SPACE Zs ` ` in HTML, looks like a space
U+00AD SOFT HYPHEN Cf Invisible unless a line breaks there
U+200B ZERO WIDTH SPACE Cf Pasted from web pages, breaks string equality
U+200D ZERO WIDTH JOINER Cf Glues emoji into one sequence
U+2028 LINE SEPARATOR Zl Extra line terminator beyond LF and CR
U+0301 COMBINING ACUTE ACCENT Mn Adds an accent to the letter before it
U+FE0F VARIATION SELECTOR-16 Mn Requests the emoji look of a character
U+202E RIGHT-TO-LEFT OVERRIDE Cf Reverses display order, used in spoofed filenames

Normalization forms

Form Does Use for
NFC Canonical decomposition, then canonical composition Storage and comparison (the usual choice)
NFD Canonical decomposition Stripping accents
NFKC Compatibility decomposition, then composition Search keys, identifiers
NFKD Compatibility decomposition Aggressive folding

Common patterns

Look up a code point's name, category and bytes

import unicodedata as u
for c in ["A", "€", "中", "\U0001D11E", "​", " "]:
    print(f"U+{ord(c):04X}", u.name(c, "?"), u.category(c), c.encode("utf-8").hex(" "))
U+0041 LATIN CAPITAL LETTER A Lu 41
U+20AC EURO SIGN Sc e2 82 ac
U+4E2D CJK UNIFIED IDEOGRAPH-4E2D Lo e4 b8 ad
U+1D11E MUSICAL SYMBOL G CLEF So f0 9d 84 9e
U+200B ZERO WIDTH SPACE Cf e2 80 8b
U+00A0 NO-BREAK SPACE Zs c2 a0

The zero-width space and no-break space are the two usual suspects when a string "looks equal" but is not.

Compare strings that look identical

import unicodedata as u
a = "é"; b = "é"
print(a == b, len(a), len(b), u.normalize("NFC", b) == a, [hex(ord(x)) for x in u.normalize("NFD", a)])
False 1 2 True ['0x65', '0x301']

Both strings show as the same accented letter, but one is a single code point and the other is two. Normalize to NFC before comparing or hashing.

Convert between a character and its code point

print(chr(0x20AC), ord("\u20ac"), hex(ord("\u20ac")))
€ 8364 0x20ac

JavaScript uses String.fromCodePoint(0x20AC) and "\u20ac".codePointAt(0). PHP uses mb_chr(8364) and mb_ord("\u{20AC}"), both from the mbstring extension. Avoid fromCharCode for anything above U+FFFF, because it takes UTF-16 units.

Split a surrogate pair in JavaScript

const s = "\u{1D11E}";
console.log(s.length, s.codePointAt(0).toString(16), s.charCodeAt(0).toString(16), s.charCodeAt(1).toString(16));
2 1d11e d834 dd1e

length counts UTF-16 units, so any character above U+FFFF counts as 2. codePointAt returns the whole code point.

Count code points, UTF-16 units and visible characters

const fam = "\u{1F468}‍\u{1F469}‍\u{1F467}";
console.log(fam.length, [...fam].length, [...new Intl.Segmenter().segment(fam)].length);
8 5 1

A family sequence is 8 units, 5 code points and 1 grapheme cluster. Use Intl.Segmenter when you need what the user sees as one character.

Strip accents

console.log("café".normalize("NFD").replace(/[̀-ͯ]/g, ""));
cafe

Decompose, then delete the combining marks in U+0300 to U+036F.

Match by script or category in a regex

console.log(/^\p{L}+$/u.test("héllo"), /\p{Script=Han}/u.test("中"), "\u{1D11E}".match(/./g).length, "\u{1D11E}".match(/./gu).length);
true true 2 1

Without the u flag, . matches one UTF-16 unit and splits astral characters. Add u whenever you use \p{...} or need whole code points.

Fold compatibility characters

import unicodedata as u
print(u.normalize("NFKC", "fi"), u.normalize("NFKC", "①"), u.normalize("NFC", "fi"))
fi 1 fi

NFKC turns the ligature U+FB01 into fi and the circled digit U+2460 into 1. NFC leaves the ligature alone.

Pitfalls

  • String length is not character count: in JavaScript "\u{1D11E}".length is 2. In Python it is 1. Know whether your language counts UTF-16 units, code points or bytes.
  • Lone surrogates are not valid text: encodeURIComponent("\ud800") throws URIError: URI malformed. Use isWellFormed() to test a string and toWellFormed() to replace lone surrogates with U+FFFD.
  • Case mapping changes length: "ß".upper() in Python gives SS, and "ß".casefold() gives ss. Do not assume upper and lower case strings have the same length, and compare with casefold instead of lower.
  • Locale-blind sorting: JavaScript's default sort() orders by UTF-16 unit, so ["ä","z","a"] becomes [ 'a', 'z', 'ä' ]. Pass localeCompare or Intl.Collator to get [ 'a', 'ä', 'z' ]. Swedish rules put ä after z.
  • Invisible characters in data: U+200B (zero-width space) and U+00A0 survive copy and paste and break equality checks, trimming and parsers. Log code points with ord when a string misbehaves.
  • Regex without the u flag: /./ matches half of a character above U+FFFF, and \p{L} silently matches the literal text p{L} instead of any letter. Always add u to patterns that handle user text.
  • Assuming the BMP is enough: code written for 16-bit units works until the first character above U+FFFF, such as a musical symbol or a rare ideograph. Test with at least one four-byte UTF-8 character.
  • Comparing without normalizing: the same visible text can arrive as NFC from one system and NFD from another, for example from file names on different operating systems. Normalize both sides to NFC before comparing or hashing.
  • Using "Unicode" for one encoding: Unicode is the character set. UTF-8, UTF-16 and UTF-32 are encodings of it. On Windows the label "Unicode" in a save dialog often means UTF-16.
  • Truncating by code point still splits emoji: cutting a string after N code points can slice a family sequence or a flag in the middle. Cut on grapheme boundaries.

Related ZipKit tools

Related cheatsheets