Unicode assigns every character a number called a code point, written like U+20AC. This sheet is for developers who need the ranges, the byte layouts, the escape syntax of each language and the string-length traps. The confusion it clears up: a code point, a UTF-16 unit, a byte and what a user sees as one character are four different things.
| Fact | Value |
|---|---|
| Code point range | U+0000 to U+10FFFF |
| Total code points | 1,114,112 |
| Planes | 17, numbered 0 to 16, each holding 65,536 code points |
| Characters assigned in Unicode 17.0 | 159,801 |
| Surrogate range | U+D800 to U+DFFF (2,048 code points, never valid characters on their own) |
| Replacement character | U+FFFD |
| Plane | Range | Name | Holds |
|---|---|---|---|
| 0 | U+0000 to U+FFFF | Basic Multilingual Plane (BMP) | Most modern scripts, common symbols |
| 1 | U+10000 to U+1FFFF | Supplementary Multilingual Plane (SMP) | Historic scripts, musical symbols, emoji |
| 2 | U+20000 to U+2FFFF | Supplementary Ideographic Plane (SIP) | Rare CJK ideographs |
| 3 | U+30000 to U+3FFFF | Tertiary Ideographic Plane (TIP) | More CJK ideographs |
| 14 | U+E0000 to U+EFFFF | Supplementary Special-purpose Plane | Tag characters and variation selectors |
| 15 and 16 | U+F0000 to U+10FFFF | Private Use planes | Defined by agreement between parties |
| Code point range | Bytes | Bit pattern |
|---|---|---|
| U+0000 to U+007F | 1 | `0xxxxxxx` |
| U+0080 to U+07FF | 2 | `110xxxxx 10xxxxxx` |
| U+0800 to U+FFFF | 3 | `1110xxxx 10xxxxxx 10xxxxxx` |
| U+10000 to U+10FFFF | 4 | `11110xxx 10xxxxxx 10xxxxxx 10xxxxxx` |
| Code point | Encoded as |
|---|---|
| Below U+10000 | One 16-bit unit, equal to the code point |
| U+10000 and above | Subtract 0x10000, then add the top 10 bits to 0xD800 and the low 10 bits to 0xDC00 |
| Language | Syntax for U+1D11E | Notes |
|---|---|---|
| JavaScript | `"\u{1D11E}"` | Braces allow up to six hex digits |
| Python | `"\U0001D11E"` or `"\N{MUSICAL SYMBOL G CLEF}"` | `\U` needs exactly eight hex digits |
| PHP 7+ | `"\u{1D11E}"` | Braces, as in JavaScript |
| HTML | `𝄞` or `𝄞` | Hex or decimal reference |
| CSS | `\1D11E` | Up to six hex digits, then a space if more hex follows |
| Code point | Name (Python 3.11 data) | Category | Why you meet it |
|---|---|---|---|
| U+FEFF | ZERO WIDTH NO-BREAK SPACE | Cf | Doubles as the byte order mark at the start of a file |
| U+FFFD | REPLACEMENT CHARACTER | So | Shown where a decoder found invalid bytes |
| U+00A0 | NO-BREAK SPACE | Zs | ` ` in HTML, looks like a space |
| U+00AD | SOFT HYPHEN | Cf | Invisible unless a line breaks there |
| U+200B | ZERO WIDTH SPACE | Cf | Pasted from web pages, breaks string equality |
| U+200D | ZERO WIDTH JOINER | Cf | Glues emoji into one sequence |
| U+2028 | LINE SEPARATOR | Zl | Extra line terminator beyond LF and CR |
| U+0301 | COMBINING ACUTE ACCENT | Mn | Adds an accent to the letter before it |
| U+FE0F | VARIATION SELECTOR-16 | Mn | Requests the emoji look of a character |
| U+202E | RIGHT-TO-LEFT OVERRIDE | Cf | Reverses display order, used in spoofed filenames |
| Form | Does | Use for |
|---|---|---|
| NFC | Canonical decomposition, then canonical composition | Storage and comparison (the usual choice) |
| NFD | Canonical decomposition | Stripping accents |
| NFKC | Compatibility decomposition, then composition | Search keys, identifiers |
| NFKD | Compatibility decomposition | Aggressive folding |
import unicodedata as u
for c in ["A", "€", "中", "\U0001D11E", "", " "]:
print(f"U+{ord(c):04X}", u.name(c, "?"), u.category(c), c.encode("utf-8").hex(" "))
U+0041 LATIN CAPITAL LETTER A Lu 41
U+20AC EURO SIGN Sc e2 82 ac
U+4E2D CJK UNIFIED IDEOGRAPH-4E2D Lo e4 b8 ad
U+1D11E MUSICAL SYMBOL G CLEF So f0 9d 84 9e
U+200B ZERO WIDTH SPACE Cf e2 80 8b
U+00A0 NO-BREAK SPACE Zs c2 a0
The zero-width space and no-break space are the two usual suspects when a string "looks equal" but is not.
import unicodedata as u
a = "é"; b = "é"
print(a == b, len(a), len(b), u.normalize("NFC", b) == a, [hex(ord(x)) for x in u.normalize("NFD", a)])
False 1 2 True ['0x65', '0x301']
Both strings show as the same accented letter, but one is a single code point and the other is two. Normalize to NFC before comparing or hashing.
print(chr(0x20AC), ord("\u20ac"), hex(ord("\u20ac")))
€ 8364 0x20ac
JavaScript uses String.fromCodePoint(0x20AC) and "\u20ac".codePointAt(0). PHP uses mb_chr(8364) and mb_ord("\u{20AC}"), both from the mbstring extension. Avoid fromCharCode for anything above U+FFFF, because it takes UTF-16 units.
const s = "\u{1D11E}";
console.log(s.length, s.codePointAt(0).toString(16), s.charCodeAt(0).toString(16), s.charCodeAt(1).toString(16));
2 1d11e d834 dd1e
length counts UTF-16 units, so any character above U+FFFF counts as 2. codePointAt returns the whole code point.
const fam = "\u{1F468}\u{1F469}\u{1F467}";
console.log(fam.length, [...fam].length, [...new Intl.Segmenter().segment(fam)].length);
8 5 1
A family sequence is 8 units, 5 code points and 1 grapheme cluster. Use Intl.Segmenter when you need what the user sees as one character.
console.log("café".normalize("NFD").replace(/[̀-ͯ]/g, ""));
cafe
Decompose, then delete the combining marks in U+0300 to U+036F.
console.log(/^\p{L}+$/u.test("héllo"), /\p{Script=Han}/u.test("中"), "\u{1D11E}".match(/./g).length, "\u{1D11E}".match(/./gu).length);
true true 2 1
Without the u flag, . matches one UTF-16 unit and splits astral characters. Add u whenever you use \p{...} or need whole code points.
import unicodedata as u
print(u.normalize("NFKC", "fi"), u.normalize("NFKC", "①"), u.normalize("NFC", "fi"))
fi 1 fi
NFKC turns the ligature U+FB01 into fi and the circled digit U+2460 into 1. NFC leaves the ligature alone.
"\u{1D11E}".length is 2. In Python it is 1. Know whether your language counts UTF-16 units, code points or bytes.encodeURIComponent("\ud800") throws URIError: URI malformed. Use isWellFormed() to test a string and toWellFormed() to replace lone surrogates with U+FFFD."ß".upper() in Python gives SS, and "ß".casefold() gives ss. Do not assume upper and lower case strings have the same length, and compare with casefold instead of lower.sort() orders by UTF-16 unit, so ["ä","z","a"] becomes [ 'a', 'z', 'ä' ]. Pass localeCompare or Intl.Collator to get [ 'a', 'ä', 'z' ]. Swedish rules put ä after z.ord when a string misbehaves.u flag: /./ matches half of a character above U+FFFF, and \p{L} silently matches the literal text p{L} instead of any letter. Always add u to patterns that handle user text.normalize, codePointAt and related string methods.