Glossary

Unicode

Unicode is a character encoding standard that assigns a unique numeric identifier — a code point, written U+0041 for the letter "A" — to essentially every character used in human writing systems, plus symbols, emoji, and control characters. It's maintained by the Unicode Consortium and underlies almost all modern text processing, replacing the patchwork of incompatible regional encodings (ASCII, Latin-1, Shift-JIS, and dozens more) that came before it.

How it works

Unicode itself only defines which number means which character — it doesn't say how those numbers are stored as bytes. Code points range from U+0000 to U+10FFFF, organized into 17 "planes" of 65,536 code points each; the first plane (the Basic Multilingual Plane) covers most living scripts, while later planes hold historic scripts, emoji, and specialized symbols. Turning code points into actual bytes is the job of a Unicode Transformation Format — UTF-8, UTF-16, or UTF-32 — each trading off byte size and compatibility differently.

Example:

'A'  → U+0041
'é'  → U+00E9
'😀' → U+1F600  (outside the Basic Multilingual Plane)

Common pitfalls

  • "Unicode" is not an encoding by itself — saying a file is "encoded in Unicode" is ambiguous; the actual byte encoding (UTF-8, UTF-16, etc.) needs to be specified.
  • A single visible character can sometimes be represented by more than one code point sequence (e.g. an accented letter as one precomposed code point vs. a base letter plus a combining accent) — string comparison without normalization can treat visually identical text as different.
  • Counting "characters" is ambiguous: a code point, a UTF-16 code unit, and a user-perceived "grapheme" (like a flag emoji built from two code points) can all give different lengths for the same string.
  • Sorting text correctly requires locale-aware collation — naive code-point ordering doesn't match how any human language actually alphabetizes.

Related terms

  • UTF-8 — the most common way Unicode code points are actually encoded as bytes.
  • URL encoding — operates on the UTF-8 bytes of Unicode text, one byte at a time, when a URL needs to carry non-ASCII characters.

See also