Glossary

URL encoding

URL encoding, also called percent-encoding, is a scheme for representing characters inside a URL that would otherwise be unsafe or ambiguous — spaces, &, ?, non-ASCII text — by replacing each such byte with a % followed by its two-digit hexadecimal value. The mechanism is defined in RFC 3986 (URI Generic Syntax) and refined for browsers by the WHATWG URL Standard.

How it works

A URL is restricted to a small set of ASCII characters, split into reserved characters (:/?#[]@!$&'()*+,;=, which have special meaning as URL syntax) and unreserved characters (letters, digits, -._~). Any byte outside the unreserved set — including a reserved character used literally rather than as syntax, and any non-ASCII byte — gets replaced with %XX, where XX is its hex value:

"hello world!"  →  "hello%20world%21"
"café"          →  "caf%C3%A9"   (UTF-8 bytes for é: 0xC3 0xA9)

Non-ASCII text must first be converted to bytes (almost always UTF-8) before each byte is percent-encoded — that's why "é" becomes two %XX triplets, not one. There's one notable inconsistency to be aware of: the generic URL spec encodes a space as %20, but the older application/x-www-form-urlencoded format used by HTML form submissions encodes a space as + instead — the two are not interchangeable, and decoding with the wrong rule silently corrupts spaces or literal + characters.

Common pitfalls

  • Encoding an already-encoded string (%2520 from double-encoding %20) is a common bug when a value passes through more than one URL-encoding layer.
  • Confusing +-for-space (form encoding) with %20-for-space (generic percent-encoding) corrupts either spaces or literal plus signs, depending on which rule the decoder expects.
  • Forgetting to encode a value before inserting it into a query string lets characters like & or = be misread as query syntax, silently breaking or hijacking other parameters.
  • Percent-encoding is not encryption or obfuscation — a %XX-encoded value is trivially decodable and shouldn't be relied on to hide data.

Related terms

  • UTF-8 — the byte encoding non-ASCII text is converted to before each byte can be percent-encoded.
  • Unicode — the character set behind the text being encoded; the code point itself is never percent-encoded directly, only its UTF-8 bytes.

See also

  • Tool: URL Encode / Decode — encode or decode URLs and query strings, handling special characters and Unicode correctly.