Cheatsheet

URL Encoding Reference

# URL Encoding Reference

URL (percent) encoding represents characters that aren't safe in a URL as % followed by two hex digits. This sheet covers RFC 3986's reserved/unreserved character sets and the JavaScript functions that implement them — which disagree with each other more than people expect.

Quick reference

Unreserved characters (never encoded)

A-Z a-z 0-9 - _ . ~

These are the only characters RFC 3986 guarantees will never be percent-encoded by a conforming encoder.

Reserved characters and their roles

Char Role in a URL Encoded form
`:` Scheme separator (`https:`) `%3A`
`/` Path separator `%2F`
`?` Query string start `%3F`
`#` Fragment start `%23`
`&` Query parameter separator `%26`
`=` Query key/value separator `%3D`
`+` Historically means *space* inside `application/x-www-form-urlencoded` bodies `%2B`
` ` (space) Never allowed literally `%20` (URLs) or `+` (form bodies only)
`@` Userinfo separator `%40`
`%` The escape character itself `%25`

JavaScript's four (not-quite-consistent) encoders

Function Encodes Leaves untouched
`encodeURIComponent()` Everything except unreserved chars `A-Za-z0-9 - _ . ! ~ * ' ( )`
`encodeURI()` Only unsafe chars, preserves URL structure Also leaves `! # $ & ' ( ) * + , - . / : ; = ? @ ~` untouched
`decodeURIComponent()` Decodes `%XX` sequences Throws `URIError` on malformed sequences
`decodeURI()` Decodes `%XX`, except reserved chars used structurally Throws on malformed sequences

The practical rule: use encodeURIComponent() on a single value going into a query string or path segment; use encodeURI() only when you have a whole URL and want to leave its /, ?, &, = structure intact.

Common patterns

Building a query string safely

const params = new URLSearchParams({ q: 'a & b', page: '2' });
params.toString(); // "q=a+%26+b&page=2"

URLSearchParams encodes spaces as + (form-encoding convention), not %20 — that's correct for query strings, but don't reuse its output for a path segment.

Encoding one value for a path segment

const slug = encodeURIComponent('café / résumé');
// "caf%C3%A9%20%2F%20r%C3%A9sum%C3%A9"
`/articles/${slug}`;

Round-tripping safely

try {
  decodeURIComponent(input);
} catch (e) {
  // malformed percent-escape, e.g. a lone "%" or "%zz"
}

PHP equivalents

rawurlencode($value);   // RFC 3986 -- space becomes %20 (use this for paths/components)
urlencode($value);      // application/x-www-form-urlencoded -- space becomes +
rawurldecode($encoded);
urldecode($encoded);

PHP's urlencode()/urldecode() are for form bodies (+ for space); rawurlencode()/rawurldecode() match encodeURIComponent()'s behavior. Mixing the two up is a frequent source of stray + characters showing up as literal pluses instead of spaces.

Pitfalls

  • encodeURI() will not escape &, =, ?, or #: it's designed to encode a complete URL, so it deliberately leaves structural characters alone. Passing a single form value through encodeURI() instead of encodeURIComponent() lets an & in the value silently inject an extra query parameter.
  • + means space only in form-encoded bodies, not in a raw path or fragment: decodeURIComponent('a+b') returns "a+b" — the literal plus — not "a b". Only application/x-www-form-urlencoded parsers (like URLSearchParams) treat + as space.
  • Double-encoding corrupts data silently: encoding an already-encoded string turns %20 into %2520 (the % itself gets re-encoded to %25). This is a common bug when a value passes through two layers that both assume it's still raw.
  • Malformed percent sequences throw, they don't return null: decodeURIComponent('%') or decodeURIComponent('%zz') throws a URIError — always wrap external/user-supplied encoded input in a try/catch before decoding.
  • Unicode characters expand to multiple percent-triplets: a single character like é becomes 2–4 encoded bytes (%C3%A9 in UTF-8), which can silently blow past a fixed-length URL/field limit that was sized assuming 1 character = 1 encoded unit.

Related ZipKit tools

  • URL Encode / Decode — encode/decode strings and inspect exactly which characters changed.
  • Base64 Encode / Decode — the other common text-safe encoding, used when +// need to survive inside a URL via the URL-safe variant.

Related cheatsheets

  • Base64 Reference — the URL-safe variant exists specifically to avoid a second layer of percent-encoding.
  • JSON Syntax Reference — query-string values and JSON string values need overlapping but distinct escaping rules.