Cheatsheet

HTML Entities Cheatsheet

An HTML entity, formally a character reference, is a way to write a character in markup as &name; or &#number;. This sheet is for developers who escape user text, paste special symbols into templates, or decode & messes. The confusion it clears up: named, decimal and hex references all produce the same character, and only a handful of characters must be escaped.

Quick reference

The characters you must escape

Character Entity Escape it in Why
`&` `&` Text and attribute values Starts every reference
`<` `&lt;` Text Starts a tag
`>` `&gt;` Text, as a habit Closes a tag, rarely dangerous alone
`"` `&quot;` Double-quoted attribute values Ends the value
`'` `&#39;` or `&apos;` Single-quoted attribute values Ends the value

Reference syntax

Form Example Result
Named `&copy;` ©
Decimal `&#169;` ©
Hexadecimal `&#xA9;` ©

The hex form accepts x or X, and the digits are case-insensitive. A decimal or hex reference can name any code point up to U+10FFFF, subject to the rewrites in the last table below.

Common named entities

Entity Character Decimal Hex Works without the semicolon
`&amp;` `&` `&#38;` `&#x26;` yes
`&lt;` `<` `&#60;` `&#x3C;` yes
`&gt;` `>` `&#62;` `&#x3E;` yes
`&quot;` `"` `&#34;` `&#x22;` yes
`&apos;` `'` `&#39;` `&#x27;` no
`&nbsp;` no-break space `&#160;` `&#xA0;` yes
`&copy;` © `&#169;` `&#xA9;` yes
`&reg;` ® `&#174;` `&#xAE;` yes
`&trade;` ™ `&#8482;` `&#x2122;` no
`&euro;` € `&#8364;` `&#x20AC;` no
`&pound;` £ `&#163;` `&#xA3;` yes
`&yen;` ¥ `&#165;` `&#xA5;` yes
`&cent;` ¢ `&#162;` `&#xA2;` yes
`&sect;` § `&#167;` `&#xA7;` yes
`&deg;` ° `&#176;` `&#xB0;` yes
`&plusmn;` ± `&#177;` `&#xB1;` yes
`&times;` × `&#215;` `&#xD7;` yes
`&divide;` ÷ `&#247;` `&#xF7;` yes
`&frac12;` ½ `&#189;` `&#xBD;` yes
`&hellip;` … `&#8230;` `&#x2026;` no
`&mdash;` — `&#8212;` `&#x2014;` no
`&ndash;` – `&#8211;` `&#x2013;` no
`&lsquo;` ‘ `&#8216;` `&#x2018;` no
`&rsquo;` ’ `&#8217;` `&#x2019;` no
`&ldquo;` “ `&#8220;` `&#x201C;` no
`&rdquo;` ” `&#8221;` `&#x201D;` no
`&laquo;` « `&#171;` `&#xAB;` yes
`&raquo;` » `&#187;` `&#xBB;` yes
`&larr;` ← `&#8592;` `&#x2190;` no
`&rarr;` → `&#8594;` `&#x2192;` no
`&bull;` • `&#8226;` `&#x2022;` no
`&middot;` · `&#183;` `&#xB7;` yes

The last column is read from the WHATWG entities list: only 106 legacy names are recognized without the closing semicolon.

Numbers that HTML rewrites

Reference Result Reason
`&#0;` U+FFFD replacement character Null is not allowed
`&#128;` € (U+20AC) Numbers 0x80 to 0x9F are remapped to Windows-1252 characters
`&#xD800;` U+FFFD Surrogate code points are not characters
`&#x110000;` U+FFFD Above the Unicode maximum

Common patterns

Escape text for HTML in JavaScript

const esc = s => s.replace(/&/g, "&amp;").replace(/</g, "&lt;").replace(/>/g, "&gt;").replace(/"/g, "&quot;").replace(/'/g, "&#39;");
console.log(esc(`<a href="x">Tom & 'Jerry'</a>`));
&lt;a href=&quot;x&quot;&gt;Tom &amp; &#39;Jerry&#39;&lt;/a&gt;

Replace the ampersand first. If you do it last, you double-escape the entities you just inserted. In a browser, assigning to textContent avoids the problem because no HTML parsing happens.

Escape and decode in Python

import html
print(html.escape("<a href=\"x\">Tom & 'Jerry'</a>"))
print(html.escape("<a href=\"x\">Tom & 'Jerry'</a>", quote=False))
print(html.unescape("&lt;p&gt; &amp;amp; &copy; &#169; &#xA9; &euro; &#128;"))
&lt;a href=&quot;x&quot;&gt;Tom &amp; &#x27;Jerry&#x27;&lt;/a&gt;
&lt;a href="x"&gt;Tom &amp; 'Jerry'&lt;/a&gt;
<p> &amp; © © © € €

html.escape quotes both quote characters by default. Pass quote=False for text nodes only.

Escape and decode in PHP

echo htmlspecialchars("<a href=\"x\">Tom & 'Jerry'</a>"), "\n";
echo html_entity_decode("&copy; &euro; &apos;", ENT_QUOTES | ENT_HTML5, "UTF-8"), "\n";
echo htmlspecialchars("it's"), " ", htmlspecialchars("it's", ENT_QUOTES | ENT_HTML5), "\n";
&lt;a href=&quot;x&quot;&gt;Tom &amp; &#039;Jerry&#039;&lt;/a&gt;
© € '
it&#039;s it&apos;s

On PHP 8.3 the default flags escape both quote types and use the HTML 4.01 form &#039;. Add ENT_HTML5 to get &apos;.

Decode entities one level at a time

import html
print(html.unescape("&amp;lt;"))
print(html.escape(html.escape("a & b")))
&lt;
a &amp;amp; b

unescape decodes once. If you see &amp;lt; in stored data, the text was escaped twice, and the fix is to stop one of the escapes, not to decode in a loop.

Write a symbol from its code point

const cp = 0x20AC;
console.log(`&#${cp};`, `&#x${cp.toString(16).toUpperCase()};`, String.fromCodePoint(cp));
&#8364; &#x20AC; €

Use a numeric reference when the character has no named entity or when the file might not be saved as UTF-8. This works for any code point, including ones above U+FFFF.

Pitfalls

  • Missing semicolon: &copy still decodes to ©, but &notit; becomes ¬it;. Python's html.unescape("&notit;") returns ¬it; because &not is one of the legacy names. Always end references with a semicolon.
  • &apos; has no legacy form: it is not in the list of 106 semicolon-less names, so &apos without the semicolon stays as literal text. Use &#39; if you target old parsers.
  • Double escaping: escaping on input and again on output turns & into &amp;amp; and shows literal &amp; on the page. Escape once, at output, for the context you write into.
  • &nbsp; is not a space: it is U+00A0. It does not collapse, it is not matched by a plain space in a regex, and a plain split(" ") does not split on it, although JavaScript's trim() does remove it.
  • Escaping outside text and quoted attributes: HTML entities do not protect text placed in script, style, URLs or unquoted attributes. Use the escape that fits that context, such as percent-encoding for URLs.
  • Entities in the wrong place: a reference inside a script element is not decoded, and a reference inside textarea or title is decoded but never creates tags. Know which element you write into.
  • Entity names are case-sensitive: &Aacute; is Á and &aacute; is á. A few uppercase aliases such as &AMP; and &COPY; exist, but do not rely on them in new markup.
  • Relying on &#128; to mean control code 128: HTML remaps numbers 0x80 to 0x9F to Windows-1252 characters, so &#128; shows €. Use the real code point, &#8364;.
  • Two-character entities: 93 of the named references expand to two code points, for example &NotEqualTilde; is U+2242 followed by U+0338. A table that maps one entity to one character loses the second.

Related ZipKit tools

Related cheatsheets