An HTML entity, formally a character reference, is a way to write a character in markup as &name; or &#number;. This sheet is for developers who escape user text, paste special symbols into templates, or decode & messes. The confusion it clears up: named, decimal and hex references all produce the same character, and only a handful of characters must be escaped.
| Character | Entity | Escape it in | Why |
|---|---|---|---|
| `&` | `&` | Text and attribute values | Starts every reference |
| `<` | `<` | Text | Starts a tag |
| `>` | `>` | Text, as a habit | Closes a tag, rarely dangerous alone |
| `"` | `"` | Double-quoted attribute values | Ends the value |
| `'` | `'` or `'` | Single-quoted attribute values | Ends the value |
| Form | Example | Result |
|---|---|---|
| Named | `©` | © |
| Decimal | `©` | © |
| Hexadecimal | `©` | © |
The hex form accepts x or X, and the digits are case-insensitive. A decimal or hex reference can name any code point up to U+10FFFF, subject to the rewrites in the last table below.
| Entity | Character | Decimal | Hex | Works without the semicolon |
|---|---|---|---|---|
| `&` | `&` | `&` | `&` | yes |
| `<` | `<` | `<` | `<` | yes |
| `>` | `>` | `>` | `>` | yes |
| `"` | `"` | `"` | `"` | yes |
| `'` | `'` | `'` | `'` | no |
| ` ` | no-break space | ` ` | ` ` | yes |
| `©` | © | `©` | `©` | yes |
| `®` | ® | `®` | `®` | yes |
| `™` | ™ | `™` | `™` | no |
| `€` | € | `€` | `€` | no |
| `£` | £ | `£` | `£` | yes |
| `¥` | ¥ | `¥` | `¥` | yes |
| `¢` | ¢ | `¢` | `¢` | yes |
| `§` | § | `§` | `§` | yes |
| `°` | ° | `°` | `°` | yes |
| `±` | ± | `±` | `±` | yes |
| `×` | × | `×` | `×` | yes |
| `÷` | ÷ | `÷` | `÷` | yes |
| `½` | ½ | `½` | `½` | yes |
| `…` | … | `…` | `…` | no |
| `—` | — | `—` | `—` | no |
| `–` | – | `–` | `–` | no |
| `‘` | ‘ | `‘` | `‘` | no |
| `’` | ’ | `’` | `’` | no |
| `“` | “ | `“` | `“` | no |
| `”` | ” | `”` | `”` | no |
| `«` | « | `«` | `«` | yes |
| `»` | » | `»` | `»` | yes |
| `←` | ← | `←` | `←` | no |
| `→` | → | `→` | `→` | no |
| `•` | • | `•` | `•` | no |
| `·` | · | `·` | `·` | yes |
The last column is read from the WHATWG entities list: only 106 legacy names are recognized without the closing semicolon.
| Reference | Result | Reason |
|---|---|---|
| `�` | U+FFFD replacement character | Null is not allowed |
| `€` | € (U+20AC) | Numbers 0x80 to 0x9F are remapped to Windows-1252 characters |
| `�` | U+FFFD | Surrogate code points are not characters |
| `�` | U+FFFD | Above the Unicode maximum |
const esc = s => s.replace(/&/g, "&").replace(/</g, "<").replace(/>/g, ">").replace(/"/g, """).replace(/'/g, "'");
console.log(esc(`<a href="x">Tom & 'Jerry'</a>`));
<a href="x">Tom & 'Jerry'</a>
Replace the ampersand first. If you do it last, you double-escape the entities you just inserted. In a browser, assigning to textContent avoids the problem because no HTML parsing happens.
import html
print(html.escape("<a href=\"x\">Tom & 'Jerry'</a>"))
print(html.escape("<a href=\"x\">Tom & 'Jerry'</a>", quote=False))
print(html.unescape("<p> &amp; © © © € €"))
<a href="x">Tom & 'Jerry'</a>
<a href="x">Tom & 'Jerry'</a>
<p> & © © © € €
html.escape quotes both quote characters by default. Pass quote=False for text nodes only.
echo htmlspecialchars("<a href=\"x\">Tom & 'Jerry'</a>"), "\n";
echo html_entity_decode("© € '", ENT_QUOTES | ENT_HTML5, "UTF-8"), "\n";
echo htmlspecialchars("it's"), " ", htmlspecialchars("it's", ENT_QUOTES | ENT_HTML5), "\n";
<a href="x">Tom & 'Jerry'</a>
© € '
it's it's
On PHP 8.3 the default flags escape both quote types and use the HTML 4.01 form '. Add ENT_HTML5 to get '.
import html
print(html.unescape("&lt;"))
print(html.escape(html.escape("a & b")))
<
a &amp; b
unescape decodes once. If you see &lt; in stored data, the text was escaped twice, and the fix is to stop one of the escapes, not to decode in a loop.
const cp = 0x20AC;
console.log(`&#${cp};`, `&#x${cp.toString(16).toUpperCase()};`, String.fromCodePoint(cp));
€ € €
Use a numeric reference when the character has no named entity or when the file might not be saved as UTF-8. This works for any code point, including ones above U+FFFF.
© still decodes to ©, but ¬it; becomes ¬it;. Python's html.unescape("¬it;") returns ¬it; because ¬ is one of the legacy names. Always end references with a semicolon.' has no legacy form: it is not in the list of 106 semicolon-less names, so &apos without the semicolon stays as literal text. Use ' if you target old parsers.& into &amp; and shows literal & on the page. Escape once, at output, for the context you write into. is not a space: it is U+00A0. It does not collapse, it is not matched by a plain space in a regex, and a plain split(" ") does not split on it, although JavaScript's trim() does remove it.script, style, URLs or unquoted attributes. Use the escape that fits that context, such as percent-encoding for URLs.script element is not decoded, and a reference inside textarea or title is decoded but never creates tags. Know which element you write into.Á is Á and á is á. A few uppercase aliases such as & and © exist, but do not rely on them in new markup.€ to mean control code 128: HTML remaps numbers 0x80 to 0x9F to Windows-1252 characters, so € shows €. Use the real code point, €.≂̸ is U+2242 followed by U+0338. A table that maps one entity to one character loses the second.%26 replaces &.&#x...;.