Cheatsheet

ASCII Table Cheatsheet

This sheet is the lookup table for ASCII, the 128-code character set defined for network use in RFC 20. Use it when you need the number behind a character, the name of a control code, or a quick way to test whether a string is pure ASCII. The usual confusion it clears up: ASCII stops at 127, and "extended ASCII" is not part of the standard.

Quick reference

Code ranges

Decimal Hex Contents
0-31 00-1F Control characters
32 20 Space (SP)
33-47 21-2F `! " # $ % & ' ( ) * + , - . /`
48-57 30-39 Digits `0` to `9`
58-64 3A-40 `: ; < = > ? @`
65-90 41-5A Uppercase `A` to `Z`
91-96 5B-60 `[ \ ] ^ _` and the backtick
97-122 61-7A Lowercase `a` to `z`
123-126 7B-7E Left brace, vertical bar, right brace, tilde
127 7F DEL

Control codes you will actually meet

Dec Hex Name Escape in most languages Typical use
0 00 NUL `\0` String terminator in C
7 07 BEL `\a` Terminal bell
8 08 BS `\b` Backspace
9 09 HT `\t` Horizontal tab
10 0A LF `\n` Unix line ending
11 0B VT `\v` Vertical tab
12 0C FF `\f` Form feed
13 0D CR `\r` Part of the Windows line ending
27 1B ESC `\x1b` Starts terminal escape sequences
127 7F DEL `\x7f` Delete

RFC 20 names the rest: SOH 1, STX 2, ETX 3, EOT 4, ENQ 5, ACK 6, SO 14, SI 15, DLE 16, DC1 17 to DC4 20, NAK 21, SYN 22, ETB 23, CAN 24, EM 25, SUB 26, FS 28, GS 29, RS 30, US 31.

Bit tricks

Task Rule Example
Digit to number subtract 48 (0x30) `'7'` (55) gives 7
Lower to upper clear bit 0x20, or subtract 32 `a` (97) gives `A` (65)
Upper to lower set bit 0x20, or add 32 `A` (65) gives `a` (97)
Toggle case XOR with 0x20 `97 ^ 32` is 65
Control from letter AND with 0x1F `C` (67) gives 3 (ETX, Ctrl+C)

Write a code point in source code

Language Syntax for `A` (65) Notes
Python `"\x41"` or `chr(65)` `\x` takes exactly two hex digits
JavaScript `"\x41"` or `"\u0041"` `\u` takes four hex digits
C and PHP `"\x41"` or `"\101"` `\101` is octal
HTML `&#65;` or `&#x41;` Decimal and hex numeric references
URL `%41` Percent-encoding of the byte

Common patterns

Convert between a character and its code

print(ord("A"), chr(97), f"{ord('A'):08b} {ord('A'):#x} {ord('A'):#o}")
65 a 01000001 0x41 0o101

ord and chr work on any Unicode code point, so they match ASCII only for values 0 to 127.

Print every printable ASCII character

print("".join(chr(i) for i in range(32, 127)))
 !"#$%&'()*+,-./0123456789:;<=>?@ABCDEFGHIJKLMNOPQRSTUVWXYZ[\]^_`abcdefghijklmnopqrstuvwxyz{|}~

The output starts with a space (32) and ends at the tilde (126).

Test whether text is pure ASCII

print("Hello".isascii(), "café".isascii())
True False

str.isascii() is built into Python. In JavaScript, use a range test, shown next.

Check for ASCII in JavaScript

console.log(/^[\x00-\x7F]*$/.test("hello"), /^[\x00-\x7F]*$/.test("héllo"));
true false

The class covers codes 0 to 127. Use [\x20-\x7E] to allow printable characters only.

Flip letter case with bit operations

console.log("a".charCodeAt(0)^32, String.fromCharCode("a".charCodeAt(0)&~32));
65 A

This only works for the 52 ASCII letters. Applied to other characters it corrupts them.

Replace non-ASCII characters with a placeholder

console.log("a-b".replace(/[^\x20-\x7E]/g, "?"), "x€y".replace(/[^\x20-\x7E]/g, "?"));
a-b x?y

This replaces anything outside the printable range, tabs and newlines included. Widen the class if you need to keep them.

Strip accents to get an ASCII-friendly string

console.log("naïve café".normalize("NFD").replace(/[̀-ͯ]/g, ""));
naive cafe

Normalization splits each accented letter into a base letter plus a combining mark, and the regex removes the marks. It does nothing for letters like ß or ø.

Convert text to hex codes and back

print(' '.join(f'{b:02x}' for b in b'Hi!'), bytes.fromhex('48 69 21').decode('ascii'))
48 69 21 Hi!

Each ASCII character is one byte, so three characters give three hex pairs. decode('ascii') raises an error on any byte above 127, which makes it a strict validator.

Dump the raw bytes of a string

printf 'A\tB\r\n' | od -c
0000000   A  \t   B  \r  \n
0000005

Use od -An -tu1 -tx1 to print decimal and hex bytes instead. This is the quickest way to spot a stray CR in a file that looks fine in an editor, because CR and LF both render as a line break in many viewers.

Pitfalls

  • Treating 128 to 255 as ASCII: those values are not defined by ASCII. Windows-1252 and ISO 8859-1 both put characters there, and they differ. Call the data by its real encoding name.
  • Node's "ascii" encoding does not reject non-ASCII: Buffer.from("é", "ascii") prints <Buffer e9>, the same byte as "latin1". It silently truncates instead of failing, so validate with a regex first.
  • Browsers read the ascii label as Windows-1252: the WHATWG Encoding Standard maps the labels us-ascii, latin1 and iso-8859-1 to windows-1252. A page labeled ASCII can still show curly quotes.
  • Byte-order sorting: uppercase letters (65 to 90) sort before lowercase (97 to 122). sorted(["apple","Zebra","Banana"]) gives ['Banana', 'Zebra', 'apple'] in Python. Sort with key=str.lower for human order.
  • Punctuation is not contiguous: the symbols sit in four runs (33 to 47, 58 to 64, 91 to 96, 123 to 126). A range check from ! to ~ includes digits and letters too.
  • Counting printable characters: 95 characters are printable, from space (32) to tilde (126). Excluding the space leaves 94: 52 letters, 10 digits and 32 punctuation marks, which is the length of Python's string.punctuation.
  • DEL is a control code: 127 sits next to ~ but is not printable. RFC 20 itself says DEL is not strictly a control character, so libraries disagree on how to classify it.

Related ZipKit tools

Related cheatsheets