Glossary

ASCII

ASCII is a character encoding that maps 128 characters, such as letters, digits, punctuation and control codes, to the numbers 0 through 127. ASCII stands for American Standard Code for Information Interchange. RFC 20 (October 1969) adopted it for network use, based on the USA standard X3.4-1968. Today it is the foundation of UTF-8, HTTP headers, source code and most text protocols.

How it works

ASCII uses 7 bits per character, which gives 2 to the power of 7, or 128, codes. In practice each code is stored in one 8-bit byte whose top bit is 0, which is what RFC 20 suggests for network interchange.

  • 0 to 31: 32 control characters, such as NUL (0), HT or horizontal tab (9), LF (10), CR (13) and ESC (27). They control devices and are not drawn.
  • 127: DEL. It sits outside the 0–31 control block but prints nothing, so it is usually counted with the controls, which gives 33 non-printing codes.
  • 32: the space character.
  • 33 to 126: the visible characters, 94 of them. With the space, that makes 95 printable characters.
  • 48 to 57: the digits 0 to 9.
  • 65 to 90: uppercase A to Z.
  • 97 to 122: lowercase a to z.

Uppercase and lowercase letters differ by exactly 32, which is a single bit (the 0x20 bit). So A is 65 and a is 97, and you can flip case with a bit operation.

This run prints the decimal, binary and hex value of a few characters, the case offset, and what happens to a non-ASCII letter:

for c in "A","a","0"," ","\n","~":
    print(repr(c), ord(c), format(ord(c),"07b"), hex(ord(c)))
print(ord("A")^ord("a"), chr(ord("a")-32))
print("é".encode("ascii","replace"))
try: "é".encode("ascii")
except UnicodeEncodeError as e: print(e)
'A' 65 1000001 0x41
'a' 97 1100001 0x61
'0' 48 0110000 0x30
' ' 32 0100000 0x20
'\n' 10 0001010 0xa
'~' 126 1111110 0x7e
32 A
b'?'
'ascii' codec can't encode character '\xe9' in position 0: ordinal not in range(128)

How many characters are in ASCII?

ASCII has 128 characters, numbered 0 to 127. Of these, 95 are printable, and 33 are non-printing: the 32 control codes from 0 to 31 plus DEL at 127. Anything above 127 is not ASCII. Extended ASCII is an informal name for various 8-bit code pages, such as ISO 8859-1 and Windows-1252, that add characters from 128 to 255 and disagree with each other.

ASCII vs UTF-8

UTF-8 encodes the first 128 Unicode code points as the same single bytes as ASCII, so a pure ASCII file is already valid UTF-8. Any character beyond 127 takes two to four bytes in UTF-8, as defined in RFC 3629. ASCII cannot represent accented letters, other scripts or emoji, which is why modern text should be Unicode.

Common pitfalls

  • Encoding non-ASCII text as ASCII: Python raises UnicodeEncodeError: 'ascii' codec can't encode character. Encode as UTF-8 instead, or choose an explicit error policy only when losing data is acceptable.
  • Decoding UTF-8 bytes as Latin-1 or ASCII: you get garbled text such as "é" for "é". Declare the encoding at both ends.
  • Calling Windows-1252 "ASCII": bytes 128 to 255 are not ASCII. Smart quotes and the euro sign live in that range.
  • Forgetting the line-ending difference: text files use LF (10), CR (13) or both, depending on the system. Match the format your tool expects.
  • Treating control characters as safe: NUL (0) ends strings in C, and stray control codes in logs or filenames can break parsers. Validate or strip them at input.
  • Assuming sort order is alphabetical: in ASCII every uppercase letter sorts before every lowercase one, so "Zebra" comes before "apple" in a plain byte sort.

Related terms

  • UTF-8 — the encoding that keeps ASCII bytes unchanged and adds the rest of Unicode.
  • Unicode — the character set that contains ASCII as its first 128 code points.
  • CRLF — the CR and LF control characters used as a line ending.
  • Base64 — turns arbitrary bytes into a 64-character ASCII alphabet.
  • URL encoding — escapes bytes outside a safe ASCII subset using percent signs.

See also