ASCII is a character encoding that maps 128 characters, such as letters, digits, punctuation and control codes, to the numbers 0 through 127. ASCII stands for American Standard Code for Information Interchange. RFC 20 (October 1969) adopted it for network use, based on the USA standard X3.4-1968. Today it is the foundation of UTF-8, HTTP headers, source code and most text protocols.
ASCII uses 7 bits per character, which gives 2 to the power of 7, or 128, codes. In practice each code is stored in one 8-bit byte whose top bit is 0, which is what RFC 20 suggests for network interchange.
Uppercase and lowercase letters differ by exactly 32, which is a single bit (the 0x20 bit). So A is 65 and a is 97, and you can flip case with a bit operation.
This run prints the decimal, binary and hex value of a few characters, the case offset, and what happens to a non-ASCII letter:
for c in "A","a","0"," ","\n","~":
print(repr(c), ord(c), format(ord(c),"07b"), hex(ord(c)))
print(ord("A")^ord("a"), chr(ord("a")-32))
print("é".encode("ascii","replace"))
try: "é".encode("ascii")
except UnicodeEncodeError as e: print(e)
'A' 65 1000001 0x41
'a' 97 1100001 0x61
'0' 48 0110000 0x30
' ' 32 0100000 0x20
'\n' 10 0001010 0xa
'~' 126 1111110 0x7e
32 A
b'?'
'ascii' codec can't encode character '\xe9' in position 0: ordinal not in range(128)
ASCII has 128 characters, numbered 0 to 127. Of these, 95 are printable, and 33 are non-printing: the 32 control codes from 0 to 31 plus DEL at 127. Anything above 127 is not ASCII. Extended ASCII is an informal name for various 8-bit code pages, such as ISO 8859-1 and Windows-1252, that add characters from 128 to 255 and disagree with each other.
UTF-8 encodes the first 128 Unicode code points as the same single bytes as ASCII, so a pure ASCII file is already valid UTF-8. Any character beyond 127 takes two to four bytes in UTF-8, as defined in RFC 3629. ASCII cannot represent accented letters, other scripts or emoji, which is why modern text should be Unicode.
UnicodeEncodeError: 'ascii' codec can't encode character. Encode as UTF-8 instead, or choose an explicit error policy only when losing data is acceptable.