Every letter you type is secretly a number, and agreeing on which number was hard
A computer cannot store the letter g, only numbers. Character encoding is the agreement about which number stands for which symbol. Getting everyone to share one scheme proved surprisingly hard, partly because in 1985 a typical personal computer's hard drive held only about 10 megabytes, and wasting even a single bit per character felt expensive.
The idea predates computers. Bacon's cipher, Braille and maritime signal flags all turned symbols into patterns, and in 1869 Hans Schjellerup devised a four-digit numeric code for sending Chinese characters by telegraph. Morse code, introduced in the 1840s, built letters from short and long signals and short and long gaps, producing codes of varying length. It is still heard in amateur radio and aviation.
Machines brought fixed-length codes. Émile Baudot created a five-bit telegraph code in 1870, later modified by Donald Murray and standardised in 1930. Herman Hollerith encoded census data as holes in punched cards in the late 19th century, and IBM's early electronic computers from 1953 used six-bit codes tied to those cards, enough for digits, capital letters and a few symbols. IBM's eight-bit EBCDIC followed in 1963 for the System/360, adding lower case. That same year the first version of ASCII appeared, a simpler seven-bit code that industry widely embraced; a 1967 revision added lower-case letters.
Those schemes served English well and most other languages poorly. By the 1980s researchers faced a dilemma. More characters needed more bits, yet for the majority of users, writing in the Latin alphabet, those extra bits would always sit at zero, squandering storage when a drive of roughly 10 megabytes cost about 250 dollars wholesale.
The escape, which grew into Unicode, abandoned the old telegraph assumption that each character must map straight to a fixed bit pattern. Instead every character gets an abstract number called a code point, and separate encodings decide how to store it. UTF-8 uses one to four bytes as needed, so plain English stays compact while any script remains possible. Unicode has since replaced most earlier encodings, and UTF-8 dominates the web.
Source: Character encoding