Predicting the next symbol well is the same skill as compressing data
A language model built for text, DeepMind's Chinchilla 70B, squeezed images to 43.4 percent and audio to 16.4 percent of their original size, beating the PNG and FLAC formats. That is less strange than it sounds. Any system that predicts what comes next can be turned into a compressor, and a perfect compressor doubles as a predictor.
Compression comes in two kinds. Lossless methods exploit repetition and can be reversed exactly: rather than storing red pixel, red pixel, red pixel, an image file can simply say 279 red pixels. The Lempel–Ziv family builds a table of strings already seen and swaps repeats for short references; Terry Welch's variant became the standard for general-purpose compression in the mid-1980s, turning up in GIF images, PKZIP and modems.
Lossy methods throw away detail people barely notice. Human eyes register small shifts in brightness more keenly than shifts in colour, and hearing has similar blind spots, so formats round off what goes unperceived. Most rely on the discrete cosine transform, proposed by Nasir Ahmed in 1972 and introduced in January 1974 after work with T. Natarajan and K. R. Rao. It sits behind JPEG photos, MPEG video and MP3 audio, while cameras, DVDs, Blu-ray and streaming all lean on lossy coding. Repeated lossy saving causes generation loss.
Claude Shannon laid the theoretical foundations in papers from the late 1940s and early 1950s. Arithmetic coding, which turns probability estimates into bits without forcing each symbol into a whole number of bits, often outperforms the older Huffman method and is used in modern video standards such as H.264 and HEVC. That is where prediction comes in: better probability estimates mean fewer bits.
The link runs deep enough that compression has been proposed as a benchmark for general intelligence. One theoretical view, tied to the Hutter Prize, defines the best compression of a file as the shortest program that regenerates it, which means a zip archive's true size should include the unzipping software. The Chinchilla results carry a caveat: its test data may have overlapped with material it was trained on.
Source: Data compression