Finding something worth knowing…

Technology

HAL's dying song in 2001 came from a real Bell Labs computer demo

In 1961 an IBM 7090 at Bell Labs sang Daisy Bell in a synthetic voice. Arthur C. Clarke happened to be visiting a friend there and heard it. He was so struck that he wrote the moment into 2001: A Space Odyssey, where the computer HAL 9000 sings that same tune while it is being shut down.

People had chased talking machines for centuries. Medieval legends credited scholars such as Roger Bacon with brass heads that could speak. In 1779 Christian Kratzenstein won a prize from the Russian academy for model vocal tracts that produced five long vowels, and Wolfgang von Kempelen's bellows-driven machine, described in 1791, added a tongue and lips so it could manage consonants too. At the 1939 New York World's Fair, Homer Dudley of Bell Labs showed the Voder, a voice synthesiser played from a keyboard.

Computers took over from the late 1950s. The Bell Labs singing demo was the work of physicist John Larry Kelly and Louis Gerstman, with musical backing arranged by Max Mathews, and in 1968 Noriko Umeda's team in Japan produced the first general text-to-speech system for English. Speech chips then shrank into consumer gadgets: a talking calculator for blind users in 1976, the Speak and Spell toy in 1978, and talking arcade games by 1980. Unix had already shipped a speak utility in 1974. Synthetic voices were nearly always male until Ann Syrdal of AT&T built a female one in 1990.

A modern text-to-speech engine works in two stages. A front end first spells out numbers and abbreviations as words, works out how each word is pronounced, and marks phrases and sentences so the rhythm sounds natural. A back end then turns that description into sound, either by stitching together fragments of recorded speech or by modelling the vocal tract to build a wholly artificial voice.

Deep learning transformed the field after DeepMind's WaveNet in 2016, which generated raw audio waveforms directly. Early neural systems needed tens of hours of recordings to copy a voice. By 2024 OpenAI confirmed that about 15 seconds could be enough, and judged its own voice-cloning tool too risky to release publicly.

Source: Speech synthesis

Related

More in Technology · All topics