Recurrent neural networks gave machines a memory by looping outputs back in
Read a sentence one word at a time and you carry the earlier words in your head. Recurrent neural networks do something similar: each step feeds part of its output back into the next, building a running memory. That loop powered voice search and machine translation until transformers took over.
An ordinary feedforward network treats every input on its own. A recurrent network instead keeps a hidden state that is updated at each step from the new input and the previous state, so the order of words, sounds or readings matters. That made RNNs a natural fit for handwriting recognition, speech, language modelling and translation. Diagrams that seem to show many layers are often one network unrolled across time.
The idea has two roots. In neuroscience, Santiago Ramón y Cajal described looping structures in the cerebellum in 1901, and by the 1940s researchers were arguing that the brain had feedback circuits, with Donald Hebb suggesting reverberating loops could hold short-term memory. In physics, the Ising model of magnets from the 1920s eventually fed into John Hopfield's 1982 network, which became a standard way to study neural networks through statistical mechanics. Frank Rosenblatt had already published perceptrons with recurrent connections in 1960.
Early RNNs struggled to learn links between distant parts of a sequence, because training signals faded as they travelled back through many steps, the vanishing gradient problem. In 1997 Sepp Hochreiter and Jürgen Schmidhuber introduced long short-term memory, or LSTM, which became the standard fix; gated recurrent units later offered a lighter alternative. The same year saw bidirectional networks that read sequences forwards and backwards at once.
By the mid-2000s bidirectional LSTMs were improving speech recognition and were used in Google voice search. In 2010 Tomáš Mikolov's team showed recurrent language models beating older n-gram methods, and in 2014 encoder–decoder RNNs set the pace in translation. Those systems helped inspire attention mechanisms and the transformer, which now dominates because it handles long-range context better and runs in parallel. RNNs still earn their keep where efficiency and real-time processing matter.
Source: Recurrent neural network