How attention lets a translation model look back at the right word
Ask a neural network to translate I love you into French and watch where it looks. To produce je it puts 94% of its focus on I. For the next word it jumps ahead to you and writes t'. Only then does it return to love and choose aime. That shifting focus is called attention.
Attention is a way for a model to decide how much each piece of a sequence matters relative to all the others. In language, those judgements become soft weights attached to every word. Unlike the ordinary weights a network learns slowly during training, soft weights are recalculated on each forward pass, so they change with every new input.
It was invented to fix a weakness in recurrent neural networks, which read a sentence one word at a time and tend to favour the latest words while earlier ones fade. Attention gives each token direct access to any other part of the sentence instead of relying on a chain of previous states. Soft weighting matters because alignment is not always one to one: the English look it up maps onto the single French phrase cherchez-le, so the model blends several hidden vectors rather than picking one winner.
The big leap was self-attention, in which every element of the input attends to every other element and so captures relationships across the whole sequence. Introduced in its highly parallel form in 2017, it powered the Transformer, which dropped slow step-by-step recurrence entirely. Models such as BERT, T5 and GPT are built on it. The same idea now helps computer vision systems focus on the relevant regions of an image and aids speech recognition.
There is a cost. The attention matrix grows with the square of the number of tokens, so long inputs eat GPU memory. Flash attention eases this by splitting the calculation into blocks small enough for the chip's fast on-board memory, without losing accuracy. Attention maps drawn as heat maps are a popular way to peek inside vision transformers, although higher scores do not always mean a part of the input mattered more.
Source: Attention (machine learning)