ANOVA grew out of crop plots, and a dog show explains it best
Analysis of variance asks one clever question: is the spread between groups much bigger than the spread inside them? If so, the groups probably really differ. Ronald Fisher built the method around a century ago while studying crop yields, and it became one of the workhorses of experimental science.
The core idea rests on the law of total variance: all the scatter in a data set can be split into parts with different sources. ANOVA divides it into variation between group averages and variation within each group, then compares the two with an F-test. In its simplest form it checks whether several population means are equal, extending the familiar t-test beyond two groups.
A dog show makes the logic concrete. Suppose you try to predict weight by sorting dogs into young or old and short- or long-haired. Each group still contains big and small animals, the averages look alike, and the grouping explains almost nothing. Splitting by working breed versus pet, and by athleticism, does better, though the groups still overlap. Sorting by breed works best of all, since every Chihuahua is light and every St Bernard heavy. ANOVA supplies formal tools for these intuitions, and unlike correlation it does not need all the data to be numeric.
Its roots run deep. Laplace was testing hypotheses in the 1770s, and around 1800 he and Gauss developed least squares for combining astronomical and survey measurements. By 1827 Laplace was applying such methods to atmospheric tides. Astronomers also studied the personal equation, errors caused by observers' reaction times, and those experimental methods later passed into psychology.
Fisher introduced the word variance in a 1918 paper on genetics, first applied the analysis to real data in 1921 in a study of crop variation, and in 1923, with Winifred Mackenzie, examined yields from plots given different varieties and fertilisers. His 1925 book Statistical Methods for Research Workers made the method famous. Jerzy Neyman published the first randomisation model, in Polish, in 1923. Statisticians still disagree about how exactly to define fixed and random effects.
Source: Analysis of variance