A random forest gets smarter by making each tree slightly ignorant
One decision tree trained on data tends to memorise noise and predict badly. Grow hundreds of them, show each a different random slice of the examples and hide some features from every split, then let them vote. The crowd of deliberately handicapped trees beats any single expert tree, and adding more rarely makes it worse.
Decision trees are popular because they cope with irrelevant inputs, ignore changes of scale and can be inspected by eye. The drawback, as the statisticians Hastie and colleagues put it, is that they are seldom accurate. A tree grown very deep learns quirky patterns specific to its training set: low bias but huge variance. Averaging many such trees cancels much of that variance, at the price of a little extra bias and some loss of interpretability.
Two tricks keep the trees from simply agreeing with each other. The first is bagging: each tree trains on a sample drawn with replacement from the data, so no two see quite the same examples. The second is feature bagging: at every candidate split, only a random subset of variables may be considered, so one dominant predictor cannot shape every tree identically. Classification answers come from a majority vote, regression answers from the average, and the spread of the trees' guesses gives a rough measure of uncertainty. Typical forests contain from a few hundred to several thousand trees.
The lineage is layered. Salzberg and Heath proposed combining randomised trees by majority vote in 1993. Tin Kam Ho built the first random decision forest algorithm in 1995 and showed something surprising: forests could keep gaining accuracy as they grew without overfitting, contradicting the usual belief that complexity eventually hurts. Leo Breiman and Adele Cutler extended the method and trademarked the name Random Forests in 2006, and Breiman's paper on it became one of the most cited in the world.
Breiman also supplied two practical tools. Out-of-bag error estimates accuracy using, for each example, only the trees that never saw it during training, so no separate test set is needed. Permutation importance scrambles one feature at a time and measures how much worse the forest performs, ranking which inputs matter most. Variants include ExtraTrees, which pick split points at random rather than optimising them.
Source: Random forest