How statisticians charge a model for every extra knob it adds
Give a statistical model more adjustable parameters and it will almost always hug your data more tightly, even when it is just memorising noise. The Japanese statistician Hirotugu Akaike devised a scoring rule that rewards a good fit but fines every added parameter, so the simplest adequate explanation tends to win.
His criterion, known as AIC, starts from information theory. No model captures the process that really produced the data, so every model throws some information away. You cannot measure that loss directly, because the true process is unknown. What Akaike showed, in work from 1974, is that you can estimate how much more information one candidate loses than another. The score is simple: twice the number of estimated parameters, minus twice the logarithm of the model's best achievable likelihood. Lower is better.
The scores become intuitive when turned into odds. Say three models score 100, 102 and 110. Raising e to half the gap between the best score and each rival gives its relative likelihood: the second is about 0.368 times as likely as the first to be the least lossy, the third a mere 0.007. You would drop the third, then either gather more data, admit the evidence cannot separate the top two, or blend them in a weighted average.
The method has clear limits. It only ranks models against one another and says nothing about absolute quality, so if every candidate is poor it gives no warning, and checking the chosen model afterwards is wise. With small samples it tends to reward too many parameters, which prompted a corrected version called AICc. Unlike the classic likelihood-ratio test, though, it can compare models that are not nested versions of each other.
Its reach turned out to be philosophical as well as practical. Any hypothesis test can be recast as a comparison of models, so any test can be run through AIC, and estimation fits the framework too. That means it can underpin statistical inference without significance levels or Bayesian priors, a third foundation beside the frequentist and Bayesian schools.
Source: Akaike information criterion