Finding something worth knowing…

Science

Sample means turn normal—even when raw data do not

The central limit theorem says that a properly scaled sample average drifts toward a bell curve as the sample grows, even if each observation comes from a skewed or jagged law. That single fact lets normal-based tools travel far beyond truly Gaussian data.

In probability theory the CLT asserts that, under suitable conditions, a normalized sample mean converges in distribution to a standard normal. Several precise versions exist for different assumptions, but the shared message is huge: methods built for normals often remain usable for averages drawn from other populations. Ancestral forms reach back to 1811; the modern statements crystallized in the 1920s. The de Moivre–Laplace theorem—that normals approximate binomials—is the earliest recognizable version.

In statistical language, draw i.i.d. observations with mean μ and form their arithmetic mean X̄_n. The law of large numbers already says X̄_n settles on μ. The classical CLT describes the jitter around that settling: √n (X̄_n − μ) approaches a normal with mean 0 and variance equal to the common variance σ². Equivalently, repeating a large-sample average many times yields a histogram of those averages that looks increasingly Gaussian. The shape of each single X_i can be almost anything; the average's fluctuations still normalize.

Independence and identical distribution are the textbook hypotheses, yet weaker theorems allow non-identical laws or limited dependence if moment growth is controlled. Lyapunov's condition is a convenient sufficient check; satisfying it implies Lindeberg's 1920 condition, though the converse fails. Under Lindeberg, standardized sums of independent (not necessarily identical) centered variables still converge to a standard normal. Uniform convergence of the cdf to the limiting Φ is part of the classical story.

Practically, the theorem is why poll margins, measurement averages and many estimators come wrapped in normal confidence talk once n is "large enough"—with the honest caveat that how large depends on the parent distribution's tails and on which CLT variant you invoke.

Source: Central limit theorem

Related

More in Science · All topics