Finding something worth knowing…

Science

Pearson's χ² asks if table frequencies match a null story

A chi-squared test compares what you counted in categories with what a null hypothesis predicted. For large contingency tables it checks whether two categorical factors act independently; tiny samples usually switch to Fisher's exact test instead of Pearson's asymptotic χ².

When observations fall into mutually exclusive classes, Pearson's chi-squared statistic sums (observed − expected)² / expected across cells. Under a true null and large samples, that sum behaves like a χ² random variable, so extreme values flag that the observed pattern would be unlikely if the null held. Independence tests for paired categorical variables and goodness-of-fit checks both live under this umbrella; many so-called χ² procedures are asymptotic—the sampling distribution only approaches χ² as n grows.

Nineteenth-century biologists often assumed normality; Karl Pearson, seeing skew in biological data, built the Pearson distribution family (1893–1916) and, in 1900, published the χ² goodness-of-fit paper now treated as a foundation of modern statistics. With known expected counts mi = npi he showed X² tends to χ² with k − 1 degrees of freedom. When expectations themselves are estimated from the sample he argued the same degrees of freedom still worked for practice—an approximation contested until R. A. Fisher's 1922 and 1924 papers settled the degrees-of-freedom debate.

Variants abound: tests of a normal population's variance (exact χ² with n − 1 df), Cochran–Mantel–Haenszel and McNemar procedures, portmanteau autocorrelation checks, and likelihood-ratio nested-model tests. Frank Yates's continuity correction subtracts 0.5 from each absolute residual in 2 × 2 tables, shrinking X² and raising p-values. Cryptanalysts even score plaintext-versus-ciphertext letter distributions with χ², favouring the decryption with the smallest statistic.

A textbook city of one million with four neighbourhoods illustrates the independence test: from a sample of 650, expected white-collar residents in neighbourhood A are about 80.54 if residence and job class are independent, contributing roughly 1.11 to the total X², which is referred to χ² with (3−1)(4−1) = 6 degrees of freedom. Fixed neighbourhood quotas turn the same arithmetic into a homogeneity test of equal job-class proportions.

Source: Chi-squared test

Related

More in Science · All topics