Finding something worth knowing…

Science

A small p is not the chance your hypothesis is true

In null-hypothesis significance testing, a p-value is the probability of seeing a result at least as extreme as the one observed if the null hypothesis were true. Tiny p-values make that extreme outcome look rare under the null—yet the American Statistical Association warns they do not measure how likely the hypothesis itself is.

Researchers posit a null—often that a mean difference or correlation is zero—then compute a test statistic from the data. The p-value asks how surprising that statistic would be under the null’s sampling distribution, for one-sided or two-sided tails as chosen. All else equal, smaller p-values count as stronger evidence against the null; crossing a pre-set alpha level (commonly 0.05, as Ronald Fisher floated in 1925) licenses rejection. Alpha is fixed before peeking at the data; it is not estimated from the sample.

Misuse is widespread enough to occupy metascience. The ASA’s 2016 statement stressed that p-values do not give the probability the studied hypothesis is true, do not measure effect size or importance, and are poor standalone evidence for a model. A 2019 ASA task force nevertheless held that properly applied significance tests can increase the rigor of conclusions. Rejecting a null never names which alternative is best; larger samples make tiny deviations detectable, so scientific relevance still needs separate judgment.

Under a simple continuous null, p-values are uniformly distributed between zero and one; repeat the experiment and the number moves. Collections of published p-values form p-curves used to probe publication bias and p-hacking. Fisher’s combined probability test can pool independent p-values. For composite nulls the p-value is taken at the least favorable boundary case so that rejecting when p ≤ alpha still controls the Type I error rate the theorist set out to bound.

Source: P-value

Related

More in Science · All topics