Finding something worth knowing…

Science

The Mann-Whitney test is a tournament between two groups of numbers

Want to know whether one group tends to score higher than another without assuming a bell curve? Pair every value in the first group against every value in the second, award a point for each win and half a point for each tie, and add up the tally.

That tally is U, the heart of the test devised by Henry Mann and Donald Ransom Whitney and also known as the Wilcoxon rank-sum test. It is nonparametric: rather than assuming data follow a particular distribution, it asks only whether randomly chosen values from two populations share the same distribution. The requirements are modest. Observations must be independent, and it must be possible to say which of any two values is larger, so ordinal data such as rankings qualify.

For larger samples, counting pairwise contests gets tedious, so the usual recipe pools everything and ranks it from 1 upwards. Tied values share the midpoint of the ranks they would have occupied, so in the list 3, 5, 5, 5, 5, 8 each five gets rank 3.5. Summing the ranks of one sample is enough, because all the ranks together must total N(N + 1)/2.

U also converts neatly into an effect size. Dividing it by its largest possible value, the product of the two sample sizes, gives the probability that a random observation from one group beats a random observation from the other. That quantity is identical to the common language effect size, and it is the same idea used to score how well a classifier ranks one class above another, which lets the method stretch to problems with more than two classes.

It is often misread as a comparison of medians. That interpretation holds only under strict assumptions, such as continuous data where one distribution is simply shifted relative to the other. When shapes and spreads differ, the test can reject the null hypothesis with a tiny p-value even though the medians are equal. It should also not be confused with the Wilcoxon signed-rank test, which likewise sums ranks but is designed for matched pairs rather than independent samples.

Source: Mann–Whitney U test

Related

More in Science · All topics