Percentiles quietly set internet bills, speed limits and children's growth charts
Your internet provider may ignore your heaviest five per cent of usage when it bills you, and the speed limit on your road may have been set by watching how fast most drivers already go. Both rely on percentiles, a simple idea that statisticians still cannot agree how to calculate.
A percentile is the value below which a given share of the data falls. The 97th percentile, for example, is the score that 97 per cent of observations sit beneath. Percentiles divide data into 100 slices, and a few have their own names: the 25th is the first quartile, the 50th is the median, and the 75th is the third quartile. They are measured in the same units as the data, so a weight percentile is expressed in kilograms, not per cent. The related percentile rank works the other way round, starting with a score and reporting the percentage of scores below it.
For large populations following a bell curve, each step of standard deviation corresponds to a fixed percentile. The mean sits at the 50th; one standard deviation above is roughly the 84th; three above is nearly the 99.9th. That is why hardly anyone lies more than three standard deviations from average height.
The practical uses are everywhere. Providers of burstable bandwidth often bill at the 95th or 98th percentile of usage, discarding rare spikes so customers pay for typical demand. Traffic engineers use the 85th percentile of drivers' speeds to judge whether a limit is set sensibly. Doctors plot children's height and weight on growth charts against national percentiles, and bankers use percentile-based value at risk to estimate how far a portfolio might fall.
Surprisingly, there is no single standard formula. Rob Hyndman and Yanning Fan catalogued nine, and most software picks one of them. Some methods always return a value that actually appears in the data, while others interpolate between neighbouring values, as spreadsheet functions typically do. With large samples from a continuous distribution the methods converge, but on small datasets they can give noticeably different answers.
Source: Percentile