Finding something worth knowing…

Health

Why 20 kelvin is twice 10 kelvin, but Celsius can't make that claim

The coefficient of variation divides a data set's standard deviation by its mean, turning spread into a pure, unitless ratio. That makes it ideal for comparing wildly different measurements, from lab assays to investment returns, but it only works when zero genuinely means nothing, which is why temperature in Celsius breaks it.

A standard deviation on its own says little until you know the average it surrounds: a wobble of 10 matters far more around 20 than around 10,000. Dividing by the mean fixes that, and because the units cancel, the result is dimensionless and often written as a percentage, sometimes labelled relative standard deviation. Chemists use it to report how repeatable an assay is; engineers apply it in quality checks, and economists, epidemiologists and neuroscientists rely on it too.

Simple examples show the scale. Three readings of 100 have no spread, so the ratio is 0. Readings of 90, 100 and 110, treated as a sample, give a standard deviation of 10 around a mean of 100, a coefficient of 0.1. A scattered set running from 1 up to 88 has a mean of 27.9 and a standard deviation of 32.9, pushing the ratio to 1.18. Treat those numbers as a whole population rather than a sample and the figures shrink slightly, to 0.0816 and 1.10.

The measure has traps. It only makes sense on ratio scales with a true zero. Celsius and Fahrenheit set zero arbitrarily, so the same temperatures give different coefficients depending on the scale, whereas kelvin starts at the total absence of thermal energy. As the mean approaches zero, the ratio races towards infinity and becomes jumpy. It also cannot directly produce a confidence interval for the mean, and the straightforward sample estimate runs low for small samples, which is why statisticians apply a correction factor for normally distributed data. For roughly log-normal data, a separate formula known as the geometric CV is often preferred.

In queueing and reliability theory the benchmark is the exponential distribution, whose standard deviation equals its mean, giving a coefficient of exactly 1. Distributions below that, such as the Erlang, count as low-variance; those above it, such as the hyper-exponential, count as high-variance.

Source: Coefficient of variation

Related

More in Health · All topics