Finding something worth knowing…

Science

Smoke alarms chase recall while courtrooms chase precision

A smoke detector that shrieks at burnt toast is doing its job: missing a real fire would be far worse than a false alarm. A court that lets ten guilty people walk rather than jail one innocent is making the opposite bet. Two simple ratios, precision and recall, capture that trade-off.

Picture software hunting for dogs in a photo that contains twelve dogs and ten cats. It flags eight animals as dogs, but only five really are; the other three are cats. Its precision, the share of its picks that were right, is 5 out of 8. Its recall, the share of all real dogs it found, is 5 out of 12, because seven dogs slipped past. Precision measures quality of what you return; recall measures how much of what matters you caught.

Each is easy to game on its own. Return everything and recall is perfect, since nothing is missed, but precision collapses. Return only the one or two items you are surest of and precision can be flawless while recall is dismal. So the two usually pull against each other, and improving one tends to cost the other. In statistical terms, recall is one minus the rate of misses, known as type II errors, while precision ties to false alarms, or type I errors, though it also depends on how common the target is to begin with.

Which matters more depends on the price of each mistake. Smoke detectors are tuned for recall. Blackstone's famous ratio pushes the justice system toward precision. Fraud detection leans toward recall, since an undetected fraudulent payment can be very costly, while a diagnostic test whose false positives trigger needless treatment may favour precision.

Because a single number is handy, the two are often blended. The F-measure takes a weighted harmonic mean of them, and a precision-recall curve shows how one falls as the other rises. Critics note that both ignore the true negatives, the cats correctly left alone, and can be manipulated by biasing predictions; chance-corrected measures such as the Matthews correlation coefficient try to fix that. A classifier guessing blindly scores a precision equal to how common the target class is.

Source: Precision and recall

Related

More in Science · All topics