The value that shows up most—even for names, not numbers
The mode is the most frequent value in a sample, or the peak of a probability mass or density. Unlike mean and median, it works for categories: among Korean surnames, “Kim” can be the mode. In a normal bell it coincides with mean and median; in skewed wealth data it often does not.
For a discrete random variable the mode is any point that maximizes the probability mass function—the outcome most likely to be sampled. Continuous densities treat local maxima as modes, so multimodal laws have several peaks. In a sample the mode is the most common element: six dominates a list with four sixes among ones, threes, sevens, twelves, and a seventeen, while a short list with two ones and two fours is bimodal. Uniform discrete laws make every value a mode. Continuous samples almost never repeat exactly, so analysts bin into histograms or use kernel density estimates; histograms for modest samples are sensitive to bin width, with practice often aiming for a sizable share of data in roughly five to ten busy bins.
Mean, median, and mode agree in symmetric unimodal laws such as the normal. Karl Pearson introduced the term “mode” in 1895 for the abscissa of maximum frequency. Pearson’s rough rule for mildly skewed continuous unimodal shapes places the median about one third of the way from mean toward mode—median ≈ (2×mean + mode)/3—but the three statistics can appear in any order, and the rule is not always true. Van Zwet gave conditions for the ordering mode ≤ median ≤ mean. For unimodal laws the mode lies within √3 standard deviations of the mean; median and mean are within about 0.77 standard deviations of each other.
A log-normal example skews the story. If Y = e^X with X normal of mean zero, Y’s median is always one. With σ = 0.25, mean ≈ 1.032 and mode ≈ 0.939; with σ = 1, mean ≈ 1.649 and mode ≈ 0.368—strong skew that still roughly obeys Pearson’s thirding in the mild case. Practically, the mode (and median) shrug off outliers that yank the mean; the mean of a finite sample is always defined, while some pathological distributions lack a mode entirely. Affine transforms aX + b carry mean, median, and mode along when they are defined and unique.
Source: Mode (statistics)