Squared spread that algebra loves, units hate
Variance measures how far numbers sit from their average: the expected squared gap from the mean. Take its square root and you get standard deviation—friendlier everyday units, fussier algebra. Many probability distributions lack a finite variance at all.
In probability language it is the second central moment and the covariance of a variable with itself, often written sigma squared. One payoff is clean arithmetic: for uncorrelated variables, the variance of a sum equals the sum of the variances—something absolute-deviation measures do not grant so easily. The cost is awkward units that are squares of the original scale, which is why finished reports usually quote the root instead. Many distributions simply lack a finite variance. Theory defines one variance from a distribution equation; data yield another—population variance if every observation is in hand, sample variance if only a subset is, used to estimate the whole.
Generate infinitely many draws from a theoretical law and the sample figure converges on the distribution's own variance. That centrality shows up across descriptive summaries, inference, hypothesis tests, goodness of fit, and Monte Carlo work. Computationally, an expanded identity says variance equals the mean of squares minus the square of the mean—but floating-point code should avoid that form when the two terms nearly cancel. For equally likely finite lists one may also write variance through all pairwise squared distances, averaging half the squared gaps across every pair.
Discrete laws weight each squared deviation by its probability mass; continuous densities integrate (x minus mu) squared against the density or, equivalently, integrate x squared and subtract mu squared via Lebesgue or Riemann forms. The exponential distribution with rate lambda, density lambda e to the minus lambda x on the positive half-line, has mean 1/lambda—a standard worked example for pushing the integral machinery through. Whether discrete, continuous, mixed, or neither, the same squared-deviation expectation remains the definition that ties the abstract parameter to the spread you see in data.
Source: Variance