KL divergence measures the cost of believing the wrong model
Suppose you design a code for messages assuming one pattern of letter frequencies, but the real messages follow another. On average you will waste some bits per symbol. That wasted amount is the Kullback–Leibler divergence, a measure of how far an approximation strays from reality that turns up everywhere from neuroscience to machine learning.
Formally, it compares a true probability distribution, usually called P, with an approximating one, Q. Often P stands for observed data and Q for a theory or model, though sometimes the roles are flipped, with P as the model and Q as simulated data meant to match it. The divergence is the expected extra surprise from using Q when outcomes actually follow P, or equivalently the average number of additional bits needed to encode samples of P with a code built for Q.
It behaves like a distance in some ways but not others. It is never negative and equals zero only when the two distributions are identical. Yet it is not symmetric: the divergence of P from Q generally differs from that of Q from P. Nor does it obey the triangle inequality, so mathematicians do not count it as a true metric. In the language of information geometry it acts more like a squared distance, and for some important families of distributions it even satisfies a version of the Pythagorean theorem.
Solomon Kullback and Richard Leibler introduced it in 1951, describing it as the mean information for discriminating between two hypotheses. A symmetric version had already been used by Harold Jeffreys in 1948. Kullback himself preferred the phrase discrimination information and called the one-way quantities directed divergences. The one-way version now carries the names of both authors, while the symmetric form is known as the Jeffreys divergence.
Its uses range widely: measuring relative entropy in information theory, gauging information gained when comparing statistical models, and practical work in bioinformatics, fluid mechanics and neuroscience. In machine learning, algorithms such as expectation–maximization sometimes minimise the reversed direction instead, because it is easier to compute and shrinking one direction usually shrinks the other too.
Source: Kullback–Leibler divergence