Published
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
KL divergence
You will learn: Inspect directional excess log-loss, signed contributions, and zero-support failures.
Start with: Cross-entropy
How much additional expected log-loss do you incur by using Q when outcomes follow P? Cross-entropy minus source entropy gives Kullback–Leibler divergence:
The order matters: the expectation uses P. This lesson is about finite labeled distributions, with explicit support checks. It does not measure the physical separation of two random samples.
Set relative weights for three labeled outcomes. Each vector is normalized by its own total. They specify probabilities, not collected observations. An all-zero vector is invalid and is never replaced with a uniform distribution.
Filled blue bars: P, the distribution used for averaging. Outlined gold bars: Q, the probabilities used for prediction. Both use the same outcome labels.
| Quantity | bits per outcome |
|---|---|
| Entropy H(P) | 1.500000 |
| Cross-entropy H(P,Q) | 1.584963 |
| KL(P ∥ Q) | 0.08496250 |
| Reverse KL(Q ∥ P) | 0.08170417 |
Signed contributions to KL(P ∥ Q)
Individual KL contributions can be negative where the prediction puts more probability than the source. Their total is nonnegative. Reversing direction changes both the logarithmic ratio and the averaging weights.
| Outcome | P(outcome) | Q(outcome) | p log(p/q) |
|---|---|---|---|
| A | 0.5000000 | 0.3333333 | 0.2924813 |
| B | 0.2500000 | 0.3333333 | -0.1037594 |
| C | 0.2500000 | 0.3333333 | -0.1037594 |
An outcome with source probability zero contributes zero, including when its prediction probability is also zero. For fixed source probabilities, matching the prediction makes cross-entropy equal entropy and KL equal zero. Cross-entropy is an expected loss, not classification accuracy.
Negative pieces can have a positive sum
For the default P=(1/2,1/4,1/4) and uniform Q, the A term is 0.5 log₂(3/2)≈0.292481. Each of the other two terms is 0.25 log₂(3/4)≈−0.103759. Their sum is 0.084963 bits.
A negative term means Q assigned that outcome more probability than P did. Normalization prevents Q from doing so for every outcome with positive P while matching or exceeding P’s mass everywhere else. The result is a nonnegative total, not necessarily nonnegative bars.
Set Q=P: all ratios on positive-probability outcomes become one and the total is zero. The extra expected loss disappears, while the source entropy remains.
Reverse the question
Set P weights to 1,1,0 and Q weights to 9,1,0. In the forward direction the two occurring labels receive equal averaging weight:
In reverse, Q supplies weights 0.9 and 0.1, and the ratios reverse. The result is approximately 0.531004 bits. The dropdown changes both parts of that calculation. It does not merely negate the first result.
KL is not symmetric in general. Some pairs do have equal forward and reverse values, including identical distributions. For example, swapping the probabilities of a binary distribution can also give equal values. “Never symmetric” would be an incorrect rule.
Make a prediction
Can KL(P ∥ Q) be finite while KL(Q ∥ P) is infinite?
Explore the answer
Yes. Let P=(1,0,0) and Q=(1/2,1/2,0). The forward value is one bit. In reverse, Q’s B outcomes receive prediction probability zero under P, so the reverse value is infinite.
Why the total cannot be negative
If some pᵢ>0 has qᵢ=0, divergence is infinite. Otherwise apply Jensen’s inequality to the convex function −log₂ x on the positive-P support S:
The last sum cannot exceed one. Equality requires a constant qᵢ/pᵢ on S and no Q mass outside S, which together give Q=P. Zero-P terms are defined as zero, including when Q also gives zero. There is no 0/0 calculation to perform for an outcome omitted from the expectation.
Nonnegativity and equality at P=Q are not enough to make a distance metric. KL is generally asymmetric and also need not obey a triangle inequality. For example, with binary probabilities P=(0.1,0.9), Q=(0.5,0.5), R=(0.9,0.1), KL(P∥R)≈2.535940 exceeds KL(P∥Q)+KL(Q∥R)≈1.267970 bits.
Labels and information loss
Permuting both distributions’ labels together leaves KL unchanged. Merging categories can reduce it because the merged observation hides distinctions. For P=(1/2,1/4,1/4) and Q=(1/2,3/8,1/8), KL is positive; after merging B and C, both become (1/2,1/2), so the divergence becomes zero. Agreement on a coarse summary need not mean agreement on the original outcomes.
Changing from bits to nats multiplies all finite values by ln 2. It does not change direction, equality, or support failures. The finite controls calculate exact model expectations up to floating-point arithmetic; they do not estimate KL reliably from sparse observations merely by assigning zero to unseen labels.
Sources
Shalizi’s entropy and divergence notes give the expectation definition, support convention, and nonnegativity argument. SciPy’s entropy reference provides a numerical cross-check for the fair-versus-biased binary example. The signed terms, reverse calculation, triangle counterexample, and category-merging example above can be checked independently.