Published
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
Cross-entropy
You will learn: Separate a source's uncertainty from the expected log-loss of a chosen prediction.
Start with: Shannon entropy
A prediction assigns probabilities before an outcome is revealed. How should we score it? Under log-loss, an observed label i incurs −log₂ qᵢ bits. The score penalizes confidently assigning very small probability to what actually occurs. Review Shannon entropy first; this lesson keeps that source distribution P and introduces a separate prediction Q.
P supplies the averaging weights; Q supplies the predicted probabilities. Both must refer to the same labels. This expected score is cross-entropy, not the fraction of labels classified correctly.
Set relative weights for three labeled outcomes. Each vector is normalized by its own total. They specify probabilities, not collected observations. An all-zero vector is invalid and is never replaced with a uniform distribution.
Filled blue bars: P, the distribution used for averaging. Outlined gold bars: Q, the probabilities used for prediction. Both use the same outcome labels.
| Quantity | bits per outcome |
|---|---|
| Entropy H(P) | 1.500000 |
| Cross-entropy H(P,Q) | 1.584963 |
| KL(P ∥ Q) | 0.08496250 |
Each outcome contributes to expected log-loss
| Outcome | P(outcome) | Q(outcome) | −p log q |
|---|---|---|---|
| A | 0.5000000 | 0.3333333 | 0.7924813 |
| B | 0.2500000 | 0.3333333 | 0.3962406 |
| C | 0.2500000 | 0.3333333 | 0.3962406 |
An outcome with source probability zero contributes zero, including when its prediction probability is also zero. For fixed source probabilities, matching the prediction makes cross-entropy equal entropy and KL equal zero. Cross-entropy is an expected loss, not classification accuracy.
The same source, two predictions
Keep P=(1/2,1/4,1/4). Its entropy is 1.5 bits. The default uniform prediction Q=(1/3,1/3,1/3) assigns every possible outcome log₂ 3≈1.584963 bits of loss, so its expected loss is also 1.584963. Matching Q to P reduces the expected loss by 0.084963 bits per outcome.
This improvement does not mean every realized label gets a better score. Under the matched prediction, A costs one bit while B and C each cost two. A realized B scores worse than under the uniform prediction. The comparison averages those outcomes under P; a single draw cannot establish which probability model is better.
For a more concentrated prediction, set Q weights to 6,1,1. Then Q=(3/4,1/8,1/8), and expected loss is 0.5 log₂(4/3)+0.5·3≈1.707519 bits. Predicting A too confidently makes the rarer outcomes expensive. Both P and this Q would choose A as their most probable label, yet their log-loss differs.
Make a prediction
Does minimizing cross-entropy require making every prediction certain?
Explore the answer
No. For a fixed source P, the minimum is achieved by Q=P. The default source remains uncertain; its best expected log-loss is 1.5 bits, not zero. Forcing all probability onto A makes the score infinite because B and C still occur under P.
The part you can improve
Insert and subtract the source log-probability:
For fixed P, its entropy is constant. Reducing expected cross-entropy is therefore equivalent to reducing KL divergence. If a restricted prediction family cannot represent P, its best attainable score may exceed H(P).
This identity does not justify comparing raw losses across datasets with different label frequencies or different conditioning information as if they had the same baseline difficulty. Nor does a low average loss alone prove subgroup calibration or acceptable decisions; those are separate questions.
A zero prediction is an explicit claim
Set Q’s C weight to zero while leaving P(C)>0. Cross-entropy becomes infinite. We display that result without replacing zero with a small number. If P(C)=0 too, C contributes zero instead: it is absent from this expectation. An all-zero weight vector is not a probability distribution and pauses calculation.
Smoothing can keep predictions positive, but it changes Q and its score. In a fitted model, smoothing or a prior should be a stated modeling choice. It should not be silently inserted just to make a graph finite.
From expectations to observed data
Suppose a sample has counts nᵢ and total N. Its average log-loss is −Σᵢ(nᵢ/N)log₂ qᵢ, the cross-entropy of the empirical frequencies with Q. For independent observations under a common categorical model, multiplying their probabilities gives the likelihood. Taking its negative logarithm and dividing by N gives the same observed loss.
This algebraic equivalence does not make a training score an unbiased assessment of a model selected using that training set. Held-out evaluation answers a different question. For predictions that depend on input features, each observation uses its own conditional Q, rather than one common three-category vector.
Make a prediction
If C was absent from a small dataset, is predicting Q(C)=0 always harmless?
Explore the answer
It gives no C contribution to that dataset’s observed average. It can still give infinite population expected loss if C has positive source probability, and it fails immediately if a future C occurs. Unseen and impossible are different claims.
Explore maximum likelihood to hold observations fixed and inspect the optimizing parameter.
Continue to KL divergence to reverse the averaging direction and inspect the excess-loss terms.
Sources
SciPy’s entropy reference states the cross-entropy decomposition. Shalizi’s notes on entropy and divergence develop expected log-density comparisons. The three-outcome predictions, their scores, and the count-to-likelihood calculation above can be reproduced directly.