law #18
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Cross-entropy

You will learn: Separate a source's uncertainty from the expected log-loss of a chosen prediction.

Start with: Shannon entropy

A prediction assigns probabilities before an outcome is revealed. How should we score it? Under log-loss, an observed label i incurs −log₂ qᵢ bits. The score penalizes confidently assigning very small probability to what actually occurs. Review Shannon entropy first; this lesson keeps that source distribution P and introduces a separate prediction Q.

H(P,Q)=EP[−log⁡2Q(X)]=−∑ipilog⁡2qi.H(P,Q)=E_P[-\log_2 Q(X)]=-\sum_i p_i\log_2 q_i.

P supplies the averaging weights; Q supplies the predicted probabilities. Both must refer to the same labels. This expected score is cross-entropy, not the fraction of labels classified correctly.

Set relative weights for three labeled outcomes. Each vector is normalized by its own total. They specify probabilities, not collected observations. An all-zero vector is invalid and is never replaced with a uniform distribution.

Source P and prediction Q probabilities00.20.40.60.81ABCoutcomeprobability

Filled blue bars: P, the distribution used for averaging. Outlined gold bars: Q, the probabilities used for prediction. Both use the same outcome labels.

Quantitybits per outcome
Entropy H(P)1.500000
Cross-entropy H(P,Q)1.584963
KL(P ∥ Q)0.08496250

Each outcome contributes to expected log-loss

Per-outcome contributions to the selected information quantity00.20.40.60.8ABCoutcomebits contribution
OutcomeP(outcome)Q(outcome)−p log q
A0.50000000.33333330.7924813
B0.25000000.33333330.3962406
C0.25000000.33333330.3962406

An outcome with source probability zero contributes zero, including when its prediction probability is also zero. For fixed source probabilities, matching the prediction makes cross-entropy equal entropy and KL equal zero. Cross-entropy is an expected loss, not classification accuracy.

Exact calculations for a finite three-outcome model. Base 2 gives bits and the natural logarithm gives nats; changing units does not change which prediction minimizes expected log-loss.

The same source, two predictions

Keep P=(1/2,1/4,1/4). Its entropy is 1.5 bits. The default uniform prediction Q=(1/3,1/3,1/3) assigns every possible outcome log₂ 3≈1.584963 bits of loss, so its expected loss is also 1.584963. Matching Q to P reduces the expected loss by 0.084963 bits per outcome.

This improvement does not mean every realized label gets a better score. Under the matched prediction, A costs one bit while B and C each cost two. A realized B scores worse than under the uniform prediction. The comparison averages those outcomes under P; a single draw cannot establish which probability model is better.

For a more concentrated prediction, set Q weights to 6,1,1. Then Q=(3/4,1/8,1/8), and expected loss is 0.5 log₂(4/3)+0.5·3≈1.707519 bits. Predicting A too confidently makes the rarer outcomes expensive. Both P and this Q would choose A as their most probable label, yet their log-loss differs.

Make a prediction

Does minimizing cross-entropy require making every prediction certain?

Explore the answer

No. For a fixed source P, the minimum is achieved by Q=P. The default source remains uncertain; its best expected log-loss is 1.5 bits, not zero. Forcing all probability onto A makes the score infinite because B and C still occur under P.

The part you can improve

Insert and subtract the source log-probability:

−∑ipilog⁡2qi=−∑ipilog⁡2pi+∑ipilog⁡2piqi=H(P)+DKL(P∥Q).-\sum_i p_i\log_2 q_i=-\sum_i p_i\log_2 p_i+\sum_i p_i\log_2\frac{p_i}{q_i}=H(P)+D_{\mathrm{KL}}(P\Vert Q).

For fixed P, its entropy is constant. Reducing expected cross-entropy is therefore equivalent to reducing KL divergence. If a restricted prediction family cannot represent P, its best attainable score may exceed H(P).

This identity does not justify comparing raw losses across datasets with different label frequencies or different conditioning information as if they had the same baseline difficulty. Nor does a low average loss alone prove subgroup calibration or acceptable decisions; those are separate questions.

A zero prediction is an explicit claim

Set Q’s C weight to zero while leaving P(C)>0. Cross-entropy becomes infinite. We display that result without replacing zero with a small number. If P(C)=0 too, C contributes zero instead: it is absent from this expectation. An all-zero weight vector is not a probability distribution and pauses calculation.

Smoothing can keep predictions positive, but it changes Q and its score. In a fitted model, smoothing or a prior should be a stated modeling choice. It should not be silently inserted just to make a graph finite.

From expectations to observed data

Suppose a sample has counts nᵢ and total N. Its average log-loss is −Σᵢ(nᵢ/N)log₂ qᵢ, the cross-entropy of the empirical frequencies with Q. For independent observations under a common categorical model, multiplying their probabilities gives the likelihood. Taking its negative logarithm and dividing by N gives the same observed loss.

This algebraic equivalence does not make a training score an unbiased assessment of a model selected using that training set. Held-out evaluation answers a different question. For predictions that depend on input features, each observation uses its own conditional Q, rather than one common three-category vector.

Make a prediction

If C was absent from a small dataset, is predicting Q(C)=0 always harmless?

Explore the answer

It gives no C contribution to that dataset’s observed average. It can still give infinite population expected loss if C has positive source probability, and it fails immediately if a future C occurs. Unseen and impossible are different claims.

Explore maximum likelihood to hold observations fixed and inspect the optimizing parameter.

Continue to KL divergence to reverse the averaging direction and inspect the excess-loss terms.

Sources

SciPy’s entropy reference states the cross-entropy decomposition. Shalizi’s notes on entropy and divergence develop expected log-density comparisons. The three-outcome predictions, their scores, and the count-to-likelihood calculation above can be reproduced directly.

Reset all settings