law #14
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Law of total variance

You will learn: Decompose a population variance without changing its outcomes when groups are relabeled.

Start with: Expectation and variance · Joint and marginal distributions

Why can two individually consistent groups produce a widely spread mixture? Total variance separates variation within groups from variation between their means. Changing the information used to define groups can move variance between the terms while leaving the population unchanged.

This lesson builds on expectation and variance and joint and marginal distributions. Assume Y has a finite second moment. If G identifies a group, then:

Var⁡(Y)=E[Var⁡(Y∣G)]+Var⁡(E[Y∣G]).\operatorname{Var}(Y)=E[\operatorname{Var}(Y\mid G)]+\operatorname{Var}(E[Y\mid G]).

The first term is an average of conditional variances, not their standard deviations. The second is the variance of conditional means, not the variance of group labels.

Within-group and between-group contributions to total variance0246one populationvariance (squared units)
Population quantityValue
Mean E[Y]3.000000
Filled lower segment: expected conditional variance3.250000
Outlined upper segment: variance of conditional means3.000000
Total variance (sum)6.250000
Direct sum of p(y − mean)²6.250000

Merging the groups leaves the population unchanged. All variance then lies within the single merged group. With separate labels, the filled lower term averages group variances using population weights; the outlined upper term measures spread of group means around the population mean. Neither is a causal effect.

Inspect the whole joint distribution

Choose A with probability 0.25, otherwise B. Within a chosen group, its two outcomes have equal probability. The group SD control places them at mean ± SD. Coincident values remain separate joint outcomes; zero-probability rows contribute nothing.

GroupYJoint probabilityContribution to variance
A-1.0000.12502.00000
A1.0000.12500.500000
B2.0000.37500.375000
B6.0000.37503.37500
Exact probabilities over four joint outcomes. No sampling noise or fitted causal model is involved.

Work through the four outcomes

The default model chooses group A with probability 1/4 and B with probability 3/4. Group A takes values −1 and 1 equally often: mean 0, variance 1. Group B takes values 2 and 6 equally often: mean 4, variance 4. These four group/value pairs specify the complete joint distribution.

The population mean is (1/4)·0+(3/4)·4=3. The within-group contribution is (1/4)·1+(3/4)·4=3.25. Group means differ from 3 by −3 and 1, so their contribution is (1/4)·9+(3/4)·1=3. Total variance is 6.25, giving population SD 2.5.

You can also calculate it directly from the four rows: probabilities 1/8,1/8,3/8,3/8 multiply squared distances 16,4,1,9 from the population mean. Their weighted sum is 6.25. This independent calculation is displayed alongside the decomposition.

Make a prediction

If both groups have SD zero but different means, must the whole population have zero variance?

Explore the answer

No. Within each group Y is constant, so the first term vanishes. With both groups having positive probability and different means, the between-group term is positive. Try SDs zero with means 0 and 4: at weight 1/4, total variance is 3.

Why the identity holds

Let m(G)=E[Y|G] and μ=E[Y]. Write each deviation as:

Y−μ=(Y−m(G))+(m(G)−μ).Y-\mu=(Y-m(G))+(m(G)-\mu).

Square and average. The residual square contributes E[Var(Y|G)], and the squared mean difference contributes Var(m(G)). The cross term vanishes because the residual has conditional mean zero:

E[(Y−m(G))(m(G)−μ)]=E[(m(G)−μ)E[Y−m(G)∣G]]=0.E[(Y-m(G))(m(G)-\mu)]=E[(m(G)-\mu)E[Y-m(G)\mid G]]=0.

The averaging step uses the tower property, E[E[Z|G]]=E[Z], for integrable Z. Finite E[Y²] supplies the moments needed here. Subtracting infinite quantities is not a valid way to extend this finite-variance calculation to a Cauchy population.

Change what you condition on

Select “No group information.” The four possible outcomes and their probabilities stay the same. The single merged group’s conditional variance is now 6.25, and its sole mean is the population mean. Between-group variance becomes zero.

More informative grouping cannot increase the average remaining conditional variance when the groupings are nested. But it can increase variance in a particular subgroup. Nor does a larger between-group term prove that the grouping causes the variation. It measures the predictive information in conditional means under the chosen joint distribution.

When both group means are equal, the between term is zero even if the group variances differ. At weight zero or one, only one group occurs; zero-probability rows make no contribution. The demo retains the specified mechanism for the unused group without pretending its conditional distribution was learned from observations.

Make a prediction

Is averaging two group variances equally correct when one group contains three quarters of the population?

Explore the answer

Generally no. Conditional variance must be averaged using the group’s probability. In the default example an equal average would give 2.5 instead of the correct within contribution 3.25. The target population determines the weights.

Application: heterogeneous counts

Suppose a count Y is conditionally Poisson with a random rate Λ for a fixed one-unit window. Both its conditional mean and variance equal Λ. Total expectation and total variance give:

E[Y]=E[Λ],Var⁡(Y)=E[Λ]+Var⁡(Λ).E[Y]=E[\Lambda],\qquad \operatorname{Var}(Y)=E[\Lambda]+\operatorname{Var}(\Lambda).

If half the windows have rate 1 and half rate 5, the mean count is 3, but its variance is 3+4=7. A single Poisson(3) model would have variance 3. This synthetic construction explains one possible source of overdispersion; an observed variance greater than the mean does not uniquely identify rate heterogeneity.

See Poisson processes for fixed-rate arrivals and conjugate priors for prediction that averages over posterior rate uncertainty. Gaussian mixtures provide a continuous example of within- and between-component spread.

Source

Random Services: conditional expected value develops conditional expectation, its averaging property, and conditional variance. The four-outcome model above gives a complete finite example whose identity can be checked by direct summation.

Return to learning paths and foundations.

Reset all settings