distribution #25
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Gaussian mixture

You will learn: Separate latent-group probabilities from the mean and spread of their mixture.

Start with: Normal distribution · Law of total variance

New to the notation? Start with the connected foundation for the underlying definitions and a worked example.

A Gaussian mixture is a weighted sum of Normal densities. Each component has its own mean and standard deviation; each sample is drawn from exactly one component, chosen at random according to the mixing weight.

f(x)=w N(x∣μ1,σ1)+(1−w) N(x∣μ2,σ2)f(x)=w\,\mathcal{N}(x\mid\mu_{1},\sigma_{1})+(1-w)\,\mathcal{N}(x\mid\mu_{2},\sigma_{2})
density by x 00.10.20.3−4−2024xdensity
mixture (solid) w · N(μ₁, σ₁) (long dashes) (1 − w) · N(μ₂, σ₂) (short dashes) E[X] = 0.60, Var = 3.63

Choose a latent component before drawing a value

Draw 1: uniform choice 0.62707 is at least w=0.4, so it selects component 2. The resulting normal draw is -0.36918. Component 1 supplies 593/1500 observations in this sample.

Each σ in this lesson denotes a standard deviation. The component label is latent when fitting ordinary unlabelled data; the simulator exposes it because it generated it. Two components can overlap into one peak, and identical component distributions give exactly one normal density.

Separate variation within and between components

Within-component variance: 0.694000. Between-component variance: 2.940000. Their sum is the mixture variance 3.634000. Simply averaging σ₁² and σ₂² misses the second term when component means differ.

Choosing a component differs from adding independent normal variables
ConstructionMeanVarianceP(value≤0)
Choose X₁ or X₂0.60003.63400.374560
Sum X₁+X₂0.50001.49000.341044
Average wX₁+(1−w)X₂0.60000.33640.150455

If an observed value were exactly x=0, the model's conditional component-1 probability, defined through its densities, would be 0.899754. This is different from the prior mixing weight. A point has zero probability under a continuous density; the conditional label probability is obtained from the density ratio.

0/1500 mixture draws are outside the density window. Their values remain in the sampling construction and denominators.

f(x)=w N(x∣μ1,σ1)+(1−w) N(x∣μ2,σ2)f(x)=w\,\mathcal{N}(x\mid\mu_{1},\sigma_{1})+(1-w)\,\mathcal{N}(x\mid\mu_{2},\sigma_{2}). Slide μ₁ and μ₂ apart relative to σ's until two modes appear — the hallmark of a bimodal mixture.

What to notice

  • Bimodality is not automatic. Even with μ1≠μ2\mu_{1}\ne\mu_{2}, the mixture may be unimodal. For equal weights and equal standard deviations, separation greater than two standard deviations produces two modes. Unequal weights or scales change that threshold.
  • Weight controls which mode dominates. At w=0.5w=0.5 the components contribute equally; push the slider and one peak drowns out the other.
  • Overall mean and variance decompose. The mean is the weighted average of component means. The variance picks up an extra term from the squared distance of each component from the overall mean — a special case of the law of total variance.

Why it matters

A mixture can describe heterogeneity, but fitted components require interpretation:

  • Clustering. A fitted mixture supplies model-based probabilities of component membership. It does not prove that its components correspond to real groups.
  • Density estimation. Additional components can make a density more flexible, with corresponding fitting and model-selection costs.
  • Heavy tails from simple ingredients. Two Normals with wildly different σ\sigma produce a leptokurtic “outlier-prone” mixture — a possible model for contamination. A finite normal mixture still has finite moments and does not become a power-law tail.
E[X]=wμ1+(1−w)μ2E[X]=w\mu_{1}+(1-w)\mu_{2}
Var(X)=w(σ12+(μ1−μ)2)+(1−w)(σ22+(μ2−μ)2)\mathrm{Var}(X)=w(\sigma_{1}^{2}+(\mu_{1}-\mu)^{2})+(1-w)(\sigma_{2}^{2}+(\mu_{2}-\mu)^{2})

Calculate within-group and between-group variation

Let C be a component label, with P(C=1)=w. Conditional on C, draw from its normal distribution. The law of total variance gives Var(X)=E[Var(X|C)]+Var(E[X|C]). For two components, these terms are wσ₁²+(1−w)σ₂² and w(1−w)(μ₁−μ₂)².

Take equal weights, means −2 and 2, and both SDs equal to 1. The mixture mean is zero; within-component variance is 1 and between-component variance is 4, totaling 5. Averaging the two component variances alone would miss most of the spread.

A mixture is different from a sum

If independent X₁ and X₂ have those two normal laws, their sum is normal with mean zero and variance 2. Their equally weighted average is normal with variance 0.5. Selecting one of them at random gives the mixture with variance 5. The same component descriptions have led to three different random experiments.

The demo’s table compares all three constructions using the same cutoff. It also displays the latent uniform choice for a selected mixture draw. The components’ dashed density curves integrate to w and 1−w; their pointwise sum integrates to one.

Update a component probability after observing a value

For an observed x, the conditional probability of component 1 is wf₁(x)/(wf₁(x)+(1−w)f₂(x)). With equal weights, means −1 and 1, and unit SDs, x=0 leaves the probability at 1/2. At x=1, it becomes 1/(1+e²)≈0.119203 because the observation is closer to component 2. The mixing weight is a prior probability; the density ratio updates it.

Make a prediction

If the two means and standard deviations are identical, does changing w change the marginal density?

Explore the answer

No. Both components have the same density, so their weighted sum is that density for every w. The label frequencies can change while the distribution of observed values stays the same. Unlabelled values alone cannot distinguish those identical components.

References

Stan’s finite-mixture guide describes latent component assignments, marginal mixture densities, and the distinction from vectorized component selection. The variance and sum calculations above follow from conditioning and independence. Continue to joint and conditional distributions or normal sums.

Reset all settings