Published
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
Multivariate normal distribution
You will learn: Connect covariance with conditional normals, projections, and singular joint support.
Start with: Normal distribution · Joint and marginal distributions
Two measurements can each look normal while their relationship carries additional information. A multivariate normal model specifies that relationship through a mean vector and covariance matrix. Every linear combination of its coordinates is normal, allowing a constant as a degenerate normal. Normal marginal shapes alone do not guarantee this joint model.
This experiment uses two variables with zero means, positive standard deviations σX and σY, and correlation ρ between −1 and 1. Review the normal distribution and joint and marginal distributions first.
The diagonal entries are variances; the off-diagonal entries are covariances. They have squared or cross-product units. Correlation removes the scale factors and is dimensionless.
Blue: all 300 simulated pairs, with zero population means. Gold: a 95% population ellipse, containing 286/300 simulated pairs. Red: conditioning value X=1 and conditional mean. Axes rescale to retain all points; visual angle also depends on axis scales.
The ellipse is a joint population probability region. It is not a confidence region for estimated means, nor two separate marginal 95% intervals.
Conditioning changes the distribution
Solid blue is the marginal CDF of Y; dashed red is the conditional CDF given X=1. This is a model calculation, not a subset of observations that happen to equal that exact continuous value. CDFs retain their full probability scale outside the displayed window.
| Population quantity | Value |
|---|---|
| Cov(X,Y) | 1.200000 |
| E[Y | X=1] | 1.200000 |
| Var(Y | X=1) | 2.560000 |
| Marginal Var(Y) | 4.000000 |
A linear combination is still normal
Define U=X+wY. Its population mean is zero and variance is 2.600000. The probability below uses the full normal CDF.
| Event | Population probability | Simulated fraction |
|---|---|---|
| U ≤ 1 | 0.7324283 | 228/300 = 0.760000 |
Inspect the first ten simulated pairs
| Pair | X | Y | X+wY |
|---|---|---|---|
| 1 | 1.98307 | 1.85632 | 0.12675 |
| 2 | 0.33422 | 2.17357 | -1.83935 |
| 3 | 0.61727 | 1.08116 | -0.46389 |
| 4 | 2.20461 | 1.71330 | 0.49130 |
| 5 | 0.72745 | -0.25768 | 0.98513 |
| 6 | 0.03928 | 0.16208 | -0.12280 |
| 7 | 0.37609 | -0.21475 | 0.59083 |
| 8 | -1.40549 | -1.63612 | 0.23063 |
| 9 | 0.40274 | 1.04806 | -0.64532 |
| 10 | 0.16970 | -2.22788 | 2.39757 |
Build dependence from independent normals
Let Z₁ and Z₂ be independent standard normals. The experiment constructs:
The shared Z₁ term creates covariance ρσXσY. The two independent contributions to Y’s variance add to σY²[ρ²+(1−ρ²)]=σY². Changing ρ preserves both marginal distributions while changing their relationship. For nonzero means, add μX and μY after this construction.
At the defaults σX=1, σY=2, and ρ=0.6, the covariance matrix has entries 1, 1.2, 1.2, 4. Given X=1, the shared component contributes 2·0.6·1=1.2 to Y, while its independent residual has variance 4·(1−0.36)=2.56. Thus the conditional SD is 1.6, smaller than the marginal SD 2.
A conditional distribution is a model calculation
The same construction gives:
Here N’s second argument is variance. Conditioning on another x changes the mean but not this conditional variance. For |ρ|=1 the residual vanishes and the formula describes a point mass. The demo shows its right-continuous CDF rather than trying to draw a zero-width density.
For continuous X, a particular value has probability zero. The conditional normal is supplied by the specified joint model, not by counting simulated rows with exact equality. Estimating a conditional distribution from real data requires additional choices and evidence.
Make a prediction
At ρ=−1 with σX=1 and σY=2, what does observing X=1 imply?
Explore the answer
Y must equal −2 under this model. The conditional variance is zero, even though the marginal variance of Y is four. All joint observations lie on the line Y=−2X.
Joint probability is not two marginal intervals
With |ρ|<1, Z₁²+Z₂² has a chi-squared distribution with two degrees of freedom. Its 95th percentile is −2 log(0.05). Transforming that disk with the construction above gives the gold 95% population ellipse. Its axes are determined by covariance; its visible angle also depends on the chart scales.
This is not a confidence region for an estimated mean. At |ρ|=1 the joint distribution has only one random direction, so the demo uses the standard-normal 95% segment instead. A two-dimensional density and its ellipse formula no longer apply. The sampled fraction inside the region fluctuates; the population probability remains 0.95.
Project the pair onto one number
For U=X+wY, expand its square and take expectations:
At the defaults with w=−1, Var(X−Y)=1+4−2·1.2=2.6. Its mean is zero, so P(X−Y≤1)=Φ(1/√2.6), about 0.732. The table compares that full-distribution probability with the seeded sample fraction. At ρ=1, equal SDs, and w=−1, U is identically zero, so its inclusive CDF at zero is one.
Make a prediction
Do normal marginals and zero correlation always imply independence?
Explore the answer
No. Let Z be standard normal and S an independent fair sign, then set X=Z and Y=SZ. Each marginal is standard normal and E[XY]=E[S]E[Z²]=0, but |X|=|Y| always. Their joint distribution lies on two crossing lines and is not jointly normal. Independence at zero correlation relies on the joint-normal assumption.
Regression to the mean applies the conditional-mean idea to repeated measurements. Total variance explains how marginal spread combines conditional spread and changing conditional means.
Sources
The Book of Statistical Proofs: conditional multivariate normals derives the general conditional mean and covariance. Random Services: normal samples develops normal-vector transformations and sums of squares. The explicit two-normal construction above also lets you verify the displayed covariance, conditional moments, and singular cases directly.
Return to learning paths and foundations.