distribution #24
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Multivariate normal distribution

You will learn: Connect covariance with conditional normals, projections, and singular joint support.

Start with: Normal distribution · Joint and marginal distributions

Two measurements can each look normal while their relationship carries additional information. A multivariate normal model specifies that relationship through a mean vector and covariance matrix. Every linear combination of its coordinates is normal, allowing a constant as a degenerate normal. Normal marginal shapes alone do not guarantee this joint model.

This experiment uses two variables with zero means, positive standard deviations σX and σY, and correlation ρ between −1 and 1. Review the normal distribution and joint and marginal distributions first.

Σ=(σX2ρσXσYρσXσYσY2).\Sigma=\begin{pmatrix}\sigma_X^2 & \rho\sigma_X\sigma_Y\\ \rho\sigma_X\sigma_Y & \sigma_Y^2\end{pmatrix}.

The diagonal entries are variances; the off-diagonal entries are covariances. They have squared or cross-product units. Correlation removes the scale factors and is dimensionless.

Seeded bivariate normal observations and a 95 percent population probability regionBlue points are simulated observations. The gold region is computed from the specified population covariance, not estimated from these points.−6−4−20246−3−2−10123XY

Blue: all 300 simulated pairs, with zero population means. Gold: a 95% population ellipse, containing 286/300 simulated pairs. Red: conditioning value X=1 and conditional mean. Axes rescale to retain all points; visual angle also depends on axis scales.

The ellipse is a joint population probability region. It is not a confidence region for estimated means, nor two separate marginal 95% intervals.

Conditioning changes the distribution

Marginal and conditional normal cumulative distributions00.20.40.60.81−8−6−4−202468Y valuecumulative probability

Solid blue is the marginal CDF of Y; dashed red is the conditional CDF given X=1. This is a model calculation, not a subset of observations that happen to equal that exact continuous value. CDFs retain their full probability scale outside the displayed window.

Population quantityValue
Cov(X,Y)1.200000
E[Y | X=1]1.200000
Var(Y | X=1)2.560000
Marginal Var(Y)4.000000

A linear combination is still normal

Define U=X+wY. Its population mean is zero and variance is 2.600000. The probability below uses the full normal CDF.

EventPopulation probabilitySimulated fraction
U ≤ 10.7324283228/300 = 0.760000
Inspect the first ten simulated pairs
PairXYX+wY
11.983071.856320.12675
20.334222.17357-1.83935
30.617271.08116-0.46389
42.204611.713300.49130
50.72745-0.257680.98513
60.039280.16208-0.12280
70.37609-0.214750.59083
8-1.40549-1.636120.23063
90.402741.04806-0.64532
100.16970-2.227882.39757
A two-variable member of the multivariate normal family. Increasing the number of simulated pairs preserves the existing seeded prefix; it does not change the population model.

Build dependence from independent normals

Let Z₁ and Z₂ be independent standard normals. The experiment constructs:

X=σXZ1,Y=σY(ρZ1+1−ρ2Z2).X=\sigma_X Z_1,\qquad Y=\sigma_Y\bigl(\rho Z_1+\sqrt{1-\rho^2}Z_2\bigr).

The shared Z₁ term creates covariance ρσXσY. The two independent contributions to Y’s variance add to σY²[ρ²+(1−ρ²)]=σY². Changing ρ preserves both marginal distributions while changing their relationship. For nonzero means, add μX and μY after this construction.

At the defaults σX=1, σY=2, and ρ=0.6, the covariance matrix has entries 1, 1.2, 1.2, 4. Given X=1, the shared component contributes 2·0.6·1=1.2 to Y, while its independent residual has variance 4·(1−0.36)=2.56. Thus the conditional SD is 1.6, smaller than the marginal SD 2.

A conditional distribution is a model calculation

The same construction gives:

Y∣X=x∼N(ρσYσXx, σY2(1−ρ2)).Y\mid X=x\sim N\left(\rho\frac{\sigma_Y}{\sigma_X}x,\ \sigma_Y^2(1-\rho^2)\right).

Here N’s second argument is variance. Conditioning on another x changes the mean but not this conditional variance. For |ρ|=1 the residual vanishes and the formula describes a point mass. The demo shows its right-continuous CDF rather than trying to draw a zero-width density.

For continuous X, a particular value has probability zero. The conditional normal is supplied by the specified joint model, not by counting simulated rows with exact equality. Estimating a conditional distribution from real data requires additional choices and evidence.

Make a prediction

At ρ=−1 with σX=1 and σY=2, what does observing X=1 imply?

Explore the answer

Y must equal −2 under this model. The conditional variance is zero, even though the marginal variance of Y is four. All joint observations lie on the line Y=−2X.

Joint probability is not two marginal intervals

With |ρ|<1, Z₁²+Z₂² has a chi-squared distribution with two degrees of freedom. Its 95th percentile is −2 log(0.05). Transforming that disk with the construction above gives the gold 95% population ellipse. Its axes are determined by covariance; its visible angle also depends on the chart scales.

This is not a confidence region for an estimated mean. At |ρ|=1 the joint distribution has only one random direction, so the demo uses the standard-normal 95% segment instead. A two-dimensional density and its ellipse formula no longer apply. The sampled fraction inside the region fluctuates; the population probability remains 0.95.

Project the pair onto one number

For U=X+wY, expand its square and take expectations:

Var⁡(U)=σX2+w2σY2+2wρσXσY.\operatorname{Var}(U)=\sigma_X^2+w^2\sigma_Y^2+2w\rho\sigma_X\sigma_Y.

At the defaults with w=−1, Var(X−Y)=1+4−2·1.2=2.6. Its mean is zero, so P(X−Y≤1)=Φ(1/√2.6), about 0.732. The table compares that full-distribution probability with the seeded sample fraction. At ρ=1, equal SDs, and w=−1, U is identically zero, so its inclusive CDF at zero is one.

Make a prediction

Do normal marginals and zero correlation always imply independence?

Explore the answer

No. Let Z be standard normal and S an independent fair sign, then set X=Z and Y=SZ. Each marginal is standard normal and E[XY]=E[S]E[Z²]=0, but |X|=|Y| always. Their joint distribution lies on two crossing lines and is not jointly normal. Independence at zero correlation relies on the joint-normal assumption.

Regression to the mean applies the conditional-mean idea to repeated measurements. Total variance explains how marginal spread combines conditional spread and changing conditional means.

Sources

The Book of Statistical Proofs: conditional multivariate normals derives the general conditional mean and covariance. Random Services: normal samples develops normal-vector transformations and sums of squares. The explicit two-normal construction above also lets you verify the displayed covariance, conditional moments, and singular cases directly.

Return to learning paths and foundations.

Reset all settings