distribution #15
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Chi-squared distribution

You will learn: Connect squared normal variables with variance estimation and reversed interval limits.

Start with: Normal distribution

For a positive integer k, take kk independent standard normals, square each, and add them up. The result is chi-squared with kk degrees of freedom.

X=Z12+Z22+⋯+Zk2X=Z_{1}^{2}+Z_{2}^{2}+\cdots+Z_{k}^{2}

This construction gives mean k and variance 2k. Its density is:

f(x;k)=xk/2−1e−x/22k/2 Γ(k/2)f(x;k)=\frac{x^{k/2-1}e^{-x/2}}{2^{k/2}\,\Gamma(k/2)}
density by x00.050.10.150.2024681012141618xdensity
Chi-squared (mean = k = 4.0) Normal: mean k, SD √(2k)

Build one squared-normal sum

A separate seeded construction with 4 independent standard normals
TermZZ²
1-0.49510.2452
20.46470.2159
3-1.15391.3316
40.05470.0030

The sum of squares is 1.79566. Each term has mean 1 and variance 2. Independence gives total mean 4 and variance 8.

Skewness = √(8/k) = 1.4142. 2 of 1500 draws exceed the plot's upper edge 18.4668. The curve starts at 0.000001 and the vertical axis is capped at 0.2115. At k<2, the density diverges at zero; clipping the drawing does not remove those values from calculations.

Invert a variance sampling distribution

Assume 5 independent normal observations with unknown mean and variance. Their sample variance uses n−1=4 degrees of freedom. With the entered S², the 95% interval for the population variance is [0.538441, 12.3860]. It divides (n−1)S² by the upper chi-squared quantile 11.14329 for the lower limit, and by the lower quantile 0.48442 for the upper limit. This is a repeated-sampling coverage statement, not a posterior probability about this particular interval.

PDF f(x;k)=xk/2−1e−x/22k/2Γ(k/2)(x≥0)f(x;k)=\frac{x^{k/2-1}e^{-x/2}}{2^{k/2}\Gamma(k/2)}\quad(x\ge 0). For integer k, the sum of k independent squared standard normals: X=Z12+Z22+⋯+Zk2,Zi∼N(0,1)X=Z_1^{2}+Z_2^{2}+\cdots+Z_k^{2},\quad Z_i\sim\mathcal{N}(0,1). The dashed normal approximation has mean k and standard deviation √(2k). For fractional k, sampling uses the gamma-family extension. All draws remain in histogram denominators.

What to notice

  • Small k is strongly right-skewed. At k=1k=1 the mode sits at zero and the density blows up there — a single Normal squared spends most of its time near zero.
  • Large k goes Normal. The dashed normal curve has mean k and standard deviation √(2k). The approximation improves with k, but at k=30 the skewness is still √(8/30) ≈ 0.516. Similar-looking central regions do not guarantee accurate tail probabilities.
  • Mean and variance are both linear in k. Adding one more squared Normal adds 1 to the mean and 2 to the variance.

Why it matters

Chi-squared reference distributions arise in several specific constructions. Squaring a statistic alone does not make it chi-squared:

  • Sample variance. If your data are iid Normal, the scaled sample variance is exactly chi-squared:
    (n−1)S2σ2∼χn−12\frac{(n-1)S^{2}}{\sigma^{2}}\sim\chi^{2}_{n-1}
    This is the denominator of Student’s t.
  • Pearson’s χ² test. Summing squared residuals of observed vs. expected counts produces an approximate chi-squared.
  • Likelihood ratios. Under regularity conditions, −2 log(LR) is chi-squared in the sample-size limit (Wilks’ theorem).

The F distribution uses a ratio of independent scaled chi-squared variables.

The squared-normal construction uses integer k. The density extends to every real k>0 through the Gamma(k/2, rate 1/2) family.

Squaring requires a distributional assumption

For a fixed illustrative triple Z=(1,−2,0.5), the squares sum to 1+4+0.25=5.25. If the three components were generated independently from standard normals, the sum would have a chi-squared reference with k=3, mean 3, and variance 6. The table constructs one such seeded sum; its particular value need not equal the expectation.

Contrast independent variables taking −1 or +1 with equal probability. Each has mean zero and variance one, but every square equals one, so their sum is exactly k. Matching the first two moments is not enough to obtain a chi-squared law.

Why an estimated mean costs one degree of freedom

For n independent normal observations, estimating the mean forces the residuals to sum to zero. Their squared sum, divided by the true variance, has a chi-squared law with n−1 degrees of freedom. The residuals themselves are dependent; an orthogonal set of n−1 normal contrasts explains the squared-normal representation. Using n in place of n−1 changes the sample-variance calculation.

To invert this law, start from lower≤(n−1)S²/σ²≤upper and solve for σ². The upper chi-squared quantile belongs in the denominator of the lower variance limit. Reversing the quantiles is a common error; the demo displays both explicitly.

At n=10 and S²=1.5, the 2.5% and 97.5% chi-squared quantiles with nine degrees of freedom are approximately 2.700389 and 19.022768. A nominal 95% variance interval is [0.709676, 4.999279], in squared measurement units. Taking square roots of both endpoints gives the corresponding standard-deviation interval. It is asymmetric because the reference distribution is skewed.

Make a prediction

If measurements are independent but strongly non-normal, is this variance interval still guaranteed to have exactly 95% coverage?

Explore the answer

No. Its exact pivot uses a normal population. Independence and a finite variance alone do not make a squared-residual sum chi-squared. Model checking and an appropriate alternative procedure are separate from computing these reference quantiles.

References

NIST: chi-square distribution gives the squared-normal construction and moments. NIST: variance test states the scaled sample-variance statistic. Compare the gamma waiting-time interpretation to see why a fractional k is a valid density parameter without being a literal count of normals.

Reset all settings