Updated
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
Central Limit Theorem
You will learn: Distinguish source observations from sample means and check when a normal approximation is useful.
Start with: Normal distribution · Law of Large Numbers
New to the notation? Start with the connected foundation for the underlying definitions and a worked example.
One waiting time can be strongly skewed. Average several independent waiting times, repeat that experiment, and the distribution of those averages often looks more symmetric. The Central Limit Theorem explains a precise version of this pattern.
It concerns the distribution of an average across hypothetical repeated samples. It does not turn the original observations into normal data.
Follow one sample into many means
The demo shows a sample and its mean, then repeats that experiment. n is the number of observations per mean. k is how many independent means appear in the histogram. Those controls answer different questions.
The first sample: 0.31. Its mean is 0.308. Repeat that experiment 2000 times to make the histogram.
Population mean = 0; standard error of the mean = 2.062. Doubling n multiplies this standard error by 1/√2 ≈ 0.707. Increasing k only makes the simulated histogram more precise.
Start with exponential data and n = 1. Increase n while keeping k fixed. Then keep n fixed and raise k. The first change alters the sampling distribution; the second gives a less noisy picture of that same distribution.
The theorem and its assumptions
For independent, identically distributed observations with mean μ and finite positive variance σ²:
The arrow describes convergence of distribution functions to the standard normal CDF. Subtract the center and divide by the standard error. Without that normalization, the mean concentrates at μ; it does not converge to a fixed nondegenerate bell curve.
For a sufficiently large finite n, the corresponding approximation on the original scale is:
Here the second parameter denotes variance. The standard error is σ/√n. This expression is a useful approximation, not the convergence statement itself.
Standardization makes the comparison fair
Switch the demo to the standardized scale. When population moments exist, the reference curve is always N(0,1). Differences in the shapes are easier to see without simultaneously changing the center and width.
For a normal source the mean is exactly normal for every n. For a skewed source the approximation may improve with n, but there is no universal sample size at which all relevant probabilities become accurate.
Make a prediction
How many observations are needed to halve the standard error: twice as many or four times as many?
Show a hint
The standard error is proportional to 1/√n.
Explore the answer
Four times as many. Doubling n multiplies the standard error by 1/√2 ≈ 0.707; quadrupling multiplies it by 1/2. This describes repeated-sampling spread, not a promise that a particular sample mean will move closer to μ.
Worked example: an average waiting time
Suppose independent waiting times are exponential with mean 2 minutes and standard deviation 2 minutes. For n = 100, the sample mean has standard error 2/√100 = 0.2 minutes.
A mean between 1.6 and 2.4 minutes is within two standard errors of the population mean. Standardization turns that event into a value between −2 and 2. The normal approximation assigns probability about 95.45%.
The individual waits remain exponential. An exact calculation is also possible here because their sum is gamma-distributed; the CLT is useful when an exact sampling distribution is inconvenient or unavailable.
Slow convergence is different from failure
Choose rare Bernoulli observations, with success probability 0.01. At n = 30, the expected count is only 0.3 and the probability of no successes is 0.99³⁰ ≈ 73.97%. A smooth bell cannot describe that large point mass well. The variance is finite; the theorem applies asymptotically, but this sample size is not large enough for every purpose.
Now choose Cauchy. The population mean and variance do not exist. Independent standard Cauchy averages remain standard Cauchy, even at large n. Waiting longer does not repair a missing assumption. The demo therefore displays no standard error or normal reference for this case.
Dependence can change the variance formula too. Repeated measurements from one strongly correlated process do not generally have the same information as independent observations.
LLN and CLT answer different questions
The Law of Large Numbers describes averages approaching a population mean under its assumptions. The CLT describes the limiting shape of standardized fluctuations around that mean. Neither says errors improve monotonically, and neither says a run must “compensate” for earlier outcomes.
Check your understanding
Make a prediction
A histogram of 5,000 sample means looks smooth. Does that prove n was large enough for a normal approximation?
Explore the answer
No. Increasing the number of means k can reveal a smooth but non-normal sampling distribution. Check the source, n, assumptions, and the specific event you want to approximate. For a rare Bernoulli source, more simulated means can make the point masses clearer rather than make them disappear.
References and next steps
- ProbabilityCourse: the Central Limit Theorem, for the standardized statement and CDF interpretation.
- Confidence intervals use sampling distributions to quantify uncertainty; Cauchy explores the failure of averaging in more detail.