law #2
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Law of Large Numbers

You will learn: Check convergence of averages without assuming errors improve at every step.

Start with: Expectation and variance

New to the notation? Start with the connected foundation for the underlying definitions and a worked example.

An average can become more precise as it uses more independent observations, while its particular path still moves closer to and farther from the target. The law of large numbers describes a limit under stated conditions. It does not promise exact balance after a finite number of coin flips.

Xˉn=1n∑i=1nXi  ⟶  E[X]\bar{X}_n = \frac{1}{n}\sum_{i=1}^{n} X_i \;\longrightarrow\; \mathbb{E}[X]
running fraction of heads by trial0.000.200.400.600.801.002004006008001000trialrunning fraction of heads

The dashed reference is the finite population mean p=0.5. Its sample-mean SD is √(p(1−p)/n)=0.0158114. Changing n extends each run's existing observations; changing runs adds independent sequences.

At n=1000, 20/20 simulated averages are within 0.05 of the reference. The model probability is 0.998608.

Along the dark focus run, the distance from the reference increased on 504 of 999 updates. Convergence is not monotone improvement.

Compare every final sample average
All runs at the selected sample size
RunAverageWithin chosen distance?
10.5050000Yes
20.5120000Yes
30.5040000Yes
40.4840000Yes
50.4980000Yes
60.4750000Yes
70.4800000Yes
80.4820000Yes
90.5030000Yes
100.5260000Yes
110.4910000Yes
120.4910000Yes
130.5070000Yes
140.5080000Yes
150.5080000Yes
160.4760000Yes
170.4830000Yes
180.5070000Yes
190.4820000Yes
200.4870000Yes
Light blue lines: 20 independent sequences, each containing 1000 observations. The dark line is a focus sequence. The reference is the mean for the coin model and the symmetry center for Cauchy.

What to try

  • With default settings (p = 0.5, n = 1000, 20 runs), compare the final spread around 0.5 and count how often the focus run moves farther from it.
  • Drop n to 50 and resample a few times. The curves end up all over the place — one run might finish at 0.7, another at 0.3.
  • Crank p to 0.2. All lines still converge, but to 0.2 instead of 0.5. The destination is the parameter; the journey is noise.

How fast?

For independent observations with finite variance σ², the variance of their average is exactly σ²/n. The standard error is therefore σ/n\sigma / \sqrt{n}. This describes a distribution of averages, not the realized absolute error on every path. A normal approximation to those fluctuations is an additional CLT question.

For a fair coin, the standard error is 0.5/√n: 0.05 at n=100 and 0.025 at n=400. Quadrupling independent sample size halves that scale. Increasing the number of plotted runs estimates the distribution more precisely, but does not reduce the variability of an individual fixed-size average.

Why it matters

Averaging is useful only in relation to the sampling model and target. Copying one observation many times does not create independent information. Changing the data-generating probability partway through a run changes the target; this demo instead holds the selected probability fixed within each complete run.

Work a finite event rather than promise convergence

With ten fair flips, being within 0.2 of p=0.5 means observing 3 through 7 heads, inclusive. Summing those five binomial probabilities gives 0.890625. It is a high probability, not certainty. The demo performs the corresponding exact binomial calculation at the selected n, p, and tolerance, alongside the observed fraction of runs.

A path can temporarily worsen. After six heads in ten flips its fraction is 0.6, at distance 0.1 from 0.5. One more head gives 7/11≈0.636364, farther from the target. That step is compatible with long-run convergence. The experiment counts such increases along its focus run.

Check the finite-mean condition

For iid observations, finite E[|X|] is sufficient for the usual strong law. Finite variance is not required for that result: a Pareto variable with minimum 1 and shape 1.5 has mean 3 and infinite variance, yet its independent averages converge to 3. The simple variance-based rate calculation above does not apply there.

Standard Cauchy observations do not have an expectation. Their independent average remains standard Cauchy at every sample size, so P(|average|≤ε)=2 arctan(ε)/π does not approach one. Zero in this view is a symmetry center, not a mean. More runs estimate that persistent distribution; more observations per average do not make it collapse to zero.

Make a prediction

If the first 100 fair flips contain unusually many heads, does the LLN make the next flip more likely to be tails?

Explore the answer

No. Under independent fair flips its chance of tails remains 1/2. Earlier imbalance becomes a smaller fraction of a longer sequence without any compensating change in the next-trial probability.

References

Random Services: the law of large numbers distinguishes convergence modes, derives the finite-variance calculation, and states the finite-absolute-mean extension. Compare the Cauchy average experiment, Pareto moment thresholds, and gambler’s fallacy.

Reset all settings