law #12
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Confidence intervals

You will learn: Interpret repeated-sample coverage without assigning a posterior probability to a fixed parameter.

Start with: Central Limit Theorem

A 95% confidence interval comes from a procedure designed to cover the true parameter in 95% of repeated experiments under its stated assumptions. The data and the interval change from experiment to experiment; the population parameter stays fixed.

That definition does not guarantee that every procedure labeled “95%” achieves 95% coverage for every population. The experiment below lets you investigate both the interpretation and its limitations.

Repeated experiments

Each row starts with a fresh independent sample, computes its mean and sample standard deviation, and constructs a t interval. Solid intervals cover the population mean; dashed intervals miss it.

true μ = 0sample mean ± 95% CI, 100 repeated experiments
covers μ (98/100) misses μ (2/100) Empirical coverage: 98.0% (nominal 95%)

Average interval width: 0.7358. Approximate 95% Wilson interval for the procedure's achieved coverage: [93.0%, 99.4%]. This measures Monte Carlo uncertainty across the repeated experiments.

Interval 1, highlighted with a thicker line and larger dot
Sample meanLower limitUpper limitContains true mean
-0.1043-0.50980.3012Yes
Each row is one independent experiment with a nominal 95% t interval. Exact coverage requires independent normal observations. Other sources can under-cover, especially with small samples. The displayed coverage proportion fluctuates from run to run. Its Wilson interval concerns simulation precision, not the unknown mean in one sample.

The observed coverage fraction fluctuates. One picture of 100 intervals cannot establish a procedure’s true coverage. Resample, increase the number of experiments, and distinguish simulation error from systematic undercoverage.

Constructing the interval

For independent normal observations with unknown variance, the usual two-sided interval is:

Xˉ±t1−α/2, n−1sn\bar X\pm t_{1-\alpha/2,\,n-1}\frac{s}{\sqrt n}

Here t denotes the lower-tail quantile at probability 1−α/2 with n−1 degrees of freedom. The sample standard deviation s uses denominator n−1. The t reference accounts for estimating the population standard deviation; exact nominal coverage depends on normal sampling.

A worked example

Suppose 25 independent normal observations give a mean of 10 and a sample standard deviation of 2. A 95% interval uses t₀.₉₇₅,₂₄ ≈ 2.0639. Its half-width is 2.0639 × 2/√25 ≈ 0.8256, giving [9.1744, 10.8256].

If the sampling distribution is badly misspecified, this arithmetic still produces numbers, but the 95% label may no longer describe the repeated coverage accurately.

When nominal coverage fails

Switch the source to exponential and set n = 10. The sample mean and estimated standard error can both be small in samples that miss the population’s long right tail. Their t intervals then miss the true mean too often.

A reproducible calculation with this site’s sampler, seed 1, and 100,000 independent Exponential(1) experiments gave coverage of 89.911% at n=10, 92.563% at n=30, and 94.097% at n=100 for nominal 95% intervals. Monte Carlo standard errors were about 0.095, 0.083, and 0.075 percentage points. These are simulation estimates, not exact guarantees for those sample sizes.

Increasing n can therefore affect both width and coverage. More experiments in the display only estimate the coverage of the chosen procedure more precisely; they do not fix it.

What a particular interval means

Once an interval has been calculated, it either contains the fixed parameter or does not. Frequentist confidence describes the procedure’s repeated behavior; it is not a posterior probability for this particular interval.

A Bayesian credible interval answers a different question using a posterior distribution, conditional on a model and prior. Both approaches need explicit assumptions; neither interpretation should be substituted for the other.

Make a prediction

Exactly 95 of 100 intervals cover the mean. Does that prove this method has 95% coverage?

Show a hint

Imagine repeating the whole display.

Explore the answer

No. The count is random. Even a procedure with exact 95% coverage will produce different counts across runs. Assess the procedure mathematically when possible, or estimate its coverage over enough independent experiments to distinguish systematic error from simulation noise.

Continue exploring

Student’s t explains the reference distribution. The CLT motivates approximations at larger n. The bootstrap offers alternative constructions, but resampling also has assumptions and can fail with small or unrepresentative samples.

References

  • NIST: Student’s t distribution.
  • The coverage experiment uses the sample-mean t construction above; its regression check lives in this project’s numerical tests.

Read width and coverage together

Use the interval-number control to inspect one sample mean and its limits, then compare that width with the average width across repetitions. Increasing the confidence level widens intervals for the same sampled data. Increasing sample size usually narrows them, but realized widths also fluctuate because the estimated standard deviation changes.

The interval for achieved coverage answers a second question: how precisely do these repeated experiments estimate the procedure’s coverage rate? It uses the count of covered intervals as binomial successes. That Monte Carlo interval is not a confidence interval for the population mean, and it does not turn nominal confidence into guaranteed coverage for a skewed source.

Make a prediction

You double the number of repeated experiments without changing observations per sample. Should each sample's interval become half as wide?

Explore the answer

No. Repetitions estimate the procedure’s behavior more precisely. Individual interval width depends on the sample size, estimated spread, and critical value, not how many other samples were simulated.

Reset all settings