Updated
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
Confidence intervals
You will learn: Interpret repeated-sample coverage without assigning a posterior probability to a fixed parameter.
Start with: Central Limit Theorem
A 95% confidence interval comes from a procedure designed to cover the true parameter in 95% of repeated experiments under its stated assumptions. The data and the interval change from experiment to experiment; the population parameter stays fixed.
That definition does not guarantee that every procedure labeled “95%” achieves 95% coverage for every population. The experiment below lets you investigate both the interpretation and its limitations.
Repeated experiments
Each row starts with a fresh independent sample, computes its mean and sample standard deviation, and constructs a t interval. Solid intervals cover the population mean; dashed intervals miss it.
Average interval width: 0.7358. Approximate 95% Wilson interval for the procedure's achieved coverage: [93.0%, 99.4%]. This measures Monte Carlo uncertainty across the repeated experiments.
| Sample mean | Lower limit | Upper limit | Contains true mean |
|---|---|---|---|
| -0.1043 | -0.5098 | 0.3012 | Yes |
The observed coverage fraction fluctuates. One picture of 100 intervals cannot establish a procedure’s true coverage. Resample, increase the number of experiments, and distinguish simulation error from systematic undercoverage.
Constructing the interval
For independent normal observations with unknown variance, the usual two-sided interval is:
Here t denotes the lower-tail quantile at probability 1−α/2 with n−1 degrees of freedom. The sample standard deviation s uses denominator n−1. The t reference accounts for estimating the population standard deviation; exact nominal coverage depends on normal sampling.
A worked example
Suppose 25 independent normal observations give a mean of 10 and a sample standard deviation of 2. A 95% interval uses t₀.₉₇₅,₂₄ ≈ 2.0639. Its half-width is 2.0639 × 2/√25 ≈ 0.8256, giving [9.1744, 10.8256].
If the sampling distribution is badly misspecified, this arithmetic still produces numbers, but the 95% label may no longer describe the repeated coverage accurately.
When nominal coverage fails
Switch the source to exponential and set n = 10. The sample mean and estimated standard error can both be small in samples that miss the population’s long right tail. Their t intervals then miss the true mean too often.
A reproducible calculation with this site’s sampler, seed 1, and 100,000 independent Exponential(1) experiments gave coverage of 89.911% at n=10, 92.563% at n=30, and 94.097% at n=100 for nominal 95% intervals. Monte Carlo standard errors were about 0.095, 0.083, and 0.075 percentage points. These are simulation estimates, not exact guarantees for those sample sizes.
Increasing n can therefore affect both width and coverage. More experiments in the display only estimate the coverage of the chosen procedure more precisely; they do not fix it.
What a particular interval means
Once an interval has been calculated, it either contains the fixed parameter or does not. Frequentist confidence describes the procedure’s repeated behavior; it is not a posterior probability for this particular interval.
A Bayesian credible interval answers a different question using a posterior distribution, conditional on a model and prior. Both approaches need explicit assumptions; neither interpretation should be substituted for the other.
Make a prediction
Exactly 95 of 100 intervals cover the mean. Does that prove this method has 95% coverage?
Show a hint
Imagine repeating the whole display.
Explore the answer
No. The count is random. Even a procedure with exact 95% coverage will produce different counts across runs. Assess the procedure mathematically when possible, or estimate its coverage over enough independent experiments to distinguish systematic error from simulation noise.
Continue exploring
Student’s t explains the reference distribution. The CLT motivates approximations at larger n. The bootstrap offers alternative constructions, but resampling also has assumptions and can fail with small or unrepresentative samples.
References
- NIST: Student’s t distribution.
- The coverage experiment uses the sample-mean t construction above; its regression check lives in this project’s numerical tests.
Read width and coverage together
Use the interval-number control to inspect one sample mean and its limits, then compare that width with the average width across repetitions. Increasing the confidence level widens intervals for the same sampled data. Increasing sample size usually narrows them, but realized widths also fluctuate because the estimated standard deviation changes.
The interval for achieved coverage answers a second question: how precisely do these repeated experiments estimate the procedure’s coverage rate? It uses the count of covered intervals as binomial successes. That Monte Carlo interval is not a confidence interval for the population mean, and it does not turn nominal confidence into guaranteed coverage for a skewed source.
Make a prediction
You double the number of repeated experiments without changing observations per sample. Should each sample's interval become half as wide?
Explore the answer
No. Repetitions estimate the procedure’s behavior more precisely. Individual interval width depends on the sample size, estimated spread, and critical value, not how many other samples were simulated.