law #11
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

The bootstrap

You will learn: Distinguish resampling uncertainty from the original sample and examine failure cases.

Start with: Empirical CDF · Expectation and variance

Got one sample. Need a standard error. Traditional approach: derive the sampling distribution of your statistic, plug in the variance formula, hope you didn’t mess up the algebra. Bootstrap approach: resample your data with replacement a few thousand times, take the statistic of each resample, and use the spread of those numbers as the standard error.

density by bootstrap sample mean00.511.520.20.40.60.811.21.41.61.8bootstrap sample meandensity
bootstrap means (1000 resamples) Normal approximation: mean x̄, SD s/√n original x̄ = 0.957

Original sample: n = 30, x̄ = 0.957, s = 1.206. Bootstrap SE = 0.225; CLT SE = s/√n = 0.220. 95% percentile CI: [0.566, 1.439].

Exact conditional bootstrap SE of a mean: 0.216514. Bootstrap bias estimate (mean resampled mean − observed mean): -0.00154279; its exact conditional target is zero for the mean. That does not imply other statistics are unbiased. The drawing clips 2/1000 means; all remain in the interval and SE.

Inspect one resample from the empirical distribution

Each original row has probability 1/30 per independent draw. Repeated row indices are allowed, and some original rows can be absent. Inspected resample mean: 1.02446.

Original observations and selected resample indices
Original rowValueTimes selected
10.9863751
20.002739470
30.7496050
43.966003
53.453900
60.3300384
70.9489140
81.275620
90.5547722
105.263521
110.6074671
120.6709720
130.1495841
140.5170861
150.2843861
160.1678191
170.6716670
180.06974662
190.5018981
201.457180
210.3366790
220.2118001
230.04366953
240.5575360
250.8925310
261.602882
270.3493730
280.2300041
291.073323
300.7898401

Draw order (row numbers): 16, 6, 4, 28, 13, 29, 9, 22, 26, 23, 26, 4, 18, 1, 15, 19, 4, 10, 6, 11, 6, 30, 18, 6, 23, 9, 14, 29, 29, 23.

Repeat the entire study

For this synthetic source, the known mean is 1.00000. Of 60 independently generated datasets of size 30, 55 percentile intervals contain that mean. Each uses 300 bootstrap resamples regardless of B above. Nominal 95% is not guaranteed observed coverage; 60 studies also have Monte Carlo variation. Increasing B refines one dataset's conditional calculation, whereas increasing n changes how much original information is available.

All 60 study intervals
StudyLowerUpperContains true mean?
10.6149941.01199Yes
20.7002571.53169Yes
30.4984771.00278Yes
40.3989980.863542No
50.6003081.32508Yes
60.6912961.54414Yes
70.7630671.49678Yes
80.7646671.64315Yes
90.7811101.88340Yes
100.4851931.05470Yes
110.8555951.54894Yes
120.6920971.29363Yes
130.4211860.963309No
140.6908981.23769Yes
150.6323381.31737Yes
160.5917221.25415Yes
170.5334801.03882Yes
180.6794101.23217Yes
190.6531611.70533Yes
200.8268121.71789Yes
210.7129781.31312Yes
220.7187261.35364Yes
230.4192851.32559Yes
240.8231281.50018Yes
250.5600331.21785Yes
260.6739181.18260Yes
270.6071371.38642Yes
280.6744371.34360Yes
290.6337891.13350Yes
300.6843461.22024Yes
310.7822391.43043Yes
320.4640791.09348Yes
330.7157151.30901Yes
340.8917672.10807Yes
350.6890931.24719Yes
361.057721.81553No
370.6340621.34415Yes
380.7004891.48767Yes
390.5578101.27743Yes
400.7230321.26953Yes
410.6329371.45468Yes
420.6937761.41343Yes
430.7794321.24493Yes
440.8235391.59524Yes
450.5637671.00771Yes
460.6466941.40146Yes
470.6335291.27172Yes
480.6749721.70911Yes
490.7453401.84181Yes
500.7712621.49717Yes
510.7436971.24132Yes
520.5787161.24064Yes
530.4258810.946723No
541.024951.62202No
550.5407041.13831Yes
560.6512231.29432Yes
570.6652661.26498Yes
580.6851891.50123Yes
590.7947031.51474Yes
600.9638101.70031Yes
Resample with replacement from the single observed dataset, compute the statistic on each resample, repeat B times. For the sample mean: SE^boot≈s/n\widehat{\mathrm{SE}}_{\mathrm{boot}}\approx s/\sqrt{n}, matching the CLT. The resampling design and statistic determine whether the bootstrap is valid; there is no universal guarantee for quantiles, maxima, or dependent data.

What to notice

  • Bootstrap SE ≈ CLT SE. For the sample mean, the bootstrap standard error approaches √((n−1)/n) × s/ns/\sqrt{n} as the number of resamples grows, when s uses the n−1 variance denominator. The small finite-sample factor comes from treating the empirical distribution as the population. That’s the sanity check: for easy statistics the bootstrap reproduces the analytic answer.
  • Skewness does not disappear automatically. A small lognormal sample can miss influential tail observations. Its bootstrap distribution may look smooth while understating the uncertainty in a new sample from the population.
  • Percentile intervals have nominal coverage. The 2.5th and 97.5th bootstrap percentiles define a nominal 95% interval. Actual repeated-sample coverage need not be 95%, especially with small samples, bias, or heavy tails.

Work through one resample

Suppose your data are [1, 2, 6], with mean 3. A bootstrap sample takes three independent draws with replacement from those three values. [6, 6, 1] is allowed and has mean 13/3. [2, 2, 2] is also allowed. Repeating this generates a distribution of resampled means around the observed mean, not around an independently known population mean.

The empirical variance is (4+1+9)/3 = 14/3. Therefore the exact conditional variance of a size-three bootstrap mean is 14/9, and its standard error is about 1.247. By comparison s/√3 is √(7/3), about 1.528. More bootstrap repetitions reduce simulation noise; they do not turn three observations into a larger original sample.

Why it matters

The bootstrap shines when analytic SEs don’t exist or are too fragile:

  • Median, quantiles, IQR. Resampling can approximate uncertainty under suitable regularity conditions, such as a positive density at the target quantile; extreme quantiles and small samples need particular care.
  • Non-linear functions of the data. Correlation, R², ratios, ROC-AUC. Delta-method approximations get hairy; bootstrap can help, but consistency and coverage need checking for the statistic and sampling design.
  • Regression coefficients with heteroskedasticity. Cluster resampling preserves within-cluster dependence; a wild bootstrap can address heteroskedastic regression errors. The resampling scheme must match the design.

Variants worth knowing

  • BCₐ (bias-corrected and accelerated). Adjusts the percentile CI for skewness and bias in the bootstrap distribution. It can improve coverage under suitable regularity conditions; improvement is not universal.
  • Parametric bootstrap. Fit a model, then resample from the fitted model instead of from the data. Relies on the fitted model being an adequate description of the sampling process.
  • Block bootstrap. For time series, resample blocks of consecutive observations to preserve autocorrelation.
SE^boot=1B−1∑b=1B(θ^∗b−θ^∗ˉ)2\widehat{\mathrm{SE}}_{\mathrm{boot}} = \sqrt{\frac{1}{B-1}\sum_{b=1}^{B}\left(\hat\theta^{*b}-\bar{\hat\theta^{*}}\right)^{2}}
[θ^(α/2)∗,  θ^(1−α/2)∗][\hat\theta^{*}_{(\alpha/2)},\;\hat\theta^{*}_{(1-\alpha/2)}]

Read the resampling experiment

Open the original-observation table and select a bootstrap replicate. Every row is assigned mass 1/n in the empirical distribution, even when two rows have the same value. The table gives selection counts and the draw order, making duplicates and omitted rows visible. Changing B keeps the original data and earlier resamples; changing the source or original n changes the dataset being analyzed.

The interval uses the empirical inverse CDF without interpolation: the p-quantile is the sorted value in position ceil(Bp), counted from one. Different percentile conventions can give slightly different endpoints at finite B. The separate repeated-study experiment generates 60 new datasets from the known synthetic source and uses 300 resamples per dataset, so it asks about coverage rather than just conditional simulation precision.

Make a prediction

Would raising B from 100 to 5,000 fix a dataset that contains no rare high observations?

Explore the answer

No. Each resample still uses only the observed values. More B stabilizes that empirical resampling distribution; it cannot add missing population information. The 60-study comparison can expose a coverage problem even when one histogram looks stable.

A failure you can prove without simulation

Suppose the target is the unknown upper endpoint θ of Uniform(0, θ), estimated by the sample maximum M. Every ordinary resample has maximum at most M. For observations [0.2, 0.5, 0.8], all 27 ordered size-three resamples have maxima 0.2, 0.5, or 0.8. Their respective counts are 1, 7, and 19, giving a 95% percentile interval [0.2, 0.8]. If the actual endpoint is 1, that interval misses it.

More generally, M is strictly below θ with probability one under this continuous model. Every ordinary percentile bootstrap interval for the endpoint has its upper endpoint at most M, so its coverage of θ is zero. A bootstrap that works for a smooth mean need not work for a support endpoint. A justified parametric or specialized procedure is needed for this different target.

Keep the sampling units intact

Independent row resampling assumes the rows act as independent sampling units. Ten measurements of one person are not ten independent people. Resampling whole people can preserve within-person dependence; time-series settings need an appropriate block or model-based scheme. Selecting a block length or cluster definition is part of the statistical design, not something a larger B resolves.

Make a prediction

Why is the exact bootstrap SE for the mean slightly smaller than s/√n?

Explore the answer

The empirical population variance divides the squared deviations by n, while the usual unbiased sample variance divides by n−1. The resample mean has variance (n−1)s²/n², giving SE √((n−1)/n) · s/√n. The simulation fluctuates around this conditional target.

Reference

CMU Advanced Data Analysis: Bootstrapping explains model-based and empirical resampling and the distinction between percentile and pivotal intervals. The three-value resample and endpoint counterexamples above are finite calculations, rather than claims of universal bootstrap validity.

The empirical CDF lesson makes the resampling distribution explicit, including ties and its inverse-quantile convention.

Reset all settings