Updated
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
The bootstrap
You will learn: Distinguish resampling uncertainty from the original sample and examine failure cases.
Start with: Empirical CDF · Expectation and variance
Got one sample. Need a standard error. Traditional approach: derive the sampling distribution of your statistic, plug in the variance formula, hope you didn’t mess up the algebra. Bootstrap approach: resample your data with replacement a few thousand times, take the statistic of each resample, and use the spread of those numbers as the standard error.
Original sample: n = 30, x̄ = 0.957, s = 1.206. Bootstrap SE = 0.225; CLT SE = s/√n = 0.220. 95% percentile CI: [0.566, 1.439].
Exact conditional bootstrap SE of a mean: 0.216514. Bootstrap bias estimate (mean resampled mean − observed mean): -0.00154279; its exact conditional target is zero for the mean. That does not imply other statistics are unbiased. The drawing clips 2/1000 means; all remain in the interval and SE.
Inspect one resample from the empirical distribution
Each original row has probability 1/30 per independent draw. Repeated row indices are allowed, and some original rows can be absent. Inspected resample mean: 1.02446.
Original observations and selected resample indices
| Original row | Value | Times selected |
|---|---|---|
| 1 | 0.986375 | 1 |
| 2 | 0.00273947 | 0 |
| 3 | 0.749605 | 0 |
| 4 | 3.96600 | 3 |
| 5 | 3.45390 | 0 |
| 6 | 0.330038 | 4 |
| 7 | 0.948914 | 0 |
| 8 | 1.27562 | 0 |
| 9 | 0.554772 | 2 |
| 10 | 5.26352 | 1 |
| 11 | 0.607467 | 1 |
| 12 | 0.670972 | 0 |
| 13 | 0.149584 | 1 |
| 14 | 0.517086 | 1 |
| 15 | 0.284386 | 1 |
| 16 | 0.167819 | 1 |
| 17 | 0.671667 | 0 |
| 18 | 0.0697466 | 2 |
| 19 | 0.501898 | 1 |
| 20 | 1.45718 | 0 |
| 21 | 0.336679 | 0 |
| 22 | 0.211800 | 1 |
| 23 | 0.0436695 | 3 |
| 24 | 0.557536 | 0 |
| 25 | 0.892531 | 0 |
| 26 | 1.60288 | 2 |
| 27 | 0.349373 | 0 |
| 28 | 0.230004 | 1 |
| 29 | 1.07332 | 3 |
| 30 | 0.789840 | 1 |
Draw order (row numbers): 16, 6, 4, 28, 13, 29, 9, 22, 26, 23, 26, 4, 18, 1, 15, 19, 4, 10, 6, 11, 6, 30, 18, 6, 23, 9, 14, 29, 29, 23.
Repeat the entire study
For this synthetic source, the known mean is 1.00000. Of 60 independently generated datasets of size 30, 55 percentile intervals contain that mean. Each uses 300 bootstrap resamples regardless of B above. Nominal 95% is not guaranteed observed coverage; 60 studies also have Monte Carlo variation. Increasing B refines one dataset's conditional calculation, whereas increasing n changes how much original information is available.
All 60 study intervals
| Study | Lower | Upper | Contains true mean? |
|---|---|---|---|
| 1 | 0.614994 | 1.01199 | Yes |
| 2 | 0.700257 | 1.53169 | Yes |
| 3 | 0.498477 | 1.00278 | Yes |
| 4 | 0.398998 | 0.863542 | No |
| 5 | 0.600308 | 1.32508 | Yes |
| 6 | 0.691296 | 1.54414 | Yes |
| 7 | 0.763067 | 1.49678 | Yes |
| 8 | 0.764667 | 1.64315 | Yes |
| 9 | 0.781110 | 1.88340 | Yes |
| 10 | 0.485193 | 1.05470 | Yes |
| 11 | 0.855595 | 1.54894 | Yes |
| 12 | 0.692097 | 1.29363 | Yes |
| 13 | 0.421186 | 0.963309 | No |
| 14 | 0.690898 | 1.23769 | Yes |
| 15 | 0.632338 | 1.31737 | Yes |
| 16 | 0.591722 | 1.25415 | Yes |
| 17 | 0.533480 | 1.03882 | Yes |
| 18 | 0.679410 | 1.23217 | Yes |
| 19 | 0.653161 | 1.70533 | Yes |
| 20 | 0.826812 | 1.71789 | Yes |
| 21 | 0.712978 | 1.31312 | Yes |
| 22 | 0.718726 | 1.35364 | Yes |
| 23 | 0.419285 | 1.32559 | Yes |
| 24 | 0.823128 | 1.50018 | Yes |
| 25 | 0.560033 | 1.21785 | Yes |
| 26 | 0.673918 | 1.18260 | Yes |
| 27 | 0.607137 | 1.38642 | Yes |
| 28 | 0.674437 | 1.34360 | Yes |
| 29 | 0.633789 | 1.13350 | Yes |
| 30 | 0.684346 | 1.22024 | Yes |
| 31 | 0.782239 | 1.43043 | Yes |
| 32 | 0.464079 | 1.09348 | Yes |
| 33 | 0.715715 | 1.30901 | Yes |
| 34 | 0.891767 | 2.10807 | Yes |
| 35 | 0.689093 | 1.24719 | Yes |
| 36 | 1.05772 | 1.81553 | No |
| 37 | 0.634062 | 1.34415 | Yes |
| 38 | 0.700489 | 1.48767 | Yes |
| 39 | 0.557810 | 1.27743 | Yes |
| 40 | 0.723032 | 1.26953 | Yes |
| 41 | 0.632937 | 1.45468 | Yes |
| 42 | 0.693776 | 1.41343 | Yes |
| 43 | 0.779432 | 1.24493 | Yes |
| 44 | 0.823539 | 1.59524 | Yes |
| 45 | 0.563767 | 1.00771 | Yes |
| 46 | 0.646694 | 1.40146 | Yes |
| 47 | 0.633529 | 1.27172 | Yes |
| 48 | 0.674972 | 1.70911 | Yes |
| 49 | 0.745340 | 1.84181 | Yes |
| 50 | 0.771262 | 1.49717 | Yes |
| 51 | 0.743697 | 1.24132 | Yes |
| 52 | 0.578716 | 1.24064 | Yes |
| 53 | 0.425881 | 0.946723 | No |
| 54 | 1.02495 | 1.62202 | No |
| 55 | 0.540704 | 1.13831 | Yes |
| 56 | 0.651223 | 1.29432 | Yes |
| 57 | 0.665266 | 1.26498 | Yes |
| 58 | 0.685189 | 1.50123 | Yes |
| 59 | 0.794703 | 1.51474 | Yes |
| 60 | 0.963810 | 1.70031 | Yes |
What to notice
- Bootstrap SE ≈ CLT SE. For the sample mean, the bootstrap standard error approaches √((n−1)/n) × as the number of resamples grows, when s uses the n−1 variance denominator. The small finite-sample factor comes from treating the empirical distribution as the population. That’s the sanity check: for easy statistics the bootstrap reproduces the analytic answer.
- Skewness does not disappear automatically. A small lognormal sample can miss influential tail observations. Its bootstrap distribution may look smooth while understating the uncertainty in a new sample from the population.
- Percentile intervals have nominal coverage. The 2.5th and 97.5th bootstrap percentiles define a nominal 95% interval. Actual repeated-sample coverage need not be 95%, especially with small samples, bias, or heavy tails.
Work through one resample
Suppose your data are [1, 2, 6], with mean 3. A bootstrap sample takes three independent draws with replacement from those three values. [6, 6, 1] is allowed and has mean 13/3. [2, 2, 2] is also allowed. Repeating this generates a distribution of resampled means around the observed mean, not around an independently known population mean.
The empirical variance is (4+1+9)/3 = 14/3. Therefore the exact conditional variance of a size-three bootstrap mean is 14/9, and its standard error is about 1.247. By comparison s/√3 is √(7/3), about 1.528. More bootstrap repetitions reduce simulation noise; they do not turn three observations into a larger original sample.
Why it matters
The bootstrap shines when analytic SEs don’t exist or are too fragile:
- Median, quantiles, IQR. Resampling can approximate uncertainty under suitable regularity conditions, such as a positive density at the target quantile; extreme quantiles and small samples need particular care.
- Non-linear functions of the data. Correlation, R², ratios, ROC-AUC. Delta-method approximations get hairy; bootstrap can help, but consistency and coverage need checking for the statistic and sampling design.
- Regression coefficients with heteroskedasticity. Cluster resampling preserves within-cluster dependence; a wild bootstrap can address heteroskedastic regression errors. The resampling scheme must match the design.
Variants worth knowing
- BCₐ (bias-corrected and accelerated). Adjusts the percentile CI for skewness and bias in the bootstrap distribution. It can improve coverage under suitable regularity conditions; improvement is not universal.
- Parametric bootstrap. Fit a model, then resample from the fitted model instead of from the data. Relies on the fitted model being an adequate description of the sampling process.
- Block bootstrap. For time series, resample blocks of consecutive observations to preserve autocorrelation.
Read the resampling experiment
Open the original-observation table and select a bootstrap replicate. Every row is assigned mass 1/n in the empirical distribution, even when two rows have the same value. The table gives selection counts and the draw order, making duplicates and omitted rows visible. Changing B keeps the original data and earlier resamples; changing the source or original n changes the dataset being analyzed.
The interval uses the empirical inverse CDF without interpolation: the p-quantile is the sorted value in position ceil(Bp), counted from one. Different percentile conventions can give slightly different endpoints at finite B. The separate repeated-study experiment generates 60 new datasets from the known synthetic source and uses 300 resamples per dataset, so it asks about coverage rather than just conditional simulation precision.
Make a prediction
Would raising B from 100 to 5,000 fix a dataset that contains no rare high observations?
Explore the answer
No. Each resample still uses only the observed values. More B stabilizes that empirical resampling distribution; it cannot add missing population information. The 60-study comparison can expose a coverage problem even when one histogram looks stable.
A failure you can prove without simulation
Suppose the target is the unknown upper endpoint θ of Uniform(0, θ), estimated by the sample maximum M. Every ordinary resample has maximum at most M. For observations [0.2, 0.5, 0.8], all 27 ordered size-three resamples have maxima 0.2, 0.5, or 0.8. Their respective counts are 1, 7, and 19, giving a 95% percentile interval [0.2, 0.8]. If the actual endpoint is 1, that interval misses it.
More generally, M is strictly below θ with probability one under this continuous model. Every ordinary percentile bootstrap interval for the endpoint has its upper endpoint at most M, so its coverage of θ is zero. A bootstrap that works for a smooth mean need not work for a support endpoint. A justified parametric or specialized procedure is needed for this different target.
Keep the sampling units intact
Independent row resampling assumes the rows act as independent sampling units. Ten measurements of one person are not ten independent people. Resampling whole people can preserve within-person dependence; time-series settings need an appropriate block or model-based scheme. Selecting a block length or cluster definition is part of the statistical design, not something a larger B resolves.
Make a prediction
Why is the exact bootstrap SE for the mean slightly smaller than s/√n?
Explore the answer
The empirical population variance divides the squared deviations by n, while the usual unbiased sample variance divides by n−1. The resample mean has variance (n−1)s²/n², giving SE √((n−1)/n) · s/√n. The simulation fluctuates around this conditional target.
Reference
CMU Advanced Data Analysis: Bootstrapping explains model-based and empirical resampling and the distinction between percentile and pivotal intervals. The three-value resample and endpoint counterexamples above are finite calculations, rather than claims of universal bootstrap validity.
The empirical CDF lesson makes the resampling distribution explicit, including ties and its inverse-quantile convention.