Updated
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
Chebyshev’s inequality
You will learn: Use finite variance to bound deviations and compare exact tails with the bound.
Start with: Markov's inequality
For a distribution with finite mean μ and finite positive variance σ², the probability of landing standard deviations away from the mean is bounded by :
Here k > 0. No symmetry or density is required. With zero variance, X equals μ almost surely; use a positive absolute-error threshold instead of dividing by its zero standard deviation.
| Same event: |X| ≥ k | Probability |
|---|---|
| Selected continuous model | 0.0455003 |
| Chebyshev bound, capped at 1 | 0.250000 |
| Uncapped 1/k² | 0.250000 |
Every model has mean zero and variance one. Probability 1.973e-9 lies outside the drawing at ±6 and remains in the full calculation. A zero uniform tail is an exact support result.
Attain the bound with three points or fewer
This separate distribution has mean zero and variance one. Its inclusive tail |X| ≥ 2 is exactly 0.250000. The boundary atoms count; replacing ≥ with > changes this construction's probability.
| Value | Probability | In event? | Contribution to E[X²] |
|---|---|---|---|
| -2 | 0.125000 | Yes | 0.500000 |
| 0 | 0.750000 | No | 0.00000 |
| 2 | 0.125000 | Yes | 0.500000 |
What to notice
- At k = 2, the bound says ≤ 25%. For the Normal it’s actually 4.6% — loose by a factor of five. For a variance-one t(3), it is about 4.1%. Tail ordering can change with the threshold: “heavier-tailed” does not mean more mass past every fixed number of standard deviations.
- Simple discrete distributions can make Chebyshev tight. A three-point distribution (for k > 1) that puts probability on each of and the rest at the mean actually achieves the bound. Chebyshev is the price of distribution-freeness.
- Below k = 1 the bound is trivial — it exceeds 1, so it tells you nothing. One-sided versions (Cantelli) tighten this somewhat.
Why it matters
Chebyshev is the shortest path from “finite variance” to a weak form of the Law of Large Numbers. Apply it to the sample mean , from n independent observations with common mean and variance, whose variance is , and you get:
which goes to zero as . That proves convergence in probability under these assumptions. Perfectly copied observations do not have variance σ²/n; their average retains the variance of one observation. The general finite-absolute-mean law of large numbers needs a different argument.
Proof
Apply Markov’s inequality to :
The squaring trick converts the one-sided non-negativity requirement of Markov into a two-sided concentration bound.
An exact boundary example
At k = 2, put probability 1/8 at −2, probability 3/4 at zero, and probability 1/8 at +2. The mean is zero and the variance is . The event |X| ≥ 2 includes both outer atoms, giving 1/4 exactly. The strict event |X| > 2 has probability zero. Keep the event’s boundary convention when reading a tail bound.
For any k at least one, the demo places probability 1/(2k²) at each of ±k. This is a different distribution for each k, so sharpness does not mean that a single distribution attains the formula for every threshold. For k below one, equal probabilities at ±1 attain the capped bound 1.
Make a prediction
Does Chebyshev guarantee that 95% of observations lie within two standard deviations?
Explore the answer
No. At k = 2 it guarantees at least 75% in the open interval (μ − 2σ, μ + 2σ). Approximately 95.45% in that interval is a normal-model result. The three-point example attains only 75%.
A conservative sample-size calculation
For independent observations with common standard deviation 2, require a sample mean to be within 0.5 of its population mean except with probability at most 0.05. Chebyshev gives , so n = 320 is sufficient. It is not claimed to be necessary or optimal, and replacing the known variance with a small sample estimate does not preserve this exact guarantee.
Make a prediction
Can the same bound be applied to a Cauchy population using its sample standard deviation?
Explore the answer
No. The population has no finite variance or mean. A finite sample statistic cannot supply those missing population assumptions.
References and connections
Stanford CS265 lecture 4 develops Markov and Chebyshev with explicit moment assumptions. The atom calculations above verify sharpness directly. Compare Hoeffding when observations are bounded, and confidence intervals for procedures that estimate uncertainty from data.