Updated
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
Markov’s inequality
You will learn: Apply a nonnegative-variable bound and inspect cases where it is sharp or loose.
Start with: Expectation and variance
For any non-negative random variable and any positive threshold :
A finite mean gives a finite upper bound. A nonnegative variable may instead have infinite expectation, in which case the inequality gives no useful finite constraint. No variance or shape assumption is needed.
| Same event: X ≥ a | Probability |
|---|---|
| Selected continuous model | 0.135335 |
| Markov bound, capped at 1 | 0.500000 |
| Uncapped mean / threshold | 0.500000 |
The shading stops at 5.00. Probability 0.006738 beyond that point remains in the full tail calculation. χ²₄ and Gamma(shape 2, scale 2) are the same distribution, hence give identical results.
A distribution that attains the bound
Keep mean 1, but move probability to the values below. This is a different distribution for each threshold. Its probability of X ≥ 2 equals 0.500000 exactly.
| Value | Probability | In event? | Mean contribution |
|---|---|---|---|
| 0 | 0.500000 | No | 0.00000 |
| 2 | 0.500000 | Yes | 1.00000 |
What if negative values are allowed?
Give −2 and +2 probability 1/2 each. The mean is zero, but P(X ≥ 2) = 1/2. The attempted bound zero is false: nonnegativity is essential.
What to notice
- The bound is usually loose. Pick Exponential(1) and slide a to 5. The true tail is . The Markov bound is 0.2 — thirty times too large. That looseness is the price of demanding almost no assumptions about X.
- It can be trivial. Whenever , Markov says — a bound bigger than 1, which says nothing. The demo shows those cases clamped at 1.
- A smaller bound need not be relatively tighter. For Exponential(1), the ratio of the Markov bound to the true tail is exp(a)/a, which grows for a > 1. The absolute bound decreases while its relative error worsens.
Why it matters
Markov is a useful starting point for several concentration inequalities. Apply it to and you get Chebyshev. Apply it to and you get the Chernoff bound, which leads to Hoeffding and friends. The moment you square, exponentiate, or otherwise transform , Markov’s bound on the transformed variable becomes a bound on the original.
Proof sketch
Since (the indicator is 1 when X is at least a, 0 otherwise), taking expectations on both sides gives . Divide by a. That’s the whole argument.
Build the worst case
Suppose the mean is 1 and the threshold is 2. An exponential wait has probability of reaching two time units. But a variable that equals 2 half the time and 0 otherwise has the same mean and reaches 2 with probability 0.5, exactly the Markov bound. The extra exponential assumption provides information a mean alone cannot supply.
For a general finite mean m and threshold a at least m, put probability m/a at a and the remainder at zero. The mean is still m. When a is below m, the constant variable X = m attains the probability cap 1. The demo constructs these alternatives automatically; none is an exponential density. Atoms on the threshold count in the event ≥.
Make a prediction
The average nonnegative delay is 3 minutes. Must the probability of a delay of at least 12 minutes be 25%?
Explore the answer
No. It is at most 3/12 = 25%. A delay of exactly 3 minutes every time gives probability zero. A delay of 12 minutes with probability 1/4 and zero otherwise attains 25%. Both have mean 3.
Break an assumption
Let X be −2 or +2, equally likely. Its mean is zero, but P(X ≥ 2) is 1/2. Applying the displayed formula would falsely give zero. The proof fails at the negative outcome: X is then smaller than the zero-valued indicator term. Applying Markov to |X| or X² is valid, but requires the expectation of that transformed variable, not E[X].
Make a prediction
If a nonnegative X has mean zero, can it be positive with nonzero probability?
Explore the answer
No. Markov gives P(X ≥ a) = 0 for every a > 0. Taking thresholds 1, 1/2, 1/3, … covers all positive values, so X = 0 almost surely.
References and next steps
Random Services: properties of expectation states Markov and its indicator argument. The two-point alternatives above provide directly checkable sharpness examples. Continue with Chebyshev to use variance, or Hoeffding to add independent bounded observations.