law #7
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Markov’s inequality

You will learn: Apply a nonnegative-variable bound and inspect cases where it is sharp or loose.

Start with: Expectation and variance

For any non-negative random variable XX and any positive threshold aa:

P(X≥a)≤E[X]aP(X\ge a) \le \frac{E[X]}{a}

A finite mean gives a finite upper bound. A nonnegative variable may instead have infinite expectation, in which case the inequality gives no useful finite constraint. No variance or shape assumption is needed.

density by x00.20.40.60.81012345xdensity
Same event: X ≥ aProbability
Selected continuous model0.135335
Markov bound, capped at 10.500000
Uncapped mean / threshold0.500000

The shading stops at 5.00. Probability 0.006738 beyond that point remains in the full tail calculation. χ²₄ and Gamma(shape 2, scale 2) are the same distribution, hence give identical results.

A distribution that attains the bound

Keep mean 1, but move probability to the values below. This is a different distribution for each threshold. Its probability of X ≥ 2 equals 0.500000 exactly.

ValueProbabilityIn event?Mean contribution
00.500000No0.00000
20.500000Yes1.00000
What if negative values are allowed?

Give −2 and +2 probability 1/2 each. The mean is zero, but P(X ≥ 2) = 1/2. The attempted bound zero is false: nonnegativity is essential.

The density curve belongs to the selected continuous model. The discrete construction below it shares the mean and attains P(X≥a)≤E[X]aP(X\ge a)\le \frac{E[X]}{a} after capping at one. A mean alone cannot distinguish these tails.

What to notice

  • The bound is usually loose. Pick Exponential(1) and slide a to 5. The true tail is e−5≈0.0067e^{-5}\approx 0.0067. The Markov bound is 0.2 — thirty times too large. That looseness is the price of demanding almost no assumptions about X.
  • It can be trivial. Whenever a<E[X]a<E[X], Markov says P(X≥a)≤E[X]/a>1P(X\ge a)\le E[X]/a > 1 — a bound bigger than 1, which says nothing. The demo shows those cases clamped at 1.
  • A smaller bound need not be relatively tighter. For Exponential(1), the ratio of the Markov bound to the true tail is exp(a)/a, which grows for a > 1. The absolute bound decreases while its relative error worsens.

Why it matters

Markov is a useful starting point for several concentration inequalities. Apply it to (X−μ)2(X-\mu)^{2} and you get Chebyshev. Apply it to etXe^{tX} and you get the Chernoff bound, which leads to Hoeffding and friends. The moment you square, exponentiate, or otherwise transform XX, Markov’s bound on the transformed variable becomes a bound on the original.

Proof sketch

Since X≥a⋅1{X≥a}X\ge a\cdot \mathbb{1}\{X\ge a\} (the indicator is 1 when X is at least a, 0 otherwise), taking expectations on both sides gives E[X]≥a P(X≥a)E[X] \ge a\, P(X\ge a). Divide by a. That’s the whole argument.

Build the worst case

Suppose the mean is 1 and the threshold is 2. An exponential wait has probability e−2≈0.135335e^{-2}\approx0.135335 of reaching two time units. But a variable that equals 2 half the time and 0 otherwise has the same mean and reaches 2 with probability 0.5, exactly the Markov bound. The extra exponential assumption provides information a mean alone cannot supply.

For a general finite mean m and threshold a at least m, put probability m/a at a and the remainder at zero. The mean is still m. When a is below m, the constant variable X = m attains the probability cap 1. The demo constructs these alternatives automatically; none is an exponential density. Atoms on the threshold count in the event ≥.

Make a prediction

The average nonnegative delay is 3 minutes. Must the probability of a delay of at least 12 minutes be 25%?

Explore the answer

No. It is at most 3/12 = 25%. A delay of exactly 3 minutes every time gives probability zero. A delay of 12 minutes with probability 1/4 and zero otherwise attains 25%. Both have mean 3.

Break an assumption

Let X be −2 or +2, equally likely. Its mean is zero, but P(X ≥ 2) is 1/2. Applying the displayed formula would falsely give zero. The proof fails at the negative outcome: X is then smaller than the zero-valued indicator term. Applying Markov to |X| or X² is valid, but requires the expectation of that transformed variable, not E[X].

Make a prediction

If a nonnegative X has mean zero, can it be positive with nonzero probability?

Explore the answer

No. Markov gives P(X ≥ a) = 0 for every a > 0. Taking thresholds 1, 1/2, 1/3, … covers all positive values, so X = 0 almost surely.

References and next steps

Random Services: properties of expectation states Markov and its indicator argument. The two-point alternatives above provide directly checkable sharpness examples. Continue with Chebyshev to use variance, or Hoeffding to add independent bounded observations.

Reset all settings