Updated
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
Jensen’s inequality
You will learn: Relate curvature to the expectation of a transformed variable and its equality cases.
Start with: Expectation and variance
For a convex function, the function of an average is at most the average of the function:
The variable must lie in the function’s convex domain, with finite E[X]. Start with finite transformed expectations; for nonnegative convex transforms the right side may also be positive infinity, yielding a valid but nonfinite comparison. For concave functions such as log and square root, the direction reverses when the expectations are defined.
Exact two-point experiment
X equals a with probability w and b with probability 1 − w. Both points are positive so all four functions are defined. These plotted expectations are exact sums, with no simulation.
| Exact two-point quantity | Value |
|---|---|
| E[X] | 5.000000 |
| f(E[X]), indigo curve point | 25.00000 |
| E[f(X)], amber chord point | 41.00000 |
| E[f(X)] − f(E[X]) | 16.00000 |
The function is convex: the chord point is at or above the curve. Set w to 0 or 1, or make a = b, to see equality. Moving w need not increase the gap monotonically.
A separate log-normal sample
Here log X is normal with location μ and standard deviation σ. The table separates model expectations from finite sample calculations. These controls do not move the two-point drawing.
| Quantity | Population model | Seeded sample |
|---|---|---|
| Mean of X | 4.241852 | 4.177822 |
| Function of the mean | 17.99331 | 17.45420 |
| Mean of the function | 29.37077 | 28.28404 |
Read the chord as a probability calculation
The first experiment has only two possible outcomes: a with probability w, and b with probability 1 − w. Its mean is wa + (1 − w)b. The point on the straight chord at that horizontal coordinate has height wf(a) + (1 − w)f(b); the point on the curve has height f(wa + (1 − w)b). Convexity puts the chord on or above the curve.
With equal chances of 1 and 9, E[X] = 5. For squaring, E[X²] = 41 while E[X]² = 25. Their difference, 16, is exactly the variance. Here changing the probability can first enlarge and then shrink the gap; “more variation gives a larger gap” is not a universal comparison between arbitrary distributions and transforms.
Make a prediction
For X equal to 1 or 9, which settings make the strictly convex square function attain equality?
Explore the answer
Set w to zero or one, so X is constant. Making a = b also makes it constant. With both probabilities positive and distinct outcomes, the gap is positive. For a general affine function, equality holds even for nonconstant X.
A log-utility calculation
Keep equal chances of 1 and 9 and switch to log. The arithmetic mean is 5, but the mean log is . Exponentiating it gives the geometric mean 3. Thus a person whose utility is log wealth would be indifferent between this lottery and certain wealth 3, while certain wealth 5 has higher expected utility. This is a stated utility model, not a claim about everyone’s preferences.
The positivity assumption matters: log is undefined at zero and at negative real values. Square root is concave on its nonnegative domain; the demo uses positive endpoints for a shared comparison.
Make a prediction
Does log E[X] equal E[log X] when X is positive?
Explore the answer
Only in equality cases. Positivity makes the expressions meaningful, but does not let a nonlinear function pass through an expectation. With equal outcomes 1 and 9, the two values are log 5 and log 3.
Samples and populations are different averages
The second experiment uses X = exp(Z), with Z normal of mean μ and variance σ². Its exact moments give:
A finite sample has its own equal-weight Jensen inequality. The sample mean of f(X) and f(sample mean) must obey the relevant ordering, but neither is automatically a population expectation. The table keeps the two columns separate.
The exponential transform exposes a stronger limitation: for every σ > 0, a log-normal X has E[exp(X)] = ∞. Writing Z’s density in the expectation produces an integrand proportional to . As z increases, the eᶻ term eventually dominates the negative quadratic, so the integral diverges. This happens even though every positive integer power moment of X is finite.
A particular finite sample still has a finite mathematical average of exp(X). The demo computes its logarithm by subtracting the largest exponent before exponentiation. If the result is too large for ordinary numeric display, it retains that logarithmic value instead of dropping observations or drawing infinite coordinates. Increasing n cannot make this sample average estimate a finite population target that does not exist.
Why the inequality holds
For a differentiable convex function and an interior point m = E[X], its tangent is a supporting line:
Taking expectations eliminates the linear term because E[X − m] = 0. Nondifferentiable convex functions use a supporting slope instead. Concavity reverses the supporting-line inequality. Strict convexity gives equality only for a constant X almost surely, under the finite-expectation conditions used in that comparison.
References and next steps
Random Services: Jensen’s inequality develops the supporting-line argument. Statlect: log-normal distribution provides its moments and explains why its moment-generating function does not exist. The finite two-point calculations above can be checked directly.
Continue with expectation and variance for the squaring identity, log-normal variables for multiplicative models, and Markov or Chebyshev for tail bounds with different assumptions.