Updated
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
Log-normal distribution
You will learn: Transform a normal variable and interpret skew, quantiles, and upper-tail probabilities.
Start with: Normal distribution
If X is normally distributed then Y = e^X is log-normal: its logarithm is normal. A product of independent positive factors can be approximately lognormal when a CLT applies to their log values. For example, iid log factors with finite mean and variance admit a standardized-sum CLT. Independence of arbitrary positive factors alone is insufficient.
0 of 500 draws lie outside the displayed original range [0.000, 7.389]. Every draw stays in the histogram denominator. Exact mass above the upper edge: 0.0032%.
| Draw | ln X | X |
|---|---|---|
| 1 | 0.4830 | 1.62091 |
| 2 | 0.5616 | 1.75339 |
| 3 | -0.0246 | 0.975687 |
| 4 | -0.0904 | 0.913524 |
| 5 | 0.6530 | 1.92135 |
Measure upper-tail risk
P(X>3) = 0.0140022. This is a full-tail calculation, including values beyond the chart.
The 95% quantile is 2.276017. The mode is 0.77880 and the variance is 0.364696 in squared original units.
What to notice
- Increasing σ stretches the tail dramatically to the right while the mode moves left. The mean pulls far above the median.
- Short dashes mark the median , long dashes the mean . The gap between them grows with σ — a diagnostic for skew.
- The mode (peak of the density) sits at , lower still.
Why multiplicative processes produce log-normals
Taking logarithms changes a product into a sum: log(X₁ · X₂ · … · Xₙ) = log X₁ + … + log Xₙ. If the log factors satisfy an appropriate CLT, their centered, scaled sum approaches normal. This can motivate a lognormal approximation for a finite product; it does not mean the unscaled product converges to one fixed lognormal law.
Multiplicative growth motivates this model. Dependence, time-varying parameters, and extreme tails can invalidate it; inspect the log-scale data before relying on the approximation.
Mean vs median divergence
The mean is pulled right by rare extreme values — a direct consequence of the lognormal mean formula. This is a statement about the lognormal model, not evidence that a particular income dataset follows it. The median locates the halfway point; the mean is needed for expected totals.
The two scales describe the same observations
At μ=0 and σ=1, the log values have mean zero and standard deviation one. On the original scale, the median is 1, the mean is exp(1/2)≈1.648721, and the mode is exp(−1)≈0.367879. The original-scale variance is (e−1)e≈4.670774. The parameters μ and σ belong to ln X, not X.
Switch the scale selector. The first-five-draw table remains the same sample, while the histogram and its density transform together. Simply relabeling a horizontal axis without changing the density would give the wrong areas. The log view uses density per log unit; the original view uses density per original unit.
An exact multiplicative construction
Let independent log factors A and B both be Normal(0,0.5), where 0.5 denotes variance. Then A+B is Normal(0,1), and exp(A)×exp(B)=exp(A+B) is exactly Lognormal(0,1). The product’s median is one and its mean is exp(1/2). This exact result relies on normal log factors; an approximate CLT argument for many other factors needs separate assumptions.
For a fixed worked pair A=0.4 and B=−0.1, the factors are approximately 1.491825 and 0.904837. Their product is exp(0.3)≈1.349859. Adding the original factors instead would define another quantity and generally would not be lognormal.
A central plot can hide important probability
At μ=0, σ=1, P(X>exp(2))=P(Z>2)≈0.022750, and the 95th percentile is exp(1.644854)≈5.180252. Raise μ to 2 and σ to 2: the model mean becomes exp(4)≈54.598, beyond the original plot’s upper limit of 25. The tail calculation and omitted-count notice retain that distinction.
Make a prediction
Two positive variables have lognormal-looking histograms. Is their product necessarily lognormal with log variances added?
Explore the answer
No. Adding variances requires independence here, and exact normality of the summed logs needs an appropriate joint model. Dependence can change the variance; marginal shapes alone do not specify the distribution of a sum or product.
Reference
Random Services: lognormal distribution gives the log transformation, moments, quantiles, and independent-product relationship. The normal lesson explains the sum rule, while Pareto gives a different positive-tail model.