distribution #8
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Log-normal distribution

You will learn: Transform a normal variable and interpret skew, quantiles, and upper-tail probabilities.

Start with: Normal distribution

If X is normally distributed then Y = e^X is log-normal: its logarithm is normal. A product of independent positive factors can be approximately lognormal when a CLT applies to their log values. For example, iid log factors with finite mean and variance admit a standardized-sum CLT. Independence of arbitrary positive factors alone is insufficient.

f(x;μ,σ)=1xσ2πexp⁡ ⁣(−(ln⁡x−μ)22σ2)f(x;\mu,\sigma)=\frac{1}{x\sigma\sqrt{2\pi}}\exp\!\left(-\frac{(\ln x-\mu)^2}{2\sigma^2}\right)
density per original unit by x00.20.40.60.8101234567xdensity per original unit
Median = e^μ
1.000
Mean = e^(μ+σ²/2)
1.133

0 of 500 draws lie outside the displayed original range [0.000, 7.389]. Every draw stays in the histogram denominator. Exact mass above the upper edge: 0.0032%.

The first five draws, expressed on both scales
Drawln XX
10.48301.62091
20.56161.75339
3-0.02460.975687
4-0.09040.913524
50.65301.92135

Measure upper-tail risk

P(X>3) = 0.0140022. This is a full-tail calculation, including values beyond the chart.

The 95% quantile is 2.276017. The mode is 0.77880 and the variance is 0.364696 in squared original units.

ln X is Normal with mean μ and standard deviation σ. Short dashes mark the original-scale median; long dashes mark the original-scale mean. In log view, the single marker is μ, the mean and median of ln X. Changing scale transforms the data and its density together.

What to notice

  • Increasing σ stretches the tail dramatically to the right while the mode moves left. The mean pulls far above the median.
  • Short dashes mark the median eμe^{\mu}, long dashes the mean eμ+σ2/2e^{\mu+\sigma^2/2}. The gap between them grows with σ — a diagnostic for skew.
  • The mode (peak of the density) sits at eμ−σ2e^{\mu-\sigma^2}, lower still.

Why multiplicative processes produce log-normals

Taking logarithms changes a product into a sum: log(X₁ · X₂ · … · Xₙ) = log X₁ + … + log Xₙ. If the log factors satisfy an appropriate CLT, their centered, scaled sum approaches normal. This can motivate a lognormal approximation for a finite product; it does not mean the unscaled product converges to one fixed lognormal law.

Multiplicative growth motivates this model. Dependence, time-varying parameters, and extreme tails can invalidate it; inspect the log-scale data before relying on the approximation.

Mean vs median divergence

The mean is pulled right by rare extreme values — a direct consequence of the lognormal mean formula. This is a statement about the lognormal model, not evidence that a particular income dataset follows it. The median locates the halfway point; the mean is needed for expected totals.

The two scales describe the same observations

At μ=0 and σ=1, the log values have mean zero and standard deviation one. On the original scale, the median is 1, the mean is exp(1/2)≈1.648721, and the mode is exp(−1)≈0.367879. The original-scale variance is (e−1)e≈4.670774. The parameters μ and σ belong to ln X, not X.

Switch the scale selector. The first-five-draw table remains the same sample, while the histogram and its density transform together. Simply relabeling a horizontal axis without changing the density would give the wrong areas. The log view uses density per log unit; the original view uses density per original unit.

An exact multiplicative construction

Let independent log factors A and B both be Normal(0,0.5), where 0.5 denotes variance. Then A+B is Normal(0,1), and exp(A)×exp(B)=exp(A+B) is exactly Lognormal(0,1). The product’s median is one and its mean is exp(1/2). This exact result relies on normal log factors; an approximate CLT argument for many other factors needs separate assumptions.

For a fixed worked pair A=0.4 and B=−0.1, the factors are approximately 1.491825 and 0.904837. Their product is exp(0.3)≈1.349859. Adding the original factors instead would define another quantity and generally would not be lognormal.

A central plot can hide important probability

At μ=0, σ=1, P(X>exp(2))=P(Z>2)≈0.022750, and the 95th percentile is exp(1.644854)≈5.180252. Raise μ to 2 and σ to 2: the model mean becomes exp(4)≈54.598, beyond the original plot’s upper limit of 25. The tail calculation and omitted-count notice retain that distinction.

Make a prediction

Two positive variables have lognormal-looking histograms. Is their product necessarily lognormal with log variances added?

Explore the answer

No. Adding variances requires independence here, and exact normality of the summed logs needs an appropriate joint model. Dependence can change the variance; marginal shapes alone do not specify the distribution of a sum or product.

Reference

Random Services: lognormal distribution gives the log transformation, moments, quantiles, and independent-product relationship. The normal lesson explains the sum rule, while Pareto gives a different positive-tail model.

Reset all settings