distribution #19
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Laplace distribution

You will learn: Connect exponential differences, absolute error, and the role of the median.

Start with: Exponential distribution

Also called the double exponential. Take two mirrored exponential tails and glue them at the mode: that’s Laplace. Independent zero-centered Laplace priors on regression coefficients produce an L1 coefficient penalty in a MAP estimate. With a Gaussian error likelihood, this gives the LASSO objective. Laplace errors instead produce absolute-residual regression; these are different uses of the same density.

f(x;μ,b)=12bexp⁡ ⁣(−∣x−μ∣b)f(x;\mu,b)=\frac{1}{2b}\exp\!\left(-\frac{|x-\mu|}{b}\right)
density by x 00.10.20.30.40.5−8−6−4−202468xdensity
Laplace Normal, same variance

Subtract independent exponential waits

First three histogram draws: X=μ+A−B, with independent exponential A and B of rate 1/b
ABX
0.46675.9014-5.4347
0.63970.01910.6206
0.03211.2690-1.2369

Each wait has mean b=1 and variance b²=1.000. Subtraction centers X at μ=0; independence makes its variance 2b²=2.000. Reusing the same wait on both sides would give the constant μ instead.

At a distance of 3 matched standard deviations, P(|X−μ|≥4.2426) is 0.0143696 for Laplace and 0.00269980 for normal. Far tails are heavier for Laplace; that does not imply a larger outside probability at every distance.

0 of 1000 observations are outside the density window. Exact omitted Laplace mass is 0.00033546; full samples remain in the denominator.

Fit a location by absolute residuals

Use the illustrative observations 1, 2, 2, 3, 12. Their median is 2 and mean is 4. This fit is separate from the simulated histogram.

Sum of absolute residuals: 12.000. Sum of squared residuals: 102.000. The first is in observation units, the second in squared units; compare each across candidate centers, not directly with one another.

With fixed Laplace scale b, negative log likelihood is 5 log(2b) + (absolute residual sum)/b. Its minimum is at the median. A normal likelihood instead minimizes squared residuals at the mean. This is a statement about the error likelihood; a Laplace coefficient prior is a separate modeling choice.

PDF f(x;μ,b)=12bexp⁡ ⁣(−∣x−μ∣b)f(x;\mu,b)=\frac{1}{2b}\exp\!\left(-\frac{|x-\mu|}{b}\right). The cusp at μ\mu and the heavier tails are what set it apart from the matched-variance Normal.

What to notice

  • A sharp cusp at μ\mu. Unlike the Normal’s smooth peak, the density has a kink where the two exponential tails meet. That non-differentiability is what drives L1 solutions to exact zeros.
  • Heavier tails than the matched Normal. Decay is exponential in ∣x−μ∣|x-\mu| instead of squared, so the same standard deviation implies much more probability far from centre.
  • Median is the MLE for μ\mu, not the mean. This concerns a location likelihood with Laplace errors, rather than a coefficient prior:
μ^MLE=median(x1,…,xn)\hat{\mu}_{\mathrm{MLE}}=\mathrm{median}(x_1,\dots,x_n)

Why it matters

A likelihood describes observations given model parameters; a prior describes uncertainty about those parameters. Keep that distinction when connecting a density to an optimization objective. Absolute residuals from a Laplace likelihood and absolute coefficient penalties from a Laplace prior can appear in different models.

Construct the two tails

Let A and B be independent exponential waits, each with rate 1/b. Then X=μ+A−B is Laplace(μ,b). For a nonnegative difference d, integrating their joint density along A=B+d gives ∫₀∞ b⁻² exp(−(2u+d)/b) du=exp(−d/b)/(2b). Swapping A and B gives the negative half by symmetry. The table shows the actual waits used for the histogram.

For μ=1 and b=2, the example A=5 and B=1.5 produces X=4.5. Its population mean is 1 and variance is 8. If A and B were the same random wait, X would always equal 1; independence is essential to the claimed distribution.

Compare tails with the same variance

The standard deviation is √2 b. At three matched standard deviations, the Laplace outside probability is exp(−3√2)≈0.014370, compared with approximately 0.002700 for normal. But at one standard deviation the probabilities are about 0.243117 and 0.317311: a sharper center and farther tails leave less probability at some intermediate distances. “Heavier tails” describes sufficiently remote behavior, not every cutoff.

Work an absolute-error fit

For 1,2,2,3,12, a fitted location of 2 gives absolute residual sum 12 and squared residual sum 102. A location of 4 gives 16 and 82. The sample median minimizes the first, while the sample mean minimizes the second. The large observation pulls the mean farther than the median; neither optimization rule proves that its assumed likelihood is appropriate.

With known b, multiplying independent Laplace densities and taking the negative logarithm gives n log(2b) + Σ|xᵢ−μ|/b. The first term does not depend on μ, which explains the median optimizer. With an even sample size, an entire interval between the two middle observations can minimize the absolute loss.

Make a prediction

If a Laplace prior is continuous and centered at zero, does a random draw from it equal zero with positive probability?

Explore the answer

No. A single point has probability zero. An exact zero can occur in a penalized mode estimate because the optimization objective has a cusp; that is different from drawing from the prior.

References

Random Services: Laplace distribution gives the exponential construction, density, and moments. The likelihood calculation above follows directly by multiplying the displayed densities. Continue to normal for squared-error geometry and logistic for another smooth, exponential-tailed model.

Reset all settings