paradox #20
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Lindley’s paradox

You will learn: Separate a null tail area from a Bayes factor and posterior model probability.

Start with: Bayes' theorem

A two-sided test reports a p-value near 0.05, yet a Bayesian model comparison assigns a large posterior probability to the point null. These numbers can coexist because they ask different questions and use different assumptions. Lindley’s paradox makes the difference especially visible as sample size changes.

Consider n independent observations from a Normal distribution with unknown mean μ and known standard deviation one. The null hypothesis fixes μ at zero. The alternative gives μ a proper Normal prior with mean zero and standard deviation τ. A separate prior probability q is assigned to the null model itself.

Hold the test statistic fixed

Observations are independent Normal(μ, 1), with known standard deviation 1. H₀ fixes μ = 0. Under H₁, μ has a proper Normal(0, τ²) prior with τ = 1.000000. Observed mean: 0.019600000; standard error: 0.010000000. Changing n holds z fixed, so it changes the observed mean.

QuantityValueMeaning
Two-sided p-value0.049995790Null tail area beyond |z|
Bayes factor BF₀₁14.652519Null marginal density / alternative marginal density
Posterior P(H₀ | data)0.93611252Uses the displayed prior model probability
Fixed-z comparison across sample sizes00.20.40.60.810123456log10 sample sizeprobability / tail area

Solid: posterior null probability. Dashed: p-value, unchanged because z is fixed. The shared vertical scale compares magnitudes; the two quantities answer different questions. Log10 sample sizes 2, 4, and 6 mean 100, 10,000, and 1,000,000 observations.

An exact normal-model Bayes factor, not an estimated-variance t-test. Prior width and prior model probability are distinct inputs. No automatic accept/reject label is inferred from either number; a decision also needs a rule and error costs.

The experiment holds the observed statistic z fixed when you change n. Since z is the sample mean divided by its standard error, this means the observed mean changes to z/√n. It does not hold a nonzero mean fixed while adding more and more observations.

The p-value is the probability, under the null, of a statistic at least as far from zero as the observed one: 2P(Z ≥ |z|). This tail area does not depend on n once z is specified. It is not the probability that the null hypothesis is true, and it does not use the alternative prior width or the prior probability q.

Compare predictive densities

Under the null, the sample mean has distribution Normal(0, 1/n). Under the alternative, integrating over μ gives Normal(0, τ² + 1/n). The Bayes factor BF₀₁ is the null predictive density at the observed mean divided by the alternative predictive density there. It compares densities at the observation, rather than adding a tail area.

Taking the ratio of the two normal densities gives

BF₀₁ = sqrt(1 + nτ²) exp[−z² nτ² / (2(1 + nτ²))].

A factor greater than one means these data increase the odds of the null relative to this particular alternative. Posterior odds equal prior odds times the factor. Thus the posterior null probability is q BF₀₁ / (q BF₀₁ + 1 − q).

The prior width τ and the prior model probability q play different roles. Width describes which means the alternative predicts. Model probability describes the relative weight assigned to the two models before seeing data. Changing q changes the posterior probability without changing the Bayes factor or p-value.

Work through the fixed-z sequence

Take z = 1.96, τ = 1, and equal prior model probabilities. The p-value remains approximately 0.04999579. At n = 100, the observed mean is 0.196, the Bayes factor is about 1.50047, and the posterior null probability is 0.60008.

At n = 10,000, the mean is 0.0196, the factor is about 14.65252, and the posterior null probability is 0.93611. At one million observations, the mean is 0.00196 and the posterior probability is about 0.99322. The p-value remains unchanged throughout.

For fixed nonzero τ and fixed z, the exponential term approaches a positive constant while the square-root factor grows like √n. The alternative spreads predictive mass over a much broader range than the null, and increasingly small observed means can favor the concentrated null in a density comparison even while retaining the same standardized tail position.

Prior sensitivity is part of the result

An extremely broad alternative prior is not automatically neutral. Spreading its mass across large positive and negative means lowers its predictive density near small observed means. The width control makes that effect visible. Conversely, as a proper alternative prior becomes concentrated near zero, its predictions approach the null’s, and the Bayes factor approaches one.

Do not replace the proper prior with an unspecified flat density and expect the same model comparison: arbitrary normalization constants do not generally cancel between distinct models. Likewise, assigning zero prior probability to a model prevents finite evidence from reviving it in this setup; the endpoint controls preserve that fact.

Make a prediction

If the observed mean stayed fixed at a nonzero value as n increased, would this same fixed-z argument apply?

Explore the answer

No. The statistic’s magnitude would grow like √n instead of staying fixed. The p-value would shrink, and the normal-model Bayes factor would eventually favor the alternative. The paradoxical sequence depends on specifying what is held constant.

The normal-prior derivation appears in Szabó and van der Vaart’s Bayesian Statistics lecture notes, example 1.15. Continue to Bayes’ theorem for odds updating and the normal distribution for the distinction between densities and tail areas.

Reset all settings