law #22
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Conjugate priors

You will learn: Update beta and gamma priors and distinguish posterior parameters from future predictions.

Start with: Beta distribution · Gamma distribution

How does an uncertain parameter become a prediction for new data? Conjugacy makes the update algebra simple: a chosen prior family remains the posterior family under a particular likelihood. The posterior prediction still requires averaging over that uncertainty.

Start with Bayes’ rule, the beta distribution, and the gamma distribution. The two models below use independent observations conditional on one shared parameter. They do not assume the unknown parameter is redrawn separately for every observed trial.

Prior and posterior over the unknown parameter00.10.20.30.40.5024681012unknown rate λ per time unitdensity

Gold solid: prior. Blue dashed: posterior. Gamma parameters use shape and rate, not scale. Its infinite right tail continues beyond this drawing; calculations use the full distribution. Shapes below one are outside this demo's controls.

Parameter distributionParametersMeanVariance
PriorGamma(2, 1)2.000002.00000
PosteriorGamma(6, 3)2.000000.666667

Observed 4 events in 2 time units. Add events to shape and exposure to rate. Splitting these totals into successive independent batches gives the same posterior. Reusing a batch would count its evidence twice.

Predict a future observation

Predictive count probabilities after integrating parameter uncertainty00.050.10.150.20.250.30246810121416future event countpredictive probability

Future exposure: 1 time units. The negative-binomial predictive distribution has infinite support; probability above 16 is 3.999e-7.

PredictionValue
Mean2.000000
Variance, including parameter uncertainty2.666667
P(zero), integrating the posterior0.1779785
P(zero), plugging in the posterior mean0.1353353
Inspect the predictive probability table
CountProbability
00.1779785
10.2669678
20.2335968
30.1557312
40.08759880
50.04379940
60.02007473
70.008603454
80.003495153
90.001359226
100.0005097098
110.0001853490
120.00006564445
130.00002272308
140.000007709616
150.000002569872
168.432392e-7
Exact conjugate updates under a stated likelihood and prior. Computational convenience does not validate either modeling choice.

Gamma prior, Poisson observations

Suppose a count c comes from a constant-rate Poisson process observed for exposure t. Its likelihood as a function of λ is proportional to λᶜe⁻ᵗλ. Choose a Gamma(α,β) prior in the shape–rate convention, with mean α/β and variance α/β². Multiplying prior and likelihood gives:

λα−1e−βλ λce−tλ ∝ λα+c−1e−(β+t)λ.\lambda^{\alpha-1}e^{-\beta\lambda}\,\lambda^c e^{-t\lambda}\ \propto\ \lambda^{\alpha+c-1}e^{-(\beta+t)\lambda}.

The posterior is Gamma(α+c,β+t). Count adds to shape; exposure adds to rate. Observing zero events for ten hours differs from observing zero events for one hour. The zero counts add no shape, but the longer absence adds more exposure evidence.

In the default example, Gamma(2,1) followed by four events in two hours gives Gamma(6,3). The rate mean stays 2 per hour while its variance falls from 2 to 2/3 in squared rate units. Matching the prior mean does not make the data uninformative.

The posterior mean can be written as a weighted average of the prior mean α/β and observed rate c/t, with weights β/(β+t) and t/(β+t). This is an algebraic interpretation of the parameters, not a claim that the prior came from actual historical events.

Make a prediction

Would four events in four hours give the same posterior as four events in two hours?

Explore the answer

No. With the same Gamma(2,1) prior the four-hour posterior is Gamma(6,5), mean 1.2 per hour. Counts without exposure do not specify the evidence about a rate. If time units change, the rate parameter and all exposure quantities must be transformed consistently.

Predict counts by averaging rates

For future exposure h, first condition on λ: K|λ is Poisson(hλ). Integrating its probability against the gamma posterior produces a negative-binomial count distribution. Writing a,b for posterior shape and rate:

P(K=k∣data)=Γ(a+k)Γ(a)k!(bb+h)a(hb+h)k,k=0,1,….P(K=k\mid\text{data})=\frac{\Gamma(a+k)}{\Gamma(a)k!}\left(\frac b{b+h}\right)^a\left(\frac h{b+h}\right)^k,\quad k=0,1,\ldots.

This parameterization counts future events. It does not require a to be an integer or an actual sequence stopped after a successes. With Gamma(6,3) and h=1, the predictive mean is 2 and variance is 8/3. Its zero-count probability is (3/4)⁶=0.1779785. Plugging the posterior mean 2 into a Poisson distribution instead gives e⁻²=0.1353353.

Total variance explains the extra spread: the predictive variance is hE[λ|data]+h²Var(λ|data). Both models have the same predictive mean, but the plug-in model omits uncertainty about λ.

Beta prior, Bernoulli trials

For s successes and f failures under one fixed success probability p, a Beta(α,β) prior becomes Beta(α+s,β+f). Here β is a second shape, not an exposure rate. Multiplying pˢ(1−p)ᶠ into the prior adds the two exponents directly.

Predicting m new trials gives a beta-binomial distribution: a binomial probability averaged over the beta posterior. For a Beta(2,2) prior and 4 successes/2 failures, the posterior is Beta(6,4). In three future trials, the mean number of successes is 1.8. The probability of none is:

410511612=111≈0.0909091.\frac4{10}\frac5{11}\frac6{12}=\frac1{11}\approx0.0909091.

Plugging in p=0.6 gives 0.4³=0.064 instead. Future trials are independent conditional on p, but are positively associated after averaging over a shared uncertain p. Seeing one future success would also teach you something about the probability of the next.

Sequential updates and their limits

With a common fixed parameter and the stated conditional independence, updating after each batch or pooling their sufficient totals gives the same posterior. For the gamma model, one event in half an hour followed by three in 1.5 hours adds the same count/exposure as four in two hours. Do not use the first batch again when forming the second likelihood.

Conjugacy guarantees a family, not a steadily decreasing numerical variance after every possible observation. Try Gamma(1,10), fifty events, and exposure 0.1. The surprising high count moves the posterior to a much larger rate, and its absolute variance increases. Nor does conjugacy prove that the prior, constant-rate assumption, or independence model is appropriate.

Make a prediction

If the actual rate changes between batches, does pooling their totals prove a single-rate model is correct?

Explore the answer

No. The same-parameter model is an assumption made before the algebra. Pooling can discard temporal evidence of change. Inspect arrivals or rates across windows, and consider a model that represents variation when the evidence calls for it.

Reproduce the calculations in Python

The downloadable script pins Python 3.13, SciPy 1.16.2, and NumPy 2.3.3. Save it and run uv run conjugate-priors.py. SciPy’s gamma uses scale=1/rate. Its negative-binomial arguments here are n=posterior shape and p=rate/(rate+future exposure); the output count is the future event count.

Show the executable example
# /// script
# requires-python = ">=3.13,<3.14"
# dependencies = ["scipy==1.16.2", "numpy==2.3.3"]
# ///

import math

from scipy.stats import betabinom, gamma, nbinom, poisson

shape, rate = 2 + 4, 1 + 2
posterior = gamma(a=shape, scale=1 / rate)
future_exposure = 1
predictive = nbinom(n=shape, p=rate / (rate + future_exposure))

assert math.isclose(posterior.mean(), 2)
assert math.isclose(posterior.var(), 2 / 3)
assert math.isclose(predictive.pmf(0), (3 / 4) ** 6)
assert math.isclose(predictive.var(), 8 / 3)
print(f"Gamma posterior rate mean: {posterior.mean():.7f}")
print(f"Predictive P(zero): {predictive.pmf(0):.7f}")
print(f"Plug-in P(zero): {poisson(mu=2).pmf(0):.7f}")

coin_prediction = betabinom(n=3, a=2 + 4, b=2 + 2)
assert math.isclose(coin_prediction.pmf(0), 1 / 11)
assert math.isclose(coin_prediction.mean(), 1.8)
print(f"Beta-binomial P(zero): {coin_prediction.pmf(0):.7f}")

assert math.isclose(poisson.sf(2, mu=2), gamma.cdf(1, a=3, scale=1 / 2))
print(f"Third arrival within one hour: {poisson.sf(2, mu=2):.7f}")

The assertions check the worked probabilities and the count/wait identity. See SciPy’s gamma, negative-binomial, and beta-binomial parameter documentation.

Sources and next steps

The Book of Statistical Proofs shows gamma conjugacy for a Poisson likelihood. Charles Geyer’s Poisson inference notes discuss gamma priors and parameterization. The exposure and predictive expressions above follow by multiplying or integrating the displayed densities. The beta lesson adds credible intervals; the Poisson process lesson shows what the constant-rate arrival model assumes.

Return to learning paths and foundations.

Reset all settings