Updated
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
Beta distribution
You will learn: Update a probability distribution and distinguish parameter uncertainty from future observations.
Start with: Bayes' theorem · Bernoulli distribution
The Beta distribution lives on the interval (0, 1), making it useful for modelling uncertainty about a probability or a continuous proportion. Rates per unit of time need not be below 1 and generally require a different model. Its two shape parameters, and , make it extraordinarily flexible.
Posterior: Beta(2, 5), mean 0.2857. This is also the posterior predictive chance of heads on the next independent trial with the same unknown probability.
95.0% equal-tailed credible interval: [0.04327187, 0.64123458]. Each outside tail contains 2.50% posterior probability.
Posterior probability that p exceeds 0.5: 0.10938.
The shaded area represents posterior probability inside the interval. This statement is conditional on the prior, data, and independent Bernoulli model with a shared, fixed p.
Predict future observations
| Before observing these data | After observing these data |
|---|---|
| 87.500% | 87.500% |
These predictions average over uncertainty in the shared p. For several flips, substituting the mean p into a binomial model generally gives a different answer.
Shape guide
| α, β | Shape |
|---|---|
| α = β = 1 | Uniform density on this probability scale |
| α = β > 1 | Symmetric bell — certainty near the centre |
| α > β > 1 | Unimodal and left-skewed; more weight toward higher values |
| α < 1 or β < 1 | U-shaped or J-shaped — mass at the extremes |
The Bayesian story
If you start with a flat prior (Beta(1, 1)) and observe h heads and t tails from a coin, the posterior probability of the coin’s bias is Beta(1 + h, 1 + t). The observation controls are separate from the prior shapes: set α = β = 1 first, then enter h heads and t tails.
This is Bayesian updating in closed form. Each new flip shifts the distribution, and with enough evidence the posterior narrows to a spike around the true bias — under independent Bernoulli sampling and a fixed positive Beta prior. The prior can still matter substantially with little data.
An interval is an area
A 95% equal-tailed credible interval leaves 2.5% of the posterior probability below its lower limit and 2.5% above its upper limit. It need not be symmetric about the mean or be the shortest interval containing 95% probability. For a U-shaped beta density, it can include a low-density middle region.
The chart shades that posterior area. The threshold control asks a different question: how much posterior probability lies above a particular value of p? A density’s height at the threshold cannot answer that question; its area can.
Make a prediction
With a uniform prior and no observations, what is the 95% equal-tailed interval?
Explore the answer
[0.025, 0.975]. Density equals 1 across the unit interval, so a length of 0.95 contains probability 0.95. Set α = β = 1 and both observation counts to zero to check.
This is a statement about uncertainty in p under the selected prior and likelihood. It is not an automatic guarantee that this procedure covers a fixed true p in 95% of repeated samples. Compare the sampling interpretation in confidence intervals.
Hold the data fixed and change the prior
The presets use 8 heads and 2 tails. A Beta(1, 1) prior gives Beta(9, 3), with posterior mean 0.75. A Beta(20, 20) prior gives Beta(28, 22), with mean 0.56. Both prior means are 0.5; their strengths differ.
Write the prior mean as m and its strength as s = α + β. The posterior mean is a weighted average of m and the observed heads fraction: the prior receives weight s, the data receive weight h + t. When there are no observations, it simply equals m.
Make a prediction
Does choosing a stronger prior always produce a better estimate?
Explore the answer
No. A concentrated prior can help if its information is relevant, or pull estimates away from the truth if it is poorly chosen. Compare plausible priors and state what informed them.
Predict outcomes, not just the parameter
The next-flip probability of heads equals the posterior mean. For several future flips, first imagine drawing a shared p from the posterior, then making independent flips conditional on that p. Integrating over p introduces dependence in the predictive distribution.
For a uniform prior and no observations, the chance of at least one head in two future flips is 2/3. Substituting the mean p = 1/2 into a binomial formula would instead give 3/4. The predictive table averages over uncertainty and reports both the prior and posterior answers.
For posterior shapes a and b, the probability of no heads in m future flips is the product of (b + i)/(a + b + i), for i = 0 through m − 1. Subtracting from one gives the table’s event probability.
When this model stops fitting
The simple update assumes a shared probability and independent Bernoulli observations conditional on it. Changing conditions, dependence between flips, or selectively recorded outcomes require a different likelihood. A beta distribution on a continuous proportion is also not automatically a beta-binomial observation model.
Continue with reading distributions for densities and areas, or Bayes’ theorem for the update behind the shortcut.
Reference
Marco Taboga’s beta distribution derivation gives its density, moments, and binomial conjugacy. The interval uses numerical beta quantiles; the future-flip calculation integrates the Bernoulli likelihood over that beta density.
Compare conjugate models
Conjugate priors compares beta updating with gamma updating for count rates, then shows why posterior predictions differ from plugging in a posterior mean.
For three category probabilities constrained to sum to one, continue with the Dirichlet distribution and its full predictive count table.