Published
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
Dirichlet distribution
You will learn: Update category probabilities jointly and integrate uncertainty into future counts.
Start with: Beta distribution · Joint and marginal distributions
A coin has one unknown probability. What if each observation belongs to one of three categories? Their probabilities must be nonnegative and sum to one, so learning one affects the others. The Dirichlet extends the beta distribution to such probability vectors. This lesson builds on Bayes’ theorem and joint distributions.
Separate outcomes, counts, and probabilities
A categorical observation is one label, such as A, B, or C, with fixed probabilities θA, θB, θC. For n conditionally independent observations sharing that vector, the multinomial distribution describes their counts:
The factorial coefficient counts label sequences with those totals. A Dirichlet distribution instead describes uncertainty about θ itself. One dot below is a possible probability vector, not a category label or a count vector. The triangle’s corners assign all probability to A, B, or C; its center assigns one third to each.
Each blue point is one possible probability vector, not one observed category. The larger outlined gold point marks the posterior mean. The triangle contains all 400 simulated vectors. Its surface represents two free coordinates; probabilities sum to one.
| Category | Prior shape | Observed count | Posterior shape | Prior mean | Posterior mean |
|---|---|---|---|---|---|
| A | 1.00 | 4 | 5.00 | 0.3333 | 0.5000 |
| B | 1.00 | 2 | 3.00 | 0.3333 | 0.3000 |
| C | 1.00 | 1 | 2.00 | 0.3333 | 0.2000 |
The common multiplier holds the prior mean fixed and changes its concentration. Posterior means can change because prior strength changes relative to the same observed counts. No observations means posterior equals prior.
One probability has a beta marginal
For the displayed posterior, θA has Beta(5.00, 5.00). P(θA ≤ 0.50) = 0.5000000; simulated fraction 220/400. This probability about θA differs from the probability that the next observation is A.
| Displayed distribution | Exact value | Simulated mean |
|---|---|---|
| E[θA] | 0.500000 | 0.496829 |
| E[θB] | 0.300000 | 0.290288 |
| E[θC] | 0.200000 | 0.212882 |
| Total concentration | 10.00 | — |
Inspect the covariance matrix and first ten vectors
Covariances concern the probability vector, not future count covariances. Off-diagonal values are negative and every row sums to zero because θA+θB+θC is fixed.
| Covariance | θA | θB | θC |
|---|---|---|---|
| θA | 0.0227273 | -0.0136364 | -0.00909091 |
| θB | -0.0136364 | 0.0190909 | -0.00545455 |
| θC | -0.00909091 | -0.00545455 | 0.0145455 |
| Vector | θA | θB | θC | Sum |
|---|---|---|---|---|
| 1 | 0.47762 | 0.19410 | 0.32829 | 1.00000 |
| 2 | 0.36841 | 0.32850 | 0.30308 | 1.00000 |
| 3 | 0.45329 | 0.38095 | 0.16576 | 1.00000 |
| 4 | 0.36047 | 0.29143 | 0.34810 | 1.00000 |
| 5 | 0.56253 | 0.22569 | 0.21179 | 1.00000 |
| 6 | 0.63764 | 0.26535 | 0.09701 | 1.00000 |
| 7 | 0.34342 | 0.19290 | 0.46368 | 1.00000 |
| 8 | 0.69221 | 0.13970 | 0.16810 | 1.00000 |
| 9 | 0.44407 | 0.29883 | 0.25710 | 1.00000 |
| 10 | 0.51025 | 0.40573 | 0.08402 | 1.00000 |
Predict future counts by averaging over the posterior
This table always uses the posterior after the observed counts, regardless of the scatterplot view. One shared probability vector governs all future observations; they are conditionally independent given that vector. Integrated probabilities include uncertainty in it; plug-in probabilities fix it at its posterior mean.
| Future counts (A, B, C) | Integrated probability | Plug-in probability |
|---|---|---|
| (0, 0, 2) | 0.0545455 | 0.0400000 |
| (0, 1, 1) | 0.109091 | 0.120000 |
| (0, 2, 0) | 0.109091 | 0.0900000 |
| (1, 0, 1) | 0.181818 | 0.200000 |
| (1, 1, 0) | 0.272727 | 0.300000 |
| (2, 0, 0) | 0.272727 | 0.250000 |
| Sum over all 6 count outcomes | 1.00000000 | 1.00000000 |
Shape, concentration, and a shared constraint
Write θ∼Dirichlet(αA,αB,αC), with all shapes positive and total T=αA+αB+αC. In coordinates θA,θB, with θC=1−θA−θB, its density is:
The coordinate domain is θA≥0, θB≥0, θA+θB≤1. This is a two-dimensional density, not a density over ordinary three-dimensional volume. At shapes below one, densities can diverge near boundaries while total probability remains finite. The point experiment samples the distribution without treating a tall density as point probability.
Integrating the density after multiplying by θj shifts its j-th exponent by one. Taking the ratio of normalization constants and using Γ(a+1)=aΓ(a) yields E[θj]=αj/T. Shifting twice similarly gives E[θj²]=αj(αj+1)/[T(T+1)] and E[θiθj]=αiαj/[T(T+1)] for distinct indices. Subtract products of means to obtain:
A common multiplier preserves the prior mean and changes concentration T. It does not preserve posterior means after fixed observations because it changes the relative weight of prior information. The coordinates are not independent beta draws: their sum is exactly one, so every covariance-matrix row sums to zero.
Update the synthetic counts
Multiply the prior density by the multinomial likelihood. Each exponent αj−1 gains cj, giving posterior shapes αj+cj. At the default prior (1,1,1), observed counts (4,2,1) produce Dirichlet(5,3,2). Its mean vector is (0.5,0.3,0.2) and total concentration is 10.
The exact posterior variances are 0.25/11, 0.21/11, and 0.16/11. Cov(θA,θB)=−0.15/11. In the A row, 0.25/11−0.15/11−0.10/11=0, as the constraint requires. With no observed counts, the posterior equals the prior. The default counts are a teaching example, not measured category frequencies.
Make a prediction
Does drawing 2,000 probability vectors instead of 400 make the posterior more concentrated?
Explore the answer
No. Those draws approximate the same distribution. Only changing the prior or observed counts changes its shape. More simulated vectors can make the plotted shape and simulated summaries less noisy, but they do not supply more observations of A, B, or C.
Predict while retaining parameter uncertainty
The next observation is A with probability E[θA|data]=0.5. But two future observations both being A has probability E[θA²|data]=5·6/(10·11)=3/11, about 0.272727. Fixing θA at its mean would give 0.25. Their difference is the posterior variance of θA.
All future observations here share the same uncertain θ. They are independent conditional on it; averaging over it induces predictive dependence. For a future count vector k totaling m, integrating the multinomial likelihood gives:
where a denotes posterior shapes and (a)₀=1. The table includes every count vector for the selected future total, so both its integrated and plug-in columns sum to one. With one future observation the columns agree; joint predictions for several observations can differ.
Each coordinate’s marginal is beta: θA∼Beta(aA,aB+aC). Therefore a probability such as P(θA≤0.5|data) concerns the uncertain parameter; it is not the chance that the next label is A. Conjugate priors applies the same integration principle to beta/binomial and gamma/Poisson predictions.
Make a prediction
If the prior is uniform on the triangle, is the probability vector fixed at (1/3,1/3,1/3)?
Explore the answer
No. Dirichlet(1,1,1) spreads probability over the triangle; its mean is at the center. For two future observations, it predicts both A with probability 1/6. Fixing all probabilities at one third instead gives 1/9. A distribution over vectors differs from its mean vector.
Compare the resulting category probabilities with cross-entropy: a prediction can represent uncertainty and still receive a precise expected log-loss.
Sources
StatLect: the Dirichlet distribution develops its density, moments, and beta marginals. Stanford Stats 366: modeling mixtures connects normalized independent gamma draws, multinomial observations, and conjugate updating. The demo uses normalized gamma draws for the point cloud and exact finite count probabilities for prediction; the latter are independently checked by enumerating sequential predictive outcomes.
Return to learning paths and foundations.