distribution #23
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Dirichlet distribution

You will learn: Update category probabilities jointly and integrate uncertainty into future counts.

Start with: Beta distribution · Joint and marginal distributions

A coin has one unknown probability. What if each observation belongs to one of three categories? Their probabilities must be nonnegative and sum to one, so learning one affects the others. The Dirichlet extends the beta distribution to such probability vectors. This lesson builds on Bayes’ theorem and joint distributions.

Separate outcomes, counts, and probabilities

A categorical observation is one label, such as A, B, or C, with fixed probabilities θA, θB, θC. For n conditionally independent observations sharing that vector, the multinomial distribution describes their counts:

P(CA=cA,CB=cB,CC=cC∣θ)=n!cA!cB!cC!θAcAθBcBθCcC,cA+cB+cC=n.P(C_A=c_A,C_B=c_B,C_C=c_C\mid\theta)=\frac{n!}{c_A!c_B!c_C!}\theta_A^{c_A}\theta_B^{c_B}\theta_C^{c_C},\quad c_A+c_B+c_C=n.

The factorial coefficient counts label sequences with those totals. A Dirichlet distribution instead describes uncertainty about θ itself. One dot below is a possible probability vector, not a category label or a count vector. The triangle’s corners assign all probability to A, B, or C; its center assigns one third to each.

Dirichlet probability vectors inside a three-category simplexEach blue point is a vector of three nonnegative probabilities adding to one. The corners assign all probability to one category. The larger outlined gold point is the population mean.A = 1B = 1C = 1

Each blue point is one possible probability vector, not one observed category. The larger outlined gold point marks the posterior mean. The triangle contains all 400 simulated vectors. Its surface represents two free coordinates; probabilities sum to one.

CategoryPrior shapeObserved countPosterior shapePrior meanPosterior mean
A1.0045.000.33330.5000
B1.0023.000.33330.3000
C1.0012.000.33330.2000

The common multiplier holds the prior mean fixed and changes its concentration. Posterior means can change because prior strength changes relative to the same observed counts. No observations means posterior equals prior.

One probability has a beta marginal

For the displayed posterior, θA has Beta(5.00, 5.00). P(θA ≤ 0.50) = 0.5000000; simulated fraction 220/400. This probability about θA differs from the probability that the next observation is A.

Displayed distributionExact valueSimulated mean
E[θA]0.5000000.496829
E[θB]0.3000000.290288
E[θC]0.2000000.212882
Total concentration10.00—
Inspect the covariance matrix and first ten vectors

Covariances concern the probability vector, not future count covariances. Off-diagonal values are negative and every row sums to zero because θA+θB+θC is fixed.

CovarianceθAθBθC
θA0.0227273-0.0136364-0.00909091
θB-0.01363640.0190909-0.00545455
θC-0.00909091-0.005454550.0145455
VectorθAθBθCSum
10.477620.194100.328291.00000
20.368410.328500.303081.00000
30.453290.380950.165761.00000
40.360470.291430.348101.00000
50.562530.225690.211791.00000
60.637640.265350.097011.00000
70.343420.192900.463681.00000
80.692210.139700.168101.00000
90.444070.298830.257101.00000
100.510250.405730.084021.00000

Predict future counts by averaging over the posterior

This table always uses the posterior after the observed counts, regardless of the scatterplot view. One shared probability vector governs all future observations; they are conditionally independent given that vector. Integrated probabilities include uncertainty in it; plug-in probabilities fix it at its posterior mean.

Future counts (A, B, C)Integrated probabilityPlug-in probability
(0, 0, 2)0.05454550.0400000
(0, 1, 1)0.1090910.120000
(0, 2, 0)0.1090910.0900000
(1, 0, 1)0.1818180.200000
(1, 1, 0)0.2727270.300000
(2, 0, 0)0.2727270.250000
Sum over all 6 count outcomes1.000000001.00000000
Synthetic category counts and seeded Dirichlet draws. Sampling more probability vectors improves only simulation precision; it does not add observed category data.

Shape, concentration, and a shared constraint

Write θ∼Dirichlet(αA,αB,αC), with all shapes positive and total T=αA+αB+αC. In coordinates θA,θB, with θC=1−θA−θB, its density is:

f(θA,θB)=Γ(T)Γ(αA)Γ(αB)Γ(αC)∏j∈{A,B,C}θjαj−1.f(\theta_A,\theta_B)=\frac{\Gamma(T)}{\Gamma(\alpha_A)\Gamma(\alpha_B)\Gamma(\alpha_C)}\prod_{j\in\{A,B,C\}}\theta_j^{\alpha_j-1}.

The coordinate domain is θA≥0, θB≥0, θA+θB≤1. This is a two-dimensional density, not a density over ordinary three-dimensional volume. At shapes below one, densities can diverge near boundaries while total probability remains finite. The point experiment samples the distribution without treating a tall density as point probability.

Integrating the density after multiplying by θj shifts its j-th exponent by one. Taking the ratio of normalization constants and using Γ(a+1)=aΓ(a) yields E[θj]=αj/T. Shifting twice similarly gives E[θj²]=αj(αj+1)/[T(T+1)] and E[θiθj]=αiαj/[T(T+1)] for distinct indices. Subtract products of means to obtain:

Var⁡(θj)=mj(1−mj)T+1,Cov⁡(θi,θj)=−mimjT+1,mj=αjT.\operatorname{Var}(\theta_j)=\frac{m_j(1-m_j)}{T+1},\qquad \operatorname{Cov}(\theta_i,\theta_j)=-\frac{m_i m_j}{T+1},\quad m_j=\frac{\alpha_j}{T}.

A common multiplier preserves the prior mean and changes concentration T. It does not preserve posterior means after fixed observations because it changes the relative weight of prior information. The coordinates are not independent beta draws: their sum is exactly one, so every covariance-matrix row sums to zero.

Update the synthetic counts

Multiply the prior density by the multinomial likelihood. Each exponent αj−1 gains cj, giving posterior shapes αj+cj. At the default prior (1,1,1), observed counts (4,2,1) produce Dirichlet(5,3,2). Its mean vector is (0.5,0.3,0.2) and total concentration is 10.

The exact posterior variances are 0.25/11, 0.21/11, and 0.16/11. Cov(θA,θB)=−0.15/11. In the A row, 0.25/11−0.15/11−0.10/11=0, as the constraint requires. With no observed counts, the posterior equals the prior. The default counts are a teaching example, not measured category frequencies.

Make a prediction

Does drawing 2,000 probability vectors instead of 400 make the posterior more concentrated?

Explore the answer

No. Those draws approximate the same distribution. Only changing the prior or observed counts changes its shape. More simulated vectors can make the plotted shape and simulated summaries less noisy, but they do not supply more observations of A, B, or C.

Predict while retaining parameter uncertainty

The next observation is A with probability E[θA|data]=0.5. But two future observations both being A has probability E[θA²|data]=5·6/(10·11)=3/11, about 0.272727. Fixing θA at its mean would give 0.25. Their difference is the posterior variance of θA.

All future observations here share the same uncertain θ. They are independent conditional on it; averaging over it induces predictive dependence. For a future count vector k totaling m, integrating the multinomial likelihood gives:

P(K=k∣data)=m!∏jkj!∏j(aj)kj(∑jaj)m,(a)r=a(a+1)⋯(a+r−1),P(K=k\mid\text{data})=\frac{m!}{\prod_j k_j!}\frac{\prod_j (a_j)_{k_j}}{(\sum_j a_j)_m},\qquad (a)_r=a(a+1)\cdots(a+r-1),

where a denotes posterior shapes and (a)₀=1. The table includes every count vector for the selected future total, so both its integrated and plug-in columns sum to one. With one future observation the columns agree; joint predictions for several observations can differ.

Each coordinate’s marginal is beta: θA∼Beta(aA,aB+aC). Therefore a probability such as P(θA≤0.5|data) concerns the uncertain parameter; it is not the chance that the next label is A. Conjugate priors applies the same integration principle to beta/binomial and gamma/Poisson predictions.

Make a prediction

If the prior is uniform on the triangle, is the probability vector fixed at (1/3,1/3,1/3)?

Explore the answer

No. Dirichlet(1,1,1) spreads probability over the triangle; its mean is at the center. For two future observations, it predicts both A with probability 1/6. Fixing all probabilities at one third instead gives 1/9. A distribution over vectors differs from its mean vector.

Compare the resulting category probabilities with cross-entropy: a prediction can represent uncertainty and still receive a precise expected log-loss.

Sources

StatLect: the Dirichlet distribution develops its density, moments, and beta marginals. Stanford Stats 366: modeling mixtures connects normalized independent gamma draws, multinomial observations, and conjugate updating. The demo uses normalized gamma draws for the point cloud and exact finite count probabilities for prediction; the latter are independently checked by enumerating sequential predictive outcomes.

Return to learning paths and foundations.

Reset all settings