Published
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
Pólya’s urn
You will learn: Track reinforcement while distinguishing a martingale mean from a random limiting fraction.
Start with: Conditioning and independence
Begin with one red ball and one ball of another color. Draw uniformly, return the selected ball, and add one extra ball of that same color. A red first draw changes the next-red probability from one half to two thirds. Another red changes it to three quarters. Past outcomes alter future probabilities.
This positive reinforcement does not force the urn back toward equal proportions. Yet before observing any draws, its expected red fraction stays one half. The experiment makes those two statements visible together: individual trajectories can settle near different fractions while their population mean remains fixed.
Follow a history and its evolving prediction
Solid path 1: red-draw count 1. The urn contains 2 red and 30 other-color balls. Next-red probability 0.0625000. Dashed paths are eleven other realizations, not confidence limits.
| Population quantity | Exact value |
|---|---|
| Expected red draws | 15.0000000 |
| Variance of red draws | 80.0000000 |
| Expected red fraction in urn | 0.5000000 |
The limiting urn fraction has Beta(1, 1) distribution. The bars above describe the finite count K, not that continuous limit.
Inspect the selected path
| Completed draws | Red draws | Next-red probability |
|---|---|---|
| 0 | 0 | 0.5000000 |
| 1 | 1 | 0.6666667 |
| 2 | 1 | 0.5000000 |
| 3 | 1 | 0.4000000 |
| 4 | 1 | 0.3333333 |
| 5 | 1 | 0.2857143 |
| 6 | 1 | 0.2500000 |
| 7 | 1 | 0.2222222 |
| 8 | 1 | 0.2000000 |
| 9 | 1 | 0.1818182 |
| 10 | 1 | 0.1666667 |
| 11 | 1 | 0.1538462 |
| 12 | 1 | 0.1428571 |
| 13 | 1 | 0.1333333 |
| 14 | 1 | 0.1250000 |
| 15 | 1 | 0.1176471 |
| 16 | 1 | 0.1111111 |
| 17 | 1 | 0.1052632 |
| 18 | 1 | 0.1000000 |
| 19 | 1 | 0.0952381 |
| 20 | 1 | 0.0909091 |
| 21 | 1 | 0.0869565 |
| 22 | 1 | 0.0833333 |
| 23 | 1 | 0.0800000 |
| 24 | 1 | 0.0769231 |
| 25 | 1 | 0.0740741 |
| 26 | 1 | 0.0714286 |
| 27 | 1 | 0.0689655 |
| 28 | 1 | 0.0666667 |
| 29 | 1 | 0.0645161 |
| 30 | 1 | 0.0625000 |
The controls specify initial counts a and b, extra balls c added per draw, and the number n of completed draws. The selected ball is always returned before reinforcement. With c = 0, sampling leaves the urn unchanged. The displayed paths use repeatable random seeds, and increasing n continues their previous histories.
The upper plot shows the fraction physically in the urn, not the fraction of observed draws that were red. After k red draws among t total draws, the urn contains a + ck red balls and b + c(t − k) other-color balls. Its next-red probability is (a + ck)/(a + b + ct). The selected path’s table exposes both its red-draw count and that probability at every time.
Calculate a small experiment
Set a = b = c = 1 and n = 2. The sequences RR and OO each have probability (1/2)(2/3) = 1/3. The sequences RO and OR each have probability (1/2)(1/3) = 1/6. Thus the number of red draws is zero, one, or two with equal probabilities 1/3.
Without reinforcement, those count probabilities would be 1/4, 1/2, 1/4. The mean stays one red draw, but the variance changes from one half to two thirds. Reinforcement moves probability toward the extreme counts. For the default thirty draws with unit initial counts and unit reinforcement, all thirty-one possible red counts have probability 1/31; the count variance is 80, compared with 7.5 for thirty independent fair draws.
The bars are calculated by propagating count probabilities. If fₜ(k) is the chance of k red draws after t draws, send fraction (a + ck)/(a + b + ct) of that mass to count k + 1 and the remainder to k. Start with all mass at count zero. This finite recurrence also handles zero draws, one-color urns, and no reinforcement without approximating a continuous distribution.
Why the mean does not drift
Suppose the current urn has R red balls out of T, so its red fraction is z = R/T. After the next draw the fraction is (R + c)/(T + c) with probability z, and R/(T + c) otherwise. Its conditional expectation is
z(R + c)/(T + c) + (1 − z)R/(T + c) = (R + cz)/(T + c) = z.
So the urn fraction is a martingale: given its current history, its expected next value equals its current value. That says nothing about every sample staying near the initial value. A first red draw still raises the conditional forecast. Averaging over both possible first draws restores the original expectation.
Make a prediction
After five red draws in a row, should the other color be more likely next to restore balance?
Explore the answer
No. With a = b = c = 1 the urn contains six red balls and one other ball, so the next-red probability is 6/7. The process has reinforcement, not a compensating force.
A random limit
With positive a, b, and c, the urn fraction converges to a random value with Beta(a/c, b/c) distribution. Its limiting value varies between experiments. In the unit-count case that limit is uniform on the interval from zero to one. With no reinforcement the urn fraction is constant instead; with an absent color it stays at its deterministic endpoint. A beta law with zero shape parameters is not used for those cases.
Kyle Siegrist’s treatment develops the exchangeability and limiting-distribution results. The finite calculations above can be checked directly in the experiment. Continue to Dirichlet probabilities for the connection between reinforced counts and predictions over several categories, or martingales for the meaning of a conditionally unchanged expectation.