puzzle #17
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Pólya’s urn

You will learn: Track reinforcement while distinguishing a martingale mean from a random limiting fraction.

Start with: Conditioning and independence

Begin with one red ball and one ball of another color. Draw uniformly, return the selected ball, and add one extra ball of that same color. A red first draw changes the next-red probability from one half to two thirds. Another red changes it to three quarters. Past outcomes alter future probabilities.

This positive reinforcement does not force the urn back toward equal proportions. Yet before observing any draws, its expected red fraction stays one half. The experiment makes those two statements visible together: individual trajectories can settle near different fractions while their population mean remains fixed.

Follow a history and its evolving prediction

Twelve seeded paths of the urn's red fraction00.20.40.60.81051015202530draws completedred fraction in urn

Solid path 1: red-draw count 1. The urn contains 2 red and 30 other-color balls. Next-red probability 0.0625000. Dashed paths are eleven other realizations, not confidence limits.

Exact distribution of the number of red draws00.010.020.03051015202530red draws, Kprobability
Population quantityExact value
Expected red draws15.0000000
Variance of red draws80.0000000
Expected red fraction in urn0.5000000

The limiting urn fraction has Beta(1, 1) distribution. The bars above describe the finite count K, not that continuous limit.

Inspect the selected path
Completed drawsRed drawsNext-red probability
000.5000000
110.6666667
210.5000000
310.4000000
410.3333333
510.2857143
610.2500000
710.2222222
810.2000000
910.1818182
1010.1666667
1110.1538462
1210.1428571
1310.1333333
1410.1250000
1510.1176471
1610.1111111
1710.1052632
1810.1000000
1910.0952381
2010.0909091
2110.0869565
2210.0833333
2310.0800000
2410.0769231
2510.0740741
2610.0714286
2710.0689655
2810.0666667
2910.0645161
3010.0625000
Return the selected ball before adding c extra balls of its color. Increasing the draw count retains each path's history. A zero-total initial urn is corrected visibly because no draw could be defined.

The controls specify initial counts a and b, extra balls c added per draw, and the number n of completed draws. The selected ball is always returned before reinforcement. With c = 0, sampling leaves the urn unchanged. The displayed paths use repeatable random seeds, and increasing n continues their previous histories.

The upper plot shows the fraction physically in the urn, not the fraction of observed draws that were red. After k red draws among t total draws, the urn contains a + ck red balls and b + c(t − k) other-color balls. Its next-red probability is (a + ck)/(a + b + ct). The selected path’s table exposes both its red-draw count and that probability at every time.

Calculate a small experiment

Set a = b = c = 1 and n = 2. The sequences RR and OO each have probability (1/2)(2/3) = 1/3. The sequences RO and OR each have probability (1/2)(1/3) = 1/6. Thus the number of red draws is zero, one, or two with equal probabilities 1/3.

Without reinforcement, those count probabilities would be 1/4, 1/2, 1/4. The mean stays one red draw, but the variance changes from one half to two thirds. Reinforcement moves probability toward the extreme counts. For the default thirty draws with unit initial counts and unit reinforcement, all thirty-one possible red counts have probability 1/31; the count variance is 80, compared with 7.5 for thirty independent fair draws.

The bars are calculated by propagating count probabilities. If fₜ(k) is the chance of k red draws after t draws, send fraction (a + ck)/(a + b + ct) of that mass to count k + 1 and the remainder to k. Start with all mass at count zero. This finite recurrence also handles zero draws, one-color urns, and no reinforcement without approximating a continuous distribution.

Why the mean does not drift

Suppose the current urn has R red balls out of T, so its red fraction is z = R/T. After the next draw the fraction is (R + c)/(T + c) with probability z, and R/(T + c) otherwise. Its conditional expectation is

z(R + c)/(T + c) + (1 − z)R/(T + c) = (R + cz)/(T + c) = z.

So the urn fraction is a martingale: given its current history, its expected next value equals its current value. That says nothing about every sample staying near the initial value. A first red draw still raises the conditional forecast. Averaging over both possible first draws restores the original expectation.

Make a prediction

After five red draws in a row, should the other color be more likely next to restore balance?

Explore the answer

No. With a = b = c = 1 the urn contains six red balls and one other ball, so the next-red probability is 6/7. The process has reinforcement, not a compensating force.

A random limit

With positive a, b, and c, the urn fraction converges to a random value with Beta(a/c, b/c) distribution. Its limiting value varies between experiments. In the unit-count case that limit is uniform on the interval from zero to one. With no reinforcement the urn fraction is constant instead; with an absent color it stays at its deterministic endpoint. A beta law with zero shape parameters is not used for those cases.

Kyle Siegrist’s treatment develops the exchangeability and limiting-distribution results. The finite calculations above can be checked directly in the experiment. Continue to Dirichlet probabilities for the connection between reinforced counts and predictions over several categories, or martingales for the meaning of a conditionally unchanged expectation.

Reset all settings