paradox #4
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Gambler’s fallacy

You will learn: Distinguish independence from the belief that a streak must be compensated.

Start with: Bernoulli distribution · Conditioning and independence

Five heads in a row. Surely tails is overdue — the coin needs to “balance out.” This intuition is wrong, and it is wrong in a precise way: a fair coin has no memory.

Each flip is independent. The probability of tails on the next flip is exactly 12\tfrac{1}{2}, regardless of what came before. The coin cannot “know” it owes you tails.

running proportion of heads by flip number00.20.40.60.8150100150200250300flip numberrunning proportion of heads

Condition on a specified initial streak

ModelChance of these initial headsNext head after that history
Independent coin, current p0.03125000.500000
Independent fair coin0.03125000.500000
20 shuffled tokens: 10 H, 10 T, no replacement0.01625390.333333

The token history removes heads from a finite pool. Its next-head probability is (10 − h)/(20 − h), and its initial-streak probability multiplies successive changing fractions. The coin history does not change a known p. These are separate probability models; the running curves above use the independent coin.

Dilution without repayment

After 5 initial heads, the expected heads in the next 100 flips are 50.0000. The expected final proportion is (h + mp)/(h + m) = 0.523810. The expected excess count above p(h+m) stays h(1−p) = 2.50000, while its proportion shrinks. A realized run fluctuates around this conditional prediction.

Running proportion of heads across 15 independent sequences. Dashed line: true p = 0.50. The first sequence is highlighted in blue.

What to notice

  • Convergence with probability one. Under the independent, fixed-p model, the running proportion converges toward the true probability p, not because the coin corrects itself, but simply because new flips dilute old history. With 1000 flips, one outlier streak is a tiny fraction of the total.
  • Short runs diverge wildly. In the first 20 flips, the running proportion swings dramatically. This is where gamblers are tempted to “see patterns.”
  • Biased coins (p ≠ 0.5). Drag p away from 0.5. The running proportion still converges with probability one — now to the true p, not 0.5. The gambler’s error remains the same.

Law of Large Numbers vs the fallacy

The Law of Large Numbers says the running average converges to the true mean. It does not say past outcomes must be corrected. Long-run balance happens through dilution, not compensation.

The gambler’s fallacy incorrectly adds a compensation mechanism to an independent fixed-probability model. Real financial or sporting sequences need not satisfy that model; independence must be justified rather than assumed from this coin example.

Before a streak and after observing it

For an independent fair coin, a specified initial sequence of five heads has probability 1/32. Conditional on those five heads already occurring, a sixth head still has probability 1/2. The probability of six initial heads, computed before any toss, is 1/64. Those are three distinct events and information sets.

The table concerns an initial block starting at a specified position. The probability of finding a streak anywhere in a long sequence is different because many overlapping starting positions are possible; multiplying one streak probability by the number of positions generally double-counts outcomes.

Make a prediction

After five heads, is HHHHHT a more likely complete six-flip string than HHHHHH was before tossing?

Explore the answer

No. Each specified fair six-flip string has probability 1/64. After observing the first five heads, the two possible next outcomes each have conditional probability 1/2.

Change the model: no replacement

Shuffle 10 H tokens and 10 T tokens, then draw without replacement. After five heads, five H and ten T remain, so the next-head probability is 5/15 = 1/3. After all ten H tokens have appeared, another H is impossible. The change follows from removing objects from a known finite population, not from a coin remembering its history.

The original probability of five initial heads in this urn is (10/20)(9/19)(8/18)(7/17)(6/16)≈0.016254(10/20)(9/19)(8/18)(7/17)(6/16)\approx0.016254, distinct from the fair independent coin’s 0.03125. Both the prior streak probability and the conditional next draw depend on the protocol.

Work the dilution calculation

Given five initial heads and 100 future fair independent flips, expected future heads are 50. The expected final proportion is 55/105 ≈ 0.523810. The expected excess heads above half the total remains 2.5; dividing that fixed expected excess by 105 makes the proportional discrepancy small. No compensating bias toward tails has been introduced.

Make a prediction

If the coin's bias is unknown, can a run of heads change a rational next-head prediction?

Explore the answer

Yes, through learning about the bias, rather than through repayment. A posterior predictive probability can rise after heads under a specified prior. That is a different model from this experiment’s known, fixed p; see the beta lesson.

References

Random Services: independence formalizes what conditioning leaves unchanged. Its urn sampling chapter distinguishes replacement protocols. Compare geometric waiting and beta learning for two different ways to reason about the next observation.

Reset all settings