paradox #5
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

St. Petersburg paradox

You will learn: Compare an infinite expected payoff with finite samples and a capped game.

Start with: Geometric distribution · Expectation and variance

Flip a fair coin until it comes up tails. You win 2k2^k coins, where kk is the number of heads before the first tails. How much would you pay to play?

The probability of each outcome: P(K=k)=12k+1P(K=k) = \tfrac{1}{2^{k+1}}. Multiply payoff by probability, sum over every outcome:

E[X]=∑k=0∞2k⋅12k+1=∑k=0∞12=∞E[X] = \sum_{k=0}^{\infty} 2^k \cdot \tfrac{1}{2^{k+1}} = \sum_{k=0}^{\infty} \tfrac{1}{2} = \infty

The expected payoff is infinite. Yet almost no one would pay more than $20 or $30 to play. This is the St. Petersburg paradox, posed by Nicolas Bernoulli to Pierre Rémond de Montmort in 1713 and named for the journal in which his cousin Daniel published the famous resolution from St. Petersburg in 1738.

Play it yourself

Click through a few rounds. Most pay $1 or $2; every so often you’ll hit a $32, $64, or higher. Those rare jackpots explain the gap between typical payoffs and the mathematical expectation.

Press Play a round to start. The pot starts at $1 and doubles on every heads.
pot: $1
— no flips yet —
Rounds appear immediately, without an animation delay. Reset replays the same seeded sequence. Most rounds pay $1, $2, or $4. Every so often a streak of heads sends the pot to $32, $64, or higher. Those rare jackpots are exactly what makes the expected value diverge — the bigger they get, the rarer they get, but each tier contributes the same $0.50 to the mean.

What it looks like at scale

Run thousands of rounds at a chosen ticket price. The cumulative net climbs steadily downward — until a single jackpot lifts the curve, or doesn’t. Each tier of payoff contributes the same 12\tfrac{1}{2} to the mean, so doubling the games you simulate increases the running mean by roughly a constant. The empirical mean drifts up logarithmically and never settles.

Buying 1000 tickets at $10: net −$3099 • biggest win $1024 • 69 of 1000 games beat the ticket price
cumulative profit / loss ($) by ticket #−3000−2000−100002004006008001000ticket #cumulative profit / loss ($)
Payoff frequency — theoretical (filled) vs your run (outlined)
frequency by log₂(payoff) = k (win 2ᵏ coins) 00.10.20.30.40.50.6012345678910log₂(payoff) = k (win 2ᵏ coins)frequency
Running mean payoff — never settles
running mean payoff by game number05101520252004006008001000game numberrunning mean payoff
Sample mean after 1000 games
$6.9
Theoretical E[X]
∞
Top: the running net of payoff − ticket price. Try $5, $10, $25 — almost every price looks "obviously losing" until a single huge jackpot lifts the curve. Middle: empirical payoff frequencies match the theoretical 2⁻⁽ᵏ⁺¹⁾ shape closely. Bottom: the mean drifts up logarithmically with games and never converges.

Why the sum diverges

Each tier of payoff — $1, $2, $4, $8, … — contributes exactly 12\tfrac{1}{2} to the expected value. Doubling the prize compensates for halving the probability, infinitely many times.

This is the signature of a fat-tailed distribution. Rare events grow large enough, fast enough, that they dominate the average. The St. Petersburg payoff has no finite variance and no finite mean. A single catastrophic payoff can dwarf the cumulative result of every game played before it — which is what your running-mean curve keeps demonstrating.

In the empirical world, this regime is everywhere. Insurance losses, financial crashes, viral epidemics, war casualties, book sales, city sizes — all live in fat-tailed territory. Nassim Taleb popularized the term black swan for the rare extreme events that drive long-run aggregates in such distributions. St. Petersburg is the cleanest mathematical example.

Bernoulli’s resolution: utility, not cash

Daniel Bernoulli’s 1738 fix: people don’t maximize expected dollars, they maximize expected utility. If utility grows as log⁡W\log W, then doubling wealth feels equally good no matter where you start — your second million matters far less than your first hundred dollars.

Apply that to St. Petersburg. The payoffs grow exponentially, but their utility grows only linearly in kk, which the halving probabilities cancel. The expected utility, E[log⁡(W+X)]E[\log(W + X)], is finite when final wealth stays positive. To evaluate an entry price c at starting wealth w, compare E[log(w−c+X)] with log(w); the illustrative expectation of log(X) alone does not determine willingness to pay. A price must respect the domain of the utility function.

This was the birth of expected utility theory and modern decision theory under uncertainty. It also reframes risk in a way that connects to the rest of probability: insurance, hedging, and the Kelly criterion (which sizes bets to maximize expected log-wealth) all rest on the same concavity that resolves the paradox.

The casino doesn’t have infinite money

Bernoulli’s argument was philosophical. There’s a more mundane resolution: no real casino can pay an infinite prize. Cap the payoff at the casino’s bankroll 2N2^N and the expected value collapses to something modest:

E ⁣[min⁡(2k, 2N)]=N2+1E\!\left[\min(2^k,\, 2^N)\right] = \tfrac{N}{2} + 1
Casino bankroll $1.0M (220) → fair ticket $11.00
fair price ($) by log₂(bankroll cap)051015202530102030405060log₂(bankroll cap)fair price ($)
jump to
With any finite bankroll the expected payoff is just N/2 + 1 dollars, where the casino can pay at most 2N. At a cap near one quadrillion dollars, the risk-neutral expected payout is about $26. This calculation depends on the payout cap; it does not establish a real casino's willingness to offer the game or a player's willingness to pay.

A casino with a million-dollar bankroll can offer the game fairly for about $11 a ticket. Even one with the entire world economy as backing only gets to around $26. The “infinity” lives entirely in the tail beyond any plausible bankroll — exactly where intuition has been refusing to look the whole time.

Karl Menger’s twist

Menger’s construction shows a limitation of unbounded utility: for a given increasing, unbounded utility function, one can choose payoffs that grow quickly enough to make expected utility diverge again. The payoff sequence depends on the chosen utility. A fixed doubly exponential sequence does not outrun every concave utility function. Bounded utility avoids this particular divergence; it is a modeling choice, not a conclusion established about all real preferences.

Gabriel Cramér had already anticipated part of this in a 1728 letter to Nicolas Bernoulli, proposing both a square-root utility and a payoff cap as resolutions — predating Daniel’s published account by a decade. The paradox has been a stress-test for decision theory ever since.

Practical implications

  • Heavy-tailed payoffs demand more than the sample mean. Use medians, trimmed means, or full distribution summaries when the variance might not exist.
  • Insurance is the social institution built on concave utility — pooling fat tails so each participant pays a small premium against rare ruin.
  • The Kelly criterion for sizing risky bets maximizes expected log-wealth, the practical descendant of Bernoulli’s utility argument.
  • Whenever someone quotes an expected value for a payoff with unbounded upside — lottery jackpots, venture investments, viral content — pay attention to the tail before paying for the ticket.
Reset all settings