paradox #9
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Sleeping Beauty

You will learn: Distinguish trial and awakening measures without treating simulated counts as a philosophical resolution.

Start with: Conditioning and independence

A coin determines an awakening schedule. Heads produces one awakening, on Monday. Tails produces two, on Monday and Tuesday. Between awakenings, Beauty’s memory of the previous awakening is erased, and the awakenings provide no other distinguishing information. When awake, what probability should she assign to heads?

For a fair coin, the best-known answers are one half and one third. The disagreement concerns self-locating belief: what does learning “I am awake now” warrant, given that the number of indistinguishable occasions depends on the coin? A simulation can verify a specified sampling protocol. It cannot, by counting alone, establish that this protocol is the correct representation of Beauty’s evidence.

Choose the unit being sampled

AwakeningProbability under selected protocol
Heads / Monday0.33333333
Tails / Monday0.33333333
Tails / Tuesday0.33333333

P(heads): 0.33333333. Conditional on the sampled awakening being Monday: 0.50000000. The conditioning calculation uses the selected sampling protocol.

Heads probability under trial and pooled-awakening sampling00.20.40.60.8100.20.40.60.81coin P(heads)sampled P(heads)

Dashed: uniform trial (also its Monday). Solid: long-run awakening pool. The open circle marks the selected protocol. These are different sampling experiments.

Observed: 47 heads trials, 53 tails trials, 153 awakenings. Empirical heads fraction per trial: 47/100. Per awakening: 47/153. Ratios from a finite pool fluctuate; the exact curve gives the limiting pool proportion.

Inspect first 30 trials

1: T → Monday and Tuesday; 2: T → Monday and Tuesday; 3: T → Monday and Tuesday; 4: T → Monday and Tuesday; 5: H → Monday; 6: H → Monday; 7: T → Monday and Tuesday; 8: H → Monday; 9: T → Monday and Tuesday; 10: H → Monday; 11: T → Monday and Tuesday; 12: H → Monday; 13: H → Monday; 14: H → Monday; 15: H → Monday; 16: H → Monday; 17: T → Monday and Tuesday; 18: T → Monday and Tuesday; 19: T → Monday and Tuesday; 20: H → Monday; 21: T → Monday and Tuesday; 22: H → Monday; 23: H → Monday; 24: T → Monday and Tuesday; 25: T → Monday and Tuesday; 26: H → Monday; 27: T → Monday and Tuesday; 28: T → Monday and Tuesday; 29: H → Monday; 30: T → Monday and Tuesday

Purchase schedule (ticket pays 1 on heads)Expected net per trial
One ticket per trial0.100000
One ticket every awakening-0.100000
Heads produces one Monday awakening; tails produces Monday and Tuesday, with no memory distinguishing them. Sampling protocols and purchase schedules are explicitly imposed here. Their arithmetic does not decide which credence Beauty should adopt in the original philosophical problem.

Let the coin’s heads probability be p. A trial is the complete experiment generated by one coin toss. An awakening is one occasion within a trial. There are three possible labeled occasions: heads Monday, tails Monday, and tails Tuesday. Listing three labels does not by itself make them equally probable.

The first protocol samples uniformly from a large pool of awakenings generated by repeated trials. Heads contributes about p awakenings per trial, while tails contributes about 2(1 − p). The limiting fraction associated with heads is therefore p/(2 − p). At a fair coin, all three labeled occasions receive weight one third.

The second protocol first chooses a trial uniformly, then chooses uniformly among that trial’s awakenings. Heads still has probability p. Its Monday receives all that mass, while tails’ mass is split equally between its two days. At a fair coin the occasion weights are one half, one quarter, and one quarter. The two-stage selection gives a tails awakening half as much weight as the heads awakening within a selected trial.

A third protocol samples just the Monday awakening from a uniform trial. Both coin outcomes contribute exactly one such occasion, so its heads probability is again p. These are three well-defined experiments; their differences are intentional.

Conditioning on Monday still needs the protocol

Under pooled-awakening sampling, conditioning on Monday removes tails Tuesday. Heads Monday and tails Monday then have relative weights p and 1 − p, giving heads probability p. Under trial-then-awakening sampling, their weights were p and (1 − p)/2, giving 2p/(1 + p). At a fair coin those Monday-conditioned probabilities are one half and two thirds.

Thus even the sentence “you learn that it is Monday” must be interpreted within the probability measure already in use. Changing the sampling measure midway through a calculation can produce an apparent contradiction that is simply inconsistent conditioning.

Check the counting experiment

The seeded trial list records one coin outcome per trial and expands tails into two awakenings. With h heads among n trials, the trial fraction is h/n and the pooled-awakening fraction is h/(2n − h). They are empirical ratios, not exact finite-sample expectations of those ratios. As the number of trials grows, they approach p and p/(2 − p), respectively.

Increasing the trial count retains the prior coin sequence. Changing the seed changes that example but leaves the theoretical weights unchanged. With zero trials, the empirical fractions are undefined; no evidence has been sampled. With a deterministic coin, both principal protocols agree because the coin outcome no longer varies.

A bet needs a purchase schedule

Suppose a ticket costs c tokens and pays one token if the trial’s coin is heads. Buying once per trial gives expected net p − c. Buying once at every awakening gives expected net p − c(2 − p) per trial: the heads payout occurs once, but a tails trial requires two purchases. For p = 1/2 and c = 0.4, the expectations are +0.1 and −0.1 tokens. These are different contracts, so different recommendations are unsurprising.

Make a prediction

Does a simulation with twice as many tails awakenings as heads awakenings refute every halfer account?

Explore the answer

No. It verifies the pooled-awakening frequency. A philosophical argument must still justify why that frequency is the probability Beauty should use for the original evidence and question. Likewise, observing half heads among trials does not establish that trials are the appropriate sampling unit for every awakening-level decision.

Elga’s original paper argues for the thirder position. Bostrom’s alternative analysis challenges both standard positions. The experiment makes their underlying sampling issues inspectable while leaving the philosophical dispute explicit. Continue to the inspection paradox to study a related change of sampling weights with an unambiguous observation procedure.

Reset all settings