law #3
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Bayes’ theorem

You will learn: Update conditional probabilities and odds while checking how evidence was generated.

Start with: Conditioning and independence

Bayes’ theorem tells you how to update a belief when new evidence arrives. You start with a prior probability P(H)P(H) that some hypothesis H is true. You observe evidence E. The posterior is:

P(H∣E)=P(E∣H) P(H)P(E)P(H \mid E) = \frac{P(E \mid H)\,P(H)}{P(E)}

The denominator is the total probability of seeing the evidence at all:

P(E)=P(E∣H)P(H)+P(E∣¬H)P(¬H)P(E) = P(E \mid H)P(H) + P(E \mid \neg H)P(\neg H)
The left region has probability 0.3; the right has probability 0.7. Their shaded evidence areas are 0.24 and 0.13999999999999999. Exact expected counts are provided in the table below.P(H) = 0.30P(¬H) = 0.70P(H∩E)P(¬H∩E)
Posterior P(H|E)
63.2%
was 30% prior
P(E) — evidence probability
38.0%
P(¬H|E)
36.8%

Read the same calculation as expected counts

Expected counts in a hypothetical population of 10,000
GroupEvidence ENo evidence ETotal
H true2400.0600.03000.0
H false1400.05600.07000.0

Among the 3800.0 expected E cases, 2400.0 have H. These are model-based expected counts, so fractional values are allowed.

Update with several observations

Assume the same H stays true or false, and observations are independent conditional on H, with the same likelihoods each time. Re-reading the same evidence does not count as a new observation.

Posterior after this evidence: 63.1579%. Each positive multiplies the odds by 4.0000; each negative by 0.2500. With no observations, the posterior equals the prior.

Left region: H; right region: H false. The colored upper areas contain evidence E. Narrow-region labels are omitted; the table retains every value. Posterior = P(H∣E)=P(E∣H) P(H)P(E)P(H|E) = \frac{P(E|H)\,P(H)}{P(E)}.

Reading the area diagram

The rectangle represents the entire sample space. Width encodes the prior split: left = P(H), right = P(¬H). Height encodes the likelihoods: shaded top fraction of each column = probability of evidence E given that column’s hypothesis.

The posterior P(H|E) is the proportion of the total shaded area that falls in the left (H) column. Drag the sliders and watch the ratio shift.

What to notice

  • Strong evidence isn’t enough. Set P(H) very low (rare hypothesis) and P(E|¬H) only a little lower than P(E|H). Even 95% sensitivity gives a low posterior when prevalence is 1%. This is base rate neglect.
  • False positives matter when the hypothesis is rare. Reducing P(E|¬H) can substantially raise the posterior because the large ¬H population may contribute many of the positive results. The size of the change depends on all three input probabilities.
  • The prior can be overwhelmed. With strong enough evidence (high P(E|H), low P(E|¬H)), the posterior can approach 1 for a fixed positive prior as the likelihood ratio grows. A zero prior stays zero when the evidence has positive marginal probability. Convergence also depends on the model and how the evidence is generated.

Make the denominator visible

At the default prior 0.3, likelihood P(E|H)=0.8 and alternative likelihood P(E|not H)=0.2, a hypothetical 10,000 cases contain 3,000 with H and 7,000 without it. Expected E counts are 2,400 and 1,400. Conditioning on E leaves 3,800 cases, of which 2,400 have H: 63.1579%. The likelihood 80% answers a different conditional question from this posterior.

Keep the likelihoods fixed and reduce the prior to 0.01. Expected E counts become 80 and 1,980, giving posterior 3.8835%. These synthetic numbers illustrate sensitivity to the prior; they are not estimates for a real screening system.

Sequential evidence needs a dependence model

Posterior odds equal prior odds times the likelihood ratio. With the default likelihoods, each positive observation multiplies odds by four and each negative by one quarter. Two positives give odds (3/7)×16=48/7 and posterior 87.2727%. One positive and one negative cancel their likelihood ratios, returning the posterior to 30%.

This multiplication requires observations that are independent conditional on the same H, with the specified likelihoods. Two copies of the same result are perfectly dependent, so counting both as independent would exaggerate the evidence. The order of observations does not matter in this particular fixed-likelihood model; other models can require additional state.

Make a prediction

If P(E|H) equals P(E|not H), how should observing E change the prior?

Explore the answer

It should not. The likelihood ratio is one, so the posterior odds equal the prior odds. Frequent evidence is not necessarily informative evidence.

Reference

ProbabilityCourse: Bayes’ rule derives the reversed conditional probability and the total-probability denominator. The repeated-evidence control applies that identity under the additional conditional-independence assumption stated above.

Reset all settings