Updated
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
Bayes’ theorem
You will learn: Update conditional probabilities and odds while checking how evidence was generated.
Start with: Conditioning and independence
Bayes’ theorem tells you how to update a belief when new evidence arrives. You start with a prior probability that some hypothesis H is true. You observe evidence E. The posterior is:
The denominator is the total probability of seeing the evidence at all:
Read the same calculation as expected counts
| Group | Evidence E | No evidence E | Total |
|---|---|---|---|
| H true | 2400.0 | 600.0 | 3000.0 |
| H false | 1400.0 | 5600.0 | 7000.0 |
Among the 3800.0 expected E cases, 2400.0 have H. These are model-based expected counts, so fractional values are allowed.
Update with several observations
Assume the same H stays true or false, and observations are independent conditional on H, with the same likelihoods each time. Re-reading the same evidence does not count as a new observation.
Posterior after this evidence: 63.1579%. Each positive multiplies the odds by 4.0000; each negative by 0.2500. With no observations, the posterior equals the prior.
Reading the area diagram
The rectangle represents the entire sample space. Width encodes the prior split: left = P(H), right = P(¬H). Height encodes the likelihoods: shaded top fraction of each column = probability of evidence E given that column’s hypothesis.
The posterior P(H|E) is the proportion of the total shaded area that falls in the left (H) column. Drag the sliders and watch the ratio shift.
What to notice
- Strong evidence isn’t enough. Set P(H) very low (rare hypothesis) and P(E|¬H) only a little lower than P(E|H). Even 95% sensitivity gives a low posterior when prevalence is 1%. This is base rate neglect.
- False positives matter when the hypothesis is rare. Reducing P(E|¬H) can substantially raise the posterior because the large ¬H population may contribute many of the positive results. The size of the change depends on all three input probabilities.
- The prior can be overwhelmed. With strong enough evidence (high P(E|H), low P(E|¬H)), the posterior can approach 1 for a fixed positive prior as the likelihood ratio grows. A zero prior stays zero when the evidence has positive marginal probability. Convergence also depends on the model and how the evidence is generated.
Make the denominator visible
At the default prior 0.3, likelihood P(E|H)=0.8 and alternative likelihood P(E|not H)=0.2, a hypothetical 10,000 cases contain 3,000 with H and 7,000 without it. Expected E counts are 2,400 and 1,400. Conditioning on E leaves 3,800 cases, of which 2,400 have H: 63.1579%. The likelihood 80% answers a different conditional question from this posterior.
Keep the likelihoods fixed and reduce the prior to 0.01. Expected E counts become 80 and 1,980, giving posterior 3.8835%. These synthetic numbers illustrate sensitivity to the prior; they are not estimates for a real screening system.
Sequential evidence needs a dependence model
Posterior odds equal prior odds times the likelihood ratio. With the default likelihoods, each positive observation multiplies odds by four and each negative by one quarter. Two positives give odds (3/7)×16=48/7 and posterior 87.2727%. One positive and one negative cancel their likelihood ratios, returning the posterior to 30%.
This multiplication requires observations that are independent conditional on the same H, with the specified likelihoods. Two copies of the same result are perfectly dependent, so counting both as independent would exaggerate the evidence. The order of observations does not matter in this particular fixed-likelihood model; other models can require additional state.
Make a prediction
If P(E|H) equals P(E|not H), how should observing E change the prior?
Explore the answer
It should not. The likelihood ratio is one, so the posterior odds equal the prior odds. Frequent evidence is not necessarily informative evidence.
Reference
ProbabilityCourse: Bayes’ rule derives the reversed conditional probability and the total-probability denominator. The repeated-evidence control applies that identity under the additional conditional-independence assumption stated above.