Updated
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
Base rate neglect
You will learn: Combine prevalence with conditional evidence without reversing the conditioning.
Start with: Conditioning and independence
New to the notation? Start with the connected foundation for the underlying definitions and a worked example.
In this hypothetical example, a test has 99% sensitivity and 99% specificity. You test positive. How likely is it that you’re actually sick?
The instinct says ~99%. Under these assumptions and a prevalence of 1%, the answer is closer to 50%.
Squares represent condition present; circles represent condition absent. Filled marks test positive; open marks test negative. Use the table for precise values in this dense grid.
Exact expected frequencies per 10,000
| Model status | Positive | Negative | Total |
|---|---|---|---|
| Condition present | 99.0000 | 1.0000 | 100.0000 |
| Condition absent | 99.0000 | 9801.0000 | 9900.0000 |
| All people | 198.0000 | 9802.0000 | 10000 |
PPV = true positives / all positives = 99.0000 / 198.0000 = 50.0000%. NPV = true negatives / all negatives = 99.9898%. Expected counts can be fractional; the colored grid alone rounds them to whole people. Neither is an observed clinical dataset.
Keep the test; change prevalence
| Prevalence | Expected true positives | Expected false positives | PPV |
|---|---|---|---|
| 0.1% | 9.900 | 99.900 | 9.016% |
| 1.0% | 99.000 | 99.000 | 50.000% |
| 10.0% | 990.000 | 90.000 | 91.667% |
Why so low?
Because the healthy population is enormous relative to the sick one. Out of 10 000 people:
- 100 are sick. The 99% sensitive test catches 99 of them — those are the true positives.
- 9 900 are healthy. The 99% specific test misidentifies 1% of them as positive — that’s 99 false positives.
Among the 198 people who tested positive, only 99 are actually sick. .
What to try
- Drop prevalence to 0.1%. Watch collapse — you’d need near-perfect specificity for a positive result to mean much.
- Lift prevalence to 10%. Now the same test is far more informative — you’re starting from a less surprising prior.
- Drag specificity from 95% to 99.9%. With the other inputs fixed, compare the resulting false-positive counts. A follow-up test requires its own conditional error model.
Read the denominators
At prevalence 0.1%, sensitivity 99%, and specificity 99%, the expected counts per 10,000 are 9.9 true positives and 99.9 false positives. PPV is 9.9/109.8 ≈ 9.0164%. These fractional values describe repeated-population expectations. The grid rounds counts to whole people, and its count ratio is labeled separately rather than silently substituted into the model calculation.
Sensitivity is TP/(TP+FN); specificity is TN/(TN+FP). Positive predictive value is TP/(TP+FP), while negative predictive value is TN/(TN+FN). A highly sensitive test can still have low PPV when the positive group contains many false positives. The experiment specifies both error rates; a single unspecified “accuracy” number cannot determine these cells.
Make a prediction
Keep both error rates fixed. At prevalence 10% rather than 1%, does PPV stay 50%?
Explore the answer
No. With sensitivity and specificity 99%, expected true and false positives per 10,000 are 990 and 90. PPV is 990/1080 = 91.6667%. The test’s conditional rates are unchanged; the population mixture changes the positive-result denominator.
What the model does and does not establish
This is a hypothetical probability model, not clinical evidence about a particular test. Its calculations assume the stated sensitivity and specificity apply to the population under discussion. Changing who is tested may change those conditional rates as well as prevalence; the comparison table deliberately holds them fixed to isolate the base-rate calculation.
The same mathematics applies to a rare-event alert system if its inputs are defined and measured consistently. It does not by itself establish what action to take after an alert or a positive test. Follow-up evidence also need not be independent: multiplying two likelihood ratios requires the relevant conditional-independence assumptions.
Make a prediction
If specificity is exactly one and prevalence is positive, what is PPV in this model?
Explore the answer
One: no condition-absent person tests positive, and sensitivity is positive in the supported settings. Perfect specificity does not imply perfect sensitivity; false negatives can remain.
Reference
Random Services: conditional probability develops the conditioning and Bayes identities. The table here derives every cell by multiplying the population fraction by its conditional test probability. Continue with Bayes for sequential evidence and the role of conditional independence.