paradox #2
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Base rate neglect

You will learn: Combine prevalence with conditional evidence without reversing the conditioning.

Start with: Conditioning and independence

New to the notation? Start with the connected foundation for the underlying definitions and a worked example.

In this hypothetical example, a test has 99% sensitivity and 99% specificity. You test positive. How likely is it that you’re actually sick?

The instinct says ~99%. Under these assumptions and a prevalence of 1%, the answer is closer to 50%.

If the test comes back positive, the patient is sick with probability
50.0%
Sick, tests positive 99
Sick, tests negative 1
Healthy, tests positive 99
Healthy, tests negative 9801
In the rounded illustration, 99 of 198 positive tests are true positives (50.00%).

Squares represent condition present; circles represent condition absent. Filled marks test positive; open marks test negative. Use the table for precise values in this dense grid.

Exact expected frequencies per 10,000

Model statusPositiveNegativeTotal
Condition present99.00001.0000100.0000
Condition absent99.00009801.00009900.0000
All people198.00009802.000010000

PPV = true positives / all positives = 99.0000 / 198.0000 = 50.0000%. NPV = true negatives / all negatives = 99.9898%. Expected counts can be fractional; the colored grid alone rounds them to whole people. Neither is an observed clinical dataset.

Keep the test; change prevalence

PrevalenceExpected true positivesExpected false positivesPPV
0.1%9.90099.9009.016%
1.0%99.00099.00050.000%
10.0%990.00090.00091.667%
The probability uses the exact model inputs. The 10000-person illustration rounds expected cell counts to whole people; its count ratio can differ slightly. A rare condition can produce more false positives than true positives, depending on sensitivity and specificity.

Why so low?

Because the healthy population is enormous relative to the sick one. Out of 10 000 people:

  • 100 are sick. The 99% sensitive test catches 99 of them — those are the true positives.
  • 9 900 are healthy. The 99% specific test misidentifies 1% of them as positive — that’s 99 false positives.

Among the 198 people who tested positive, only 99 are actually sick. 99/198=50%99/198 = 50\%.

P(sick∣+)=P(+∣sick)⋅P(sick)P(+∣sick)P(sick)+P(+∣healthy)P(healthy)P(\text{sick} \mid +) = \frac{P(+\mid\text{sick}) \cdot P(\text{sick})}{P(+\mid\text{sick}) P(\text{sick}) + P(+\mid\text{healthy}) P(\text{healthy})}

What to try

  • Drop prevalence to 0.1%. Watch P(sick∣+)P(\text{sick} \mid +) collapse — you’d need near-perfect specificity for a positive result to mean much.
  • Lift prevalence to 10%. Now the same test is far more informative — you’re starting from a less surprising prior.
  • Drag specificity from 95% to 99.9%. With the other inputs fixed, compare the resulting false-positive counts. A follow-up test requires its own conditional error model.

Read the denominators

At prevalence 0.1%, sensitivity 99%, and specificity 99%, the expected counts per 10,000 are 9.9 true positives and 99.9 false positives. PPV is 9.9/109.8 ≈ 9.0164%. These fractional values describe repeated-population expectations. The grid rounds counts to whole people, and its count ratio is labeled separately rather than silently substituted into the model calculation.

Sensitivity is TP/(TP+FN); specificity is TN/(TN+FP). Positive predictive value is TP/(TP+FP), while negative predictive value is TN/(TN+FN). A highly sensitive test can still have low PPV when the positive group contains many false positives. The experiment specifies both error rates; a single unspecified “accuracy” number cannot determine these cells.

Make a prediction

Keep both error rates fixed. At prevalence 10% rather than 1%, does PPV stay 50%?

Explore the answer

No. With sensitivity and specificity 99%, expected true and false positives per 10,000 are 990 and 90. PPV is 990/1080 = 91.6667%. The test’s conditional rates are unchanged; the population mixture changes the positive-result denominator.

What the model does and does not establish

This is a hypothetical probability model, not clinical evidence about a particular test. Its calculations assume the stated sensitivity and specificity apply to the population under discussion. Changing who is tested may change those conditional rates as well as prevalence; the comparison table deliberately holds them fixed to isolate the base-rate calculation.

The same mathematics applies to a rare-event alert system if its inputs are defined and measured consistently. It does not by itself establish what action to take after an alert or a positive test. Follow-up evidence also need not be independent: multiplying two likelihood ratios requires the relevant conditional-independence assumptions.

Make a prediction

If specificity is exactly one and prevalence is positive, what is PPV in this model?

Explore the answer

One: no condition-absent person tests positive, and sensitivity is positive in the supported settings. Perfect specificity does not imply perfect sensitivity; false negatives can remain.

Reference

Random Services: conditional probability develops the conditioning and Bayes identities. The table here derives every cell by multiplying the population fraction by its conditional test probability. Continue with Bayes for sequential evidence and the role of conditional independence.

Reset all settings