paradox #3
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Simpson’s paradox

You will learn: Separate within-group comparisons from changes in the weights used to combine groups.

Start with: Joint and marginal distributions

New to the notation? Start with the connected foundation for the underlying definitions and a worked example.

Hospital B has a higher survival rate in both severe and mild cases — yet Hospital A’s overall survival rate looks better. How is that possible?

The answer: Hospital A treats mostly mild cases (which have high survival), while Hospital B takes on mostly severe cases. When you pool the groups, the case mix drowns out the per-group advantage.

Severe cases
A: 30%
B: 40%
A B B higher
Mild cases
A: 70%
B: 80%
A B B higher
Overall
A: 62%
B: 48%
A B A higher ← paradox!
Hospital A case mix: 20% severe, 80% mild
Hospital B case mix: 80% severe, 20% mild

Strict reversal: B has a higher rate in both subgroups, while A has a higher overall rate. The two totals use different subgroup weights.

Inspect the joint table

Each hypothetical hospital has 1,000 cases. These are expected frequencies calculated from the controls, not observed hospital data; displayed decimals are rounded.

Hospital and severityCasesSurviveDo not surviveSurvival rate
A, severe200.00060.000140.00030.00%
A, mild800.000560.000240.00070.00%
B, severe800.000320.000480.00040.00%
B, mild200.000160.00040.00080.00%

Use a common target case mix

RateAB
Actual separate mixes62.000%48.000%
Same chosen mix50.000%60.000%

A common weighted average cannot reverse a strict advantage in both subgroups. It answers a standardized descriptive question. Interpreting that difference as a treatment effect requires additional causal and comparability assumptions; a table alone cannot establish them.

What to notice

  • Default settings show the paradox. Hospital B is 10 percentage points higher in both subgroups, yet Hospital A’s pooled rate is higher — because A’s patients are 80% mild, B’s are 80% severe.
  • Equalise the case mix. Drag both ”% severe cases” sliders to the same value and the paradox disappears. The per-group winner becomes the overall winner.
  • Flip the advantage. Make A better in both subgroups but give A the harder case mix. The paradox flips direction.

Why it matters

Simpson’s paradox is not just an academic curiosity. It has misled researchers in:

  • Admissions — Different application mixes across departments can change aggregate admission rates. An aggregate disparity alone does not identify a causal explanation.
  • Sports — A player can have a higher batting average than a rival in every individual season, yet a lower career average.
  • Epidemiology — Crude rates and age-specific rates can point in opposite directions when age distributions differ. A causal treatment claim needs more than those associations.

Decide whether to adjust using the question and a defensible causal model. Conditioning on a mediator or collider can create a different bias; stratifying is not automatically a correction.

Work from the joint table

At the default settings, A has 200 severe cases with 60 expected survivors, plus 800 mild cases with 560 survivors. Its total is 620/1000 = 62%. B has 800 severe cases with 320 survivors, plus 200 mild cases with 160 survivors, totaling 480/1000 = 48%. Within severe cases B is 40% versus A’s 30%; within mild cases B is 80% versus 70%.

The denominators change when you pool. A’s overall rate is 0.2 × 0.3 + 0.8 × 0.7; B’s is 0.8 × 0.4 + 0.2 × 0.8. Averaging the subgroup percentages equally would answer a different question, because the observed mixtures are not equal.

Standardize to one target mix

Set the common severe fraction to 50%. A’s standardized rate becomes 50%, B’s 60%. These retain the original subgroup rates but give them the same weights. With common weight w, the B−A difference is:

w(bs−as)+(1−w)(bm−am).w(b_s-a_s)+(1-w)(b_m-a_m).

If both differences are strictly positive, every common mixture is positive. That is why equal weighting eliminates a strict reversal. It does not prove that the chosen standard population is appropriate for every decision.

Make a prediction

If the two hospitals tie in both subgroups but have different case mixes, is a different pooled rate a strict Simpson reversal?

Explore the answer

No. Different pooled rates are possible because the subgroup rates differ from each other, but neither hospital has a strict within-subgroup advantage to reverse. The demo now labels ties separately and requires both subgroup differences and the pooled difference to have opposing strict signs.

Decide which question is meaningful

The separate totals describe the observed case mixtures. Standardized totals describe a chosen common mixture. Neither alone identifies what would happen if the same person received care in the other hospital. That interpretation requires assumptions about selection, baseline severity, measurement, and other common causes.

If a grouping variable is affected by the exposure, conditioning may remove part of the effect of interest or induce selection bias. More stratification is not automatically more correct. State the comparison and a defensible causal structure before selecting the adjustment set.

Make a prediction

Can a reversal run in the opposite direction, with A higher within both groups but B higher overall?

Explore the answer

Yes. Give A the larger severe-case fraction while keeping A’s two subgroup rates above B’s. The same weighted-average mechanism operates in either direction; the demo’s message follows the current values.

Reference

Judea Pearl: Understanding Simpson’s Paradox distinguishes the numerical pattern from the causal question. The hypothetical joint table above supplies a self-contained arithmetic example; it does not claim to measure real hospital performance.

Reset all settings