paradox #14
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Berkson’s paradox

You will learn: Calculate how a shared selection rule creates association between independent traits.

Start with: Conditioning and independence

Imagine two independent binary traits, A and B. A rule admits anyone with at least one trait. Among admitted observations, seeing A makes B less likely. Nothing in the original model made one trait suppress the other: the association was created by restricting the sample.

This is a simple form of Berkson’s paradox. Abstract traits and a transparent admission rule let you inspect every population and selected cell. The experiment demonstrates what selection can do; it does not establish that a particular real-world correlation has this explanation.

Start with independent traits

Evidence probability under the original sampling scheme: 0.7500000.

Original and observation-weighted probabilities00.20.40.60.810,00,11,01,1underlying outcomeprobability

Filled left bars: original probability. Open right bars: probability after the stated observation. The same underlying outcomes remain in the table, including excluded ones.

OutcomePriorP(observation | outcome)Joint massPosterior
0,00.25000000.0000000.0000000.000000
0,10.25000001.0000000.25000000.3333333
1,00.25000001.0000000.25000000.3333333
1,10.25000001.0000000.25000000.3333333
QuantityWhole populationSelected population
P(A=1)0.50000000.6666667
P(B=1)0.50000000.6666667
Covariance0-0.1111111
Correlation0-0.5000000

The two traits are generated independently. Selection changes the population being described. A constant trait has zero variance and an undefined correlation; it does not have correlation zero.

Exact enumeration and conditional probability. There is no Monte Carlo error and no unstated reporting rule.

At the default, all four trait pairs have population probability 1/4 and population covariance zero. Keeping A = 1 or B = 1 removes only 0,0. The selected pairs 0,1; 1,0; and 1,1 each have probability 1/3.

Among selected observations with A = 0, B must be one: otherwise that observation would be excluded. Among those with A = 1, B can be zero or one, equally likely at the default. Thus P(B = 1 | A = 0, selected) is one, while P(B = 1 | A = 1, selected) is one half.

Compare the population statement P(B = 1 | A) = 1/2. The denominators differ. Conditioning on admission has made the traits dependent.

Compute the association

The selected means are both 2/3. AB equals one only in cell 1,1, so its selected expectation is 1/3. Covariance is 1/3 − (2/3)² = −1/9. Each binary trait has variance 2/9, giving correlation −1/2.

More generally, let p and q be the independent population probabilities and s = p + q − pq the “or” selection probability. Selected means are p/s and q/s, and the joint probability is pq/s. Covariance is pq(s − 1)/s² when s is positive. It is negative for nondegenerate interior probabilities. That sign follows this rule, not a general law for all selection mechanisms.

Change the admission rule

Keep everyone recovers the independent population. Keep exactly one trait removes 0,0 and 1,1. Whenever both remaining cells have positive mass, B = 1 − A and correlation is −1.

Keep A = 1 and B = 1 admits only observations with both traits. Covariance is zero, but correlation is undefined because both variances are zero. Zero covariance in this constant sample is not an informative estimate of the population relationship.

Make a prediction

Does a negative correlation after selection prove that A causally reduces B?

Explore the answer

No. The experiment starts with independent traits and creates negative association solely by conditioning on selection. A causal claim needs a model of how variables and selection arise, plus evidence supporting that model.

The collider connection

Admission can be a common effect of two variables: A affects admission and B affects admission. A common effect is called a collider in a causal diagram. Conditioning on that effect can associate otherwise unrelated causes. Numerical consequences still depend on probabilities and the selection mechanism.

In applied data, unseen population cells and admission probabilities may be unknown. This arithmetic cannot identify them from selected observations alone. It assumes independence before selection; changing that premise can strengthen, weaken, or reverse an existing association. Empty selected samples and constant traits remain explicitly undefined in the experiment.

References

Cole and colleagues, Illustrating bias due to conditioning on a collider explains the causal-diagram connection. Berkson’s bias, selection bias, and missing data examines selection with probability tables. Continue with Simpson’s paradox for a different mechanism: changing subgroup aggregation weights.

Reset all settings