Updated
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
Hypergeometric distribution
You will learn: Compute counts from finite sampling without replacement and check feasible outcomes.
Start with: Binomial distribution
An urn with balls, of them red. You draw without replacement, with each K-item subset equally likely. How many reds do you expect to pull?
Count an exact event
P(X=1) = 0.0278558; P(X≤1) = 0.030781. The possible counts run from 0 to 10.
Mean = 4.0000. Variance = 1.9592. On the first draw P(marked)=20/50. After a marked draw it becomes 19/49. After an unmarked draw it becomes 20/49.
Use a population of ten or fewer to list every subset. The exact probability calculation above remains available for larger populations.
What to notice
- Binomial convergence. The dashed red caps show Binomial(K, m/N). When the population is large relative to the sample size, the with-replacement and without-replacement distributions are nearly indistinguishable. Increase N while keeping the draw count fixed and the marked fraction near a fixed p to compare the limit.
- The finite-population correction. The hypergeometric variance carries an extra factor of that the Binomial lacks. At K = N (you drew everything) the variance is zero — there’s no randomness left.
- Support is narrower than it looks. You can only draw between and successes. Extreme slider values make the support shrink to a single point.
Where it matters
- Quality inspection. Pulling K items from a batch of N and counting defects.
- Card games. “What’s the chance of drawing 3 aces in a 5-card hand?” is Hypergeometric(52, 4, 5).
- Capture-recapture models. A closed population with a fixed marked count and a uniformly selected second sample gives a hypergeometric overlap count. Unequal catchability, lost marks, births, or migration can break that model; estimating population size requires additional assumptions.
The replacement comparison keeps both full supports visible and computes the same event in each model.
Enumerate before approximating
Choose the five-item preset: items 1 and 2 are marked, and three distinct items are sampled. There are C(5,3)=10 equally likely subsets. Six contain exactly one marked item, three contain both, and one contains none. Thus P(X=1)=0.6 and P(X≤1)=0.7. The table lists every subset and records whether it contributes to the cumulative event.
The formula counts favorable subsets in two stages: choose k marked items from m, then K−k unmarked items from N−m. Divide their product by the total number of K-item subsets. Treating different orders of the same subset as additional distinct hands would be inconsistent unless every hand is expanded by the same K! factor.
Dependence changes variance without changing the mean
In the five-item example, the first draw has marked probability 2/5. After a marked item it becomes 1/4; after an unmarked item it becomes 2/4. Before seeing any draws, each position still has marginal probability 2/5 of being marked. Linearity of expectation therefore gives E[X]=3×2/5=1.2, even though the indicators are dependent.
For two distinct draw positions, P(both marked)=(2/5)(1/4)=0.1. Their covariance is 0.1−0.4²=−0.06. Three individual variances contribute 3×0.24, while six ordered cross terms contribute 6×(−0.06), leaving Var(X)=0.36. This matches the finite-population correction 0.72×(5−3)/(5−1).
Make a prediction
If you inspect all five items, should increasing the number of simulation runs reveal any uncertainty in the marked count?
Explore the answer
No. Every complete sample contains exactly two marked items, so its variance is zero. More repetitions do not introduce randomness into that fixed count. With replacement, repeated selections can instead produce counts other than two.
Reference
Random Services: hypergeometric distribution derives uniform-subset sampling, its count law, and the covariance correction. Use the replacement comparison to keep the mean fixed while comparing event probabilities.